Remote sensing image cloud recognition model construction, cloud recognition method, device and medium
By constructing a multimodal deep learning model and comprehensively utilizing multi-source and multi-dimensional information such as spectrum, spatial texture, and time, the problems of misjudgment, missed detection, and boundary blurring in cloud identification in remote sensing images were solved, achieving high-precision and robust cloud identification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHONGKE YIYUAN TECHNOLOGY (BEIJING) CO LTD
- Filing Date
- 2025-12-30
- Publication Date
- 2026-05-29
AI Technical Summary
Existing cloud recognition methods for remote sensing images rely on a single optical image input. They are limited by differences in image time, changes in illumination, and the spectral similarity between clouds and bright ground features such as snow and deserts. This makes them prone to misidentification and omission. Furthermore, traditional convolutional neural network structures have insufficient receptive fields when capturing large-scale cloud boundaries, making it difficult to simultaneously represent the features of large-scale cloud clusters and small-scale thin clouds. As a result, the recognition accuracy is insufficient to meet the actual needs of large-scale remote sensing production.
A multimodal deep learning model is constructed by acquiring multi-temporal and multispectral remote sensing images, extracting three complementary features: spectral, temporal, and spatial. A proprietary three-branch network is designed for deep feature extraction, and an adaptive interaction and weighted integration of cross-modal information is achieved through a deep feature fusion module. Finally, a segmentation module outputs high-precision recognition results.
It significantly improves the accuracy, boundary detail fidelity, and overall robustness of cloud identification, enabling high-precision cloud identification in complex surface environments and under varying lighting conditions, thus meeting the needs of large-scale remote sensing production.
Smart Images

Figure CN122116159A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of remote sensing image cloud detection technology, specifically to the construction of remote sensing image cloud recognition models, cloud recognition methods, devices, and media. Background Technology
[0002] In the production of digital orthophoto maps (DOMs) from optical satellite remote sensing imagery, cloud cover is one of the most common and challenging problems. Because satellite imagery is inevitably affected by meteorological conditions, cloud cover and cloud shadows frequently exist within the images. These cloud features directly obscure ground features, leading to significant errors in subsequent geometric correction, mosaicking, and color balancing. For example, ground feature information cannot be correctly extracted from cloud-covered areas, and cloud shadows can cause abrupt color changes and abnormal brightness, thus affecting the overall accuracy and visual quality of the resulting map.
[0003] Existing cloud recognition semantic segmentation methods mostly rely on a single optical image input, which is limited by image temporal differences, illumination variations, and the spectral similarity between clouds and bright features such as snowfields and deserts, making them prone to misidentification and missed identification. In addition, traditional convolutional neural network structures suffer from insufficient receptive field when capturing large-scale cloud boundaries, making it difficult to simultaneously represent the features of large-scale cloud clusters and small-scale thin clouds. This limits the model's generalization ability, and the recognition accuracy is insufficient to meet the practical needs of large-scale remote sensing production. Summary of the Invention
[0004] This invention provides a remote sensing image cloud recognition model construction, cloud recognition method, device and medium to solve the problems of insufficient utilization of multi-source information, limited feature expression ability and low accuracy of cloud boundary recognition in the process of remote sensing image cloud detection and cloud mask extraction, and improve the comprehensiveness, precision and automation level of cloud layer recognition.
[0005] In a first aspect, the present invention provides a method for constructing a cloud recognition model for remote sensing images, comprising: acquiring multi-temporal multispectral remote sensing images and extracting spectral, temporal, and spatial features of the multi-temporal multispectral remote sensing images; constructing a multimodal deep learning model, the multimodal deep learning model comprising three deep feature extraction sub-networks for processing spectral, temporal, and spatial features respectively, a deep feature fusion module for fusing the deep features output by the three deep feature extraction sub-networks, and a segmentation module for semantically segmenting the fused deep features output by the deep feature fusion module and outputting cloud recognition results; training the multimodal deep learning model using spectral, temporal, and spatial features to obtain a cloud recognition model for remote sensing images, the cloud recognition model for outputting cloud recognition results.
[0006] The remote sensing image cloud recognition model construction method provided in this invention extracts three complementary features—spectral, temporal, and spatial—from multi-temporal and multispectral remote sensing images to construct a comprehensive feature representation. A proprietary three-branch network is then designed to perform deep feature extraction, and a deep feature fusion module enables adaptive interaction and weighted integration of cross-modal information. Finally, a segmentation module outputs high-precision recognition results, effectively overcoming the inherent defects of traditional methods that rely on single features, such as misjudgment of highly reflective ground objects, missed detection of thin clouds, and blurred boundaries. This significantly improves the accuracy, boundary detail fidelity, and overall robustness of cloud recognition under complex surface environments and varying lighting conditions. This invention comprehensively utilizes multi-source, multi-dimensional information such as spectral, spatial texture, and temporal data to improve the accuracy and detail fidelity of cloud recognition, meeting the needs of cloud monitoring and remote sensing data analysis in complex surface environments.
[0007] In one optional implementation, the steps of acquiring multi-temporal multispectral remote sensing images and extracting spectral, temporal, and spatial features from the multi-temporal multispectral remote sensing images include: acquiring multi-temporal multispectral remote sensing images with different spatial resolutions within a geographic area; preprocessing the multi-temporal multispectral remote sensing images to obtain spatially aligned multi-temporal multispectral remote sensing images with consistent resolution; and extracting multi-dimensional features from the preprocessed multi-temporal multispectral remote sensing images.
[0008] This implementation method acquires and fuses multi-resolution, multi-temporal multispectral remote sensing images, and performs unified spatial alignment and resolution normalization preprocessing to achieve high-quality integration of multi-source heterogeneous data. This process effectively overcomes the information limitations and alignment errors caused by traditional methods due to single data sources and varying scales. Based on this, spectral, temporal, and spatial three-dimensional features are systematically extracted, constructing a complete and spatiotemporally consistent feature representation system. This provides high-quality, multi-view input data for subsequent deep learning models, fundamentally enhancing the model's ability to identify complex cloud layers, significantly improving the accuracy, robustness, and automation level of cloud detection, and providing a reliable data foundation for large-scale, high-precision remote sensing image cloud processing.
[0009] In one optional implementation, the step of preprocessing multi-temporal multispectral remote sensing images to obtain spatially aligned and resolution-consistent multi-temporal multispectral remote sensing images includes: performing geometric correction and coordinate unification on the multi-temporal multispectral remote sensing images to obtain spatially aligned multi-temporal multispectral remote sensing images; and performing spatial resolution normalization on the spatially aligned multi-temporal multispectral remote sensing images using a resampling algorithm based on weighted neighborhood averaging to obtain spatially aligned and resolution-consistent multi-temporal multispectral remote sensing images.
[0010] In one optional implementation, the step of normalizing spatial resolution for spatially aligned multi-temporal multispectral remote sensing images using a weighted neighborhood averaging resampling algorithm includes: determining the neighborhood range of each target pixel; calculating the spatial distance between each neighboring pixel and the target pixel within the neighborhood range; determining the weight of each neighboring pixel based on the spatial distance between each neighboring pixel and the target pixel; and averaging the pixel values of each neighboring pixel according to the weights to obtain the normalized value of the target pixel, thus achieving spatial resolution normalization.
[0011] This implementation method fuses local spatial information from original images at different resolutions based on their geographical proximity to the target pixel using a distance-weighted approach. This not only achieves scale uniformity but, more importantly, effectively preserves key spatial structures and texture trends during downsampling by differentially weighting neighboring high-resolution pixels. During upsampling, smooth interpolation enhances spatial continuity. This adaptive weight allocation based on spatial distance ensures that the normalized image, while maintaining a unified resolution, maximizes the fusion and retention of effective information from the original multi-scale data. This provides subsequent multimodal deep learning models with higher information density and stronger spatial consistency, directly improving the model's ability to perceive and recognize subtle cloud features and complex boundaries.
[0012] In one optional implementation, the step of extracting multi-dimensional features from the preprocessed multi-temporal multispectral remote sensing image includes: directly extracting multi-band reflectance data as spectral dimension features from the preprocessed multi-temporal multispectral remote sensing image; extracting temporal dimension features from the preprocessed multi-temporal multispectral remote sensing image by calculating the reflectance difference of the same pixel between adjacent temporal phases; and extracting spatial dimension features from the preprocessed multi-temporal multispectral remote sensing image by calculating the local statistics and spatial gradient of the reflectance values of each pixel in a single temporal phase.
[0013] This implementation method directly extracts multi-band reflectance as spectral features, fully preserving the most essential spectral fingerprint information of ground objects and clouds. By calculating pixel-level temporal reflectance differences to obtain temporal features, it effectively captures spectral transient signals caused by the dynamic appearance and movement of clouds, thus reliably distinguishing stable, bright ground objects from transient cloud cover. Simultaneously, based on a single temporal phase, it extracts local statistics and spatial gradients as spatial features, quantitatively describing the internal texture uniformity, edge sharpness, and overall morphological structure of cloud clusters. These three types of features characterize remote sensing images from three dimensions: material properties, temporal variation patterns, and spatial morphology, forming a comprehensive and in-depth description of clouds. This provides a hierarchical and discriminative input for subsequent multimodal deep learning models, greatly enhancing the model's ability to finely identify and segment thin clouds, fragmented clouds, and cloud clusters with spectra similar to ground objects in complex scenes.
[0014] In one optional implementation, the step of constructing a multimodal deep learning model includes a spectral deep feature extraction sub-network, a spatial deep feature extraction sub-network, a temporal deep feature extraction sub-network, a deep feature fusion module, and a segmentation module. Specifically: the spectral deep feature extraction sub-network includes a deep convolutional neural network and a channel attention mechanism to extract deep features from the input spectral dimension features, obtaining spectral deep features; the spatial deep feature extraction sub-network includes a convolutional neural network and a pyramid pooling module to extract deep features from the input spatial dimension features, obtaining spatial deep features; the temporal deep feature extraction sub-network includes a convolutional long short-term memory network to extract deep features from the input temporal dimension features, obtaining temporal deep features; the deep feature fusion module concatenates and fuses the spectral deep features, temporal deep features, and spatial deep features to obtain fused deep features; and the segmentation module performs semantic segmentation on the fused deep features, outputting cloud recognition results.
[0015] This implementation utilizes a multimodal deep learning model to achieve efficient collaborative processing and deep fusion of spectral, spatial, and temporal features. It equips each modality with an optimal network structure: the spectral deep feature extraction sub-network, combining deep convolution and channel attention mechanisms, automatically focuses on the key bands most discriminative for cloud recognition; the spatial deep feature extraction sub-network effectively captures multi-scale spatial structures from local texture to global morphology through convolution and pyramid pooling; and the temporal deep feature extraction sub-network accurately models the temporal dynamic evolution of pixel reflectance using a convolutional long short-term memory network. Subsequently, the deep feature fusion module intelligently fuses the three deep features, and the segmentation module performs semantic segmentation accordingly. This model fully leverages the unique advantages of each modality, significantly enhancing its ability to distinguish complex cloud formations, variable environments, and spectral confusion scenarios. Ultimately, it achieves high-precision and robust cloud recognition, effectively solving the bottleneck problems of limited receptive field and insufficient feature representation capabilities in traditional single-network systems.
[0016] Secondly, the present invention provides a remote sensing image cloud recognition method, comprising: acquiring multi-temporal multispectral remote sensing images of a target geographic area, and extracting spectral dimension features, temporal dimension features, and spatial dimension features of the multi-temporal multispectral remote sensing images; inputting the spectral dimension features, temporal dimension features, and spatial dimension features into a remote sensing image cloud recognition model to obtain cloud recognition results of the target geographic area, wherein the remote sensing image cloud recognition model is constructed according to the method of the first aspect or any corresponding embodiment thereof.
[0017] Thirdly, the present invention provides a remote sensing image cloud recognition model construction device, comprising: a multi-dimensional feature extraction module for acquiring multi-temporal multispectral remote sensing images and extracting spectral, temporal, and spatial features of the multi-temporal multispectral remote sensing images; a model construction module for constructing a multimodal deep learning model, the multimodal deep learning model comprising three deep feature extraction sub-networks for processing spectral, temporal, and spatial features respectively, a deep feature fusion module for fusing the deep features output by the three deep feature extraction sub-networks, and a segmentation module for semantically segmenting the fused deep features output by the deep feature fusion module and outputting cloud recognition results; and a model training module for training the multimodal deep learning model using spectral, temporal, and spatial features to obtain a remote sensing image cloud recognition model, the remote sensing image cloud recognition model being used to output cloud recognition results.
[0018] Fourthly, the present invention provides a remote sensing image cloud recognition device, comprising: a data acquisition module for acquiring multi-temporal multispectral remote sensing images of a target geographic area and extracting spectral dimension features, temporal dimension features, and spatial dimension features of the multi-temporal multispectral remote sensing images; and a cloud recognition module for inputting the spectral dimension features, temporal dimension features, and spatial dimension features into a remote sensing image cloud recognition model to obtain cloud recognition results for the target geographic area, wherein the remote sensing image cloud recognition model is constructed according to the method of the first aspect or any corresponding embodiment thereof.
[0019] Fifthly, the present invention provides a computer-readable storage medium storing computer instructions, which are used to cause a computer to execute the remote sensing image cloud recognition model construction method of the first aspect or any corresponding embodiment thereof, or to execute the remote sensing image cloud recognition method of the second aspect. Attached Figure Description
[0020] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0021] Figure 1 This is a schematic diagram of the first process of constructing a remote sensing image cloud recognition model according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the second process of constructing a remote sensing image cloud recognition model according to an embodiment of the present invention; Figure 3 This is a flowchart illustrating a remote sensing image cloud recognition method according to an embodiment of the present invention; Figure 4 This is a structural block diagram of a remote sensing image cloud recognition model construction device according to an embodiment of the present invention; Figure 5 This is a structural block diagram of a remote sensing image cloud recognition device according to an embodiment of the present invention; Figure 6 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0023] It is understood that before using the technical solutions disclosed in the various embodiments of the present invention, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in the present invention and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.
[0024] In the production process of digital orthophoto maps from optical satellite remote sensing images, cloud cover and cloud shadow interference are core challenges that seriously affect data quality and the accuracy of subsequent processing. Existing cloud recognition semantic segmentation methods based on single image input are prone to the following drawbacks due to limitations such as spectral obfuscation, temporal variations, and the inadequacy of traditional network structures in representing multi-scale cloud features: The automatic recognition accuracy is low. Traditional methods rely on threshold segmentation, spectral feature ratio and other means, which are difficult to accurately distinguish between highly reflective ground objects and clouds, resulting in misjudgment and omission. The cloud boundary is blurred and details are lost. In complex surface or multi-cloud scenarios, the distinction between cloud and ground object boundaries is low. The cloud outline is prone to breakage, fusion or discontinuity in recognition, which affects the spatial consistency of the mask.
[0025] The problem of insufficient utilization of multidimensional features is that existing cloud detection methods generally rely on single spectral or brightness features, making it difficult to comprehensively utilize multidimensional information such as spectrum, spatial texture and temporal changes, resulting in insufficient accuracy of cloud recognition.
[0026] The problem is that the temporal variation characteristics have not been fully explored. Clouds in multi-temporal remote sensing data have obvious dynamic change patterns, but existing methods are unable to capture their temporal characteristics and cannot effectively distinguish between short-term cloud shadows, stable cloud layers, and changes in ground features.
[0027] Poor boundary details and spatial consistency, coupled with the complex morphology of cloud edges, make traditional methods prone to false positives or false negatives when dealing with high-contrast boundaries and thin cloud regions. This results in discontinuous, hollow, or jagged edges. Furthermore, insufficient spatial feature modeling can lead to breaks or fusion of cloud masks in local areas, distorting the overall cloud morphology and affecting further meteorological, environmental, or surface analysis.
[0028] Due to insufficient automation and adaptability, existing cloud detection methods often rely on manual parameter settings based on experience or threshold adjustments for specific scenarios, making it difficult to achieve universal and fully automated processing. Detection results vary significantly across different satellite sensors, resolutions, or surface environments, limiting the widespread application of large-scale, long-term remote sensing image cloud detection and analysis.
[0029] Based on this, the present invention provides a remote sensing image cloud recognition model construction, cloud recognition method, device and medium to solve the problems of insufficient utilization of multi-source information, limited feature expression ability and low accuracy of cloud boundary recognition in the process of remote sensing image cloud detection and cloud mask extraction, and improve the comprehensiveness, precision and automation level of cloud layer recognition.
[0030] According to an embodiment of the present invention, a method for constructing a cloud recognition model for remote sensing images is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0031] This embodiment provides a method for constructing a cloud recognition model for remote sensing images. Figure 1 This is a flowchart of a remote sensing image cloud recognition model construction method according to an embodiment of the present invention, such as... Figure 1 As shown, the process includes the following steps: Step S101: Acquire multi-temporal multispectral remote sensing images and extract the spectral dimension features, temporal dimension features and spatial dimension features of the multi-temporal multispectral remote sensing images.
[0032] This step involves systematically mining preliminary features from the acquired images, extracting information from three physical dimensions. Spectral features reflect the absorption and reflection characteristics of electromagnetic waves of different wavelengths by the inherent material composition of ground features and clouds, serving as the fundamental basis for distinguishing different ground feature types such as vegetation, water bodies, clouds, and snow. Temporal features capture the dynamic changes in land cover over time. The appearance, movement, or dissipation of clouds causes rapid abrupt changes in reflectivity, while stable surfaces show gradual changes. Therefore, this feature effectively distinguishes between "transient clouds" and "persistently bright ground features." Spatial features reflect the spatial distribution, texture structure, and boundary clarity of ground features and clouds. For example, thick clouds are usually uniform inside, while cloud edges or complex surfaces exhibit high variance and high gradients. This step, by extracting spectral, temporal, and spatial features in parallel, overcomes the limitations of traditional methods that rely on single spectral features, significantly improving the discriminative power of features for complex scenes and laying a solid data foundation for subsequent models to achieve high-precision recognition.
[0033] Step S102: Construct a multimodal deep learning model. The multimodal deep learning model includes three deep feature extraction sub-networks for processing spectral dimension features, temporal dimension features, and spatial dimension features, respectively; a deep feature fusion module for fusing the deep features output by the three deep feature extraction sub-networks; and a segmentation module for semantic segmentation of the fused deep features output by the deep feature fusion module and outputting cloud recognition results.
[0034] The three deep feature extraction sub-networks in this step are used to perform deep feature extraction on the preliminary features of each dimension in step S101, and then perform fusion and segmentation. By making full use of the combined advantages of spectral, spatial and temporal three-dimensional information, the recognition performance and robustness of the model far exceed those of traditional single network models in complex environments.
[0035] Step S103: The multimodal deep learning model is trained using spectral dimension features, temporal dimension features and spatial dimension features to obtain the remote sensing image cloud recognition model. The remote sensing image cloud recognition model is used to output the cloud recognition results.
[0036] This step is the training step of the model. By inputting multi-dimensional features such as spectrum, time, and space into a multimodal deep learning architecture for training, a highly generalizable and deployable high-precision cloud recognition model is finally generated, which can achieve fully automatic, pixel-level accurate cloud detection for new remote sensing images.
[0037] The remote sensing image cloud recognition model construction method provided in this embodiment extracts three complementary features—spectral, temporal, and spatial—from multi-temporal and multispectral remote sensing images to construct a comprehensive feature representation. Then, a proprietary three-branch network is designed to perform deep feature abstraction, and a deep feature fusion module achieves adaptive interaction and weighted integration of cross-modal information. Finally, a segmentation module outputs high-precision recognition results, effectively overcoming the inherent defects of traditional methods that rely on single features, such as misjudgment of highly reflective ground objects, missed detection of thin clouds, and blurred boundaries. This significantly improves the accuracy, boundary detail fidelity, and overall robustness of cloud recognition in complex surface environments and under varying lighting conditions. This invention comprehensively utilizes multi-source and multi-dimensional information such as spectral, spatial texture, and temporal data to improve the accuracy and detail fidelity of cloud recognition, meeting the needs of cloud monitoring and remote sensing data analysis in complex surface environments.
[0038] This embodiment provides another method for constructing a remote sensing image cloud recognition model. Figure 2 This is a flowchart of a remote sensing image cloud recognition model construction method according to an embodiment of the present invention, such as... Figure 2 As shown, the process includes the following steps: Step S201: Acquire multi-temporal multispectral remote sensing images and extract the spectral dimension features, temporal dimension features and spatial dimension features of the multi-temporal multispectral remote sensing images.
[0039] This step optimizes the input data through multi-dimensional and multi-modal fusion, combining optical remote sensing imagery with multi-source auxiliary data to compensate for the limitations of single spectral channels in cloud or ground feature discrimination. The remote sensing imagery used contains multispectral, multi-resolution, and multi-temporal features to support the training and validation of the cloud semantic segmentation and recognition model. Each image uses pixels as the smallest unit, and each pixel records the reflectance value at that location in different spectral bands.
[0040] Specifically, step S201 above includes: Step S2011: Acquire multi-temporal multispectral remote sensing images of different spatial resolutions within the geographic area.
[0041] In this step, "multispectral" refers to the presence of multiple spectral bands in the remote sensing image, including visible light, near-infrared, short-wave infrared, and thermal infrared. Blue light band (0.45–0.51 μm): used to identify water bodies, cloud edges, and atmospheric scattering characteristics; Green light band (0.53–0.59 μm): reflects the difference between vegetation and bare land; Red band (0.64–0.67 μm): used for ground feature identification and vegetation monitoring; Near-infrared band (0.85–0.88 μm): Vegetation has high reflectivity and cloud reflectivity is stable, which helps to distinguish ground features from clouds; Shortwave infrared band (1.57–2.29 μm): sensitive to differences in clouds, water vapor, and highly reflective ground features (snow, desert); Thermal infrared band (10.6–12.5 μm): reflects the surface radiation temperature and is used to distinguish between cold clouds and hot surfaces.
[0042] "Multi-temporal images" refers to remote sensing images of the same geographic area taken at different times. Multi-temporal images of the same area can be used to analyze the dynamic changes of clouds. By combining time-series reflectance differences, it is possible to distinguish short-lived clouds from stable ground backgrounds, thus improving the robustness of identification.
[0043] "Different spatial resolutions" refers to remote sensing images of the same geographic area acquired from different satellite sensors or processed at different levels by the same sensor. The actual ground size represented by a single pixel in these images, i.e., their spatial resolution, varies. The images are stored as a two-dimensional matrix, with each pixel corresponding to a fixed spatial resolution, such as 30 m × 30 m, and containing multi-band reflectance values. The boundary morphology, connectivity, and texture distribution of clouds can be extracted through spatial neighborhood relationships. Cloud regions exhibit high brightness non-uniformity and blurred edges in space. By calculating indicators such as the gray-level co-occurrence matrix, local variance, and gradient direction, the texture complexity of clouds can be quantified, which is used to identify thin and thick cloud regions.
[0044] Step S2012: Preprocess the multi-temporal multispectral remote sensing images to obtain spatially aligned multi-temporal multispectral remote sensing images with consistent resolution.
[0045] To ensure the compatibility and computational consistency of multi-source remote sensing data in subsequent deep learning models, all image data need to undergo standardized preprocessing. In some optional implementations, step S2012 includes: Step a1 involves performing geometric correction and coordinate unification on the multi-temporal multispectral remote sensing images to obtain spatially aligned multi-temporal multispectral remote sensing images.
[0046] All image data undergoes geometric correction to eliminate spatial distortions caused by imaging pose and terrain undulations, and is uniformly projected to the same geographic coordinate system, such as WGS84. Images from different sensors are processed through control point matching and resampling, with accuracy controlled at the sub-pixel level, ensuring strict alignment of images across bands and multiple temporal phases at the pixel level.
[0047] Step a2: For the spatially aligned multi-temporal multispectral remote sensing images, a resampling algorithm based on weighted neighborhood averaging is used to normalize the spatial resolution, resulting in spatially aligned multi-temporal multispectral remote sensing images with consistent resolution.
[0048] In some optional implementations, the neighborhood range of each target pixel in the spatially aligned multi-temporal multispectral remote sensing image is first determined; the spatial distance between each neighboring pixel and the target pixel within the neighborhood range is calculated; the weight of each neighboring pixel is determined based on the spatial distance between each neighboring pixel and the target pixel; the pixel values of each neighboring pixel are weighted and summed according to the weights, and then averaged to obtain the normalized value of the target pixel, thus achieving spatial resolution normalization.
[0049] In some specific implementations, spatial resolution normalization can be achieved using the following process: Remote sensing images exist at different spatial resolutions, such as 15 m, 30 m, and 100 m. To ensure consistent pixel scale when inputting into the deep learning model, this implementation method performs spatial resolution normalization processing on all image bands. A resampling algorithm based on weighted neighborhood averaging is used to unify images of different resolutions to a spatial scale of 30 m. The core calculation formula is as follows:
[0050] Where R'(x,y) is the target resolution pixel value, Ω(x,y) is the neighborhood window corresponding to the target pixel, R(i,j) is the original pixel value of the neighboring pixel, and w ij Distance weighting coefficients for neighboring pixels.
[0051] dij The spatial distance between the original neighbor cell (i,j) and the target cell (x,y) (in pixel spacing):
[0052] σ is a spatial smoothing parameter used to adjust the decay rate of the weighting function and control the range of influence from the neighborhood. When σ is large, more distant pixels are included in the calculation, resulting in a smoother result; when σ is small, the resampling result is closer to the original resolution.
[0053] When performing downsampling from high resolution to low resolution, the formula behaves as a weighted aggregation process, which can smooth noise while preserving the main spatial structure information and obtaining a more stable spectral response. In the case of upsampling from low resolution to high resolution, the formula is equivalent to smoothing interpolation of low-resolution pixels, which can enhance spatial continuity. However, due to insufficient original information, the generated details are mainly estimation results.
[0054] This normalization process unifies images from different sources and scales in space, providing a highly consistent data input foundation for the multi-source feature fusion of subsequent deep learning models.
[0055] Step S2013: Extract multi-dimensional features from the preprocessed multi-temporal multispectral remote sensing images.
[0056] In some optional implementations, step S2013 above includes: Step b1: For the preprocessed multi-temporal multispectral remote sensing image, directly extract multi-band reflectance data as spectral dimension features.
[0057] Step b2: For the preprocessed multi-temporal multispectral remote sensing image, extract the temporal dimension features by calculating the reflectance difference of the same pixel between adjacent temporal phases.
[0058] In the process of cloud identification, in order to reveal the dynamic changes of clouds using multi-temporal images, temporal feature normalization processing is performed on multi-source images of the same area acquired at different times. First, the images of each time phase are registered in the order of shooting time to ensure complete spatial alignment; then, reflectance difference is calculated for each pixel in the time dimension to enhance the transient characteristics of the clouds and reduce the interference of stable ground objects.
[0059] Let the same region at time t k With t k+1 The normalized reflectance at time t is R tk (x,y) and R tk+1( (x, y), then the time difference feature is defined as:
[0060] Where, ΔR t (x,y) reflects the intensity of spectral changes of a pixel between adjacent time phases. When |ΔR t When (x,y)| increases significantly, it indicates that there is a sudden change in reflectivity in the region, which is usually related to the appearance, dissipation or movement of clouds.
[0061] To eliminate illumination shift caused by differences in observation time periods, the time difference results are further standardized to obtain the spectral variation intensity ΔR' of the standardized pixels between adjacent time phases. t (x,y):
[0062] Where, μΔR t With σΔR t These are the mean and standard deviation of the time difference for this region, respectively.
[0063] Ultimately, the standardized temporal features and spatial features of each band are combined to form the model input, enabling the deep learning network to simultaneously perceive spatial structure and temporal dynamic changes, thus achieving higher robustness and spatiotemporal identification capabilities in cloud recognition.
[0064] Step b3: For the preprocessed multi-temporal multispectral remote sensing images, spatial dimension features are extracted by calculating the local statistics and spatial gradient of the reflectance values of each pixel in a single temporal phase.
[0065] In cloud identification using remote sensing imagery, spatial features are a crucial foundation for reflecting cloud distribution patterns and ground structure. The imagery is stored as a two-dimensional matrix, with each pixel corresponding to a fixed spatial location and resolution (e.g., 30 m × 30 m), recording multi-band reflectance values. By analyzing the spatial relationships between pixels and their neighboring pixels, cloud boundaries, connected regions, and texture features can be effectively extracted, providing spatial contextual information for deep learning models.
[0066] In the specific processing, the image matrix is first scanned to extract a local window for each pixel, typically a 3×3 or 5×5 neighborhood. The brightness variance, gray-level co-occurrence matrix (GLCM) features, and gradient direction information of this region are then calculated. The local variance σ... 2 (x, y) is used to measure the degree of change in pixel brightness, and the calculation formula is:
[0067] Where Ω(x,y) is the neighborhood window centered at pixel (x,y), and R(i,j) is the reflectance of the pixel within the neighborhood. Let N be the neighborhood mean, and N be the total number of pixels in the local region Ω(x,y).
[0068] In addition, to highlight the cloud edges and structural features, gradient values are calculated to obtain a gradient magnitude map:
[0069] in, R is the gradient value of pixel (x,y) and R is the reflectance of pixel (x,y). Areas with large gradient values usually correspond to the edge of thick clouds or the transition zone between cloud shadows, while areas with gentle gradients are mostly homogeneous clouds or the ground surface.
[0070] After the above calculations, a set of spatial feature maps is generated:
[0071] Where R′ is the normalized reflectivity matrix, σ 2 G represents the brightness variance map, and G represents the gradient magnitude map. Together, these three maps reflect the spatial brightness distribution, local texture complexity, and boundary variations of the image.
[0072] Ultimately, the spatial feature map, along with temporal difference features and radiation features, are used to construct a multi-channel input tensor for the deep learning model, enabling multi-dimensional semantic segmentation and recognition of clouds. This process effectively improves the model's ability to distinguish complex cloud types (such as thin clouds, cirrus clouds, and highly reflective ground features).
[0073] In some alternative implementations, after obtaining the spectral, temporal, and spatial features through the above process, it is necessary to uniformly integrate the features of each dimension in order to form a multimodal feature vector that can be directly input into a multimodal deep learning model.
[0074] For spectral features, spectral feature alignment is required. After geometric correction and spatial resolution normalization of all band images, the reflectance values of each band are arranged into a spectral feature vector according to pixel location:
[0075] For time-dimensional features, time-dimensional feature fusion is required to obtain a time-dimensional feature vector:
[0076] For spatial dimension features, they need to be fused to obtain spatial dimension feature vectors:
[0077] The above processing yields the vector data used for inputting the model.
[0078] It is worth noting that step S201, "acquiring multi-temporal and multispectral remote sensing images with different spatial resolutions within a geographic region," refers to the standard sample preparation process performed to construct the model training dataset. In practical applications, this process needs to be repeated across numerous different geographic regions. This involves acquiring corresponding multi-source and multi-temporal images from multiple independent geographic regions with wide distribution and diverse surface types, ultimately aggregating them into a large-scale training sample set with sufficient diversity and representativeness. Then, using the sample data corresponding to different geographic regions, a training sample set is constructed, and the model is trained through step S203. This design aims to ensure that the trained model can learn universal cloud recognition patterns, rather than being limited to a specific region, thereby guaranteeing that the model exhibits strong generalization ability and high recognition accuracy when facing various unknown geographic images during actual deployment. Specifically, following step S201, spectral, temporal, and spatial features corresponding to multiple geographic regions are acquired, and these features are used to train the multimodal deep learning model, resulting in a remote sensing image cloud recognition model.
[0079] Step S202: Construct a multimodal deep learning model. The multimodal deep learning model includes three deep feature extraction sub-networks for processing spectral dimension features, temporal dimension features, and spatial dimension features, respectively; a deep feature fusion module for fusing the output features of the three deep feature extraction sub-networks; and a segmentation module for semantic segmentation of the fused deep features output by the deep feature fusion module and outputting cloud recognition results.
[0080] To fully utilize spectral, spatial, and temporal features, this invention proposes a self-developed multimodal cloud semantic segmentation deep learning model. This model employs specialized processing modules for different features and implements information interaction in the fusion layer, ultimately outputting a high-precision cloud mask. Specifically, the aforementioned multimodal deep learning model includes a spectral deep feature extraction subnetwork, a spatial deep feature extraction subnetwork, a temporal deep feature extraction subnetwork, a deep feature fusion module, and a segmentation module.
[0081] The deep spectral feature extraction subnetwork comprises a deep convolutional neural network and a channel attention mechanism, used to extract deep spectral features from the input spectral dimension features, thus obtaining deep spectral features. Specifically, the input to this subnetwork is a multi-band spectral reflectance matrix, i.e., the spectral dimension feature vector F. spectral ∈R H×W×Cspectral A deep convolutional neural network (DCNN) is used to extract local spectral patterns, and a channel attention mechanism is introduced to weight and enhance key bands. The output is a deep spectral feature vector F'. spectral ∈R H×W×'C’spectral .
[0082] The deep spatial feature extraction subnetwork comprises a convolutional neural network and a pyramid pooling module, used to perform deep feature extraction on the input spatial dimension features to obtain deep spatial features. Specifically, the input is a spatial dimension feature vector F. spatial ∈R H×W×Cspatial Spatial structure information is extracted using a lightweight convolutional network. A multi-scale feature fusion module (Pyramid Pooling Module, PPM) is added after the convolutional output to enhance cloud boundary and connectivity representation. The output is a deep spatial feature vector F'. spatial ∈R H×W×C’spatial .
[0083] The temporal deep feature extraction subnetwork includes a convolutional long short-term memory network, used to perform deep feature extraction on the input temporal dimension features, obtaining the temporal deep features. Specifically, the input is the temporal dimension feature vector F after temporal difference normalization. temporal ∈R H×W×Ctemporal For the temporal dimension features of each pixel location, ConvLSTM is used to capture the dynamic changes in the cloud layer while keeping the spatial dimension H×W constant. The output is a deep temporal feature vector F'. temporal ∈R H×W×C’temporal .
[0084] The deep feature fusion module is used to concatenate and fuse spectral deep features, temporal deep features, and spatial deep features to obtain fused deep features. Specifically, through the sub-networks of the above three channels, vector forms of three-dimensional deep features are obtained. The deep feature fusion module concatenates the three-dimensional deep features in vector form through the channels to obtain fused deep features in vector form. F fused =Concat[F' spectral ,F' spatial ,F' temporal ]∈R H×W×C’fused Then, a lightweight self-attention mechanism is introduced to perform spatial channel weighting on the fused features, thereby enhancing the feature response of key cloud regions.
[0085] The segmentation module performs semantic segmentation on the fused deep features, outputting cloud recognition results. Specifically, the fused features are input into a UNet++ or improved UNet structure for high-resolution semantic segmentation, preserving low-level spatial details through skip connections while combining high-level semantic information. The output is a cloud mask probability map Y∈R. H×W×1 The probability map is classified by threshold or maximum probability to generate the final cloud mask, thus achieving semantic segmentation.
[0086] Step S203: The multimodal deep learning model is trained using spectral dimension features, temporal dimension features and spatial dimension features to obtain the remote sensing image cloud recognition model. The remote sensing image cloud recognition model is used to output the cloud recognition results.
[0087] The multimodal cloud recognition deep learning model proposed in this invention is optimized through end-to-end joint training after the design of spectral, spatial, and temporal feature modules is completed. Each feature module (spectral feature extraction module, spatial texture feature extraction module, and temporal series dynamic feature module) is a learnable neural network structure, and its parameters participate in backpropagation and weight update together with the subsequent fusion layer and semantic segmentation module, thereby achieving the overall optimal feature representation and recognition performance.
[0088] The training process is as follows: Step 1: Multimodal feature sample organization: The training dataset is divided into training, validation, and test sets, typically in a ratio of 8:1:1. To improve generalization ability, data augmentation strategies such as image rotation, brightness perturbation, and random spectral channel masking are introduced during training to simulate observational changes under different sensor and meteorological conditions.
[0089] Preprocessed multispectral and multitemporal remote sensing image data are input into the various feature branches of the model, and then fused. The fused deep features are used as input to the segmentation module. The labeled data are manually annotated cloud masks Y∈R. H×W×1 This is the corresponding cloud mask truth map, where 0 represents a non-cloud pixel and 1 represents a cloud pixel.
[0090] The second step involves joint optimization and gradient propagation mechanisms. All modules of the model share a unified loss target, and end-to-end weight updates are achieved through backpropagation. θ={θspectral,θspatial,θtemporal,θfusion,θseg} Each part represents a set of trainable parameters for different functional modules in the network, with the following meanings: θspectral represents the module parameters of the spectral deep feature extraction subnetwork, including the convolutional kernel weights, bias terms, and weighting coefficients of the channel attention mechanism in the deep convolutional neural network (DCNN) layer. It is used to learn the spectral response relationship between multi-band reflectance, highlighting key bands that contribute significantly to cloud recognition. θspatial represents the module parameters of the spatial deep feature extraction subnetwork, including the weights and biases of the lightweight convolutional network and the multi-scale feature fusion module (PPM). It is used to capture the spatial structure, boundary morphology, and connectivity features of clouds, realizing the expression of spatial context. θtemporal represents the module parameters of the temporal deep feature extraction subnetwork, corresponding to the gate weights and state update parameters of the ConvLSTM network. It is used to model the reflectance variation of the same pixel in multi-temporal images, capturing the dynamic change patterns of clouds. θfusion represents the parameters of the deep feature fusion module, including the weight matrix and normalization parameters used for self-attention calculation after feature concatenation. It is responsible for cross-modal information interaction and channel weighting, achieving adaptive fusion of spectral, spatial, and temporal features. θseg represents the parameters of the segmentation module, including the encoder and decoder convolutional layer weights and skip connection weights in the UNet++ or improved UNet backbone structure. It is used to perform pixel-level semantic segmentation of the fused features and output the final cloud mask probability map.
[0091] During training, the gradient signal is not only transmitted from the output layer to the semantic segmentation module, but also propagates forward layer by layer to each feature branch. This means that the parameters of each sub-network branch participate in joint optimization.
[0092] During training, the three branches of the network front end receive input synchronously and extract feature representations respectively. Then, the deep feature fusion module and the segmentation module perform joint backpropagation, so that the gradient can update the parameters of the three feature extraction modules at the same time, realizing feature collaborative optimization.
[0093] The third step is to design the loss function: The overall objective function for model training is a weighted combination of the losses from multiple tasks: Ltotal=αL CE +βL Dice +γL IoU +δL Boundary Among them, L CE The weighted cross-entropy loss is used to distinguish between cloud and non-cloud pixels; L Dice To improve consistency loss in cloud regions; L IoU To optimize the overall segmentation accuracy loss; L BoundaryTo enhance the continuity and sharpness of cloud edges, various loss terms are weighted and fused together to affect gradient updates, simultaneously optimizing the model in terms of accuracy and boundary details. α, β, γ, and δ are weight coefficients, adaptively adjusted based on validation set performance in experiments.
[0094] Step 4: Optimization and training strategies: The model uses the Adam optimizer for parameter updates, with an initial learning rate of 1×10⁻⁶. -4 Furthermore, a learning rate decay scheduler (CosineAnnealing) is used to accelerate the convergence process. To prevent overfitting, an early stopping mechanism is introduced, terminating training when the validation set loss shows no improvement for several consecutive rounds.
[0095] During training, a dynamic weight averaging (DWA) strategy is used to balance the gradient contributions of the spectral, spatial, and temporal branches, avoiding the dominance of a single modality in training.
[0096] In addition, the model weights with the best performance on the validation set are automatically saved through callback functions to ensure that the final output network is in the optimal fusion state.
[0097] Step 5: Model Validation and Convergence Evaluation After training, the model was evaluated using an independent test set. Key metrics included overall accuracy (OA), intersection-over-union (IoU), cloud mask boundary accuracy (BDE), and thin cloud recognition rate (TCA). Results show that the proposed multimodal fusion network achieves high-precision and continuous cloud recognition in complex terrains, high-reflectivity areas, and multi-cloud scenarios.
[0098] The remote sensing image cloud recognition model construction method provided in this embodiment has the following specific advantages: (i) Multimodal feature fusion to improve the completeness of feature representation: This invention introduces three types of features simultaneously: spectral, spatial, and temporal, to achieve a full-dimensional expression of cloud information. The spectral dimension enhances the difference in spectral response between ground objects and clouds through multi-band reflectance data. The spatial dimension utilizes gray-level co-occurrence matrix, local variance, and gradient information to extract texture and morphological features, thereby strengthening the structural expression of clouds. The temporal dimension employs temporal difference normalization to capture multi-temporal change trends, significantly improving the sensitivity to dynamic changes in clouds.
[0099] This data feature decomposition strategy enables the model to understand cloud features from multiple perspectives and scales, providing a high-quality input foundation for subsequent deep learning models.
[0100] (II) Advantages of the model structure: The model adopts a modular and multi-path structure, introducing specialized processing networks for different types of features: the spectral deep feature extraction sub-network uses a deep convolutional network (DCNN) combined with a channel attention mechanism to achieve adaptive enhancement of key bands; the spatial deep feature extraction sub-network uses a lightweight convolutional network + multi-scale feature fusion pyramid pooling (PPM) to effectively extract cloud boundaries and connectivity; and the temporal deep feature extraction sub-network models the dynamic laws of pixel changes over time through a long short-term memory network (LSTM).
[0101] The three types of sub-network modules work together to achieve an organic combination of "spectral discriminability + spatial continuity + temporal dynamics", which significantly improves the model's comprehensive expressive ability.
[0102] (III) Advantages of Multimodal Fusion and Semantic Segmentation: In this embodiment, during the multimodal information fusion stage, a channel concat method is used to deeply fuse spectral, spatial, and temporal deep features. A lightweight self-attention mechanism is introduced to achieve cross-modal feature information interaction and adaptive weight allocation. This mechanism can dynamically focus on regions and feature channels that contribute significantly to cloud recognition, effectively mitigating the mismatch problem caused by differences in the distribution of features across different modalities, thereby improving the overall fusion effect.
[0103] The fused multimodal features are input into the U-Net++ semantic segmentation decoding structure. This structure achieves pixel-by-pixel prediction of high-resolution features through dense skip connections and multi-scale feature aggregation, significantly enhancing the detail restoration and continuity of cloud boundaries. The model described in this invention exhibits high robustness and generalization ability under complex climate, multi-source sensor, and multi-temporal conditions, effectively improving the accuracy and stability of cloud detection and cloud mask generation in remote sensing images.
[0104] This embodiment provides a method for cloud recognition in remote sensing images. Figure 3 This is a flowchart of a remote sensing image cloud recognition method according to an embodiment of the present invention, such as... Figure 3 As shown, the process includes the following steps: Step S301: Acquire multi-temporal multispectral remote sensing images of the target geographic area, and extract the spectral, temporal, and spatial dimensional features of the multi-temporal multispectral remote sensing images. See [link to specific method] for details. Figure 2 Step S201 in the illustrated embodiment will not be described again in this embodiment.
[0105] Step S302: Input the spectral dimension features, temporal dimension features and spatial dimension features into the remote sensing image cloud recognition model to obtain the cloud recognition result of the target geographic area. The remote sensing image cloud recognition model is constructed according to the remote sensing image cloud recognition model construction method of any embodiment of the present invention.
[0106] The remote sensing image cloud recognition method provided in this embodiment fully utilizes various information such as spectral density, spatial texture, and temporal variation. Through feature extraction, temporal modeling, and multi-scale fusion techniques, it achieves high-precision cloud identification and cloud mask generation. The proposed technical solution has strong adaptability and automation capabilities, and can effectively improve the insufficient recognition accuracy of existing technologies in cloud boundary identification, thin cloud detection, and complex surface environments.
[0107] This embodiment also provides a remote sensing image cloud recognition model construction device, which is used to implement the above embodiments and preferred embodiments, and will not be repeated for details already described. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0108] This embodiment provides a device for constructing a remote sensing image cloud recognition model, such as... Figure 4 As shown, it includes: The multi-dimensional feature extraction module 401 is used to acquire multi-temporal multispectral remote sensing images and extract the spectral dimension features, temporal dimension features and spatial dimension features of the multi-temporal multispectral remote sensing images. The model building module 402 is used to build a multimodal deep learning model. The multimodal deep learning model includes three deep feature extraction sub-networks for processing spectral dimension features, temporal dimension features and spatial dimension features respectively, a deep feature fusion module for fusing the output features of the three deep feature extraction sub-networks, and a segmentation module for performing semantic segmentation on the fused deep features output by the deep feature fusion module and outputting cloud recognition results. The model training module 403 is used to train a multimodal deep learning model using spectral, temporal, and spatial features to obtain a remote sensing image cloud recognition model, which is used to output cloud recognition results.
[0109] The remote sensing image cloud recognition model construction apparatus provided in this embodiment of the invention can execute the remote sensing image cloud recognition model construction method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method. Further functional descriptions of the above modules and units are the same as in the corresponding embodiments described above, and will not be repeated here.
[0110] This embodiment also provides a remote sensing image cloud recognition device, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0111] This embodiment provides a remote sensing image cloud recognition device, such as... Figure 5 As shown, it includes: The data acquisition module 501 is used to acquire multi-temporal multispectral remote sensing images of the target geographic area and extract the spectral dimension features, temporal dimension features and spatial dimension features of the multi-temporal multispectral remote sensing images. The cloud recognition module 502 is used to input spectral dimension features, temporal dimension features and spatial dimension features into the remote sensing image cloud recognition model to obtain the cloud recognition result of the target geographic area. The remote sensing image cloud recognition model is constructed according to the remote sensing image cloud recognition model construction method of any embodiment.
[0112] The remote sensing image cloud recognition device provided in this embodiment of the invention can execute the remote sensing image cloud recognition method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects for executing the method. Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.
[0113] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.
[0114] The following is a detailed reference. Figure 6 This diagram illustrates a suitable structural design for implementing an electronic device according to embodiments of the present invention. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 601, which can perform various appropriate actions and processes based on a program stored in read-only memory (ROM) 602 or a program loaded from memory 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of the electronic device. The processor 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0115] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.
[0116] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a memory 608, or installed from a ROM 602. When the computer program is executed by the processor 601, it performs the functions defined in the remote sensing image cloud recognition model construction method or remote sensing image cloud recognition method of the embodiments of the present invention.
[0117] Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.
[0118] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the remote sensing image cloud recognition model construction method or remote sensing image cloud recognition method shown in the above embodiments is implemented.
[0119] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0120] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A method for constructing a cloud recognition model for remote sensing images, characterized in that, The method includes: Acquire multi-temporal multispectral remote sensing images and extract the spectral dimension features, temporal dimension features, and spatial dimension features of the multi-temporal multispectral remote sensing images; A multimodal deep learning model is constructed, comprising three deep feature extraction sub-networks for processing the spectral dimension features, temporal dimension features and spatial dimension features respectively, a deep feature fusion module for fusing the deep features output by the three deep feature extraction sub-networks, and a segmentation module for performing semantic segmentation on the fused deep features output by the deep feature fusion module and outputting cloud recognition results. The multimodal deep learning model is trained using the spectral, temporal, and spatial features to obtain a remote sensing image cloud recognition model, which is used to output cloud recognition results.
2. The method according to claim 1, characterized in that, The steps of acquiring multi-temporal multispectral remote sensing images and extracting the spectral, temporal, and spatial features of the multi-temporal multispectral remote sensing images include: Acquire multi-temporal, multispectral remote sensing images of different spatial resolutions within a geographic region; The multi-temporal multispectral remote sensing images are preprocessed to obtain spatially aligned and consistent multi-temporal multispectral remote sensing images. Multidimensional feature extraction is performed on the preprocessed multi-temporal multispectral remote sensing images.
3. The method according to claim 2, characterized in that, The step of preprocessing the multi-temporal multispectral remote sensing images to obtain spatially aligned and consistent-resolution multi-temporal multispectral remote sensing images includes: Geometric correction and coordinate unification are performed on the multi-temporal multispectral remote sensing images to obtain spatially aligned multi-temporal multispectral remote sensing images. For spatially aligned multi-temporal multispectral remote sensing images, a resampling algorithm based on weighted neighborhood averaging is used to normalize the spatial resolution, resulting in spatially aligned multi-temporal multispectral remote sensing images with consistent resolution.
4. The method according to claim 3, characterized in that, The step of normalizing the spatial resolution of the spatially aligned multi-temporal multispectral remote sensing images using a resampling algorithm based on weighted neighborhood averaging includes: Determine the neighborhood range of each target pixel in spatially aligned multi-temporal multispectral remote sensing images; Calculate the spatial distance between each neighboring pixel within the neighborhood and the target pixel; The weight of each neighboring pixel is determined based on the spatial distance between each neighboring pixel and the target pixel. The pixel values of each neighboring pixel are weighted and summed according to the weights, and then averaged to obtain the normalized value of the target pixel, thus achieving spatial resolution normalization.
5. The method according to claim 2, characterized in that, The step of extracting multi-dimensional features from the preprocessed multi-temporal multispectral remote sensing images includes: For preprocessed multi-temporal multispectral remote sensing images, multi-band reflectance data are directly extracted as spectral dimension features; For preprocessed multi-temporal multispectral remote sensing images, temporal dimension features are extracted by calculating the reflectance difference of the same pixel between adjacent temporal phases. For preprocessed multi-temporal multispectral remote sensing images, spatial dimension features are extracted by calculating the local statistics and spatial gradient of the reflectance values of each pixel in a single temporal phase.
6. The method according to claim 1, characterized in that, In the step of constructing the multimodal deep learning model, the multimodal deep learning model includes a spectral deep feature extraction subnetwork, a spatial deep feature extraction subnetwork, a temporal deep feature extraction subnetwork, a deep feature fusion module, and a segmentation module, wherein: The deep spectral feature extraction subnetwork includes a deep convolutional neural network and a channel attention mechanism, which is used to extract deep features from the input spectral dimension features to obtain deep spectral features; The spatial deep feature extraction subnetwork includes a convolutional neural network and a pyramid pooling module, which is used to extract deep features from the input spatial dimension features to obtain spatial deep features; The temporal deep feature extraction subnetwork contains a convolutional long short-term memory network, which is used to perform deep feature extraction on the input temporal dimension features to obtain temporal deep features; The deep feature fusion module is used to splice and fuse spectral deep features, temporal deep features, and spatial deep features to obtain fused deep features; The segmentation module is used to perform semantic segmentation on the fused deep features and output cloud recognition results.
7. A method for cloud recognition in remote sensing images, characterized in that, The method includes: Acquire multi-temporal multispectral remote sensing images of the target geographic area, and extract the spectral dimension features, temporal dimension features, and spatial dimension features of the multi-temporal multispectral remote sensing images; The spectral, temporal, and spatial features are input into the remote sensing image cloud recognition model to obtain the cloud recognition result of the target geographic area. The remote sensing image cloud recognition model is constructed according to the method described in any one of claims 1-6.
8. A device for constructing a cloud recognition model for remote sensing images, characterized in that, The device includes: A multi-dimensional feature extraction module is used to acquire multi-temporal multispectral remote sensing images and extract the spectral dimension features, temporal dimension features and spatial dimension features of the multi-temporal multispectral remote sensing images. The model building module is used to build a multimodal deep learning model. The multimodal deep learning model includes three deep feature extraction sub-networks for processing the spectral dimension features, temporal dimension features and spatial dimension features respectively, a deep feature fusion module for fusing the deep features output by the three deep feature extraction sub-networks, and a segmentation module for performing semantic segmentation on the fused deep features output by the deep feature fusion module and outputting cloud recognition results. The model training module is used to train the multimodal deep learning model using the spectral dimension features, temporal dimension features and spatial dimension features to obtain a remote sensing image cloud recognition model, which is used to output cloud recognition results.
9. A remote sensing image cloud recognition device, characterized in that, The device includes: The data acquisition module is used to acquire multi-temporal multispectral remote sensing images of the target geographic area and extract the spectral dimension features, temporal dimension features and spatial dimension features of the multi-temporal multispectral remote sensing images. The cloud recognition module is used to input the spectral dimension features, temporal dimension features and spatial dimension features into the remote sensing image cloud recognition model to obtain the cloud recognition result of the target geographic area. The remote sensing image cloud recognition model is constructed according to the method of any one of claims 1-6.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to execute the remote sensing image cloud recognition model construction method according to any one of claims 1 to 6, or to execute the remote sensing image cloud recognition method according to claim 7.