Multi-resolution mapping joint learning Landsat-MODIS spatio-temporal data fusion method and device
Through the multi-resolution mapping joint learning method, the problem of poor fusion effect caused by the large difference in resolution between Landsat and MODIS satellite data is solved, high-precision spatiotemporal data fusion is achieved, and reconstructed data with high spatial and temporal resolution is generated.
Patent Information
- Application Number
- CN202510769098.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-23
AI Technical Summary
Existing technologies make it difficult to accurately depict the huge resolution differences between Landsat and MODIS satellite data, resulting in insufficient spatiotemporal data fusion effect and accuracy.
The multi-resolution mapping joint learning method is adopted. By acquiring Landsat and MODIS data and preprocessing them, an initial spatiotemporal data fusion model of multi-resolution mapping joint learning is constructed. The training dataset is used to train the target spatiotemporal data fusion model to reconstruct the Landsat data of the target date.
It improves the effect and accuracy of spatiotemporal data fusion, can more accurately depict complex scale differences, and generate reconstructed data with high spatial and temporal resolution.
Smart Images

Figure CN120689216A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing data processing, and in particular to a Landsat-MODIS spatiotemporal data fusion method and device for multi-resolution mapping joint learning. Background Art
[0002] Satellite remote sensing is an important means of monitoring the surface ecological environment over a large area. However, due to the limitations of sensor performance, a single satellite system usually cannot achieve both high spatial and high temporal resolution. Spatiotemporal data fusion technology based on Landsat and MODIS satellite data generates reconstructed data with both high spatial (30 meters) and high temporal resolution (1 day) by integrating the complementary spatiotemporal advantages of the two observation systems. The fine-scale, high-frequency data generated by spatiotemporal data fusion technology has effectively promoted research and development in areas such as crop phenology monitoring, aboveground biomass estimation, and disaster emergency response.
[0003] Scholars have conducted research on the spatiotemporal fusion of Landsat and MODIS satellite data. Specific methods can be categorized into three main categories: weighted function methods, spatial unmixing methods, and machine learning methods. Weighted function and spatial unmixing models use linear relationships to model the scale mapping relationship between the two data types. While these methods are widely used, their linear assumptions make it difficult to accurately fit complex scale-degradation relationships. Consequently, the reconstruction results often suffer from detail loss and reconstruction artifacts, leading to significant spectral and texture deviations. Machine learning methods employ a data-driven fusion framework, modeling scale mapping relationships through nonlinear relationships, and more accurately characterizing complex relationships. Therefore, compared to the first two methods, machine learning methods typically achieve better fusion results.
[0004] However, the current spatiotemporal fusion models of Landsat and MODIS satellite data are designed based on 30-meter resolution Landsat data and 500-meter resolution MOD09GA data. The spatial resolution difference is as high as about 16 times. Even through machine learning models, it is difficult to accurately depict such a huge scale difference. Summary of the Invention
[0005] In view of this, the main purpose of the embodiments of the present invention is to provide a Landsat-MODIS spatiotemporal data fusion method and device for multi-resolution mapping joint learning, in order to solve at least one of the problems of the existing technology. The present invention can improve the effect and accuracy of spatiotemporal data fusion.
[0006] To achieve the above objectives, an embodiment of the present invention provides a Landsat-MODIS spatiotemporal data fusion method for multi-resolution mapping joint learning, the method comprising:
[0007] Acquire Landsat data and MODIS data; MODIS data includes first surface reflectance data and second surface reflectance data;
[0008] Preprocess Landsat data and MODIS data to obtain training data sets;
[0009] Construct an initial spatiotemporal data fusion model for joint learning of multi-resolution mapping;
[0010] Training the initial spatiotemporal data fusion model using the training data set to obtain a target spatiotemporal data fusion model;
[0011] The Landsat data of the target date is reconstructed by using the target spatiotemporal data fusion model to obtain reconstructed data.
[0012] In some embodiments, the preprocessing of Landsat data and MODIS data to obtain a training data set comprises the following steps:
[0013] Performing spectral band extraction operations on the Landsat data, the first surface reflectance data, and the second surface reflectance data, respectively, to obtain first spectral band data, second spectral band data, and third spectral band data;
[0014] performing projection transformation, geometric registration, radiometric normalization, and resampling operations on the first spectral band data, the second spectral band data, and the third spectral band data, respectively, to obtain first resampled data, second resampled data, and third resampled data;
[0015] The training data set is constructed according to the first resampled data, the second resampled data, and the third resampled data.
[0016] In some embodiments, constructing the training dataset based on the first resampled data, the second resampled data, and the third resampled data comprises the following steps:
[0017] Matching the first resampled data, the second resampled data, and the third resampled data observed on the same day to obtain a Landsat-MODIS data pair;
[0018] The training data set is obtained according to a plurality of Landsat-MODIS data pairs of different surface areas and different time intervals.
[0019] In some embodiments, constructing an initial spatiotemporal data fusion model for joint learning of multi-resolution mappings comprises the following steps:
[0020] Based on the channel attention mechanism and the spatial attention mechanism, a joint attention mechanism module is constructed;
[0021] Construct cross-resolution and cross-band feature calibration modules;
[0022] Construct a multi-resolution progressive fusion module;
[0023] Construct a dual-stream time-phase differential fusion module;
[0024] Construct reconstruction data generation module and loss function;
[0025] The initial spatiotemporal data fusion model is obtained according to the joint attention mechanism module, the feature calibration module, the multi-resolution progressive fusion module, the dual-stream temporal differential fusion module, the reconstructed data generation module and the loss function.
[0026] In some embodiments, reconstructing the Landsat data of the target date using the target spatiotemporal data fusion model to obtain reconstructed data includes the following steps:
[0027] Select a reference date and a target date;
[0028] Using the Landsat data of the reference date, the MODIS data of the reference date, and the MODIS data of the target date as network input data;
[0029] The network input data is input into the target spatiotemporal data fusion model, and the reconstructed data is output.
[0030] In some embodiments, inputting the network input data into the target spatiotemporal data fusion model and outputting the reconstructed data comprises the following steps:
[0031] Performing cross-resolution and cross-band feature mapping fusion on the first surface reflectance data and the second surface reflectance data observed on the same day through the joint attention mechanism module and the cross-resolution and cross-band feature calibration module of the target spatiotemporal data fusion model;
[0032] Extract the variation characteristics of MODIS data and the first multi-scale features of Landsat data;
[0033] fusing the change feature and the first multi-scale feature through a multi-resolution progressive fusion module to obtain a target date Landsat feature;
[0034] Performing bidirectional feature extraction on the change features through a dual-stream time-phase difference fusion module to obtain forward difference features and backward difference features;
[0035] fusing the target date Landsat feature, the forward difference feature, and the backward difference feature to obtain a composite feature;
[0036] Performing convolution processing on the composite feature through a reconstruction data generation module of the target spatiotemporal data fusion model to obtain a second multi-scale feature;
[0037] Performing a cascade operation on the composite feature and the second multi-scale feature to obtain a multi-scale aggregated feature;
[0038] Residual connection and dimensionality reduction processing are performed on the target date Landsat features and the multi-scale aggregation features to obtain the reconstructed data.
[0039] To achieve the above objectives, another aspect of the present invention provides a Landsat-MODIS spatiotemporal data fusion device for multi-resolution mapping joint learning, the device comprising:
[0040] The first module is used to obtain Landsat data and MODIS data; the MODIS data includes first surface reflectance data and second surface reflectance data;
[0041] The second module is used to preprocess Landsat data and MODIS data to obtain training data sets;
[0042] The third module is used to obtain the initial spatiotemporal data fusion model for building multi-resolution mapping joint learning;
[0043] The fourth module is used to obtain the training data set, train the initial spatiotemporal data fusion model, and obtain the target spatiotemporal data fusion model;
[0044] The fifth module is used to reconstruct the Landsat data of the target date through the target spatiotemporal data fusion model to obtain reconstructed data.
[0045] To achieve the above-mentioned purpose, another aspect of an embodiment of the present invention provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the Landsat-MODIS spatiotemporal data fusion method of multi-resolution mapping joint learning described above.
[0046] To achieve the above-mentioned purpose, another aspect of an embodiment of the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the Landsat-MODIS spatiotemporal data fusion method of multi-resolution mapping joint learning described above.
[0047] To achieve the above objectives, another aspect of an embodiment of the present invention provides a computer program product or computer program, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium and execute the computer instructions, causing the computer device to perform the aforementioned Landsat-MODIS spatiotemporal data fusion method for joint learning of multi-resolution mapping.
[0048] Embodiments of the present invention include at least the following beneficial effects: the present invention provides a Landsat-MODIS spatiotemporal data fusion method and apparatus for multi-resolution mapping joint learning, the solution acquiring Landsat data and MODIS data; preprocessing the Landsat data and MODIS data to obtain a training data set; constructing an initial spatiotemporal data fusion model for multi-resolution mapping joint learning; training the initial spatiotemporal data fusion model using the training data set to obtain a target spatiotemporal data fusion model; and reconstructing the Landsat data of a target date using the target spatiotemporal data fusion model to obtain reconstructed data, thereby improving the effect and accuracy of spatiotemporal data fusion. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0050] Figure 1 Flowchart of a Landsat-MODIS spatiotemporal data fusion method for multi-resolution mapping joint learning provided by an embodiment of the present invention;
[0051] Figure 2 1 is an overall schematic diagram of a spatiotemporal data fusion model for joint learning of multi-resolution mapping provided by an embodiment of the present invention;
[0052] Figure 3 Schematic diagram of the network structure of the cross-resolution and cross-band feature calibration module provided by an embodiment of the present invention;
[0053] Figure 4 Schematic diagram of the network structure of the multi-resolution progressive fusion module provided by an embodiment of the present invention;
[0054] Figure 5 Schematic diagram of the network structure of the dual-stream time-phase differential fusion module provided by an embodiment of the present invention;
[0055] Figure 6 It is a schematic diagram of the hardware structure of the electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0056] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the present invention and are not intended to limit the present invention. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present invention. They are merely examples of devices and methods consistent with some aspects of the embodiments of the present invention as detailed in the appended claims.
[0057] It should be noted that although the functional modules are divided in the system schematic and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the system or the order in the flowchart. The terms "first / S100" and "second / S200" in the specification and claims and the above-mentioned figures may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein may be interpreted as "at the time of" or "when" or "in response to a determination".
[0058] The terms "at least one", "plurality", "each", "any", etc. used in the present invention include at least one, two or more, multiple, two or more, each refers to each of the corresponding multiple, and any refers to any one of the multiple.
[0059] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention pertains. The terms used herein are for the purpose of describing embodiments of the present invention only and are not intended to limit the present invention.
[0060] like Figure 1 As shown, an embodiment of the present invention provides a Landsat-MODIS spatiotemporal data fusion method for multi-resolution mapping joint learning, which may include but is not limited to steps S100 to S500:
[0061] Step S100, acquiring Landsat data and MODIS data; the MODIS data includes first surface reflectance data and second surface reflectance data;
[0062] Step S200, preprocessing the Landsat data and MODIS data to obtain a training data set;
[0063] Step S300, constructing an initial spatiotemporal data fusion model for multi-resolution mapping joint learning;
[0064] Step S400, training the initial spatiotemporal data fusion model using the training data set to obtain a target spatiotemporal data fusion model;
[0065] Step S500: reconstructing the Landsat data of the target date by using the target spatiotemporal data fusion model to obtain reconstructed data.
[0066] In step S100 of some embodiments, Landsat data and MODIS data are collected. For example, Landsat data, i.e., surface reflectance data, is acquired via the Landsat series of satellites; MODIS data, including two types of surface reflectance data, is acquired via the MODIS sensor: first surface reflectance data MOD09GA and second surface reflectance data MOD09GQ. The Landsat data has a spatial resolution of 30 meters and a temporal resolution of 16 days; the MOD09GA data has a spatial resolution of 500 meters and a temporal resolution of 1 day; and the MOD09GQ data has a spatial resolution of 250 meters and a temporal resolution of 1 day.
[0067] In some embodiments, step S200 may include but is not limited to steps S210 to S230:
[0068] Step S210, performing spectral band extraction operations on the Landsat data, the first surface reflectance data, and the second surface reflectance data, respectively, to obtain first spectral band data, second spectral band data, and third spectral band data;
[0069] Step S220, performing projection conversion, geometric registration, radiometric normalization, and resampling operations on the first spectral band data, the second spectral band data, and the third spectral band data, respectively, to obtain first resampled data, second resampled data, and third resampled data;
[0070] Step S230: construct the training data set according to the first resampled data, the second resampled data, and the third resampled data.
[0071] In steps S210 to S220 of some embodiments, spectral bands representing surface reflectance are extracted from Landsat data, MOD09GA data, and MOD09GQ data; geometric deviations of the extracted spectral band data are eliminated through projection transformation and geometric alignment; then radiometric normalization is performed to eliminate radiometric deviations; and then, through resampling processing, the MOD09GA and MOD09GQ data are sampled to a resolution of 30 meters for pixel-by-pixel matching with the Landsat data.
[0072] In some embodiments, the spectral band extraction operation requires extracting bands representing surface reflected light information from Landsat data, MOD09GA data, and MOD09GQ data. For example, six spectral bands covering visible light, near-infrared, and short-wave infrared are extracted from Landsat and MOD09GA data, while two bands, red light and near-infrared, are extracted from MOD09GQ data.
[0073] In some embodiments, step S230 may include but is not limited to steps S231 to S232:
[0074] Step S231, matching the first resampled data, the second resampled data, and the third resampled data observed on the same day to obtain a Landsat-MODIS data pair;
[0075] Step S232: obtaining the training data set based on a plurality of Landsat-MODIS data pairs of different surface areas and different time intervals.
[0076] In steps S231 to S232 of some embodiments, the Landsat and MODIS data observed on the same day are matched (referred to as data pairs), and at least two Landsat-MODIS data pairs are collected for the same area. According to the dates, the observation date of the Landsat-MODIS data pair of the previous date is used as the reference date, and the observation date of the Landsat-MODIS data pair of the next date is used as the target date. Multiple sets of data covering different surface areas and different time intervals are collected to construct a training data set. Optionally, in resolution resampling, in order to facilitate the resolution push in the subsequent fusion model, the spatial resolution of the MODIS data is resampled to an integer multiple of the spatial resolution of the Landsat data. MOD09GA is sampled to a resolution of 480 meters, and MOD09GQ is sampled to a resolution of 240 meters. On this basis, MOD09GA and MOD09GQ are further sampled to a resolution of 30 meters to complete pixel-by-pixel spatial matching.
[0077] In some embodiments, step S300 may include but is not limited to steps S310 to S360:
[0078] Step S310: constructing a joint attention mechanism module based on the channel attention mechanism and the spatial attention mechanism;
[0079] Step S320: constructing a cross-resolution and cross-band feature calibration module;
[0080] Step S330, constructing a multi-resolution progressive fusion module;
[0081] Step S340: construct a dual-stream time-phase differential fusion module;
[0082] Step S350, constructing a reconstruction data generation module and a loss function;
[0083] Step S360, obtaining the initial spatiotemporal data fusion model according to the joint attention mechanism module, the feature calibration module, the multi-resolution progressive fusion module, the dual-stream temporal differential fusion module, the reconstructed data generation module and the loss function.
[0084] In steps S310 to S360 of some embodiments, Figure 2 As shown in the figure, an initial spatiotemporal data fusion model for joint learning of multi-resolution mapping is constructed, which uses a deep neural network to characterize the complex mapping relationship between multiple resolutions. The spatiotemporal data fusion model for joint learning of multi-resolution mapping includes a joint attention mechanism module, a cross-resolution and cross-band feature calibration module (such as Figure 3 As shown), multi-resolution progressive fusion module (as shown Figure 4 As shown), dual-stream time-phase differential fusion module (as shown Figure 5 As shown), reconstruction data generation module and loss function.
[0085] In step S400 of some embodiments, the constructed multi-resolution mapping joint learning spatiotemporal data fusion model is trained using the constructed training data set, the objective function of the spatiotemporal data fusion model network is optimized and solved using the stochastic gradient algorithm, and the network parameters are adjusted using the back propagation mechanism until the model converges to obtain the target spatiotemporal data fusion model.
[0086] In some embodiments, step S500 may include but is not limited to steps S510 to S530:
[0087] Step S510, selecting a reference date and a target date;
[0088] Step S520, using the Landsat data of the reference date, the MODIS data of the reference date, and the MODIS data of the target date as network input data;
[0089] Step S530: input the network input data into the target spatiotemporal data fusion model, and output the reconstructed data.
[0090] In step S510 of some embodiments, the acquired Landsat data and MODIS data are matched with the Landsat data and MODIS data observed on the same day to obtain a Landsat-MODIS data pair. At least two Landsat-MODIS data pairs are collected for the same area, and the actual observation date of the Landsat satellite is used as the target date, and the observation date of the most recent Landsat-MODIS data pair before the target date is used as the reference date.
[0091] In step S520 of some embodiments, the target spatiotemporal data fusion model network input includes five data, namely: Landsat data of a reference date, MODIS data of a reference date (including MOD09GQ data of a reference date, MOD09GA data of a reference date), and MODIS data of the target date (including MOD09GQ data of a target date, MOD09GA data of a target date).
[0092] In step S530 of some embodiments, Landsat data for a reference date, MOD09GQ data for a reference date, MOD09GA data for a reference date, MOD09GQ data for a target date, and MOD09GA data for a target date are input into a target spatiotemporal data fusion model, which outputs Landsat reconstructed data for the target date. For example, in terms of the network structure of the target spatiotemporal data fusion model, first, a joint attention mechanism module constructed by combining spatial and channel features, combined with a cross-resolution and cross-band feature calibration module, performs cross-resolution and cross-band feature mapping fusion on the MOD09GQ and MOD09GA data of the same day, thereby obtaining a 240-meter resolution, full-band MODIS feature map. Subsequently, a multi-resolution progressive fusion module is used to sequentially inject the change features extracted by MODIS into the multi-scale features extracted by Landsat, achieving a progressive increase in spatial scale. Second, a dual-stream temporal difference module is used to perform bidirectional extraction of change features, achieving a robust expression of change features. Finally, based on the pyramid expansion structure and dense residual structure, the fusion features are further mined and extracted. After feature dimensionality reduction, the output of the network is obtained, namely the Landsat reconstruction data of the target date.
[0093] In some embodiments, step S530 may include but is not limited to steps S531 to S538:
[0094] Step S531: performing cross-resolution and cross-band feature mapping fusion on the first surface reflectance data and the second surface reflectance data observed on the same day through the joint attention mechanism module and the cross-resolution and cross-band feature calibration module of the target spatiotemporal data fusion model;
[0095] Step S532, extracting the change characteristics of the MODIS data and extracting the first multi-scale characteristics of the Landsat data;
[0096] Step S533: fusing the change feature and the first multi-scale feature through a multi-resolution progressive fusion module to obtain the Landsat feature of the target date;
[0097] Step S534: performing bidirectional feature extraction on the change features through a dual-stream time-phase difference fusion module to obtain forward difference features and backward difference features;
[0098] Step S535, fusing the target date Landsat feature, the forward difference feature, and the backward difference feature to obtain a composite feature;
[0099] Step S536, performing convolution processing on the composite feature through the reconstruction data generation module of the target spatiotemporal data fusion model to obtain a second multi-scale feature;
[0100] Step S537: performing a cascade operation on the composite feature and the second multi-scale feature to obtain a multi-scale aggregated feature;
[0101] Step S538 : performing residual connection and dimensionality reduction processing on the target date Landsat features and the multi-scale aggregation features to obtain the reconstructed data.
[0102] In step S531 of some embodiments, the purpose of the cross-resolution and cross-band feature calibration module is to perform channel and spatial weighted optimization on the first surface reflectance data and the second surface reflectance data observed on the same day. Exemplarily, based on the established joint attention module, the MOD09GA feature of the reference date and the MOD09GQ feature of the reference date are fused to obtain the MODIS feature of the reference date (including both the spectral information of the MOD09GA of the reference date and the spatial information of the MOD09GQ of the reference date); the MOD09GA feature of the target date and the MOD09GQ feature of the target date are fused to obtain the MODIS feature of the target date (including both the spectral information of the MOD09GA of the target date and the spatial information of the MOD09GQ of the target date).
[0103] Optionally, for the first surface reflectance data MOD09GA and the second surface reflectance data MOD09GQ of the same date, first, the original MOD09GA and MOD09GQ data are projected into the feature space based on a series of convolution operations. Then, a joint attention mechanism module is used to generate channel weights based on the MOD09GA features rich in band information, and spatial weights based on the MOD09GQ features rich in spatial information. The two feature maps are mutually calibrated to generate a feature map. The resulting feature map combines optimized spatial and spectral information. In addition, when processing MODIS data for the reference date and the target date, the consistency of the model feature expression is ensured by sharing module parameters.
[0104] In some embodiments, the input of the joint attention mechanism module is a set of (relatively) low-resolution features and a set of (relatively) high-resolution features. Low-resolution features have less spatial information but rich spectral information; high-resolution features have rich spatial information but less spectral information. The joint attention mechanism module is based on the channel attention mechanism, using the channel weights of the low-resolution features for weighting and recalibrating the high-resolution features; based on the spatial attention mechanism, using the spatial weights of the high-resolution features for weighting and recalibrating the low-resolution features. Optionally, the joint attention mechanism module (AttentionBlock, referred to as AB) is defined as:
[0105] F cA =(conv 3×3 (F ms ))⊙σ(conv 1×1 (P avg (F cs ))) (1)
[0106] F SA =(conv 3×3 (F cs ))⊙σ(conv 1×1 (F ms )) (2)
[0107] AB(F cs ,F ms )=conv 3×3 (F CA +F SA ) (3)
[0108] Among them, AB(F cs ,F ms ) represents the feature map obtained by the attention mechanism module, which integrates F cs Spectral information and F ms spatial information; F cs and Fms Represent the low-resolution and high-resolution feature maps of the input respectively; F CA and F SA Represents the output feature maps of the channel and spatial attention modules respectively; conv 1×1 (·) and conv 3×3 (·) denotes the convolution layers with kernel sizes of 1×1 and 3×3, respectively; P avg represents the average pooling operation; σ represents the Sigmoid activation function; ⊙ represents the element-wise multiplication operation.
[0109] In steps S532 to S533 of some embodiments, feature extraction is performed on MODIS data to obtain change features, and feature extraction is performed on Landsat data to obtain first multi-scale features. The change features and the first multi-scale features are fused through a multi-resolution progressive fusion module to obtain target date Landsat features. For example, the purpose of the multi-resolution progressive fusion module is to fuse 250-meter resolution MODIS features (resampled to 240-meter resolution on the model) with 30-meter resolution Landsat features in a step-by-step fusion process to achieve a progressive increase in spatial scale. Ideally, the target date Landsat features should be equal to the reference date Landsat features plus the temporal change features. However, there is a large spatial resolution difference between the change features extracted by MODIS and the required Landsat change features, so directly adding MODIS change features to Landsat features may result in loss of detail and artifacts in the fusion result. In this embodiment of the present invention, MODIS change features are injected into Landsat features through four fusion processes using a step-by-step fusion process with the help of an attention mechanism module. The process can be defined as:
[0110]
[0111] Where D 240m It represents the difference between the 240-meter MODIS characteristic maps of two phases, which is used to characterize the temporal change characteristics; and I represents the features extracted from the reference phase Landsat data using dilated convolution, with the corresponding dilation rates being 1, 2, 4, and 8, respectively, and extracting feature maps at scales of 30 meters, 60 meters, 120 meters, and 240 meters; 240m , I 120m , I 60m and I 30m Represent the features obtained by the four fusion processes, where I 240m , I 120m and I 60mAs the intermediate results, the scale of the fusion features is pushed to 240 meters, 120 meters and 60 meters respectively, and I 30m It is the final result of the multi-resolution progressive fusion module output, represented as a Landsat scale (i.e., 30 meters) feature map.
[0112] In some embodiments, first, formula (4) fuses the 240-meter resolution MODIS change features and the 240-meter resolution Landsat features through the attention mechanism module to obtain the 240-meter resolution output features; then, formula (5) further fuses the 240-meter features with the 120-meter Landsat features, pushing the spatial scale to 120 meters; then, formula (6) fuses the 120-meter features with the 60-meter Landsat features, pushing the spatial scale to 60 meters; finally, formula (7) fuses the 60-meter features with the 30-meter Landsat features, pushing the spatial scale to 30-meter resolution. Since the Landsat features are injected four times, the Landsat features need to be multiplied by a coefficient of 1 / 4. The 30-meter resolution features finally obtained by formula (7) are the output features of the multi-resolution progressive fusion module, which are recorded as the target date Landsat features.
[0113] In steps S534 to S535 of some embodiments, the purpose of the dual-stream temporal differential fusion module is to extract the change features in both directions and achieve a robust expression of the change features. Based on the MODIS features of the reference date and the target date obtained by the established joint attention module fusion, the forward and backward differential features are further calculated. Among them, the forward differential feature D f It is obtained by subtracting the reference phase MODIS feature from the target time MODIS feature and processing it through a convolution layer with a Tanh activation function; the backward difference feature D b is obtained by reverse subtraction. Subsequently, the Landsat feature map is estimated by the dual-stream structure. In the forward stream, the reference time Landsat feature and the forward differential feature D are compared by the residual structure. f The generated output features can be fused to approximately represent the target time Landsat features. For example, the formula used includes:
[0114]
[0115] in, represents the reference time Landsat feature map, D f represents the forward differential feature, O fRepresents the output features of the forward flow. Similarly, in the backward flow, the target date Landsat features are fused with the backward difference features. The generated output features can approximately represent the reference time Landsat features. The formulas used include:
[0116] O b =conv 3×3 (I 30m +D b ⊙I 30m ) (9)
[0117] Among them, I 30m represents the target time Landsat feature map, D b represents the backward difference feature, O b Represents the output features of the backward stream. Finally, the dual-stream feature maps are superimposed in the channel dimension and reduced in dimension through the convolution layer to generate a fused differential feature that can fully characterize the dynamic changes in both directions. The target date Landsat feature map is fused with the forward differential feature map and the backward differential feature map to obtain the composite feature F c The obtained composite features can fully describe the spatiotemporal information of the fusion process.
[0118] In some embodiments, in steps S536 to S538, the purpose of the reconstruction data generation module is to further mine and extract fusion features based on the pyramid expansion structure and dense residual structure, and obtain the Landsat reconstruction data of the target date through feature dimensionality reduction. The multi-scale feature map is extracted from the fusion feature through the pyramid expansion convolution structure. For example, first, the composite feature map F c Perform a convolution operation with an expansion rate of 1 to obtain the feature F d1 ; Further, F d1 Apply convolution processing with a dilation rate of 2 to generate features with a larger receptive field F d2 ; By analogy, convolutional layers with expansion rates of 4 and 8 are used to extract larger-scale features F d4 With F d8 Finally, the original features and the derived multi-scale features are concatenated to generate multi-scale aggregate features through the convolution layer. This structure can effectively map objects of different scales and project multi-scale features into a unified space. The process can be expressed as:
[0119] F d1 =Dilate(F c ,1),F d2 =Dilate(F d1 ,2),F d4 =Dilate(F d2 ,4),F d8 =Dilate(Fd4 ,8) (10)
[0120] F PDN =conv 3×3 ([F d1 ,F d2 ,F d4 ,F d8 ]) (11)
[0121] Among them, Dilate(·,i) represents the convolution operation with dilation rate i; F PDN represents the multi-scale aggregated features. The output features are then further processed by a residual dense module to enhance the reuse and propagation of key feature information. Finally, the fusion stage combines the target date Landsat features and the multi-scale aggregated features via a residual connection. This is then processed through a convolutional layer for dimensionality reduction to generate the reconstructed data.
[0122] In some embodiments, in order to effectively balance the numerical distribution and texture information of the fusion result, a loss function L(Θ) is constructed for the spatiotemporal data fusion model. Optionally, the loss function L(Θ) is derived from the Charbonnier loss function L char (Θ) and marginal loss function L edge (Θ) together, its expression is:
[0123] L(Θ)=λ1L char (Θ)+λ2L edge (Θ) (12)
[0124] Where Θ represents the model parameters, λ1 and λ2 represent the regularization parameters. The Charbonnier loss function L char (Θ) is a variant of the L1 norm, which is used to maintain the consistency of the numerical distribution of the reconstruction results with the real Landsat image and is defined as:
[0125]
[0126] in, represents the output of the network, Represents the label data for model training, i.e., the real observed target time Landsat data. N represents the number of training samples, and ε is a constant term used to prevent the gradient from disappearing. The marginal loss function L edge (Θ) is used to ensure the consistency of edge information between the reconstruction result and the real Landsat data by extracting the edge information of the two and minimizing their difference, which is expressed as:
[0127]
[0128] Where Δ represents the Laplace edge operator.
[0129] The embodiment of the present invention further provides a Landsat-MODIS spatiotemporal data fusion device for multi-resolution mapping joint learning, which can implement the above-mentioned Landsat-MODIS spatiotemporal data fusion method for multi-resolution mapping joint learning. The device includes:
[0130] The first module is used to obtain Landsat data and MODIS data; the MODIS data includes first surface reflectance data and second surface reflectance data;
[0131] The second module is used to preprocess Landsat data and MODIS data to obtain training data sets;
[0132] The third module is used to obtain the initial spatiotemporal data fusion model for building multi-resolution mapping joint learning;
[0133] The fourth module is used to obtain the training data set, train the initial spatiotemporal data fusion model, and obtain the target spatiotemporal data fusion model;
[0134] The fifth module is used to reconstruct the Landsat data of the target date through the target spatiotemporal data fusion model to obtain reconstructed data.
[0135] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0136] An embodiment of the present invention further provides an electronic device comprising a processor and a memory, wherein the memory stores a computer program. When the processor executes the computer program, it implements the Landsat-MODIS spatiotemporal data fusion method for multi-resolution mapping joint learning. The electronic device can be any intelligent terminal, including a tablet computer and an in-vehicle computer.
[0137] It can be understood that the contents of the above method embodiments are applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0138] refer to Figure 6 , Figure 6 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:
[0139] The processor 601 may be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided by the embodiments of the present invention.
[0140] The memory 602 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 602 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 602 and is called by the processor 601 to execute the Landsat-MODIS spatiotemporal data fusion method for multi-resolution mapping joint learning according to the embodiment of the present invention.
[0141] Input / output interface 603, used to implement information input and output;
[0142] Communication interface 604, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);
[0143] Bus 605 , which transmits information between various components of the device (e.g., processor 601 , memory 602 , input / output interface 603 , and communication interface 604 );
[0144] The processor 601 , the memory 602 , the input / output interface 603 and the communication interface 604 are connected to each other in communication within the device via a bus 605 .
[0145] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the computer program implements the Landsat-MODIS spatiotemporal data fusion method of multi-resolution mapping joint learning.
[0146] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiment, the functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0147] An embodiment of the present invention further provides a computer program product or computer program, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium and execute the computer instructions, causing the computer device to perform the aforementioned Landsat-MODIS spatiotemporal data fusion method for joint learning of multi-resolution mapping.
[0148] In summary, the Landsat-MODIS spatiotemporal data fusion method and device for multi-resolution mapping joint learning in an embodiment of the present invention achieves bridging of 500-meter and 30-meter resolutions by introducing 250-meter MODIS data (MOD09GQ). Compared with the traditional fusion method based on 30-meter resolution Landsat data and 500-meter resolution MOD09GA data, the embodiment of the present invention transforms the fusion problem from a "500-meter → 30-meter" mapping relationship to a "500-meter → 250-meter → 30-meter" multi-resolution mapping relationship, significantly improving the fusion effect, facilitating the accurate reconstruction of spectral features and spatial details, and having wide application value. Specifically, based on 30-meter Landsat data, 250-meter MODIS data (MOD09GQ) and 500-meter MODIS data (MOD09GA), a deep learning model spanning multiple resolution levels was constructed. The model integratedly established a multi-scale mapping relationship of "500 meters → 250 meters → 30 meters", accurately described the above cross-resolution relationship through deep learning technology, and combined multiple network modules to characterize the spatiotemporal transformation relationship in the fusion process. Multiple network modules were used to optimize the spatiotemporal characteristics implicit in the data, thereby improving the effect and accuracy of the spatiotemporal data fusion model and achieving accurate reconstruction of spectral characteristics and spatial details.
[0149] In some optional embodiments, the function / operation mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the function / operation involved, the two boxes shown in succession can actually be executed substantially simultaneously or the boxes can sometimes be executed in reverse order. In addition, the embodiment presented and described in the flow chart of the present invention is provided in an exemplary manner for the purpose of providing a more comprehensive understanding of the technology. The disclosed method is not limited to the operation and logic flow presented herein. Optional embodiments are contemplated in which the order of the various operations is changed and the sub-operations described as a part of a larger operation are performed independently.
[0150] Furthermore, although the present invention is described in the context of functional modules, it should be understood that, unless otherwise indicated, one or more of the functions and / or features described may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It will also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present invention. More specifically, given the properties, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the module will be understood within the ordinary skill of an engineer. Therefore, a person skilled in the art using ordinary skill will be able to implement the present invention set forth in the claims without undue experimentation. It will also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.
[0151] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0152] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0153] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering, or processing in another suitable manner as necessary, and then stored in a computer memory.
[0154] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0155] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0156] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.
[0157] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the embodiments. Those skilled in the art can make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present invention.
Claims
1. A Landsat-MODIS spatiotemporal data fusion method based on multi-resolution mapping joint learning, characterized in that: The following steps are involved: Acquire Landsat data and MODIS data; MODIS data includes first surface reflectance data and second surface reflectance data; Preprocess Landsat data and MODIS data to obtain training data sets; Construct an initial spatiotemporal data fusion model for joint learning of multi-resolution mapping; Training the initial spatiotemporal data fusion model using the training data set to obtain a target spatiotemporal data fusion model; The Landsat data of the target date is reconstructed by using the target spatiotemporal data fusion model to obtain reconstructed data.
2. The Landsat-MODIS spatiotemporal data fusion method for multi-resolution mapping joint learning according to claim 1, characterized in that: The Landsat data and MODIS data are preprocessed to obtain a training data set, comprising the following steps: Performing spectral band extraction operations on the Landsat data, the first surface reflectance data, and the second surface reflectance data, respectively, to obtain first spectral band data, second spectral band data, and third spectral band data; performing projection transformation, geometric registration, radiometric normalization, and resampling operations on the first spectral band data, the second spectral band data, and the third spectral band data, respectively, to obtain first resampled data, second resampled data, and third resampled data; The training data set is constructed according to the first resampled data, the second resampled data, and the third resampled data.
3. The Landsat-MODIS spatiotemporal data fusion method of multi-resolution mapping joint learning according to claim 2, characterized in that: The step of constructing the training data set according to the first resampled data, the second resampled data, and the third resampled data comprises the following steps: Matching the first resampled data, the second resampled data, and the third resampled data observed on the same day to obtain a Landsat-MODIS data pair; The training data set is obtained according to a plurality of Landsat-MODIS data pairs of different surface areas and different time intervals.
4. The Landsat-MODIS spatiotemporal data fusion method for multi-resolution mapping joint learning according to claim 1, It is characterized by: The construction of the initial spatiotemporal data fusion model for multi-resolution mapping joint learning includes the following steps: Based on the channel attention mechanism and the spatial attention mechanism, a joint attention mechanism module is constructed; Construct cross-resolution and cross-band feature calibration modules; Construct a multi-resolution progressive fusion module; Construct a dual-stream time-phase differential fusion module; Construct reconstruction data generation module and loss function; The initial spatiotemporal data fusion model is obtained according to the joint attention mechanism module, the feature calibration module, the multi-resolution progressive fusion module, the dual-stream temporal differential fusion module, the reconstructed data generation module and the loss function.
5. The Landsat-MODIS spatiotemporal data fusion method of multi-resolution mapping joint learning according to claim 1, characterized in that: The method of reconstructing the Landsat data of the target date by using the target spatiotemporal data fusion model to obtain reconstructed data includes the following steps: Select a reference date and a target date; Using the Landsat data of the reference date, the MODIS data of the reference date, and the MODIS data of the target date as network input data; The network input data is input into the target spatiotemporal data fusion model, and the reconstructed data is output.
6. The Landsat-MODIS spatiotemporal data fusion method for multi-resolution mapping joint learning according to claim 5, characterized in that: The step of inputting the network input data into the target spatiotemporal data fusion model and outputting the reconstructed data comprises the following steps: Performing cross-resolution and cross-band feature mapping fusion on the first surface reflectance data and the second surface reflectance data observed on the same day through the joint attention mechanism module and the cross-resolution and cross-band feature calibration module of the target spatiotemporal data fusion model; Extract the variation characteristics of MODIS data and the first multi-scale features of Landsat data; fusing the change feature and the first multi-scale feature through a multi-resolution progressive fusion module to obtain a target date Landsat feature; Performing bidirectional feature extraction on the change features through a dual-stream time-phase difference fusion module to obtain forward difference features and backward difference features; fusing the target date Landsat feature, the forward difference feature, and the backward difference feature to obtain a composite feature; Performing convolution processing on the composite feature through a reconstruction data generation module of the target spatiotemporal data fusion model to obtain a second multi-scale feature; Performing a cascade operation on the composite feature and the second multi-scale feature to obtain a multi-scale aggregated feature; Residual connection and dimensionality reduction processing are performed on the target date Landsat features and the multi-scale aggregation features to obtain the reconstructed data.
7. A Landsat-MODIS spatiotemporal data fusion device for multi-resolution mapping joint learning, characterized in that: include: The first module is used to obtain Landsat data and MODIS data; The MODIS data includes first surface reflectance data and second surface reflectance data; The second module is used to preprocess Landsat data and MODIS data to obtain training data sets; The third module is used to obtain the initial spatiotemporal data fusion model for building multi-resolution mapping joint learning; The fourth module is used to obtain the training data set, train the initial spatiotemporal data fusion model, and obtain the target spatiotemporal data fusion model; The fifth module is used to reconstruct the Landsat data of the target date through the target spatiotemporal data fusion model to obtain reconstructed data.
8. An electronic device, characterized in that: including a processor and a memory; The memory is used to store programs; The processor executes the program to implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that The storage medium stores a program, and the program is executed by a processor to implement the method according to any one of claims 1 to 6.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Landsat-MODIS remote sensing image space-time fusion method based on double-layer cross attention mechanism
CN121010871A