Remote sensing image space-time fusion method and system based on geometric algebraic generative adversarial network

Through the method of generating an adversarial network by geometric algebra, combining geometric algebra generator and multi-dimensional discriminator, the problem of spectrum and structural information loss in the spatiotemporal fusion of remote sensing images is solved, and high-quality high-spatiotemporal resolution remote sensing images are generated.

CN120339089APending Publication Date: 2025-07-18SHANGHAI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510498304.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When the existing remote sensing image spatiotemporal fusion method processes high spatial and temporal resolution images, there is a problem of loss of spectral and structural information between channels, making it difficult to generate high-quality fusion images.

Method used

The method of generating an adversarial network based on geometric algebra is adopted, and the remote sensing image space-time fusion is performed through geometric algebra generator and multi-dimensional discriminator, and multi-source remote sensing data is processed using geometric algebra residual encoder and upsampling decoder, retaining the spectral and structural information of the image, and training and optimization is performed through multi-dimensional discriminator.

Benefits of technology

It significantly improves the accuracy and robustness of generating remote sensing images, can better cope with complex spatiotemporal changes and variable data quality, and generates high spatiotemporal resolution images that are closer to reality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339089A_ABST
    Figure CN120339089A_ABST
Patent Text Reader

Abstract

The invention relates to a remote sensing image space-time fusion method and system based on a geometric algebraic generative adversarial network, and the method employs a trained geometric algebraic generative adversarial model to carry out the remote sensing image space-time fusion, and the geometric algebraic generative adversarial model comprises a geometric algebraic generator. The geometric algebraic generator comprises a geometric algebraic residual encoder and a geometric algebraic up-sampling decoder, and the method comprises the following steps: acquiring multi-source remote sensing data, cutting and preprocessing the multi-source remote sensing data, and processing the preprocessed multi-source remote sensing data by using the geometric algebraic residual encoder to obtain time characteristics and spatial characteristics; carrying out feature fusion on the time features and the space features to obtain fusion features; and inputting the fused features into a geometric algebraic up-sampling decoder to output a first reconstructed remote sensing image. Compared with the prior art, the geometric algebraic convolution is utilized to improve the traditional generative adversarial model, so that the detail features of the original image can be reserved to the greatest extent during feature extraction, and the reconstructed image structure is closer to the reality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of remote sensing technology, and particularly to a spatio-temporal fusion method and system for remote sensing images based on geometric algebra generative adversarial network. Background Art

[0002] With the development of remote sensing technology, remote sensing images have been widely used in fields such as environmental monitoring, resource management, and disaster prediction. However, due to technical and budget limitations, obtaining remote sensing images with high spatial resolution and high temporal resolution still faces huge challenges. This problem greatly restricts the further development of remote sensing technology in practical applications. To solve this problem, spatio-temporal fusion technology has emerged, which can provide higher-quality spatio-temporal data by combining remote sensing images with different temporal and spatial resolutions. Spatio-temporal fusion technology overcomes the limitations of a single sensor in spatial resolution and temporal resolution by integrating information from remote sensing data of different sources. For example, the Landsat-8 satellite provides a spatial resolution of 30 meters and a revisit period of 16 days, while the MODIS satellite provides a lower spatial resolution (such as 250m, 500m, 1000m) but a short revisit period of only 1 day. By using the spatio-temporal fusion technology to utilize the temporal variation information of MODIS images and the spatial details of Landsat images, their advantages can be combined to generate fused images with high spatial and high temporal resolutions. However, the fusion process is relatively complex and usually relies on relatively simple rules, making it difficult to cope with complex land cover changes and irregular temporal differences. With the rapid development of deep learning, deep learning-based spatio-temporal fusion methods are used to solve the defects of existing spatio-temporal fusion technologies. As an effective deep learning method, the generative adversarial network (GAN) has been widely applied in fields such as image generation and image super-resolution. Due to its powerful generation ability, GAN has been introduced into the spatio-temporal fusion field to process and generate high-quality remote sensing images. The generation network generates more accurate fusion results by learning the mapping relationship between images of different resolutions. Compared with traditional methods, the GAN-based spatio-temporal fusion method can better capture the potential spatial and temporal features in remote sensing images, thereby improving the quality of the fused images. For example, Chinese Patent Application 《CN115131637A》 discloses a multi-level feature spatio-temporal remote sensing image fusion method based on a generative adversarial network, which inputs the high-resolution real image at the target reference time and the low-resolution real image at the target reference time into a trained conditional generator network for multi-level feature spatio-temporal remote sensing image fusion to generate a high-resolution image at the target prediction time. Although it solves the defects in existing remote sensing spatio-temporal fusion that it is difficult to predict the sudden change areas of land cover types, the overall differences brought by sensors are large, and the generated images are not realistic enough, it mainly focuses on the statistical distribution of global and local features and does not explicitly model the geometric relationship between different spectral channels, resulting in the loss of some spectral information and structural information between channels during the fusion process, thus affecting the fidelity of the generated images.

[0003] Therefore, it is a technical problem to be solved to provide a spatio-temporal fusion method for remote sensing images that can retain the spectral and structural information between channels. Summary of the Invention

[0004] The object of the present invention is to overcome the defects of the above-mentioned existing technologies, and provide a method and system for remote sensing image spatio-temporal fusion based on geometric algebra generative adversarial network, which overcomes the limitations of existing spatio-temporal fusion methods in dealing with spatial and temporal dependencies, and proposes a geometric algebra generative adversarial network model combined with domain knowledge enhancement for the loss of information between channels, so as to generate remote sensing images with high spatio-temporal resolution more completely. By integrating images with high temporal resolution and low spatial resolution and images with low temporal resolution and high spatial resolution into the model, the generation accuracy and the robustness of the model are significantly improved.

[0005] The object of the present invention can be achieved by the following technical solutions:

[0006] According to the first aspect of the present invention, there is provided a method for remote sensing image spatio-temporal fusion based on geometric algebra generative adversarial network. The method uses a trained geometric algebra generative adversarial model for remote sensing image spatio-temporal fusion. The geometric algebra generative adversarial model includes a geometric algebra generator, and the geometric algebra generator includes a geometric algebra residual encoder and a geometric algebra upsampling decoder. The steps include:

[0007] Obtain multi-source remote sensing data and perform cropping and preprocessing on it, and use the geometric algebra residual encoder to process the preprocessed multi-source remote sensing data to obtain temporal features and spatial features; the multi-source remote sensing data includes a low-spatial-resolution image on the prediction date and a random high-spatial-resolution image;

[0008] Fuse the temporal features and spatial features to obtain fused features;

[0009] Input the fused features into the geometric algebra upsampling decoder to output a first reconstructed remote sensing image.

[0010] As a preferred technical solution, using the geometric algebra residual encoder, the following steps are performed on both the low-spatial-resolution image on the prediction date and the random high-spatial-resolution image to obtain temporal features and spatial features. The geometric algebra residual encoder includes multiple geometric convolutional layers:

[0011] Map each pixel in the multi-source remote sensing data to a multi-vector. Among them, when obtaining temporal features, the multi-source remote sensing data processed is the low-spatial-resolution image on the prediction date, and when obtaining spatial features, the multi-source remote sensing data processed is the random high-spatial-resolution image;

[0012] Perform geometric algebra residual convolution operations on the multi-vector to obtain temporal features or spatial features; the geometric algebra residual convolution operations include convolution processing and residual processing.

[0013] As a preferred technical solution, the method for obtaining the multi-vectors is as follows:

[0014]

[0015] where [I(x,y)] i 、[I(x,y)] ij and [I(x,y)] 1,...,ne1,...,n all represent the spectral components of the hyperspectral image pixel point I at the position (x,y).

[0016] As a preferred technical solution, the method for the convolution processing includes:

[0017] Each of the geometric convolution layers includes a plurality of neurons, and each of the neurons performs geometric algebraic convolution, and its expression is:

[0018]

[0019] where represents the output of the j-th neuron in the l-th geometric convolution layer; b represents the index set of all basis vectors in the geometric algebraic space; A represents the elements in the index set; g (l) represents the activation function of the l-th geometric convolution layer; represents the convolution kernel of the j-th neuron in the l-th geometric convolution layer; x l-1 represents the output of the (l - 1)-th geometric convolution layer; represents the bias term; e A represents the basis vector;

[0020] The outputs of all neurons in the same geometric convolution layer are aggregated to obtain the output of the corresponding set convolution layer.

[0021] As a preferred technical solution, the method for the residual processing is as follows:

[0022] y l+1 = F(x l+1 ,W l+1 ) + x l ,

[0023] where F(x l+1 ,W l+1 ) represents the multi-vectors after geometric algebraic convolution; W l+1 represents the corresponding geometric algebraic convolution kernel weights; when l = 1, x l represents the multi-vectors, and when l ≠ 1, x l represents the output of the l-th geometric algebraic layer.

[0024] As a preferred technical solution, the method for fusing the temporal features and the spatial features is as follows:

[0025]

[0026] Among them, X represents the spatial feature; μ X represents the mean of the spatial feature; σ X represents the standard deviation of the spatial feature; σ y represents the standard deviation of the temporal feature; μ y represents the mean of the temporal feature.

[0027] As a preferred technical solution, the geometric algebra generative adversarial model further includes a multi-dimensional discriminator, and the multi-dimensional discriminator includes a plurality of binary class discriminators. The method for training the geometric algebra generative adversarial model includes:

[0028] Obtain a low-spatial-resolution image, a random high-spatial-resolution image, and a high-spatial-resolution image of the predicted date for model training. After cropping and preprocessing the low-spatial-resolution image and the random high-spatial-resolution image of the predicted date, use the geometric algebra generator to generate a second reconstructed remote sensing image;

[0029] Take the high-spatial-resolution image of the predicted date and the second reconstructed remote sensing image as a detail group, perform different scale scalings, and then input them into the corresponding binary class discriminator to output a first discrimination result; among them, the high-spatial-resolution image of the predicted date is the conditional input;

[0030] Take the random high-spatial-resolution image and the high-spatial-resolution image of the predicted date as a shooting group, perform different scale scalings, and then input them into the corresponding binary class discriminator to output a second discrimination result; among them, the random high-spatial-resolution image is the conditional input;

[0031] Take the random high-spatial-resolution image and the second reconstructed remote sensing image as an effect group, perform different scale scalings, and then input them into the corresponding binary class discriminator to output a third discrimination result; among them, the random high-spatial-resolution image is the conditional input;

[0032] Perform weighted averaging on the first discrimination result, the second discrimination result, and the third discrimination result to obtain a discrimination result;

[0033] Based on the discrimination result, calculate the loss function values of the geometric algebra generator and each discriminator respectively;

[0034] Based on the loss function values of the geometric algebra generator and each discriminator, optimize the model parameters of the geometric algebra generative adversarial model.

[0035] As a preferred technical solution, the loss function of the geometric algebra generator is:

[0036]

[0037] Among them, L feature represents the L1 loss; L spectrum represents the spectral loss; L vision represents the structural loss; L gan represents the geometric algebra generator loss function, and Among them, D(G(z)) represents the discrimination result output by the discriminator when the input is the second reconstructed remote sensing image, c represents the hyperparameter of the loss function; α, β, and λ all represent weight coefficients.

[0038] As a preferred technical solution, the loss function of each of the binary class discriminators is:

[0039]

[0040] Among them, x~p(x) represents the data distribution of the conditional input; D(x) represents the discrimination result; G(z') represents the high spatial resolution image of the predicted date or the second reconstructed remote sensing image; a and b both represent hyperparameters.

[0041] According to the second aspect of the present invention, a remote sensing image spatio-temporal fusion system based on a geometric algebra generative adversarial network, the system is used to implement the above method.

[0042] Compared with the prior art, the present invention has the following beneficial effects:

[0043] 1), The present invention improves the traditional generative adversarial model using geometric algebra convolution, constructs a high-dimensional geometric algebra space in the non-Euclidean domain, maps each pixel of the acquired multi-source remote sensing image into a multi-vector, and then convolves the multi-dimensional vector using an encoder with a residual structure, ensuring that the spectral information between channels and the structural information between channels in the multi-source remote sensing image can be maximally retained during the convolution process, and using an encoder with a residual structure to process the image can maximally retain the time characteristics of the spatial resolution image, making the reconstructed remote sensing image closer to the real result, capable of coping with complex spatio-temporal changes and variable data quality, and having a wide range of application prospects.

[0044] 2), Because the present invention uses multiple binary class discriminators, when the present invention performs model training and actual application, only two images need to be input. Compared with the prior art where three images need to be input for ordinary testing and four images need to be input during training, the present invention can generate images with higher accuracy using less input data, and by forming a multi-dimensional discriminator from multiple discriminators to process the input image, problems such as distortion and error of the reconstructed image caused by reducing the number of input data can be avoided. Description of the Drawings

[0045] Figure 1 is the method flowchart of the present invention;

[0046] Figure 2 is the feature extraction flowchart of the present invention;

[0047] Figure 3 is the structure diagram of the geometric algebra generative adversarial model of the present invention;

[0048] Figure 4 is the discrimination flowchart of the multi-dimensional discriminator of the present invention. Detailed implementation manners

[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0050] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the ordinary meanings understood by those with ordinary skills in the technical field to which this application belongs. The words such as "a", "one", "kind", "the" and the like involved in this application do not represent a quantity limit and may represent a single or plural number. The terms "including", "comprising", "having" and any variations thereof involved in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device including a series of steps or modules (units) is not limited to the listed steps or units, but may further include unlisted steps or units, or may further include other steps or units inherent to these processes, methods, products or devices. The terms "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "plurality" involved in this application refers to two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific order for the objects.

[0051] Embodiment 1

[0052] Aiming at the limitations in dealing with spatial and temporal dependencies and the loss of information between channels in the spatio-temporal fusion method of remote sensing images in the prior art, the present invention proposes a spatio-temporal fusion method and system for remote sensing images based on geometric algebra generative adversarial network (GAN). Through the powerful information retention ability of geometric algebra and combined with the optimization mechanism of the generative adversarial network, spatio-temporal fusion of multi-source remote sensing data is carried out to generate high-quality remote sensing images. Compared with traditional spatio-temporal fusion methods, the present invention has higher image quality and stronger adaptability, and can cope with complex spatio-temporal changes and variable data quality.

[0053] Specifically, the method of the present invention uses a trained geometric algebra generative adversarial model to integrate images with high temporal resolution and low spatial resolution and images with low temporal resolution and high spatial resolution into the model for spatio-temporal fusion of remote sensing images, which can significantly improve the generation accuracy and the robustness of the model. The geometric algebra generative adversarial model includes a geometric algebra generator and a multi-dimensional discriminator composed of nine binary classifiers with the same structure. Among them, the geometric algebra generator is responsible for generating realistic images through spatio-temporal feature extraction and feature fusion, and inputting the generated images into the multi-dimensional discriminator to confuse the multi-dimensional discriminator. The multi-dimensional discriminator discriminates whether the images generated by the geometric algebra generator are real or fake. The process is as Figure 1 shown, including:

[0054] S1. Obtain multi-source remote sensing data, crop and preprocess it, and use a geometric algebra residual encoder to process the preprocessed multi-source remote sensing data to obtain temporal features and spatial features.

[0055] The geometric algebra generator includes a geometric algebra residual encoder and a geometric algebra upsampling decoder, which are symmetric structures and use residual blocks as basic blocks. Specifically, use the geometric algebra residual encoder to process the low-spatial-resolution image of the prediction date to obtain temporal features Use the geometric algebra residual encoder to process the random high-spatial-resolution image to obtain spatial features

[0056] S11. Select characteristic and commonly used satellites at home and abroad, and select GF satellites, landsat-8 satellites and modis satellites in descending order of spatial resolution. Use the above satellites to take multi-source remote sensing images and perform cropping and preprocessing on them. Among them, the multi-source remote sensing data includes the low-spatial-resolution image of the prediction date and the random high-spatial-resolution image; the preprocessing includes atmospheric correction, cloud mask removal, reprojection, resampling, image registration, and null value part removal.

[0057] S12. Map each pixel in the multi-source remote sensing data into a multi-vector.

[0058] According to the multi-spectral channel characteristics of multi-dimensional remote sensing images, a high-dimensional geometric algebra (GA) space is constructed in a non-Euclidean domain, and geometric algebra convolution is constructed to map each pixel in multi-source remote sensing images into a multi-vector, and its expression is:

[0059]

[0060] where, [I(x,y)] i 、[I(x,y)] ij and [I(x,y)] 1,...,ne1,...,n all represent the spectral components of the hyperspectral image pixel I at the position (x,y);

[0061] Each multi-vector is defined by a linear combination of grading elements related to the size of the geometric algebra space. In the geometric algebra space, a multi-vector consists of a real part and multiple imaginary parts, and the number of imaginary parts depends on the size of the geometric algebra space. The input multi-source remote sensing images are expressed and calculated through geometric algebra products, sharing geometric algebra components during the multiplication process, thereby reducing the number of network parameters. The output is very sensitive to geometric algebra components. Even a small change in the components can lead to a significant change in the results, which helps to capture potential relationships in the hidden layer, thereby improving the structural representation quality between spectra and channels, and realizing the modeling and feature extraction of complex spatial dependence relationships.

[0062] S13. Utilize geometric algebra residual convolution operations on the multi-vector to obtain temporal features or spatial features; the geometric algebra residual convolution operations include convolution processing and residual processing.

[0063] Specifically, when processing a low-spatial-resolution image of the predicted date, first, the multi-vector converted from it is processed by an activation function and then undergoes N geometric algebra residual convolution processes with a convolution kernel of 3×3, and then undergoes a geometric algebra residual convolution process with a convolution kernel of 1×1, and finally, after being processed by the activation function, spatial features are output; when processing a random high-spatial-resolution image, only the multi-vector converted from it needs to be processed by a geometric algebra residual convolution with a convolution kernel of 1×1 to obtain temporal features, and its process is as Figure 2 shown. Specifically, the processes of convolution processing and residual processing are as shown in steps S131~S132.

[0064] S131. Convolution processing:

[0065] Compared with the ordinary convolution operation which is carried out separately on the channels of each pixel, in the geometric algebra residual encoder provided by the present invention, the channels are relatively independent. After pixelizing into multivectors, the geometric algebra convolution kernel will act on the multivectors expressed by the pixels and perform convolution operations. Each geometric algebra convolution kernel not only performs weighted summation on the input, but also involves rotation and scaling transformations, so it can retain the geometric structure information of the input data. Specifically, the geometric algebra residual encoder provided by the present invention includes multiple multi-layer geometric convolution layers. Each geometric convolution layer includes multiple neurons, and each neuron performs geometric algebra convolution. Its expression is:

[0066]

[0067] Wherein, represents the output of the j-th neuron in the l-th geometric convolution layer; b represents the index set of all basis vectors in the geometric algebra space, which is used to describe the combination of basis vectors in the space, so as to represent the multivector composed of different basis element combinations; A represents the element in the index set; g (l) represents the activation function of the l-th geometric convolution layer; represents the convolution kernel of the j-th neuron in the l-th geometric convolution layer; x l-1 represents the output of the (l - 1)-th geometric convolution layer; represents the bias term; e A represents the basis vector.

[0068] Finally, the outputs of all neurons in the same geometric convolution layer are aggregated to obtain the output of the corresponding set convolution layer.

[0069] S132. Residual processing:

[0070] Residual connections are introduced in each convolution layer of the geometric algebra residual encoder. For the output y of a certain layer, a residual structure is formed by adding the input of the corresponding layer to the output of the layer. This structure can retain the original detailed features of the input image, that is, adding a residual structure means adding an offset on the original basis. Its expression is:

[0071] y l+1 = F(x l+1 , W l+1 ) + x l ,

[0072] Wherein, F(x l+1 , W l+1 ) represents the multivector after geometric algebra convolution; W l+1 represents the corresponding geometric algebra convolution kernel weight; when l = 1, x l represents the multivector, and when l ≠ 1, x l represents the output of the l-th geometric algebra layer.

[0073] When using residual structures to connect convolutional layers, convolutional layers with multiple convolutional kernels can extract spectral and structural information between multiple channels, retain the detailed features of the original image, and also prevent problems such as gradient explosion.

[0074] S2. Perform feature fusion on the temporal features and spatial features to obtain fused features.

[0075] During the spatio-temporal feature fusion process, remote sensing images involve not only spatial features (such as image texture, object boundaries), but also temporal features in the time dimension (such as seasonal changes, climate changes, etc.). At the feature fusion stage, the obtained features are stitched using ADAIN to ensure the accuracy of global information, eliminate spectral errors or satellite imaging differences, retain the main structure in the temporal features, and align the feature details in the spatial features. Specifically, its expression is:

[0076]

[0077] where X represents the spatial feature; μ X represents the mean of the spatial feature; σ X represents the standard deviation of the spatial feature; σ y represents the standard deviation of the temporal feature; μ y represents the mean of the temporal feature.

[0078] So far, the geometric algebra generator uses geometric algebra convolution to process the spatial and temporal features in the image. The convolutional kernel based on geometric algebra can capture high-order feature relationships in space. Compared with traditional convolutional neural networks (CNNs), GA convolution can more efficiently express complex geometric shapes and spatial interactions within the image space. Each layer of the convolutional kernel can not only process conventional image features, but also enhance the geometric information at different scales and in different directions in the image; using the residual network, the problem of gradient disappearance and information loss in deep networks is solved by introducing skip connections; in the spatio-temporal fusion task, the residual connection helps the generator better retain spatial and temporal features, avoiding performance degradation that may occur in the training of deep networks, and using ADAIN instead of conventional weighted summation for feature fusion to ensure the accuracy of global information.

[0079] S3. Input the fused features into the geometric algebra upsampling decoder to output the first reconstructed remote sensing image.

[0080] After the operations of steps S1 to S3, a reconstructed high-precision remote sensing image is obtained. In addition, the present invention also provides a training method for a geometric algebra generative adversarial model. During the training process, a multi-dimensional discriminator discriminates the images generated by the geometric algebra generator in multiple aspects and outputs a discrimination result of yes or no. Based on this result, the geometric algebra generative adversarial network is trained. Each discriminator in the multi-dimensional discriminator provided by the present invention is composed of a plurality of residual blocks, and spectral normalization is adopted during the convolution process of each residual block to stabilize the adversarial process, ensure that the gradient change of the model will not be too large, avoid the problems of gradient explosion or gradient disappearance during the training process, improve the training stability, and reduce the training imbalance between the generator and the discriminator.

[0081] The training process is as Figure 3 shown and includes:

[0082] A1. Obtain a low-spatial-resolution image, a random high-spatial-resolution image, and a high-spatial-resolution image of the prediction date for model training. After cropping and preprocessing the low-spatial-resolution image and the random high-spatial-resolution image of the prediction date, use the geometric algebra generator to generate a second reconstructed remote sensing image.

[0083] Among them, in order to ensure the authenticity of the ground object information contained in the rough information, the multi-dimensional discriminator adopts three groups of inputs as Figure 4 shown. Specifically, the high-spatial-resolution image of the prediction date and the second reconstructed remote sensing image are used as the detail group; the random high-spatial-resolution image and the high-spatial-resolution image of the prediction date are used as the shooting group; the random high-spatial-resolution image and the second reconstructed remote sensing image are used as the effect group to execute steps A2 to A4.

[0084] A2. After scaling the detail group to three different scales of 0.25, 0.5, and 1, input it into the corresponding discriminator. Among them, the high-spatial-resolution image of the prediction date is the conditional input, and finally, after being processed by the sigmoid activation function, the first discrimination result is obtained.

[0085] A3. After scaling the shooting group to three different scales of 0.25, 0.5, and 1, input it into the corresponding discriminator. Among them, the random high-spatial-resolution image is the conditional input, and finally, after being processed by the sigmoid activation function, the second discrimination result is output.

[0086] A4. After scaling the effect group to three different scales of 0.25, 0.5, and 1, input it into the corresponding discriminator. Among them, the random high-spatial-resolution image is the conditional input, and finally, after being processed by the sigmoid activation function, the third discrimination result is output.

[0087] After the scaling in steps A2 - A3, training the multi - dimensional discriminator with multi - scale inputs can enhance the discrimination ability of the multi - dimensional discriminator, ultimately promoting the optimization process of the generative adversarial network, thereby generating higher - quality images.

[0088] A5. Perform weighted averaging on the first discrimination result, the second discrimination result, and the third discrimination result to obtain the discrimination result.

[0089] A6. Calculate the loss function values of the geometric algebra generator and each discriminator respectively based on the discrimination result.

[0090] Specifically, the loss function of the geometric algebra generator is calculated based on the output result of the discriminator. By minimizing the probability that its output samples are misclassified as "false" samples by the discriminator, its generation ability is optimized. The loss function of the geometric algebra generator is usually represented by the negative log - likelihood function, and its goal is to maximize the probability that the discriminator predicts the generated samples as "true". In the present invention, the loss function of the geometric algebra generator includes the geometric algebra generator loss function, L1 loss, spectral loss, and structure loss, with the geometric algebra generator loss function as the main body. Among them, the L1 loss is used to ensure the minimization of the difference between the generated image and the real image at each pixel point; the spectral loss (Spectral Loss) is used to evaluate the similarity between the generated image and the real image in the frequency spectrum; the structure loss (Structure Loss) measures the similarity between the generated image and the real image in terms of structure by calculating the multi - scale structural similarity (MS - SSIM) of the image. Its expression is:

[0091]

[0092] where, L feature represents the L1 loss; L spectrum represents the spectral loss; L vision represents the structure loss; L gan represents the geometric algebra generator loss, and where, D(G(z)) represents the discrimination result output by the discriminator when the input is the second reconstructed remote - sensing image, c represents the hyper - parameter of the loss function; α, β, and λ all represent weight coefficients.

[0093] The loss function of the multi-dimensional discriminator is defined based on the classification results of real samples and generated samples. The goal is to maximize the probabilities of its "true" prediction for real samples and "false" prediction for generated samples. In this way, the discriminator and the generator continuously conduct adversarial training until an equilibrium state is reached, enabling the generator to generate high-quality images that are difficult to distinguish from real samples, while the discriminator can accurately distinguish between generated samples and real samples. Therefore, the least squares GAN (LSGAN) is used as its loss function. By minimizing the Euclidean distance, the output of the multi-dimensional discriminator becomes smoother, thereby improving the training stability. The loss function of each discriminator in the multi-dimensional discriminator is as follows:

[0094]

[0095] Among them, \(x\sim p(x)\) represents the data distribution of conditional input; \(D(x)\) represents the discrimination result; \(G(z')\) represents the high-spatial-resolution image of the predicted date or the second reconstructed remote sensing image; both \(a\) and \(b\) represent hyperparameters.

[0096] Through the above game process, the generative adversarial network can continuously improve the performance of the generator and finally generate samples close to the real data distribution.

[0097] A7. Optimize the model parameters of the geometric algebra generative adversarial model based on the loss function values of the geometric algebra generator and each discriminator.

[0098] Specifically, evaluate the geometric algebra generative adversarial model trained through steps A1 - A7. Using the cross-validation method, divide the data set into a training set and a test set. Repeat the training and testing processes multiple times. Each time, calculate evaluation metrics such as the root mean square error, structural similarity index, correlation coefficient, and mean absolute error to evaluate the model performance, and take their average values as the final evaluation results.

[0099] The calculation methods of each metric are as follows:

[0100] Root Mean Square Error (RMSE):

[0101]

[0102] Among them, \(N\) represents the total number of image pixels (samples); \(y\) i represents the real image pixel of the \(i\)-th multi-source remote sensing image; represents the \(i\)-th reconstructed remote sensing image pixel generated by the geometric algebra generator.

[0103] Mean Absolute Error (MAE):

[0104]

[0105] Structural Similarity Index (SSIM):

[0106]

[0107] where μ x and μ y represent the means of the real remote sensing image and the reconstructed remote sensing image generated by the geometric algebra generator respectively; σ xy represents the covariance of the real remote sensing image and the reconstructed remote sensing image generated by the geometric algebra generator; σ x and σ y represent the variances of the real remote sensing image and the reconstructed remote sensing image generated by the geometric algebra generator respectively; C1 and C2 are constants, and their function is to avoid the denominator being 0.

[0108] Correlation Coefficient (CC):

[0109]

[0110] where x i and y i represent the pixel values of the reconstructed remote sensing image generated by the geometric algebra generator, and represent the mean of the reconstructed remote sensing image generated by the geometric algebra generator.

[0111] Enhanced Relative Global Error to the Average Spectral (ERGAS):

[0112]

[0113] where D represents the image band; RMSE i represents the root mean square error value of the i-th remote sensing image; μ i represents the mean of the i-th band.

[0114] Example 2

[0115] In order to verify the feasibility and superiority of the method provided by the present invention, three groups of data on different dates were selected from the LGC dataset in this example, and comparative experiments were conducted with ESTARFM and FastVSDF.

[0116] Among them, FastVSDF adopts technologies such as FAVC classification, global guided filtering residual compensation, and Gaussian weighted local optimization. Different from the method provided by the present invention, this model will ignore the edge processing of images and does not have complex regression analysis. In order to verify the superiority of the method provided by the present invention, the root mean square error and correlation coefficient are calculated for each adopted method. Specifically, the root mean square error (RMSE) can directly test the accuracy of the generated image; the correlation coefficient (CC) is used to measure the linear correlation between the captured image and the generated image.

[0117] For each model, its performance was tested by repeating the experiment eight times on 14 pairs of MODIS-Landsat 6-band images with a size of 3200×2720 captured from 2004 to 2005 in the data, and 3 pairs were selected as the test set. Table 1 shows the mean values of each type of index.

[0118]

[0119] As can be seen from Table 1, the method provided by the present invention performs well in the above three groups of test data sets. Compared with the commonly used methods in the prior art, the generated images are more realistic, that is, the method provided by the present invention has feasibility and superiority.

[0120] Embodiment 3

[0121] The present invention also provides a remote sensing image spatio-temporal fusion system based on a geometric algebra generative adversarial network. The system includes a central processing unit (CPU), which can execute various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) or computer program instructions loaded from a storage unit into a random access memory (RAM). In the RAM, various programs and data required for device operation can also be stored. The CPU, ROM, and RAM are connected to each other through a bus. An input / output (I / O) interface is also connected to the bus.

[0122] Multiple components in the device are connected to the I / O interface, including: an input unit, such as a keyboard, a mouse, etc.; an output unit, such as various types of displays, speakers, etc.; a storage unit, such as a disk, an optical disc, etc.; and a communication unit, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit allows the device to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0123] The processing unit executes the various methods and processes described above, such as method S1-S3 and method A1-A7. For example, in some embodiments, method S1-S3 and method A1-A7 may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device via the ROM and / or the communication unit. When the computer program is loaded into the RAM and executed by the CPU, one or more steps of method S1-S3 and method A1-A7 described above may be executed. Alternatively, in other embodiments, the CPU may be configured to execute method S1-S3 and method A1-A7 by any other suitable means (e.g., by means of firmware).

[0124] The functions described above herein may be performed, at least in part, by one or more hardware logic components. By way of example, and without limitation, the types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0125] The program code for implementing the methods of the present invention may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing apparatus such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on the remote machine or server.

[0126] In the context of the present invention, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be either a machine-readable signal medium or a machine-readable storage medium. Machine-readable media may include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0127] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A spatio-temporal fusion method for remote sensing images based on geometric algebra generative adversarial network, characterized in that, The described method uses a trained geometric algebra generative adversarial model for spatio-temporal fusion of remote sensing images. The geometric algebra generative adversarial model includes a geometric algebra generator, and the geometric algebra generator includes a geometric algebra residual encoder and a geometric algebra upsampling decoder. The steps are as follows: Obtain multi-source remote sensing data, crop and preprocess it, and use the geometric algebra residual encoder to process the preprocessed multi-source remote sensing data to obtain temporal features and spatial features. The multi-source remote sensing data includes a low-spatial-resolution image of the prediction date and a random high-spatial-resolution image. Fuse the temporal features and spatial features to obtain fused features. Input the fused features into the geometric algebra upsampling decoder to output a first reconstructed remote sensing image.

2. The spatio-temporal fusion method for remote sensing images based on geometric algebra generative adversarial network according to claim 1, characterized in that Using the geometric algebra residual encoder, perform the following steps on both the low-spatial-resolution image of the prediction date and the random high-spatial-resolution image to obtain temporal features and spatial features. The geometric algebra residual encoder includes multiple geometric convolutional layers: Map each pixel in the multi-source remote sensing data to a multivector. Among them, when obtaining temporal features, the multi-source remote sensing data processed is the low-spatial-resolution image of the prediction date, and when obtaining spatial features, the multi-source remote sensing data processed is the random high-spatial-resolution image. Perform geometric algebra residual convolution operations on the multivector to obtain temporal features or spatial features. The geometric algebra residual convolution operation includes convolution processing and residual processing.

3. A spatio-temporal fusion method for remote sensing images based on geometric algebra generative adversarial network according to claim 2, characterized in that, The method for obtaining the multivector is: Among them, [I(x, y)] i , [I(x, y)] ij and [I(x, y)] 1,...,ne1,...,n all represent the spectral components of the hyperspectral image pixel point I at the position (x, y).

4. A spatio-temporal fusion method for remote sensing images based on geometric algebra generative adversarial network according to claim 2, characterized in that The method for convolution processing is: Each geometric convolutional layer includes multiple neurons, and each neuron performs geometric algebra convolution. Its expression is: Among them, represents the output of the j-th neuron in the l-th geometric convolutional layer; b represents the index set of all basis vectors in the geometric algebra space; A represents the elements in the index set; g (l) represents the activation function of the l-th geometric convolutional layer; represents the convolutional kernel of the j-th neuron in the l-th geometric convolutional layer; x l-1 represents the output of the (l - 1)-th geometric convolutional layer; represents the bias term; e A represents the basis vector; Collect the outputs of all neurons in the same geometric convolutional layer to obtain the output of the corresponding set convolutional layer.

5. A spatio-temporal fusion method for remote sensing images based on geometric algebra generative adversarial network according to claim 2, characterized in that, The method for residual processing is: y l+1 = F(x l+1 , W l+1 ) + x l , Among them, F(x l+1 , W l+1 ) represents the multivector after geometric algebra convolution; W l+1 represents the corresponding geometric algebra convolution kernel weight; when l = 1, x l represents the multivector, and when l ≠ 1, x l represents the output of the l-th geometric algebra layer.

6. A spatio-temporal fusion method for remote sensing images based on geometric algebra generative adversarial network according to claim 1, characterized in that, The method for fusing the temporal features and spatial features is: Among them, X represents the spatial feature; μ X represents the mean of the spatial feature; σ X represents the standard deviation of the spatial feature; σ y represents the standard deviation of the temporal feature; μ y represents the mean of the temporal feature.

7. A spatio-temporal fusion method for remote sensing images based on geometric algebra generative adversarial network according to claim 1, characterized in that The geometric algebra generative adversarial model further includes a multi-dimensional discriminator. The multi-dimensional discriminator includes multiple binary class discriminators. The method for training the geometric algebra generative adversarial model includes: Obtain the low-spatial-resolution image of the prediction date, the random high-spatial-resolution image, and the high-spatial-resolution image of the prediction date for model training. After cropping and preprocessing the low-spatial-resolution image of the prediction date and the random high-spatial-resolution image, use the geometric algebra generator to generate a second reconstructed remote sensing image. Use the high-spatial-resolution image of the prediction date and the second reconstructed remote sensing image as a detail group, perform different scale scalings, and input them into the corresponding binary class discriminator to output a first discrimination result. Among them, the high-spatial-resolution image of the prediction date is the conditional input. Use the random high-spatial-resolution image and the high-spatial-resolution image of the prediction date as a shooting group, perform different scale scalings, and input them into the corresponding binary class discriminator to output a second discrimination result. Among them, the random high-spatial-resolution image is the conditional input. Use the random high-spatial-resolution image and the second reconstructed remote sensing image as the effect group, perform different scale scalings on them, and then input them into the corresponding binary discriminator to output the third discrimination result; wherein, the random high-spatial-resolution image is the conditional input; Perform weighted averaging on the first discrimination result, the second discrimination result, and the third discrimination result to obtain the discrimination result; Calculate the loss function values of the geometric algebra generator and each discriminator respectively based on the discrimination result; Optimize the model parameters of the geometric algebra generative adversarial model based on the loss function values of the geometric algebra generator and each discriminator.

8. A spatio-temporal fusion method for remote sensing images based on geometric algebra generative adversarial network according to claim 7, characterized in that, The loss function of the geometric algebra generator is: Among them, L feature represents the L1 loss; L spectrum represents the spectral loss; L vision represents the structural loss; L gan represents the geometric algebra generator loss function, and where D(G(z)) represents the discrimination result output by the discriminator when the input is the second reconstructed remote sensing image, c represents the hyperparameter of the loss function; α, β, and λ all represent weight coefficients.

9. A spatio-temporal fusion method for remote sensing images based on geometric algebra generative adversarial network according to claim 7, characterized in that The loss function of each binary discriminator is: Wherein, x~p(x) represents the data distribution of the conditional input; D(x) represents the discrimination result; G(z') represents the high-spatial-resolution image of the predicted date or the second reconstructed remote sensing image; both a and b represent hyperparameters.

10. A spatio-temporal fusion system for remote sensing images based on geometric algebra generative adversarial network, characterized in that, The system is used to implement the method described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Multi-level feature space-time remote sensing image fusion method based on generative adversarial network

    CN115131637A