Remote sensing image fusion inversion method, device, equipment and medium

By resampling and stacking learning inversion models from high-resolution UAV imagery and low-resolution satellite imagery, the problem of insufficient accuracy in remote sensing image fusion in existing technologies is solved, and high-precision environmental monitoring results are achieved.

CN121746955BActive Publication Date: 2026-05-01NORTHEASTERN UNIV CHINA +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NORTHEASTERN UNIV CHINA
Filing Date
2026-02-26
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing remote sensing image fusion methods mostly remain at the level of simple data combination, resulting in the fusion image accuracy being insufficient to meet the requirements of high-precision quantitative environmental parameter inversion, especially how to make full use of the rich ground feature details contained in high-resolution images to correct or over-resolution low-resolution images and establish robust nonlinear mapping relationships.

Method used

By acquiring high-resolution UAV imagery and low-resolution satellite imagery, training samples are constructed after resampling. A stacked learning inversion model is then used for model training, including a base learner layer and a meta learner layer. The base learner layer uses multiple machine learning regression models for preliminary prediction, and the meta learner layer performs weighted fusion to finally generate target-accurate inversion images.

Benefits of technology

It achieves high-precision environmental monitoring, effectively suppresses uncertainties in the fusion process, ensures the quantitative accuracy of the inversion results, and meets the needs of high-precision quantitative environmental parameter inversion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121746955B_ABST
    Figure CN121746955B_ABST
Patent Text Reader

Abstract

The application relates to a remote sensing image fusion inversion method and device, equipment and medium, and relates to the technical fields of remote sensing and mine monitoring. The method comprises the following steps: resampling first remote sensing images and second remote sensing images to construct training samples, performing model training on a stacked learning inversion model, the stacked learning inversion model comprising a base learner layer for predicting and outputting respective preliminary prediction results of the second remote sensing images, and a meta learner layer for performing weighted fusion processing on the multiple preliminary prediction results, thereby obtaining a trained stacked learning inversion model; and inputting the second remote sensing images into the stacked learning inversion model to obtain a target inversion image. The application can suppress uncertainty generated in the fusion process, avoid loss of key detail information, improve the precision of the fused image, meet the demand for high-precision quantitative environmental parameter inversion, fully utilize rich ground object details contained in high-resolution images to correct or super-resolution low-resolution images, and establish a robust nonlinear mapping relationship between the two.
Need to check novelty before this filing date? Find Prior Art

Description

Remote sensing image fusion and inversion methods, devices, equipment and media Technical Field

[0001] This application relates to the fields of remote sensing technology and mining area monitoring technology, specifically to a remote sensing image fusion and inversion method, device, equipment and medium. Background Technology

[0002] The mining of resources such as coal is characterized by its large scale and deep disturbance, significantly impacting ecological elements in mining areas, including surface structure, vegetation cover, soil quality, and water resources. Therefore, dynamic, continuous, and precise monitoring of the ecological environment in mining areas is crucial for guiding rational resource development, implementing effective ecological restoration, and evaluating the effectiveness of environmental governance.

[0003] Remote sensing technology, with its high efficiency and wide coverage, has become an important tool for ecological and environmental monitoring. However, single remote sensing data sources have inherent limitations. On the one hand, satellite remote sensing imagery, such as Sentinel-2, can provide free, long-term global coverage data, but its spatial resolution (usually 10 meters or lower) is limited, making it difficult to capture small-scale, detailed changes in land features such as surface subsidence and vegetation degradation in mining areas. On the other hand, unmanned aerial vehicle (UAV) remote sensing technology can acquire centimeter-level high-resolution imagery, with flexible data acquisition and good real-time performance. However, limited by endurance and operating costs, its monitoring range is usually small, making it difficult to conduct large-scale continuous observations, and it lacks traceable historical imagery data.

[0004] To combine the advantages of both high-resolution and low-resolution images and compensate for the shortcomings of a single data source, academia and industry have begun to explore multi-source data collaborative monitoring technologies. However, existing fusion methods mostly remain at the level of simple data "combination," such as simply georegistering two images and then overlaying them, or using traditional spatiotemporal data fusion algorithms (such as STARFM). These methods often introduce new uncertainties during the fusion process, or lose key details due to model simplification, resulting in the accuracy of the fused image still failing to meet the requirements for high-precision quantitative environmental parameter inversion (such as vegetation index NDVI, biomass, etc.). In particular, how to fully utilize the rich ground feature details contained in high-resolution images to "correct" or "over-resolution" low-resolution images, and establish a robust nonlinear mapping relationship between the two, is the core challenge currently facing the technology.

[0005] Therefore, there is an urgent need for a new fusion inversion method that can truly leverage the high spatial resolution advantage of UAV imagery to satellite imagery, in order to generate remote sensing data products that combine high precision and long time series characteristics, providing reliable data support for refined environmental monitoring in mining areas and other regions. Summary of the Invention

[0006] In view of this, this application provides a remote sensing image fusion and inversion method, apparatus, equipment, and medium. The main purpose is to solve the problem that existing fusion methods mostly remain at the level of simple data "combination". In the fusion process, new uncertainties are often introduced, or key details are lost due to model simplification. As a result, the accuracy of the fused image is still difficult to meet the requirements of high-precision quantitative environmental parameter inversion. In particular, it addresses the technical problem of how to make full use of the rich ground feature details contained in high-resolution images to "correct" or "over-resolution" low-resolution images and establish a robust nonlinear mapping relationship between the two.

[0007] Firstly, this application provides a remote sensing image fusion and inversion method, including:

[0008] Acquire a first remote sensing image and a second remote sensing image of the target area, and perform resampling processing on the first remote sensing image and the second remote sensing image to unify the spatial resolution of the first remote sensing image and the second remote sensing image, wherein the first remote sensing image is a high-resolution UAV image and the second remote sensing image is a low-resolution satellite image;

[0009] Training samples are constructed using the resampled first and second remote sensing images to train the stacked learning inversion model. The stacked learning inversion model includes a base learner layer and a meta learner layer. The base learner layer contains at least two machine learning regression models, which are used to predict the second remote sensing image and output their respective preliminary prediction results. The meta learner layer is used to perform weighted fusion processing on the multiple preliminary prediction results output by the base learner layer. The first remote sensing image is used as the ground truth to optimize the model and obtain the trained stacked learning inversion model.

[0010] The second remote sensing image to be processed is input into the stacked learning inversion model that has been trained to perform image inversion processing, thereby obtaining the target-accuracy inversion image.

[0011] Secondly, this application provides a remote sensing image fusion and inversion device, comprising:

[0012] The acquisition module is used to acquire a first remote sensing image and a second remote sensing image of the target area, and to resample the first remote sensing image and the second remote sensing image to unify the spatial resolution of the first remote sensing image and the second remote sensing image, wherein the first remote sensing image is a high-resolution UAV image and the second remote sensing image is a low-resolution satellite image.

[0013] The training module is used to construct training samples using the resampled first remote sensing image and the second remote sensing image, and to train the stacked learning inversion model using the training samples. The stacked learning inversion model includes a base learner layer and a meta learner layer. The base learner layer contains at least two machine learning regression models, which are used to predict the second remote sensing image and output their respective preliminary prediction results. The meta learner layer is used to perform weighted fusion processing on the multiple preliminary prediction results output by the base learner layer, and to optimize the model using the first remote sensing image as the ground truth, so as to obtain the trained stacked learning inversion model.

[0014] The processing module is used to input the second remote sensing image to be processed into the trained stacked learning inversion model for image inversion processing to obtain the target accuracy inversion image.

[0015] Thirdly, this application provides an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the computer program to implement the remote sensing image fusion and inversion method described in the first aspect.

[0016] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the remote sensing image fusion and inversion method described in the first aspect.

[0017] By employing the above technical solution, this application provides a remote sensing image fusion and inversion method, apparatus, device, and medium. Compared with existing technologies, this application can acquire a first remote sensing image and a second remote sensing image of a target area, and resample the first and second remote sensing images to unify their spatial resolution. The first remote sensing image is a high-resolution UAV image, and the second remote sensing image is a low-resolution satellite image. Training samples are constructed using the resampled first and second remote sensing images to train a stacked learning inversion model. The stacked learning inversion model includes a base learner layer and a meta learner layer. The base learner layer contains at least two machine learning regression models, each used to predict the second remote sensing image and output its own preliminary prediction results. The meta learner layer performs weighted fusion processing on the multiple preliminary prediction results output by the base learner layer, using the first remote sensing image as the ground truth for model optimization, resulting in a trained stacked learning inversion model. The second remote sensing image to be processed is then input into the trained stacked learning inversion model for image inversion processing to obtain a target-accuracy inversion image.

[0018] Using the above technical solution, the stacked learning approach of this application employs a two-level structure of "base learner layer" and "meta learner layer." Instead of directly combining pixel values, it allows the base learners to capture different feature patterns in the data, which are then intelligently fused by the meta learners. This represents a deep fusion at the feature and decision levels, uncovering the underlying logic within the data and avoiding superficial data "patchwork." Traditional simple fusion often results in spectral distortion. This application utilizes high-resolution UAV imagery as ground truth for supervised training, with the model aiming to continuously approximate the high-precision ground truth. This "ground truth calibration" method effectively suppresses uncertainties generated during the fusion process, ensuring the quantitative accuracy of the inversion results.

[0019] This application uses "quantile correction" and "ecological zoning correction" in the basic processing layer to make targeted and refined adjustments to different land features (such as water bodies, vegetation, and buildings), avoiding the loss of local details caused by global uniform processing; and extracts "multidimensional statistical features (skewness, kurtosis, etc.)" and "nonlinear transformation features" in the advanced processing layer to make the rich details hidden in the high-resolution image explicit for the model to learn.

[0020] The base learner layer employs at least two heterogeneous machine learning regression models, such as Random Forest, XGBoost, and SVR. Different models exhibit varying sensitivities to data distribution, nonlinear relationships, and noise. This ensemble strategy avoids underfitting caused by simplistic assumptions in a single model, and captures ground feature details from different perspectives, thus meeting the requirements for high-precision quantitative inversion.

[0021] During training, the model directly uses high-resolution UAV imagery (first remote sensing imagery) as the supervision label (ground truth). When processing low-resolution satellite imagery (second remote sensing imagery), the optimization objective is to predict results that are infinitely close to the UAV ground truth. Essentially, this process utilizes the rich ground feature details and texture information in the high-resolution imagery to "guide" and "correct" the pixel values ​​of the low-resolution imagery, achieving "super-resolution" reconstruction of the low-resolution imagery. By resampling to unify the spatial resolution, the high-resolution and low-resolution images are strictly aligned at the pixel level, ensuring that each low-resolution pixel can find corresponding high-resolution details for learning and establishing a robust non-linear mapping relationship between the two.

[0022] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0023] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0024] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 is a flowchart illustrating a remote sensing image fusion and inversion method provided in an embodiment of this application;

[0026] Figure 2 is a flowchart illustrating another remote sensing image fusion and inversion method provided in an embodiment of this application;

[0027] Figure 3 is a comparison diagram of an original satellite NDVI image, an inverted NDVI image, and an UAV NDVI image provided in an embodiment of this application;

[0028] Figure 4 is a comparison of the inversion results of different models of the base learner layer under different resampled data according to an embodiment of this application;

[0029] Figure 5 is a comparison chart of MAPE values ​​after processing by different resampling methods according to an embodiment of this application;

[0030] Figure 6 is a schematic diagram illustrating the effect of a stacked learning model on improving inversion accuracy according to an embodiment of this application;

[0031] Figure 7 is a schematic diagram of a remote sensing image fusion and inversion device provided in an embodiment of this application. Detailed Implementation

[0032] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of this application, including various details to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0033] The remote sensing image fusion and inversion method, apparatus, device, and medium of this application are described below with reference to the accompanying drawings.

[0034] This application provides a remote sensing image fusion and inversion method, apparatus, equipment, and medium. The main purpose is to address the problem that existing fusion methods often remain at the level of simple data "combination," which often introduces new uncertainties during the fusion process or loses key details due to model simplification. As a result, the accuracy of the fused image is still insufficient to meet the requirements of high-precision quantitative environmental parameter inversion. In particular, it addresses the technical problem of how to fully utilize the rich ground feature details contained in high-resolution images to "correct" or "over-resolution" low-resolution images and establish a robust nonlinear mapping relationship between the two.

[0035] As shown in Figure 1, an embodiment of this application provides a remote sensing image fusion and inversion method, including:

[0036] Step 101: Acquire the first and second remote sensing images of the target area, and resample the first and second remote sensing images to unify their spatial resolution.

[0037] The first remote sensing image is a high-resolution UAV image, and the second remote sensing image is a low-resolution satellite image. In this embodiment, as shown in Figure 2, the Sentinel-2 L2A level image (spatial resolution 10 meters) of the study area can be acquired as the second remote sensing image, and the image acquired on the same day by a DJI M210 RTK UAV equipped with an X5S multispectral camera (ground sampling interval 1.8 cm) can be acquired as the first remote sensing image.

[0038] In this embodiment of the disclosure, resampling processing is performed on the first remote sensing image and the second remote sensing image to unify their spatial resolution. Specifically, this may include:

[0039] First, atmospheric correction, radiometric calibration, and georegistration are performed on the first and second remote sensing images, respectively. The normalized vegetation index (NDVI) values ​​of the first and second remote sensing images are calculated to eliminate the influence of atmospheric and sensor attitude factors on the data and to ensure that the two images are strictly aligned in geospatial space. Atmospheric correction can eliminate the influence of the atmosphere on the remote sensing images, radiometric calibration can convert the digital quantization values ​​of the images into radiance values, and georegistration ensures that the images have accurate geographic coordinate information.

[0040] The formula for calculating the normalized vegetation index is shown below:

[0041]

[0042] In the formula, The normalized vegetation index value is used. For near-infrared reflectivity, Reflectivity in the red band;

[0043] Then, a resampling algorithm can be used to unify the spatial resolution of the normalized vegetation index (NDI) values ​​of the first remote sensing image and the second remote sensing image; wherein, the resampling algorithm may include at least one of nearest neighbor sampling, bilinear interpolation, or cubic convolution interpolation.

[0044] In this embodiment, considering that cubic convolution interpolation performs better in preserving image details and smoothing edges, the cubic convolution resampling method is preferred. Specifically, the NDVI data of the Sentinel-2 image is downsampled and the NDVI data of the UAV image is upsampled so that the spatial resolution of both reaches 0.1 meters.

[0045] In this embodiment of the disclosure, after resampling the first remote sensing image and the second remote sensing image to unify the spatial resolution of the first remote sensing image and the second remote sensing image, the method further includes: constructing a two-layer preprocessing model to improve data quality and model performance.

[0046] Specifically, the two-layer preprocessing model may include a first layer of basic processing and a second layer of advanced processing.

[0047] The basic processing layer can be used to perform statistical analysis and quantile correction on the normalized vegetation index (NDI) values ​​of the resampled first remote sensing image and the NDI values ​​of the second remote sensing image, as well as ecology-based zoning correction. This step specifically includes:

[0048] The statistical analysis values ​​of the Normalized Difference Vegetation Index (NDVI) of the first remote sensing image and the NDVI of the second remote sensing image are calculated. The statistical analysis values ​​may include the minimum, maximum, mean, median and standard deviation. For example, the minimum value of the UAV NDVI data is 0.1, the maximum value is 0.8 and the mean is 0.5. Satellite NDVI data also has corresponding statistical analysis values ​​to give a preliminary understanding of the distribution characteristics of the data.

[0049] The preset quantile values ​​of the normalized vegetation index (NDVI) of the second remote sensing image are adjusted to match the corresponding quantile values ​​of the NDVI of the first remote sensing image. Specifically, the 10%, 25%, 50%, 75%, and 90% quantiles of the satellite NDVI data can be precisely adjusted to match the values ​​of the UAV NDVI data at the corresponding quantiles. For example, if the 50% quantile of the satellite NDVI data is originally 0.3, it can be adjusted to match the 50% quantile of the UAV NDVI data (0.4), thereby reducing the difference in distribution between the two types of data.

[0050] Based on the range of Normalized Difference Vegetation Index (NDVI) values, the first and second remote sensing images are divided into different ecological regions. Each ecological region includes at least water bodies (lower NDVI values, e.g., less than 0.2), built-up areas (NDVI values ​​close to 0), and vegetation areas (higher NDVI values, e.g., greater than 0.3). A preset multiplier is applied to each ecological region for regional adjustment to make targeted adjustments based on the actual characteristics of different land features and improve the physical rationality of the data. For example, for vegetation areas, a multiplier of 1.1 can be applied for adjustment to more accurately reflect the true vegetation situation.

[0051] Finally, advanced processing layers can be used to extract multidimensional features from the normalized vegetation index (NDI) values ​​of the first remote sensing image and the second remote sensing image after processing by the basic processing layer.

[0052] Multidimensional features may include: advanced statistical features and nonlinear transformation features;

[0053] Advanced statistical features can be used to describe the shape and spatial correlation of data distribution, including skewness, kurtosis and autocorrelation coefficient. Skewness can be used to describe the asymmetry of data distribution, kurtosis can be used to describe the sharpness of data distribution, and autocorrelation coefficient can reflect the correlation of data at different locations.

[0054] Nonlinear transformation features can be used to enhance the model's ability to handle complex nonlinear relationships, enabling the model to capture patterns that are difficult to discover using only the original pixels. These features can include those obtained by applying logarithmic transformation, square root transformation, and reciprocal transformation to the data. For example, after performing a logarithmic transformation on NDVI data, the originally nonlinear data may exhibit more obvious linear characteristics, which is easier for subsequent model processing.

[0055] Training samples are constructed using the multidimensional features obtained through the above processing. These multidimensional features can more comprehensively reflect the characteristics of the data and provide richer information for model training.

[0056] Step 102: Construct training samples using the resampled first and second remote sensing images, and use the training samples to train the stacked learning inversion model.

[0057] The stacked learning inversion model can include a base learner layer and a meta learner layer. The base learner layer can contain at least two machine learning regression models, which can be used to predict the second remote sensing image (i.e., low-resolution image) and output their respective preliminary prediction results. The meta learner layer can be used to perform weighted fusion processing on the multiple preliminary prediction results output by the base learner layer, and use the first remote sensing image (i.e., high-resolution image) as the ground value to optimize the model and obtain the trained stacked learning inversion model.

[0058] Specifically, the base learner layer of the stacked learning inversion model may include at least two machine learning regression models among the following: Random Forest (RF) model, XGBoost (XGB) model, Support Vector Regression (SVR) model, Gradient Boosting (GB) model, LightGBM model, or CatBoost model.

[0059] In this embodiment, the base learner layer can use the six machine learning models described above to process data in parallel. Each machine learning regression model determines its optimal configuration through optimal parameter search to ensure that each model achieves its best performance. For example, for the random forest model, the optimal parameter combination on the validation set is found by adjusting parameters such as the number of trees and the maximum depth. These different types and configurations of machine learning regression models can capture complex relationships in the data from different perspectives and give full play to their respective advantages.

[0060] The meta-learner layer of the stacked learning inversion model employs a regularized linear regression model, which can be at least one of Lasso regression, Ridge regression, or ElasticNet regression. In this embodiment, the six prediction results output by the base learner layer are weighted and fused using the above three regularized linear regression models. The meta-learner layer aims to effectively integrate the prediction results of the base model, prevent overfitting, and improve the model's generalization ability. In this embodiment, experimental verification shows that ElasticNet regression performs optimally.

[0061] Step 103: Input the second remote sensing image to be processed into the trained stacked learning inversion model for image inversion processing to obtain the target accuracy inversion image.

[0062] In this embodiment of the disclosure, the Sentinel-2 satellite imagery (the second remote sensing imagery) to be processed is input into a trained stacked learning inversion model. The stacked learning inversion model processes the input satellite imagery, making predictions through multiple machine learning regression models in the base learner layer, and outputting their respective preliminary prediction results. Then, the meta-learner layer performs weighted fusion processing on these preliminary prediction results to generate the final target-precision inversion image. This inversion image has high accuracy and can more accurately reflect environmental information such as vegetation conditions in the mining area, as shown in Figure 3.

[0063] To verify the effectiveness of this method, a result verification experiment was conducted in this embodiment. The dataset was divided into training and test sets at 80% and 20% respectively. Using the UAV NDVI as the ground truth, the inversion results of the original Sentinel-2NDVI, the Sentinel-2NDVI after only resampling, and the Sentinel-2NDVI after processing with the complete method of this application were compared, as shown in Figure 4.

[0064] The results show that, after employing triple convolutional resampling and the ElasticNet stacked learning model, as shown in Figure 5, the mean absolute percentage error (MAPE) between the NDVI inversion values ​​of Sentinel-2 imagery and the ground truth NDVI values ​​from the UAV decreased from the initial 54.31% to 10.01%, a reduction of 44.30%, and the accuracy was significantly improved, as shown in Figure 6. This demonstrates that the method proposed in this application can effectively integrate two data sources to achieve high-precision environmental monitoring, providing a reliable data foundation for the long-term evolution analysis of the mining area's ecological environment.

[0065] In summary, according to the remote sensing image fusion and inversion method provided in this application, compared with the existing technology, this application can acquire a first remote sensing image and a second remote sensing image of the target area, and resample the first and second remote sensing images to unify their spatial resolution. The first remote sensing image is a high-resolution UAV image, and the second remote sensing image is a low-resolution satellite image. Training samples are constructed using the resampled first and second remote sensing images to train a stacked learning inversion model. The stacked learning inversion model includes a base learner layer and a meta-learner layer. The base learner layer contains at least two machine learning regression models, which are used to predict the second remote sensing image and output their respective preliminary prediction results. The meta-learner layer performs weighted fusion processing on the multiple preliminary prediction results output by the base learner layer, using the first remote sensing image as the ground truth for model optimization, resulting in a trained stacked learning inversion model. The second remote sensing image to be processed is then input into the trained stacked learning inversion model for image inversion processing to obtain the target-accuracy inversion image.

[0066] Using the above technical solution, the stacked learning approach of this application employs a two-level structure of "base learner layer" and "meta learner layer." Instead of directly combining pixel values, it allows the base learners to capture different feature patterns in the data, which are then intelligently fused by the meta learners. This represents a deep fusion at the feature and decision levels, uncovering the underlying logic within the data and avoiding superficial data "patchwork." Traditional simple fusion often results in spectral distortion. This application utilizes high-resolution UAV imagery as ground truth for supervised training, with the model aiming to continuously approximate the high-precision ground truth. This "ground truth calibration" method effectively suppresses uncertainties generated during the fusion process, ensuring the quantitative accuracy of the inversion results.

[0067] This application uses "quantile correction" and "ecological zoning correction" in the basic processing layer to make targeted and refined adjustments to different land features (such as water bodies, vegetation, and buildings), avoiding the loss of local details caused by global uniform processing; and extracts "multidimensional statistical features (skewness, kurtosis, etc.)" and "nonlinear transformation features" in the advanced processing layer to make the rich details hidden in the high-resolution image explicit for the model to learn.

[0068] The base learner layer employs at least two heterogeneous machine learning regression models, such as Random Forest, XGBoost, and SVR. Different models exhibit varying sensitivities to data distribution, nonlinear relationships, and noise. This ensemble strategy avoids underfitting caused by simplistic assumptions in a single model, and captures ground feature details from different perspectives, thus meeting the requirements for high-precision quantitative inversion.

[0069] During training, the model directly uses high-resolution UAV imagery (first remote sensing imagery) as the supervision label (ground truth). When processing low-resolution satellite imagery (second remote sensing imagery), the optimization objective is to predict results that are infinitely close to the UAV ground truth. Essentially, this process utilizes the rich ground feature details and texture information in the high-resolution imagery to "guide" and "correct" the pixel values ​​of the low-resolution imagery, achieving "super-resolution" reconstruction of the low-resolution imagery. By resampling to unify the spatial resolution, the high-resolution and low-resolution images are strictly aligned at the pixel level, ensuring that each low-resolution pixel can find corresponding high-resolution details for learning and establishing a robust non-linear mapping relationship between the two.

[0070] Based on the specific implementation of the method shown in Figure 1 above, this embodiment provides a remote sensing image fusion and inversion device, as shown in Figure 7. The device includes: an acquisition module 31, a training module 32, and a processing module 33.

[0071] The acquisition module 31 is used to acquire a first remote sensing image and a second remote sensing image of the target area, and to perform resampling processing on the first remote sensing image and the second remote sensing image to unify the spatial resolution of the first remote sensing image and the second remote sensing image, wherein the first remote sensing image is a high-resolution UAV image and the second remote sensing image is a low-resolution satellite image.

[0072] Training module 32 is used to construct training samples using the resampled first remote sensing image and second remote sensing image, and to train the stacked learning inversion model using the training samples. The stacked learning inversion model includes a base learner layer and a meta learner layer. The base learner layer contains at least two machine learning regression models, which are used to predict the second remote sensing image and output their respective preliminary prediction results. The meta learner layer is used to perform weighted fusion processing on the multiple preliminary prediction results output by the base learner layer, and to optimize the model using the first remote sensing image as the ground truth, so as to obtain the trained stacked learning inversion model.

[0073] The processing module 33 is used to input the second remote sensing image to be processed into the trained stacked learning inversion model for image inversion processing to obtain the target accuracy inversion image.

[0074] In specific application scenarios, the acquisition module 31 can be used to perform atmospheric correction, radiometric calibration, and georegistration on the first remote sensing image and the second remote sensing image respectively, and calculate the normalized vegetation index (NDI) values ​​of the first remote sensing image and the second remote sensing image; and use a resampling algorithm to unify the spatial resolution of the NDI values ​​of the first remote sensing image and the second remote sensing image, wherein the resampling algorithm includes at least one of nearest neighbor sampling, bilinear interpolation, or cubic convolution interpolation.

[0075] In specific application scenarios, as shown in Figure 7, the device also includes: a preprocessing module 34;

[0076] The preprocessing module 34 is used to construct a two-layer preprocessing model, which includes a first-layer basic processing layer and a second-layer advanced processing layer. The basic processing layer performs statistical analysis, quantile correction, and ecological zoning correction on the normalized vegetation index (NVI) values ​​of the resampled first and second remote sensing images. The advanced processing layer extracts multidimensional features from the NVI values ​​of the first and second remote sensing images after processing by the basic processing layer. These multidimensional features include advanced statistical features and nonlinear transformation features. The advanced statistical features include skewness, kurtosis, and autocorrelation coefficient, while the nonlinear transformation features include features obtained by applying logarithmic, square root, and reciprocal transformations to the data.

[0077] In specific application scenarios, the training module 32 can be used to construct the training samples using the multidimensional features.

[0078] In specific application scenarios, the preprocessing module 34 can be used to calculate the statistical analysis values ​​of the normalized vegetation index (NDI) values ​​of the first remote sensing image and the second remote sensing image. The statistical analysis values ​​include the minimum, maximum, mean, median, and standard deviation of the NDI values ​​of the first and second remote sensing images. It can also adjust the preset quantile values ​​of the NDI values ​​of the second remote sensing image to match the corresponding quantile values ​​of the NDI values ​​of the first remote sensing image. Furthermore, it can divide the first and second remote sensing images into different ecological regions based on the range of the NDI values, and apply a preset multiplier to each ecological region for regional adjustment. The ecological regions include at least water bodies, building areas, and vegetation areas.

[0079] In specific application scenarios, the training module 32 can be used for the base learner layer of the stacked learning inversion model, including at least two machine learning regression models among the random forest model, XGBoost model, support vector regression model, gradient boosting model, LightGBM model, or CatBoost model, and the optimal configuration of each machine learning regression model is determined through optimal parameter search.

[0080] In specific application scenarios, the training module 32 can be used to employ a regularized linear regression model for the meta-learner layer of the stacked learning inversion model. The regularized linear regression model is at least one of Lasso regression, Ridge regression, or ElasticNet regression.

[0081] It should be noted that other corresponding descriptions of the functional units involved in the remote sensing image fusion and inversion device provided in this embodiment can be found in the corresponding descriptions in Figure 1, and will not be repeated here.

[0082] Based on the method shown in Figure 1, this embodiment also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the method shown in Figure 1.

[0083] Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause a computer device (such as personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of this application.

[0084] Based on the method shown in FIG1 and the virtual device embodiment shown in FIG7, in order to achieve the above objectives, this application embodiment also provides an electronic device, which includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to implement the method shown in FIG1.

[0085] Optionally, the aforementioned physical devices may also include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, input units such as a keyboard, etc., and optional user interfaces may also include USB interfaces, card reader interfaces, etc. The network interface may optionally include standard wired interfaces, wireless interfaces (such as Wi-Fi interfaces), etc.

[0086] Those skilled in the art will understand that the physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.

[0087] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the aforementioned physical device, supporting the operation of the remote sensing image fusion and inversion program, as well as other software and / or programs. The network communication module is used to enable communication between the various components within the storage medium, as well as communication with other hardware and software in the remote sensing image fusion and inversion physical device.

[0088] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platform, or it can be implemented by hardware. By applying the scheme of this embodiment, compared with the existing technology, this application can acquire a first remote sensing image and a second remote sensing image of the target area, and resample the first and second remote sensing images to unify their spatial resolution. The first remote sensing image is a high-resolution UAV image, and the second remote sensing image is a low-resolution satellite image. Training samples are constructed using the resampled first and second remote sensing images to train a stacked learning inversion model. The stacked learning inversion model includes a base learner layer and a meta learner layer. The base learner layer contains at least two machine learning regression models, which are used to predict the second remote sensing image and output their respective preliminary prediction results. The meta learner layer is used to perform weighted fusion processing on the multiple preliminary prediction results output by the base learner layer, and the model is optimized using the first remote sensing image as the ground truth to obtain a trained stacked learning inversion model. The second remote sensing image to be processed is input into the trained stacked learning inversion model for image inversion processing to obtain the target accuracy inversion image.

[0089] Using the above technical solution, the stacked learning approach of this application employs a two-level structure of "base learner layer" and "meta learner layer." Instead of directly combining pixel values, it allows the base learners to capture different feature patterns in the data, which are then intelligently fused by the meta learners. This represents a deep fusion at the feature and decision levels, uncovering the underlying logic within the data and avoiding superficial data "patchwork." Traditional simple fusion often results in spectral distortion. This application utilizes high-resolution UAV imagery as ground truth for supervised training, with the model aiming to continuously approximate the high-precision ground truth. This "ground truth calibration" method effectively suppresses uncertainties generated during the fusion process, ensuring the quantitative accuracy of the inversion results.

[0090] This application uses "quantile correction" and "ecological zoning correction" in the basic processing layer to make targeted and refined adjustments to different land features (such as water bodies, vegetation, and buildings), avoiding the loss of local details caused by global uniform processing; and extracts "multidimensional statistical features (skewness, kurtosis, etc.)" and "nonlinear transformation features" in the advanced processing layer to make the rich details hidden in the high-resolution image explicit for the model to learn.

[0091] The base learner layer employs at least two heterogeneous machine learning regression models, such as Random Forest, XGBoost, and SVR. Different models exhibit varying sensitivities to data distribution, nonlinear relationships, and noise. This ensemble strategy avoids underfitting caused by simplistic assumptions in a single model, and captures ground feature details from different perspectives, thus meeting the requirements for high-precision quantitative inversion.

[0092] During training, the model directly uses high-resolution UAV imagery (first remote sensing imagery) as the supervision label (ground truth). When processing low-resolution satellite imagery (second remote sensing imagery), the optimization objective is to predict results that are infinitely close to the UAV ground truth. Essentially, this process utilizes the rich ground feature details and texture information in the high-resolution imagery to "guide" and "correct" the pixel values ​​of the low-resolution imagery, achieving "super-resolution" reconstruction of the low-resolution imagery. By resampling to unify the spatial resolution, the high-resolution and low-resolution images are strictly aligned at the pixel level, ensuring that each low-resolution pixel can find corresponding high-resolution details for learning and establishing a robust non-linear mapping relationship between the two.

[0093] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the term "comprising" or any other variations thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

[0094] The above are merely specific embodiments of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to these embodiments, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A remote sensing image fusion and inversion method, characterized in that, include: First and second remote sensing images of the target area are acquired, and resampling is performed on the first and second remote sensing images to unify their spatial resolution. The first remote sensing image is a high-resolution UAV image, and the second remote sensing image is a low-resolution satellite image. Training samples are constructed using the resampling first and second remote sensing images to train a stacked learning inversion model. The stacked learning inversion model includes a base learner layer and a meta learner layer. The base learner layer contains at least two machine learning regression models, each used to predict the second remote sensing image and output its own preliminary prediction results. The meta learner layer performs weighted fusion processing on the multiple preliminary prediction results output by the base learner layer, using the first remote sensing image as the ground truth for model optimization, resulting in the trained stacked learning inversion model. The second remote sensing image to be processed is then input into the training... The stacked learning inversion model, once trained, is used for image inversion processing to obtain a target-precision inverted image. After resampling the first and second remote sensing images to unify their spatial resolution, the method further includes: constructing a two-layer preprocessing model, comprising a first-layer basic processing layer and a second-layer advanced processing layer; using the basic processing layer to perform statistical analysis, quantile correction, and ecological zoning correction on the normalized vegetation index (NVI) values ​​of the resampled first and second remote sensing images; using the advanced processing layer to extract multidimensional features from the NVI values ​​of the first and second remote sensing images processed by the basic processing layer; and constructing training samples using the resampled first and second remote sensing images, specifically including constructing the training samples using the multidimensional features.

2. The method according to claim 1, characterized in that, The resampling process for the first and second remote sensing images to unify their spatial resolution specifically includes: performing atmospheric correction, radiometric calibration, and georegistration on the first and second remote sensing images respectively, and calculating the normalized vegetation index (NVI) values ​​of the first and second remote sensing images; and using a resampling algorithm to unify the spatial resolution of the NVI values ​​of the first and second remote sensing images, wherein the resampling algorithm includes at least one of nearest neighbor sampling, bilinear interpolation, or cubic convolution interpolation.

3. The method according to claim 1, characterized in that, The process of using the basic processing layer to perform statistical analysis, quantile correction, and ecological zoning correction on the normalized vegetation index (NDI) values ​​of the resampled first and second remote sensing images specifically includes: calculating the statistical analysis values ​​of the NDI values ​​of the first and second remote sensing images, where the statistical analysis values ​​include the minimum, maximum, mean, median, and standard deviation of the NDI values ​​of the first and second remote sensing images; adjusting the preset quantile values ​​of the NDI values ​​of the second remote sensing image to match the corresponding quantile values ​​of the NDI values ​​of the first remote sensing image; dividing the first and second remote sensing images into different ecological regions according to the range of the NDI values, and applying a preset multiplier to each ecological region for regional adjustment, wherein the ecological region includes at least water bodies, built-up areas, and vegetation areas.

4. The method according to claim 1, characterized in that, The multidimensional features include: advanced statistical features and nonlinear transformation features. The advanced statistical features include skewness, kurtosis, and autocorrelation coefficient. The nonlinear transformation features include features obtained by applying logarithmic transformation, square root transformation, and reciprocal transformation to the data.

5. The method according to claim 1, characterized in that, The base learner layer of the stacked learning inversion model includes at least two machine learning regression models selected from the following: random forest model, XGBoost model, support vector regression model, gradient boosting model, LightGBM model, or CatBoost model. The optimal configuration of each machine learning regression model is determined through optimal parameter search.

6. The method according to claim 1, characterized in that, The meta-learner layer of the stacked learning inversion model adopts a regularized linear regression model, which is at least one of Lasso regression, Ridge regression, or ElasticNet regression.

7. A remote sensing image fusion and inversion device, characterized in that, include: An acquisition module is used to acquire a first remote sensing image and a second remote sensing image of the target area, and to resample the first and second remote sensing images to unify their spatial resolution. The first remote sensing image is a high-resolution UAV image, and the second remote sensing image is a low-resolution satellite image. A training module is used to construct training samples using the resampled first and second remote sensing images to train a stacked learning inversion model. The stacked learning inversion model includes a base learner layer and a meta-learner layer. The base learner layer contains at least two machine learning regression models, each used to predict the second remote sensing image and output its own preliminary prediction results. The meta-learner layer performs weighted fusion processing on the multiple preliminary prediction results output by the base learner layer. The first remote sensing image is used as the ground truth for model optimization to obtain the trained stacked learning inversion model. A processing module is used to input the second remote sensing image to be processed into the trained stacked learning inversion model for image inversion processing, obtaining a target accuracy inversion image. A construction module is used to construct a two-layer preprocessing model, which includes a first basic processing layer and a second advanced processing layer. The basic processing layer performs statistical analysis, quantile correction, and ecological zoning correction on the normalized vegetation index (NVI) values ​​of the resampled first and second remote sensing images. The advanced processing layer extracts multidimensional features from the NVI values ​​of the first and second remote sensing images after processing by the basic processing layer. The multidimensional features are then used to construct the training samples.

8. An electronic device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 6.

Citation Information

Patent Citations

  • Pathological image super-resolution modeling method based on deep learning

    CN112785498A

  • Remote sensing image space-time fusion method and system based on wavelet domain cross pairing

    CN117036987A