Remote sensing image fusion inversion method, device, equipment and medium
By stacking the base learner layer and meta learner layer of the learning inversion model and using high-resolution UAV imagery as ground truth for supervised training, the uncertainty and loss of detail information in existing remote sensing image fusion methods are solved, achieving high-precision environmental parameter inversion and meeting the needs of ecological environment monitoring in mining areas.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-26
- Publication Date
- 2026-03-27
AI Technical Summary
Existing remote sensing image fusion methods mostly remain at the level of simple data combination, which leads to uncertainty and loss of key details during the fusion process, making it difficult to meet the requirements of high-precision quantitative environmental parameter inversion, especially how to make full use of ground feature details in high-resolution UAV imagery to correct or super-resolution low-resolution satellite imagery.
A stacked learning inversion model is adopted, which uses a two-level structure of base learner layer and meta learner layer to supervise training with high-resolution UAV imagery as ground truth, construct training samples, optimize the model, achieve feature-level and decision-level deep fusion of high-resolution imagery and low-resolution imagery, and establish a robust nonlinear mapping relationship.
It effectively suppresses uncertainties in the fusion process, ensures the quantitative accuracy of the inversion results, and achieves high-precision environmental parameter inversion, meeting the needs of refined environmental monitoring in mining areas and other regions.
Smart Images

Figure CN121746955A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of remote sensing and mining area monitoring, and particularly relates to a remote sensing image fusion inversion method, device, equipment and medium. BACKGROUND
[0002] The mining activities of resources such as coal have the characteristics of large scale and deep disturbance, which has a significant impact on the ecological elements such as the surface structure, vegetation coverage, soil quality and water resources of the mining area. Therefore, dynamic, continuous and accurate monitoring of the ecological environment of the mining area is crucial for guiding reasonable resource development, implementing effective ecological restoration and evaluating the effectiveness of environmental governance.
[0003] Remote sensing technology has become an important tool for ecological environment monitoring due to its efficiency and wide range. However, a single remote sensing data source has inherent limitations. On the one hand, satellite remote sensing images represented by Sentinel-2 can provide free, long-term global coverage data, but their spatial resolution (usually 10 meters or lower) is limited, making it difficult to capture small-scale, fine-scale surface feature changes such as surface subsidence and vegetation degradation in mining areas. On the other hand, unmanned aerial vehicle (UAV) remote sensing technology can obtain centimeter-level high-resolution images, with flexible data collection and good real-time performance, but is limited by endurance and operating costs, making it difficult to monitor a large area continuously and lacking historical image data for tracing.
[0004] In order to combine the advantages of both and make up for the shortcomings of a single data source, the academic and industrial communities have begun to explore multi-source data collaborative monitoring technology. However, existing fusion methods mostly stop at the level of simple “combination” of data, for example, simply georeferencing and displaying two images, or using traditional spatio-temporal data fusion algorithms (such as STARFM) for processing. These methods often introduce new uncertainties in the fusion process, or lose key details due to model simplification, making it difficult to meet the needs of high-precision quantitative environmental parameter inversion (such as vegetation index NDVI, biomass, etc.) after fusion. In particular, how to fully utilize the rich surface details contained in high-resolution images to “correct” or “super-resolution” low-resolution images and establish a robust non-linear mapping relationship between the two is a core problem faced by current technology.
[0005] Therefore, there is an urgent need for a new fusion inversion method that can truly assign the high spatial resolution advantage of unmanned aerial vehicle images to satellite images to generate remote sensing data products with high precision and long-term characteristics, providing reliable data support for fine environmental monitoring in mining areas and other regions. SUMMARY
[0006] Therefore, the application provides a remote sensing image fusion inversion method and device, equipment and medium, and mainly aims to solve the problem that the existing fusion methods are mostly simple "combination" of data, new uncertainty is introduced in the fusion process, or key detail information is lost due to model simplification, so that the precision of the fused image is still difficult to meet the demand of high-precision quantitative environmental parameter inversion, especially, how to fully utilize the rich ground object details in high-resolution images to "correct" or "super-resolution" low-resolution images, and establish a robust nonlinear mapping relationship between the two.
[0007] In a first aspect, the application provides a remote sensing image fusion inversion method, comprising: obtaining a first remote sensing image and a second remote sensing image of a target area, and performing resampling processing on the first remote sensing image and the second remote sensing image to unify the spatial resolution of the first remote sensing image and the second remote sensing image, wherein the first remote sensing image is a high-resolution unmanned aerial vehicle image, and the second remote sensing image is a low-resolution satellite image; constructing a training sample by using the resampled first remote sensing image and second remote sensing image, to perform model training on a stacked learning inversion model by using the training sample, wherein the stacked learning inversion model comprises a base learner layer and a meta learner layer, the base learner layer comprises at least two machine learning regression models, which are respectively used to predict the second remote sensing image and output respective preliminary prediction results, the meta learner layer is used to perform weighted fusion processing on the plurality of preliminary prediction results output by the base learner layer, and the first remote sensing image is used as a true value to perform model optimization, to obtain the trained stacked learning inversion model; inputting a second remote sensing image to be processed into the trained stacked learning inversion model to perform image inversion processing, to obtain a target precision inversion image.
[0008] In a second aspect, the application provides a remote sensing image fusion inversion device, comprising: an acquisition module, configured to obtain a first remote sensing image and a second remote sensing image of a target area, and perform resampling processing on the first remote sensing image and the second remote sensing image to unify the spatial resolution of the first remote sensing image and the second remote sensing image, wherein the first remote sensing image is a high-resolution unmanned aerial vehicle image, and the second remote sensing image is a low-resolution satellite image; The training module is configured to construct training samples by using the resampled first remote sensing image and the resampled second remote sensing image, and to perform model training on a stacked learning inversion model by using the training samples, wherein the stacked learning inversion model comprises a base learner layer and a meta learner layer, the base learner layer comprises at least two machine learning regression models, and each of the machine learning regression models is configured to predict the second remote sensing image and output a respective preliminary prediction result, the meta learner layer is configured to perform weighted fusion processing on the multiple preliminary prediction results output by the base learner layer, and the stacked learning inversion model is optimized by using the first remote sensing image as a true value to obtain a trained stacked learning inversion model. The processing module is configured to input the second remote sensing image to be processed into the trained stacked learning inversion model to perform image inversion processing to obtain a target-precision inversion image.
[0009] In a third aspect, the present application provides an electronic device, which comprises a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, and the processor implements the remote sensing image fusion inversion method of the first aspect when executing the computer program.
[0010] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the remote sensing image fusion inversion method of the first aspect.
[0011] By means of the above technical solution, the remote sensing image fusion inversion method, device, equipment and medium provided by the present application can obtain a first remote sensing image and a second remote sensing image of a target area, and perform resampling processing on the first remote sensing image and the second remote sensing image to unify the spatial resolutions of the first remote sensing image and the second remote sensing image, wherein the first remote sensing image is a high-resolution unmanned aerial vehicle image, and the second remote sensing image is a low-resolution satellite image; training samples are constructed by using the resampled first remote sensing image and the resampled second remote sensing image, and a stacked learning inversion model is trained by using the training samples, wherein the stacked learning inversion model comprises a base learner layer and a meta learner layer, the base learner layer comprises at least two machine learning regression models, each of which is configured to predict the second remote sensing image and output a respective preliminary prediction result, the meta learner layer is configured to perform weighted fusion processing on the multiple preliminary prediction results output by the base learner layer, the stacked learning inversion model is optimized by using the first remote sensing image as a true value to obtain a trained stacked learning inversion model, and the second remote sensing image to be processed is input into the trained stacked learning inversion model to perform image inversion processing to obtain a target-precision inversion image.
[0012] By adopting the technical scheme, the stacking learning of the application is through two-level structures of a base learner layer and a meta learner layer, and instead of directly combining pixel values, the base learners capture different feature rules in data respectively, and then the meta learner performs intelligent fusion, which is a deep fusion of feature level and decision level, excavates potential logic inside data, and avoids staying at surface data "patchwork". Traditional simple fusion often produces spectral distortion. The application uses high-resolution unmanned aerial vehicle images as true values for supervised training, and the target of the model is to constantly approach high-precision true values. In this way of "calibration with true values", uncertainty generated in the fusion process can be effectively inhibited, and the accuracy of the inversion result in the quantitative aspect is ensured.
[0013] The application performs targeted fine adjustment on different ground objects (such as water, vegetation, and buildings) through "quantile correction" and "ecological zoning correction" of the basic processing layer, avoids local detail loss caused by global uniform processing, extracts "multi-dimensional statistical features (skewness, kurtosis, etc.)" and "nonlinear transformation features" through the advanced processing layer, and makes rich detail information in high-resolution images explicit for model learning.
[0014] The base learner layer adopts at least two heterogeneous machine learning regression models such as random forest, XGBoost, and SVR. Different models have different sensitivities to data distribution, nonlinear relationship, and noise. This integration strategy avoids underfitting of a single model caused by hypothesis simplification, can capture ground object details from different angles, and thus meets the needs of high-precision quantitative inversion.
[0015] In the training process, the model directly uses high-resolution unmanned aerial vehicle images (first remote sensing images) as supervised labels (true values). When processing low-resolution satellite images (second remote sensing images), the optimization target of the model is to predict a result infinitely close to the unmanned aerial vehicle true value. This process essentially uses rich ground object details and texture information in high-resolution images to "guide" and "correct" pixel values of low-resolution images, realizes "super-resolution" reconstruction of low-resolution images, and unifies the spatial resolution through resampling, so that high-resolution and low-resolution images are strictly aligned at the pixel level, and each low-resolution pixel can find corresponding high-resolution details for learning, and a stable nonlinear mapping relationship between them is established.
[0016] The above description is only a summary of the technical scheme of the application, in order to more clearly understand the technical means of the application, the specific embodiments of the application can be implemented in accordance with the content of the specification, and in order to make the above and other purposes, features and advantages of the application more obvious and easy to understand, the following specific embodiments of the application are described. BRIEF DESCRIPTION OF DRAWINGS
[0017] The accompanying drawings, which are incorporated herein and constitute part of the specification, illustrate embodiments consistent with the application and, together with the description, function to explain the principles of the application.
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings required by the embodiments or the prior art description will be briefly introduced as follows. Obviously, those skilled in the art can obtain other drawings according to these drawings without any creative effort.
[0019] Figure 1 A flowchart of a remote sensing image fusion inversion method provided by an embodiment of the present application is shown in the figure. Figure 2 A flowchart of another remote sensing image fusion inversion method provided by an embodiment of the present application is shown in the figure. Figure 3 A comparison chart of a raw satellite NDVI image, an inverted NDVI image and a UAV NDVI image provided by an embodiment of the present application is shown in the figure. Figure 4 A comparison chart of inversion results of different models of a base learner layer under different resampling data provided by an embodiment of the present application is shown in the figure. Figure 5 A comparison chart of MAPE values processed by different resampling methods provided by an embodiment of the present application is shown in the figure. Figure 6 A schematic diagram of the effect of improving inversion accuracy of a stacked learning model provided by an embodiment of the present application is shown in the figure. Figure 7 A structural schematic diagram of a remote sensing image fusion inversion device provided by an embodiment of the present application is shown in the figure. DETAILED DESCRIPTION
[0020] The exemplary embodiments of the present application are described below with reference to the accompanying drawings, which include various details of the embodiments of the present application to assist in understanding, and should be considered as merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Also, in order to be clear and concise, the description of well-known functions and structures is omitted in the following description. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.
[0021] The remote sensing image fusion inversion method, device, equipment and medium of the embodiments of the present application are described below with reference to the accompanying drawings.
[0022] The application provides a remote sensing image fusion inversion method, device, equipment and medium, and mainly aims to solve the problem that current existing fusion methods are mostly limited to simple "combination" of data, new uncertainty is introduced in the fusion process, or key detail information is lost due to model simplification, so that the precision of the fused image is still difficult to meet the needs of high-precision quantitative environmental parameter inversion, and in particular, how to fully utilize the rich ground object details contained in high-resolution images to "correct" or "super-resolution" low-resolution images, and establish a robust nonlinear mapping relationship between the two.
[0023] As shown in Figure 1 , the embodiment of the application provides a remote sensing image fusion inversion method, comprising: Step 101, acquiring a first remote sensing image and a second remote sensing image of a target area, and performing resampling processing on the first remote sensing image and the second remote sensing image to unify the spatial resolutions of the first remote sensing image and the second remote sensing image.
[0024] In this embodiment, as shown in Figure 2 , the Sentinel-2 L2A level image (spatial resolution 10 meters) of the study area (such as the Erlin Rabbit mining area in Inner Mongolia) can be acquired as the second remote sensing image, and the image acquired by the DJI M210 RTK unmanned aerial vehicle carrying the X5S multispectral camera (ground sampling interval 1.8 cm) on the same day can be acquired as the first remote sensing image.
[0025] For the embodiment of the present disclosure, the first remote sensing image and the second remote sensing image are resampled to unify the spatial resolutions of the first remote sensing image and the second remote sensing image, which can specifically include: First, the first remote sensing image and the second remote sensing image are respectively subjected to atmospheric correction, radiation calibration and geographic registration processing, the normalized vegetation index value (NDVI) of the first remote sensing image and the normalized vegetation index value of the second remote sensing image are calculated, so as to eliminate the influence of factors such as atmosphere and sensor posture on the data, and ensure that the two images are strictly aligned in geographical space, wherein the atmospheric correction can eliminate the influence of atmosphere on the remote sensing image, the radiation calibration can convert the digital quantization value of the image into the radiation brightness value, and the geographic registration can ensure that the image has accurate geographic coordinate information; The calculation formula of the normalized vegetation index value is as follows:
[0026] In the formula, is the normalized vegetation index value, is the near-infrared band reflectivity, is the red band reflectivity; Then, a resampling algorithm can be used to unify the spatial resolutions of the normalized vegetation index values of the first remote sensing image and the normalized vegetation index values of the second remote sensing image; wherein the resampling algorithm can include at least one of nearest neighbor sampling, bilinear interpolation, or cubic convolution interpolation. In the present embodiment, considering that cubic convolution interpolation performs better in preserving image details and smoothing edges, a cubic convolution resampling method is preferably used; specifically, the NDVI data of the Sentinel-2 image is down-sampled, and the NDVI data of the unmanned aerial vehicle image is up-sampled, so that the spatial resolutions of both are 0.1 meters.
[0027] For the present embodiment, after the first remote sensing image and the second remote sensing image are resampled to unify the spatial resolutions of the first remote sensing image and the second remote sensing image, the method further includes constructing a double-layer preprocessing model to improve data quality and model performance.
[0028] Specifically, the double-layer preprocessing model can include a first layer of basic processing layer and a second layer of advanced processing layer. The basic processing layer can be used to perform statistical analysis and quantile correction processing and ecology-based partition correction processing on the normalized vegetation index values of the resampled first remote sensing image and the normalized vegetation index values of the second remote sensing image; specifically, this step can include: calculating statistical analysis values of the normalized vegetation index values of the first remote sensing image and the normalized vegetation index values of the second remote sensing image, the statistical analysis values can include minimum value, maximum value, mean value, median value, and standard deviation value, for example, by calculating the minimum value of the unmanned aerial vehicle NDVI data is 0.1, the maximum value is 0.8, the mean value is 0.5, etc., and the satellite NDVI data also has corresponding statistical analysis values, to preliminarily understand the distribution characteristics of the data; adjusting the preset quantile values of the normalized vegetation index values of the second remote sensing image to match the quantile values of the normalized vegetation index values of the first remote sensing image; specifically, the 10%, 25%, 50%, 75%, and 90% quantiles of the satellite NDVI data can be accurately adjusted to match the values of the corresponding quantiles of the unmanned aerial vehicle NDVI data; for example, the 50% quantile of the satellite NDVI data is originally 0.3, which is adjusted to match the 50% quantile of the unmanned aerial vehicle NDVI data, 0.4, thereby reducing the difference in distribution between the two types of data; Based on the range of Normalized Difference Vegetation Index (NDVI) values, the first and second remote sensing images are divided into different ecological regions. Each ecological region includes at least water bodies (lower NDVI values, e.g., less than 0.2), built-up areas (NDVI values close to 0), and vegetation areas (higher NDVI values, e.g., greater than 0.3). A preset multiplier is applied to each ecological region for regional adjustment to make targeted adjustments based on the actual characteristics of different land features and improve the physical rationality of the data. For example, for vegetation areas, a multiplier of 1.1 can be applied for adjustment to more accurately reflect the true vegetation situation. Finally, advanced processing layers can be used to extract multidimensional features from the normalized vegetation index (NDI) values of the first remote sensing image and the second remote sensing image after processing by the basic processing layer.
[0029] Multidimensional features may include: advanced statistical features and nonlinear transformation features; Advanced statistical features can be used to describe the shape and spatial correlation of data distribution, including skewness, kurtosis and autocorrelation coefficient. Skewness can be used to describe the asymmetry of data distribution, kurtosis can be used to describe the sharpness of data distribution, and autocorrelation coefficient can reflect the correlation of data at different locations. Nonlinear transformation features can be used to enhance the model's ability to handle complex nonlinear relationships, enabling the model to capture patterns that are difficult to discover using only the original pixels. These features can include those obtained by applying logarithmic transformation, square root transformation, and reciprocal transformation to the data. For example, after performing a logarithmic transformation on NDVI data, the originally nonlinear data may exhibit more obvious linear characteristics, which is easier for subsequent model processing.
[0030] Training samples are constructed using the multidimensional features obtained through the above processing. These multidimensional features can more comprehensively reflect the characteristics of the data and provide richer information for model training.
[0031] Step 102: Construct training samples using the resampled first and second remote sensing images, and use the training samples to train the stacked learning inversion model.
[0032] The stacked learning inversion model can include a base learner layer and a meta learner layer. The base learner layer can contain at least two machine learning regression models, which can be used to predict the second remote sensing image (i.e., low-resolution image) and output their respective preliminary prediction results. The meta learner layer can be used to perform weighted fusion processing on the multiple preliminary prediction results output by the base learner layer, and use the first remote sensing image (i.e., high-resolution image) as the ground value to optimize the model and obtain the trained stacked learning inversion model.
[0033] Specifically, the base learner layers of the stacked learning inversion model can include at least two machine learning regression models in random forest (RF) model, XGBoost (XGB) model, support vector regression (SVR) model, gradient boosting (GB) model, LightGBM model, or CatBoost model.
[0034] In this embodiment, the base learner layers can process data in parallel using the above six machine learning models. Each machine learning regression model is determined by optimal parameter search to determine the best configuration to ensure that each model achieves the best performance. For example, for the random forest model, by adjusting the number of trees, maximum depth, and other parameters, the optimal combination of parameters is found on the validation set. These different types and configurations of machine learning regression models can capture complex relationships in data from different angles and fully leverage their respective advantages.
[0035] The meta learner layer of the stacked learning inversion model adopts a regularized linear regression model, which can be at least one of Lasso regression, Ridge regression, or ElasticNet regression. In this embodiment, the above three regularized linear regression models are used to weight and fuse the six prediction results output by the base learner layer. The meta learner layer aims to effectively integrate the prediction results of the base layer model and prevent overfitting to improve the generalization ability of the model. In this embodiment, it has been verified through experiments that ElasticNet regression performs best.
[0036] Step 103, input the second remote sensing image to be processed into the trained stacked learning inversion model for image inversion processing to obtain a target precision inversion image.
[0037] For the embodiments of the present disclosure, the Sentinel-2 satellite image (second remote sensing image) to be processed is input into the trained stacked learning inversion model. The stacked learning inversion model processes the input satellite image, makes predictions through multiple machine learning regression models of the base learner layer, and outputs respective preliminary prediction results. Then, the meta learner layer performs weighted fusion processing on these preliminary prediction results to generate a final target precision inversion image. This inversion image has high precision and can more accurately reflect environmental information such as the vegetation condition of the mining area, as shown in Figure 3 .
[0038] To verify the effectiveness of the method, the present embodiment performs a result verification experiment. The data set is divided into a training set and a test set according to 80% / 20%. The unmanned aerial vehicle NDVI is taken as the true value, and the inversion results of the original Sentinel-2 NDVI, the Sentinel-2 NDVI only after resampling processing, and the Sentinel-2 NDVI after the complete method of the present application are compared, as shown in Figure 4As shown.
[0039] The results show that after using triple convolutional resampling and ElasticNet stacked learning model, as Figure 5 As shown, the mean absolute percentage error (MAPE) between the NDVI inverted values of Sentinel-2 imagery and the true NDVI values from the UAV decreased from an initial 54.31% to 10.01%, a reduction of 44.30%, resulting in a significant improvement in accuracy. Figure 6 As shown, this demonstrates that the method proposed in this application can effectively integrate two data sources to achieve high-precision environmental monitoring, providing a reliable data foundation for the long-term evolution analysis of the mining area's ecological environment.
[0040] In summary, according to the remote sensing image fusion and inversion method provided in this application, compared with the existing technology, this application can acquire a first remote sensing image and a second remote sensing image of the target area, and resample the first and second remote sensing images to unify their spatial resolution. The first remote sensing image is a high-resolution UAV image, and the second remote sensing image is a low-resolution satellite image. Training samples are constructed using the resampled first and second remote sensing images to train a stacked learning inversion model. The stacked learning inversion model includes a base learner layer and a meta-learner layer. The base learner layer contains at least two machine learning regression models, which are used to predict the second remote sensing image and output their respective preliminary prediction results. The meta-learner layer performs weighted fusion processing on the multiple preliminary prediction results output by the base learner layer, using the first remote sensing image as the ground truth for model optimization, resulting in a trained stacked learning inversion model. The second remote sensing image to be processed is then input into the trained stacked learning inversion model for image inversion processing to obtain the target-accuracy inversion image.
[0041] Using the above technical solution, the stacked learning approach of this application employs a two-level structure of "base learner layer" and "meta learner layer." Instead of directly combining pixel values, it allows the base learners to capture different feature patterns in the data, which are then intelligently fused by the meta learners. This represents a deep fusion at the feature and decision levels, uncovering the underlying logic within the data and avoiding superficial data "patchwork." Traditional simple fusion often results in spectral distortion. This application utilizes high-resolution UAV imagery as ground truth for supervised training, with the model aiming to continuously approximate the high-precision ground truth. This "ground truth calibration" method effectively suppresses uncertainties generated during the fusion process, ensuring the quantitative accuracy of the inversion results.
[0042] This application uses "quantile correction" and "ecological zoning correction" in the basic processing layer to make targeted and refined adjustments to different land features (such as water bodies, vegetation, and buildings), avoiding the loss of local details caused by global uniform processing; and extracts "multidimensional statistical features (skewness, kurtosis, etc.)" and "nonlinear transformation features" in the advanced processing layer to make the rich details hidden in the high-resolution image explicit for the model to learn.
[0043] The base learner layer employs at least two heterogeneous machine learning regression models, such as Random Forest, XGBoost, and SVR. Different models exhibit varying sensitivities to data distribution, nonlinear relationships, and noise. This ensemble strategy avoids underfitting caused by simplistic assumptions in a single model, and captures ground feature details from different perspectives, thus meeting the requirements for high-precision quantitative inversion.
[0044] During training, the model directly uses high-resolution UAV imagery (first remote sensing imagery) as the supervision label (ground truth). When processing low-resolution satellite imagery (second remote sensing imagery), the optimization objective is to predict results that are infinitely close to the UAV ground truth. Essentially, this process utilizes the rich ground feature details and texture information in the high-resolution imagery to "guide" and "correct" the pixel values of the low-resolution imagery, achieving "super-resolution" reconstruction of the low-resolution imagery. By resampling to unify the spatial resolution, the high-resolution and low-resolution images are strictly aligned at the pixel level, ensuring that each low-resolution pixel can find corresponding high-resolution details for learning and establishing a robust non-linear mapping relationship between the two.
[0045] Based on the above Figure 1 The specific implementation of the method shown in this embodiment provides a remote sensing image fusion and inversion device, such as... Figure 7 As shown, the device includes: an acquisition module 31, a training module 32, and a processing module 33; The acquisition module 31 is used to acquire a first remote sensing image and a second remote sensing image of the target area, and to perform resampling processing on the first remote sensing image and the second remote sensing image to unify the spatial resolution of the first remote sensing image and the second remote sensing image, wherein the first remote sensing image is a high-resolution UAV image and the second remote sensing image is a low-resolution satellite image. Training module 32 is used to construct training samples using the resampled first remote sensing image and second remote sensing image, and to train the stacked learning inversion model using the training samples. The stacked learning inversion model includes a base learner layer and a meta learner layer. The base learner layer contains at least two machine learning regression models, which are used to predict the second remote sensing image and output their respective preliminary prediction results. The meta learner layer is used to perform weighted fusion processing on the multiple preliminary prediction results output by the base learner layer, and to optimize the model using the first remote sensing image as the ground truth, so as to obtain the trained stacked learning inversion model. The processing module 33 is used to input the second remote sensing image to be processed into the trained stacked learning inversion model for image inversion processing to obtain the target accuracy inversion image.
[0046] In specific application scenarios, the acquisition module 31 can be used to perform atmospheric correction, radiometric calibration, and georegistration on the first remote sensing image and the second remote sensing image respectively, and calculate the normalized vegetation index (NDI) values of the first remote sensing image and the second remote sensing image; and use a resampling algorithm to unify the spatial resolution of the NDI values of the first remote sensing image and the second remote sensing image, wherein the resampling algorithm includes at least one of nearest neighbor sampling, bilinear interpolation, or cubic convolution interpolation.
[0047] In specific application scenarios, such as Figure 7 As shown, the device also includes: a preprocessing module 34; The preprocessing module 34 is used to construct a two-layer preprocessing model, which includes a first-layer basic processing layer and a second-layer advanced processing layer. The basic processing layer performs statistical analysis, quantile correction, and ecological zoning correction on the normalized vegetation index (NVI) values of the resampled first and second remote sensing images. The advanced processing layer extracts multidimensional features from the NVI values of the first and second remote sensing images after processing by the basic processing layer. These multidimensional features include advanced statistical features and nonlinear transformation features. The advanced statistical features include skewness, kurtosis, and autocorrelation coefficient, while the nonlinear transformation features include features obtained by applying logarithmic, square root, and reciprocal transformations to the data.
[0048] In specific application scenarios, the training module 32 can be used to construct the training samples using the multidimensional features.
[0049] In a specific application scenario, the preprocessing module 34 can be configured to calculate statistical analysis values of the normalized vegetation index values of the first remote sensing image and the normalized vegetation index values of the second remote sensing image, the statistical analysis values including minimum values, maximum values, mean values, median values, and standard deviation values of the normalized vegetation index values of the first remote sensing image and the normalized vegetation index values of the second remote sensing image; adjust a preset quantile value of the normalized vegetation index values of the second remote sensing image to match a quantile value corresponding to the normalized vegetation index values of the first remote sensing image; divide the first remote sensing image and the second remote sensing image into different ecological regions according to a normalized vegetation index value range, and apply a preset multiplier to each of the ecological regions for regional adjustment, where the ecological regions at least include a water body region, a building region, and a vegetation region.
[0050] In a specific application scenario, the training module 32 can be configured to include at least two machine learning regression models in the base learner layer of the stacked learning inversion model, such as a random forest model, an XGBoost model, a support vector regression model, a gradient boosting model, a LightGBM model, or a CatBoost model, each of which is determined by optimal parameter searching to have a best configuration.
[0051] In a specific application scenario, the training module 32 can be configured to use a regularized linear regression model in the meta learner layer of the stacked learning inversion model, the regularized linear regression model being at least one of a Lasso regression, a Ridge regression, or an ElasticNet regression.
[0052] It should be noted that other corresponding descriptions of the various functional units involved in the remote sensing image fusion inversion device provided in this embodiment can be referred to the corresponding descriptions in the Figure 1 , which will not be described here again.
[0053] Based on the method as shown in Figure 1 , correspondingly, the present embodiment also provides a computer readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the method as shown in Figure 1 .
[0054] Based on such an understanding, the technical solution of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a U disk, a mobile hard disk, etc.), and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of various implementation scenarios of the present application.
[0055] Based on the method as shown in Figure 1 , and Figure 7In order to achieve the above-mentioned purposes, the virtual device embodiment shown also provides an electronic device, which comprises a storage medium and a processor; the storage medium is used for storing a computer program; the processor is used for executing the computer program to realize the above-mentioned method. Figure 1 The method shown.
[0056] Optionally, the above-mentioned entity device can also comprise a user interface, a network interface, a camera, a radio frequency (RF) circuit, a sensor, an audio circuit, a WI-FI module, etc. The user interface can comprise a display screen (Display), an input unit such as a keyboard (Keyboard), etc. The optional user interface can also comprise a USB interface, a card reader interface, etc. The network interface can optionally comprise a standard wired interface, a wireless interface (such as a WI-FI interface), etc.
[0057] Those skilled in the art can understand that the above-mentioned entity device structure provided by the embodiment does not constitute a limitation on the entity device, and can comprise more or fewer components, or combine certain components, or different component arrangements.
[0058] The storage medium can also comprise an operating system and a network communication module. The operating system is a program for managing the hardware and software resources of the above-mentioned entity device, and supports the running of the remote sensing image fusion inversion program and other software and / or programs. The network communication module is used for realizing the communication between the components in the storage medium, and the communication between the storage medium and other hardware and software in the remote sensing image fusion inversion entity device.
[0059] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be realized by means of software and necessary general hardware platforms, or by hardware. By applying the scheme of the present embodiment, compared with the prior art, the present application can obtain a first remote sensing image and a second remote sensing image of a target area, and perform resampling processing on the first remote sensing image and the second remote sensing image to unify the spatial resolution of the first remote sensing image and the second remote sensing image, wherein the first remote sensing image is a high-resolution unmanned aerial vehicle image, and the second remote sensing image is a low-resolution satellite image; a training sample is constructed by using the resampled first remote sensing image and second remote sensing image, so as to perform model training on a stacked learning inversion model by using the training sample, wherein the stacked learning inversion model includes a base learner layer and a meta learner layer, the base learner layer includes at least two machine learning regression models, which are respectively used to predict the second remote sensing image and output respective preliminary prediction results, the meta learner layer is used to perform weighted fusion processing on the multiple preliminary prediction results output by the base learner layer, the model is optimized by taking the first remote sensing image as a true value, and a trained stacked learning inversion model is obtained; the second remote sensing image to be processed is input into the trained stacked learning inversion model for image inversion processing, and a target precision inversion image is obtained.
[0060] By adopting the above technical scheme, the stacked learning of the present application is through a two-level structure of a base learner layer and a meta learner layer, and is not directly combined with pixel values, but makes the base learner capture different feature laws in data, and then makes the meta learner perform intelligent fusion, which is a deep fusion at a feature level and a decision level, excavates the potential logic inside the data, and avoids staying at the surface of data "patching". The traditional simple fusion often produces spectral distortion. The present application uses high-resolution unmanned aerial vehicle images as true values for supervised training, and the target of the model is to constantly approach high-precision true values. This "calibration with true values" can effectively suppress the uncertainty generated in the fusion process and ensure the accuracy of the inversion result in quantity.
[0061] Through the "quantile correction" and "ecological zoning correction" of the basic processing layer, the present application performs targeted fine adjustment on different ground objects (such as water, vegetation, and buildings), avoiding the loss of local details caused by global uniform processing; through the extraction of "multidimensional statistical features (skewness, kurtosis, etc.)" and "nonlinear transformation features" by the advanced processing layer, the rich detailed information hidden in the high-resolution image is made explicit for model learning.
[0062] The base learner layer adopts at least two heterogeneous machine learning regression models such as random forest, XGBoost, SVR, etc. Different models have different sensitivities to the distribution of data, nonlinear relationships, and noise. This ensemble strategy avoids underfitting caused by the simplification of a single model and can capture feature details from different angles to meet the needs of high-precision quantitative inversion.
[0063] During the training process, the model directly takes the high-resolution unmanned aerial vehicle image (first remote sensing image) as the supervised label (true value). When processing the low-resolution satellite image (second remote sensing image), the optimization goal of the model is to predict a result that is infinitely close to the unmanned aerial vehicle true value. This process essentially uses the rich feature details and texture information in the high-resolution image to "guide" and "correct" the pixel values of the low-resolution image, achieving "super-resolution" reconstruction of the low-resolution image. By resampling to unify the spatial resolution, the high-resolution and low-resolution images are strictly aligned at the pixel level, ensuring that each low-resolution pixel can find corresponding high-resolution details for learning and establishing a robust nonlinear mapping relationship between them.
[0064] It should be noted that, in this document, relational terms such as "first" and "second", and the like can be used solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any actual such relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", "includes", "including", or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a", "comprising... a", "includes... a", "including... a", or the like with no mutually exclusive relationship is understood to imply that anything other than the specifically recited elements are not included.
[0065] The above merely provides specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A remote sensing image fusion inversion method, characterized in that, The method comprises the following steps: acquiring a first remote sensing image and a second remote sensing image of a target area, and performing resampling processing on the first remote sensing image and the second remote sensing image to unify the spatial resolutions of the first remote sensing image and the second remote sensing image, wherein the first remote sensing image is a high-resolution unmanned aerial vehicle image, and the second remote sensing image is a low-resolution satellite image; constructing a training sample by using the resampled first remote sensing image and second remote sensing image, and performing model training on a stacked learning inversion model by using the training sample, wherein the stacked learning inversion model comprises a base learner layer and a meta learner layer, the base learner layer comprises at least two machine learning regression models, and each of the machine learning regression models is used for predicting the second remote sensing image and outputting a respective preliminary prediction result, the meta learner layer is used for performing weighted fusion processing on the plurality of preliminary prediction results output by the base learner layer, and the first remote sensing image is used as a true value to perform model optimization, so as to obtain a trained stacked learning inversion model; inputting a second remote sensing image to be processed into the trained stacked learning inversion model to perform image inversion processing, and obtaining a target accuracy inversion image.
2. The method of claim 1, wherein, The resampling processing on the first remote sensing image and the second remote sensing image to unify the spatial resolutions of the first remote sensing image and the second remote sensing image comprises the following steps: performing atmospheric correction, radiation calibration and geographical registration processing on the first remote sensing image and the second remote sensing image respectively, and calculating a normalized vegetation index value of the first remote sensing image and a normalized vegetation index value of the second remote sensing image; unifying the spatial resolutions of the normalized vegetation index value of the first remote sensing image and the normalized vegetation index value of the second remote sensing image by using a resampling algorithm, wherein the resampling algorithm comprises at least one of nearest neighbor sampling, bilinear interpolation or cubic convolution interpolation.
3. The method of claim 2, wherein, After the resampling processing on the first remote sensing image and the second remote sensing image to unify the spatial resolutions of the first remote sensing image and the second remote sensing image, the method further comprises the following steps: constructing a double-layer preprocessing model, wherein the double-layer preprocessing model comprises a first layer of a basic processing layer and a second layer of an advanced processing layer; performing statistical analysis and quantile correction processing and ecology-based partition correction processing on the normalized vegetation index value of the resampled first remote sensing image and the normalized vegetation index value of the second remote sensing image by using the basic processing layer; extracting multi-dimensional features from the normalized vegetation index value of the first remote sensing image and the normalized vegetation index value of the second remote sensing image processed by the basic processing layer by using the advanced processing layer; The construction of the training sample by using the resampled first remote sensing image and second remote sensing image comprises the following steps: constructing the training sample by using the multi-dimensional features.
4. The method of claim 3, wherein, The statistical analysis and quantile correction processing and ecology-based partition correction processing on the normalized vegetation index value of the resampled first remote sensing image and the normalized vegetation index value of the second remote sensing image by using the basic processing layer comprise the following steps: calculating statistical analysis values of the normalized vegetation index values of the first remote sensing image and the second remote sensing image, the statistical analysis values including minimum values, maximum values, mean values, median values and standard deviation values of the normalized vegetation index values of the first remote sensing image and the second remote sensing image; adjusting preset quantile values of the normalized vegetation index values of the second remote sensing image to match quantile values corresponding to the normalized vegetation index values of the first remote sensing image; dividing the first remote sensing image and the second remote sensing image into different ecological regions according to normalized vegetation index value ranges, and applying a preset multiplier to each ecological region for regional adjustment, wherein the ecological regions at least include water bodies, building areas and vegetation areas.
5. The method of claim 3, wherein, The multi-dimensional features include advanced statistical features and nonlinear transformation features, wherein the advanced statistical features include skewness, kurtosis and autocorrelation coefficients, and the nonlinear transformation features include features obtained by applying logarithmic transformation, square root transformation and reciprocal transformation to data.
6. The method of claim 1, wherein, The base learner layers of the stacked learning inversion model include at least two machine learning regression models among random forest models, XGBoost models, support vector regression models, gradient boosting models, LightGBM models and CatBoost models, and each machine learning regression model is determined by optimal parameter search.
7. The method of claim 1, wherein, The meta learner layer of the stacked learning inversion model adopts a regularized linear regression model, which is at least one of Lasso regression, Ridge regression or ElasticNet regression.
8. A remote sensing image fusion inversion device, characterized in that, The method comprises: an acquisition module, configured to acquire a first remote sensing image and a second remote sensing image of a target region, and perform resampling processing on the first remote sensing image and the second remote sensing image to unify spatial resolutions of the first remote sensing image and the second remote sensing image, wherein the first remote sensing image is a high-resolution unmanned aerial vehicle image, and the second remote sensing image is a low-resolution satellite image; a training module, configured to construct training samples by using the resampled first remote sensing image and second remote sensing image, and perform model training on a stacked learning inversion model by using the training samples, wherein the stacked learning inversion model includes a base learner layer and a meta learner layer, the base learner layer includes at least two machine learning regression models, which are respectively used to predict the second remote sensing image and output respective preliminary prediction results, and the meta learner layer is used to perform weighted fusion processing on the multiple preliminary prediction results output by the base learner layer, and the first remote sensing image is used as true values to perform model optimization, so as to obtain a trained stacked learning inversion model; a processing module, configured to input a second remote sensing image to be processed into the trained stacked learning inversion model to perform image inversion processing, so as to obtain a target-precision inversion image.
9. An electronic device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, The processor implements the method in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program, which is executed by a processor, implements the method of any one of claims 1 to 7.
Citation Information
Patent Citations
Pathological image super-resolution modeling method based on deep learning
CN112785498A
Remote sensing image space-time fusion method and system based on wavelet domain cross pairing
CN117036987A
Reservoir-oriented remote sensing image water quality parameter inversion method and device
CN118485925A
Water quality monitoring method based on multiple spectrums of unmanned aerial vehicle
CN119418228A
Multi-resolution geological data conversion method and system based on ensemble learning
CN121562863A