Registration and fusion method of electronic sand table in spatial geographic image data

Through 3D convolutional neural network and quality enhancement network, the multi-source images are preprocessed, registered and fused, solving the problem of inconsistent quality of multi-source image data under high compression ratio, and achieving high-precision and high-quality image fusion effect.

CN120259095AInactive Publication Date: 2025-07-04XINJIANG ZHONGKE WEICHUANG DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510259928.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-07-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When existing deep learning methods deal with multi-source spatial geographic image data, there are problems of inconsistent quality and inaccurate image registration, especially in the case of high compression ratios, image details and clarity are easily lost.

Method used

A 3D convolutional neural network (3D-CNN) is used for feature point extraction and matching, and a quality enhancement network (QE-Net) is used to preprocess, register and fusion images of different data sources. The PROJ.4 algorithm is used to unify the coordinate system, and image quality is optimized through weighted average and least squares method.

Benefits of technology

The precise registration and high-quality fusion of images from different data sources are achieved, especially at low compression ratio, which significantly improves the PSNR and SSIM indicators of the images, and improves the image details and clarity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259095A_ABST
    Figure CN120259095A_ABST
Patent Text Reader

Abstract

The invention discloses an electronic sand table spatial geographic image data registration fusion method, which performs feature extraction and matching on image data by adopting a 3D convolutional neural network (3D-CNN), can realize accurate registration of images from different data sources, enables the spatial positions and scales of the images to be consistent, and improves the registration accuracy. Therefore, the image fusion precision and effect are improved. According to the method, the fused image is further optimized by introducing the quality enhancement network (QE-Net), and the details and definition of the image can be effectively improved especially under the conditions of video compression and low-quality data sources. Compared with a traditional method, the scheme shows more excellent performance in the aspect of video quality enhancement, and quality indexes such as the PSNR and the SSIM are remarkably improved especially under the low compression ratio (small QP value).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of geospatial image data registration, and specifically relates to a method for registering and fusing spatial geospatial image data with an electronic sand table. Background Art

[0002] Currently, with the rapid development of geographic information systems (GIS) and remote sensing technology, the processing and analysis of spatial geospatial image data have become an important technical field. Especially in various complex environments, such as disaster rescue, field surveys, and urban management scenarios, high-quality geospatial image data is required. Traditional video compression and image processing technologies, such as H.264 / AVC, H.265 / HEVC, etc., compress image data through quantization, encoding, and decoding, thereby reducing the pressure of storage and transmission. However, these technologies often lead to a reduction in the quality of videos and images. Especially in the case of high compression ratios, details and clarity in the images are easily lost.

[0003] To improve the quality of compressed images, many deep learning-based image processing methods have emerged. These methods train neural networks to learn the features of compressed videos, thereby enhancing their visual quality. However, most of the existing deep learning methods focus on processing images from a single data source and often ignore the collaborative processing of multi-source image data, resulting in problems such as inconsistent quality and inaccurate image registration when fusing across data sources.

[0004] In view of this, the present invention is specifically proposed. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to overcome the deficiencies of the prior art and provide a method for registering and fusing spatial geospatial image data with an electronic sand table, solving the problems raised in the above background art.

[0006] To solve the above technical problem, the basic concept of the technical solution adopted by the present invention is as follows:

[0007] A method for registering and fusing spatial geospatial image data with an electronic sand table includes the following steps:

[0008] S1: Preprocess the spatial geospatial image data from different data sources, including format conversion, denoising processing, and radiometric correction, to ensure data quality and consistency. Then, use a coordinate transformation algorithm to uniformly transform the image data from different data sources into the same geographic coordinate system or projection coordinate system;

[0009] S2: Register the images from different data sources through a feature point extraction and matching algorithm based on a 3D convolutional neural network to make the spatial positions and scales of the image data consistent;

[0010] S3: Pixel - level fusion is performed on the registered image data through weighted averaging to enhance the quality of the fused image, retain more detailed information, and further optimize the image quality in combination with a quality enhancement network;

[0011] S4: Use a loss function to evaluate the quality of the fused image, calculate the difference between the fused result and the original image, and correct the error through the least - squares method.

[0012] Optionally, the steps for pre - processing spatial geographic image data from different data sources, including format conversion, denoising, and radiometric correction, to ensure data quality and consistency are as follows:

[0013] Unify the image data from different data sources into a standard format, GeoTIFF, to ensure data consistency.

[0014] Use the median filtering algorithm to denoise the image data. The formula is as follows: I′(x, y) = median{I(x + i, y + j)| - k ≤ i, j ≤ k}, where I(x, y) is the original image data, I′(x, y) is the denoised image data, and k is the radius of the filtering window.

[0015] Perform sensor correction on the image to eliminate the radiometric bias caused by the sensor and convert it to a reflectance value. The formula is: where I is the digital value of the original image, b is the offset, g is the gain, and R is the corrected reflectance value.

[0016] Optionally, the steps for using the PROJ.4 algorithm to unify the image data from different data sources into the same geographic coordinate system or projection coordinate system are as follows:

[0017] Determine the source coordinate system and the target coordinate system. The source coordinate system is WGS84, and the target coordinate system is the UTM projection coordinate system. Then, load the PROJ.4 library and use its coordinate system definition string. The source coordinate system is defined as + proj = longlat + datum = WGS84 + nodefs, and the target coordinate system is defined as + proj = utm + zone = 33 + datum = WGS84;

[0018] Perform coordinate transformation through the PROJ.4 algorithm. Use the transformation formula to convert the image data points (X source , Y source ) in the source coordinate system to the coordinate points (X target , Y target ) in the target coordinate system. The formula is: where T is the transformation matrix, and [X, Y, Z] and [X′, Y′, Z′] are the coordinates in the original coordinate system and the target coordinate system respectively.

[0019] Optionally, the steps of registering images from different data sources by a feature point extraction and matching algorithm based on a 3D convolutional neural network (3D-CNN) to make the spatial positions and scales of the image data consistent are as follows:

[0020] Process the image data using a 3D convolutional neural network to automatically extract key feature points in the image. The formula is: f(x, y, z) = CNN(I(x, y, z)), where f(x, y, z) is the feature point extracted from the image data I(x, y, z), and CNN represents the 3D convolutional neural network;

[0021] Match the feature points extracted by the convolutional neural network in images from different data sources, and calculate the similarity of the feature points. The formula is: where, f1(x i , y i , z i ) and f2(x i , y i , z i ) are the feature points extracted at the position (x i , y i , z i ) in two images respectively, N is the total number of matched feature points, and L match is the matching loss function;

[0022] According to the feature point matching result, register the image through geometric transformation to make its spatial position and scale consistent. The formula is: I target = T match (I source ), where, I target is the target image, I source is the source image, and T match is the registration transformation matrix obtained from the feature point matching.

[0023] Optionally, the steps of pixel-level fusion of the registered image data by weighted average to enhance the quality of the fused image, retain more detailed information, and further optimize the image quality in combination with a quality enhancement network are as follows:

[0024] Perform pixel-level fusion on the registered image. The formula is: I fused (x, y) = αI1(x, y) + (1 - α)I2(x, y), where, I fused (x, y) is the pixel value of the fused image, I1(x, y) and I2(x, y) are the pixel values of the two registered images, and α is the weighting coefficient;

[0025] Use the quality enhancement network to further optimize the fused image, ensure the retention of detailed information and improve the image quality. The formula is: Ienhanced = QE Net (I fused ), where I enhanced is the image after quality enhancement, and QE Net represents the quality enhancement network.

[0026] Optionally, the steps of using a loss function to evaluate the quality of the fused image and calculating the difference between the fused result and the original image are as follows:

[0027] Use the mean squared error loss function to calculate the difference between the fused image and the original image. The formula is: where I fused (x i , y i ) is the pixel value of the fused image, I true (x i , y i ) is the pixel value of the real image, N is the total number of pixels, and L MSE is the mean squared error loss function, which is used to measure the difference between the fused image and the original image.

[0028] Optionally, the steps of correcting the error by the least squares method to ensure the accuracy of the final output image are as follows:

[0029] According to the least squares loss function, optimize the fused result to minimize the error between the images, thereby improving the image quality and ensuring the accuracy of the final output image. The formula is: where is the optimized image, and the objective function is minimized to correct the error between the fused image and the real image to ensure the image accuracy.

[0030] Optionally, the spatial data from different sources includes but is not limited to satellite images, UAV aerial images, and topographic maps.

[0031] After adopting the above technical solutions, the present invention has the following beneficial effects compared with the prior art. Of course, any product implementing the present invention does not necessarily need to achieve all the advantages described below:

[0032] 1. By using a 3D convolutional neural network (3D-CNN) to extract and match the features of the image data, accurate registration of images from different data sources can be achieved, enabling their spatial positions and scales to be consistent, thereby improving the accuracy and effect of image fusion.

[0033] 2. The present invention further optimizes the fused image by introducing a Quality Enhancement Network (QE-Net). Especially in the case of video compression and low-quality data sources, it can effectively enhance the image details and clarity. Compared with traditional methods, this solution shows better performance in video quality enhancement. Especially at low compression ratios (small QP values), it significantly improves quality metrics such as PSNR and SSIM.

[0034] 3. The present invention uses the PROJ.4 algorithm to uniformly transform the image data from different data sources into the same geographic coordinate system or projection coordinate system, ensuring the consistency of multi-source image data and providing a stable data basis for subsequent registration and fusion.

[0035] The following further describes in detail the specific embodiments of the present invention with reference to the accompanying drawings. Description of the Drawings

[0036] The following drawings in the description are only some embodiments. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the attached

[0037] In the figures:

[0038] Figure 1 It is a flow chart of the registration and fusion method for spatial geographic image data.

[0039] It should be noted that these drawings and textual descriptions are not intended to limit the scope of the concept of the present invention in any way, but to illustrate the concept of the present invention to those skilled in the art by referring to specific embodiments. Specific Embodiments

[0040] Now, the present invention will be further described in detail with reference to the accompanying drawings.

[0041] Please refer to Figure 1 As shown, in this embodiment, a method for registering and fusing spatial geographic image data in an electronic sand table is provided, including the following steps:

[0042] S1: Preprocess the spatial geographic image data from different data sources, including format conversion, denoising, and radiometric correction, to ensure data quality and consistency. Then, use a coordinate transformation algorithm to uniformly transform the image data from different data sources into the same geographic coordinate system or projection coordinate system;

[0043] S2: Register the images from different data sources through a feature point extraction and matching algorithm based on a 3D Convolutional Neural Network (3D-CNN) to make the spatial positions and scales of the image data consistent;

[0044] S3: Pixel - level fusion is performed on the registered image data through weighted averaging to enhance the quality of the fused image, retain more detailed information, and further optimize the image quality in combination with the Quality Enhancement Network (QE - Net);

[0045] S4: The loss function (Mean Squared Error MSE) is used to evaluate the quality of the fused image, calculate the difference between the fusion result and the original image, and correct the error through the least - squares method.

[0046] By using a 3D Convolutional Neural Network (3D - CNN) to extract features and match the image data, accurate registration of images from different data sources can be achieved, making their spatial positions and scales consistent, thereby improving the accuracy and effect of image fusion.

[0047] The present invention further optimizes the fused image by introducing the Quality Enhancement Network (QE - Net). Especially in the case of video compression and low - quality data sources, it can effectively improve image details and clarity. Compared with traditional methods, this solution shows better performance in video quality enhancement. Especially at low compression ratios (small QP values), it significantly improves quality metrics such as PSNR and SSIM.

[0048] In this embodiment, the steps for pre - processing spatial geographic image data from different data sources, including format conversion, denoising processing, and radiometric correction, to ensure data quality and consistency are as follows:

[0049] The image data from different data sources are uniformly converted into a standard format, GeoTIFF, to ensure data consistency. Spatial geographic images from different data sources may use different file formats (such as shp, tiff, jpeg, png, kml, mbtiles, etc.);

[0050] The median filtering algorithm is used to denoise the image data. The formula is as follows: I′(x, y) = median{I(x + i, y + j)| - k ≤ i, j ≤ k}, where I(x, y) is the original image data, I′(x, y) is the denoised image data, and k is the radius of the filtering window;

[0051] Sensor correction is performed on the image to eliminate the radiation deviation caused by the sensor and convert it into a reflectance value. The formula is: where I is the digital value of the original image, b is the offset, g is the gain, and R is the corrected reflectance value.

[0052] In this embodiment, the steps for using the PROJ.4 algorithm to uniformly convert the image data from different data sources into the same geographic coordinate system or projection coordinate system are as follows:

[0053] Determine the source coordinate system and the target coordinate system. The source coordinate system is WGS84, and the target coordinate system is the UTM projection coordinate system. Then, load the PROJ.4 library and use its coordinate system definition string. The source coordinate system is defined as +proj=longlat +datum=WGS84 +no d efs, and the target coordinate system is defined as +proj=utm +zone=33 +datum=WGS84;

[0054] Perform coordinate transformation through the PROJ.4 algorithm. Use the transformation formula to convert the image data points (X source , Y source ) in the source coordinate system to the coordinate points (X target , Y target ) in the target coordinate system. The formula is: where T is the transformation matrix, and [X, Y, Z] and [X′, Y′, Z′] are the coordinates in the original coordinate system and the target coordinate system respectively.

[0055] In this embodiment, the steps to register images from different data sources and make the spatial positions and scales of the image data consistent through the feature point extraction and matching algorithm based on the 3D convolutional neural network (3D-CNN) are as follows:

[0056] Use the 3D convolutional neural network to process the image data and automatically extract the key feature points in the image. The formula is: f(x, y, z) = CNN(I(x, y, z)). Where f(x, y, z) is the feature point extracted from the image data I(x, y, z), and CNN represents the 3D convolutional neural network;

[0057] Match the feature points extracted by the convolutional neural network in the images from different data sources, and calculate the similarity of the feature points. The formula is: where f1(x i , y i , z i ) and f2(x i , y i , z i ) are the feature points extracted at the position (x i , y i , z i ) in two images respectively, N is the total number of matched feature points, and L match is the matching loss function;

[0058] According to the feature point matching result, register the image through geometric transformation (such as affine transformation, perspective transformation, etc.) to make its spatial position and scale consistent. The formula is: I target = T match (I source ), where Itarget is the target image, I source is the source image, T match is the registration transformation matrix obtained from feature point matching.

[0059] In this embodiment, the steps of performing pixel-level fusion on the registered image data through weighted averaging to enhance the quality of the fused image, retain more detailed information, and further optimize the image quality in combination with the quality enhancement network are as follows:

[0060] Perform pixel-level fusion on the registered image. The formula is: I fused (x, y) = αI1(x, y) + (1 - α)I2(x, y), where I fused (x, y) is the pixel value of the fused image, I1(x, y) and I2(x, y) are the pixel values of the two registered images, and α is the weighting coefficient;

[0061] Use the quality enhancement network to further optimize the fused image to ensure the retention of detailed information and improve the image quality. The formula is: I enhanced = QE Net (I fused ), where I enhanced is the image after quality enhancement, and QE Net represents the quality enhancement network.

[0062] In this embodiment, the steps of using a loss function to evaluate the quality of the fused image and calculate the difference between the fusion result and the original image are as follows:

[0063] Use the mean squared error loss function to calculate the difference between the fused image and the original image. The formula is: where I fused (x i , y i ) is the pixel value of the fused image, I true (x i , y i ) is the pixel value of the real image, N is the total number of pixels, and L MSE is the mean squared error loss function, which is used to measure the difference between the fused image and the original image.

[0064] In this embodiment, the steps of correcting the error through the least squares method to ensure the accuracy of the final output image are as follows:

[0065] Optimize the fusion result according to the least squares loss function to minimize the error between the images, thereby improving the image quality and ensuring the accuracy of the final output image. The formula is: where For the optimized image, the objective function is minimized to correct the error between the fused image and the real image, ensuring image accuracy.

[0066] In this embodiment, the spatial data from different sources includes but is not limited to satellite images, UAV aerial images, and topographic maps.

[0067] This experiment uses the dataset in MFQE

[14] for training and testing. The real video dataset used to evaluate the model includes 108 video sequences from the Xiph (Xiph.org)

[18] and VQEG

[19] databases. The resolutions of the videos cover a variety of specifications, including: 352×240, 352×288, 720×486, 704×576, 416×240, 640×360, 832×480, 1280×720, 1920×1080, and 2560×1600. The video dataset for testing contains 18 videos, which are selected from the Video Coding Experts Group (VCEG) and are widely used standard test sequences in video quality assessment.

[0068] All videos are compressed by the H.265 / HEVC reference software HM16.5 and processed under four different quantization parameters (QP). The specific QP values are: 22, 27, 32, and 37. The performance of the model under different compression conditions is evaluated by the quality of the compressed videos generated at different compression levels.

[0069] Implementation details

[0070] The model adopted in this paper is implemented based on the PyTorch framework. During the training process, images of size 64×64 are randomly cropped from the original videos and compressed videos as training samples, and data augmentation is performed using methods such as rotation and flipping. The Adam optimizer is used during training, and the learning rate is set to 0.0005 and remains unchanged throughout the training process. The model is trained and tested under four different QP values.

[0071] For video data, the Y channel (luminance component) in the YUV / YCbCr color space contains the main information of the video. Therefore, the method in this paper only performs quality enhancement on the luminance component and ignores the chrominance component.

[0072] Performance evaluation

[0073] The evaluation of video quality usually uses the peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) as metrics. In this study, we use the PSNR increment (ΔPSNR) and SSIM increment (ΔSSIM) relative to the HEVC compressed video frames to quantify the effect of quality enhancement.

[0074] Existing technical solutions typically involve evaluating the quality of a video after compression based on traditional video quality assessment metrics, such as Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity (SSIM). These technical solutions generally rely on the following methods:

[0075] Based on traditional video compression techniques: For example, video coding techniques such as H.264 / AVC and HEVC (H.265). These methods use different quantization parameters (QP values) to compress the video and measure the quality of the compressed video through common video quality assessment metrics such as PSNR and SSIM. These methods can process the compressed bitstream and judge the visual quality after compression, but often ignore the specific enhancement effects.

[0076] Video quality enhancement methods based on deep learning: Some studies have proposed video quality enhancement models based on deep learning, which use Convolutional Neural Networks (CNNs) or Generative Adversarial Networks (GANs) to enhance the quality of compressed videos. Deep learning models usually learn the characteristics of compressed videos through training datasets and generate enhanced images or video frames through the network to improve the performance of metrics such as PSNR and SSIM.

[0077] Enhancement methods based on traditional image processing: Some traditional video quality enhancement methods use image processing techniques, such as image denoising, sharpening, and contrast enhancement, to optimize the quality of compressed videos.

[0078] 2. Comparison experiments under the same conditions

[0079] To ensure the fairness and comparability of the experiments, the following same conditions are set when comparing the present invention with existing technical solutions:

[0080] Datasets and video sequences: The same test video dataset (such as the MFQE dataset) is used, and it is ensured that the video resolution and format are consistent. All videos are compressed using the same video coding algorithm (HEVC) and tested with the same quantization parameters (QP values: 22, 27, 32, 37).

[0081] Consistent compression level: Ensure that the same compression level is used for video compression in different methods to ensure that the comparison results are not affected by differences in compression conditions.

[0082] Same evaluation metrics: The same video quality assessment metrics, such as PSNR increment (ΔPSNR) and SSIM increment (ΔSSIM), are used to compare the enhancement effects of each solution.

[0083] The training and testing environments are consistent: For deep learning-based methods, ensure that the same network architecture, optimizer, training methods (such as Adam optimizer, learning rate, etc.), and data augmentation strategies (such as rotation, flipping, etc.) are used for the training and testing of each method.

[0084] 3. After comparison, demonstrate the beneficial effects of the present invention through data or charts.

[0085] In the comparative experiments, the PSNR increment (ΔPSNR) and SSIM increment (ΔSSIM) are used to quantify the performance improvement of each method. The following are the comparative charts and data presentations obtained from the experimental results:

[0086] Table 1: PSNR Increment Comparison Table

[0087]

[0088] Table 2: SSIM Increment Comparison Table

[0089]

[0090] The PSNR increment of the present invention under all quantization parameters (QP values) is significantly better than that of traditional HEVC compression and existing deep learning methods. Especially at lower QP values (22, 27), the PSNR increment is particularly prominent, indicating that the present invention can more effectively improve the video quality at low compression rates.

[0091] In the comparison of SSIM increment, the present invention also demonstrates better performance. Especially at high compression quantization parameters (such as QP = 22 and QP = 27), the present invention performs more excellently in maintaining the quality of the visual structure.

[0092] The present invention is not limited to the above embodiments. Anyone should know that structural changes made under the inspiration of the present invention, as long as they have the same or similar technical solutions as the present invention, fall within the protection scope of the present invention. The technologies, shapes, and structures not detailed in the present invention are all well-known technologies.

Claims

1. A method for registering and fusing spatial geographic image data of an electronic sand table, characterized in that, It includes the following steps: S1: Preprocess the spatial geographic image data from different data sources, including format conversion, denoising, and radiometric correction, to ensure data quality and consistency. Then, use a coordinate transformation algorithm to uniformly transform the image data from different data sources into the same geographic coordinate system or projection coordinate system; S2: Register the images from different data sources through a feature point extraction and matching algorithm based on a 3D convolutional neural network to make the spatial positions and scales of the image data consistent; S3: Perform pixel-level fusion on the registered image data through weighted averaging to enhance the quality of the fused image, retain more detailed information, and further optimize the image quality in combination with a quality enhancement network; S4: Use a loss function to evaluate the quality of the fused image, calculate the difference between the fused result and the original image, and correct the error through the least squares method.

2. The method for registering and fusing spatial geographic image data of an electronic sand table according to claim 1, wherein, The steps for preprocessing the spatial geographic image data from different data sources, including format conversion, denoising, and radiometric correction, to ensure data quality and consistency are: Uniformly convert the image data from different data sources into a standard format, GeoTIFF, to ensure data consistency; Use a median filtering algorithm to denoise the image data. The formula is as follows: I′(x, y) = median{I(x + i, y + j)| -k ≤ i, j ≤ k}, where I(x, y) is the original image data, I′(x, y) is the denoised image data, and k is the radius of the filtering window; Perform sensor correction on the image to eliminate the radiation deviation caused by the sensor and convert it to a reflectance value. The formula is as follows: Where I is the digital value of the original image, b is the offset, g is the gain, and R is the reflectance value after correction.

3. A method for registering and fusing spatial geographic image data of an electronic sand table according to claim 1, characterized in that, The steps for using the PROJ.4 algorithm to uniformly transform the image data from different data sources into the same geographic coordinate system or projection coordinate system are: Determine the source coordinate system and the target coordinate system. The source coordinate system is WGS84, and the target coordinate system is the UTM projection coordinate system. Then, load the PROJ.4 library and use its coordinate system definition string. The source coordinate system is defined as +proj=longlat +datum=WGS84 +no d efs, and the target coordinate system is defined as +proj=utm +zone=33 datum=WGS84; Coordinate transformation is performed through the PROJ.4 algorithm, and the image data points (X source , Y source ) in the source coordinate system are converted into the coordinate points (X traget , Y target ) in the target coordinate system using the conversion formula: where T is the transformation matrix, and [X, Y, Z] and [X′, Y′, Z′] are the coordinates in the original coordinate system and the target coordinate system, respectively.

4. A method for registering and fusing spatial geographical image data of an electronic sand table according to claim 1, characterized in that, The steps for registering the images from different data sources through a feature point extraction and matching algorithm based on a 3D convolutional neural network to make the spatial positions and scales of the image data consistent are: Use a 3D convolutional neural network to process the image data and automatically extract the key feature points in the image. The formula is: f(x, y, z) = CNN(I((x, y, z)) where f(x, y, z) are the feature points extracted from the image data I(x, y, z), and CNN represents a 3D convolutional neural network; The feature points extracted using the convolutional neural network are matched in images from different data sources, and the similarity of the feature points is calculated. The formula is as follows: where f1(x i , y i , z i ) and f2(x i , y i , z i ) are the feature points extracted at the position (x i , y i , z i ) in two images respectively, N is the total number of matched feature points, and L match is the matching loss function; According to the feature point matching results, the images are registered through geometric transformation to make their spatial positions and scales consistent. The formula is: I target = T match (I source ), where I target is the target image, I source is the source image, and T match is the registration transformation matrix obtained from the feature point matching.

5. A method for registering and fusing spatial geographic image data of an electronic sand table according to claim 1, characterized in that, The steps for performing pixel-level fusion on the registered image data through weighted averaging to enhance the quality of the fused image, retain more detailed information, and further optimize the image quality in combination with a quality enhancement network are: Perform pixel-level fusion on the registered images. The formula is: I fused (x, y) = αI1(x, y) + (1 - α)I2(x, y), where I fused (x, y) is the pixel value of the fused image, I1(x, y) and I2(x, y) are the pixel values of the two registered images, and α is the weighting coefficient; Further optimize the fused image using a quality enhancement network to ensure the retention of detailed information and improve the image quality. The formula is: I enhanced = QE Net (I fused ), where I enhanced is the image after quality enhancement, and QE Net represents the quality enhancement network.

6. A method for registering and fusing spatial geographical image data of an electronic sand table according to claim 1, characterized in that, The steps for using a loss function to evaluate the quality of the fused image and calculate the difference between the fused result and the original image are: The difference between the fused image and the original image is calculated using the mean squared error loss function, and the formula is: where I fused (x i , y i ) is the pixel value of the fused image, I true (x i , y i ) is the pixel value of the real image, N is the total number of pixels, and L MSE is the mean squared error loss function, which is used to measure the difference between the fused image and the original image.

7. A method for registering and fusing spatial geographic image data of an electronic sand table according to claim 1, characterized in that The steps for correcting the error through the least squares method to ensure the accuracy of the final output image are: Optimize the fusion result according to the least squares loss function, minimize the error between images, thereby improving the image quality and ensuring the accuracy of the final output image. The formula is as follows: where is the optimized image. Minimize the objective function to correct the error between the fused image and the real image and ensure the image accuracy.

8. A method for registering and fusing spatial geographic image data of an electronic sand table according to claim 1, characterized in that, Spatial data from different sources includes but is not limited to satellite images, drone aerial images, and topographic maps.