A semantic-based remote sensing image transmission and reconstruction system, method and product
By transmitting low-resolution remote sensing images and semantic text generated by satellite and combining them with multimodal fusion reconstruction technology at the ground end, the problems of high cost and low efficiency in remote sensing image transmission are solved, and efficient and high-fidelity image reconstruction is achieved under bandwidth-limited conditions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- THE CHINESE UNIV OF HONG KONG (SHENZHEN)
- Filing Date
- 2025-12-26
- Publication Date
- 2026-05-29
AI Technical Summary
Transmission and reconstruction of high-resolution remote sensing images are costly and inefficient due to limited satellite hardware resources and downlink bandwidth. Furthermore, existing compression transmission methods may result in the loss of critical information, affecting the accuracy of subsequent applications.
By generating low-resolution remote sensing images and compact semantic text on satellite, and combining them with multimodal fusion reconstruction technology on the ground, semantic text is generated using a visual language model and then multimodal encoding and text-guided reconstruction are performed on the ground to ultimately restore high-resolution images.
While reducing transmission bandwidth usage and communication costs, it ensures high fidelity of reconstruction results to meet the quality requirements of applications such as ground feature classification and target detection.
Smart Images

Figure CN122116158A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of satellite remote sensing image processing, and in particular to a semantic-based remote sensing image transmission and reconstruction system, method, and product. Background Technology
[0002] Remote sensing (RS) imagery supports numerous critical applications, including environmental monitoring, urban planning, precision agriculture, disaster response, and national security. These tasks heavily rely on the rich spatial and spectral information of high-resolution (HR) imagery to support reliable land cover classification, target detection, and temporal change analysis. However, the acquisition and transmission of HR imagery are severely constrained by factors such as limited satellite hardware, onboard storage space, and downlink bandwidth. With the continuous growth of remote sensing data volume, the efficient transmission of high-fidelity imagery information has become a bottleneck problem that urgently needs to be solved by global observation systems.
[0003] In related technologies, the transmission and reconstruction of high-resolution remote sensing images mainly rely on two methods: one is to directly transmit complete high-resolution image data, sending the pixel-dense original image to the ground end via satellite downlink, and then the ground end can use it after simple format parsing; the other is to use lossless or lossy compression technology to compress the high-resolution image data before transmission, and the ground end receives the compressed data and performs decompression to restore the original image information.
[0004] However, the aforementioned technologies have significant drawbacks: Firstly, directly transmitting complete high-resolution images requires a huge amount of satellite downlink bandwidth, while satellite hardware resources and downlink bandwidth themselves are physically limited, resulting in high transmission latency, high communication costs, and difficulty in adapting to scenarios where the volume of remote sensing data continues to grow. Secondly, although compression transmission methods can reduce the amount of data to some extent, both lossless and lossy compression may lead to the loss of key structural and semantic information in the images. Lossy compression, in particular, cannot recover the lost information through decompression, thus affecting the accuracy of subsequent tasks such as ground feature classification and target detection, and failing to meet the requirements of high-fidelity applications. In addition, atmospheric interference, cloud cover, and other factors further exacerbate the transmission efficiency problem, often resulting in incomplete or delayed data received at the ground end, seriously hindering the implementation of remote sensing applications with high timeliness requirements. Summary of the Invention
[0005] The purpose of this application is to provide a semantic-based remote sensing image transmission and reconstruction system, method, and product, which can at least solve the problems of high transmission cost and low transmission efficiency caused by the increase in the volume of remote sensing data in related technologies.
[0006] To address the aforementioned technical problems, the first aspect of this application provides a semantic-based remote sensing image transmission and reconstruction system, including a satellite end and a ground end; The satellite is configured to process the acquired first remote sensing image based on a first processing model to obtain a second remote sensing image and semantic text describing the first remote sensing image, and to send the second remote sensing image and the semantic text to the ground terminal; wherein the resolution of the second remote sensing image is lower than the resolution of the first remote sensing image; The ground terminal is configured to perform multimodal fusion reconstruction of the received second remote sensing image and the semantic text based on a second processing model to obtain a third remote sensing image.
[0007] In some embodiments, the first processing model includes a downsampling module and a visual language module; The satellite terminal is also configured to: The second remote sensing image is obtained by downsampling the first remote sensing image based on the downsampling module. The semantic information of the first remote sensing image is extracted based on the visual language module to obtain the semantic text.
[0008] In some embodiments, the second processing model includes a multimodal coding module, a text-guided reconstruction module, and an image decoding module; The ground terminal is also configured to: Based on the multimodal coding module, feature extraction is performed on the second remote sensing image and the semantic text respectively to obtain visual embedding information and text embedding information; Based on the text-guided reconstruction module, feature optimization processing is performed on the visual embedding information and the text embedding information to obtain the final embedding information; The final embedded information is decoded based on the image decoding module to obtain the third remote sensing image.
[0009] In some embodiments, the multimodal coding module includes a CNN image encoder, a CLIP image encoder, and a CLIP text encoder; The ground terminal is also configured to: The second remote sensing image is encoded based on the CNN image encoder to obtain CNN visual embedding information; The second remote sensing image is encoded based on the CLIP image encoder to obtain CLIP visual embedding information; The semantic text is encoded based on the CLIP text encoder to obtain CLIP text embedding information.
[0010] In some embodiments, the ground terminal is further configured to: Based on the CLIP visual embedding information and the CLIP text embedding information, the semantic feature alignment of the CNN visual embedding information is performed multiple times by the text-guided reconstruction module to obtain the final embedding information.
[0011] In some embodiments, the text-guided reconstruction module includes a semantic guidance module and a feature alignment module; The ground terminal is also configured to: Based on the semantic guidance module, the cosine similarity between the global semantic embedding information and the local semantic embedding information in the CLIP visual embedding information is calculated, and a semantic guidance graph is generated; The CLIP text embedding information is mapped using a multilayer perceptron to generate a channel attention weight vector; Based on the preset learnable alignment parameters, the channel attention weight vector, and the semantic guidance map, the semantic feature alignment of the CNN visual embedding information is performed on multiple iterations by the feature alignment module to obtain the final embedding information.
[0012] In some embodiments, the image decoding module includes a first decoding module and a second decoding module; The ground terminal is also configured to: Based on the first decoding module, the final embedded information is convolutionally processed and upsampled to obtain the initial decoded image; The third remote sensing image is obtained by performing a second decoding on the initial decoded image based on the second decoding module.
[0013] The second aspect of this application provides a semantic-based remote sensing image transmission and reconstruction method, applied to a semantic-based remote sensing image transmission and reconstruction system, including a satellite end and a ground end; the method includes: The satellite terminal processes the acquired first remote sensing image based on a first processing model to obtain a second remote sensing image and semantic text describing the first remote sensing image, and then sends the second remote sensing image and the semantic text to the ground terminal; wherein, the resolution of the second remote sensing image is lower than the resolution of the first remote sensing image; The ground terminal performs multimodal fusion reconstruction on the received second remote sensing image and the semantic text based on the second processing model to obtain a third remote sensing image.
[0014] A third aspect of this application provides an electronic device, including a memory and a processor, wherein the processor is configured to execute a computer program stored in the memory, and when the processor executes the computer program, it implements the steps of the remote sensing image transmission and reconstruction method described in the second aspect of the embodiments of this application.
[0015] The fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the remote sensing image transmission and reconstruction method described in the second aspect of the embodiments of this application.
[0016] As can be seen from the above, this application provides a semantic-based remote sensing image transmission and reconstruction system, including a satellite end and a ground end. The satellite end is configured to process a first remote sensing image acquired based on a first processing model to obtain a second remote sensing image and semantic text for describing the first remote sensing image, and to send the second remote sensing image and the semantic text to the ground end. The resolution of the second remote sensing image is lower than that of the first remote sensing image. The ground end is configured to perform multimodal fusion reconstruction on the received second remote sensing image and the semantic text based on a second processing model to obtain a third remote sensing image. This application proposes a semantically efficient remote sensing image transmission system. This system replaces the transmission of complete high-resolution data by sending low-resolution images and their concise text descriptions. At the satellite end, the system uses a visual language model to generate text representations, summarizing spatial and semantic information, thereby greatly compressing the amount of transmitted data. While significantly reducing transmission bandwidth usage and communication costs, the system retains key semantic information of high-resolution images through semantic text. Combined with multimodal fusion reconstruction at the ground end, high-fidelity image restoration is achieved. Thus, in bandwidth-constrained scenarios, it balances transmission efficiency and reconstruction quality, and effectively solves the pain point of difficulty in balancing transmission and fidelity in existing technologies.
[0017] It should be understood that the description in this section is not intended to identify key or important features of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description
[0018] To more clearly illustrate the related technologies or the technical solutions in the embodiments of this application, the drawings used in the description of the related technologies or the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application, and not all embodiments. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 A schematic diagram of the structure of a semantic-based remote sensing image transmission and reconstruction system provided in an embodiment of this application; Figure 2 A detailed structural diagram of a semantic-based remote sensing image transmission and reconstruction system provided in an embodiment of this application; Figure 3A detailed structural diagram of a semantic-based remote sensing image transmission and reconstruction system provided in an embodiment of this application; Figure 4 A conceptual structure diagram of a semantic-based remote sensing image transmission and reconstruction system provided in the embodiments of this application; Figure 5 A detailed conceptual structure diagram of a semantic-based remote sensing image transmission and reconstruction system provided in the embodiments of this application; Figure 6 The experimental results of the semantic-based remote sensing image transmission and reconstruction system provided in this application embodiment on the Alsat-2B dataset are shown in the figure. Figure 7 Figure showing the experimental results of the semantic-based remote sensing image transmission and reconstruction system provided in this application embodiment on the UC MercedLand Use dataset; Figure 8 The experimental results of the semantic-based remote sensing image transmission and reconstruction system provided in this application embodiment on the AID dataset are shown in the figure. Figure 9 Figure showing the experimental results of the semantic-based remote sensing image transmission and reconstruction system provided in this application embodiment on the ILSVRC2012 dataset; Figure 10 A comparison of the visualization results of the semantic-based remote sensing image transmission and reconstruction system provided in this application embodiment on four datasets; Figure 11 A visualization of the experimental results of the semantic-based remote sensing image transmission and reconstruction system provided in this application embodiment on the Alsat-2B dataset; Figure 12 A visualization of the experimental results of the semantic-based remote sensing image transmission and reconstruction system provided in this application embodiment on the Merced LandUse dataset; Figure 13 A visualization of the experimental results of the semantic-based remote sensing image transmission and reconstruction system provided in this application embodiment on the AID dataset; Figure 14 A visualization of the experimental results of the semantic-based remote sensing image transmission and reconstruction system provided in this application embodiment on the ILSVRC2012 dataset; Figure 15 A flowchart illustrating the semantic-based remote sensing image transmission and reconstruction method provided in this application embodiment; Figure 16 A module block diagram of the electronic device provided in the embodiments of this application; Figure 17 A block diagram of a computer-readable storage medium provided in an embodiment of this application. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of this application more apparent and understandable, this application will be clearly and completely described below in conjunction with its embodiments and accompanying drawings. Throughout, the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions. It should be understood that the various embodiments of this application described below are merely illustrative of this application and are not intended to limit this application. That is, all other embodiments obtained by those skilled in the art based on the various embodiments of this application without creative effort are within the scope of protection of this application. Furthermore, the technical features involved in the various embodiments of this application described below can be combined with each other as long as they do not conflict with each other.
[0021] Remote sensing (RS) imagery supports numerous critical applications, including environmental monitoring, urban planning, precision agriculture, disaster response, and national security. These tasks heavily rely on the rich spatial and spectral information of high-resolution (HR) imagery to support reliable land cover classification, target detection, and temporal change analysis. However, the acquisition and transmission of HR imagery are severely constrained by factors such as limited satellite hardware, onboard storage space, and downlink bandwidth. With the continuous growth of remote sensing data volume, the efficient transmission of high-fidelity imagery information has become a bottleneck problem that urgently needs to be solved by global observation systems.
[0022] In related technologies, the transmission and reconstruction of high-resolution remote sensing images mainly rely on two methods: one is to directly transmit complete high-resolution image data, sending the pixel-dense original image to the ground end via satellite downlink, and then the ground end can use it after simple format parsing; the other is to use lossless or lossy compression technology to compress the high-resolution image data before transmission, and the ground end receives the compressed data and performs decompression to restore the original image information.
[0023] However, the aforementioned technologies have significant drawbacks: Firstly, directly transmitting complete high-resolution images requires a huge amount of satellite downlink bandwidth, while satellite hardware resources and downlink bandwidth themselves are physically limited, resulting in high transmission latency, high communication costs, and difficulty in adapting to scenarios where the volume of remote sensing data continues to grow. Secondly, although compression transmission methods can reduce the amount of data to some extent, both lossless and lossy compression may lead to the loss of key structural and semantic information in the images. Lossy compression, in particular, cannot recover the lost information through decompression, thus affecting the accuracy of subsequent tasks such as ground feature classification and target detection, and failing to meet the requirements of high-fidelity applications. In addition, atmospheric interference, cloud cover, and other factors further exacerbate the transmission efficiency problem, often resulting in incomplete or delayed data received at the ground end, seriously hindering the implementation of remote sensing applications with high timeliness requirements.
[0024] Therefore, there is a need for a system that can significantly reduce transmission bandwidth usage and communication costs, and balance transmission efficiency and reconstruction quality in bandwidth-constrained scenarios, in order to solve the problems of high transmission costs and low transmission efficiency caused by the increasing volume of remote sensing data in existing technologies.
[0025] This application provides a semantic-based remote sensing image transmission and reconstruction system. Specifically, please refer to... Figure 1 , Figure 1 This is a schematic diagram of the structure of a semantic-based remote sensing image transmission and reconstruction system provided in an embodiment of this application. The remote sensing image transmission and reconstruction system includes a satellite terminal 10 and a ground terminal 20. The satellite terminal 10 is configured to process a first remote sensing image acquired based on a first processing model 101 to obtain a second remote sensing image and semantic text describing the first remote sensing image, and to send the second remote sensing image and semantic text to the ground terminal 20. The resolution of the second remote sensing image is lower than that of the first remote sensing image. The ground terminal 20 is configured to perform multimodal fusion reconstruction on the received second remote sensing image and semantic text based on a second processing model 201 to obtain a third remote sensing image.
[0026] Specifically, the satellite first acquires a high-resolution first remote sensing image. This first remote sensing image is a high-resolution raw image containing rich spatial and semantic information such as ground features and terrain. Because of the high resolution of the first remote sensing image, direct transmission would consume a large amount of satellite downlink bandwidth and significantly increase communication costs. Therefore, the satellite processes the first remote sensing image using a pre-configured first processing model: on the one hand, it generates a second remote sensing image with a lower resolution than the first remote sensing image. The second remote sensing image retains the basic spatial structure information of the first remote sensing image, and the data volume is significantly reduced, adapting to bandwidth-constrained transmission scenarios; on the other hand, it extracts the core semantic information of the first remote sensing image and generates semantic text to describe it. This semantic text, in a compact form, supplements any semantic details that may be missing in the second remote sensing image, ensuring that key information is not lost. After completing the above processing, the satellite sends the second remote sensing image and the semantic text together to the ground station.
[0027] After receiving the second remote sensing image and semantic text transmitted from the satellite, the ground station performs multimodal fusion reconstruction using a pre-configured second processing model. This model fully utilizes the spatial structure information of the second remote sensing image and the semantic information of the semantic text, compensating for the insufficient resolution of the second remote sensing image through cross-modal information fusion and reconstruction, ultimately outputting a third remote sensing image. This third remote sensing image highly matches the original first remote sensing image in terms of visual detail and semantic consistency, meeting the data quality requirements for subsequent applications such as land cover classification and target detection.
[0028] This system replaces the traditional transmission of complete high-resolution images by combining the transmission of low-resolution images and semantic text. This significantly reduces the amount of data transmitted, saves bandwidth resources and communication costs, while ensuring the fidelity of the reconstruction results through the supplementary information of semantic text, thus achieving a balance between transmission efficiency and reconstruction quality.
[0029] Figure 2 A detailed structural diagram of a semantic-based remote sensing image transmission and reconstruction system provided in this application embodiment is shown below. Figure 2 As shown, in some embodiments of this application, the first processing model 101 includes a downsampling module 1011 and a visual language module 1012; the satellite end 10 is also configured to: perform downsampling processing on the first remote sensing image based on the downsampling module 1011 to obtain a second remote sensing image; and extract semantic information from the first remote sensing image based on the visual language module 1012 to obtain semantic text.
[0030] Specifically, for the first remote sensing image acquired by the satellite, the downsampling module first performs resolution compression on the first remote sensing image using a preset downsampling strategy. While preserving the basic spatial structure of the first remote sensing image (such as the overall outline of ground features and the spatial layout of the scene), it generates a second remote sensing image with a lower resolution than the first. The core purpose of this processing is to significantly reduce the image data volume, making it compatible with the bandwidth limitations of the satellite downlink. Although the resolution of the second remote sensing image is reduced, it still retains the core spatial topology of the first remote sensing image, avoiding spatial structure distortion during ground-based reconstruction.
[0031] Simultaneously, the visual language module performs semantic information extraction processing on the first remote sensing image. This module identifies, summarizes, and structures key semantic information such as land cover types, terrain features, and scene structure in the first remote sensing image, ultimately generating semantic text to represent the core semantics of the first remote sensing image. For example, when the first remote sensing image is a mountain scene, the semantic text can be described as "mountainous terrain, with some areas covered by vegetation and significant topographic relief"; when the first remote sensing image is an urban area, the semantic text can be described as "urban built-up area, containing dense buildings and road networks." This semantic text condenses the high-level semantic information of the first remote sensing image in a compact form, supplementing the detailed semantics lost in the second remote sensing image due to its reduced resolution.
[0032] Figure 3 A detailed structural diagram of a semantic-based remote sensing image transmission and reconstruction system provided in this application embodiment is shown below. Figure 3 As shown, in some embodiments of this application, the second processing model 201 includes a multimodal coding module 2011, a text-guided reconstruction module 2012, and an image decoding module 2013; the ground terminal 20 is further configured to: extract features from the second remote sensing image and semantic text based on the multimodal coding module 2011 to obtain visual embedding information and text embedding information; perform feature optimization processing on the visual embedding information and text embedding information based on the text-guided reconstruction module 2012 to obtain final embedding information; and decode the final embedding information based on the image decoding module 2013 to obtain a third remote sensing image.
[0033] Specifically, after receiving the second remote sensing image and semantic text at the ground station, the multimodal coding module first performs feature extraction. For the second remote sensing image, this module extracts its basic spatial structure and visual features to generate visual embedding information, which preserves the pixel distribution patterns and core spatial topological relationships of the second remote sensing image. For the semantic text, this module extracts its high-level semantic features to generate text embedding information, which transforms the compact text description into a structured semantic representation that can be processed by the model. Through this process, the input data from the two different modalities are uniformly transformed into a feature form that the model can recognize.
[0034] Subsequently, the text-guided reconstruction module initiates feature optimization processing. This module uses visual embedding information as a foundation and text embedding information as semantic guidance. Through cross-modal information fusion and feature optimization strategies, it refines and enhances the visual embedding information. On one hand, the text embedding information supplements the semantic details missing from the visual embedding information due to insufficient resolution of the second remote sensing image; on the other hand, through the collaborative optimization of both, the visual embedding information maintains its original spatial structure while remaining highly consistent with the semantic text description, avoiding semantic deviations during reconstruction, and ultimately outputting optimized final embedding information.
[0035] Finally, the image decoding module performs decoding operations on the final embedded information. This module uses specific decoding logic to restore the abstract final embedded information into concrete image pixel data, gradually recovering the image's resolution and visual details, and ultimately outputting the third remote sensing image. This third remote sensing image, through a complete process of multimodal coding, feature optimization, and decoding, compensates for the resolution deficiencies of the second remote sensing image while maintaining semantic consistency with the original first remote sensing image. This meets the requirements of downstream applications such as land cover classification and target detection for high-resolution, high-fidelity remote sensing imagery.
[0036] In some embodiments of this application, the multimodal coding module 2011 includes a CNN image encoder, a CLIP image encoder, and a CLIP text encoder; the ground terminal 20 is further configured to: encode a second remote sensing image based on the CNN image encoder to obtain CNN visual embedding information; encode the second remote sensing image based on the CLIP image encoder to obtain CLIP visual embedding information; and encode semantic text based on the CLIP text encoder to obtain CLIP text embedding information.
[0037] Specifically, for the second remote sensing image, the multimodal coding module employs a dual-image encoder structure to extract features from different dimensions. On one hand, a CNN image encoder encodes the second remote sensing image. CNN image encoders excel at capturing local spatial features and texture details, fully extracting basic visual information such as the contour structure and pixel distribution patterns of ground objects in the second remote sensing image, ultimately generating CNN visual embedding information. On the other hand, a CLIP image encoder simultaneously encodes the second remote sensing image. The CLIP image encoder possesses powerful semantic understanding capabilities, extracting high-level semantic features from the second remote sensing image to generate CLIP visual embedding information. Since remote sensing images often contain repetitive texture patterns and similar scene types, relying solely on a single encoder is insufficient to fully cover spatial and semantic features. CLIP visual embedding information serves as supplementary semantic information, providing a basis for subsequent feature alignment and semantic calibration. Simultaneously, semantic text transmitted from the satellite is encoded using a CLIP text encoder. The CLIP text encoder transforms compact natural language descriptions into structured semantic representations that the model can process, generating CLIP text embedding information. This embedded information accurately maps the core semantics contained in the semantic text, such as land cover types and terrain features.
[0038] In some embodiments of this application, the ground terminal 20 is further configured to: perform multiple iterations of semantic feature alignment on the CNN visual embedding information based on CLIP visual embedding information and CLIP text embedding information, and obtain the final embedding information by the text-guided reconstruction module 2012.
[0039] Specifically, while the CNN visual embedding information, serving as the primary feature carrier for subsequent reconstruction, already contains the local spatial structure and texture details of the second remote sensing image, its semantic completeness and accuracy are insufficient due to the low resolution of the second remote sensing image. CLIP visual embedding information, on the other hand, possesses high-level semantic understanding capabilities, enabling it to extract supplementary semantic features related to scene categories and core features from the second remote sensing image; CLIP text embedding information directly maps to the core semantic description of the original first remote sensing image. Furthermore, to achieve in-depth feature optimization, the text-guided reconstruction module employs a multi-iterative semantic feature alignment strategy: in each iteration, the supplementary semantics of the CLIP visual embedding information and the core semantics of the CLIP text embedding information are used as supervision to perform semantic calibration and feature refinement on the current CNN visual embedding information. The advantage of this iterative processing is that it can gradually correct potential semantic biases in the CNN visual embedding information while continuously supplementing spatial detail relationships missing due to low resolution, allowing the CNN visual embedding information to continuously approach the semantic features of the original first remote sensing image while preserving the integrity of the original spatial structure.
[0040] In some embodiments of this application, the text-guided reconstruction module 2012 includes a semantic guidance module and a feature alignment module; the ground terminal 20 is further configured to: calculate the cosine similarity between global semantic embedding information and local semantic embedding information in CLIP visual embedding information based on the semantic guidance module, and generate a semantic guidance map; map the CLIP text embedding information through a multilayer perceptron to generate a channel attention weight vector; and perform multiple iterations of semantic feature alignment on the CNN visual embedding information through the feature alignment module according to preset learnable alignment parameters, channel attention weight vectors and semantic guidance map, to obtain the final embedding information.
[0041] Specifically, the ground-based module first generates semantic spatial guidance information through a semantic guidance module. For the CLIP visual embedding information obtained during the multimodal coding stage, the semantic guidance module first extracts its global and local semantic embedding information. The global semantic embedding information reflects the overall scene semantics of the second remote sensing image, while the local semantic embedding information corresponds to the local feature semantics of different regions in the image. Then, the cosine similarity between the global and local semantic embedding information is calculated. This similarity measure quantifies the semantic association between different local regions and the overall scene, and these similarity values are arranged according to the spatial location of each local region to generate a semantic guidance map. This semantic guidance map has a pixel-level weight distribution, which can explicitly mark the semantically key regions in the second remote sensing image, providing clear spatial guidance for subsequent feature alignment and ensuring that the model focuses on optimizing the features of semantically key regions during reconstruction.
[0042] Concurrently, for the CLIP text embedding information obtained in the multimodal encoding stage, feature mapping processing is performed through a multilayer perceptron. This maps the CLIP text embedding information from the original semantic representation dimension to a dimension matching the number of channels in the CNN visual embedding information, generating a channel attention weight vector. This vector can accurately capture the correlation between the core semantics in the semantic text and the features of each channel in the CNN visual embedding, providing a semantic filtering basis for feature alignment in the channel dimension. This helps the model focus on feature channels that are highly relevant to the semantic description and filter out redundant channel information.
[0043] After completing the above preparations, the feature alignment module initiates the semantic feature alignment operation, using preset learnable alignment parameters as the adjustment basis. These parameters can be dynamically adapted according to the reconstruction effect during training, balancing the fusion ratio of channel-level semantic guidance and pixel-level spatial guidance. The feature alignment module fuses the channel attention weight vector with the semantic guidance map to form a multi-dimensional semantic guidance signal. Subsequently, based on this signal, fine-grained semantic feature alignment is performed on the CNN visual embedding information to correct feature components in the CNN visual embedding information that are inconsistent with the semantic description, and to supplement the semantic association details missing due to the low resolution of the second remote sensing image.
[0044] To achieve in-depth feature optimization, the semantic feature alignment process described above is executed iteratively a preset number of times: each iteration is based on the latest CNN visual embedding information, combined with a fixed semantic guidance map and channel attention weight vectors, and the feature alignment module continuously optimizes the semantic consistency and spatial details of the features; during the iteration, the preset learnable alignment parameters are dynamically adjusted according to the feedback of the loss function to ensure that the fusion ratio of channel-level and pixel-level guidance always adapts to the current feature optimization needs. Through multiple rounds of iterative semantic feature alignment, the final embedding information that combines accurate semantic expression and complete spatial structure is finally obtained.
[0045] In detail, Figure 4 A conceptual structure diagram of a semantic-based remote sensing image transmission and reconstruction system provided in this application embodiment is shown below. Figure 4 As shown, the process from satellite remote sensing data acquisition to ground-based high-resolution image reconstruction consists of five sequentially linked stages, with the data flow logic as follows: 1. Remote Sensing Data Acquisition and Satellite Processing The satellite first acquires high-resolution remote sensing images (HR), and then performs two types of processing at the satellite end: (1) Extract semantic information from HR through the Visual Language Module (VLM) to generate semantic text containing scene feature descriptions; (2) The resolution of the HR is compressed by the downsampling module to generate a low-resolution image (LR).
[0046] 2. Data Preparation The system receives semantic text and LR output from the satellite, encapsulates them into multimodal transmission data, and transmits them to the subsequent ground processing stage.
[0047] 3. Multimodal Encoding Feature encoding is performed on the output of the data preparation stage to obtain three types of embedding information: (1) Encode the semantic text using the CLIP text encoder to generate CLIP text embedding information; (2) Encode LR using a CNN image encoder to generate CNN visual embedding information; (3) Encode LR using CLIP image encoder to generate CLIP visual embedding information.
[0048] 4. Text-Guided Iterative Resolution Using the three types of embedding information obtained from multimodal encoding as input, the CNN visual embedding information is subjected to multiple semantic feature alignment and feature enhancement through a text-guided iterative optimization strategy to obtain the optimized final embedding information.
[0049] 5. Image Decoding The optimized final embedded information is input into the image decoder, and resolution restoration and detail optimization processes are performed sequentially to finally output a high-resolution reconstructed image (corresponding to the third remote sensing image in the scheme), thus completing the entire remote sensing image transmission and reconstruction process.
[0050] In some embodiments of this application, the image decoding module 2013 includes a first decoding module and a second decoding module; the ground terminal 20 is further configured to: perform convolution processing and upsampling on the final embedded information based on the first decoding module to obtain an initial decoded image; and perform secondary decoding on the initial decoded image based on the second decoding module to obtain a third remote sensing image.
[0051] Specifically, the ground-based system first performs basic decoding and resolution restoration on the final embedded information through a first decoding module. The final embedded information, as an abstract representation integrating spatial structure and semantic features, is first input into the first decoding module for convolution processing. This convolution operation aims to refine the local details of the feature map and strengthen the representation of key visual information such as ground feature outlines and textures, laying a clear feature foundation for subsequent resolution restoration. Subsequently, the first decoding module performs upsampling operations at a ratio adapted to the downsampling ratio of the satellite end, gradually increasing the resolution of the feature map and mapping the abstract low-dimensional features into image data with actual spatial dimensions, ultimately outputting the initial decoded image. Further, a second decoding module performs secondary decoding optimization on the initial decoded image. Through specific convolutional decoding logic, the second decoding module finely adjusts the pixel relationships, texture details, and color consistency in the initial decoded image. This fine optimization through secondary decoding significantly improves the visual details and semantic fidelity of the initial decoded image, ultimately outputting a third remote sensing image that meets the requirements of downstream applications.
[0052] In detail, Figure 5 A detailed conceptual structure diagram of a semantic-based remote sensing image transmission and reconstruction system provided in this application embodiment is shown below. Figure 5 As shown, the specific process is as follows: 1. Pre-input: This module receives three types of inputs (all of which are outputs from the multimodal coding stage): (1) Text embedding information: Semantic representation of semantic text generated by CLIP text encoder; (2) CNN visual embedding information: main visual features generated by a CNN image encoder from a low-resolution (LR) image; (3) CLIP visual embedding information: supplementary semantic features generated by the LR image encoder via CLIP.
[0053] 2. Initial Feature Alignment: The above three types of inputs are fed into the initial feature alignment module to complete the initial fusion of multimodal features. With CNN visual embedding information as the basic feature carrier, the supplementary semantics of CLIP visual embedding information and the core semantic guidance of text embedding information are introduced. Through operations such as feature splicing and dimension unification, the semantic association of the three types of features is initially calibrated, and pre-aligned visual features are output.
[0054] 3. Text-Guided Iterative Super-Resolution This is the core iterative process. Each iteration includes three steps: "context aggregation → super-resolution enhancement → text calibration," and embeds normalization components such as LayerNorm, TNA, and WNA. (1) Aggregate multi-scale global context After receiving the pre-aligned visual features, execute the following steps sequentially: LayerNorm (layer normalization): Normalizes the channel dimension of features to stabilize the feature distribution and avoid gradient explosion / vanishing in subsequent calculations; TNA (Token Normalization Attention): Performs normalization and attention operations on sub-feature blocks after features are divided by tokens to strengthen the local semantic association between tokens; WNA (Window Normalization Attention): Divides features into fixed-size windows, performs normalization and attention operations within the windows, and captures the spatial dependencies of local regions of the image. The final output is a feature that integrates multi-scale context, achieving a unification of global semantics and local spatial information of the scene.
[0055] (2) Super-resolution unit (SR Unit) Enhancements to features that incorporate multi-scale context are performed in detail, primarily relying on TTSA (Top-k TokenSelective Attention): TTSA is the specific mechanism for implementing the Residual Token Selection Group (RTSG) architecture in the disclosure process: first, the features are divided into multiple tokens, and then the tokens with more semantic / spatial information are selected through Top-k selection to enhance their representation capabilities; Perform residual connection operations on the filtered tokens to retain the basic information of the original features, while superimposing enhanced detailed features to output super-resolution enhanced features.
[0056] (3) Text-Guided Aligning After each round of super-resolution enhancement, the text embedding information is invoked again for semantic calibration. Using the core semantics of the text embedding as a reference, the current semantic representation of the super-resolution enhanced features is compared, correcting any components in the features that are inconsistent with the semantic text description. This ensures that the features after each iteration do not deviate from the semantic logic of the original remote sensing image. The aforementioned "context aggregation → super-resolution enhancement → text calibration" process iterates a preset number of times, continuously improving the spatial detail and semantic consistency of the features.
[0057] 4. Text-Guided Feature Alignment: This stage is the parallel fine-alignment module of the iterative process, focusing on the precise optimization of CNN visual embedding information. Its core components include TTSA and WSA. Token partitioning: Divides the input CNN visual embedding information into feature tokens of fixed size to adapt to the processing granularity of the attention mechanism; WSA (Window Self-Attention): Performs self-attention operations within a window that divides tokens, strengthening the spatial correlation of tokens within the same window and improving the representation accuracy of local textures; TTSA (Top-k Token Selection Attention): Re-filters key tokens within the window, prioritizing token features that enhance semantic / spatial information; Aligning Operator: Combining the supplementary semantics of CLIP visual embedding information and the guiding signals of text embedding information, it performs dimension mapping and weight adjustment on the optimized token features, and outputs finely aligned CNN visual embedding information; The aligned features are fed back into the iteration process, further enhancing the optimization effect of each iteration.
[0058] 5. Output and Joining (Joining Image Decoding) After multiple rounds of iteration and fine alignment, the final optimized visual features are output and passed to the image decoder. After convolution, upsampling, and secondary decoding, a high-resolution reconstructed image (SR) is generated, completing the entire super-resolution reconstruction process of the remote sensing image.
[0059] This application also provides a specific implementation based on PyTorch. All experiments were conducted on a single hardware platform: an Intel Core i9-10900X processor, 128 GiB of memory, and a single NVIDIA GeForce RTX 4090 graphics card; the operating system was Ubuntu 22.04 (Jammy), and the CUDA version was 12.4. The Text-RSSR framework and the baseline models used for comparison all employed L1 loss as the training objective and used the AdamW optimizer with an initial learning rate of 1×10⁻⁶. -4 The learning rate is halved halfway through the training process. Low-resolution images are obtained by downsampling high-resolution images by a factor of 4. To simulate computationally limited scenarios, the number of training epochs is set to a moderate scale. Except for CLIP-related components, all models are trained from scratch. Text-RSSR uses LLaVa-1.5 7B as the VLM for generating image descriptions; in the multimodal encoding stage, both the visual encoder and text encoder of CLIP use pre-trained weights and remain frozen during training. The initial alignment coefficient α in text-guided feature alignment is set to 0.5; the number of super-resolution units l is set to 6; RTSG super-resolution units are initialized according to the TTST original configuration. Benchmark evaluations are performed on three public RSSR datasets (Alsat-2B, UC Merced Land Use, AID) and a common object dataset ILSVRC2012. Evaluation metrics are Peak Signal-to-Noise Ratio (PSNR, calculated in RGB channels) and Structural Similarity (SSIM). Compared to existing methods, Text-RSSR performs robustly on RSSR tasks.
[0060] Furthermore, this application conducts experiments on three public RSSR datasets: Alsat-2B, UC Merced Land Use, and AID, as shown below. Figure 6 , Figure 7 and Figure 8 As shown, the super-resolution capability of Text-RSSR for common objects was further evaluated at ILSVRC2012, such as... Figure 9 As shown in the results, the method performs well on all datasets. The datasets are described below: 1. Alsat-2B: Proposed by A. Djerida et al., this dataset contains 2182 training samples and three test subsets (agriculture, city, and special scenarios), corresponding to 56, 282, and 239 image pairs, respectively. Each high-resolution image is 256×256, and each low-resolution image is 64×64, with a spatial resolution of 2.5 meters. The dataset exhibits significant color style differences between high-resolution (HR) and low-resolution (LR) models, increasing the difficulty of RSSR. All models were trained for 50 epochs, and the batch size for Text-RSSR was set to 4. Table "Alsat-2B Quantitative Results" presents the parameter count and FLOPs for each model, and reports PSNR and SSIM across four dimensions: "Average," "Agriculture," "Special Scenario," and "City." Best values are indicated in bold, and second-best values are indicated in underline.
[0061] 2. UC Merced Land Use: Proposed by Y. Yang et al., this method comprises 21 categories, with 100 high-resolution images per category, approximate resolution of 256×256, and spatial resolution of 0.3 meters. For uniformity, this application adjusts all HR images to 256×256. 75 images are randomly selected from each category for training, and the remaining 25 are used for testing. All models are trained for 100 epochs, and the batch size for Text-RSSR is set to 4. Table "UC Merced Quantitative Results" reports the PSNR / SSIM ratio by category and overall average, marking the best and second-best results.
[0062] 3. AID: Proposed by G.-S. Xia et al., this dataset contains 10,000 600×600 RGB images, covering 30 scene categories, with a spatial resolution of approximately 0.5 meters. In this application, 80% of the images in each category are randomly selected for training and 20% for testing (a total of 8000 / 2000 images). To evaluate adaptability to different resolutions, the HR (Highest Resolution) is uniformly adjusted to 576×576, and the LR (Lower Resolution) is 144×144 (different from other datasets). All models are trained for 20 epochs, and the batch size for Text-RSSR is set to 1. Table "AID Quantitative Results" provides a comparison between each category and the average.
[0063] 4. ILSVRC2012: Proposed by O. Russakovsky et al., this method includes 1000 categories, approximately 1300 training images per category, and 50,000 validation and 100,000 test images. This application randomly selects 20 images from each category to form the training set (20,000 images in total), and randomly samples 1000 images from the original test set. The HR is uniformly set to 256×256, and LR is 64×64. All models are trained for 20 epochs, and the batch size for Text-RSSR is 4. Table "ILSVRC2012 Quantitative Results" provides the overall PSNR / SSIM for each method.
[0064] To evaluate the performance of the proposed Text-RSSR, several representative methods from 2024–2025 were selected for comparison, including FAT, FMSR, GCRDN, TSFNet, and TTST. PSNR and SSIM were used as evaluation metrics.
[0065] Based on the quantitative results of the various datasets (corresponding to the tables above), Text-RSSR achieved the highest overall PSNR and SSIM across the three remote sensing datasets. The parameter count and FLOPs for each model are listed in the comparison table with Alsat-2B (FLOPs are calculated using Python's thop library based on 64×64 RGB input). Compared to TTST, Text-RSSR improved PSNR by approximately 0.0878 dB and SSIM by approximately 0.0160 dB with almost no increase in computation (FLOPs are only 0.02 G more), demonstrating high computational efficiency. Due to the domain shift between LR and HR in Alsat-2B, FAT and GCRDN showed a significant performance degradation on this dataset (PSNR not exceeding 16 dB), indicating limited robustness to domain shift. On the other hand, UC Merced and AID include a richer range of scene categories, making the advantages of text guidance more significant: On UC Merced (21 categories), Text-RSSR achieves a PSNR of approximately 27.4904 dB and an SSIM of approximately 0.7775, surpassing the suboptimal method by approximately 0.669 dB and 0.0037, respectively; on AID (30 categories), Text-RSSR achieves a PSNR of approximately 27.4089 dB and an SSIM of 0.7116, respectively, improving upon the suboptimal method by approximately 0.0742 dB and 0.0024. These quantitative results further validate the effectiveness of introducing text guidance.
[0066] From the visualization results, such as Figure 10As shown in the comparison charts (corresponding to four datasets): Text-RSSR not only has higher PSNR / SSIM, but also better visual quality. In Alsat-2B, when there is an LR / HR domain offset, Text-RSSR can recover the HR color gamut better; in UC Merced, when there are repeating line textures, Text-RSSR tends to "reconstruct lines" rather than just "enhance lines"; in categories such as parking lots in AID, Text-RSSR achieves the current best performance; on general objects in ILSVRC2012, Text-RSSR can recover repeating textures (such as guardrails) and also better reconstruct contextual details (such as bells).
[0067] Furthermore, this application provides a visual analysis of Text-RSSR, including automatically generated text descriptions and heatmaps of CLIP visual embeddings (such as...). Figures 11 to 14 As shown in the figure. It can be seen that text descriptions can often supplement the details not covered by CLIP visual embeddings; at the same time, the model’s focus on the image area will change in stages during the text-guided iteration process. For example, in the 3rd to 5th iterations of some samples, the model focuses more on the sea surface, while in the remaining iteration stages, it focuses more on the beach and high-frequency context, thus jointly promoting high-resolution reconstruction.
[0068] To further evaluate cross-domain adaptability, this application conducted bidirectional domain transfer experiments between Alsat-2B and UC Merced, setting up a "training on ILSVRC2012, testing on Alsat-2B / UC Merced" configuration. Quantitative results show that Text-RSSR has strong cross-domain generalization ability: it achieves the highest SSIM (approximately 0.5364) when "training on Alsat-2B → testing on UC Merced"; it ranks second in PSNR when "training on UC Merced → testing on Alsat-2B"; and it achieves the best performance in both PSNR and SSIM when "training on ILSVRC2012 → testing on UC Merced". Overall, Text-RSSR is highly adaptable to domain changes, and training on datasets with higher scene diversity can further improve its cross-domain performance.
[0069] To evaluate scalability under small sample conditions, this application constructed three reduced training sets (1 / 2, 1 / 4, and 1 / 8) on UC Merced and compared them with the full training set. The results show that Text-RSSR exhibits the fastest performance improvement with increasing training data; it remains competitive even with very limited data: achieving the best performance in the 1 / 2 and Full settings; ranking second in SSIM in the 1 / 8 setting; and ranking second in PSNR and first in SSIM in the 1 / 4 setting. This demonstrates that text-guided training not only improves reconstruction quality with full data but also provides robustness in low-data scenarios, showcasing good generalization ability and practical value.
[0070] In addition, to analyze the role of each component in text-guided feature alignment, this paper conducted a systematic ablation study on UC Merced. The specific steps are as follows: 1) Set text-guided feature alignment only once before the first super-resolution unit, use only CLIP visual embeddings, and remove text embeddings; 2) Use only text embeddings and remove CLIP visual embeddings; 3) Use CLIP for both visual and text embedding; 4) Text-guided feature alignment is added after each super-resolution unit to form a hierarchical structure; 5) Introduce learnable alignment coefficients for each text-guided feature alignment to obtain the complete Text-RSSR.
[0071] Quantitative results show that introducing CLIP visual embedding (pixel-level guidance) or text embedding (channel-level guidance) alone can improve performance, with text embedding providing a slightly larger improvement. Combining the two can further improve reconstruction quality, indicating that pixel-level and channel-level guidance are complementary. After introducing a hierarchical structure, PSNR and SSIM continue to improve slightly, indicating that multi-layer alignment is more conducive to the fusion of visual and textual cues. Finally, adding learnable alignment coefficients to each alignment process can further improve flexibility and performance, enabling the model to utilize textual guidance information more effectively.
[0072] In summary, this application provides a semantic-based remote sensing image transmission and reconstruction system, including a satellite end and a ground end. The satellite end is configured to process a first remote sensing image acquired based on a first processing model to obtain a second remote sensing image and semantic text describing the first remote sensing image, and to transmit the second remote sensing image and semantic text to the ground end. The resolution of the second remote sensing image is lower than that of the first remote sensing image. The ground end is configured to perform multimodal fusion reconstruction on the received second remote sensing image and semantic text based on a second processing model to obtain a third remote sensing image. This application proposes a semantically efficient remote sensing image transmission system. This system replaces the transmission of complete high-resolution data by sending low-resolution images and their concise text descriptions. At the satellite end, the system uses a visual language model to generate text representations, summarizing spatial and semantic information, thereby greatly compressing the amount of transmitted data. While significantly reducing transmission bandwidth usage and communication costs, the system retains key semantic information of high-resolution images through semantic text. Combined with multimodal fusion reconstruction at the ground end, high-fidelity image restoration is achieved. Thus, in bandwidth-constrained scenarios, it balances transmission efficiency and reconstruction quality, and effectively solves the pain point of difficulty in balancing transmission and fidelity in existing technologies.
[0073] Figure 15 A flowchart illustrating a semantic-based remote sensing image transmission and reconstruction method provided in a third aspect of this application is shown. The semantic-based remote sensing image transmission and reconstruction method is applied to a semantic-based remote sensing image transmission and reconstruction system, including a satellite end and a ground end; the method includes: Step 1501: The satellite processes the acquired first remote sensing image based on the first processing model to obtain a second remote sensing image and semantic text describing the first remote sensing image, and sends the second remote sensing image and semantic text to the ground terminal; wherein, the resolution of the second remote sensing image is lower than that of the first remote sensing image. Step 1502: The ground end performs multimodal fusion reconstruction on the received second remote sensing image and semantic text based on the second processing model to obtain the third remote sensing image.
[0074] The embodiments of each step have been described in detail above and will not be repeated here.
[0075] Please see Figure 16 , Figure 16 A block diagram of an electronic device provided in an embodiment of this application.
[0076] like Figure 16As shown, this application embodiment also provides an electronic device that can be used to implement the remote sensing image transmission and reconstruction method in the foregoing embodiments. The electronic device includes a memory 1601 and at least one processor 1602. The memory 1601 is used to store at least one program, and when the at least one program is executed by the at least one processor 1602, the at least one processor 1602 executes the remote sensing image transmission and reconstruction method provided in this application embodiment.
[0077] Please see Figure 17 , Figure 17 A block diagram of a computer-readable storage medium provided in an embodiment of this application.
[0078] like Figure 17 As shown, this application embodiment also provides a computer-readable storage medium 1700, which stores executable instructions 1710. When the executable instructions 1710 are executed, they perform the remote sensing image transmission and reconstruction method provided in this application embodiment.
[0079] The steps of the methods or algorithms described in conjunction with the embodiments disclosed in this application can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, compact disc read-only memory (CD-ROM), or any other form of storage medium known in the art.
[0080] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., Digital Video Disk, DVD), or a semiconductor medium (e.g., Solid State Disk).
[0081] It should be noted that the various embodiments in this application are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For product-related embodiments, since they are similar to method-related embodiments, the descriptions are relatively simple, and relevant parts can be referred to the descriptions of the method-related embodiments.
[0082] It should also be noted that, in this application, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0083] The above description of the disclosed embodiments enables those skilled in the art to implement or use the content of this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined in this application may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A semantic-based remote sensing image transmission and reconstruction system, characterized in that, Including satellite and ground terminals; The satellite is configured to process the acquired first remote sensing image based on a first processing model to obtain a second remote sensing image and semantic text describing the first remote sensing image, and to send the second remote sensing image and the semantic text to the ground terminal; wherein the resolution of the second remote sensing image is lower than the resolution of the first remote sensing image; The ground terminal is configured to perform multimodal fusion reconstruction of the received second remote sensing image and the semantic text based on a second processing model to obtain a third remote sensing image.
2. The remote sensing image transmission and reconstruction system according to claim 1, characterized in that, The first processing model includes a downsampling module and a visual language module; The satellite terminal is also configured to: The second remote sensing image is obtained by downsampling the first remote sensing image based on the downsampling module. The semantic information of the first remote sensing image is extracted based on the visual language module to obtain the semantic text.
3. The remote sensing image transmission and reconstruction system according to claim 1, characterized in that, The second processing model includes a multimodal coding module, a text-guided reconstruction module, and an image decoding module; The ground terminal is also configured to: Based on the multimodal coding module, feature extraction is performed on the second remote sensing image and the semantic text respectively to obtain visual embedding information and text embedding information; Based on the text-guided reconstruction module, feature optimization processing is performed on the visual embedding information and the text embedding information to obtain the final embedding information; The final embedded information is decoded based on the image decoding module to obtain the third remote sensing image.
4. The remote sensing image transmission and reconstruction system according to claim 3, characterized in that, The multimodal coding module includes a CNN image encoder, a CLIP image encoder, and a CLIP text encoder; The ground terminal is also configured to: The second remote sensing image is encoded based on the CNN image encoder to obtain CNN visual embedding information; The second remote sensing image is encoded based on the CLIP image encoder to obtain CLIP visual embedding information; The semantic text is encoded based on the CLIP text encoder to obtain CLIP text embedding information.
5. The remote sensing image transmission and reconstruction system according to claim 4, characterized in that, The ground terminal is also configured to: Based on the CLIP visual embedding information and the CLIP text embedding information, the semantic feature alignment of the CNN visual embedding information is performed multiple times by the text-guided reconstruction module to obtain the final embedding information.
6. The remote sensing image transmission and reconstruction system according to claim 5, characterized in that, The text-guided reconstruction module includes a semantic guidance module and a feature alignment module; The ground terminal is also configured to: Based on the semantic guidance module, the cosine similarity between the global semantic embedding information and the local semantic embedding information in the CLIP visual embedding information is calculated, and a semantic guidance graph is generated; The CLIP text embedding information is mapped using a multilayer perceptron to generate a channel attention weight vector; Based on the preset learnable alignment parameters, the channel attention weight vector, and the semantic guidance map, the semantic feature alignment of the CNN visual embedding information is performed on multiple iterations by the feature alignment module to obtain the final embedding information.
7. The remote sensing image transmission and reconstruction system according to claim 3, characterized in that, The image decoding module includes a first decoding module and a second decoding module; The ground terminal is also configured to: Based on the first decoding module, the final embedded information is convolutionally processed and upsampled to obtain the initial decoded image; The third remote sensing image is obtained by performing a second decoding on the initial decoded image based on the second decoding module.
8. A semantic-based remote sensing image transmission and reconstruction method, characterized in that, An application to a semantic-based remote sensing image transmission and reconstruction system, including satellite and ground-based terminals; the method includes: The satellite terminal processes the acquired first remote sensing image based on a first processing model to obtain a second remote sensing image and semantic text describing the first remote sensing image, and then sends the second remote sensing image and the semantic text to the ground terminal; wherein, the resolution of the second remote sensing image is lower than the resolution of the first remote sensing image; The ground terminal performs multimodal fusion reconstruction on the received second remote sensing image and the semantic text based on the second processing model to obtain a third remote sensing image.
9. An electronic device, characterized in that, Includes memory and processor, of which: The processor is used to execute computer programs stored in the memory; When the processor executes the computer program, it implements the steps in the remote sensing image transmission and reconstruction method of claim 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps in the remote sensing image transmission and reconstruction method of claim 8.