Method, device and storage medium for image enhancement
Through intra-modal regression and inter-modal generation models, the sparsity problem of remote sensing image data is solved, high-quality enhanced images are generated to meet user query needs, and the utilization and quality of image data are improved.
Patent Information
- Application Number
- CN202310926679.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-13
- Filing Date
- 2023-07-26
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2043-07-26
AI Technical Summary
Due to the sparsity of satellite or aircraft remote sensing image data in terms of modality, frequency band, location and time, images that meet the query requirements may not be available, and existing technologies find it difficult to effectively generate high-quality enhanced images.
It adopts intra-modal regression and inter-modal generation models, adaptively defines the model style, and uses multi-layer perceptron and spatially adaptive normalized conditional generative adversarial network technology to generate high-quality enhanced images.
Even when image data is sparse, it can generate high-quality enhanced images to meet user query requirements and improve the utilization and quality of image data.
Smart Images

Figure CN116862808B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to U.S. patent application No. 18 / 109,240, filed on February 13, 2023, the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present disclosure relates to image processing, and more particularly, to a method, apparatus, and storage medium for image enhancement. Background Art
[0004] Remote sensing, such as satellite remote sensing or aircraft remote sensing, provides unique capabilities for multimodal, multispectral, and multitemporal image acquisition for various Earth observation tasks, including land cover classification, crop yield prediction, natural disaster monitoring, climate change analysis, etc. To this end, satellite or aircraft remote sensors can collect global-scale remote sensing images with different modalities (e.g., electro-optical (EO) and synthetic aperture radar (SAR)), spectral bands (e.g., visible and near-infrared bands), spatial resolutions (e.g., 10 m, 30 m, etc.), and revisit periods (e.g., 7 days, 30 days, etc.). Summary of the Invention
[0005] According to one aspect of the present disclosure, a method for image enhancement is provided. The method may include receiving, by at least one processor, a query for an enhanced image of a queried area obtained by a remote sensor at a queried time. The method may include obtaining, by the at least one processor, an image dataset associated with the queried area from the remote sensor. In response to the image dataset including a plurality of image data associated with the queried area obtained by the remote sensor, the method may include generating, by the at least one processor, an enhanced image of the queried area using an intra-model regression model.
[0006] According to another aspect of the present disclosure, a device for image enhancement is provided. The device may include at least one processor. The device may include a memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform a query for receiving an enhanced image of a queried area obtained by a remote sensor at a queried time. The device may include a memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform the following operations: obtain an image dataset associated with the queried area from the remote sensor. The device may include a memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform the following operations: in response to an image dataset including a plurality of image data associated with the queried area obtained by the remote sensor, use an intra-model regression model to generate an enhanced image of the queried area.
[0007] According to another aspect of the present disclosure, a non-transitory computer-readable medium storing instructions for image enhancement is provided. When executed by at least one processor, the instructions may cause the at least one processor to receive a query for an enhanced image of a queried area obtained by a remote sensor at a queried time. When executed by at least one processor, the instructions may cause the at least one processor to obtain an image dataset associated with the queried area from the remote sensor. When executed by at least one processor, the instructions may cause the at least one processor to perform the following operations: in response to an image dataset including multiple image data associated with the queried area obtained by the remote sensor, using an intra-modal regression model to generate an enhanced image of the queried area. When executed by at least one processor, the instructions may cause the at least one processor to perform the following operations: in response to an image dataset including multiple image data associated with the queried area obtained by different remote sensors, using an inter-modal generative model to generate an enhanced image of the queried area.
[0008] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention, as claimed. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The accompanying drawings, which are incorporated herein and form a part of the specification, illustrate implementations of the disclosure and, together with the description, serve to further explain the disclosure and enable one skilled in the relevant art to make and use the disclosure.
[0010] Figure 1 A block diagram illustrating an exemplary operating environment of a system configured to perform super-resolution image processing in remote sensing according to the present disclosure is shown.
[0011] Figure 2 A schematic diagram illustrating an exemplary structure of the multimodal model selection technology disclosed in the present invention is shown.
[0012] Figure 3 A schematic diagram illustrating an exemplary structure of an in-mold model for image enhancement disclosed in the present invention is shown.
[0013] Figure 4 A schematic diagram illustrating an exemplary structure of an inter-modality model for image enhancement disclosed in the present invention is shown.
[0014] Figure 5 A flowchart of an exemplary method of image enhancement disclosed in the present invention is illustrated.
[0015] Implementations of the present disclosure will be described with reference to the accompanying drawings. DETAILED DESCRIPTION
[0016] Reference will now be made in detail to exemplary embodiments, examples of which are illustrated in the accompanying drawings. Wherever possible, the same reference numerals will be used throughout the drawings to refer to the same or like parts.
[0017] Remote sensing can include the process of detecting and monitoring the physical properties of an area by measuring its reflected and emitted radiation from a distance (e.g., from a satellite or aircraft). Cameras can be mounted on satellites or aircraft to collect remotely sensed images (also known as remote sensing images).
[0018] As mentioned in the background technology section above, remote sensing (satellite or aircraft remote sensing) provides a unique capability for acquiring multi-modal, multi-spectral, and multi-temporal imagery for a variety of Earth observation tasks, including land cover classification, crop yield prediction, natural disaster monitoring, and climate change analysis. To this end, satellites can collect global-scale remote sensing imagery with different modalities (e.g., EO and SAR), spectral bands (e.g., visible and near-infrared bands), spatial resolutions (e.g., 10 meters, 30 meters, etc.), and revisit periods (e.g., 7 days, 30 days, etc.).
[0019] However, such high-dimensional remote image data may exhibit sparsity in one or more of modalities, frequency bands, locations, and times due to various limitations arising from satellite or aircraft revisit periods, cloud cover, sensor modality gaps, etc. As a result, images satisfying the requested or queried modality, spectral band, geographic location, and time may not be available.
[0020] To overcome these and other challenges, the present disclosure provides an exemplary multimodal regression and generative image enhancement technique that generates enhanced images for the modality, spectral band, geographic location, and time of the query. Due to the uncertainty of the input image dataset and the output image, the present disclosure adaptively defines the model style for each query based on the input image dataset.
[0021] For example, when the input image dataset includes temporally adjacent image data of the same location captured using the same sensor or different sensors of the same modality, intra-modal regression can be used to generate enhanced images. Inter-modal regression can be used to generate enhanced images using the input data [[satellite i ,frequency band p ,position i ,date p ]…,[satellite i ,frequency band k ,position i ,date k ]]Generate enhanced images or output queries (e.g. [satellite i ,frequency band i ,position i ,date i ]).
[0022] However, in some other embodiments, the input image dataset may not include temporally adjacent image data of the same location captured using the same sensor or different sensors of the same modality. Here, the input image dataset may include image bands acquired using multiple modalities on different dates, while the output image (e.g., enhanced image) is a single image band on a single date, and the model style is defined by the combination of the input X((m1, d1)..., (m i ,d i ),...、(m B ,d B )) and output y(m j ,d j ), where (m i ,d i ) represent the mode m i (e.g., Sentinel2_B2) and date d i (e.g., 2021_12_10). In the case where the input dataset lacks temporally adjacent data from the same satellite and location, inter-modal images from the same location can be used to generate enhanced images using the inter-modal generative model. For example, an enhanced SAR image associated with satellite signal 1-1 can be generated using an EO image that was captured using EO image data captured using satellite signal 1-2.
[0023] Corresponding to each adaptively defined model style, the exemplary image processor retrieves training and test data for model training and enhanced image generation for the query. In some embodiments, the input and output modalities in the training data match the input and output modalities of the test query, and the input and output dates are within T days of the input and output dates of the test query. This strategy can reduce the number of model styles, which limits the training load of each model style. Some queries can share the same model style and be generated using the same model. The exemplary in-model model component described herein can implement a neural network or gradient boosting on a decision tree for in-model test query regression. In a non-limiting example, this can be implemented using a multilayer perceptron (MLP) model. For each model style, all available spectral bands can be used as input features, and the output query band can be used as the regression target.
[0024] Due to the substantial modality gap between SAR and EO images, an exemplary inter-modality generation component may include a generative model based on the SPatial_ADaptivE (DE) Normalization (SPADE) technique, a variant of the Conditional Generative Adversarial Network (cGAN) with spatially adaptive normalization. SPADE uses a learnable spatially adaptive transformation to modulate the normalized activations by utilizing the input semantic segmentation layout. Unlike SPADE, which uses a downsampled image segmentation map as a conditional input, the present disclosure uses a downsampled image segmentation map from the modality mi The original multispectral image is used as a conditional input to generate the image for another modality m j The following is combined with the enhanced image Figure 1-5 Additional details of exemplary multimodal regression and image enhancement techniques are provided.
[0025] Figure 1 An exemplary operating environment 100 of a system 101 configured to perform image enhancement processing in remote sensing according to the present disclosure is illustrated. Operating environment 100 may include system 101, one or more data sources 108A, ..., 108N (also individually or collectively referred to as data sources 108), user device 112, and any other suitable components. Components of operating environment 100 may be coupled to each other via network 110.
[0026] In some embodiments, system 101 may be included on a cloud computing device. The computing device may be, for example, a server, a desktop computer, a laptop computer, a tablet computer, or any other suitable electronic device including a processor and memory. In some embodiments, system 101 may include a processor 102, a memory 103, and a storage device 104. It should be understood that system 101 may also include any other suitable components for performing the functions described herein.
[0027] In some embodiments, system 101 may have different components in a single device, such as an integrated circuit (IC) chip, or in separate devices with dedicated functions. For example, the IC may be implemented as an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA). In some embodiments, one or more components of system 101 may be located in a cloud computing environment, or may alternatively be located in a single location or distributed locations. In some embodiments, the components of system 101 may be in an integrated device or distributed in different locations, but communicate with each other via network 110.
[0028] The processor 102 may include any appropriate type of general-purpose or special-purpose microprocessor, digital signal processor, microcontroller, graphics processing unit (GPU). The processor 102 may include one or more hardware units (e.g., one or more portions of an integrated circuit) that are designed to be used in conjunction with other components or to execute part of a program. The program may be stored on a computer-readable medium and, when executed by the processor 102, it may perform one or more functions. The processor 102 may be configured as a separate processor component dedicated to image processing. Alternatively, the processor 102 may be configured as a shared processor module for performing other functions unrelated to image processing.
[0029] The processor 102 may include several components, such as a query input module 105, a model selection module 106, an intra-modal model module 107 that maintains an intra-modal regression model, and an inter-modal model module 114 that maintains an inter-modal generative model. Figure 1 The query input module 105 , model selection module 106 , intra-modal model module 107 , and inter-modal model module 114 are shown within a single processor 102 , but may alternatively be implemented on different processors that are close to or remote from each other.
[0030] The query input module 105, the model selection module 106, the intra-modal model module 107, and the inter-modal model module 114 (and any corresponding sub-modules or sub-units) can be hardware units (e.g., parts of an integrated circuit) of the processor 102, which are designed to be used with other components or software units implemented by the processor 102 by executing at least a portion of a program. The program can be stored on a computer-readable medium such as the memory 103 or the storage device 104, and when executed by the processor 102, it can perform one or more functions.
[0031] Memory 103 and storage 104 may include any suitable type of mass storage device provided for storing any type of information that processor 102 may need to operate. For example, memory 103 and storage 104 may be volatile or non-volatile, magnetic, semiconductor-based, tape-based, optical, removable, non-removable, or other types of storage devices or tangible (i.e., non-transitory) computer-readable media, including but not limited to ROM, flash memory, dynamic RAM, and static RAM. Memory 103 and / or storage 104 may be configured to store one or more computer programs that can be executed by processor 102 to perform the functions disclosed herein. For example, memory 103 and / or storage 104 may be configured to store programs that can be executed by processor 102 to perform super-resolution image processing. Memory 103 and / or storage 104 may also be configured to store information and data used by processor 102.
[0032] Each data source 108 may include one or more storage devices configured to store remote sensing images. The remote sensing images may be captured by image sensors mounted on satellites, manned or unmanned aerial vehicles such as unmanned aerial vehicles (UAVs), hot air balloons, and the like. For example, a first data source 108 may be a National Agriculture Imagery Program (NAIP) data source and may store remote sensing images having a first source resolution (e.g., 0.6 meters). Remote sensing images from a NAIP data source may be referred to as NAIP images. A second data source 108 may be a Sentinel-2 data source and may store remote sensing images having a second source resolution (e.g., 10 meters). Remote sensing images from a Sentinel-2 data source may be referred to as Sentinel-2 images. Sentinel-2 images and NAIP images are free satellite remote sensing images. Although Figure 1 System 101 and data source 108 are shown separate from each other, but in some embodiments, data source 108 and system 101 may be integrated into a single device.Each remotely sensed image may include image data along with the location, date / time of the image, and modality at which the image was captured.
[0033] The user device 112 may be a computing device including a processor and memory. For example, the user device 112 may be a desktop computer, a laptop computer, a tablet computer, a smart phone, a game console, a television (TV), a music player, a wearable electronic device such as a smart watch, an Internet appliance, a smart vehicle, or any other suitable electronic device having a processor and memory. Figure 1 The system 101 and the user device 112 are illustrated as being separate from one another, but in some embodiments, the user device 112 and the system 101 may be integrated into a single device.
[0034] In some embodiments, a user may be operating on a user device 112 and may input a user query through the user device 112. The user device 112 may send the user query to the system 101 via the network 110. The user query may include one or more parameters for requesting remote sensing images. The one or more parameters may include, for example, one or more of a location (or a geographic area of interest), a specified time (or a specified time window), and a modality (e.g., a specific optical sensor). The modality may include an indication of a specific satellite (e.g., Sentinel-1, Sentinel-2, etc.) or a specific aircraft (e.g., Aircraft-1, Aircraft-2, etc.) or any other aerial optical sensor, without departing from the scope of the present disclosure. The location may be a geographic location or a ground location on the earth. For example, the location may include longitude and latitude, an address (e.g., street, city, state, country, etc.), a place of interest, etc. The optical image captured by the requested modality or by other available modalities may depict a scene or landscape at the location as seen from above.
[0035] Still refer to Figure 1 , when the query input module 105 receives a query for an enhanced image of the queried area obtained by the remote sensor at the queried time, it can send the query to the mode selection module 106. The model selection module 106 can access at least one data source of the database 108 to obtain an image dataset associated with the queried area from the queried area and the queried time. The image dataset may include images of the queried area captured via one or more modalities within a time window associated with the queried time. For example, the image may be captured at the queried time or within a time window before or after the queried time. If there are one or more images captured using the queried modality (e.g., the optical sensor of interest) in the dataset, the model selection module 106 can activate and send the dataset including the images captured using the optical sensor of interest to the intra-modal model module 107. Otherwise, if the dataset does not include images of the area of interest captured using the queried modality, the model selection module 106 can activate the dataset and send it to the inter-modal model module 114. The following is combined with Figure 2 Additional details are provided regarding the operations performed by the model selection module 106. When a dataset obtained from the database 108 includes images with a threshold number of image occlusions, such as, for example, aerial images of a city covered in fog, these images may be omitted from the dataset. Various techniques may be used to implement the identification of occluded images.
[0036] Figure 2 The diagram shows an exemplary structure 200 of the multimodal model selection technology disclosed in the present invention. Figure 1 Detailed description of the various components to describe Figure 2 The operations described in .
[0037] Reference Figure 2 , the query input module 105 may receive (at 201) a query for images associated with, for example, a modality (e.g., satellite), a spectral band, a location, and a date. Information associated with the query may be sent to the model selection module 106, which obtains a dataset related to the queried location from the database 108. The model selection module 106 may determine (at 203) whether the dataset includes one or more images of the queried location captured using the queried modality. If there are images of the queried location captured using the queried modality within a certain time window before and / or after the queried date, the model selection module 106 may select (at 205) an intra-modal regression model. Figure 3 Additional details of the operations performed by the intra-modality regression model are described. Otherwise, if the dataset for the queried location does not include images captured using the queried modality, the model selection module 106 may select (at 207) an inter-modality generative model. Figure 4 Describes additional details of the intermodal generation model.
[0038] Figure 3 The schematic diagram of the exemplary structure 300 of the in-mold model for image enhancement disclosed in the present invention is shown. Figure 1 The in-mold model module 107 is described in detail to describe Figure 3 The operations described in .
[0039] refer to Figure 3 In-model model module 107 may receive (at 301) a query image of a location captured at a specific time using a specific modality from model selection module 106. An image dataset including images of the queried location captured at various times other than the requested time using the specific modality may be received (at 303) from model selection module 106. In-model model module 107 may generate (at 305) a regression model pattern based on the queried time and the timestamps of the images in the dataset. In some embodiments, the regression model pattern may be considered a list of timestamps.
[0040] Using the regression model style, the in-model model module 107 can select (at 307) a training dataset using the regression model style. The training dataset can include training data for each timestamp or a subset of timestamps in the regression model style. The training dataset can be pre-generated using various techniques and is related to one or more of the queried modality, location, time, spectral band, etc. Using the training dataset, the in-model model module 107 can generate (at 309) a regression model. An example of a regression model can be a boosted tree model (e.g., CatBoost). The received image dataset (at 303) can then be input into the generated regression model to generate an enhanced image 311 related to the queried image. In this way, even when the query image for the location is not available at the queried time, other images captured by the same satellite at different times can be used to generate an enhanced image for the user.
[0041] Figure 4 The schematic diagram of the exemplary structure 400 of the inter-modality model for image enhancement disclosed in the present invention is shown. Figure 1 The inter-modal model module 114 is described in detail to describe Figure 4 The operations described in .
[0042] refer to Figure 4 , the inter-modality model module 114 may receive a dataset 402 comprising various channels (e.g., feature maps) associated with images captured using one or more modalities other than the queried modality. Figure 4In the non-limiting example shown, nine channels for each image in the dataset will be input into the inter-modal generative model; however, more or fewer than nine channels may be associated with each image depending on which sensor captured the images. Figure 4 The inter-modality generative model depicted in FIG may include, for example, a first convolutional layer 404a, one or more residual blocks 406, such as a generative adversarial network (GAN), a spatially adaptive normalization (SPADE) residual block, and the like. The output from the set of residual blocks 406 may be a plurality of different channels and / or feature maps associated with the queried modality generated from the input dataset, but associated with the queried modality. Before outputting the one or more channels and / or feature maps associated with the queried modality, the output of the set of residual blocks 406 may be passed through a second convolutional layer 404b, thereby providing the user with an enhanced image 408 of the queried location, time, and modality even when an image associated with the location / time captured by the queried modality is unavailable. Although two channels are described as output from the second convolutional layer 404b, more or fewer channels may be output depending on the queried modality.
[0043] Figure 5 A flow chart of an exemplary method 500 of image enhancement disclosed in the present invention is shown.
[0044] At 502, the query input model 105 may receive a request for an image of an area acquired at a time by a remote sensor. Figure 1 , the query input module 105 may receive a query / request for images of a location captured by a particular modality (eg, satellite, aircraft, etc.). The request for the image may be received from a user device 112 via the network 110 .
[0045] At 504, the model selection module 106 can obtain an image dataset associated with the region. Figure 1 When the query input module 105 receives a query, it may send the query to the mode selection module 106. The mode selection module 106 may access at least one data source of the database 108 to obtain an image dataset associated with the queried area from the queried area and the queried time. The image dataset may include images of the queried area captured via one or more modalities within a time window associated with the queried time. For example, the images may be captured at the queried time or within a time window before or after the queried time.
[0046] At step 506, the model selection module 106 may determine whether the dataset includes one or more images captured by the queried satellite within a certain time window (also referred to as "temporally adjacent images"). In response to the image dataset including one or more temporally adjacent images captured by the satellite indicated in the request, the operation may move to 508; otherwise, if no temporally adjacent images captured by the satellite in the request are available, the operation moves to 510. For example, referring to Figure 1 If one or more images captured using the queried modality (e.g., the optical sensor of interest) are present in the dataset, the model selection module 106 may activate and send the dataset including the images captured using the optical sensor of interest to the intra-modality model module 107. Otherwise, if the dataset does not include images of the region of interest captured using the queried modality, the model selection module 106 may activate and send the dataset to the inter-modality model module 114.
[0047] At 508, the intra-model model module 107 may generate an enhanced image of the region using the intra-model regression model. Figure 3 The in-model model module 107 may receive (at 301) a query image of a location captured at a specific time using a specific modality from the model selection module 106. An image dataset including images of the queried location captured at various times other than the requested time using the specific modality may be received (at 303) from the model selection module 106. The in-model model module 107 may generate (at 305) a regression model pattern based on the queried time and the timestamps of the images in the dataset. In some embodiments, the regression model pattern may be considered a list of timestamps. Using the regression model pattern, the in-model model module 107 may select (at 307) a training dataset using the regression model pattern. The training dataset may include training data for each timestamp or a subset of the timestamps in the regression model pattern. The training dataset may be pre-generated using various techniques and may be associated with one or more of the queried modality, location, time, spectral band, etc. Using the training dataset, the in-model model module 107 may generate (at 309) a regression model. An example of a regression model may be a boosted tree model (e.g., CatBoost). The image dataset received (at 303) can then be input into the generated regression model to generate an enhanced image related to the queried image 408. In this way, even when the query image for the location is not available at the queried time, an enhanced image can be generated for the user using other images captured by the same satellite at different times.
[0048] At 510, the inter-modality model module 114 may use the inter-modality generative model to generate an enhanced image of the region. Figure 4, the inter-modality model module 114 may receive a dataset 402 of various channels (eg, feature maps) associated with images captured by one or more modalities other than the image data set. Figure 4 In the non-limiting example shown, nine channels for each image in the dataset will be input into the inter-modal generative model; however, more or fewer than nine channels may be associated with each image depending on which sensor captured the images. Figure 4 The inter-modality generative model depicted in FIG may include, for example, a first convolutional layer 404a, one or more residual blocks 406, such as a generative adversarial network (GAN), a spatially adaptive normalization (SPADE) residual block, and the like. The output from the set of residual blocks 406 may be a plurality of different channels and / or feature maps associated with the queried modality generated from the input dataset, but associated with the queried modality. Before outputting the one or more channels and / or feature maps associated with the queried modality, the output of the set of residual blocks 406 may be passed through a second convolutional layer 404b, thereby providing the user with an enhanced image 408 of the queried location, time, and modality even when an image associated with the location / time captured by the queried modality is unavailable. Although two channels are described as output from the second convolutional layer 404b, more or fewer channels may be output depending on the queried modality.
[0049] Another aspect of the present disclosure relates to a non-transitory computer-readable medium storing instructions that, when executed, cause one or more processors to perform the method described above. The computer-readable medium may include volatile or non-volatile, magnetic, semiconductor, tape, optical, removable, non-removable, or other types of computer-readable media or computer-readable storage devices. For example, as disclosed, the computer-readable medium may be a storage device or memory module having computer instructions stored thereon. In some embodiments, the computer-readable medium may be a disk or flash drive having computer instructions stored thereon.
[0050] According to one aspect of the present disclosure, a method for image enhancement is provided. The method may include receiving, by at least one processor, a query for an enhanced image of a queried area obtained by a remote sensor at a queried time. The method may include obtaining, by the at least one processor, an image dataset associated with the queried area from the remote sensor. In response to the image dataset including a plurality of image data associated with the queried area obtained by the remote sensor, the method may include generating, by the at least one processor, an enhanced image of the queried area using an intra-model regression model.
[0051] In some embodiments, generating, by the at least one processor, an enhanced image of the queried area using the intra-modal regression model may include generating a model style based on a first timestamp of the queried time and a second timestamp of each image associated with the queried area.
[0052] In some embodiments, the method may include accessing, by the at least one processor, a training database comprising a plurality of training data images of different areas acquired using different remote sensors. In some embodiments, the method may include performing, by the at least one processor, a first training data search associated with a first timestamp of the time of the query and the remote sensor and the queried area, and a second training data search associated with a second timestamp, for each image in the image dataset associated with the queried area and the remote sensor. In some embodiments, the method may include selecting, by the at least one processor, a training dataset based on the first training data search and the second training data search associated with the queried area.
[0053] In some embodiments, the method may include identifying, by at least one processor, whether any image in the training dataset includes one or more occluded images of the queried region. In some embodiments, the method may include discarding, by at least one processor, the one or more occluded images of the queried region from the training dataset.
[0054] In some embodiments, the method may include generating, by at least one processor, an intra-modal regression model based on a first training data search and a second training data search associated with the queried area.
[0055] In some embodiments, the method may include inputting, by the at least one processor, an image dataset associated with the queried area from the remote sensor into the intra-model regression model. In some embodiments, the method may include receiving, by the at least one processor, an enhanced image of the queried area associated with the queried time and associated with the remote sensor based on an output of the intra-model regression model.
[0056] In some embodiments, responsive to an image dataset comprising a plurality of image data associated with a queried area obtained by different remote sensors, the method may include generating, by the at least one processor, an enhanced image of the queried area using an intermodal generative model.
[0057] In some embodiments, at least one processor accesses a training database comprising a plurality of training data images of various areas acquired using different remote sensors. In some embodiments, the method may include performing, by the at least one processor, a training data search for a training data set associated with a queried area captured by any different remote sensor at any time. In some embodiments, the method may include selecting, by the at least one processor, a training data set comprising training data for the queried area captured by different remote sensors at different times. In some embodiments, the training data set may not be associated with the modality of the queried remote sensor.
[0058] In some embodiments, the method may include inputting, by at least one processor, a training dataset into an inter-modal generative model. In some embodiments, the inter-modal generative model may include one or more convolutional layers or residual blocks. In some embodiments, the method may include receiving, by the at least one processor, an enhanced image of the queried area associated with the queried time and associated with the modality of the remote sensor from the query as an output of the inter-modal generative model.
[0059] In some embodiments, the training dataset may be associated with the first modality or the second modality. In some embodiments, in response to the inter-modality generative model receiving the training dataset associated with the first modality, the method may include generating, by at least one processor, an enhanced image of the second modality at the queried time and region as an output of the inter-modality generative model. In some embodiments, in response to the inter-modality generative model receiving the training dataset associated with the second modality, the method may include generating, by at least one processor, an enhanced image of the first modality at the queried time and region as an output of the inter-modality generative model.
[0060] According to another aspect of the present disclosure, a device for image enhancement is provided. The device may include at least one processor. The device may include a memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform a query for receiving an enhanced image of a queried area obtained by a remote sensor at a queried time. The device may include a memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform the following operations: obtain an image dataset associated with the queried area from the remote sensor. The device may include a memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform the following operations: in response to an image dataset including a plurality of image data associated with the queried area obtained by the remote sensor, use an intra-model regression model to generate an enhanced image of the queried area.
[0061] In some embodiments, the memory stores instructions that, when executed by the at least one processor, cause the at least one processor to perform generating an enhanced image of the queried area using the intra-model regression model by generating a model style based on a first timestamp of the queried time and a second timestamp of each image associated with the queried area.
[0062] In some embodiments, the memory stores instructions that, when executed by the at least one processor, can further cause the at least one processor to access a training database containing a plurality of training data images of various areas acquired using different remote sensors. In some embodiments, the memory stores instructions that, when executed by the at least one processor, can further cause the at least one processor to: for each image in the image dataset associated with the queried area and the remote sensor, perform a first training data search associated with the queried time and a first timestamp of the remote sensor and the queried area, and a second training data search associated with the second timestamp. In some embodiments, the memory stores instructions that, when executed by the at least one processor, can further cause the at least one processor to select a training dataset based on the first training data search and the second training data search associated with the queried area.
[0063] In some embodiments, the memory stores instructions that, when executed by the at least one processor, may further cause the at least one processor to: identify whether any image in the training dataset includes one or more occluded images of the queried region. In some embodiments, the memory stores instructions that, when executed by the at least one processor, may further cause the at least one processor to discard the one or more occluded images of the queried region from the training dataset.
[0064] In some embodiments, the memory stores instructions that, when executed by the at least one processor, may further cause the at least one processor to perform generating the intra-model regression model based on a first training data search and a second training data search associated with the queried area.
[0065] In some embodiments, the memory stores instructions that, when executed by the at least one processor, may further cause the at least one processor to input an image dataset associated with the queried area from the remote sensor into the intra-model regression model. In some embodiments, the memory stores instructions that, when executed by the at least one processor, may further cause the at least one processor to receive an enhanced image of the queried area associated with the queried time and associated with the remote sensor based on an output of the intra-model regression model.
[0066] In some embodiments, the memory stores instructions that, when executed by the at least one processor, may further cause the at least one processor to generate an enhanced image of the queried area using an intermodal generation model in response to an image dataset comprising a plurality of image data associated with the queried area acquired by different remote sensors. In some embodiments, the memory stores instructions that, when executed by the at least one processor, may further cause the at least one processor to access a training database comprising a plurality of training data images of various areas acquired using different remote sensors. In some embodiments, the memory stores instructions that, when executed by the at least one processor, may further cause the at least one processor to perform a training data search of training datasets associated with the queried area captured by any of the different remote sensors at any time. In some embodiments, the memory stores instructions that, when executed by the at least one processor, may further cause the at least one processor to select a training dataset comprising training data for the queried area captured by different remote sensors at different times. In some embodiments, the training dataset may not be associated with the modality of the queried remote sensor.
[0067] In some embodiments, the memory stores instructions that, when executed by the at least one processor, may further cause the at least one processor to input the training dataset into the inter-modality generation model, the inter-modality generation model comprising one or more convolutional layers or residual blocks. In some embodiments, the memory stores instructions that, when executed by the at least one processor, may further cause the at least one processor to receive, from the query, an enhanced image of the queried area associated with the queried time and associated with the modality of the remote sensor as an output of the inter-modality generation model.
[0068] In some embodiments, the training dataset is associated with the first modality or the second modality. In some embodiments, the memory stores instructions that, when executed by the at least one processor, can further cause the at least one processor to generate an enhanced image of the second modality at the queried time and region as an output of the inter-modality generation model in response to the inter-modality generation model receiving the training dataset associated with the first modality. The memory stores instructions that, when executed by the at least one processor, can further cause the at least one processor to generate an enhanced image of the first modality at the queried time and region as an output of the inter-modality generation model in response to the inter-modality generation model receiving the training dataset associated with the second modality.
[0069] According to another aspect of the present disclosure, a non-transitory computer-readable medium storing instructions for image enhancement generation is provided. When executed by at least one processor, the instructions may cause the at least one processor to perform a query for receiving an enhanced image of a queried area obtained by a remote sensor at a queried time. When executed by at least one processor, the instructions may cause the at least one processor to perform obtaining an image dataset associated with the queried area from the remote sensor. When executed by at least one processor, the instructions may cause the at least one processor to perform the following operations: in response to an image dataset including multiple image data associated with the queried area obtained by the remote sensor, using an intra-modal regression model to generate an enhanced image of the queried area. When executed by at least one processor, the instructions may cause the at least one processor to perform the following operations: in response to an image dataset including multiple image data associated with the queried area obtained by different remote sensors, using an inter-modal generative model to generate an enhanced image of the queried area.
[0070] The above description of a specific implementation can be readily modified and / or adapted for various applications. Therefore, such adaptations and modifications are intended to be within the meaning and range of equivalents of the disclosed implementations, based on the teaching and guidance presented herein.
[0071] The breadth and scope of the present disclosure should not be limited by any of the above-described exemplary implementations, but should be defined only in accordance with the following claims and their equivalents.
Claims
1. A method for image enhancement, characterized in that: The method includes: receiving, by at least one processor, a query for an enhanced image of a queried area obtained by a remote sensor at a queried time; obtaining, by the at least one processor, from the remote sensor an image dataset associated with the queried area; and generating, by the at least one processor, the enhanced image of the queried area using an intra-modal regression model in response to the image dataset obtained by the remote sensor and comprising a plurality of image data associated with the queried area; The generating, by the at least one processor, the enhanced image of the queried area using an intra-model regression model comprises: generating a model pattern based on a first timestamp of the queried time and a second timestamp of each image associated with the queried area; The method further includes: accessing, by the at least one processor, a training database comprising a plurality of training data images of respective areas acquired using different remote sensors; performing, by the at least one processor, for each image in the image dataset associated with the queried area and the remote sensor, a first training data search associated with the queried time and a first timestamp of the remote sensor and the queried area, and a second training data search associated with the second timestamp; A training data set is selected, by the at least one processor, based on the first training data search and the second training data search associated with the queried area.
2. The method according to claim 1, characterized in that The method further includes: identifying, by the at least one processor, whether any image in the training dataset includes one or more occluded images of the queried region; and The one or more occluded images of the queried region are discarded from the training dataset by the at least one processor.
3. The method according to claim 1, characterized in that The method further includes: The intra-modal regression model is generated by at least one processor based on the first training data search and the second training data search associated with the queried area.
4. The method according to claim 1, wherein The method further includes: inputting, by the at least one processor, the image dataset from the remote sensor associated with the queried area into the intra-model regression model; and The enhanced image of the queried area associated with the queried time and associated with the remote sensor is received by the at least one processor based on an output of the intra-modal regression model.
5. The method according to claim 1, wherein The method further includes: The enhanced image of the queried area is generated by the at least one processor using an inter-modal generative model in response to the image data set including a plurality of image data associated with the queried area obtained by different remote sensors.
6. The method according to claim 5, characterized in that The method further includes: accessing, by the at least one processor, a training database comprising a plurality of training data images of respective areas acquired using the different remote sensors; performing, by the at least one processor, a training data search on a training data set associated with the queried area captured at any time by any of the different remote sensors; and selecting, by the at least one processor, a training data set comprising training data for the queried area captured by the different remote sensors at different times, The training data set is not associated with a modality of the remote sensor being queried.
7. The method according to claim 6, characterized in that The method further includes: Inputting, by the at least one processor, the training data set into the inter-modal generative model, the inter-modal generative model comprising one or more convolutional layers or residual blocks; and The enhanced image of the queried area associated with the queried time and associated with a modality of the remote sensor is received from the query by the at least one processor as an output of the inter-modality generative model.
8. The method according to claim 7, wherein: The training data set is associated with the first modality or the second modality, In response to the inter-modality generative model receiving the training dataset associated with the first modality, generating, by the at least one processor, the enhanced image of the second modality at the queried time and region as an output of the inter-modality generative model, and In response to the inter-modality generative model receiving the training dataset associated with the second modality, the at least one processor generates the enhanced image of the first modality at the queried time and region as the output of the inter-modality generative model.
9. A device for image enhancement, used to implement the method according to any one of claims 1 to 8, characterized in that: The device includes: at least one processor; a memory storing instructions that, when executed by the at least one processor, cause the at least one processor to: receiving a query for an enhanced image of a queried area obtained by a remote sensor at a queried time; obtaining an image dataset associated with the queried area from the remote sensor; and The enhanced image of the queried area is generated using an intra-modal regression model in response to an image dataset including a plurality of image data associated with the queried area obtained by the remote sensor.
10. A non-transitory computer-readable medium storing program instructions, characterized in that: When the program instructions are executed by at least one processor, the method according to any one of claims 1 to 8 is implemented.