Method for generating high-resolution maps using high-resolution image and low-resolution labels

The deep learning-based method iteratively refines spatial resolution and uses noise-robust loss functions to generate high-accuracy high-resolution building maps from low-resolution inputs, addressing the inefficiencies of manual labeling and existing deep learning models.

WO2026074115A1PCT designated stage Publication Date: 2026-04-09LUXEMBOURG INSTITUTE OF SCIENCE AND TECHNOLOGY (LIST)
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-10-02
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Existing methods struggle to generate high-resolution building maps efficiently, as they require manual labeling of high-resolution images, are costly, and existing deep learning models fail to accurately differentiate building shapes due to complex structures and noise in low-resolution maps, leading to inaccurate results.

Method used

A deep learning-based method iteratively downsamples high-resolution images to intermediary resolutions, trains label prediction models using paired patches, and employs noise-robust loss functions to generate high-resolution building maps, leveraging low-resolution maps for improved accuracy.

Benefits of technology

The method effectively produces high-resolution building maps with improved accuracy and reduced costs by iteratively refining spatial resolution and handling noise, achieving precision scores up to 74% compared to 59% without the quality improvement phase.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025078423_09042026_PF_FP_ABST
    Figure EP2025078423_09042026_PF_FP_ABST
Patent Text Reader

Abstract

A method for generating a high spatial resolution - HR map (ŶHR) of an area, comprising, after acquiring an initial HR image (XHR) and an initial low spatial resolution - LR map (YLR) of the area, the iteration of a loop of: selecting (110) at least one nth intermediary spatial resolution - IR; downsampling (112) the initial HR image to the nth IR to obtain an nth IR image (X̃IR); upsampling (114) an n-1th IR map at the IR to obtain a resized n-1th IR map, the 0th IR map being the initial LR map; extracting (130) a plurality of pairs of patches, from the nth IR image and the resized n-1th IR map; training (140) a nth IR label prediction model with a sample of pairs of patches; and, using (150) the trained nth IR label prediction model in inference on the nth IR image to estimate an nth IR map, the loop being iterated until the Nth IR is equal to the resolution of the initial HR image, the HR map corresponding the Nth IR map.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] TITLE: METHOD FOR GENERATING HIGH-RESOLUTION MAPS USING HIGH- RESOLUTION IMAGE AND LOW-RESOLUTION LABELS

[0002] The present invention relates to the technical field of methods for generating high- resolution maps of a geographical area of interest, particularly high-resolution building maps.

[0003] In the following, the term image designates a bi-dimensional data matrix whose pixels are associated with one or several channels (three channels in the case of RGB optical images or one channel in the case of SAR images), each channel being a numerical variable, for example coded on 8 bits.

[0004] The term map designates a bi-dimensional data matrix whose pixels are associated with one channel, this channel being usually a binary label. For example, for a built-up map, value 1 represents a built geographical area and value 0 represents a not-built geographical area. For example, for a building map, value 1 represents an area with a building, and value 0 represents an area with no building.

[0005] An image or a map is geolocated so that each pixel is associated with a geographical area at the surface of the earth, i.e. the ground.

[0006] In the following, HR is used for “high-resolution”, IR for “intermediary resolution”, and LR for “low-resolution” to qualify the spatial resolution of data matrices.

[0007] The pixel size of a data matrix refers to its number of pixels, while the spatial resolution of this data matrix refers to the surface area that a pixel represents on the ground.

[0008] Building maps have a lot of different uses. For example, after a geographical area has been affected by an earthquake, emergency services compare building maps before and after the earthquake to identify the collapsed buildings. This helps to concentrate the efforts on the zones in real need by dispatching teams of rescuers on the priority zones without delay. However, for such a use, the building map must present a high spatial resolution.

[0009] A building map is usually derived from an image of the same geographical area. Such an image is acquired by an observation satellite, by a plane or by an unmanned aerial vehicle. The image may be an optical image acquired by an optical sensor, or a SAR image acquired by a radar system or other type of images acquired with a suitable sensor.

[0010] However, mapping buildings from an image is challenging due to buildings coming in different shapes, luminosity, and / or colors.

[0011] Traditionally, the image is analyzed by an expert, who labels each of its pixels depending on whether it comprises a building (house, tenement, skyscraper, ...) to obtain a building map. This building map has the same size and spatial resolution as the initial image.

[0012] This task is time consuming and costly as it requires an expert to manually annotate every pixel in an image of high resolution. In addition to statistical analyzes, methods based on deep learning / machine learning models have been proposed more recently to automatically detect buildings in optical images.

[0013] However, the training phase of such a model still requires an expert to populate a training database with labeled images in sufficient numbers.

[0014] The task of building a training database is time consuming and challenging even for low-resolution images. It is too difficult for high-resolution images.

[0015] Other proposed methods rely on weak supervision by exploiting either some available high-resolution labels or generated ones by an already trained model.

[0016] However, the quality of high-resolution labels generated by a pre-trained model is not guaranteed and often quite low, given the difficulty for pre-trained models to perform well over unseen scenarios and different building structures.

[0017] Beside, low resolution-LR building maps become more and more easily available, for example from online services, including “CityWatch-baseline”, a service developed by the applicant that makes use of SAR and optical images at low-resolution to automatically generate LR building maps.

[0018] Given the high availability of LR building maps, it would be advantageous to have a method that employs an LR map as a starting point for producing a corresponding high resolution-HR map.

[0019] The article by Zhuohong Li et al. “Breaking the resolution barrier: A low-to-high network for large-scale high-resolution land-cover mapping using low-resolution labels”, ISPRS Journal of Photogrammetry and Remote Sensing 192 (2022) 244-267, and the article by Kolya Malkin et al. “LABEL SUPER-RESOLUTION NETWORKS” conference paper at ICLR 2019 showcase deep learning-based super-resolution methods that produce HR maps using HR imagery and LR maps.

[0020] However, these methods are primarily designed for land cover mapping, which involves identifying and categorizing the physical surface cover of an area, such as vegetation, urban infrastructure, water bodies, or bare soil. Such identification involves larger and more homogeneous areas.

[0021] When these methods are applied for urban infrastructure mapping more specifically, the focus is on differentiating between various human-made structures like buildings, roads, and parking lots, but not on differentiating the shapes of individual buildings.

[0022] In contrast, building mapping focuses more on the shapes of buildings and differentiates them from other structures. Built-up environments are highly heterogeneous, with complex building structures and fine details that are challenging to map precisely. The spectral characteristics of building features, such as roofs, edges, and unusual shapes are more varied and complex compared to natural land cover types. The methods mentioned in these articles do not adequately address these spectral differences, resulting in inaccurate mapping when applied to building. In other words, they do not provide a framework for extracting high-frequency building features from HR imagery.

[0023] Furthermore, the large spatial resolution gap between LR maps and HR imagery results in significant misalignment of scene structures. This is particularly problematic in areas where precise boundaries and small-scale features are ubiquitous, which is typically the case for built areas.

[0024] Finally, LR building maps often contain noise due to the limitations of low-resolution imagery. This includes misclassifications where pixels corresponding to vegetation or roads are incorrectly labeled as buildings, and vice versa. In building mapping, where accuracy is crucial, such noise can significantly degrade the quality of the HR maps produced. The cited methods lack robust techniques to handle and correct this noise effectively.

[0025] The invention therefore aims at overcoming these problems by providing a deep learning-based method dedicated to automatically generating an HR map in particular an HR building map.

[0026] To this end, an aspect of the invention is a method for generating a high spatial resolution - HR map of a geographical area, characterized in that the method comprises: acquiring an initial HR image of the geographical area and an initial low spatial resolution - LR map of the geographical area; and, a phase of iteration of a loop, comprising, for an nthiteration, n being an integer between 1 and N, N being an integer equal to or greater than two: selecting at least one nthintermediary spatial resolution - IR ; downsampling the initial HR image to the nthintermediary resolution to obtain an nthIR image; upsampling an n-1thIR map at the intermediary resolution to obtain a resized n-1thIR map, the 0thIR map being the initial LR map; extracting a plurality of pairs of patches, each pair of patches being made of an image patch extracted from the nthIR image and a corresponding IR map patch extracted from the resized n-1thIR map and storing said plurality of pairs of patches in a nthtraining database; training a nthIR label prediction model with a sample of pairs of patches from the nthtraining database; and, using the trained nthIR label prediction model in inference on the nthIR image to estimate an nthIR map, the loop being iterated until the Nthintermediary resolution is equal to the resolution of the initial HR image, the HR map corresponding the NthIR map.

[0027] Preferably, the loop is iterated two times. Preferably, the training of the nthIR label prediction model starts with initializing a plurality of parameters of the nth IR label prediction model with values of the plurality of parameters of the trained n-1thIR label prediction model.

[0028] Preferably, a super-resolution loss function is used to train the nth IR label prediction model 0„: where the set S contains 70% of the pair of patches p of the nthtraining database; xIRand yIRare the nthIR image patch and nth resized IR map patch of patch p; the IR label prediction model 0„ is parametrized by the plurality of parameters 9n; DKLis a Kullback-Leibler divergence; and £ris an entropy regularisation term that is weighted by a scalar a.

[0029] Preferably, after the downsampling and the upsampling steps, the label generation loop comprises a step for post-processing the resized n-1thIR map.

[0030] Preferably, the initial HR image is an optical image of the geographical area.

[0031] Preferably the initial LR map is a building map of the geographical area, so that the high-resolution - HR map and / or the improved HR map is a building map of the geographical area.

[0032] Another aspect of the invention is a method of quality improvement of an HR map, preferably the HR map obtained with the previous method, to obtain an improved HR map, comprising: resampling an initial high spatial resolution - HR image in order to obtain a low spatial resolution - LR image having the same size as an initial HR image but the spatial resolution of an initial LR map; upsampling an initial LR map in order to obtain a resized LR map with the same size as the initial HR image; extracting a plurality of set of patches, each set of patches being made of: an HR image patch extracted from the initial HR image; a LR image patch extracted from the LR image; an HR map patch extracted from the HR map; and a LR map patch extracted from the resized LR map, each set of patches being stored in a quality improvement database; training a quality improvement model with a sample of sets of patches from the quality improvement database; and, using the trained quality improvement model in inference on the HR image to generate the improved HR map.

[0033] Preferably, the loss function used in the training step of the quality improvement model is of the form:

[0034] Lnr COmp + Pledge ^focal + ^1 + Pledge where Lf0CCdis the focal loss function; is the L1-norm loss function, and Ledgeis a edge loss function, the edge loss function being of the form: where xHRis the HR image patch of a set of patches; xLRis the LR image patch of said set of patches; yHRis a HR map patch predicted by the quality improvement model ' from xHR, yHRis the HR map patch of said set of patches; and yLRis the LR map patch of said set of patches; and where Gtis a gradient along a first direction and Gj is a gradient along a second direction of a data matrix.

[0035] Preferably, the method comprises, before extracting a plurality of set of patches, a step of post-processing the HR map.

[0036] Preferably, the initial HR image is an optical image of the geographical area.

[0037] Preferably the initial LR map is a building map of the geographical area, so that the high-resolution - HR map and / or the improved HR map is a building map of the geographical area.

[0038] Another aspect of the invention is a computer program product allowing, when its instructions are run by a computer system, to realize the previous methods.

[0039] The invention and its advantages will be better understood upon reading the following description of a preferred embodiment, provided solely by way of example, this description being made with reference to the accompanying drawings in which:

[0040] - Figure 1 is a schematic view of an example of a calculator for the implementation of the method according to the invention.

[0041] - Figure 2 is a flowchart of a preferred embodiment of the method according to the invention, highlighting the data flow between a first phase of cascading spatial resolution and a second phase of improving the HR resolution map output by the first phase;

[0042] - Figure 3 is a detailed flowchart of the first phase of the method of figure 2;

[0043] - Figure 4 is a detailed flowchart of the second phase of the method of figure 2; and,

[0044] - figure 5 shows images and maps used / obtained with the method of figure 2.

[0045] According to the invention, a high-resolution building map is generated given one HR image and one LR building map of the area of interest.

[0046] HR imagery provide high-resolution features of the area of interest, while LR building maps provide coarse estimations of where buildings might be located. Combined, they guide the learning process by providing the necessary information to create HR building maps.

[0047] The method according to the invention is computer implemented.

[0048] Figure 1 shows a calculator 20 and of a computer program product 22.

[0049] The calculator 20 is preferably a computer, a computing system, or similar electronic computing devices, adapted to manipulate and / or transform data (represented as physical quantities) stored within the calculator's registers into other data similarly represented as physical quantities stored within the calculator's registers.

[0050] The calculator 20 interacts with the computer program product 22. As illustrated in Figure 1 , the calculator 20 comprises a data processing unit 26, memories 28 and a reader 30 for information media. The calculator 20 also comprises a human machine interface, such as a keyboard 32 and a display 34.

[0051] The computer program product 22 comprises an information medium 36, which is readable by the calculator 20, for example by the reader 30. The readable information medium 36 is suitable for storing electronic instructions.

[0052] By way of example, the information medium 36 is a USB key, a floppy disk or flexible disk (of the English name "Floppy disc"), an optical disk, a CD-ROM, a magneto-optical disk, a ROM memory, a memory RAM, EPROM memory, EEPROM memory, magnetic card or optical card.

[0053] On the information medium 36, is stored a computer program 24 that comprises program instructions.

[0054] The computer program 24, when its instructions are executed by the processor 26, is adapted to entail an implementation of a super-resolution method according to the invention.

[0055] Figure 2 represents schematically a preferred embodiment of the method according to the invention.

[0056] The method 100 begins with step 103, where are acquires an initial HR image XHRof a geographical area of interest, together with an initial LR building map YLRof that same area.

[0057] The aim of the method 100 is to estimate, from these two entries, a final HR building map YHRof said geographical area of interest, but with a higher spatial resolution and size than the initial LR building map YLR, in fact a spatial resolution and a size identical to those of the initial HR image HR-

[0058] The spatial resolution of the initial HR image XHRis for example of the order of 1m2per pixel or less. This initial HR image XHRis also characterized by its high size. It comprises H x W pixels. For example, it comprises 10000 x 10000 pixels. It is for example an optical image of the geographical area captured by an optical camera on board of a satellite. Preferably, it is a RGB image, i.e. each pixel is associated with three channels, respectively Red, Green and Blue. Alternatively, the initial HR image XHRcan also be a multispectral I hyperspectral image or a SAR image.

[0059] The HR image XHRis for example uploaded from an image database to the calculator 20 on which the program 24 implementing the method is run.

[0060] The initial LR building map YLRhas a low spatial resolution, for example of the order of 10 m2per pixel. It is also characterized by its low size. It comprises w x h pixels. For example, it comprises 1000 x 1000 pixels. It comprises one channel, corresponding to a classification label with a few possible values. In the present embodiment, the classification label is a binary variable, with for example 1 corresponding to an area with a building structure (house, building, skyscraper, etc.) and 0 corresponding to an area with no building structure (vegetation, road, car park, etc.).

[0061] The initial LR building map TLRcan be obtained from a building map database, e.g. an online database. The LR building map YLRcan be itself computer generated.

[0062] The method 100 is based on the iteration of a loop to progressively increase the resolution of the map. The loop is iterated N times, with N an integer greater or equal to 2.

[0063] As illustrated in figure 3 with more detailed, in the embodiment here described, the total number N of iterations of the loop is equal to 2.

[0064] The method 100 starts with step 110 of defining a set of N intermediary spatial resolutions. The first intermediary resolution lies between the target resolution (i.e. in the present example the size of the initial HR image XHR) and the initial LR building map YLR. For example, the first intermediary resolution is taken equal to 2000x2000 pixels, which corresponds to an IR of 5 m2per pixel. The Nthintermediary resolution corresponds to the target resolution (i.e. the size of the initial HR image XHR, in the present embodiment). The nthintermediary resolution (for n other than 1 and N) lies between the size of the target resolution and the initial LR building map YLR. In other words, the set of intermediary spatial resolutions is ordered between a first intermediary resolution and a final intermediary resolution (equal to the target resolution).

[0065] Then, during the first iteration of the loop, in step 112, the initial HR image XHRis downscaled to that intermediary size to obtain an IR image XIR. According to the numerical example, in this operation the size of the initial image is reduced by a downsampling factor of 1 / 5 x 1 / 5 (i.e. the ratio between the size of image XIRand the size of image XHR).

[0066] For example, the values of the pixels of the initial HR image that are merged in the same pixel of the IR image are averaged to obtain the value of that same pixel.

[0067] With such a downsampling operation, the spatial resolution is also reduced by a factor of 1 / 5 x 1 / 5. In other words, the spatial resolution of the IR image XIRis now 5m2per pixel. Thus, the averaging process leads to a loss in spatial resolution.

[0068] In step 114, the initial LR building map YLRis upsampled to the intermediary size to obtain a resized LR building map YLR.

[0069] According to our numerical example, the size of the map is increased by an upsampling factor of 2 x 2 (i.e. the ratio of the resized map YLRto that of the initial map YLR)

[0070] For example, the values of the pixels of the resized LR building map that are associated to the same pixel of the initial LR building map are equal to the value of that same pixel. With such an upsampling operation, despite the increase in the number of pixels, the spatial resolution is unchanged. The resized LR building map YLRhas the same spatial resolution as the initial LR building map YLR.

[0071] In step 120, one or several post processing treatments are advantageously applied on the resized LR building map YLR.

[0072] Using, for example, the excess of green index - ExGi, pixels that likely correspond to vegetation are identified in the LR image. They are then automatically removed from the resized LR map. The excess of green index can for example be defined as : ExGi = 2.8*G - R - B, where “R”, “G”, and “B” refer to the red, green, and blue channel pixel values, respectively. Then a thresholding operation is applied to the ExGi map to identify vegetation pixels and remove them.

[0073] In step 130, in view of the training of a first model 0 for label prediction, pairs of patches are extracted simultaneously from the IR image XIRand the resized IR building map ^LR- pair of patches is composed of an image patch and a building map patch that are co-registered, i.e. they represent the same geographical area.

[0074] A patch is preferably a square, for example of 20 x 20 pixels.

[0075] For example, 10000 pairs of patches are extracted and stored in a first training database.

[0076] The extraction is preferably a random process (i.e. the central pixel is randomly selected among the pixels of the IR matrix), even if other partitioning techniques can be used without impacting the results.

[0077] Step 140 corresponds to the training of the first model, 0}(defined by a plurality of parameters 0 , with samples extracted from the first training database.

[0078] Once the first model 0}has been trained, it is used in inference in step 150, to compute, from the IR optical image XIR, an IR building map YIR.

[0079] The IR building map YIRhas the same size and spatial resolution as the IR image XIR. In other words, some features from the IR image XIRhave been used to improve the spatial resolution of the resized LR building map to get the IR building map.

[0080] The following steps of method 100 correspond to the second iteration of the loop.

[0081] Since in the present embodiment the loop is iterated only twice, the second intermediary resolution is taken equal to the target resolution, i.e. the high spatial resolution of the initial image HR-

[0082] Since the initial image XHRis already at the target resolution, in this second and final iteration of the loop, the downscaling step 212 corresponds to the mere identity. In step 214, the IR building map YIRis upscaled to obtain a resized IR building map

[0083] The upsampling factor used in this second and last iteration of the loop is the inverse of the downsampling factor used in step 112. The size of the resized IR building map YIRthus obtained is equal to the size of the initial HR image XHR.

[0084] However, its spatial resolution is unchanged by this transformation, so that the resized IR building map YIRis still characterized by features with the intermediary spatial resolution.

[0085] Advantageously, in step 220, one or several post-processing treatments are applied on of the resized IR map YIR, and / or the initial HR image XHR. These treatments are like those applied in step 120.

[0086] The next step 230, in view of the training of a second model 02(defined a plurality of parameters 02) for label prediction, consists in extracted pair of patches from the original HR image XHRand the resized IR building map YIR.

[0087] The format of the patches is identical to those extracted in step 130. For example, 10000 pairs of patches are extracted and stored in a second training database.

[0088] Step 240 corresponds to the training of the second model 02based on samples extracted from the second database.

[0089] In step 250, once the second model 02has been trained, it is used in inference on the original HR image XHRto generate an HR building map YHR.

[0090] The HR building map YHRhas the same size and spatial resolution as the initial HR image XHR.

[0091] Advantageously, after this first phase (also called cascade phase) 101 of progressive increase of the spatial resolution of the building map, method 100 goes on with a second phase 102, consisting of improving the quality of the HR building map generated by the first phase. The result of this second phase is an improved HR building map.

[0092] As illustrated in figure 4 with more detailed, in step 312 of the second phase 102, a LR image XLRis obtained by downsampling the initial HR image XHRin order to have a data matrix with the same size as the initial LR building map YLR, and then by upsampling this latter in order to have a data matrix with the same size as the initial HR image. In this process, the spatial information is reduced and the image thus obtained is of a low spatial resolution.

[0093] In step 314, a resized LR building map YLRis obtained by upsampling the initial LR building map YLRin order to have a data matrix with the same size as the initial image XHR. In step 320, one or several post-processing treatments are applied on the HR building map YHRand / or the initial image XHR. They are like the treatments applied in steps 120 and 220.

[0094] In step 330, in view of the training of an improvement model 0' (defined by a plurality of parameters 0') for label improvement, set of patches are extracted simultaneously from the original HR image XHRand the HR map YHR, but also from the LR image XLRand the resized LR building map YLR.

[0095] A patch may be of a different format than the patches of steps 130 and 230. For example, patches are now squares of 30x30 pixels.

[0096] For example, 5000 pairs of patches are extracted and stored in an improvement training database.

[0097] In step 340, the improvement model 0' is trained using samples of sets of patches extracted from the improvement training database.

[0098] The aim of the improvement model 0' is to align LR / HR imagery and simultaneously LR / HR building maps.

[0099] In step 350, once trained, the improvement model 0' is used in inference on the original HR image XHRto generate the final improved HR building map YHR.

[0100] The cascade phase 101 aims at estimating the high-resolution map based on the low- resolution map and the high-resolution image.

[0101] In the preferred embodiment, the cascade phase is made up of two successive iterations for gradually improving the size and the spatial resolution of the map. Even if two iterations of the loop are usually sufficient to efficiently handle the resolution gap that exists between the initial HR image and the initial LR map, more than two iterations may be performed. The nthiteration is then characterized by an nthintermediary resolution, which increases progressively at each new iteration to finally reach the size and spatial resolution of the initial image. Note that the input of nthiteration is the initial HR image downscaled to the nthintermediary resolution and the resized n-1thIR map upscaled to the nthintermediary resolution (with the resized IR map of rank 0 equal to the initial LR map).

[0102] In each iteration, the model 0„ (defined by a plurality of parameters 9n) used has the same architecture.

[0103] If the parameter of the model of the first iteration are initialized with random values, the parameters of the nthmodel of the nthiteration are initialized with the trained values of the parameters of the n-1thmodel of the n-1 iteration. This initialization improves the convergence speed of each training step. State-of-the-art methods perform poorly because they try to super resolve in a straightforward manner from the LR to the HR, but also because they do not efficiently handle the noise that is inherently present in the low-resolution map.

[0104] To be successful, the invention advantageously uses a loss function that is robust to noise.

[0105] In each iteration, the same loss function is used to train the model. Thus, to train the nthIR label prediction model 0„, the super resolution loss function sr is preferably of the form: here the set S contains 70% of the patches p (made up of an image patch xIRand a map patch yIR) of the database; the model 0nis parametrized by parameters 9n; DKLis a Kullback- Leibler divergence, and £ris an entropy regularisation term that is weighted by a scalar a.

[0106] Since the model fits pixels associated to small losses values, it is assumed to be more robust to noise generally associated to large loss values noisy pixels.

[0107] Beyond that, the KL divergence DKLmeasures the difference between two probability distributions rather than individual data points. This helps in smoothing out the noise as the KL divergence considers the overall distribution of pixels rather than being overly sensitive to specific noisy instances.

[0108] The second phase 101 aims at improving the quality of the estimated high-resolution map by attending to the LR / HR imagery alignment. Specifically, a novel loss function is designed to guide the model to attend to the high-frequency feature differences between HR and LR images and replicate such differences when the initial LR map and HR map are compared. Intuitively, this is done so that the spatial resolution gap from LR to HR image is also displayed from LR to HR maps. For that, a dedicated model is trained using an innovative noise-robust objective.

[0109] An edge-aware loss coupled with a noise-robust loss is implemented. The loss function Lnris preferably of the form:

[0110] Lnr COmp + Pledge

[0111] Where the loss function Lnris made up of a first term Lcompand a second term Ledge, their respective contribution being controlled by setting the value of parameter p.

[0112] Lfocai is the well-known focal loss function, which addresses class imbalance during training in tasks like object detection, focusing on hard-to-classify data by reducing the relative loss for well-classified data. L is the well-known L1-norm loss function that calculates the absolute difference between the predicted and true values, encouraging the model to produce predictions that are closer to the actual values.

[0113] Ledgeisaedge loss function that is dedicated to deal with the edges of building structure thanks to the loss computed over the gradients that are deduced from the HR and LR images and maps.

[0114] Preferably, the edge loss function is of the form: where xHRis the HR image patch of a set of patches; xLRis the LR image patch of the set of patches; yHR= 0'(0',xHR) is the HR map patch predicted by the model 0' from xHR, yHRis the HR map patch of the set of patches previously predicted by the model 0„; and yLRis the LR map patch of the set of patches; and where i is an integer indexing the columns of the data matrix; j is an integer indexing the rows of the data matrix; Gta gradient along direction i; Gj a gradient along direction j. The gradient of an image is a heatmap that represents the directional rate of change of pixel values. It is calculated by convolving the input image with a kernel whose values indicate the direction in which the gradient is to be performed. The directions i and j are conventionally set as the horizontal and vertical directions with respect to the input image.

[0115] Figure 5 shows several images and maps to illustrate the implementation of the method 100.

[0116] In figure 5, the different matrices of pixels are represented with the same dimensions even if they do not have the same size and / or spatial resolution.

[0117] Figure 5A is the initial HR optical image of an urban area of interest.

[0118] Figure 5B is the initial LR building map of the same urban area.

[0119] Figure 5C is the final HR map obtained according to the prior art super-resolution method, and

[0120] Figure 5D is the final HR building map obtained according to the label superresolution method 100.

[0121] When compared with a HR ground truth building map, the HR map of figure 5C - produced without the improvement phase - has a precision score of 59%, whereas the HR building map of figure 5D has a precision score of 74%.

[0122] The method according to the invention improves the accuracy of automatic building mapping by using a low-resolution label map in conjunction with high and low-resolution imagery. This accuracy gain comes at a much smaller cost, when compared to doing manual labelling of buildings.

[0123] Identification of individual buildings is critical for many different applications. Potential users can use this to quickly extract highly detailed building maps.

[0124] Although a convolutional neural network (CNN) with positional encodings has been used for each of the super-resolution and LR / HR imagery alignment models to test the preferred embodiment, the method is model-agnostic, and any other model architecture may accomplish satisfactory results.

[0125] The method according to the invention is particularly well suited to deal with maps representing buildings, in particular its second phase, since this type of structure is characterized by features with salient edges or contours. It could then be also applied to other type of maps representing objects characterized by delineated structures. It could also be applied on other maps relative to less delineated structures, such as land cover or urban infrastructure maps, the first phase bringing in itself some improvement.

[0126] In the preferred embodiment, the label of the building maps is a binary variable. In an alternative, the label has more than two possible values, for example three, four or five. In this case, the HR map may be obtained directly or by splitting the computation in several sub-problems, each sub-problem dealing with a reduced label of the binary type and the HR map being the aggregation of the results of each sub-problem.

Claims

CLAIMS1. A method (100) for generating a high spatial resolution - HR map (YHR, YHR) °fageographical area, characterized in that the method comprises: acquiring (103) an initial HR image XHR) of the geographical area and an initial low spatial resolution - LR map (TLR) of the geographical area; selecting (110) a set of N intermediary spatial resolutions - IR, N being an integer equal to or greater than two, the first intermediary spatial resolution of said set laying between the high spatial resolution and the low spatial resolution, the Nthintermediary spatial resolution of said set being equal to the high spatial resolution, and the nthintermediary spatial resolution, for n other than 1 and N, of said set laying between the high spatial resolution and the n-1thintermediary spatial resolution; and, a phase (101) of N iterations of a loop, the nthiteration of the loop, n being an integer between 1 and N, comprising: o downsampling (112) the initial HR image XHR) to the nthintermediary spatial resolution to obtain an nthIR image {XIR) o upsampling (114) an n-1thIR map at the nthintermediary spatial resolution to obtain a resized n-1thIR map, the 0thIR map being the initial LR map; o extracting (130) a plurality of pairs of patches, each pair of patches being made of an image patch extracted from the nthIR image and a corresponding IR map patch extracted from the resized n-1thIR map and storing said plurality of pairs of patches in a nthtraining database; o training (140) a nthIR label prediction model (0„) with a sample of pairs of patches from the nthtraining database; and, o using (150) the trained nthIR label prediction model in inference on the nthIR image to estimate an nthIR map, the HR map corresponding the NthIR map.

2. The method according to claim 1, wherein the loop is iterated two times.

3. The method according to any one of the claims 1 to 2, wherein the training of the nthIR label prediction model (02) starts with initializing a plurality of parameters of the nthIR label prediction model with values of the plurality of parameters of the trained n-1thIR label prediction model {0^.

4. The method according to any one of the claims 1 to 3, wherein a super-resolution loss function is used to train the nthIR label prediction model (0„):where the set S contains a plurality of the pairs of patches p of the nthtraining database; xIRand yIRare the nthIR image patch and nthresized IR map patch of the pair of patches p; the IR label prediction model 0nis parametrized by the plurality of parameters 9n; DKLis a Kullback-Leibler divergence; and £ris an entropy regularisation term that is weighted by a scalar a.

5. The method according to any one of the claims 1 to 4, wherein, after the downsampling and the upsampling steps, the label generation loop comprises a step (120) for post-processing the resized n-1thIR map.

6. The method according to any one of the claims 1 to 5, wherein, after the phase (101) of iteration, the method (100) further comprises a phase (102) for quality improvement of the NthHR map to obtain an improved HR map, the phase for quality improvement comprising: o resampling (312) the initial HR image (XHR) in order to obtain a LR image (XLR) having the same size as the initial HR image (XHR) but the spatial resolution of the initial LR map (YLR, o upsampling (314) the initial LR map (YLR) in order to obtain a resized LR map (YLR) with the same size as the initial HR image (XHR, o extracting (330) a plurality of set of patches, each set of patches being made of : an HR image patch extracted from the initial HR image; a LR image patch extracted from the LR image; an HR map patch extracted from the NthHR map; and a LR map patch extracted from the resized LR map, each set of patches being stored in a quality improvement database; o training a quality improvement model (0') with a sample of sets of patches from the quality improvement database; and, o using the trained quality improvement model in inference on the HR image (XHR) to generate the improved HR map (YHR).

7. The method according to claim 6, wherein the loss function used in the training step of the quality improvement model is of the form:Lnr COmp + Pledge ^focal T ^1 T Pledge where Lfoccis the focal loss function; is the L1-norm loss function, and Ledgeis a edge loss function, the edge loss function being of the form:+ Ll(yHR— HR) where xHRis the HR image patch of a set of patches; xLRis the LR image patch of said set of patches; yHRis a HR map patch predicted by the quality improvement model 0' from xHR, yHRis the HR map patch of said set of patches; and yLRis theLR map patch of said set of patches; and where Gtis a gradient along a first direction and Gj is a gradient along a second direction of a data matrix.

8. The method according to any one of claims 6 to 7, comprising, before extracting a plurality of set of patches, a step (320) of post-processing the HR map.

9. The method according to any one of claims 1 to 8, wherein the initial HR image is an optical image of the geographical area.

10. The method according to any one of claims 1 to 9, wherein the initial LR map is a building map of the geographical area, so that the HR map and / or the improved HR map is a building map of the geographical area.

11. A computer program product comprising a computer readable program for causing a computer to realize the method for generated an HR map (YHR) of a geographical area according to any one of the claims 1 to 10.

Citation Information

Patent Citations

  • Systems and methods for segmenting 3D images

    US20230386067A1