Image-based feature matching for images of a scene
The UDA-based semantic renormalization and pseudo-labels method addresses the day-night challenge in image-based feature matching, enabling efficient and accurate training without GPS, by leveraging semantic labels to bridge lighting gaps and create reliable pixel-wise correspondences.
Patent Information
- Application Number
- PCT/EP2024/051771
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-25
- Publication Date
- 2025-07-31
AI Technical Summary
Existing image-based feature matching technologies struggle to effectively handle the day-night scenario due to a significant visual domain gap, requiring costly manual annotation and GPS-dependent data collection, making it difficult to generalize well for nighttime conditions.
An unsupervised domain adaptation (UDA) approach that uses semantic-aware renormalization and pseudo-labels to create image pairs for training, allowing for device-friendly day-night feature matching without relying on GPS, by leveraging semantic labels and renormalizing images to bridge the lighting gap.
Enables efficient and cost-effective training for day-night scenarios by narrowing the appearance gap between day and night images, facilitating reliable pixel-wise correspondences without human annotation, thus improving feature matching accuracy.
Smart Images

Figure EP2024051771_31072025_PF_FP_ABST
Abstract
Description
[0001] IMAGE-BASED FEATURE MATCHING FOR IMAGES OF A SCENE
[0002] TECHNICAL FIELD
[0003] The present disclosure relates, in general, to image-based feature matching for images of a scene. Aspects of the disclosure relate to a device-friendly and semantic -guided approach for day -night feature matching.
[0004] BACKGROUND
[0005] Image-based feature matching is the foundation of many downstream computer vision tasks, such as localisation, 3D reconstruction, image retrieval, and augmented reality. While deep learning-based feature matching of images taken under satisfactory illumination conditions (i.e., at daytime) has been well researched, the day-night scenario remains an open challenge.
[0006] Existing approaches to day-night feature matching are mainly trained on daytime data (or day-style transferred data) and tested using a day-night benchmark (such as the Aachen day-night benchmark). However, the purpose of the benchmark is mainly to test the robustness of the daytime model and not to build an effective pipeline to particularly handle the day-night case.
[0007] If a feature matching algorithm is built particularly for addressing the day-night scenario, it is desirable to include the nighttime data into training. However, when the night-time data is considered for training, there exists a huge visual domain gap, making it difficult for the currently available approaches to generalise well for the night data. As such, in order to train on day- night pairs, ground-truth correspondences between the day-night pairs must be manually prepared, causing a large annotation burden. Furthermore, to capture the paired day-night data for training, researchers must be equipped with accurate GPS devices and assume that the GPS signal is always stable at the locations at which data is acquired, thus setting strict requirements for conducting such research.
[0008] SUMMARY
[0009] An objective of the present disclosure is to provide enable easy inclusion of night data for training and to build a pipeline for handling the day-night scenario in an unsupervised manner.
[0010] The foregoing and other objectives are achieved by the features of the independent claims.
[0011] Further implementation forms are apparent from the dependent claims, the description and the Figures.
[0012] A first aspect of the present disclosure provides a method of image-based feature matching for images of a scene, the method comprising acquiring a set of images comprising multiple images of the scene, respective ones of the multiple images captured under differing lighting conditions, for each image of the multiple images in the set of images, using unsupervised domain adaptation, UDA, providing respective semantic labels for features representing segmented objects and / or regions of the images, using the semantic labels, comparing the multiple images in the set of images, whereby to create, based on semantic similarity, a plurality of image pairs, wherein each of the plurality of image pairs comprises a first image of the scene captured under a first lighting condition and a second image of the scene captured under a second lighting condition, for each pair of the plurality of image pairs, generating: based on image characteristics of the second image, a renormalised first image indicative of the first image under the second lighting condition, and, based on image characteristics of the first image, a renormalised second image indicative of the second image under the first lighting condition, for each pair of the plurality of image pairs, generating, based on the semantic labels provided using the UDA, pseudo correspondences for the first image and for the second image on the basis of the results of a comparison between: the first image and the renormalised second image, and the second image and the renormalised first image, respectively, and training a neural network using the generated pseudo correspondences.
[0013] Accordingly, the UDA-assisted data processing enables easy inclusion of night data for training, allowing a pipeline to particularly handle the day-night case to be built in an unsupervised manner. Furthermore, the approach presented herein is device-friendly, as it does not rely on using GPS to create day-night pairs, therefore enabling the approach to be adopted even in privacy protected areas (or other areas where GPS signal is not available). By employing semantic-aware renormalisation, the appearance gap between day and night images can be narrowed for training. Additionally, as the training of a machine learning model is assisted by the renormalised data and the UDA semantic labels, reliable pixel-wise correspondences can be selected to create pseudo-labels.
[0014] Providing the respective semantic labels for the features representing the segmented objects and / or regions of the images may comprise training a segmentation model using the unsupervised domain adaptation, UDA, on a labelled dataset, and applying the segmentation model to the set of images, whereby to provide the respective semantic labels.
[0015] The segmented objects may comprise static objects, and comparing the multiple images in the set of images may comprise separating the static objects into a first group and a second group, wherein the first group comprises a first class of static objects, and the second group comprises a second class of static objects, calculating a label overlap ratio based on a weighted average label overlap ratio between the first group and the second group, and comparing the multiple images in the set of images based on the calculated label overlap ratio.
[0016] The first class of static objects may comprise static objects with a first predefined response, and the second class of static objects may comprise static objects with a second predefined response.
[0017] Generating the renormalized first image and the renormalized second image may comprise performing, based on the respective semantic labels associated with the first image and the second image, a channel-wise normalisation of at least one image characteristic of the first image with respect to the second image, and the second image with respect to the first image, wherein the at least one image characteristic comprises mean and / or variance.
[0018] The channel-wise normalization may be performed based on the semantic labels provided using the UDA, such that class wise characteristics of the first image are used for normalising corresponding features of the second image, and class-wise characteristics of the second image are used for normalising corresponding features of the first image.
[0019] The method may further comprise generating feature descriptors for the renormalised first image and the renormalised second image, comparing a similarity of the feature descriptors generated for the first image and the renormalised second image, and determining a first set of pixel locations of the feature descriptors based on the similarity being above a threshold value, comparing the similarity of the feature descriptors generated for the second image and the renormalised first image, and determining a second set of pixel locations of the feature descriptors based on the similarity being above the threshold value, comparing the first set of pixel locations and the second set of pixel locations to determine a set of concurrent pixel locations that are present in both the first set of pixel locations and the second set of pixel locations, generating a set of pseudo-labels for the set of concurrent pixel locations, and, based on the semantic labels provided using the UDA, removing those pseudo-labels of the set of pseudo-labels whose semantic information represented by the semantic labels does not match between the first set of pixel locations and the second set of pixel locations. A second aspect of the present disclosure provides a computer readable storage medium comprising computer program code, accessible by an apparatus comprising a processor, to provide instructions and / or data to the apparatus, the computer program code configured to, with the processor, cause the apparatus to acquire a set of images comprising multiple images of the scene, respective ones of the multiple images captured under differing lighting conditions, for each image of the multiple images in the set of images, using unsupervised domain adaptation, UDA, provide respective semantic labels for features representing segmented objects and / or regions of the images, using the semantic labels, compare the multiple images in the set of images, whereby to create, based on semantic similarity, a plurality of image pairs, wherein each of the plurality of image pairs comprises a first image of the scene captured under a first lighting condition and a second image of the scene captured under a second lighting condition, for each pair of the plurality of image pairs, generate: based on image characteristics of the second image, a renormalised first image indicative of the first image under the second lighting condition, and, based on image characteristics of the first image, a renormalised second image indicative of the second image under the first lighting condition, for each pair of the plurality of image pairs, generate, based on the semantic labels provided using the UDA, pseudo correspondences for the first image and for the second image on the basis of the results of a comparison between: the first image and the renormalised second image, and the second image and the renormalised first image, respectively, and train a neural network using the generated pseudo correspondences.
[0020] The computer program code configured to, with the processor, cause the apparatus to provide the respective semantic labels for the features representing the segmented objects and / or regions of the images may comprise program code configured to, with the processor, cause the apparatus to train a segmentation model using the unsupervised domain adaptation, UDA, on a labelled dataset, and apply the segmentation model to the set of images, whereby to provide the respective semantic labels.
[0021] The segmented objects may comprise static objects, wherein the computer program code configured to, with the processor, cause the apparatus to compare the multiple images in the set of images may comprise program code configured to, with the processor, cause the apparatus to separate the static objects into a first group and a second group, wherein the first group comprises a first class of static objects, and the second group comprises a second class of static objects, calculate a label overlap ratio based on a weighted average label overlap ratio between the first group and the second group, and compare the multiple images in the set of images based on the calculated label overlap ratio.
[0022] The computer program code configured to, with the processor, cause the apparatus to generate the renormalised first image and the renormalised second image may comprise program code configured to, with the processor, cause the apparatus to perform, based on the respective semantic labels associated with the first image and the second image a channel-wise normalisation of at least one image characteristic of the first image with respect to the second image, and the second image with respect to the first image, wherein the at least one image characteristic comprises mean and / or variance.
[0023] The computer readable storage medium may further comprise program code configured to, with the processor, cause the apparatus to generate feature descriptors for the renormalised first image and the renormalised second image, compare a similarity of the feature descriptors generated for the first image and the renormalised second image, and determining a first set of pixel locations of the feature descriptors based on the similarity being above a threshold value, compare the similarity of the feature descriptors generated for the second image and the renormalised first image, and determining a second set of pixel locations of the feature descriptors based on the similarity being above the threshold value, compare the first set of pixel locations and the second set of pixel locations to determine a set of concurrent pixel locations that are present in both the first set of pixel locations and the second set of pixel locations, generate a set of pseudo-labels for the set of concurrent pixel locations, and, based on the semantic labels provided using the UDA, remove those pseudo-labels of the set of pseudo-labels whose semantic information represented by the semantic labels does not match between the first set of pixel locations and the second set of pixel locations. The computer readable storage medium may further comprise program code configured to, with the processor, cause the apparatus to calculate a descriptor loss based on the set of pseudo-labels.
[0024] A third aspect of the present disclosure provides an apparatus for image-based feature matching for images of a scene, the apparatus comprising a processor, a memory coupled to the processor, the memory configured to store program code executable by the processor, the program code comprising one or more instructions, whereby to cause the apparatus to acquire a set of images comprising multiple images of the scene, respective ones of the multiple images captured under differing lighting conditions, for each image of the multiple images in the set of images, using unsupervised domain adaptation, UD A, provide respective semantic labels for features representing segmented objects and / or regions of the images, using the semantic labels, compare the multiple images in the set of images, whereby to create, based on semantic similarity, a plurality of image pairs, wherein each of the plurality of image pairs comprises a first image of the scene captured under a first lighting condition and a second image of the scene captured under a second lighting condition, for each pair of the plurality of image pairs, generate: based on image characteristics of the second image, a renormalised first image indicative of the first image under the second lighting condition, and, based on image characteristics of the first image, a renormalised second image indicative of the second image under the first lighting condition, for each pair of the plurality of image pairs, generate, based on the semantic labels provided using the UDA, pseudo correspondences for the first image and for the second image on the basis of the results of a comparison between: the first image and the renormalised second image, and the second image and the renormalised first image, respectively, and train a neural network using the generated pseudo correspondences.
[0025] These and other aspects of the invention will be apparent from the embodiment / s) described below.
[0026] BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order that the present invention may be more readily understood, embodiments of the invention will now be described, by way of example, with reference to the accompanying drawings, in which:
[0028] Figure 1 is a flow chart of a method of image-based feature matching for images of a scene according to an example;
[0029] Figure 2 is a schematic representation of day-night pair extraction based on semantic similarity to an example:
[0030] Figure 3 is a schematic representation of semantic renormalisation according to an example:
[0031] Figure 4 is a schematic representation of training and loss calculation of cross-illumination feature matching according to an example: and
[0032] Figure 5 is a schematic representation of an apparatus according to an example.
[0033] DETAILED DESCRIPTION
[0034] Example embodiments are described below in sufficient detail to enable those of ordinary skill in the art to embody and implement the systems and processes herein described. It is important to understand that embodiments can be provided in many alternate forms and should not be construed as limited to the examples set forth herein.
[0035] Accordingly, while embodiments can be modified in various ways and take on various alternative forms, specific embodiments thereof are shown in the drawings and described in detail below as examples. There is no intent to limit to the particular forms disclosed. On the contrary, all modifications, equivalents, and alternatives falling within the scope of the appended claims should be included. Elements of the example embodiments are consistently denoted by the same reference numerals throughout the drawings and detailed description where appropriate.
[0036] The terminology used herein to describe embodiments is not intended to limit the scope. The articles “a,” “an,” and “the” are singular in that they have a single referent, however the use of the singular form in the present document should not preclude the presence of more than one referent. In other words, elements referred to in the singular can number one or more, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises,” “comprising,” “includes,” and / or “including,” when used herein, specify the presence of stated features, items, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, items, steps, operations, elements, components, and / or groups thereof.
[0037] Unless otherwise defined, all terms (including technical and scientific terms) used herein are to be interpreted as is customary in the art. It will be further understood that terms in common usage should also be interpreted as is customary in the relevant art and not in an idealized or overly formal sense unless expressly so defined herein.
[0038] Due to the degraded visibility of night-time data, without human annotation for pixel-wise correspondences, the existing prior art solutions can only be trained on a daytime dataset and evaluated using a day-night benchmark - that is, they cannot be trained on day-night paired data to handle this challenging cross-illumination scenario.
[0039] For example, for SFD2 (semantic -guided feature detection and description), the adopted general semantic segmentation network is not particularly trained for the customised dataset, thus failing to produce high quality semantic prediction even on the unseen daytime images, let alone the more challenging night-time images. Therefore, using SFD2, it is not possible to perform day-night joint training.
[0040] Robot Car provides a dataset for cross-domain training of feature matching where human annotation for pixel-wise correspondences is needed. However, creating a pixel-wisely labelled training set for day-night feature matching is costly and laboursome. Furthermore, to collect paired day-night data for training, the researches must be equipped with accurate GPS devices and assume that the GPS signal is always stable at the locations selected for data acquisition, which may not be the case.
[0041] According to an example, there is provided a mechanism to enable easy inclusion of night data for training and to build a pipeline for handling the day-night scenario in an unsupervised manner. Advantageously, the mechanism provided is devicefriendly, as it does not rely on using GPS to create day-night pairs, therefore enabling the approach to be adopted even in privacy protected areas (or other areas where GPS signal is not available). By employing semantic-aware renormalisation, the appearance gap between day and night images can be narrowed for training. Additionally, as the training of a machine learning model is assisted by the renormalised data and the UDA semantic labels, reliable pixel-wise correspondences can be selected to create pseudo-labels.
[0042] Examples in the present disclosure can be provided as methods, systems or machine-readable instructions, such as any combination of software, hardware, firmware or the like. Such machine-readable instructions may be included on a computer readable storage medium (including but not limited to disc storage, CD-ROM, optical storage, etc.) having computer readable program codes therein or thereon.
[0043] The present disclosure is described with reference to flow charts and / or block diagrams of the method, devices and systems according to examples of the present disclosure. Although the flow diagrams described above show a specific order of execution, the order of execution may differ from that which is depicted. Blocks described in relation to one flow chart may be combined with those of another flow chart. In some examples, some blocks of the flow diagrams may not be necessary and / or additional blocks may be added. It shall be understood that each flow and / or block in the flow charts and / or block diagrams, as well as combinations of the flows and / or diagrams in the flow charts and / or block diagrams can be realized by machine readable instructions.
[0044] The machine-readable instructions may, for example, be executed by a machine such as a general-purpose computer, user equipment such as a smart device, e.g., a smart phone, a special purpose computer, an embedded processor or processors of other programmable data processing devices to realize the functions described in the description and diagrams. In particular, a processor or processing apparatus may execute the machine-readable instructions. Thus, modules of apparatus (for example, a module implementing a comparator unit, or a firewall structure and so on) may be implemented by a processor executing machine readable instructions stored in a memory, or a processor operating in accordance with instructions embedded in logic circuitry. The term 'processor' is to be interpreted broadly to include a CPU, processing unit, ASIC, logic unit, or programmable gate set etc. The methods and modules may all be performed by a single processor or divided amongst several processors.
[0045] Such machine-readable instructions may also be stored in a computer readable storage that can guide the computer or other programmable data processing devices to operate in a specific mode. For example, the instructions may be provided on a non- transitory computer readable storage medium encoded with instructions, executable by a processor.
[0046] Figure 1 is a flow chart of a method of image-based feature matching for images of a scene according to an example. The method comprises, in block 101, acquiring a set of images comprising multiple images of the scene, respective ones of the multiple images captured under differing lighting conditions. Here, the term “differing lighting conditions” may refer to, for example, an image captured under first lighting conditions, and an image captured under second lighting conditions. For example, the set of images may comprise at least one pair of images of the same scene, one captured during daytime, and one captured during night-time (or any other time at which the amount of lighting has decreased, for example, late afternoon and / or evening).
[0047] The method comprises, in block 102, for each image of the multiple images in the set of images, using unsupervised domain adaptation (UDA), providing respective semantic labels for features representing segmented objects and / or regions of the images. Here, the term “segmented objects” may refer to distinct objects present in the image.
[0048] Providing the respective semantic labels may comprise training a segmentation model using the UDA on a labelled dataset and applying the segmentation model to the set of images, whereby to provide the respective semantic labels.
[0049] In other words, in order to train a model that works for the day -night scenario, the first step would be to create a paired day- night dataset from the captured raw data (see block 103 described below). However, in order to create such a dataset, the existing solutions rely heavily on comparing GPS locations at which the images have been captured, creating day -night pairs based on the images being captured at the same location. This approach requires the use of accurate GPS devices, as well as stable GPS signal at all locations, which may not be possible.
[0050] In contrast, the present solution leverages a domain adaptation technique which trains a segmentation model on an open source labelled dataset and transfers the knowledge to the collected unlabelled data. As such, using the invention, for each day or night captured image, high quality semantic labels may be generated. The segmented objects may comprise objects present in the captured image, for example, a traffic light, a building, or similar. The segmented objects may be categorised as static objects (i.e., objects that do not change their location between images) and / or dynamic objects (i.e., objects that may change their location between images, for example, cars). The semantic labels may be generated using an open-source dataset, such as the Dark Zurich dataset. The method comprises, in block 103 , using the semantic labels, comparing the multiple images in the set of images, whereby to create, based on semantic similarity, a plurality of image pairs, wherein each of the plurality of image pairs comprises a first image of the scene captured under a first lighting condition and a second image of the scene captured under a second lighting condition.
[0051] As discussed above, the first lighting condition may comprise a lighting condition corresponding to daytime, and the second lighting condition may comprise a lighting condition other than daytime (i.e., night-time, late evening, or similar). The method may comprise create day-night pairs based on their semantic similarity, using the UDA segmented masks. Specifically, the semantic similarity of static classes only may be considered. Looping over the whole dataset, a Label Overlap Ratio (LOR) may be considered according to the below equation.
[0052] If the LOR is larger than a certain threshold tau, a day-night pair may be found. To calculate a semantic overlap ratio, the static classes may be divided into two groups: shift sensitive classes (SSC) (i.e., classes with a predefined response to light, for example, traffic lights, traffic signs or poles), and other static classes (i.e., classes less sensitive to light - for example, roads and buildings). Different to the prior art solutions, in order to ensure that all static classes contribute to the semantic similarity comparison, the LOR may be calculated based on a weighted average between SSC and non-SSC, with the judging condition being prioritised based on the SSC. If the overlap is larger than a threshold, a pair may be created.
[0053] To aid understanding of the image pair creation, reference is made to Figure 2. Figure 2 is a schematic representation of day- night pair extraction based on semantic similarity to an example. The first image 201 may comprise a day image, and the second image 202 may comprise a night image. As discussed, for semantic similarity consideration, only static classes 203 may be considered. In particular, weight may be given to the shift sensitive classes 204.
[0054] Referring back to Figure 1, in block 104, the method comprises, for each pair of the plurality of image pairs, generating: based on image characteristics of the second image, a renormalised first image indicative of the first image under the second lighting condition, and, based on image characteristics of the first image, a renormalised second image indicative of the second image under the first lighting condition.
[0055] In other words, in order to make use of the image pairs created in block 103 and the semantic information generated in block 102 to train a strong feature matching algorithm for the day-night scenario, semantic-aware renormalisation may be applied. In particular, a pair of day-night images may be considered, and channel-wise normalisation (regarding, for example, mean and variance) may be performed with respect to each other, so as to narrow the appearance gap between the first image and the second image. Figure 3 is a schematic representation of semantic renormalisation according to an example. A renormalised first image 301-2 may be generated based on the first image 301-1, with a renormalised second image 302-2 being generated based on the second image 302-1. As discussed above, the renormalised first image 301-2 may be indicative of the first image 301-1 under the second lighting condition (i.e., the lighting condition associated with the second image 302-1), and the renormalised second image 302-2 may be indicative of the second image 302-1 under the first lighting condition (i.e., the lighting condition associated with the first image 301-1). However, instead of blindly normalising the whole image using the statistics associated with the paired image, the semantic labels may be considered, such that, for example, night sky statistics are used to normalise sky regions of the day image, and vice versa. By using this approach, the resulting renormalised images 301-2 and 302-2 may more closely resemble the paired images. The semantic labels may be acquired in an unsupervised manner, transferred from a labelled open-source dataset using domain adaptation.
[0056] The method comprises, in block 105, for each pair of the plurality of image pairs, generating, based on the semantic labels provided using the UDA, pseudo correspondences for the first image and for the second image on the basis of the results of a comparison between the first image and the renormalised second image, and the second image and the renormalised first image, respectively. In block 106, the method comprises training a neural network using the generated pseudo correspondences.
[0057] Figure 4 is a schematic representation of training and loss calculation of cross-illumination feature matching according to an example, building on the method steps of blocks 105 and 106. For training the cross-illumination feature matching network, in each training iteration, all four images (i.e., first image 401-1 under the first lighting condition, second image 402-1 of the same scene under the second lighting condition, renormalised first image 401-2 and the renormalised second image 402-2) may be passed to the network and their outputs may be obtained, with a descriptor loss being computer with semantic guidance. The output feature descriptors may be obtained from the first image 401-1, the second image 402-1, the renormalised first image 401-2, and the renormalised second image 402-2. Here, the term feature descriptors may refer to the pixel-wise / location-wise raw outputs of the neural network.
[0058] Due to the fact that it may be quite challenging to directly match the features between the first image 401-1 and the second image 402-1, instead, a bridge may be introduced between the two. In particular, since the renormalised first image 401-2 and the second image 402-1 have more similar appearances (i.e., the lighting condition is the same or very similar between the two images), it may be easier to find the closest feature descriptors between them according to a threshold and to acquire a first set of pixel locations of the similar feature descriptors. Similarly, a second set of pixel locations where the feature descriptors are similar may also be found for the renormalised second image 402-2 and the first image 401-1. With the two sets of pixel locations of the similar feature descriptors, a consistency check may be performed between them to determine a set of concurrent pixel locations, with only the concurrent pixel locations being left as the pseudo-labels 403. That is, the pixel locations of similar feature descriptors found only as a result of a single comparison between the image and its renormalised version may be deleted.
[0059] The pseudo-labels 403 may then be checked in terms of their semantic consistency. In particular, those pseudo-labels 403 whose semantic information does not match between the first set of pixel locations and the second set of pixel locations may be removed from the set of pseudo-labels 403. For example, if a road pixel has been matched to a tree pixel, the pseudo-label 403 may be removed. Due to the camera pose difference, the correctly matched descriptors usually do not appear in the same pixel location of the images. Due to the illumination gap, there can also be wrongly matched descriptors. Therefore, the pseudolabels may be interpreted as the locations of the reliably matched descriptors. Once these locations are known, the neural network can be encouraged to minimise the distance between the feature descriptors between two images.
[0060] The remaining pseudo-labels 403 may be adopted to compute the descriptor loss between the first image 401-1 and the second image 402-1. In such way, a reliable descriptor loss that is applied directly to the matched pair of images can be ensured. Furthermore, as the process does not require any human annotation and is completely unsupervised, the approach is less costly and laboursome than the currently available methods.
[0061] Figure 5 is a schematic representation of an apparatus according to an example. The apparatus 500 may comprise a processor 503, and a memory 505 coupled to the processor 503 and configured to store instructions or program code 507, executable by the processor 503. The apparatus 500 may comprise the program code 507 arranged to cause the apparatus to perform the method of image-based feature matching for images of a scene as described herein.
[0062] According to an example, machine-readable instructions can be loaded onto a computer or other programmable data processing devices, so that the computer or other programmable data processing devices perform a series of operations to produce computer-implemented processing, thus the instructions executed on the computer or other programmable devices provide an operation for realizing functions specified by flow(s) in the flow charts and / or block(s) in the block diagrams.
[0063] Further, the teachings herein may be implemented in the form of a computer or software product, such as a non-transitory machine-readable storage medium, the computer software or product being stored in a storage medium and comprising a plurality of instructions, e.g., machine readable instructions, for making a computer device implement the methods recited in the examples of the present disclosure.
[0064] In some examples, some methods can be performed in a cloud-computing or network-based environment. Cloud-computing environments may provide various services and applications via the Internet. These cloud-based services (e.g., software as a service, platform as a service, infrastructure as a service, etc.) may be accessible through a web browser or other remote interface of the user equipment for example. Various functions described herein may be provided through a remote desktop environment or any other cloud-based computing environment.
[0065] While various embodiments have been described and / or illustrated herein in the context of fully functional computing systems, one or more of these exemplary embodiments may be distributed as a program product in a variety of forms, regardless of the particular type of computer-readable-storage media used to actually cany out the distribution. The embodiments disclosed herein may also be implemented using software modules that perform certain tasks. These software modules may include script, batch, or other executable files that may be stored on a computer-readable storage medium or in a computing system. In some embodiments, these software modules may configure a computing system to perform one or more of the exemplary embodiments disclosed herein. In addition, one or more of the modules described herein may transform data, physical devices, and / or representations of physical devices from one form to another.
[0066] The preceding description has been provided to enable others skilled in the art to best utilize various aspects of the exemplary embodiments disclosed herein. This exemplary description is not intended to be exhaustive or to be limited to any precise form disclosed. Many modifications and variations are possible without departing from the spirit and scope of the instant disclosure. The embodiments disclosed herein should be considered in all respects illustrative and not restrictive. Reference should be made to the appended claims and their equivalents in determining the scope of the instant disclosure.
Claims
CLAIMS1. A method of image-based feature matching for images of a scene, the method comprising: acquiring a set of images comprising multiple images of the scene, respective ones of the multiple images captured under differing lighting conditions (101); for each image of the multiple images in the set of images, using unsupervised domain adaptation, UD A, providing respective semantic labels for features representing segmented objects and / or regions of the images (102); using the semantic labels, comparing the multiple images in the set of images, whereby to create, based on semantic similarity, a plurality of image pairs, wherein each of the plurality of image pairs comprises a first image of the scene captured under a first lighting condition and a second image of the scene captured under a second lighting condition (103); for each pair of the plurality of image pairs, generating: based on image characteristics of the second image, a renormalised first image indicative of the first image under the second lighting condition, and, based on image characteristics of the first image, a renormalised second image indicative of the second image under the first lighting condition (104); for each pair of the plurality of image pairs, generating, based on the semantic labels provided using the UDA, pseudo correspondences for the first image and for the second image on the basis of the results of a comparison between: the first image and the renormalised second image, and the second image and the renormalised first image, respectively (105); and training a neural network using the generated pseudo correspondences (106).
2. The method of claim 1, wherein providing the respective semantic labels for the features representing the segmented objects and / or regions of the images (102) comprises: training a segmentation model using the unsupervised domain adaptation, UDA, on a labelled dataset; applying the segmentation model to the set of images, whereby to provide the respective semantic labels.
3. The method of claim 1 or 2, wherein the segmented objects comprise static objects, wherein comparing the multiple images in the set of images comprises: separating the static objects into a first group and a second group, wherein the first group comprises a first class of static objects, and the second group comprises a second class of static objects; calculating a label overlap ratio based on a weighted average label overlap ratio between the first group and the second group; and comparing the multiple images in the set of images based on the calculated label overlap ratio.
4. The method of claim 3, wherein the first class of static objects comprises static objects with a first predefined response, and the second class of static objects comprises static objects with a second predefined response.
5. The method of any one of claims 1 to 4, wherein generating the renormalised first image and the renormalised second image (104) comprises: performing, based on the respective semantic labels associated with the first image and the second image, a channel-wise normalisation of at least one image characteristic of the first image with respect to the second image, and the second image with respect to the first image, wherein the at least one image characteristic comprises mean and / or variance.
6. The method of claim 5, wherein the channel-wise normalisation is performed based on the semantic labels provided using the UDA, such that class wise characteristics of the first image are used for normalising corresponding features of the second image, and class-wise characteristics of the second image are used for normalising corresponding features of the first image.
7. The method of any preceding claim, further comprising: generating feature descriptors for the renormalised first image and the renormalised second image; comparing a similarity of the feature descriptors generated for the first image and the renormalised second image, and determining a first set of pixel locations of the feature descriptors based on the similarity being above a threshold value; comparing the similarity of the feature descriptors generated for the second image and the renormalised first image, and determining a second set of pixel locations of the feature descriptors based on the similarity being above the threshold value; comparing the first set of pixel locations and the second set of pixel locations to determine a set of concurrent pixel locations that are present in both the first set of pixel locations and the second set of pixel locations; generating a set of pseudo-labels for the set of concurrent pixel locations; and based on the semantic labels provided using the UDA, removing those pseudo-labels of the set of pseudolabels whose semantic information represented by the semantic labels does not match between the first set of pixel locations and the second set of pixel locations.
8. The method of claim 7, further comprising: calculating a descriptor loss based on the set of pseudo-labels.
9. A computer readable storage medium comprising computer program code, accessible by an apparatus comprising a processor, to provide instructions and / or data to the apparatus, the computer program code configured to, with the processor, cause the apparatus to: acquire a set of images comprising multiple images of the scene, respective ones of the multiple images captured under differing lighting conditions;for each image of the multiple images in the set of images, using unsupervised domain adaptation, UDA, provide respective semantic labels for features representing segmented objects and / or regions of the images; using the semantic labels, compare the multiple images in the set of images, whereby to create, based on semantic similarity, a plurality of image pairs, wherein each of the plurality of image pairs comprises a first image of the scene captured under a first lighting condition and a second image of the scene captured under a second lighting condition; for each pair of the plurality of image pairs, generate: based on image characteristics of the second image, a renormalised first image indicative of the first image under the second lighting condition, and, based on image characteristics of the first image, a renormalised second image indicative of the second image under the first lighting condition; for each pair of the plurality of image pairs, generate, based on the semantic labels provided using the UDA, pseudo correspondences for the first image and for the second image on the basis of the results of a comparison between: the first image and the renormalised second image, and the second image and the renormalised first image, respectively; and train a neural network using the generated pseudo correspondences.
10. The computer readable storage medium of claim 9, wherein the computer program code configured to, with the processor, cause the apparatus to provide the respective semantic labels for the features representing the segmented objects and / or regions of the images comprises program code configured to, with the processor, cause the apparatus to: train a segmentation model using the unsupervised domain adaptation, UDA, on a labelled dataset; apply the segmentation model to the set of images, whereby to provide the respective semantic labels.
11. The computer readable storage medium of claim 9 or 10, wherein the segmented objects comprise static objects, wherein the computer program code configured to, with the processor, cause the apparatus to compare the multiple images in the set of images comprises program code configured to, with the processor, cause the apparatus to: separate the static objects into a first group and a second group, wherein the first group comprises a first class of static objects, and the second group comprises a second class of static objects; calculate a label overlap ratio based on a weighted average label overlap ratio between the first group and the second group; and compare the multiple images in the set of images based on the calculated label overlap ratio.
12. The computer readable storage medium of any one of claims 9 to 11, wherein the computer program code configured to, with the processor, cause the apparatus to generate the renormalised first image and the renormalised second image comprises program code configured to, with the processor, cause the apparatus to: perform, based on the respective semantic labels associated with the first image and the second image a channel-wise normalisation of at least one image characteristic of the first image with respect to the second image, and the second image with respect to the first image,wherein the at least one image characteristic comprises mean and / or variance.
13. The computer readable storage medium of any one of claims 9 to 12, further comprising program code configured to, with the processor, cause the apparatus to: generate feature descriptors for the renormalised first image and the renormalised second image; compare a similarity of the feature descriptors generated for the first image and the renormalised second image, and determining a first set of pixel locations of the feature descriptors based on the similarity being above a threshold value; compare the similarity of the feature descriptors generated for the second image and the renormalised first image, and determining a second set of pixel locations of the feature descriptors based on the similarity being above the threshold value; compare the first set of pixel locations and the second set of pixel locations to determine a set of concurrent pixel locations that are present in both the first set of pixel locations and the second set of pixel locations; generate a set of pseudo-labels for the set of concurrent pixel locations; and based on the semantic labels provided using the UDA, remove those pseudo-labels of the set of pseudolabels whose semantic information represented by the semantic labels does not match between the first set of pixel locations and the second set of pixel locations.
14. The computer readable storage medium of any one of claims 9 to 13, further comprising program code configured to, with the processor, cause the apparatus to: calculate a descriptor loss based on the set of pseudo-labels.
15. An apparatus (500) for image-based feature matching for images of a scene, the apparatus comprising: a processor (503); a memory (505) coupled to the processor (503), the memory (505) configured to store program code (507) executable by the processor, the program code (507) comprising one or more instructions, whereby to cause the apparatus (500) to: acquire a set of images comprising multiple images of the scene, respective ones of the multiple images captured under differing lighting conditions; for each image of the multiple images in the set of images, using unsupervised domain adaptation, UDA, provide respective semantic labels for features representing segmented objects and / or regions of the images; using the semantic labels, compare the multiple images in the set of images, whereby to create, based on semantic similarity, a plurality of image pairs, wherein each of the plurality of image pairs comprises a first image of the scene captured under a first lighting condition and a second image of the scene captured under a second lighting condition;for each pair of the plurality of image pairs, generate: based on image characteristics of the second image, a renormalised first image indicative of the first image under the second lighting condition, and, based on image characteristics of the first image, a renormalised second image indicative of the second image under the first lighting condition; for each pair of the plurality of image pairs, generate, based on the semantic labels provided using theUDA, pseudo correspondences for the first image and for the second image on the basis of the results of a comparison between: the first image and the renormalised second image, and the second image and the renormalised first image, respectively; and train a neural network using the generated pseudo correspondences.