Image matching device, image matching method, training device, training method, and program
The image matching device enhances image resolution before feature extraction to ensure accurate matching of ground and aerial images, addressing the challenge of varying resolutions in existing systems.
Patent Information
- Application Number
- JP2026511673
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-08-21
- Publication Date
- 2026-08-26
AI Technical Summary
Existing systems for cross-view image matching between ground and aerial images do not adequately address the issue of varying image resolutions, which can hinder accurate determination of image matches.
An image matching device that enhances the resolution of lower-resolution images to a higher level before calculating feature quantities, allowing for accurate matching by comparing enhanced features with higher-resolution images.
Enables reliable matching of images with varying resolutions, even when high-resolution images are not initially available, by improving the resolution of one or both images to ensure accurate feature extraction and comparison.
Smart Images

Figure 2026529004000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to an image matching apparatus, an image matching method, a training apparatus, a training method, and a non - transient computer - readable medium.
Background Art
[0002] Computer systems for performing cross - view matching between ground and aerial views (matching between ground images and aerial images) have been developed. For example, Non - Patent Document 1 discloses a system including a set of CNNs (Convolutional Neural Networks) for matching ground images and aerial images. Specifically, one of the CNNs acquires a set of a ground image and an orientation map indicating the orientation (azimuth and altitude) for each position imaged on the ground image, and extracts features therefrom. The other CNN acquires a set of an aerial image and an orientation map indicating the orientation (azimuth and distance) for each position imaged on the aerial image, and extracts features therefrom. Then, the system determines whether the ground image matches the aerial image based on the extracted features.
Prior Art Documents
Non - Patent Documents
[0003]
Non - Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Non-patent document 1 does not mention image resolution. This disclosure aims to provide a novel technique for determining whether ground images and aerial images match each other. [Means for solving the problem]
[0005] This disclosure provides an image matching device comprising at least one memory configured to store instructions and at least one processor configured to execute instructions, wherein the instructions include acquiring a first view image and a second view image, calculating feature quantities for the first view image, calculating feature quantities for the second view image, and determining whether the first view image and the second view image match based on the feature quantities of the first view image and the feature quantities of the second view image. The calculation of features for the first view image includes improving the resolution of the first view image and calculating the features of the improved-resolution first view image as features for the first view image, in order to generate a first view image with improved resolution. The first view image and the second view image are either a ground image and an aerial image, respectively, or the first view image and the second view image are an aerial image and a ground image, respectively.
[0006] This disclosure further provides a computer-based image matching method. The image matching method includes the steps of: acquiring a first view image and a second view image; calculating feature quantities for the first view image; calculating feature quantities for the second view image; and determining whether the first view image and the second view image match based on the feature quantities for the first view image and the feature quantities for the second view image. The calculation of features for the first view image includes the steps of improving the resolution of the first view image to generate a resolution-enhanced first view image, and calculating the features of the resolution-enhanced first view image as features of the first view image. The first view image and the second view image are either a ground image and an aerial image, respectively, or the first view image and the second view image are an aerial image and a ground image, respectively.
[0007] The Disclosure further provides a non-temporary computer-readable medium for storing a program, the program enabling the computer to perform the following actions: acquire a first view image and a second view image; calculate feature quantities for the first view image and the second view image; and determine whether the first view image and the second view image match based on the feature quantities for the first and second view images. The calculation of features for the first view image includes improving the resolution of the first view image to generate a resolution-enhanced first view image, and calculating the features of the resolution-enhanced first view image as features of the first view image. The first view image and the second view image are either a ground image and an aerial image, respectively, or the first view image and the second view image are an aerial image and a ground image, respectively.
[0008] The disclosure further provides a training device comprising at least one memory configured to store instructions and at least one processor configured to execute instructions, wherein the instructions are to acquire a target image and to train one or more models. Training one or more models involves reducing the resolution of acquired target images to generate reduced-resolution target images, inputting the reduced-resolution target images into a resolution enhancement model to generate improved-resolution target images, calculating a loss using the acquired target images and improved-resolution target images, and updating the resolution enhancement model based on the loss. The target image is either a ground image or an aerial image.
[0009] This disclosure further provides a training method performed by a computer. The training method includes the steps of acquiring a target image and training one or more models. Training one or more models includes the steps of: reducing the resolution of an acquired target image to generate a reduced-resolution target image; inputting the reduced-resolution target image into a resolution enhancement model to generate a re-enhanced target image; calculating a loss using the acquired target image and the re-enhanced target image; and updating the resolution enhancement model based on the loss. The target image is either a ground image or an aerial image.
[0010] This disclosure further provides a non-temporary computer-readable medium for storing a program, which the computer uses to perform tasks such as acquiring a target image and training one or more models. Training one or more models involves reducing the resolution of acquired target images to generate reduced-resolution target images, inputting the reduced-resolution target images into a resolution enhancement model to generate improved-resolution target images, calculating a loss using the acquired target images and improved-resolution target images, and updating the resolution enhancement model based on the loss. The target image is either a ground image or an aerial image.
[0011] This disclosure provides a novel technology for determining whether ground images and aerial images match each other. [Brief explanation of the drawing]
[0012] [Figure 1] This is a schematic diagram of an image matching device. [Figure 2] This figure shows examples of ground images and aerial images. [Figure 3]It is a schematic diagram of an image matching device that handles a second view image having a first-level resolution. [Figure 4] It is a block diagram showing an example of the functional configuration of an image matching device. [Figure 5] It is a block diagram showing an example of the hardware configuration of a computer 1000 that realizes an image matching device. [Figure 6] It is a flowchart for explaining an example of the processing flow of an image matching device. [Figure 7] It is a diagram showing a geographical location estimation system including an image matching device. [Figure 8] It is a diagram showing an example of the configuration of a first feature extraction unit. [Figure 9] It is a diagram showing a first example of the configuration of a second feature extraction unit. [Figure 10] It is a diagram showing a second example of the configuration of a second feature extraction unit. [Figure 11] It is a diagram showing an example of the functional configuration of a training device. [Figure 12] It is a flowchart showing an example of the processing flow of a training device. [Figure 13] It is a diagram showing a first example of a training method for a first feature extraction unit. [Figure 14] It is a diagram showing an image matching device that handles a first view image having a second-level resolution and a second view image having a second-level resolution. [Figure 15] It is a diagram showing additional training performed on a feature extraction model. [Figure 16] It is a diagram showing an example of a method for training a resolution improvement model separately from a feature extraction model. [Figure 17] It is a diagram showing a first example of a training method for a second feature extraction unit. [Figure 18] It is a diagram showing additional training performed on a feature extraction model. [Figure 19] It is a diagram showing an example of a method for training a resolution improvement model separately from a feature extraction model.
Mode for Carrying Out the Invention
[0013] Hereinafter, exemplary embodiments of the present disclosure will be described with reference to the drawings. The same elements are assigned the same reference numerals throughout the drawings, and redundant descriptions are omitted where necessary. In addition, certain information (such as a predetermined value or threshold) is pre-stored in a storage device accessible to the computer using that information, unless otherwise specified.
[0014] Embodiment 1 <Overview> Figure 1 shows an overview of the image matching device 2000. The image matching device 2000 functions as a discriminator that performs matching between ground images and aerial images (so-called ground-to-air cross-view matching). Figure 2 shows an example of ground image 10 and aerial image 15.
[0015] The ground image 10 is a digital image that includes a ground view of a certain location, and is, for example, an RGB or grayscale image of the ground landscape. The ground image is generated by a ground camera, for example, one carried by a pedestrian or mounted on a vehicle. The ground image may be panoramic (with a 360-degree view) or limited (less than 360 degrees).
[0016] Aerial image 15 is a digital image containing a plan view of a certain location, such as an RGB or grayscale image of an aerial landscape. For example, aerial images are generated by aerial cameras mounted on drones, airplanes, or satellites.
[0017] The image matching device 2000 acquires a first view image 20 and a second view image 30 as images to be compared with each other. One is a ground image 10, and the other is an aerial image 15. In other words, if the image matching device 2000 is configured to acquire the ground image 10 as the first view image 20, the image matching device 2000 is configured to acquire the aerial image 15 as the second view image 30. On the other hand, if the image matching device 2000 is configured to acquire the aerial image 15 as the first view image 20, the image matching device 2000 is configured to acquire the ground image 10 as the second view image 30.
[0018] The image matching device 2000 is used in a situation where the first view image 20 has a first level of resolution, and the second view image 30 has either a first level or a second level of resolution. The first level of resolution may include only a specific image resolution, or it may include image resolutions within a specific range. The same applies to the second level of resolution.
[0019] The first and second level resolutions are defined such that the second level resolution is higher than the first level resolution. It is assumed that resolution is quantified such that the higher the image resolution, the smaller the value representing that image resolution. An example of this is when the unit "cm / pixel" is used to represent the image resolution.
[0020] In this case, if the first-level and second-level resolutions are defined by specific real numbers r1 and r2, respectively, then the numbers r1 and r2 satisfy the condition "r1 > r2". For example, the first-level resolution is defined as "50 cm / pixel" and the second-level resolution is defined as "25 cm / pixel". If the first-level and second-level resolutions are defined by ranges R1 and R2, respectively, then the ranges R1 and R2 satisfy the condition "the upper limit of range R2 is less than the upper limit of range R1". For example, the first-level resolution is defined as the range (45 [cm / pixel], 55 [cm / pixel]) and the second-level resolution is defined as the range (20 [cm / pixel], 30 [cm / pixel]).
[0021] Furthermore, the resolution levels of the first view image and the second view image may be defined separately from each other. This means that the first level resolution of the first view image is not necessarily equivalent to the first level resolution of the second view image. Similarly, the second level resolution of the first view image is not necessarily equivalent to the second level resolution of the second view image.
[0022] Since the resolution of the first view image 20 is at the first level, the image matching device 2000 performs a resolution enhancement on the first view image 20 to raise its resolution to the second level. The first view image obtained as a result of the resolution enhancement is called the "resolution-enhanced first view image 40". The image matching device 2000 calculates the feature quantities of the resolution-enhanced first view image 40 as the feature quantities of the first view image 20.
[0023] With respect to the second view image 30, the image matching device 2000 does not improve the resolution of the second view image 30 if its resolution is at the second level. In this case, the image matching device 2000 directly calculates the feature quantities of the second view image 30 from there. This case is shown in Figure 1.
[0024] On the other hand, if the resolution of the second view image 30 is at the first level, the image matching device 2000 improves the resolution of the second view image 30 to the second level. Figure 3 shows an overview of the image matching device 2000 handling the second view image 30 which has a first-level resolution. The second view image obtained as a result of the resolution improvement is called the "resolution-improved second view image 50". The image matching device 2000 calculates the feature quantities of the resolution-improved second view image 50 as feature quantities of the second view image 30.
[0025] <Examples of effects and benefits> As described above, the image matching device 2000 acquires the first view image 20 and the second view image 30, and determines whether the first view image 20 and the second view image 30 match by comparing their feature quantities. In order to calculate the feature quantities of the first view image 20, the first view image 20 is converted into a resolution-enhanced first view image 40 by performing a resolution enhancement on the first view image 20. The feature quantities of the resolution-enhanced first view image 40 are used as the feature quantities of the first view image 20.
[0026] Since the image matching device 2000 has a mechanism to improve the resolution of the first view image 20 before calculating its features, the image matching device 2000 can handle a first view image that is not high enough in resolution to determine whether the first view image and the second view image 30 match. Furthermore, if the image matching device 2000 has a mechanism to improve the resolution of the second view image 30 before calculating its features, the image matching device 2000 can handle a second view image that is not sufficiently high in resolution.
[0027] The image matching device 2000 is useful in various situations. One such situation is when it is difficult to reliably acquire a high-resolution first view image. Assume that the first view image 20 is a ground image provided by the user. The user can take a picture of their surroundings using their smartphone, and that picture may be provided as the first view image 20.
[0028] In this case, the resolution of the first view image 20 depends on the performance of the smartphone. Therefore, there may be smartphones that provide the image matching device 2000 with a low-resolution first view image. Even in this case, the image matching device 2000 can still determine whether the first view image 20 and the second view image 30 match.
[0029] The image matching device 2000 will be described in more detail below.
[0030] <Example of functional configuration> Figure 4 is a block diagram showing an example of the functional configuration of the image matching device 2000. The image matching device 2000 includes an acquisition unit 2020, a first feature extraction unit 2040, a second feature extraction unit 2060, and a determination unit 2080. The acquisition unit 2020 acquires the first view image 20 and the second view image 30. The first feature extraction unit 2040 calculates the feature quantities of the first view image 20. Specifically, the first feature extraction unit 2040 improves the resolution of the first view image 20 in order to generate a resolution-enhanced first view image 40. Then, the first feature extraction unit 2040 calculates the feature quantities of the resolution-enhanced first view image 40 as the feature quantities of the first view image 20.
[0031] The second feature extraction unit 2060 calculates the feature quantities of the second view image 30. The determination unit 2080 determines whether the first view image 20 and the second view image 30 match based on the feature quantities of the first view image 20 and the feature quantities of the second view image 30.
[0032] Furthermore, if the image matching device 2000 is configured to acquire a second view image with a first level of resolution, the second feature extraction unit 2060 is configured to improve the resolution of the second view image 30 in order to generate a second view image 50 with improved resolution. The second feature extraction unit 2060 is also configured to calculate the feature quantities of the second view image 50 with improved resolution as the feature quantities of the second view image 30.
[0033] <Example hardware configuration> The image matching device 2000 may be implemented by one or more computers. Each of these computers may be a dedicated computer manufactured specifically for implementing the image matching device 2000, or it may be a general-purpose computer such as a personal computer (PC), server machine, or mobile device.
[0034] The image matching device 2000 may be implemented by installing an application on one or more computers. The application is implemented by a program that causes one or more computers to function as the image matching device 2000. In other words, the program is an implementation of the functional part of the image matching device 2000.
[0035] Figure 5 is a block diagram showing an example of the hardware configuration of a computer 1000 that implements an image matching device 2000. In Figure 5, the computer 1000 includes a bus 1020, a processor 1040, memory 1060, a storage device 1080, an input / output (I / O) interface 1100, and a network interface 1120.
[0036] Bus 1020 is a data transmission channel for the processor 1040, memory 1060, storage device 1080, I / O interface 1100, and network interface 1120 to send and receive data to and from each other. The processor 1040 is a processor such as a CPU (Central Processing Unit), GPU (Graphics Processing Unit), FPGA (Field-Programmable Gate Array), or DSP (Digital Signal Processor). Memory 1060 is a primary memory component such as RAM (Random Access Memory) or ROM (Read Only Memory). Storage device 1080 is a secondary storage component such as a hard disk, SSD (Solid State Drive), or memory card. I / O interface 1100 is an interface between the computer 1000 and peripheral devices such as a keyboard, mouse, or display device. Network interface 1120 is an interface between the computer 1000 and a network. The network may be a LAN (Local Area Network) or a WAN (Wide Area Network).
[0037] The memory device 1080 can store the program described above. The processor 1040 reads the program from the memory device 1080, executes the program, and implements the various functional units of the image matching device 2000.
[0038] The hardware configuration of computer 1000 is not limited to that shown in Figure 5. For example, as described above, the image matching device 2000 may be implemented by multiple computers. In this case, these computers may be interconnected via a network.
[0039] <Processing flow> Figure 6 is a flowchart illustrating an example of the processing flow performed by the image matching device 2000. The acquisition unit 2020 acquires the first view image 20 and the second view image 30 (S102). The first feature extraction unit 2040 calculates the feature quantities of the first view image 20 (S104). The second feature extraction unit 2060 calculates the feature quantities of the second view image 30 (S106). The determination unit 2080 determines whether the first view image 20 and the second view image 30 match based on the feature quantities of the first view image 20 and the feature quantities of the second view image 30 (S108).
[0040] Note that the processing flow performed by the image matching device 2000 is not limited to that shown in Figure 6. For example, the calculation of the feature quantities of the first view image 20 (i.e., step S104) and the calculation of the feature quantities of the second view image 30 (i.e., step S106) may be performed in parallel, or in the reverse order of the order shown in Figure 6.
[0041] <Example applications of the image matching device 2000> The image matching device 2000 has a variety of possible applications. For example, the image matching device 2000 may be used as part of a system for geolocation of images (hereinafter referred to as a geolocation system). Image geolocation is a technique for determining the location where an input image was captured. The geolocation system 500 can be implemented by one or more arbitrary computers, as shown in Figure 5. It should be noted that the geolocation system is merely one example of an application for the image matching device 2000, and its applications are not limited to use in a geolocation system.
[0042] Figure 7 shows a geolocation system 500 including an image matching device 2000. The geolocation system 500 includes the image matching device 2000 and a location database 600. The location database 600 includes multiple aerial images, each with location information attached. An example of location information may be the GPS (Global Positioning System) coordinates of a location captured on the center of the corresponding aerial image.
[0043] The geolocation system 500 receives a query containing a ground image from a client (for example, a user terminal). The geolocation system 500 then searches the location database 600 for an aerial image that matches the ground image in the received query, thereby determining the location where the ground image was taken. Specifically, until an aerial image matching the ground image in the query is detected, the geolocation system 500 repeatedly performs the following: retrieving one of the aerial images from the location database 600, inputting the ground image obtained from the query and the aerial image obtained from the location database 600 into the image matching device 2000, and determining whether the output of the image matching device 2000 indicates that the ground image and the aerial image match. As described above, the image matching device 2000 may be configured to treat the ground image and the aerial image as the first view image 20 and the second view image 30, respectively, or it may be configured to treat the ground image and the aerial image as the second view image 30 and the first view image 20, respectively.
[0044] By repeatedly performing the above process, the geolocation system 500 can find an aerial image that includes the location where the ground image was taken. Since the detected aerial image is associated with location information such as GPS coordinates, the geolocation system 500 can determine that the location where the ground image was taken is the location indicated by the location information associated with the aerial image that matches the ground image.
[0045] Furthermore, ground images and aerial images are used in the reverse order in the geolocation system 500. In this case, the location database 600 stores multiple ground images with location information attached. The geolocation system 500 receives a query that includes an aerial image, searches the location database 600 for a ground image that matches the aerial image in the query, and thereby determines the location of the place captured on the aerial image.
[0046] <Data acquisition: S102> The acquisition unit 2020 acquires the first view image 20 and the second view image 30 (S102). There are various ways to acquire this data. In some implementations, the acquisition unit 2020 can receive the first view image 20, the second view image 30, or both, transmitted from another computer. In other embodiments, the acquisition unit 2020 can acquire the first view image 20, the second view image 30, or both, from a storage device accessed by the acquisition unit 2020.
[0047] The first view image 20 and the second view image 30 may be acquired in the same way or in different ways. For example, the acquisition unit 2020 may receive the first view image 20 from another computer while acquiring the second view image 30 from a storage device, or vice versa.
[0048] <Calculation of features of the first view image: S104> The first feature extraction unit 2040 calculates the feature quantities of the first view image 20 (S104). As described above, the resolution of the first view image 20 is increased in order to generate the resolution-enhanced first view image 40. Next, the first feature extraction unit 2040 calculates the feature quantities of the resolution-enhanced first view image 40 as the feature quantities of the first view image 20.
[0049] Figure 8 shows an example configuration of the first feature extraction unit 2040. The first feature extraction unit 2040 can include two machine learning-based models (such as neural networks) called the "resolution enhancement model 100" and the "feature extraction model 110".
[0050] The resolution enhancement model 100 is configured to take an image as input and output another image of the same size as the input image. In addition, the resolution enhancement model 100 is pre-trained to enhance the resolution of the input first view image to a second level in response to the input first view image having a first level of resolution, thereby generating a first view image having a second level of resolution. The training method for the resolution enhancement model 100 will be described later.
[0051] The feature extraction model 110 is configured to take an image as input and output values (such as vectors or tensors) calculated based on the input image. In addition, the feature extraction model 110 is pre-trained to compute the features of a first view image having a second level of resolution in response to the input image. The training method for the first feature extraction model 110 will be described later.
[0052] The first feature extraction unit 2040 inputs the first view image 20 acquired by the acquisition unit 2020 to the resolution enhancement model 100. As a result, the resolution enhancement model 100 outputs the resolution-enhanced first view image 40. Next, the resolution-enhanced first view image 40 is supplied to the feature extraction model 110. As a result, the feature extraction model 110 outputs the feature quantities of the resolution-enhanced first view image 40. The feature quantities of the resolution-enhanced first view image 40 are treated as feature quantities of the first view image 20 by the determination unit 2080.
[0053] <Calculation of features in the second view image: S106> The second feature extraction unit 2060 calculates the features of the second view image 30 (S106). Figure 9 shows a first example of the configuration of the second feature extraction unit 2060. In the example shown in Figure 9, the second view image 30 is assumed to have a second level of resolution. Therefore, the second feature extraction unit 2060 does not have the function of improving the second view image 30. Specifically, the second feature extraction model 2060 may include a machine learning-based model (such as a neural network) called the "feature extraction model 120".
[0054] The feature extraction model 120 is configured to take an image as input and output values (such as vectors or tensors) calculated based on the input image. In addition, the feature extraction model 120 is pre-trained to compute the features of a second view image having a second level of resolution in response to the input image. The training method for the second feature extraction model 120 will be described later.
[0055] The second feature extraction unit 2060 inputs the second view image 30 acquired by the acquisition unit 2020 into the feature extraction model 120, thereby acquiring the feature quantities of the second view image 30 that the feature extraction model 120 outputs.
[0056] The second feature extraction unit 2060 further includes a function to improve the resolution of the second view image 30 to a second level in order to generate a resolution-enhanced second view image 50, if the second view image 30 has a first level of resolution. Figure 10 shows a second example of the configuration of the second feature extraction unit 2060. In the example shown in Figure 10, the second view image 30 is assumed to have a first level of resolution. In addition to the feature extraction model 120, the second feature extraction unit 2060 includes a resolution enhancement model 130.
[0057] The resolution enhancement model 130 is configured to take an image as input and output another image of the same size as the input image. In addition, the resolution enhancement model 130 is pre-trained to enhance the resolution of the input second view image to a second level in response to the input second view image having a first level of resolution, thereby generating a second view image having a second level of resolution. The training method for the resolution enhancement model 130 will be described later.
[0058] When the image matching device 2000 is configured to handle a second view image 30 having a first level of resolution, the second feature extraction unit 2060 inputs the second view image 30 acquired by the acquisition unit 2020 to the resolution enhancement model 130. As a result, the resolution enhancement model 130 outputs a resolution-enhanced second view image 50. Next, the resolution-enhanced second view image 50 is supplied to the feature extraction model 120. As a result, the feature extraction model 120 outputs the feature quantities of the resolution-enhanced second view image 50. The feature quantities of the resolution-enhanced second view image 50 are treated as feature quantities of the second view image 30 by the determination unit 2080.
[0059] <Matching: S108> The determination unit 2080 determines whether the first view image 20 and the second view image 30 match (S108). Specifically, the determination unit 2080 makes the determination by comparing the feature quantities of the first view image 20 and the feature quantities of the second view image 30.
[0060] The determination unit 2080 can calculate a similarity score representing the similarity between the features of the first view image 20 and the features of the second view image 30. There are various metrics for quantifying the similarity between features, and any one of them can be used to calculate the similarity score. For example, the similarity score may be calculated as one of various types of distance (e.g., L2 distance), correlation, cosine similarity, or neural network (NN) based similarity between the features of the first view image 20 and the features of the second view image 30. NN-based similarity is similarity calculated by a neural network trained to calculate similarity between two input data (in this disclosure, the features of the first view image 20 and the features of the second view image 30).
[0061] The determination unit 2080 determines whether the first view image 20 and the second view image 30 match based on the calculated similarity. Conceptually, the higher the similarity between the features of the first view image 20 and the features of the second view image 30, the higher the probability that the first view image 20 and the second view image 30 match. For example, the determination unit 2080 determines whether the similarity score is above a predetermined threshold. If the similarity score is above the predetermined threshold, the determination unit 2080 determines that the first view image 20 and the second view image 30 match. On the other hand, if the similarity score is below the predetermined threshold, the determination unit 2080 determines that the first view image 20 and the second view image 30 do not match.
[0062] In the above case, the higher the similarity between features, the higher the similarity score. Therefore, if a metric is used in which the value calculated for the compared features decreases as the similarity between the compared features increases (for example, distance), the similarity score can be defined as the reciprocal of the value calculated for the compared features.
[0063] In another example, if the similarity score decreases as the similarity between the compared features increases, the determination unit 2080 can determine whether the similarity score is below a predetermined threshold. If the similarity score is below the predetermined threshold, the determination unit 2080 determines that the first view image 20 and the second view image 30 are a match. On the other hand, if the similarity score is greater than the predetermined threshold, the determination unit 2080 determines that the first view image 20 and the second view image 30 are not a match.
[0064] <Output from image matching device> The image matching device 2000 can output information regarding the determination result (hereinafter referred to as output information). For example, the output information can indicate whether the first view image 20 and the second view image 30 match. In addition, as explained with reference to Figure 7, the output information may further include location information indicating the location where the queried image was captured. The queried image is either the first view image 20 or the second view image 30. In other words, the queried image is either a ground image or an aerial image.
[0065] There are various ways to output output information. For example, the image matching device 2000 can store the output information in a storage device. In another example, the image matching device 2000 can output the output information to a display device, which then displays the contents of the output information. In yet another example, the image matching device 2000 can output the output information to another computer, such as the computer included in the geolocation system 500 shown in Figure 7.
[0066] <Training the model> As described above, machine learning-based models can be used to compute the features of the first view image 20 and the second view image 30. These models are trained before being used by the image matching device. Hereafter, the device that trains these models will be referred to as the "training device".
[0067] Figure 11 shows an example of the functional configuration of the training device 3000. The training device 3000 includes an acquisition unit 3020, a calculation unit 3040, and an update unit 3060. The acquisition unit 3020 acquires training data. The calculation unit 3020 applies the training data to the model to be trained and calculates the loss based on the data output by the model. The update unit 3060 updates the model based on the loss.
[0068] Furthermore, the model can be updated by updating its trainable parameters based on the loss. If the model is a neural network, the trainable parameters can include weights assigned to edges and biases.
[0069] The training device 3000 may have a hardware configuration similar to that of the image matching device 2000. For example, the hardware configuration of the training device 3000 may be similar to that of the image matching device 2000, as shown in Figure 5. However, the storage device of the training device 3000 includes a program that implements the functions of the training device 3000.
[0070] Figure 12 is a flowchart illustrating an example of the processing flow performed by the training device 3000. The acquisition unit 3020 acquires training data (S202). The calculation unit 3040 inputs the training data to the model to be trained (S204). The calculation unit 3040 calculates the loss based on the data output by the model (S206). The update unit 3060 updates the model based on the loss (S208).
[0071] The process shown in Figure 12 is repeated until the model is sufficiently trained.
[0072] The following describes an example of a model training method. Specifically, first, an example of a model training method in the first feature extraction unit 2040 will be described. Next, an example of a model training method in the second feature extraction unit 2060 will be described.
[0073] <<Training of the first feature extraction unit 2040>> As shown in Figure 8, the first feature extraction unit 2040 may include a resolution enhancement model 100 and a feature extraction model 110. Figure 13 shows an example of a first training method for the first feature extraction unit 2040. First, the acquisition unit 3020 acquires a first view image 150 as training data. The first view image 150 is a first view image having a second level of resolution.
[0074] Next, the calculation unit 3040 performs a resolution reduction on the first view image 150 in order to generate a resolution-reduced first view image 160, which is a first view image with a first level of resolution. For example, the calculation unit 3040 performs downsampling on the first view image 150 and resizes the resulting image to the same size as the first view image 150. In this way, the calculation unit 3040 can obtain a resolution-reduced first view image 160 that is the same size as the first view image 150 but has a lower resolution than the first view image 150.
[0075] The calculation unit 3040 inputs the resolution-reduced first view image 160 into the resolution-enhancing model 100 to obtain the resolution-enhanced first view image 170. The resolution-enhanced first view image 170 is assumed to be equivalent to the first view image 150 if the resolution-enhancing model 100 has already been sufficiently trained.
[0076] Next, the first view image 170 with improved resolution is supplied to the feature extraction model 110. As a result, the feature quantities of the first view image 170 with improved resolution are output by the feature extraction model 110.
[0077] The calculation unit 3040 also calculates the features of the first view image 150. Specifically, the calculation unit 3040 inputs the first view image 150 into a pre-trained feature extraction model 180, which is a machine learning-based model that has been pre-trained to calculate the features of the first view image at a second level of resolution.
[0078] The features of the resolution-enhanced first view image 170 and the features of the first view image 150 are assumed to be equivalent if the resolution enhancement model 100 and the feature extraction model 110 have already been sufficiently trained. Therefore, the training device 3000 calculates a loss that represents the degree of difference between the features of the first view image 150 and the features of the resolution-enhanced first view image 170.
[0079] The update unit 3060 updates the resolution enhancement model 100 and the feature extraction model 110 using the calculated loss. Specifically, the update unit 3060 updates the trainable parameters of the resolution enhancement model 100 and the trainable parameters of the feature extraction model 110 based on the loss.
[0080] According to the training device 3000 shown in Figure 13, the resolution enhancement model 100 can be trained so that it can accurately improve the resolution of the first view image from the first level to the second level. In addition, the feature extraction model 110 can be trained so that it can accurately extract features from the first view image at the second level of resolution.
[0081] As described above, the pre-trained feature extraction model 180 is pre-trained. For example, the pre-trained feature extraction model 180 is trained to be used as part of another image matching device, which is configured to acquire a first view image with a second level of resolution and a second view image with a second level of resolution, and to determine whether the acquired first view image and the acquired second view image match each other.
[0082] Figure 14 shows an image matching device 400 that handles a first view image with a second-level resolution and a second view image with a second-level resolution. The image matching device 400 acquires a first view image 410 and a second view image 420, both with second-level resolution. The image matching device 400 has a feature extraction model 430 and a feature extraction model 440. The feature extraction model 430 takes the first view image 410 as input and calculates the feature quantities of the first view image 410. The feature extraction model 440 takes the second view image 420 as input and calculates the feature quantities of the second view image 420. The image matching device 400 determines whether the first view image 410 and the second view image 420 match by comparing their feature quantities.
[0083] The feature extraction model 430 performs the same function as the pre-trained feature extraction model 180, which is to compute the features of the first view image with a resolution of level 2. Therefore, the feature extraction model 430 can be employed in the training device 3000 and used as the pre-trained feature extraction model 180.
[0084] In some embodiments, after completing the training of the first feature extraction unit 2040 shown in Figure 13, the training device 3000 can perform additional training on the feature extraction model 110. Figure 15 shows the additional training performed on the feature extraction model 110. The acquisition unit 3020 acquires the first view image 200 and the second view image 210 as training data. The first view image 200 is a first view image having a first level of resolution, and the second view image 210 is a second view image having a second level of resolution.
[0085] The calculation unit 3040 inputs the first view image 200 to the resolution enhancement model 100, which has been trained in the manner shown in Figure 13, to obtain a resolution-enhanced first view image 220. The resolution-enhanced first view image 220 is then supplied to the feature extraction model 110, which has been trained in the manner shown in Figure 13. This allows the feature quantities of the resolution-enhanced first view image 220 to be obtained.
[0086] Furthermore, the calculation unit 3040 obtains the features of the second view image 210 by inputting the second view image 210 into the pre-trained feature extraction model 230. The feature extraction model 440 shown in Figure 14 can be used as the pre-trained feature extraction model 230.
[0087] The calculation unit 3040 calculates a loss representing the degree of difference between the features of the first view image 200 and the features of the second view image 210, and updates the feature extraction model 110 based on the loss. The feature extraction model 110 can be trained to have a small loss when it is assumed that the first view image 200 and the second view image 210 are identical. On the other hand, the feature extraction model 110 can be trained to have a large loss when it is assumed that the first view image 200 and the second view image 210 are not identical.
[0088] More specifically, the training device 3000 can use a set of first view images 200, positive examples of second view images 210, and negative examples of second view images 210 as training data. Positive examples of second view images 210 are second view images 210 that are assumed to match the first view image 200. On the other hand, negative examples of second view images 210 are second view images 210 that are assumed not to match the first view image 200. In this case, the training device 3000 can calculate the triplet loss based on the features of the first view image 200, the features of the positive examples of second view images 210, and the features of the negative examples of second view images 210, and update the feature extraction model 110 based on the triplet loss.
[0089] According to the training shown in Figure 15, the feature extraction model 110 can be trained to accurately calculate the features of the first view image from the perspective of matching the first view image with the second view image.
[0090] In some embodiments, the resolution enhancement model 100 may be trained separately from the feature extraction model 110. Figure 16 shows an example of how to train the resolution enhancement model 100 separately from the feature extraction model 110.
[0091] The acquisition unit 3020 acquires the first view image 240 as training data. The first view image 240 is a first view image having a second level of resolution. The calculation unit 3040 performs a resolution reduction on the first view image 240 in order to generate a resolution-reduced first view image 250, which is a first view image with a first level of resolution. The calculation unit 3040 inputs the resolution-reduced first view image 250 into the resolution enhancement model 100 to obtain a resolution-enhanced first view image 260.
[0092] The first view image 260 with improved resolution is assumed to be equivalent to the first view image 240, provided that the resolution enhancement model 100 has already been sufficiently trained. Therefore, the calculation unit 3040 calculates a loss representing the degree of difference between the first view image 240 and the first view image 260 with improved resolution, and updates the resolution enhancement model 100 based on the loss.
[0093] <<Training of the second feature extraction unit 2060>> As shown in Figure 9, when the resolution of the second view image 30 is at the second level, the second feature extraction unit 2060 may include the feature extraction model 120 but not the resolution enhancement model 130. In this case, the feature extraction model 440 shown in Figure 14 can be used as the feature extraction model 120.
[0094] On the other hand, the second feature extraction unit 2060 may include a resolution enhancement model 130 and a feature extraction model 120 if the resolution of the second view image 30 is at the first level. In this case, the second extraction model 2060 can be trained in the same way as the first feature extraction unit 2040.
[0095] Figure 17 shows an example of the first training method for the second feature extraction unit 2060. The acquisition unit 3020 acquires the second view image 270 as training data. The second view image 270 is a second view image having a second level of resolution.
[0096] The calculation unit 3040 reduces the resolution of the second view image 270 in order to generate a resolution-reduced second view image 280, which is a second view image with a first level of resolution. The calculation unit 3040 inputs the resolution-reduced second view image 280 into the resolution enhancement model 130 to obtain a further resolution-enhanced second view image 290.
[0097] The second view image 290 with improved resolution is supplied to the feature extraction model 120. As a result, the feature quantities of the second view image 290 with improved resolution are output by the feature extraction model 120.
[0098] Furthermore, the calculation unit 3040 calculates the features of the second view image 270. Specifically, the calculation unit 3040 inputs the second view image 270 into a pre-trained feature extraction model 300, which is a machine learning-based model that has been pre-trained to calculate the features of the second view image at a second level of resolution. As the pre-trained feature extraction model 300, the feature extraction model 340 shown in Figure 14 can be used.
[0099] The feature quantities of the first view image 170 with improved resolution and the feature quantities of the first view image 150 are assumed to be equivalent if the resolution improvement model 130 and the feature extraction model 120 have already been sufficiently trained. Therefore, the calculation unit 3040 calculates a loss that represents the degree of difference between the feature quantities of the second view image 270 and the feature quantities of the second view image 290 with improved resolution.
[0100] The update unit 3060 updates the resolution enhancement model 130 and the feature extraction model 120 using the calculated loss. Specifically, the update unit 3060 updates the trainable parameters of the resolution enhancement model 130 and the trainable parameters of the feature extraction model 120 based on the loss.
[0101] According to the training device 3000 shown in Figure 17, the training device 3000 can train the resolution enhancement model 130 so that it can accurately improve the resolution of the second view image from the first level to the second level. In addition, the feature extraction model 120 can be trained so that it can accurately extract features from the second view image at the second level of resolution.
[0102] In some embodiments, after completing the training of the second feature extraction unit 2060 shown in Figure 17, the training device 3000 can perform additional training on the feature extraction model 120. Figure 18 shows the additional training performed on the feature extraction model 120. The acquisition unit 3020 acquires the first view image 310 and the second view image 320 as training data. The first view image 310 is a first view image having a second level of resolution, and the second view image 320 is a second view image having a first level of resolution.
[0103] The calculation unit 3040 inputs the second view image 320 into the resolution enhancement model 130, which has been trained in the manner shown in Figure 17, to obtain a resolution-enhanced second view image 330. The resolution-enhanced second view image 330 is then supplied to the feature extraction model 120, which has been trained in the manner shown in Figure 17. This allows the feature quantities of the resolution-enhanced second view image 330 to be obtained.
[0104] Furthermore, the calculation unit 3040 inputs the first view image 310 into the pre-trained feature extraction model 340 to obtain the features of the first view image 310. Note that the feature extraction model 430 shown in Figure 14 can be used as the pre-trained feature extraction model 340.
[0105] The calculation unit 3040 calculates a loss representing the degree of difference between the features of the first view image 310 and the features of the second view image 320, and updates the feature extraction model 120 based on the loss. The feature extraction model 120 can be trained to produce a small loss when it is assumed that the first view image 310 and the second view image 320 are identical. On the other hand, the feature extraction model 120 can be trained to produce a large loss when it is assumed that the first view image 310 and the second view image 320 are not identical.
[0106] More specifically, the training device 3000 can use a set of first view images 310, positive examples of second view images 320, and negative examples of second view images 320 as training data. Positive examples of second view images 320 are second view images 320 that are assumed to match the first view image 310. On the other hand, negative examples of second view images 320 are second view images 320 that are assumed not to match the first view image 310. In this case, the training device 3000 can calculate the triplet loss based on the features of the first view image 310, the features of the positive examples of second view images 320, and the features of the negative examples of second view images 320, and update the feature extraction model 120 based on the triplet loss.
[0107] According to the training shown in Figure 18, the feature extraction model 120 can be trained to accurately calculate the features of the second view image in terms of matching the first view image with the second view image.
[0108] In some embodiments, the resolution enhancement model 130 may be trained separately from the feature extraction model 120. Figure 19 shows an example of how to train the resolution enhancement model 130 separately from the feature extraction model 120.
[0109] The acquisition unit 3020 acquires the second view image 350 as training data. The second view image 350 is a second view image having a second level of resolution. The calculation unit 3040 performs a resolution reduction on the second view image 350 in order to generate a resolution-reduced second view image 360, which is a second view image with a first level of resolution. The calculation unit 3040 inputs the resolution-reduced second view image 360 into the resolution enhancement model 130 to obtain a further resolution-enhanced second view image 370.
[0110] The second view image 370 with improved resolution is assumed to be equivalent to the second view image 350, provided that the resolution enhancement model 130 has already been sufficiently trained. Therefore, the calculation unit 3040 calculates a loss representing the degree of difference between the second view image 350 and the second view image 370 with improved resolution, and updates the resolution enhancement model 130 based on this loss.
[0111] Programs can be stored and provided to a computer using various types of non-transitory computer-readable media. Non-transitory computer-readable media include various types of tangible storage media. Examples of non-transitory computer-readable media include magnetic recording media (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical recording media (e.g., magneto-optical disks), CD-ROMs, CD-Rs, CD-R / Ws, and semiconductor memory (e.g., mask ROMs, PROMs (Programmable ROMs), EPROMs (Erasable PROMs), flash ROMs, RAMs). Programs may also be provided to a computer using various types of transient computer-readable media. Examples of transient computer-readable media include electrical signals, optical signals, and electromagnetic waves. Transitory computer-readable media can be supplied to a computer via wired communication channels such as electric wires and optical fibers, or via wireless communication channels.
[0112] While this disclosure has been described with reference to embodiments, it is not limited to the embodiments described above. Various modifications to the structure and details of this disclosure are possible, as can be understood by those skilled in the art within the scope of this disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.
[0113] Each drawing is merely illustrative to illustrate one or more embodiments. Each drawing may be associated with one or more other embodiments rather than with only one specific embodiment. As those skilled in the art will understand, various features or steps described with reference to any one drawing can be combined with features or steps shown in one or more other drawings, for example, to create embodiments not explicitly shown or described. Not all features or steps shown in any one drawing to illustrate an exemplary embodiment are necessarily required, and some features or steps may be omitted. The order of steps shown in any of the drawings may be changed as appropriate.
[0114] Some or all of the above embodiments may also be described as follows, but are not limited to the following: (Note 1) An image matching device, At least one memory configured to store instructions, The system comprises at least one processor configured to execute the instruction, wherein the instruction is To obtain the first view image and the second view image, The calculation of the features of the first view image, To calculate the features of the second view image, The method involves determining whether the first view image and the second view image match based on the features of the first view image and the features of the second view image. The calculation of the features of the first view image is as follows: In order to generate a first view image with improved resolution, the resolution of the first view image is improved, This includes calculating the feature quantities of the first view image with improved resolution as the feature quantities of the first view image, An image matching device in which the first view image and the second view image are, respectively, a ground image and an aerial image, or the first view image and the second view image are, respectively, an aerial image and a ground image. (Note 2) The calculation of the features of the second view image is as follows: In order to generate a second view image with improved resolution, the resolution of the aforementioned second view image is improved, The image matching apparatus according to Appendix 1, comprising calculating the feature quantities of the resolution-enhanced second view image as the feature quantities of the second view image. (Note 3) A computer-based image matching method, Steps include acquiring the first view image and the second view image, The steps include: calculating the features of the first view image, The steps include: calculating the features of the second view image, The process includes the step of determining whether the first view image and the second view image match based on the features of the first view image and the features of the second view image, The calculation of the features of the first view image is as follows: To generate a first view image with improved resolution, the process involves improving the resolution of the first view image, The step includes calculating the feature quantities of the first view image with improved resolution as the feature quantities of the first view image, An image matching method wherein the first view image and the second view image are a ground image and an aerial image, respectively, or the first view image and the second view image are an aerial image and a ground image, respectively. (Note 4) The calculation of the features of the second view image is as follows: To generate a second view image with improved resolution, the process involves improving the resolution of the second view image, The image matching method according to Appendix 3, comprising the step of calculating the feature quantities of the second view image with improved resolution as the feature quantities of the second view image. (Note 5) A non-temporary computer-readable medium for storing programs, wherein a computer can, through a program, To obtain the first view image and the second view image, The calculation of the features of the first view image, To calculate the features of the second view image, Based on the features of the first view image and the features of the second view image, determine whether the first view image and the second view image match, and perform the following: The calculation of the features of the first view image is as follows: In order to generate a first view image with improved resolution, the resolution of the first view image is improved, This includes calculating the feature quantities of the first view image with improved resolution as the feature quantities of the first view image, A medium in which the first view image and the second view image are, respectively, a ground image and an aerial image, or the first view image and the second view image are, respectively, an aerial image and a ground image. (Note 6) The calculation of the features of the second view image is as follows: In order to generate a second view image with improved resolution, the resolution of the aforementioned second view image is improved, The medium according to Appendix 5, which includes calculating the feature quantities of the second view image with improved resolution as the feature quantities of the second view image. (Note 7) A training device, At least one memory configured to store instructions, The system comprises at least one processor configured to execute the instruction, wherein the instruction is Acquiring the target image, This involves training one or more models. The training of the one or more models is In order to generate a target image with reduced resolution, the resolution of the acquired target image is reduced, In order to generate a target image with improved resolution, the reduced-resolution target image is input into the resolution enhancement model, The loss is calculated using the acquired target image and the target image with improved resolution. This includes updating the resolution improvement model based on the loss, The training device is such that the target image is either a ground image or an aerial image. (Note 8) The reduction in the resolution of the acquired target image is The acquired target image is downsampled, The training apparatus according to Appendix 7, comprising generating the reduced-resolution target image by enlarging the size of the image obtained by the downsampling to the size of the acquired target image. (Note 9) The training apparatus according to Appendix 7 or 8, wherein the loss is calculated to represent the degree of difference between the acquired target image and the resolution-enhanced target image. (Note 10) The training of the one or more models is In order to calculate the feature quantities of the target image with improved resolution, the target image with improved resolution is input into a feature extraction model, The process involves calculating the feature quantities of the acquired target image, The loss is calculated to show the degree of difference between the feature quantities of the acquired target image and the feature quantities of the target image with improved resolution. The training apparatus according to Appendix 7 or 8, further comprising updating the resolution-enhanced target image and the feature extraction model based on the loss. (Note 11) A training method performed by a computer, Steps to acquire the target image, This includes the step of training one or more models, The training of the one or more models is To generate a target image with reduced resolution, the steps include reducing the resolution of the acquired target image, The steps include inputting the reduced-resolution target image into a resolution enhancement model in order to generate a target image with improved resolution, A step of calculating the loss using the acquired target image and the target image with improved resolution, The step of updating the resolution improvement model based on the loss, A training method wherein the target image is either a ground image or an aerial image. (Note 12) The reduction in the resolution of the acquired target image is The steps include downsampling the acquired target image, The training method according to Appendix 11, comprising the step of generating the reduced-resolution target image by enlarging the size of the image obtained by the downsampling to the size of the acquired target image. (Note 13) The training method according to Appendix 11 or 12, wherein the loss is calculated to represent the degree of difference between the acquired target image and the resolution-enhanced target image. (Note 14) The training of the one or more models is To calculate the feature quantities of the target image with improved resolution, the step is to input the target image with improved resolution into a feature extraction model. The steps include: calculating the feature quantities of the acquired target image, The steps include: calculating the loss that indicates the degree of difference between the feature quantities of the acquired target image and the feature quantities of the target image with improved resolution; The training method according to Appendix 11 or 12, further comprising the step of updating the resolution-improved target image and the feature extraction model based on the loss. (Note 15) A non-temporary computer-readable medium for storing programs, wherein a computer can, through a program, Acquiring the target image, This involves training one or more models, and performing the following: The training of the one or more models is In order to generate a target image with reduced resolution, the resolution of the acquired target image is reduced, In order to generate a target image with improved resolution, the reduced-resolution target image is input into the resolution enhancement model, The loss is calculated using the acquired target image and the target image with improved resolution. This includes updating the resolution improvement model based on the loss, The aforementioned target image is a medium, such as a ground image or an aerial image. (Note 16) The reduction in the resolution of the acquired target image is The acquired target image is downsampled, The medium according to Appendix 15, which includes generating the reduced-resolution target image by enlarging the size of the image obtained by the downsampling to the size of the acquired target image. (Note 17) The loss is calculated to represent the degree of difference between the acquired target image and the resolution-enhanced target image, in the medium described in Appendix 15 or 16. (Note 18) The training of the one or more models is In order to calculate the feature quantities of the target image with improved resolution, the target image with improved resolution is input into a feature extraction model, The process involves calculating the feature quantities of the acquired target image, The loss is calculated to show the degree of difference between the feature quantities of the acquired target image and the feature quantities of the target image with improved resolution. The medium according to Appendix 15 or 16, further comprising updating the resolution-enhanced target image and the feature extraction model based on the aforementioned loss.
[0115] Some or all of the elements described in any appendix may apply to various hardware, software, recording means, systems, and methods for recording software. [Explanation of Symbols]
[0116] 10 Ground images 15 Aerial images 20 First View Image 30. Second View Image 40. Improved resolution first view image 50 Improved resolution second view image 100 Resolution Improved Model 110 Feature Extraction Models 120 Feature Extraction Models 130 High-Resolution Model 150 First View Image 160 Resolution Reduced First View Image 170 Resolution Improved First View Image 180 Feature Extraction Models 200 First View Image 210 Second View Image 220 Improved resolution first view image 230 Feature Extraction Models 240 First View Image 250 first view image with reduced resolution 260 resolution improved first view image 270 Second View Image 280-resolution reduced-resolution second view image 290 Resolution improved 2nd view image 300 Feature Extraction Models 310 First View Image 320 Second View Image 330 Improved resolution second view image 340 Feature Extraction Models 350 Second View Image 360° reduced-resolution second view image 370 Resolution improved 2nd view image 400 Image Matching Device 410 First View Image 420 Second View Image 430 Feature Extraction Models 440 Feature Extraction Models 500 Geographic Location Systems 600 Location Databases 1000 computers 1020 Bus 1040 processor 1060 memory 1080 storage devices 1100 Input / Output Interface 1120 Network Interface 2000 Image Matching Device 2020 Acquisition Department 2040 First Feature Extraction Unit 2060 Second Feature Extraction Unit
Claims
1. An image matching device, At least one memory configured to store instructions, The system comprises at least one processor configured to execute the instruction, wherein the instruction is To obtain the first view image and the second view image, Calculating the feature quantities of the first view image, Calculating the features of the second view image, The method involves determining whether the first view image and the second view image match based on the features of the first view image and the features of the second view image. The calculation of the features of the first view image is as follows: In order to generate a first view image with improved resolution, the resolution of the first view image is improved, This includes calculating the feature quantities of the first view image with improved resolution as the feature quantities of the first view image, An image matching device in which the first view image and the second view image are, respectively, a ground image and an aerial image, or the first view image and the second view image are, respectively, an aerial image and a ground image.
2. The calculation of the features of the second view image is as follows: In order to generate a second view image with improved resolution, the resolution of the second view image is improved, The image matching apparatus according to claim 1, comprising calculating the feature quantities of the second view image with improved resolution as the feature quantities of the second view image.
3. A computer-based image matching method, Steps include acquiring a first view image and a second view image, The steps include: calculating the feature quantities of the first view image, The steps include calculating the feature quantities of the second view image, The process includes the step of determining whether the first view image and the second view image match based on the features of the first view image and the features of the second view image, The calculation of the features of the first view image is as follows: To generate a first view image with improved resolution, the process involves improving the resolution of the first view image, The step includes calculating the feature quantities of the first view image with improved resolution as the feature quantities of the first view image, An image matching method wherein the first view image and the second view image are a ground image and an aerial image, respectively, or the first view image and the second view image are an aerial image and a ground image, respectively.
4. The calculation of the features of the second view image is as follows: To generate a second view image with improved resolution, the process involves improving the resolution of the second view image, The image matching method according to claim 3, comprising the step of calculating the feature quantities of the second view image with improved resolution as the feature quantities of the second view image.
5. A non-temporary computer-readable medium for storing programs, wherein a computer can, through a program, To obtain the first view image and the second view image, Calculating the feature quantities of the first view image, Calculating the features of the second view image, Based on the features of the first view image and the features of the second view image, determine whether the first view image and the second view image match, and perform the following: The calculation of the features of the first view image is as follows: In order to generate a first view image with improved resolution, the resolution of the first view image is improved, This includes calculating the feature quantities of the first view image with improved resolution as the feature quantities of the first view image, A medium in which the first view image and the second view image are, respectively, a ground image and an aerial image, or the first view image and the second view image are, respectively, an aerial image and a ground image.
6. The calculation of the features of the second view image is as follows: In order to generate a second view image with improved resolution, the resolution of the second view image is improved, The medium according to claim 5, comprising calculating the feature quantities of the second view image with improved resolution as the feature quantities of the second view image.
7. A training device, At least one memory configured to store instructions, The system comprises at least one processor configured to execute the instruction, wherein the instruction is Acquiring the target image, This involves training one or more models. The training of the one or more models is In order to generate a target image with reduced resolution, the resolution of the acquired target image is reduced, In order to generate a target image with improved resolution, the reduced-resolution target image is input into the resolution enhancement model, The loss is calculated using the acquired target image and the target image with improved resolution. This includes updating the resolution improvement model based on the loss, The training device is such that the target image is either a ground image or an aerial image.
8. The reduction in the resolution of the acquired target image is The acquired target image is downsampled, The training apparatus according to claim 7, further comprising generating the reduced-resolution target image by enlarging the size of the image obtained by the downsampling to the size of the acquired target image.
9. The training apparatus according to claim 7 or 8, wherein the loss is calculated to represent the degree of difference between the acquired target image and the resolution-enhanced target image.
10. The training of the one or more models is In order to calculate the feature quantities of the target image with improved resolution, the target image with improved resolution is input into a feature extraction model, The process involves calculating the feature quantities of the acquired target image, The loss is calculated to show the degree of difference between the feature quantities of the acquired target image and the feature quantities of the target image with improved resolution. The training apparatus according to claim 7 or 8, further comprising updating the resolution-enhanced target image and the feature extraction model based on the loss.
11. A training method performed by a computer, Steps to acquire the target image, The steps include, The training of the one or more models is To generate a target image with reduced resolution, the steps include reducing the resolution of the acquired target image, The steps include inputting the reduced-resolution target image into a resolution enhancement model in order to generate a target image with improved resolution, A step of calculating the loss using the acquired target image and the target image with improved resolution, The step of updating the resolution improvement model based on the loss, A training method wherein the target image is either a ground image or an aerial image.
12. The reduction in the resolution of the acquired target image is The steps include downsampling the acquired target image, The training method according to claim 11, comprising the step of generating the reduced-resolution target image by enlarging the size of the image obtained by the downsampling to the size of the acquired target image.
13. The training method according to claim 11 or 12, wherein the loss is calculated to represent the degree of difference between the acquired target image and the resolution-enhanced target image.
14. The training of the one or more models is To calculate the feature quantities of the target image with improved resolution, the step is to input the target image with improved resolution into a feature extraction model. The steps include: calculating the feature quantities of the acquired target image, The steps include: calculating the loss that indicates the degree of difference between the feature quantities of the acquired target image and the feature quantities of the target image with improved resolution; The training method according to claim 11 or 12, further comprising the step of updating the resolution-improved target image and the feature extraction model based on the loss.
15. A non-temporary computer-readable medium for storing programs, wherein a computer can, through a program, Acquiring the target image, This involves training one or more models, and performing the following: The training of the one or more models is In order to generate a target image with reduced resolution, the resolution of the acquired target image is reduced, In order to generate a target image with improved resolution, the reduced-resolution target image is input into the resolution enhancement model, The loss is calculated using the acquired target image and the target image with improved resolution. This includes updating the resolution improvement model based on the loss, The aforementioned target image is a medium, such as a ground image or an aerial image.
16. The reduction in the resolution of the acquired target image is The acquired target image is downsampled, The medium according to claim 15, comprising generating the reduced-resolution target image by enlarging the size of the image obtained by the downsampling to the size of the acquired target image.
17. The medium according to claim 15 or 16, wherein the loss is calculated to represent the degree of difference between the acquired target image and the resolution-enhanced target image.
18. The training of the one or more models is In order to calculate the feature quantities of the target image with improved resolution, the target image with improved resolution is input into a feature extraction model, The process involves calculating the feature quantities of the acquired target image, The loss is calculated to show the degree of difference between the feature quantities of the acquired target image and the feature quantities of the target image with improved resolution. The medium according to claim 15 or 16, further comprising updating the resolution-enhanced target image and the feature extraction model based on the loss.