Training device, training method, and program

The training device and method utilize terrestrial, aerial, and map images to calculate a combination loss, enhancing the performance of feature extractors in cross-view image localization by addressing the limitations of existing systems that only use camera-captured images.

JP7794359B2Active Publication Date: 2026-01-06NEC CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025508758
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-08-25
Publication Date
2026-01-06
Estimated Expiration
2042-08-25

AI Technical Summary

Technical Problem

Existing systems for cross-view image localization do not utilize images other than those captured by a camera or their pose maps for training feature extractors.

Method used

A training device and method that uses terrestrial, aerial, and map images to train a feature extractor set by calculating a combination loss based on the features extracted from these images, updating the extractors to enhance their performance in cross-view image matching.

Benefits of technology

Provides a novel technique for training feature extractors, improving their effectiveness in cross-view image localization by leveraging a combination of terrestrial, aerial, and map images, which can accelerate training and enhance matching accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007794359000003
    Figure 0007794359000003
  • Figure 0007794359000004
    Figure 0007794359000004
  • Figure 0007794359000005
    Figure 0007794359000005
Patent Text Reader

Abstract

A training device (2000) acquires training data (10) including terrestrial images (20), aerial images (30), and a map image (40). The training device (2000) acquires feature quantities of the terrestrial images (20), the aerial images (30), and the map image (40) by inputting the terrestrial images (20), the aerial images (30), and the map image (40) to a first feature extractor (60), a second feature extractor (70), and a third feature extractor (80), respectively. The training device (2000) calculates a combination loss based on the acquired feature quantities, and updates the first feature extractor (60), the second feature extractor (70), and the third feature extractor (80) based on the combination loss.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure generally relates to training devices, training methods, and non-transitory computer-readable storage media. [Background technology]

[0002] Computer systems that perform cross-view image localization have been developed. For example, Non-Patent Document 1 discloses a system that includes a set of feature extractors implemented with a convolutional neural network (CNN) and matches ground-level images with satellite images to determine where the ground-level images were captured. Specifically, one of the feature extractors is configured to acquire a set of ground-level images and an orientation map indicating the orientation (azimuth angle and altitude) of each location captured in the ground-level images, and is trained to extract features therefrom. The other is configured to acquire a set of satellite images and an orientation map indicating the orientation (azimuth angle and distance) of each location captured in the satellite images, and is trained to extract features therefrom. The system then determines whether the ground-level image matches the satellite image based on the features extracted by the trained feature extractors. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] International Publication No. 2022 / 034678 [Patent Document 2] International Publication No. 2022 / 044105 [Non-patent literature]

[0004] [Non-Patent Document 1] Liu Liu and Hongdong Li, “Lending Orientation to Neural Networks for Cross-view Geo-localization”, [online], March 29, 2019, [Retrieved August 17, 2022],<arXiv, https: / / arxiv.org / pdf / 1903.12351.pdf> Retrieved from Summary of the Invention [Problem to be solved by the invention]

[0005] Non-Patent Document 1 does not consider using images other than those captured by a camera or their pose maps for training the feature extractor. The purpose of this disclosure is to provide a novel technique for training a feature extractor. [Means for solving the problem]

[0006] The present disclosure provides a training device comprising at least one memory configured to store instructions and at least one processor. At least one processor is configured to execute instructions to acquire training data including a first terrestrial image, a first aerial image, and a first map image; input the first terrestrial image to a first feature extractor to extract features of the first terrestrial image; input the first aerial image to a second feature extractor to extract features of the first aerial image; input the first map image to a third feature extractor to extract features of the first map image; calculate a combination loss based on the features of the first terrestrial image, the features of the first aerial image, and the features of the first map image; and update the first feature extractor, the second feature extractor, and the third feature extractor based on the combination loss.

[0007] The present disclosure further provides a training method, including: acquiring training data including a first ground image, a first aerial image, and a first map image; inputting the first ground image to a first feature extractor to extract features of the first ground image; inputting the first aerial image to a second feature extractor to extract features of the first aerial image; inputting the first map image to a third feature extractor to extract features of the first map image; calculating a combination loss based on the features of the first ground image, the features of the first aerial image, and the features of the first map image; and updating the first feature extractor, the second feature extractor, and the third feature extractor based on the combination loss.

[0008] The present disclosure also provides a non-transitory computer-readable storage medium storing a program that causes a computer to acquire training data including a first terrestrial image, a first aerial image, and a first map image, input the first terrestrial image to a first feature extractor to extract features of the first terrestrial image, input the first aerial image to a second feature extractor to extract features of the first aerial image, input the first map image to a third feature extractor to extract features of the first map image, calculate a combination loss based on the features of the first terrestrial image, the features of the first aerial image, and the features of the first map image, and update the first feature extractor, the second feature extractor, and the third feature extractor based on the combination loss. [Effects of the Invention]

[0009] According to the present disclosure, novel techniques for training a feature extractor can be provided. [Brief explanation of the drawings]

[0010] [Figure 1] 1 shows an overview of a training device according to a first embodiment. [Figure 2] An example of training data is shown below. [Figure 3]1 is a block diagram showing an example of the functional configuration of a training apparatus according to a first embodiment. [Figure 4] FIG. 2 is a block diagram showing an example of the hardware configuration of a computer that realizes the training device of the first embodiment. [Figure 5] 1 shows a flowchart illustrating an example of a processing flow by the training device according to the first embodiment. [Figure 6] 1 illustrates a geolocalization system employing all or part of a set of feature extractors. [Figure 7] An example of a method for calculating a similarity score will be described below. [Figure 8] 1 illustrates an exemplary method for calculating a similarity score. [Figure 9] An example of a method for calculating a similarity score will be described below. [Figure 10] 1 shows an overview of a training device according to a second embodiment. [Figure 11] 1 shows an example of data augmentation by a training device. [Figure 12] FIG. 10 is a block diagram showing an example of the functional configuration of a training device according to a second embodiment. [Figure 13] 10 shows a flowchart illustrating an example of a processing flow by the training device of embodiment 2. [Figure 14] This shows a case where part of a map image is replaced with an aerial image. [Figure 15] This shows a case where part of an aerial image is replaced with a map image. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In each drawing, the same elements are denoted by the same reference numerals, and duplicated descriptions are omitted as necessary. Unless otherwise specified, predetermined information (e.g., predetermined values, predetermined thresholds, etc.) is pre-stored in a storage unit accessible by a computer that uses the information. In this disclosure, the storage unit may be implemented using one or more storage devices, such as a hard disk, a solid-state drive (SSD), or random-access memory (RAM).

[0012] Embodiment 1 <Summary> Fig. 1 shows an overview of a training device 2000 according to embodiment 1. Note that Fig. 1 does not limit the operation of the training device 2000, but merely shows an example of the possible operation of the training device 2000.

[0013] The training device 2000 is a device that acquires training data 10 and uses the training data 10 to train a feature extractor set 50. The training data 10 includes terrestrial images 20, aerial images 30, and map images 40. The feature extractor set 50 includes three feature extractors, namely, a first feature extractor 60, a second feature extractor 70, and a third feature extractor 80.

[0014] The first feature extractor 60 receives the terrestrial image 20 as input and extracts feature quantities from the input terrestrial image 20. The second feature extractor 70 receives the aerial image 30 as input and extracts feature quantities from the input aerial image 30. The third feature extractor 80 receives the map image 40 as input and extracts feature quantities from the input map image 40.

[0015] The feature extractor may take various forms, and any of these forms may be applied to the first feature extractor 60, the second feature extractor 70, and the third feature extractor 80. For example, the first feature extractor 60, the second feature extractor 70, and the third feature extractor 80 may be realized as a machine learning-based model such as a neural network. It should be noted that the first feature extractor 60, the second feature extractor 70, and the third feature extractor 80 may be realized in forms different from each other.

[0016] 2 shows an example of training data 10. Terrestrial imagery 20 is a digital image (e.g., an RGB image or a grayscale image) containing a ground view of a location. Terrestrial imagery 20 is generated by a camera, called a "ground view camera," that captures a ground view of a location. The ground camera may be held by a pedestrian or mounted on a vehicle such as a car, motorcycle, or drone. Terrestrial imagery 20 may be panoramic (having a 360-degree field of view) or may have a limited field of view (less than 360 degrees).

[0017] The aerial image 30 is a digital image (e.g., an RGB image or a grayscale image) that includes an aerial view of a location. The aerial image 30 may be generated by capturing a scene from above using a camera called an aerial camera installed on a drone, airplane, satellite, or the like.

[0018] The map image 40 is a digital image (e.g., an RGB image or a grayscale image) containing a map of a location. The map image 40 may be obtained from open data such as OpenStreetMap®, or may be prepared by the provider or user of the training device 2000.

[0019] The aerial image 30 and the map image 40 in the training data 10 correspond to the same position. For example, the center position of the location indicated by the aerial image 30 and the center position of the location indicated by the map image 40 are sufficiently close, so that the aerial image 30 and the map image 40 can be associated with the same position information. The position information is information that identifies a position, such as GPS (Global Positioning System) coordinates.

[0020] The training device 2000 can train the feature extractor set 50 as follows. The training device 2000 inputs a terrestrial image 20 to a first feature extractor 60, thereby acquiring the feature quantities of the terrestrial image 20 from the first feature extractor 60. Similarly, the training device 2000 inputs an aerial image 30 to a second feature extractor 70, thereby acquiring the feature quantities of the aerial image 30 from the second feature extractor 70. Furthermore, the training device 2000 inputs a map image 40 to a third feature extractor 80, thereby acquiring the feature quantities of the map image 40 from the third feature extractor 80.

[0021] The training device 2000 calculates a combination loss based on the feature amounts of the terrestrial image 20, the feature amounts of the aerial image 30, and the feature amounts of the map image 40. The combination loss can be calculated by combining the loss between the feature amounts of the terrestrial image 20 and the feature amounts of the aerial image 30, the loss between the feature amounts of the terrestrial image 20 and the feature amounts of the map image 40, and the loss between the feature amounts of the aerial image 30 and the feature amounts of the map image 40. Then, the training device 2000 updates the feature extractor set 50 based on the combination loss. The feature extractor set 50 can be trained by updating it using multiple training data 10.

[0022] <Examples of effects> According to the training device 2000 of the first embodiment, a combination loss is calculated using the feature amounts of the terrestrial image 20, the feature amounts of the aerial image 30, and the feature amounts of the map image 40, and this combination loss amount is used to train a set of feature extractors, i.e., the feature extractor set 50. Thus, a novel technique for training feature extractors is provided.

[0023] It should be noted that, as will be explained in more detail later, feature extractor set 50 may be used for cross-view image matching, however, either second feature extractor 70 or third feature extractor 80 need not be used for cross-view image matching.

[0024] For example, suppose the third feature extractor 80 is not used for cross-view image matching. In this case, it is technically possible to exclude the third feature extractor 80 from the feature extractor set 50 when training the feature extractor set 50. However, even when the third feature extractor 80 is not used for cross-view image matching, it is advantageous to use the third feature extractor 80 in training the feature extractor set 50. Specifically, because the map image 40 is simpler than the aerial image 30 (e.g., buildings are depicted as rectangles, etc.), a loss calculated based on the features of the map image 40 can accelerate the training of the first feature extractor 60 and the second feature extractor 70.

[0025] Furthermore, by providing both the second feature extractor 70 and the third feature extractor 80, it is possible to select either one or both depending on the situation in which cross-view image matching is performed. For example, because aerial images 30 contain more information than map images 40, the second feature extractor 70 can be preferably employed for cross-view image matching as long as aerial images 30 are available. However, there are cases in which aerial images 30 cannot be used due to, for example, regulations by national or local authorities. In these situations, the third feature extractor 80 is employed for cross-view image matching.

[0026] The training device 2000 is described in more detail below.

[0027] <Example of functional configuration> 3 is a block diagram showing an example of the functional configuration of the training device 2000 according to embodiment 1. The training device 2000 includes an acquisition unit 2020, a feature extraction unit 2040, and an update unit 2060.

[0028] The acquisition unit 2020 acquires training data 10 including terrestrial images 20, aerial images 30, and a map image 40. The feature extraction unit 2040 inputs the terrestrial images 20 to a first feature extractor 60 and acquires the feature quantities of the terrestrial images 20 from the first feature extractor 60. The feature extraction unit 2040 inputs the aerial images 30 to a second feature extractor 70 and acquires the feature quantities of the aerial images 30 from the second feature extractor 70. The feature extraction unit 2040 inputs the map image 40 to a third feature extractor 80 and acquires the feature quantities of the map image 40 from the third feature extractor 80. The update unit 2060 calculates a combination loss based on the feature quantities of the terrestrial images 20, the feature quantities of the aerial images 30, and the feature quantities of the map image 40. Then, the update unit 2060 updates the feature extractor set 50 (that is, the first feature extractor 60, the second feature extractor 70, and the third feature extractor 80) based on the combination loss.

[0029] <Example of hardware configuration> Training device 2000 may be implemented by one or more computers, each of which may be a dedicated computer manufactured for implementing training device 2000, or a general-purpose computer such as a personal computer (PC), a server machine, or a mobile device.

[0030] The training device 2000 can be realized by installing an application on a computer. This application is implemented by a program that causes the computer to function as the training device 2000. In other words, the program implements each functional unit of the training device 2000. There are various ways to acquire the program. For example, the program can be acquired from a storage medium (such as a DVD disk or USB memory) on which the program is pre-stored. As another example, the program can be acquired by downloading it from a server machine that manages the storage medium on which the program is pre-stored.

[0031] Fig. 4 is a block diagram showing an example of the hardware configuration of a computer 1000 that realizes the training device 2000 of embodiment 1. In Fig. 4, the computer 1000 has a bus 1020, a processor 1040, a memory 1060, a storage device 1080, an input / output (I / O) interface 1100, and a network interface 1120.

[0032] The bus 1020 is a data transmission path through which the processor 1040, memory 1060, storage device 1080, input / output interface 1100, and network interface 1120 transmit and receive data to and from each other. The processor 1040 is a processor such as a central processing unit (CPU), a graphics processing unit (GPU), a field-programmable gate array (FPGA), or a digital signal processor (DSP). The memory 1060 is a main memory element such as a random access memory (RAM) or a read-only memory (ROM). The storage device 1080 is an auxiliary memory element such as a hard disk, a solid-state drive (SSD), or a memory card. The input / output interface 1100 is an interface between the computer 1000 and peripheral devices such as a keyboard, a mouse, or a display device. The network interface 1120 is an interface between the computer 1000 and a network. The network may be a local area network (LAN) or a wide area network (WAN). The storage device 1080 may store the above-mentioned programs. The processor 1040 executes a program to realize each functional unit of the training device 2000.

[0033] The hardware configuration of the computer 1000 is not limited to that shown in Fig. 4. For example, as described above, the training device 2000 may be realized by multiple computers. In this case, these computers may be connected to each other via a network.

[0034] <Processing flow> 5 is a flowchart showing an example of the processing flow of the training device 2000 according to the first embodiment. The acquisition unit 2020 acquires training data 10 including terrestrial images 20, aerial images 30, and a map image 40 (S102). The feature extraction unit 2040 inputs the terrestrial images 20 to the first feature extractor 60 to acquire feature quantities of the terrestrial images 20 (S104). The feature extraction unit 2040 inputs the aerial images 30 to the second feature extractor 70 to acquire feature quantities of the aerial images 30 (S106). The feature extraction unit 2040 inputs the map image 40 to the third feature extractor 80 to acquire feature quantities of the map image 40 (S108). The update unit 2060 calculates a combination loss based on the acquired features (S110). The update unit 2060 updates the first feature extractor 60, the second feature extractor 70, and the third feature extractor 80 based on the combination loss (S112).

[0035] It should be noted that the flowchart shown in Fig. 5 is merely an example of a processing flow that can be executed by the training device 2000, and the processing flow executed by the training device 2000 is not limited to that shown in Fig. 5. For example, the extraction of feature amounts from the terrestrial captured image 20 (S104), the extraction of feature amounts from the aerial captured image 30 (S106), and the extraction of feature amounts from the map image 40 (S108) may be performed in an order different from that shown in Fig. 5, or may be performed in parallel.

[0036] As described above, the training device 2000 may train the feature extractor set 50 using multiple sets of training data 10. Various methods for training a feature extractor using multiple sets of training data are known, and one of these methods may be applied to the training device 2000. For example, the training device 2000 may perform the process illustrated in FIG. 5 for each of the multiple sets of training data 10. As another example, the training device 2000 may perform batch training on the feature extractor set 50 using the multiple sets of training data 10. In this case, the training device 2000 may obtain an aggregate loss by aggregating the combination losses obtained from the multiple sets of training data 10, and update the feature extractor set 50 based on the aggregate loss. The aggregate loss may be a statistical value such as the average value of the combination losses.

[0037] <Application example of feature extractor set 50> As described above, all or part of the feature extractor set 50 may be used in a matching device that performs cross-view image matching. In the following, to facilitate understanding of the feature extractor set 50, such a matching device will be described as an application example of the feature extractor set 50.

[0038] 6 illustrates a geolocalization system 200 in which all or part of the feature extractor set 50 may be employed. The geolocalization system 200 is a system for performing image geolocalization. Image geolocalization is a technique for identifying the location where an input image was captured. The geolocalization system 200 may be implemented by one or more computers, such as those shown in FIG. 4.

[0039] The geo-localization system 200 includes a matching device 250. The matching device 250 acquires ground information 210 and aerial information 220, and determines whether the ground information 210 and the aerial information 220 match.

[0040] The ground information 210 includes images of a location captured from a ground perspective, i.e., ground-captured images 20. The aerial information 220 includes at least one type of image showing a location from a planar view. When the second feature extractor 70 is employed in the matching device 250, the aerial information 220 includes aerial images 30. When the third feature extractor 80 is employed in the matching device 250, the aerial information 220 includes map images 40.

[0041] The matching device 250 may calculate a similarity score indicating the similarity between the ground feature amount and the aeronautical feature amount to determine whether the ground information 210 and the aeronautical information 220 match. Then, the matching device 250 determines that the ground information 210 matches the aeronautical information 220 if the similarity score is sufficiently large (for example, larger than a predetermined threshold).

[0042] The ground features are a set of features extracted from the ground information 210, i.e., features extracted from the ground-captured images 20. The aerial features are a set of features extracted from the aerial information 220, i.e., features extracted from the aerial images 30, features extracted from the map images 40, or both.

[0043] 7 to 9 show exemplary methods for calculating the similarity score. In the example shown in Fig. 7, the matching device 250 employs the second feature extractor 70 but does not employ the third feature extractor 80. In this case, the matching device 250 calculates the similarity between the feature amounts of the terrestrial-photographed image 20 and the feature amounts of the aerial-photographed image 30 as the similarity score.

[0044] 8, the matching device 250 employs the third feature extractor 80 but does not employ the second feature extractor 70. In this case, the matching device 250 calculates the similarity between the feature amount of the ground captured image 20 and the feature amount of the map image 40 as a similarity score.

[0045] 9, the matching device 250 employs both the second feature extractor 70 and the third feature extractor 80. In this case, the matching device 250 calculates the similarity between the feature amounts of the terrestrial captured image 20 and the feature amounts of the aerial captured image 30, and the similarity between the feature amounts of the terrestrial captured image 20 and the feature amounts of the map image 40, and combines these (for example, calculates a weighted average) to calculate a similarity score.

[0046] Various metrics can be used to calculate the similarity between features. For example, the similarity between features can be calculated as one of various types of distance (e.g., L2 distance), correlation, cosine similarity, or neural network (NN)-based similarity between features. NN-based similarity is similarity calculated by a neural network trained to calculate the similarity between input features.

[0047] 6, the geo-localization system 200 also includes a location database 300. The location database 300 includes location information 230 associated with the aeronautical information 220 for each of the various locations. The location information 230 identifies the location of the place corresponding to the aeronautical information 220 associated with the location information 230.

[0048] If a user wants to know where a terrestrial image 20 was taken, the user may operate the user terminal to transmit ground information 210 containing the terrestrial image 20 to the geolocalization system 200. The geolocalization system 200 receives the ground information 210 and searches the location database 300 for aerial information 220 that matches the received ground information 210 to identify the location in the ground information 210 where the terrestrial image 20 was taken.

[0049] Specifically, the geo-localization system 200 repeatedly executes the following steps until aerial information 220 matching the ground information 210 is detected: acquiring one of the aerial information 220 from the location database 300, inputting the set of the ground information 210 and the aerial information 220 into the matching device 250, and determining whether the matching device 250 indicates that the ground information 210 matches the aerial information 220. If aerial information 220 matching the ground information 210 is detected, the geo-localization system 200 can determine that the location where the terrestrial image 20 in the ground information 210 was captured is the location identified by the location information 230 associated with the detected aerial information 220.

[0050] The geo-localization system 200 may transmit a response 240 to the user terminal. The response 240 may include location information 230 identified by the geo-localization system 200 to specify where the terrestrial image 20 was taken. The response 240 may also include aerial information 220 identified by the geo-localization system 200 as matching the terrestrial information 210.

[0051] It should be noted that the geo-localization system 200 may be configured to receive aerial information including aerial images 30 and identify the location where the received aerial images 30 were captured. In this case, the location database 300 includes pairs of ground information and location information. The matching device 250 also includes a first feature extractor 60 and a second feature extractor 70.

[0052] Specifically, the geolocalization system 200 repeatedly acquires one of the ground information from the location database 300, inputs the set of the aerial information and the ground information into the matching device 250, and determines whether the matching device 250 indicates that the aerial information matches the ground information, until ground information matching the received aerial information is detected. If ground information matching the received aerial information is detected, the geolocalization system 200 can identify the location where the aerial image 30 in the received aerial information was captured as the location identified by the location information associated with the detected ground information. The geolocalization system then transmits a response to the user terminal. The response may include the detected ground information, the location information associated with the detected ground information, or both.

[0053] <Acquisition of training data: S102> The acquiring unit 2020 acquires the training data 10 (S102). There are various methods for acquiring the training data 10. In some implementations, the acquiring unit 2020 may generate the training data 10, or may receive the training data 10 transmitted from another computer. In other implementations, the training data may be stored in advance in a storage unit accessible to the acquiring unit 2020. In this case, the acquiring unit 2020 reads the training data 10 from the storage unit.

[0054] <Feature extraction: S104, S106, S108> The feature extraction unit 2040 extracts features from the terrestrial image 20, the aerial image 30, and the map image 40 (S104, S106, S108). Specifically, the feature extraction unit 2040 extracts the terrestrial image 20 from the training data 10 and inputs the terrestrial image 20 to the first feature extractor 60. The first feature extractor 60 is configured to extract features from the input image, so the feature extraction unit 2040 can acquire the feature of the terrestrial image 20 from the first feature extractor 60. Similarly, the feature extraction unit 2040 extracts the aerial image 30 from the training data 10 and inputs the aerial image 30 to the second feature extractor 70, thereby acquiring the feature of the aerial image 30 from the second feature extractor 70. Furthermore, the feature extraction unit 2040 extracts the map image 40 from the training data 10 and inputs the map image 40 to the third feature extractor 80 , thereby obtaining the feature quantities of the map image 40 from the third feature extractor 80 .

[0055] <Coupling loss calculation: S110> The update unit 2060 calculates a combination loss based on the feature amounts of the terrestrial image 20, the feature amounts of the aerial image 30, and the feature amounts of the map image 40 (S110). As described above, the combination loss can be calculated by combining the loss between the feature amounts of the terrestrial image 20 and the feature amounts of the aerial image 30, the loss between the feature amounts of the terrestrial image 20 and the feature amounts of the map image 40, and the loss between the feature amounts of the aerial image 30 and the feature amounts of the map image 40. In this case, the combination loss can be calculated using the following loss function L:

number

[0056] In equation (1), f_g, f_a, and f_m represent the features of the terrestrial image 20, the aerial image 30, and the map image 40, respectively. Note that subscripts are underlined except in equation (1). L represents a loss function for calculating the combination loss. L_ga is a loss function for calculating the loss between the features of the terrestrial image 20 and the aerial image 30. L_gm is a loss function for calculating the loss between the features of the terrestrial image 20 and the map image 40. L_am is a loss function for calculating the loss between the features of the aerial image 30 and the map image 40. W_ga, W_gm, and Wam represent weights assigned to L_ga, L_gm, and Lam, respectively. If equal weights are assigned to L_ga, L_gm, and L_am, the weights W_ga, W_gm, and W_am can be removed from equation (1).

[0057] There are various types of loss functions, and one of these loss functions (e.g., contrastive loss or triplet loss) can be adopted as the loss functions L_ga, L_gm, and L_am. Since the feature extractor set 50 can be used to perform matching between the ground information 210 and the aerial information 220, as illustrated with reference to FIG. 6, the loss between the ground information 210 and the aerial information 220 should be sufficiently small when the ground information 210 and the aerial information 220 indicate the same position, and should not be sufficiently small when the ground information 210 and the aerial information 220 indicate different positions.

[0058] Specifically, the loss between the terrestrial image 20 and the aerial image 30 should be sufficiently small when the position at which the terrestrial image 20 was captured is sufficiently close to the center position of the aerial image 30, but should not be sufficiently small when the position at which the terrestrial image 20 was captured is not sufficiently close to the center position of the aerial image 30. Similarly, the loss between the terrestrial image 20 and the map image 40 should be sufficiently small when the position at which the terrestrial image 20 was captured is sufficiently close to the center position of the map image 40, but should not be sufficiently small when the position at which the terrestrial image 20 was captured is not sufficiently close to the center position of the map image 40.

[0059] To achieve this, the training device 2000 may use both positive example training data 10 and negative example training data 10. The positive example training data 10 satisfies the condition that the position at which the terrestrial image 20 was captured is sufficiently close to both the center position of the aerial image 30 and the center position of the map image 40. On the other hand, the negative example training data 10 satisfies the condition that the position at which the terrestrial image 20 was captured is not sufficiently close to the center position of either the aerial image 30 or the map image 40.

[0060] The training device 2000 may train the feature extractor set 50 using both a set of features extracted from the positive example training data 10 and a set of features extracted from the negative example training data 10. There are various methods for training feature extractors using positive and negative examples, and any of these methods can be applied to the training device 2000.

[0061] When triplet loss is employed, the training device 2000 may employ training data 10 including both positive and negative examples. Specifically, the training data 10 may include a pair of positive examples including a terrestrial image 20, a positive example aerial image 30, and a positive example map image 40, and a pair of negative examples including a negative example aerial image 30 and a negative example map image 40. The pair of positive examples satisfies the condition that the location where the terrestrial image 20 was captured is sufficiently close to both the center of the positive example aerial image 30 and the center of the positive example map image 40. On the other hand, the pair of negative examples satisfies the condition that the location where the terrestrial image 20 was captured is not sufficiently close to the center of either the negative example aerial image 30 or the negative example map image 40. Furthermore, the centers of the aerial image 30 and the map image 40 of the same pair are sufficiently close to each other.

[0062] The training device 2000 may use, as the features extracted from each image of the training data 10, features extracted from the ground-based image 20, features extracted from the positive example aerial image 30, features extracted from the negative example aerial image 30, features extracted from the positive example map image 40, and features extracted from the negative example map image 40.

[0063] If triplet loss is employed, the loss function for calculating the combining loss may be defined as follows:

number

[0064] In equation (2), f_ap, f_an, f_mp, and f_mn represent the feature values ​​of the positive example aerial image 30, the feature values ​​of the negative example aerial image 30, the feature values ​​of the positive example map image 40, and the feature values ​​of the negative example map image 40, respectively. L_ga is a triplet loss function that calculates the triplet loss between the feature values ​​of the terrestrial image 20, the feature values ​​of the positive example aerial image 30, and the feature values ​​of the negative example aerial image 30. L_gm is a triplet loss function that calculates the triplet loss between the feature values ​​of the terrestrial image 20, the feature values ​​of the positive example map image 40, and the feature values ​​of the negative example map image 40. L_gam is a triplet loss function that calculates the triplet loss between the feature values ​​of the terrestrial image 20, the feature values ​​of the positive example aerial image 30, and the feature values ​​of the negative example map image 40. Lgma is a triplet loss function that calculates the triplet loss between the feature values ​​of the terrestrial image 20, the feature values ​​of the positive example map image 40, and the feature values ​​of the negative example aerial image 30. W_gam and W_gma represent weights assigned to L_gam and L_gma, respectively.

[0065] <Update of feature extractor set 50: S112> The update unit 2060 updates the feature extractor set 50 based on the combination loss (S112). The first feature extractor 60, the second feature extractor 70, and the third feature extractor 80 are configured to have several trainable parameters, for example, weights assigned to each connection of a neural network. In this manner, the update unit 2060 updates the trainable parameters of the first feature extractor 60, the second feature extractor 70, and the third feature extractor 80 based on the combination loss, thereby updating the feature extractor set 50. Note that there are various known methods for updating the trainable parameters of the feature extractors using losses calculated based on features obtained from the feature extractors, and one of these methods can be applied to the update unit 2060.

[0066] <Output from Training Device 2000> The training device 2000 may output the training results of the feature extractor set 50. The training results may be output in any manner. For example, the training device 2000 may store the trained parameters (e.g., weights assigned to each connection in the neural network) of the first feature extractor 60, the second feature extractor 70, and the third feature extractor 80 in a storage unit. In another example, the training device 2000 may transmit the trained parameters to another device, such as the matching device 250. Note that in addition to the parameters, a program implementing the feature extractor set 50 may also be output.

[0067] When the matching device 250 is implemented in the training device 2000, the training device 2000 does not need to output the training results. In this case, from the perspective of the user of the training device 2000, it is preferable that the training device 2000 notify the user that the training of the matching device 250 has been completed.

[0068] Embodiment 2 <Summary> Fig. 10 shows an overview of a training device 2000 according to embodiment 2. It should be noted that Fig. 10 does not limit the operation of the training device 2000, but merely shows an example of possible operations of the training device 2000. Unless otherwise specified, the training device 2000 according to embodiment 2 includes all of the functions included in embodiment 1.

[0069] The training device 2000 of the second embodiment further performs data augmentation on the training data 10 to generate augmented training data 100 including a terrestrial image 110, an aerial image 120, and a map image 130. The terrestrial image 110 is the same as the terrestrial image 20 in the training data 10. On the other hand, the aerial image 120, the map image 130, or both are generated based on the aerial image 30 and the map image 40 in the training data 10, and therefore are partially different from the corresponding images in the training data 10.

[0070] The data augmentation performed by the training device 2000 includes image blending of a portion of the aerial image 30 and a portion of the map image 40. Fig. 11 shows an example of data augmentation performed by the training device 2000. In Fig. 11, the partial image 32 in the aerial image 30 and the partial image 42 in the map image 40 are the targets of image blending.

[0071] Specifically, the training device 2000 blends the partial image 32 and the partial image 42 at a blend ratio of Ra:Rm to obtain the augmented image 140 (in the case of FIG. 11 , Ra=Rm=0.5). Then, the training device 2000 generates the aerial image 120 by replacing the partial image 32 in the aerial image 30 with the augmented image 140. Similarly, the training device 2000 generates the map image 130 by replacing the partial image 42 in the map image 40 with the augmented image 140.

[0072] In the example of Figure 11, one extended image 140 is used to replace both partial image 32 and partial image 42, but the training device 2000 may use the extended image 140 to replace either partial image 32 or partial image 42.

[0073] The training device 2000 of the second embodiment trains the feature extractor set 50 using the augmented training data 100 in the same way that the training device 2000 of the first embodiment trains the feature extractor set 50 using the training data 10. Specifically, the training device 2000 inputs a terrestrial image 110, an aerial image 120, and a map image 130 to a first feature extractor 60, a second feature extractor 70, and a third feature extractor 80, respectively. As a result, the training device 2000 acquires the feature amounts of the terrestrial image 110, the aerial image 120, and the map image 130. The training device 2000 then calculates a combination loss based on the feature amounts of the terrestrial image 110, the aerial image 120, and the map image 130, and updates the feature extractor set 50 based on the combination loss.

[0074] The training device 2000 may generate the extended training data 100 by modifying either the aerial image 30 or the map image 40. In this case, either the aerial image 120 or the map image 130 is the same as the training data 10.

[0075] <Examples of effects> According to the training device 2000 of the second embodiment, augmented training data 100 is generated by performing data augmentation using training data 10. Data augmentation includes image blending, in which a portion of aerial image 30 and a portion of map image 40 (i.e., partial image 32 and partial image 42) are blended to generate augmented image 140, replacing at least one of them to generate aerial image 120, map image 130, or both, included in augmented training data 100. Thus, a novel technique is provided for performing data augmentation to generate images for training a feature extractor.

[0076] Additionally, data augmentation by the training device 2000 can increase the amount of information in the map image 40. The map image 40 does not necessarily need to be detailed. The level of detail in the map image 40 may depend on the type of mapping technique employed to generate the map image 40 or the effort taken to generate the map image 40. For example, detailed information (e.g., trees, buildings, parking lots, etc.) may be omitted from the map image 40. More detailed content can be added to the map image 40 by blending a portion of the map image 40 with a corresponding portion of the aerial image 30.

[0077] Additionally, the training device 2000 can perform data augmentation to simplify the information in the aerial image 30. Due to the many details present in the aerial image 30, it takes time for the feature extractor set 50 to learn meaningful features. Reducing the details in the aerial image 30 can prevent the feature extractor set 50 from focusing on detailed information (e.g., color) and allow the feature extractor set 50 to learn concepts, thereby shortening the training time and simplifying the feature learning process.

[0078] The training device 2000 is described in more detail below.

[0079] <Example of functional configuration> FIG. 12 is a block diagram showing an example of the functional configuration of a training device 2000 according to the second embodiment. As shown in FIG. 12, the training device 2000 according to the second embodiment includes an extension unit 2080 in addition to the functional units included in the training device 2000 according to the first embodiment. The extension unit 2080 generates extended training data 100 based on training data 10. The feature extraction unit 2040 according to the second embodiment inputs the terrestrial images 110, the aerial images 120, and the map images 130 in the extended training data 100 to the first feature extractor 60, the second feature extractor 70, and the third feature extractor 80, respectively. As a result, the feature extraction unit 2040 acquires the feature amounts of the terrestrial images 110, the aerial images 120, and the map images 130. The update unit 2060 calculates a combination loss based on the feature amount of the terrestrial image 110, the feature amount of the aerial image 120, and the feature amount of the map image 130, and updates the feature extractor set 50 based on the combination loss.

[0080] <Example of hardware configuration> The training device 2000 of the second embodiment may be realized by one or more computers, as in the first embodiment. Therefore, the hardware configuration of the training device 2000 of the second embodiment can be represented in FIG. 4, as in the first embodiment. However, the storage device 1080 of the second embodiment includes a program for implementing the training device 2000 of the second embodiment.

[0081] <Processing flow> FIG. 13 is a flowchart showing an example of the processing flow of the training device 2000 according to the second embodiment. The training device 2000 may execute the processing shown in FIG. 13 in addition to the processing shown in FIG. 5. The extension unit 2080 generates extended training data 100 from the training data 10 (S202). The feature extraction unit 2040 inputs the terrestrial image 110 to the first feature extractor 60 to acquire the feature quantities of the terrestrial image 110 (S204). The feature extraction unit 2040 inputs the aerial image 120 to the second feature extractor 70 to acquire the feature quantities of the aerial image 120 (S206). The feature extraction unit 2040 inputs the map image 130 to the third feature extractor 80 to acquire the feature quantities of the map image 130 (S208). The update unit 2060 calculates the combination loss based on the acquired feature quantities (S210). The update unit 2060 updates the first feature extractor 60, the second feature extractor 70, and the third feature extractor 80 based on the combination loss (S212).

[0082] Similar to the processing flow performed by the training device 2000 of embodiment 1, the processing flow performed by the training device 2000 of embodiment 2 is not limited to that shown in Fig. 13. For example, the extraction of features from the terrestrial captured image 110 (S204), the extraction of feature amounts from the aerial captured image 120 (S206), and the extraction of feature amounts from the map image 130 (S208) may be performed in an order different from that shown in Fig. 13, or may be performed in parallel.

[0083] The training device 2000 of embodiment 2 may use a plurality of augmented training data 100 to train the feature extractor set 50 in the same manner as when the training device 2000 uses a plurality of training data 10. Also, when the training device 2000 of embodiment 2 performs batch training on the feature extractor set 50, the training device 2000 may calculate and aggregate a combined loss from the training data 10 and the augmented training data 100.

[0084] <Data Extension: S202> The augmentation unit 2080 performs data augmentation on the training data 10 to generate augmented training data 100 from the training data 10 (S202). The data augmentation by the augmentation unit 2080 includes image blending of the aerial image 120 and the map image 130. An example of data augmentation will be described in detail below.

[0085] The expansion unit 2080 may identify one or more pairs of partial images 32 and 42, referred to as "partial image pairs," that are the target of image blending. The partial images 32 and 42 of a partial image pair are disposed in the same position as each other and have the same shape and size as each other. In this manner, the expansion unit 2080 may identify each partial image pair by identifying the position, shape, and size of each partial image pair.

[0086] A partial image pair can be represented by a tuple Ai=(Pi, SHi, SZi), where i denotes the identifier of the partial image pair, Ai denotes the i-th partial image pair, Pi denotes the position of partial image 32 and partial image 42 of the i-th partial image pair, SHi denotes the shape of partial image 32 and partial image 42 of the i-th partial image pair, and SZi denotes the size of partial image 32 and partial image 42 of the i-th partial image pair. In this case, partial image 32 of Ai is located at position Pi in the aerial image 30 and has shape SHi and size SZi. Similarly, partial image 42 of Ai is located at position Pi on the map image 40 and has shape SHi and size SZi.

[0087] The shape of a partial image may be any predetermined shape, such as a rectangle or a circle. There are various methods for expressing the position and size of a partial image, and one of these methods can be applied to a partial image pair. Assume that the shape of a partial image is rectangular. In this case, the position of a partial image can be represented by the coordinates of one of its vertices (e.g., the upper left vertex), and the size of a partial image can be represented by a pair of its width and height (i.e., the length of the long side and the length of the short side).

[0088] In another example, the shape of the partial image may be circular. In this case, the position of the partial image may be represented by the coordinates of its center, and the size may be represented by its radius or diameter.

[0089] The partial image pairs may be predefined or may be dynamically identified by the extension unit 2080. In the former case, information indicating the definition of each partial image pair, for example, a tuple (Pi, SHi, SZi), is pre-stored in a storage unit accessible to the extension unit 2080. The extension unit 2080 acquires this information from the storage unit and identifies the partial image pairs to be used for data extension. The extension unit 2080 may use all predetermined partial image pairs for data extension, or may use one or more partial image pairs from among the predetermined partial image pairs for data extension. In the latter case, the number of partial image pairs to be selected may be predefined or may be dynamically identified (for example, randomly identified).

[0090] When the partial image pairs are dynamically identified, the extension unit 2080 may dynamically identify (for example, randomly identify) the number of partial image pairs, and may dynamically identify (for example, randomly identify) the position, shape, and size of each partial image pair.

[0091] The number of partial image pairs may be predefined. In this case, the expansion unit 2080 may dynamically identify the position, shape, and size of each partial image pair to identify a predetermined number of partial image pairs.

[0092] In addition, one or more of the position, shape, and size may be defined in advance. The shape of the partial image pair is assumed to be defined in advance as a rectangle. In this case, the expansion unit 2080 identifies the rectangular partial image pair by specifying the position and size of the rectangle.

[0093] After identifying the partial image pairs, the extension unit 2080 performs image blending for each partial image pair to generate an extended image 140. In image blending, partial image 32 and partial image 42 are combined at a blend ratio Ra:Rm where Ra+Rm=1. The blend ratio may be common to all partial image pairs, or may be individually identified for each partial image pair. The blend ratio may be predefined or may be dynamically identified, such as randomly.

[0094] After generating the augmented image 140, the augmentation unit 2080 may replace the partial image 32 with the augmented image 140 to generate the aerial image 120, or may replace the partial image 42 with the augmented image 140 to generate the map image 130, or may do both. For each augmented image 140, the partial image that replaces it may be predefined or dynamically selected (e.g., randomly selected).

[0095] Note that for one or more partial image pairs, the extension unit 2080 may not use either the partial image 32 or the partial image 42 of the partial image pair to generate the extended image 140. In other words, image blending may be performed at a blend ratio of 1:0 (Ra=1, Rm=0) or 0:1 (Ra=0, Rm=1).

[0096] The augmented image 140 is generated with a blending ratio of 1:0 (i.e., the partial image 42 is not used to generate the augmented image 140), and the partial image 42 is replaced by the augmented image 140, which means that a portion of the map image 40 is completely replaced by the aerial image 30.

[0097] This process can be performed without image blending. Specifically, the augmentation unit 2080 may extract the partial image 32 as the augmented image 140 and perform image replacement on the map image 40, replacing the partial image 42 with the augmented image 140.

[0098] 14 shows a case where a part of a map image 40 is replaced with an aerial image 30. In the map image 130, a partial image 42 is replaced with an extended image 140 that is equivalent to the partial image 32.

[0099] Similarly, if the augmented image 140 is generated with a blending ratio of 0:1 (i.e., the partial image 32 is not used to generate the augmented image 140), and the partial image 32 is replaced with the augmented image 140, this means that a portion of the aerial image 30 is completely replaced with the map image 40.

[0100] This process can also be performed without image blending. Specifically, the extension unit 2080 can extract the partial image 42 as the extended image 140 and perform image replacement on the aerial image 30, replacing the partial image 32 with the extended image 140.

[0101] 15 shows a case where a part of the aerial image 30 is replaced with a map image 40. In the aerial image 120, the partial image 32 is replaced with an extended image 140 that is equivalent to the partial image 42.

[0102] It should be noted that in addition to the image blending described above, the augmentation unit 2080 may perform one or more methods for data augmentation on the aerial imagery 30, the map imagery 40, or both. Examples of these methods are disclosed in U.S. Patent No. 6,277,949 and U.S. Patent No. 6,277,949.

[0103] <Output from Training Device 2000> The training device 2000 of the second embodiment may output information similar to the information output by the training device 2000 of the first embodiment. In addition, the training device 2000 of the second embodiment may output the extended training data 100.

[0104] The program can be stored and provided to a computer using various types of non-transitory computer-readable media. Non-transitory computer-readable media include various types of tangible storage media. Examples of non-transitory computer-readable media include magnetic recording media (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical recording media (e.g., magneto-optical disks), CD-ROMs, CD-Rs, CD-R / Ws, and semiconductor memories (e.g., mask ROMs, programmable ROMs (PROMs), erasable PROMs (EPROMs), flash ROMs, and RAMs). The program may also be provided to a computer by various types of transitory computer-readable media. Examples of transitory computer-readable media include electrical signals, optical signals, and electromagnetic waves. The transitory computer-readable media can provide the program to a computer via a wired communication path such as an electric wire or optical fiber, or via a wireless communication path.

[0105] Although the present invention has been described above with reference to the embodiments, the present invention is not limited to the above embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present invention within the scope of the present invention.

[0106] All or part of the above-described embodiments can also be described as, but not limited to, the following supplementary notes. <Additional Notes> (Appendix 1) 1. A training device comprising: at least one memory configured to store instructions; at least one processor, wherein the at least one processor: acquiring training data including a first terrestrial image, a first aerial image, and a first map image; inputting the first ground image to a first feature extractor to extract a feature amount of the first ground image; inputting the first aerial image into a second feature extractor to extract features of the first aerial image; inputting the first map image to a third feature extractor to extract features of the first map image; calculating a combination loss based on the feature amount of the first ground-photographed image, the feature amount of the first aerial-photographed image, and the feature amount of the first map image; updating the first feature extractor, the second feature extractor, and the third feature extractor based on the combination loss; a training device configured to execute the instructions to: (Appendix 2) The calculation of the coupling loss is calculating a first loss based on the feature amount of the first terrestrial image and the feature amount of the first aerial image; calculating a second loss based on the feature amount of the first ground image and the feature amount of the first map image; calculating a third loss based on the feature amount of the first aerial image and the feature amount of the first map image; and combining the first loss, the second loss, and the third loss into the combined loss. (Appendix 3) 3. The training device of claim 2, wherein the combined loss is a weighted sum of the first loss, the second loss, and the third loss. (Appendix 4) The at least one process comprises: further configured to generate augmented training data based on the training data, the augmented training data including a second terrestrial image, a second aerial image, and a second map image; The generating of the augmented training data comprises: generating an augmented image by blending a portion of the first aerial image with a portion of the first map image; 4. The training device according to any one of appendices 1 to 3, comprising: replacing the portion of the first aerial image with the augmented image to obtain the second aerial image; replacing the portion of the first map image with the augmented image to obtain the second map image; or performing both of these. (Appendix 5) The generating of the augmented training data comprises: Identifying one or more partial image pairs that are pairs of a portion of the first aerial image and a portion of the first map image; generating the augmented image for each partial image pair; the portion of the first aerial image and the portion of the first map image included in the partial image pair have the same position, shape, and size; 5. The training device of claim 4, wherein the identifying the partial image pairs includes identifying the positions, the shapes, and the sizes of the partial image pairs. (Appendix 6) The at least one processor: inputting the second terrestrial image to a first feature extractor to extract features of the second terrestrial image; inputting the second aerial image into a second feature extractor to extract features of the second aerial image; inputting the second map image to a third feature extractor to extract features of the second map image; calculating a second combination loss based on the feature amount of the second ground-photographed image, the feature amount of the second aerial-photographed image, and the feature amount of the second map image; 5. The training apparatus of claim 4, further configured to: update the first feature extractor, the second feature extractor, and the third feature extractor based on the second combination loss. (Appendix 7) 1. A computer-implemented training method comprising: acquiring training data including a first terrestrial image, a first aerial image, and a first map image; inputting the first ground image to a first feature extractor to extract a feature amount of the first ground image; inputting the first aerial image into a second feature extractor to extract features of the first aerial image; inputting the first map image to a third feature extractor to extract features of the first map image; calculating a combination loss based on the feature amount of the first ground-photographed image, the feature amount of the first aerial-photographed image, and the feature amount of the first map image; updating the first feature extractor, the second feature extractor, and the third feature extractor based on the combination loss. (Appendix 8) The calculation of the coupling loss is calculating a first loss based on the feature amount of the first terrestrial image and the feature amount of the first aerial image; calculating a second loss based on the feature amount of the first ground image and the feature amount of the first map image; calculating a third loss based on the feature amount of the first aerial image and the feature amount of the first map image; and combining the first loss, the second loss, and the third loss into the combined loss. (Appendix 9) 9. The training method of claim 8, wherein the combined loss is a weighted sum of the first loss, the second loss, and the third loss. (Appendix 10) generating extended training data based on the training data, the extended training data including a second terrestrial image, a second aerial image, and a second map image; The generating of the augmented training data comprises: generating an augmented image by blending a portion of the first aerial image with a portion of the first map image; A training method described in any one of Appendices 7 to 9, comprising replacing the portion of the first aerial image with the augmented image to obtain the second aerial image, replacing the portion of the first map image with the augmented image to obtain the second map image, or performing both. (Appendix 11) The generating of the augmented training data comprises: Identifying one or more partial image pairs that are pairs of a portion of the first aerial image and a portion of the first map image; generating the augmented image for each partial image pair; the portion of the first aerial image and the portion of the first map image included in the partial image pair have the same position, shape, and size; 11. The training method of claim 10, wherein identifying the partial image pairs includes identifying the positions, the shapes, and the sizes of the partial image pairs. (Appendix 12) inputting the second ground image to a first feature extractor to extract a feature amount of the second ground image; inputting the second aerial image into a second feature extractor to extract features of the second aerial image; inputting the second map image to a third feature extractor to extract features of the second map image; calculating a second combination loss based on the feature amount of the second ground-photographed image, the feature amount of the second aerial-photographed image, and the feature amount of the second map image; 11. The training method of claim 10, further comprising updating the first feature extractor, the second feature extractor, and the third feature extractor based on the second combination loss. (Appendix 13) 1. A non-transitory computer-readable storage medium, comprising: acquiring training data including a first terrestrial image, a first aerial image, and a first map image; inputting the first ground image to a first feature extractor to extract a feature amount of the first ground image; inputting the first aerial image into a second feature extractor to extract features of the first aerial image; inputting the first map image to a third feature extractor to extract features of the first map image; calculating a combination loss based on the feature amount of the first ground-photographed image, the feature amount of the first aerial-photographed image, and the feature amount of the first map image; and updating the first feature extractor, the second feature extractor, and the third feature extractor based on the combination loss. (Appendix 14) The calculation of the coupling loss is calculating a first loss based on the feature amount of the first terrestrial image and the feature amount of the first aerial image; calculating a second loss based on the feature amount of the first ground image and the feature amount of the first map image; calculating a third loss based on the feature amount of the first aerial image and the feature amount of the first map image; and combining the first loss, the second loss, and the third loss into the combined loss. (Appendix 15) 15. The storage medium of claim 14, wherein the combined loss is a weighted sum of the first loss, the second loss, and the third loss. (Appendix 16) The program causes the computer to: generating extended training data based on the training data, the extended training data including a second terrestrial image, a second aerial image, and a second map image; The generating of the augmented training data comprises: generating an augmented image by blending a portion of the first aerial image with a portion of the first map image; A storage medium described in any one of Appendices 13 to 15, comprising replacing the portion of the first aerial image with the extended image to obtain the second aerial image, replacing the portion of the first map image with the extended image to obtain the second map image, or performing both. (Appendix 17) The generating of the augmented training data comprises: Identifying one or more partial image pairs that are pairs of a portion of the first aerial image and a portion of the first map image; generating the augmented image for each partial image pair; the portion of the first aerial image and the portion of the first map image included in the partial image pair have the same position, shape, and size; 17. The storage medium of claim 16, wherein the identifying the partial image pair includes identifying the position, the shape, and the size of the partial image pair. (Appendix 18) The program causes the computer to: inputting the second ground image to a first feature extractor to extract a feature amount of the second ground image; inputting the second aerial image into a second feature extractor to extract features of the second aerial image; inputting the second map image to a third feature extractor to extract features of the second map image; calculating a second combination loss based on the feature amount of the second ground-photographed image, the feature amount of the second aerial-photographed image, and the feature amount of the second map image; 17. The storage medium of claim 16, further causing the method to perform: updating the first feature extractor, the second feature extractor, and the third feature extractor based on the second combining loss. [Explanation of symbols]

[0107] 10 Training data 20 Ground-based images 30 Aerial Images 40 map images 50 Feature Extractor Set 60 First feature extractor 70 Second feature extractor 80 Third feature extractor 100 extended training data 110 Ground-based imagery 120 aerial images 130 map images 140 Extended Images 200 Geolocalization System 210 Ground Information 220 Aviation Information 230 Location information 240 response 250 Matching Device 300 Location Database 1000 computers 1020 Bus 1040 processor 1060 memory 1080 storage device 1100 Input / Output Interface 1120 Network Interface 2000 training equipment 2020 Acquisition Department 2040 Feature Extraction Unit 2060 Update Department 2080 Extension

Claims

1. An acquisition means for acquiring training data including a first ground image, a first aerial image, and a first map image; a feature extraction means for inputting the first ground-photographed image to a first feature extractor to extract feature quantities of the first ground-photographed image, inputting the first aerial-photographed image to a second feature extractor to extract feature quantities of the first aerial-photographed image, and inputting the first map image to a third feature extractor to extract feature quantities of the first map image; an updating means for calculating a combination loss based on the feature amounts of the first ground-photographed image, the feature amounts of the first aerial-photographed image, and the feature amounts of the first map image, and updating the first feature extractor, the second feature extractor, and the third feature extractor based on the combination loss.

2. The calculation of the coupling loss is calculating a first loss based on the feature amount of the first ground-photographed image and the feature amount of the first aerial-photographed image; calculating a second loss based on the feature amount of the first ground image and the feature amount of the first map image; calculating a third loss based on the feature amount of the first aerial image and the feature amount of the first map image; and combining the first loss, the second loss, and the third loss into the combined loss.

3. The training device of claim 2 , wherein the combined loss is a weighted sum of the first loss, the second loss, and the third loss.

4. having an extension means for generating extended training data based on the training data, the extended training data includes a second ground image, a second aerial image, and a second map image; The generating of the augmented training data comprises: blending a portion of the first aerial image with a portion of the first map image to generate an augmented image; The training device according to any one of claims 1 to 3, further comprising: replacing the portion of the first aerial image with the augmented image to obtain the second aerial image; replacing the portion of the first map image with the augmented image to obtain the second map image; or performing both of these.

5. The generating of the augmented training data comprises: Identifying one or more partial image pairs, each pair being a portion of the first aerial image and a portion of the first map image; generating the augmented image for each partial image pair; the portion of the first aerial image and the portion of the first map image included in the partial image pair have the same position, shape, and size; The training device of claim 4 , wherein the identifying the partial image pairs includes identifying the positions, the shapes, and the sizes of the partial image pairs.

6. The feature extraction means inputting the second terrestrial image to the first feature extractor to extract features of the second terrestrial image; inputting the second aerial image into the second feature extractor to extract features of the second aerial image; inputting the second map image to the third feature extractor to extract features of the second map image; calculating a second combination loss based on the feature amount of the second ground-photographed image, the feature amount of the second aerial-photographed image, and the feature amount of the second map image; The training device of claim 4 , further comprising: updating the first feature extractor, the second feature extractor, and the third feature extractor based on the second combination loss.

7. 1. A computer-implemented training method comprising: an acquisition step of acquiring training data including a first ground image, a first aerial image, and a first map image; a feature extraction step of inputting the first ground-photographed image to a first feature extractor to extract feature quantities of the first ground-photographed image, inputting the first aerial-photographed image to a second feature extractor to extract feature quantities of the first aerial-photographed image, and inputting the first map image to a third feature extractor to extract feature quantities of the first map image; an updating step of calculating a combination loss based on the feature amounts of the first ground-image, the feature amounts of the first aerial image, and the feature amounts of the first map image, and updating the first feature extractor, the second feature extractor, and the third feature extractor based on the combination loss.

8. The calculation of the coupling loss is calculating a first loss based on the feature amount of the first ground-photographed image and the feature amount of the first aerial-photographed image; calculating a second loss based on the feature amount of the first ground image and the feature amount of the first map image; calculating a third loss based on the feature amount of the first aerial image and the feature amount of the first map image; and combining the first loss, the second loss, and the third loss into the combined loss.

9. An acquisition step of acquiring training data including a first ground image, a first aerial image, and a first map image; a feature extraction step of inputting the first ground-photographed image to a first feature extractor to extract feature quantities of the first ground-photographed image, inputting the first aerial-photographed image to a second feature extractor to extract feature quantities of the first aerial-photographed image, and inputting the first map image to a third feature extractor to extract feature quantities of the first map image; a program that causes a computer to execute the following: calculating a combination loss based on the feature amounts of the first ground-image, the feature amounts of the first aerial image, and the feature amounts of the first map image; and updating the first feature extractor, the second feature extractor, and the third feature extractor based on the combination loss.

10. The calculation of the coupling loss is calculating a first loss based on the feature amount of the first ground-photographed image and the feature amount of the first aerial-photographed image; calculating a second loss based on the feature amount of the first ground image and the feature amount of the first map image; calculating a third loss based on the feature amount of the first aerial image and the feature amount of the first map image; and combining the first loss, the second loss, and the third loss into the combined loss.

Citation Information

Patent Citations

  • Image processing method, program for executing the method, storage medium, imaging apparatus, and image processing system

    JP2009258953A

  • System for specifying detection object position

    JP2019185689A

  • Image augmentation apparatus, control method, and non-transitory computer-readable storage medium

    WO2022034678A1

  • Image augmentation apparatus, control method, and non-transitory computer-readable storage medium

    WO2022044105A1