Information processing method, information processing apparatus, and program

By determining overlap points in the image and dividing overlapping areas and non-overlapping areas for training, the generated model can extract image feature quantities with higher accuracy, solving the problem of insufficient extraction accuracy of image feature quantities in the prior art, and improving the accuracy of image retrieval and visual positioning systems.

CN120266173APending Publication Date: 2025-07-04SONY GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202380080428.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-29
Filing Date
2023-10-25
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, the accuracy of extracting image feature amounts from images using models obtained through learning is insufficient, resulting in low accuracy of image retrieval and visual positioning systems.

Method used

By determining the overlap points between images, dividing the image into overlapping areas and non-overlapping areas, and training based on the feature amount of the overlapping areas and non-overlapping areas, a model that can extract the feature amount of the image with higher accuracy is generated.

Benefits of technology

Improve the accuracy of image retrieval and visual positioning systems, reduce the cost of manually labeling overlapping areas, and improve the accuracy of position and posture estimation of imaging equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0005410876540000011
    Figure HDA0005410876540000011
  • Figure HDA0005410876540000021
    Figure HDA0005410876540000021
  • Figure HDA0005410876540000022
    Figure HDA0005410876540000022
Patent Text Reader

Abstract

An information processing apparatus includes circuitry configured to: receive a first image; receiving a model from the learning device; and outputting a first position and a first posture on the basis of a first image feature amount extracted from the first image and the model, the model is obtained by determining at least one overlap point between the second image and the third image, dividing the second image into an overlap region and a non-overlap region in response to determining the at least one overlap point, the overlapping region includes at least one overlapping point, and training is performed based on a second image feature amount corresponding to the overlapping region and a third image feature amount corresponding to the non-overlapping region.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to Related Applications

[0002] This application claims the benefit of Japanese Patent Application No. JP 2022-189993, filed on Nov. 29, 2022, the entire contents of which are incorporated herein by reference. Technical Field

[0003] The present disclosure relates to an information processing method, an information processing apparatus, and a program. Background Art

[0004] Recently, techniques for extracting feature amounts from images have been used. For example, techniques for extracting feature amounts from images are used in image retrieval techniques. In image retrieval techniques, a DB image similar to a query image is retrieved from a plurality of DB images pre-registered in a database (DB). At this time, it is determined whether the query image and the DB image are similar to each other based on whether the image feature amounts extracted from the query image and the image feature amounts extracted from the DB image are close to each other.

[0005] NPL 1 discloses an embodiment of an image retrieval technique. With the image retrieval technique disclosed in NPL 1, a DB image is divided into a plurality of regions each having a fixed size, and it is checked whether each of the plurality of regions overlaps with the query image based on whether the image feature amounts extracted from each of the plurality of regions are close to the image feature amounts extracted from the query image. Then, the regions in the DB image determined to overlap with the query image (the regions will also be referred to as “overlapping regions” hereinafter) are preferentially used to assist learning.

[0006] Note that the overlapping regions in the DB image are regions similar to part or all of the regions of the query image.

[0007] A model obtained by learning is used to extract image feature amounts. For example, a model obtained by learning can be implemented using a deep neural network (DNN) or the like. Thus, as an embodiment, a model obtained by learning can extract image feature amounts from an image with higher accuracy, which can contribute to improving the accuracy of image retrieval.

[0008] Citation List

[0009] Non-Patent Literature

[0010] NPL 1: Yixiao Ge, et al. “Self-supervising Fine-grained Region Similarities for Large-scale Image Localization”, ECCV 2020 Summary of the Invention

[0011] Technical Problem

[0012] In view of the above, a technique is needed that can use a learned model to extract image feature amounts from images with higher accuracy.

[0013] Solution to the Problem

[0014] According to the present disclosure, there is provided an information processing apparatus including a circuit system configured to: receive a first image; receive a model from a learning device; and output a first position and a first pose based on a first image feature amount extracted from the first image and the model, wherein the model is obtained by: determining at least one overlapping point between a second image and a third image; dividing the second image into an overlapping region and a non-overlapping region in response to determining the at least one overlapping point, the overlapping region including the at least one overlapping point; and performing training based on a second image feature amount corresponding to the overlapping region and a third image feature amount corresponding to the non-overlapping region.

[0015] Furthermore, according to the present disclosure, an information processing method includes: receiving a first image; receiving a model from a learning device; and outputting a first position and a first pose based on a first image feature amount extracted from the first image and the model, wherein the model is obtained by: determining at least one overlapping point between a second image and a third image; dividing the second image into an overlapping region and a non-overlapping region in response to determining the at least one overlapping point, the overlapping region including the at least one overlapping point; and performing training based on a second image feature amount corresponding to the overlapping region and a third image feature amount corresponding to the non-overlapping region.

[0016] Furthermore, according to the present disclosure, there is provided a non-transitory computer-readable medium having a program thereon that, when executed by a computer, causes the computer to execute an information processing method including: receiving a first image; receiving a model from a learning device; and outputting a first position and a first pose based on a first image feature amount extracted from the first image and the model, wherein the model is obtained by: determining at least one overlapping point between a second image and a third image; dividing the second image into an overlapping region and a non-overlapping region in response to determining the at least one overlapping point, the overlapping region including the at least one overlapping point; and performing training based on a second image feature amount corresponding to the overlapping region and a third image feature amount corresponding to the non-overlapping region. Brief Description of the Drawings

[0017] Figure 1 is a diagram showing an exemplary configuration of an information processing system according to an embodiment of the present disclosure.

[0018] Figure 2It is a diagram for explaining an embodiment of image retrieval.

[0019] Figure 3 It is a diagram showing an exemplary operation of estimating device position / pose information regarding the imaging device 110 when capturing an inference query image G3.

[0020] Figure 4 It is a diagram for explaining a method for learning a DNN for extracting an image feature amount according to a comparative embodiment.

[0021] Figure 5 It is a diagram for explaining the problems of the comparative embodiment.

[0022] Figure 6 It is a diagram showing the flow of a method for learning a DNN for extracting an image feature amount according to an embodiment of the present disclosure.

[0023] Figure 7 It is a diagram for explaining a method for learning a DNN for extracting an image feature amount according to an embodiment of the present disclosure.

[0024] Figure 8 It is a diagram showing an example functional configuration of the terminal device 10 according to an embodiment of the present disclosure.

[0025] Figure 9 It is a diagram showing an example functional configuration of the inference device 30 according to an embodiment of the present disclosure.

[0026] Figure 10 It is a diagram showing a specific embodiment configuration of the image retrieval unit 310.

[0027] Figure 11 It is a diagram showing a specific embodiment configuration of the feature point matching unit 320.

[0028] Figure 12 It is a diagram showing an example functional configuration of the learning device 20 according to an embodiment of the present disclosure.

[0029] Figure 13 It is a diagram showing a specific embodiment configuration of the three-dimensional recovery unit 210 according to the first overlapping point extraction method.

[0030] Figure 14 It is a diagram for explaining the respective functions of the position / pose estimation unit 212 and the depth estimation unit 214.

[0031] Figure 15 It is a diagram showing a specific embodiment configuration of the three-dimensional recovery unit 210 according to the second overlapping point extraction method.

[0032] Figure 16It is a view of a three-dimensional point group observed obliquely from the side.

[0033] Figure 17 It is a view of a three-dimensional point group observed from above.

[0034] Figure 18 It is a view showing an embodiment of a grid.

[0035] Figure 19 It is a view for explaining the case of extracting overlapping points based on grid information.

[0036] Figure 20 It is a view showing a specific embodiment configuration of the feature amount extraction unit 530 according to a comparative example.

[0037] Figure 21 It is a view showing a specific embodiment configuration of the feature amount extraction unit 230 according to an embodiment of the present disclosure.

[0038] Figure 22 It is a view for explaining a first modification example.

[0039] Figure 23 It is a view for explaining a second modification example.

[0040] Figure 24 It is a view for explaining a third modification example.

[0041] Figure 25 It is a view showing an embodiment of an overlapping region and a non-overlapping region according to a third modification example.

[0042] Figure 26 It is a block diagram showing an exemplary hardware configuration of the information processing device 900. Detailed Description of Specific Embodiments

[0043] Hereinafter, preferred embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. Note that in this specification and the drawings, components having substantially the same functional configuration are denoted by the same reference numerals, and their descriptions will not be repeated.

[0044] In addition, in the specification and the drawings, a plurality of components having substantially the same or similar functional configurations may be distinguished by different letters added to the same reference numeral. However, in cases where it is not necessary to specifically distinguish a plurality of components having substantially the same or similar functional configurations from each other, only the same reference numeral is added thereto.

[0045] Note that the description will be made in the following order.

[0046] 0. Overview

[0047] 1. Details of Embodiments

[0048] 1.1. Exemplary functional configurations of terminal devices

[0049] 1.2. Exemplary functional configurations of inference devices

[0050] 1.3. Exemplary functional configurations of learning devices

[0051] 2. Various modifications

[0052] 3. Exemplary hardware configurations

[0053] 4. Conclusion

[0054] <0. Overview>

[0055] First, refer to Figures 1 to 7 to describe an overview of embodiments of the present disclosure.

[0056] Figure 1 is a diagram showing an exemplary configuration of an information processing system according to an embodiment of the present disclosure. As Figure 1 shown, the information processing system 1 according to an embodiment of the present disclosure includes a terminal device 10, a learning device 20, and an inference device 30. The terminal device 10, the learning device 20, and the inference device 30 are each connected to a network 40 and are designed to be able to communicate with each other via the network 40.

[0057] First, the learning device 20 generates a model (which is an image feature quantity extraction unit) by training based on a query image for learning (hereinafter also referred to as a "learning query image") and one or more DB images for learning (hereinafter also referred to as "learning DB images"), and the model extracts image feature quantities from a query image for inference (hereinafter also referred to as an "inference query image"). The inference query image may correspond to an embodiment of a first image. In an embodiment of the present disclosure, it is mainly assumed that the model generated by the learning device 20 is implemented using a DNN. However, the model can be generated by learning using some other machine learning algorithms. The learning device 20 sends the model generated by learning to the inference device 30 via the network 40. The inference device 30 receives the model sent from the learning device 20 via the network 40.

[0058] The inference device 30 includes a DB. In the DB, one or more DB images for inference (hereinafter also referred to as "inference DB images") are pre-registered. In the following description, in some cases, one or more inference DB images will be referred to as "all inference DB images".

[0059] In addition, the image feature amounts extracted from the inference DB images, and the information indicating the position and orientation of the imaging device when the inference DB images are captured are associated with the corresponding inference DB images and registered in the DB. In the following description, the position and orientation will also be referred to as "position / orientation". Further, in the following description, the position / orientation information of the imaging device when capturing an image will also be simply referred to as "device position / orientation information corresponding to the image".

[0060] The terminal device 10 includes an imaging device. The terminal device 10 sends an inference query image captured by the imaging device to the inference device 30 via the network 40. The inference device 30 receives the inference query image via the network 40. Then, the inference device 30 estimates the similarity between the inference query image and each inference DB image, and based on this similarity, performs an image retrieval for retrieving an inference DB image similar to the inference query image. Now refer to Figure 2 Briefly describe an embodiment of the image retrieval.

[0061] Figure 2 is a diagram for explaining an embodiment of the image retrieval. Refer to Figure 2 , feature points F101 to F103 exist in the real space. In addition, each inference DB image, the image feature amounts extracted from each inference DB image, and the device position / orientation information corresponding to each inference DB image are pre-registered. The inference device 30 extracts the image feature amounts from the inference query image G3, and calculates the difference between the image feature amounts extracted from the inference query image G3 and the image feature amounts extracted from each inference DB image.

[0062] Note that the difference between the feature amounts can be the difference between the vectors representing the image feature amounts. The inference device 30 ranks one or more inference DB images such that the image feature amounts with smaller differences extracted from the inference DB images are ranked higher compared to the image feature amounts extracted from the inference query image G3.

[0063] The inference query image G3 captured by the imaging device 110 in the position / orientation C3 includes feature points F301 to F303 corresponding to the feature points F101 to F103. In addition, the inference DB image G4 captured by the imaging device 814 in the position / orientation C4 includes feature points F401 to F403 corresponding to the feature points F101 to F103. Note that the imaging device 110 and the imaging device 814 can be different imaging devices, or can be the same imaging device performing imaging at different times.

[0064] At this time, the feature points F301 to F303 that appear in the inference query image G3, and the feature points F401 to F403 that appear in the inference DB image G4 correspond to the same feature points F101 to F103 existing in the real space. Therefore, the difference between the image feature amounts extracted from the inference query image G3 and the image feature amounts extracted from the inference DB image G4 is small, and the inference DB image G4 is regarded as being ranked higher in the order determined by the inference device 30.

[0065] Then, among the image feature amounts extracted from all the inference DB images, the inference device 30 designates a predetermined number of image feature amounts in ascending order of the difference from the image feature amounts extracted from the inference query image. In the following description, the inference DB images corresponding to the respective image feature amounts of this predetermined number of image feature amounts will also be referred to as "higher-order inference DB images".

[0066] Next, with reference to Figure 3 , the estimation of the device position / pose information of the imaging device 110 when capturing the inference query image G3 is described.

[0067] Figure 3 is a diagram showing an exemplary operation of estimating the device position / pose information of the imaging device 110 when capturing the inference query image G3. As Figure 3 shown, image retrieval based on the inference query image (S11) is performed. As described above, in the image retrieval, the higher-order inference DB images corresponding to the inference query image are obtained from the DB by the inference device 30.

[0068] Subsequently, the inference device 30 performs matching of the feature points between the inference query image and the higher-order inference DB image (S12). Therefore, the corresponding pixels between the inference query image and the higher-order inference DB image are obtained as corresponding point pairs.

[0069] Subsequently, the inference device 30 estimates the relative position / pose of the imaging device when capturing the inference query image based on the two-dimensional coordinates of the corresponding point pairs between the inference query image and the higher-order inference DB image, and the three-dimensional coordinates of the feature points of the corresponding point pairs in the higher-order inference DB image, based on the device position / pose of the imaging device when the higher-order inference DB image is captured (S13).

[0070] The inference device 30 estimates the device position / pose of the imaging device when capturing the inference query image based on the device position / pose information corresponding to the higher-order inference DB image and the relative position / pose of the imaging device when capturing the inference query image. For example, the series of operations for estimating the device position / pose information correspond to the relocalization process of a Simultaneous Localization and Mapping (SLAM) system.

[0071] The inference device 30 sends device position / pose information corresponding to an inference query image to the terminal device 10 via the network 40. The terminal device 10 receives the device position / pose information via the network 40. Using the received device position / pose information, the terminal device 10 can perform various processes.

[0072] The service of providing device position / pose in this way is also called a Visual Positioning System (VPS), and can be provided from the learning device 20 and the inference device 30 to the terminal device 10 as a cloud service.

[0073] Here, the terminal device 10 can be a smart phone or the like. At this time, in the terminal device 10, an Augmented Reality (AR) application can be used to superimpose AR objects on the real space with high precision based on the device position / pose information. Alternatively, the terminal device 10 can be an autonomous mobile unit (e.g., such as a drone). In this case, the autonomous mobile unit can move based on the device position / pose information.

[0074] The model generated by the learning device 20 is implemented using a DNN, and image feature amounts are extracted from the inference query image and each inference DB image through this model. Hereinafter, the DNN that extracts image feature amounts from images will also be referred to as the "image feature amount extraction DNN". Contrast learning is used in the learning of the image feature amount extraction DNN. Here, refer to Figure 4 and Figure 5 describe a method for learning the image feature amount extraction DNN according to a comparative example.

[0075] Figure 4 is a diagram for explaining a method for learning the image feature amount extraction DNN according to a comparative example. Refer to Figure 4 shows a learning query image G2 and a learning DB image G1. The learning device 20 obtains the image feature amount E2 output from the DNN based on the input of the learning query image G2 to the DNN. In addition, the learning device 20 obtains the image feature amount E1 output from the DNN based on the input of the learning DB image G1 to the DNN.

[0076] In the feature amount space E, there is an image feature amount E2 extracted from the learning query image G2. In addition, in the feature amount space E, there is an image feature amount E1 extracted from the learning DB image G1.

[0077] In a comparative example, when a ground truth label is attached to the learning DB image G1, the DNN is learned such that the image feature amount E1 extracted from the learning DB image G1 approaches the image feature amount E2 extracted from the learning query image G2 (such that the image feature amount E1 moves in the direction D1). On the other hand, when no ground truth label is attached to the learning DB image G1, the DNN is learned such that the image feature amount E1 extracted from the learning DB image G1 moves away from the image feature amount E2 extracted from the learning query image G2 (such that the image feature amount E1 moves in the direction D2).

[0078] Figure 5 is a diagram for explaining the problems of the comparative example. In Figure 5 In the illustrated example, an overlapping region G11 between the learning query image G2 and the learning DB image G1 in the learning DB image G1, and a non-overlapping region G12 that is a region other than the overlapping region G11 in the learning DB image G1 are shown. In the comparative example, it is necessary to know in advance whether the learning DB image G1 is correct. Therefore, there is a first problem of the labor cost for creating the ground truth label.

[0079] In addition, when the learning DB image G1 is correct, the DNN is learned such that the image feature amount corresponding to the non-overlapping region G12 that does not overlap with the learning query image G2 in the learning DB image G1 approaches the image feature amount corresponding to the learning query image G2. Therefore, in the comparative example, there is a second problem that confusion occurs in the learning of the DNN and the learning of the DNN cannot be effectively performed.

[0080] Next, refer to Figure 6 and Figure 7 to describe a method for learning a DNN for image feature amount extraction according to an embodiment of the present disclosure.

[0081] Figure 6 is a diagram showing a process of a method for learning a DNN for image feature amount extraction according to an embodiment of the present disclosure. Refer to Figure 6 to show the learning query image G2 and the learning DB image G1. In an embodiment of the present disclosure, for example, the learning device 20 extracts an overlapping region between the learning query image G2 and the learning DB image G1 based on three-dimensional information related to the learning query image G2 and the learning DB image G1. This can solve the problems of the comparative example.

[0082] For example, when extracting three-dimensional information only from an image, a three-dimensional restoration technique for generating a three-dimensional model from the learning query image G2 and the learning DB image G1 can be used to determine the overlapping region. That is, the overlapping region (hereinafter also referred to as "overlapping point") obtained during the three-dimensional restoration can be used to determine the overlapping region (S21).

[0083] Figure 7 is a diagram for explaining a method for learning a DNN for extracting an image feature amount according to an embodiment of the present disclosure. Refer to Figure 7 , overlapping points Q1 are extracted from the learning DB image G1, and an overlapping region G11 and a non-overlapping region G12 are extracted based on the overlapping points Q1. Then, through region division S31, the image feature amount corresponding to the learning DB image G1 is divided into an image feature amount E11 corresponding to the overlapping region G11 and an image feature amount E12 corresponding to the non-overlapping region G12.

[0084] In an embodiment of the present disclosure, the learning device 20 learns the DNN such that the image feature amount E11 corresponding to the overlapping region approaches the image feature amount E2 extracted from the learning query image G2 (such that the image feature amount E11 moves along the direction D11). Moreover, the DNN is learned such that the image feature amount E12 corresponding to the non-overlapping region moves away from the image feature amount E2 extracted from the learning query image G2 (such that the image feature amount E12 moves along the direction D12).

[0085] Therefore, in an embodiment of the present disclosure, the learning device 20 can automatically attach a ground truth label to the overlapping region G11. This solves the first problem of the labor cost for creating the ground truth label.

[0086] In addition, in an embodiment of the present disclosure, learning is performed such that the image feature amount extracted from the overlapping region G11 approaches the image feature amount extracted from the learning query image G2, and the image feature amount extracted from the non-overlapping region G12 moves away from the image feature amount extracted from the learning query image G2. This solves the second problem that confusion occurs in the learning of the DNN and the learning of the DNN cannot be effectively performed.

[0087] The above is an overview of the embodiment of the present disclosure.

[0088] <1. Details of the Embodiment>

[0089] Next, the embodiment of the present disclosure will be described in detail.

[0090] (1.1. Example Functional Configuration of the Terminal Device)

[0091] Next, mainly refer to Figure 8 to describe the example functional configuration of the terminal device 10 according to an embodiment of the present disclosure.

[0092] Figure 8 is a diagram showing the example functional configuration of the terminal device 10 according to an embodiment of the present disclosure. As Figure 8As shown, the terminal device 10 according to an embodiment of the present disclosure includes an imaging device 110, an operation unit 120, a control unit 130, a storage unit 150, and a presentation unit 160.

[0093] (Imaging device 110)

[0094] Based on a predetermined imaging start operation input by the user, the imaging device 110 obtains an inference query image by capturing an image of an imaging range determined according to the position and orientation of the imaging device 110 in the real space. The imaging device 110 outputs the inference query image to the control unit 130. When the imaging device 110 outputs the inference query image to the control unit 130, the control unit 130 executes processing corresponding to the inference query image.

[0095] (Operation unit 120)

[0096] The operation unit 120 has a function of receiving various operations input by the user. For example, the operation unit 120 may be formed of an input device such as a touch panel or a button. The operation unit 120 outputs the operation input by the user to the control unit 130. When the operation unit 120 outputs such an operation to the control unit 130, the control unit 130 executes processing corresponding to the operation.

[0097] (Control unit 130)

[0098] For example, the control unit 130 may be formed of one or more central processing units (CPUs). In the case where the control unit 130 is formed of a processing device such as a CPU, the processing device may be formed of an electronic circuit. The control unit 130 may be formed by a processing device executing a program.

[0099] For example, when an inference query image is input from the imaging device 110, the control unit 130 controls a communication unit (not shown in the figure) so that the inference query image is sent to the inference device 30. Moreover, when receiving device position / orientation information from the inference device 30 through the communication unit (not shown), the control unit 130 controls the presentation unit 160 to arrange an AR object in the augmented reality space based on the device position / orientation information.

[0100] (Storage unit 150)

[0101] The storage unit 150 is a recording medium including a memory, and stores a program to be executed by the control unit 130 and data required for executing the program. Moreover, the storage unit 150 temporarily stores data for calculation to be executed by the control unit 130. The storage unit 150 is formed of a magnetic storage device, a semiconductor storage device, an optical storage device, a magneto-optical storage device, etc.

[0102] (Presentation unit 160)

[0103] The presentation unit 160 presents various information to the user under the control of the control unit 130. For example, the presentation unit 160 is formed with a display and displays an AR object under the control of the control unit 130.

[0104] The above is a description of the exemplary functional configuration of the terminal device 10 according to an embodiment of the present disclosure.

[0105] (1.2. Exemplary functional configuration of the inference device)

[0106] Next, mainly with reference to Figures 9 to 11 describe the exemplary functional configuration of the inference device 30 according to an embodiment of the present disclosure.

[0107] Figure 9 is a diagram showing the exemplary functional configuration of the inference device 30 according to an embodiment of the present disclosure. As Figure 9 shown, the inference device 30 according to an embodiment of the present disclosure includes a control unit 300 and a memory 390. In addition, the control unit 300 includes an image retrieval unit 310, a feature point matching unit 320, a relative position / pose estimation unit 330, and a device position / pose estimation unit 340.

[0108] (Control unit 300)

[0109] For example, the control unit 300 may be formed by one or more central processing units (CPUs). In the case where the control unit 300 is formed by a processing device such as a CPU, the processing device may be formed by an electronic circuit. For the control unit 300, a program is executed by the processing device.

[0110] The control unit 300 extracts an image feature amount (first image feature amount) from the inference query image using a model updated and obtained through training. Then, the control unit 300 estimates the device position / pose information (first position / pose information) about the imaging device 110 when the inference query image is captured based on the image feature amount extracted from the inference query image.

[0111] More specifically, among the image feature amounts of the corresponding inference DB images, the control unit 300 designates a predetermined number of image feature amounts as high-order inference DB images in ascending order of the difference from the image feature amount extracted from the inference query image. Then, the control unit 300 estimates the device position / pose information about the imaging device 110 when the inference query image is captured based on the high-order inference DB images (fourth images) and the inference query image.

[0112] As an example, the control unit 300 specifies, from the high-order inference DB image, a second feature point having a pixel feature amount with the smallest difference from the pixel feature amount at the first feature point in the inference query image. Then, based on the two-dimensional coordinates of the first feature point in the inference query image, the two-dimensional coordinates of the second feature point in the high-order inference DB image, the three-dimensional position information regarding the second feature point, and the device position / pose information corresponding to the high-order inference DB image, the control unit 300 estimates the device position / pose information regarding the imaging device 110 when the inference query image is captured.

[0113] (Memory 390)

[0114] The memory 390 is a recording medium that stores programs to be executed by the control unit 300 and data (such as various databases) required for executing the programs. Moreover, the memory 390 temporarily stores data for calculation to be executed by the control unit 300. The memory 390 is formed of a magnetic storage device, a semiconductor storage device, an optical storage device, a magneto-optical storage device, or the like.

[0115] (Image retrieval unit 310)

[0116] Figure 10 is a diagram showing a specific embodiment configuration of the image retrieval unit 310. As Figure 10 shown, the image retrieval unit 310 includes an image feature amount extraction unit 312 and an image feature amount matching unit 314. Note that, as an example, the image feature amount extraction unit 312 may be a model that has been sent from the learning device 20 and received by a communication unit (not shown) of the inference device 30.

[0117] The image feature amount extraction unit 312 acquires an inference query image from the imaging device 110 included in the terminal device 10. In addition, the image feature amount extraction unit 312 extracts an image feature amount from the inference query image.

[0118] The image feature amount matching unit 314 acquires the image feature amounts extracted from each inference DB image from the memory 390. Then, the image feature amount matching unit 314 calculates the difference between the image feature amount extracted from the inference query image and the image feature amounts extracted from each inference DB image. The image feature amount matching unit 314 ranks all the inference DB images such that the image feature amounts with smaller differences extracted from the inference DB images are ranked higher compared to the image feature amount extracted from the inference query image.

[0119] Among the image feature quantities extracted from all the inference DB images, the image feature quantity matching unit 314 designates a predetermined number of image feature quantities in ascending order of the difference from the image feature quantity extracted from the inference query image. The inference DB images corresponding to the respective image feature quantities among the predetermined number of image feature quantities are high-order inference DB images.

[0120] (Feature point matching unit 320)

[0121] Figure 11 is a diagram showing the configuration of a specific embodiment of the feature point matching unit 320. As Figure 11 shown, the feature point matching unit 320 includes a pixel feature quantity extraction unit 322 and a pixel feature quantity matching unit 324.

[0122] The pixel feature quantity extraction unit 322 acquires an inference query image from the imaging device 110 included in the terminal device 10. In addition, the pixel feature quantity extraction unit 322 extracts pixel feature quantities from the inference query image. More specifically, the pixel feature quantity extraction unit 322 detects a plurality of feature points from the inference query image, and calculates the pixel feature quantity at the feature points based on the peripheral pixel information about each of the plurality of feature points. For the detection of feature points and the extraction of pixel feature quantities, for example, a known method such as Scale-Invariant Feature Transform (SIFT) can be used, or a DNN method can be used.

[0123] The pixel feature quantity matching unit 324 acquires the image feature quantities extracted from each high-order inference DB image from the memory 390. Then, the pixel feature quantity matching unit 324 designates, as a corresponding point pair, two feature points having the smallest difference in pixel feature quantity between the feature points (first feature points) extracted from the inference query image and the feature points (second feature points) extracted from the high-order inference DB images.

[0124] (Relative position / pose estimation unit 330)

[0125] The relative position / pose estimation unit 330 estimates the relative position / pose information about the imaging device 110 at the time of capturing the inference query image, based on the two-dimensional coordinates of each point of the corresponding point pair and the three-dimensional coordinates of the feature points in the high-order inference DB image in the corresponding point pair, with reference to the device position / pose information corresponding to the high-order inference DB image. As a method for estimating the relative position / pose information about the imaging device 110, a known method such as the PnP algorithm is used.

[0126] (Device position / pose estimation unit 340)

[0127] The device position / pose estimation unit 340 estimates the device position / pose information corresponding to the inference query image based on the device position / pose information corresponding to the inference DB image and the relative position / pose information about the imaging device when capturing the inference query image, with reference to the device position / pose information corresponding to the high-order inference DB image. For example, the device position / pose information corresponding to the inference query image can be provided to the terminal device 10.

[0128] The above is a description of the exemplary functional configuration of the inference device 30 according to an embodiment of the present disclosure.

[0129] (1.3. Exemplary functional configuration of the learning device)

[0130] Next, mainly with reference to Figures 12 to 21 describe the exemplary functional configuration of the learning device 20 according to an embodiment of the present disclosure.

[0131] Figure 12 is a diagram showing the exemplary functional configuration of the learning device 20 according to an embodiment of the present disclosure. As Figure 12 shown, the learning device 20 according to an embodiment of the present disclosure includes a control unit 200 and a memory 290. In addition, the control unit 200 includes a three-dimensional recovery unit 210, an overlapping point extraction unit 220, a feature quantity extraction unit 230, a learning loss calculation unit 240, a region determination unit 250, and an update unit 260.

[0132] (Control unit 200)

[0133] For example, the control unit 200 can be formed by one or more central processing units (CPUs). In the case where the control unit 200 is formed by a processing device (such as a CPU), the processing device can be formed by an electronic circuit. The control unit 200 can be formed by a processing device executing a program.

[0134] (Memory 290)

[0135] The memory 290 is a recording medium that stores programs to be executed by the control unit 200 and data (such as various databases) required for executing the programs. Moreover, the memory 290 temporarily stores data for calculation to be executed by the control unit 200. The memory 290 is formed by a magnetic storage device, a semiconductor storage device, an optical storage device, a magneto-optical storage device, etc.

[0136] The learning query image and one or more learning DB images for learning are pre-stored in the memory 290. In the following description, in some cases, the one or more learning DB images will be referred to as "all learning DB images". Further, in some cases, the image group in which the learning query image and all learning DB images are combined will be referred to as "all learning images". Each learning DB image may correspond to an embodiment of the first image. The learning query image may correspond to an embodiment of the second image.

[0137] As will be described later, the overlapping points (overlapping positions) between the learning query image and the learning DB images are extracted, and embodiments of the method for extracting the overlapping points include a first overlapping point extraction method and a second overlapping point extraction method. First, refer to Figure 13 and Figure 14 to describe the first overlapping point extraction method.

[0138] (First overlapping point extraction method)

[0139] Figure 13 is a diagram showing a specific embodiment configuration of the 3D recovery unit 210 according to the first overlapping point extraction method. The 3D recovery unit 210 generates a 3D model based on all learning images. A 3D recovery technique as a known technique can be used to generate the 3D model. Since 3D information related to all learning images is obtained during the process of generating the 3D model, the overlapping points can be extracted based on the 3D information.

[0140] By the first overlapping point extraction method, the 3D information used when extracting the overlapping points may include a 3D feature point group calculated based on sparse corresponding point pairs between every two images among all learning images.

[0141] As Figure 13 shown, the 3D recovery unit 210 includes a position / pose estimation unit 212, a depth estimation unit 214, a point cloud generation unit 216, and a mesh generation unit 217. Here, refer to Figure 14 to describe the respective functions of the position / pose estimation unit 212 and the depth estimation unit 214.

[0142] Figure 14 is a diagram for illustrating the respective functions of the position / pose estimation unit 212 and the depth estimation unit 214. The position / pose estimation unit 212 estimates the device position / pose information corresponding to each learning image based on all learning images, and calculates a 3D feature point group existing in the real space and sparse corresponding point pairs (commonality graph) between every two images among all learning images.

[0143] As a method for estimating device position / pose information corresponding to each learning image based on all learning images and calculating the three-dimensional feature point group and corresponding point pairs as described above, a known method such as Structure from Motion (SfM) can be used. Figure 14 The illustrated embodiment case is based on the assumption that the learning DB image G1 and the learning query image G2 are included in all learning images. Note that the origin C1 is the viewpoint of the imaging device 811, and the origin C2 is the viewpoint of the imaging device 812.

[0144] Here, an exemplary method in which the position / pose estimation unit 212 estimates device position / pose information corresponding to each learning image and calculates the three-dimensional feature point group and corresponding point pairs is briefly described. First, the position / pose estimation unit 212 calculates corresponding point pairs between the learning DB image G1 and the learning query image G2 based on the learning DB image G1 and the learning query image G2. Here, the method for calculating corresponding point pairs is not limited to any specific method.

[0145] As an example, the position / pose estimation unit 212 can extract pixel feature amounts from each of the learning DB image G1 and the learning query image G2. Then, the position / pose estimation unit 212 can match the pixel feature amounts of the corresponding pixels in the learning DB image G1 with the pixel feature amounts of the corresponding pixels in the learning query image G2 to calculate corresponding point pairs of the pixels having the smallest difference in the pixel feature amounts between the learning DB image G1 and the learning query image G2.

[0146] In Figure 14 the illustrated embodiment, the combination of the feature point F11 in the learning DB image G1 and the feature point F21 in the learning query image G2 is a corresponding point pair. Moreover, the combination of the feature point F12 in the learning DB image G1 and the feature point F22 in the learning query image G2 is a corresponding point pair. In addition, the combination of the feature point F13 in the learning DB image G1 and the feature point F23 in the learning query image G2 is a corresponding point pair.

[0147] In addition, based on the corresponding point pairs, the position / pose estimation unit 212 calculates, by triangulation, the position / pose information (first position / pose information) regarding the imaging device 811 (first imaging device) when capturing the learning DB image G1, the position / pose information (second position / pose information) regarding the imaging device 812 (second imaging device) when capturing the learning query image G2, and the three-dimensional feature point groups F1 to F3 as temporary calculation results. Similarly, then, the position / pose estimation unit 212 performs temporary calculations between the other two images among all the learning images, and updates the device position / pose information, the three-dimensional feature point groups, and the corresponding point pairs corresponding to the respective learning images by bundle adjustment so that the temporary calculation results are consistent between every two images. The updated corresponding point pairs correspond to the sparse corresponding point pairs described above.

[0148] The depth estimation unit 214 calculates the depth of each pixel in each learning image and the dense corresponding point pairs (consistency map) between every two images among all the learning images based on all the learning images, the device position / pose information corresponding to each learning image, the three-dimensional feature point groups, and the sparse corresponding point pairs.

[0149] As a method for calculating the depth of each pixel in each learning image based on all the learning images, the device position / pose information corresponding to each learning image, the three-dimensional feature point groups, and the sparse corresponding point pairs as described above, a known method such as multi-view stereo (MVS) can be used.

[0150] Here, an embodiment of the method by which the depth estimation unit 214 calculates the depth of each pixel in each learning image is briefly described. First, the depth estimation unit 214 selects a pair (two) of images that capture the same three-dimensional feature points based on all the learning images, the three-dimensional feature point groups, and the sparse corresponding point pairs.

[0151] At this time, if the angle between the two images is too small, it is also predicted that the depth of each pixel calculated by triangulation cannot be calculated with high accuracy. Therefore, the depth estimation unit 214 can calculate the angle between the two images based on the device position / pose information corresponding to each of the two images and the three-dimensional feature point groups, and limit the angle between the two images to an angle equal to or greater than a predetermined angle. In other words, the depth estimation unit 214 can select two images that do not form an angle smaller than the predetermined angle.

[0152] Then, the depth estimation unit 214 performs pixel matching that is block matching between the two images, and calculates the corresponding point pairs for each pixel between the two images. The corresponding point pairs calculated here correspond to the dense corresponding point pairs described above.

[0153] In Figure 14In the illustrated embodiment, the combination of the point N34 in the learning DB image G1 and the point N44 in the learning query image G2 is a corresponding point pair and corresponds to the three-dimensional point N4. In addition, the combination of the point N35 in the learning DB image G1 and the point N45 in the learning query image G2 is also a corresponding point pair and corresponds to the three-dimensional point N5. In addition, the combination of the point N36 in the learning DB image G1 and the point N46 in the learning query image G2 is also a corresponding point pair and corresponds to the three-dimensional point N6.

[0154] The depth estimation unit 214 calculates the depth of each pixel in each learning image by triangulation based on the device position / pose information corresponding to each learning image and the dense corresponding point pairs. Figure 14 The depth direction T1 from the origin C1 of the imaging device 811 toward the front of the imaging device 811 in the learning DB image G1 is shown. In addition, the depth t1 with respect to the origin C1 (reference position) of the imaging device 811 at the point N35 in the learning DB image G1 is shown.

[0155] The depth estimation unit 214 outputs the depth of each pixel in each learning image to the point cloud generation unit 216. In addition, the depth estimation unit 214 outputs the device position / pose information corresponding to each learning image to the point cloud generation unit 216. At the same time, the depth estimation unit 214 outputs the dense corresponding point pairs to the overlapping point extraction unit 220.

[0156] The point cloud generation unit 216 generates a three-dimensional point cloud by integrating the depth of the corresponding pixel in the corresponding learning image based on the depth of the corresponding pixel in the corresponding learning image and the device position / pose information corresponding to the corresponding learning image. The method of obtaining the three-dimensional point cloud by integrating the depth of the corresponding pixel in the corresponding learning image in this way is also called Fusion.

[0157] The mesh generation unit 217 generates a mesh based on the three-dimensional point cloud generated by the point cloud generation unit 216. The mesh generation unit 217 can generate a mesh using various mesh generation techniques as known methods.

[0158] The overlapping point extraction unit 220 determines the presence / absence of overlapping points between the learning query image and each learning DB image. More specifically, the overlapping point extraction unit 220 determines the presence / absence of overlapping points between the learning query image and each learning DB image based on the three-dimensional information related to all learning images.

[0159] Through the first overlapping point extraction method, the overlapping point extraction unit 220 obtains the dense corresponding point pairs from the depth estimation unit 214. Then, the overlapping point extraction unit 220 extracts overlapping points from the pixels in the learning DB image based on the corresponding points, and the overlapping points are pixels used as corresponding to the points in the pixels of the learning query image.

[0160] When extracting overlapping points from the learned DB images, the overlapping point extraction unit 220 determines that there are overlapping points between the learned DB images and the learned query image. On the other hand, when no overlapping points are extracted from the learned DB images, the overlapping point extraction unit 220 determines that there are no overlapping points between the learned DB images and the learned query image.

[0161] Note that the overlapping point extraction unit 220 may remove the erroneously detected corresponding points as noise points from the dense corresponding point pairs obtained by the depth estimation unit 214, and extract overlapping points based on the remaining corresponding points. For example, when the number of corresponding points within a predetermined range (e.g., a range of 3×3 pixels centered on the target pixel) around the target pixel is equal to or less than a threshold, the target pixel may be determined as a noise point.

[0162] (Second overlapping point extraction method)

[0163] Next, refer to Figures 15 to 19 Describe the second overlapping point extraction method.

[0164] Figure 15 FIG. is a diagram showing a specific embodiment configuration of the three-dimensional recovery unit 210 according to the second overlapping point extraction method. By the second overlapping point extraction method, the three-dimensional recovery unit 210 also generates a three-dimensional model based on all the learned images as in the operation of the first overlapping point extraction method. By the second overlapping point extraction method, three-dimensional information related to all the learned images is also obtained during the generation of the three-dimensional model, and thus, overlapping points can be extracted based on the three-dimensional information.

[0165] As Figure 15 shown, by the second overlapping point extraction method, the position / pose estimation unit 212 outputs the device position / pose information corresponding to each learned image to the overlapping point extraction unit 220. In addition, the mesh generation unit 217 outputs the generated mesh to the overlapping point extraction unit 220.

[0166] Here, by the second overlapping point extraction method, the three-dimensional information used when extracting overlapping points may include information based on the device position / pose information corresponding to each learned DB image and the device position / pose information corresponding to the learned query image.

[0167] More specifically, the information based on the device position / pose information corresponding to each learned DB image and the device position / pose information corresponding to the learned query image may include the coordinates of three-dimensional points (first three-dimensional points) in the real space and the coordinates of three-dimensional points (second three-dimensional points) in the real space. Moreover, the three-dimensional information used when extracting overlapping points may include the normal direction with respect to the object surface at these three-dimensional points.

[0168] In the grid generated by the grid generation unit 217, it includes a three-dimensional point group in the real space and the normal direction with respect to the object surface at the three-dimensional points. The overlapping point extraction unit 220 extracts overlapping points based on the grid generated by the grid generation unit 217 and the device position / pose information corresponding to each learning image.

[0169] Figure 16 is a view of the three-dimensional point group observed obliquely from the side. Refer to Figure 16 , there exist a three-dimensional point N51 and a three-dimensional point N56 in the real space. The coordinates of the three-dimensional point N51 and the coordinates of the three-dimensional point N56 are included in the grid output from the grid generation unit 217 to the overlapping point extraction unit 220.

[0170] In addition, refer to Figure 16 , the imaging device 811 that has captured the learning DB image G1 is shown. When capturing the learning DB image G1, the device position / pose information regarding the imaging device 811 is output from the position / pose estimation unit 212 to the overlapping point extraction unit 220. In addition, the imaging device 812 whose front direction is opposite to the front direction of the imaging device 811 is also shown.

[0171] Here, the learning DB image G1 is projected from the real space onto a vertical plane with respect to the front direction, and the vertical plane is located at a position separated by the focal length from the imaging device 811 in the front direction. The overlapping point extraction unit 220 calculates a quadrangular pyramid p1 (in the quadrangular pyramid, the origin C1 is the vertex, the centroid line L1 is the axis, and the surface b1 is the base) surrounded by the straight lines drawn from the origin C1 of the imaging device 811 to the four corners of the pixel g3 in the learning DB image G1.

[0172] Figure 17 is a view of the three-dimensional point group observed from above. Figure 17 The quadrangular pyramid P1 surrounded by the straight lines drawn from the origin C1 of the imaging device 811 to the four corners of the learning DB image is shown. Inside the quadrangular pyramid P1, there exist three-dimensional points N51 to N58.

[0173] At this time, it can also be considered that the overlapping point extraction unit 220 calculates all the three-dimensional points N51 to N58 that first appear inside the quadrangular pyramid P1 when observed from the origin C1 as the points visible from the imaging device 811. However, the angle between the direction from the three-dimensional point N52 to the origin C1 and the normal direction with respect to the object surface at the three-dimensional point N52 is 90 degrees or more, and therefore, it can be assumed that the three-dimensional point N52 is not visible from the origin C1. Similarly, it can be assumed that the three-dimensional points N53, N55, N57, and N58 are not visible from the origin C1.

[0174] In view of this, when the angle between the direction from the three-dimensional point toward the origin C1 and the normal direction with respect to the object surface at the three-dimensional point is 90 degrees or greater, it is expected that the overlapping point extraction unit 220 does not set the three-dimensional point as a point visible from the imaging device 811. Therefore, the possibility of erroneously extracting overlapping points based on visible points can be reduced. Note that the normal direction with respect to the object surface at the three-dimensional point may also be included in the mesh output from the mesh generation unit 217.

[0175] Figure 18 is a diagram showing an embodiment of a mesh. Refer to Figure 18 , a mesh including three-dimensional points N61 to N63 is shown. Each of the three-dimensional points N61 to N63 may also be represented as a vertex. Moreover, line segments W12, W23, and W31 connecting every two of the three-dimensional points N61 to N63 are shown. The normal directions V1 to V3 with respect to the object surface at the three-dimensional points N61 to N63 may also be included in the mesh.

[0176] Figure 19 is a diagram for explaining the case of extracting overlapping points based on mesh information. In Figure 19 the illustrated embodiment, three-dimensional points N61 to N72 exist inside the quadrangular pyramid P1. In addition, the normal directions V1 to V12 with respect to the object surface at the three-dimensional points N61 to N72 are shown.

[0177] Here, the angle between the direction from the three-dimensional point N61 toward the origin C1 and the normal direction V1 with respect to the object surface at the three-dimensional point N61 is 90 degrees or greater. Therefore, the overlapping point extraction unit 220 does not need to set the three-dimensional point N61 as a point visible from the imaging device 811. Similarly, the overlapping point extraction unit 220 may not set the three-dimensional points N62, N63, N65, N67, N69, and N70 to N72 as points visible from the imaging device 811. At the same time, the overlapping point extraction unit 220 may set the three-dimensional points N64, N66, and N68 as visible points.

[0178] Similarly, the three-dimensional points N65, N67, and N69 may be set as points visible from the imaging device 812. Then, when the points visible from the imaging device 811 and the points visible from the imaging device 812 include the same three-dimensional points, the overlapping point extraction unit 220 may extract the three-dimensional points as overlapping points between the learning DB image captured by the imaging device 811 and the learning query image captured by the imaging device 812.

[0179] In this way, the overlapping point extraction unit 220 extracts overlapping points between each learned DB image and the learned query image. The overlapping point extraction unit 220 stores the extracted overlapping points in the memory 290. The overlapping points stored in the memory 290 are acquired by the region determination unit 250, and the overlapping region corresponding to the overlapping points is determined by the region determination unit 250.

[0180] In an embodiment of the present disclosure, training is performed based on the overlapping region and non-overlapping region determined by the region determination unit 250. Using the model obtained through such learning, an image feature amount is extracted from the inference query image with higher accuracy. To facilitate understanding of the advantages of the learning method according to the embodiment of the present disclosure, first refer to Figure 20 Describe an embodiment of the learning method according to a comparative example.

[0181] (Feature amount extraction unit 530 according to the comparative example)

[0182] Figure 20 is a diagram showing a specific embodiment configuration of the feature amount extraction unit 530 according to the comparative example. The feature amount extraction unit 530 is formed by a DNN. As Figure 20 shown, the feature amount extraction unit 530 according to the comparative example includes a learned query image feature amount extraction unit 231 and a learned DB image feature amount extraction unit 534. The learned query image feature amount extraction unit 231 includes a pixel feature amount extraction unit 232 and a summation processing unit 233. At the same time, the learned DB image feature amount extraction unit 534 includes a pixel feature amount extraction unit 235 and a summation processing unit 539.

[0183] The pixel feature amount extraction unit 232 is formed by a convolutional neural network (CNN). For example, the pixel feature amount extraction unit 232 acquires the learned query image from the memory 290 and extracts the pixel feature amount of each pixel (or each pixel after resolution reduction) constituting the learned query image. Such a pixel feature amount can be represented by a vector.

[0184] The summation processing unit 233 is formed by a pooling layer. For example, the summation processing unit 233 generates an image feature amount (second image feature amount) of the learned query image by summing the pixel feature amounts of the corresponding pixels extracted from the learned query image by the pixel feature amount extraction unit 232.

[0185] Here, various methods can be assumed as the method for summing pixel feature amounts. For example, the method for summing pixel feature amounts can be a method for outputting the maximum value among the pixel feature amounts of corresponding pixels as a representative value, a method for clustering the pixel feature amounts of corresponding pixels, summing the pixel feature amounts of each cluster, and combining the sums of the corresponding clusters into a vector, or some other known method.

[0186] Similar to the pixel feature amount extraction unit 232, the pixel feature amount extraction unit 235 is formed by a CNN. For example, the pixel feature amount extraction unit 235 acquires a learned DB image from the memory 290, and extracts the pixel feature amounts of each pixel (or each pixel after resolution reduction) constituting the learned DB image. Such pixel feature amounts can be represented by vectors.

[0187] Similar to the summing processing unit 233, the summing processing unit 539 is formed by a pooling layer. For example, the summing processing unit 539 generates an image feature amount (first image feature amount) of the learned DB image by summing the pixel feature amounts of the corresponding pixels extracted from the learned DB image by the pixel feature amount extraction unit 235.

[0188] The learning loss calculation unit 240 calculates a differential value for updating the DNN forming the feature amount extraction unit 530 based on the image feature amount extracted from the learned query image and the image feature amount extracted from the learned DB image. Such a differential value can also be referred to as a "gradient". Here, the learning loss according to a known method can be used in calculating the learning loss (loss function) for calculating the differential value. As an example, a method called triplet loss can be used to calculate the learning loss.

[0189] By this method, in the case where there is an overlapping region in the learned DB image that overlaps at least a partial region of the learned query image, a differential value for updating the DNN is calculated such that the image feature amount extracted from the overlapping region in the learned DB image is close to the image feature amount extracted from the image query image, and the image feature amount extracted from the non-overlapping region in the learned DB image moves away from the image feature amount extracted from the image query image.

[0190] However, in the comparative example, it is generally assumed that information (which is a ground truth label) indicating which region in the learned DB image is the overlapping region is manually attached rather than automatically. On the other hand, in the embodiment of the present disclosure, information indicating which region is the overlapping region is automatically attached. Next, refer to Figure 21 Describe an embodiment of the learning method according to the embodiment of the present technology.

[0191] (Feature amount extraction unit 230 according to the embodiment of the present disclosure)

[0192] Figure 21 This is a diagram showing a specific embodiment configuration of the feature quantity extraction unit 230 according to an embodiment of the present disclosure. Similar to the feature quantity extraction unit 530 according to the comparative embodiment, the feature quantity extraction unit 230 according to the embodiment of the present disclosure is formed by a DNN. As Figure 21 shown, similar to the feature quantity extraction unit 530, the feature quantity extraction unit 230 includes a learning query image feature quantity extraction unit 231. In addition, instead of the learning DB image feature quantity extraction unit 534, the feature quantity extraction unit 230 includes a learning DB image feature quantity extraction unit 234. Note that the feature quantity extraction unit 230 may correspond to an embodiment of the extraction unit.

[0193] Similar to the learning DB image feature quantity extraction unit 534, the learning DB image feature quantity extraction unit 234 includes a pixel feature quantity extraction unit 235. In addition, instead of the summation processing unit 539, the learning DB image feature quantity extraction unit 234 includes a region division unit 236, a summation processing unit 237, and a summation processing unit 238. The region division unit 236 receives a determination result from the region determination unit 250 included in the learning device 20.

[0194] The region determination unit 250 obtains overlapping points from the memory 290. Then, the region determination unit 250 determines an overlapping region corresponding to the overlapping points based on the overlapping points obtained from the memory 290. As an embodiment, the region determination unit 250 may determine the pixels in the learning DB image where overlapping points exist as the overlapping region. However, in the case where the overlapping points are sparsely present, the overlapping region may also be sparsely present.

[0195] Therefore, the region determination unit 250 may determine the rectangular region including the overlapping points in the learning DB image as the overlapping region. Alternatively, when the set of pixels where overlapping points are determined to exist in the learning DB image is determined as a temporary overlapping region, the region determination unit 250 may apply a median filter to the temporary overlapping region to exclude pixels in the region where the density of the overlapping points is lower than a predetermined density from the overlapping region.

[0196] The region determination unit 250 outputs a determination result indicating which region in the learning DB image is the overlapping region to the region division unit 236.

[0197] Based on the determination result output from the region determination unit 250, the region division unit 236 divides the pixel feature amounts extracted from the learning DB images into the pixel feature amounts of the corresponding pixels belonging to the overlapping regions and the pixel feature amounts of the corresponding pixels belonging to the non-overlapping regions. Then, the region division unit 236 outputs the pixel feature amounts of the corresponding pixels belonging to the overlapping regions to the summation processing unit 237, and outputs the pixel feature amounts of the corresponding pixels belonging to the non-overlapping regions to the summation processing unit 238.

[0198] Similar to the summation processing unit 539, the summation processing unit 237 is formed by a pooling layer. For example, the summation processing unit 237 generates an image feature amount (overlapping region feature amount) corresponding to the overlapping region by summing up the pixel feature amounts of the corresponding pixels belonging to the overlapping regions output from the region division unit 236. The summation processing unit 237 outputs the image feature amount corresponding to the overlapping region to the learning loss calculation unit 240.

[0199] Similar to the summation processing unit 237, the summation processing unit 238 is formed by a pooling layer. For example, the summation processing unit 238 sums up the pixel feature amounts of the corresponding pixels belonging to the non-overlapping regions output from the region division unit 236 to generate an image feature amount (non-overlapping region feature amount) corresponding to the non-overlapping region. The summation processing unit 238 outputs the image feature amount corresponding to the non-overlapping region to the learning loss calculation unit 240.

[0200] (Learning loss calculation unit 240)

[0201] The learning loss calculation unit 240 calculates a differential value for updating the DNN forming the feature amount extraction unit 230 based on the image feature amount of the learning query image, the image feature amount corresponding to the overlapping region, and the image feature amount corresponding to the non-overlapping region. Here, the learning loss according to a known method can be used in calculating the learning loss for calculating the differential value. As an example, a method called triplet loss can be used to calculate the learning loss.

[0202] More specifically, the learning loss calculation unit 240 calculates a differential value for updating the DNN such that the image feature amount corresponding to the overlapping region and the image feature amount of the learning query image are close to each other, and the image feature amount corresponding to the non-overlapping region and the image feature amount corresponding to the learning query image are moved away from each other. The learning loss calculation unit 240 outputs the differential value to the update unit 260( Figure 12 ).

[0203] (Update unit 260)

[0204] The update unit 260 updates the DNN based on the differential value output from the learning loss calculation unit 240. More specifically, the update unit 260 updates the weight parameters forming the DNN by backpropagation based on the differential value output from the learning loss calculation unit 240. Such an update of the DNN is repeatedly performed for all the learning DB images.

[0205] The feature quantity extraction unit 230 formed by the updated DNN is sent to the inference device 30 via the network 40 and is used as the image feature quantity extraction unit 312 in the inference device 30. The image feature quantity extraction unit 312 can extract the image feature quantity from the inference query image with higher accuracy.

[0206] The above is a description of the exemplary functional configuration of the learning device 20 according to an embodiment of the present disclosure.

[0207] <2. Various modification examples>

[0208] Next, with reference to Figures 22 to 25 various modifications of the information processing system 1 according to an embodiment of the present disclosure will be described.

[0209] (First modification example)

[0210] Figure 22 FIG. is a diagram for explaining the first modification example. In the above description, an embodiment has been described in which the depth of each pixel in each learning image is estimated based on all the learning images. However, it is not necessarily the case that the depth of each pixel in each learning image is estimated only based on the image. For example, as Figure 22 shown, the depth of each pixel in each learning image can be measured by a distance measurement device 610.

[0211] Note that, regarding the type of the distance measurement device 610, any of various sensors can be used. For example, the distance measurement device 610 can be a light detection and ranging (LiDAR) sensor, a stereo depth sensor, or some other distance measurement device.

[0212] In addition, in the above description, an embodiment has been described in which the device position / pose information corresponding to each learning image is estimated based on all the learning images. However, it is not necessarily the case that the device position / pose information corresponding to each learning image is estimated only based on the image. For example, as Figure 22 shown, the device position / pose information corresponding to each learning image can be measured by a SLAM device 620 (self-positioning device).

[0213] Note that, regarding the type of the sensor forming the SLAM device 620, any of various sensors can be used. For example, the SLAM device 620 can include a camera, or can include a combination of a camera and an inertial measurement unit (IMU) sensor.

[0214] (Second modified example)

[0215] Figure 23 It is a diagram for explaining the second modified example. In the above description, an embodiment has been described in which the three-dimensional restoration unit 210 outputs a mesh and device position / pose information corresponding to each learning image to the overlapping point extraction unit 220 based on all the learning images. However, the computer graphics (CG) 710 may output a mesh and device position / pose information corresponding to each learning image to the overlapping point extraction unit 220 based on all the learning images.

[0216] The CG 710 is a program for generating a three-dimensional model. Such a three-dimensional model is arranged in a virtual space. At this time, the learning DB image may be an image based on a predetermined position / pose (first viewpoint) in the virtual space, and the learning query image may be an image based on a predetermined position / pose (second viewpoint) in the virtual space. In addition, the coordinates of the three-dimensional point group in the real space can be obtained from the three-dimensional model generated by the mesh generation unit 217.

[0217] Note that the device position / pose information corresponding to the learning DB image may correspond to the position / pose information about the virtual imaging device when the learning DB image is captured in the virtual space. Similarly, the device position / pose information corresponding to the learning query image may correspond to the position / pose information about the virtual imaging device when the learning query image is captured in the virtual space.

[0218] (Third modified example)

[0219] Figure 24 It is a diagram for explaining the third modified example. In the above description, an embodiment has been described in which learning of the feature amount extraction unit 230 is performed based on the overlapping region and the non-overlapping region. However, in the case where a predetermined object does not appear in the learning query image, for example, if the predetermined object appears in the overlapping region, there is a possibility of confusion during learning, and learning cannot be performed effectively.

[0220] Note that the predetermined object may be a moving object (for example, such as a person or a car). Alternatively, since a mirror image reflected in glass has an adverse effect on learning, the predetermined object may be glass or the like. Alternatively, the predetermined object may be a non-unique object such as the sky.

[0221] Therefore, learning of the feature quantity extraction unit 230 can be performed based on the image feature quantity corresponding to the non-object region and the image feature quantity corresponding to the non-overlapping region, where the non-object region is obtained by excluding the object region including the region in which a predetermined object is detected from the overlapping region. Note that the predetermined object can be detected by a semantic segmentation DNN (or a DNN capable of detecting the predetermined object at pixel pitch).

[0222] As Figure 24 shown, the predetermined object can be detected by the object detection unit 270. Then, the region determination unit 250 can use the result of the object detection performed by the object detection unit 270 to determine the overlapping region.

[0223] Figure 25 is a diagram showing an embodiment of the overlapping region and the non-overlapping region according to the third modification. Referring to Figure 25 , a learning DB image G1 is shown, and an overlapping region G11 and a non-overlapping region G12 included in the learning DB image G1 are shown.

[0224] In the overlapping region G11, an embodiment of a car as a predetermined object is shown. The object detection unit 270 detects the car as an embodiment of the predetermined target from the overlapping region G11. The object detection unit 270 detects a rectangular region including the region where the car is detected as the object region G13. The region determination unit 250 determines a non-object region G14 obtained by excluding the object region G13 from the overlapping region G11.

[0225] The region determination unit 250 outputs the determination result indicating which region in the learning DB image G1 is the non-object region G14 to the region division unit 236. Note that Figure 25 shows a learning DB image G5 combining the non-object region G14 and the non-overlapping region G12.

[0226] Based on the determination result output from the region determination unit 250, the region division unit 236 divides the learning DB image G1 into the object region G13, the non-object region G14, and the non-overlapping region G12. Then, the region division unit 236 outputs the non-object region G14 to the summation processing unit 237 and outputs the non-overlapping region G12 to the summation processing unit 238. Therefore, in the learning of the feature quantity extraction unit 230, the non-object region G14 from which the object region G13 that has an adverse effect on learning is excluded is used. Therefore, the learning is effectively performed.

[0227] The above is a description of various modifications of the information processing system 1 according to the embodiment of the present disclosure.

[0228] <3. Exemplary Hardware Configuration>

[0229] Now refer to Figure 26 to describe an exemplary hardware configuration of an information processing device 900 as an example of an inference device 30 according to an embodiment of the present disclosure. Figure 26 is a block diagram showing an exemplary hardware configuration of the information processing device 900. Note that the inference device 30 does not necessarily have Figure 26 all the hardware configurations shown in Figure 26 , and components of the hardware configurations shown in

[0230] do not need to exist in the inference device 30. In addition, the hardware configuration of the learning device 20 can be formed in a manner similar to the hardware configuration of the inference device 30. Figure 26 As shown in Figure 26 , the information processing device 900 includes a central processing unit (CPU) 901, a read-only memory (ROM) 902, and a random access memory (RAM) 903. The information processing device 900 may further include a host bus 907, a bridge 909, an external bus 911, an interface 913, an input device 915, an output device 917, a storage device 919, a drive 921, a connection port 923, and a communication device 925. The information processing device 900 may have a processing circuit such as a digital signal processor (DSP) or an application specific integrated circuit (ASIC) instead of or in combination with the CPU 901.

[0231] The CPU 901 serves as an arithmetic processing device and a control device, and controls all or some of the operations in the information processing device 900 according to various programs recorded in the ROM 902, the RAM 903, the storage device 919, or the removable recording medium 927. The ROM 902 stores programs, calculation parameters, etc. used by the CPU 901. The RAM 903 temporarily stores programs used by the CPU 901 during execution, parameters appropriately changed during execution, etc. The CPU 901, the ROM 902, and the RAM 903 are interconnected via a host bus 907 formed by an internal bus (such as a CPU bus). In addition, the host bus 907 is connected to an external bus 911, such as a peripheral component interconnect / interface (PCI) bus, via the bridge 909.

[0232] For example, the input device 915 is a device operated by a user, such as a button. The input device 915 may include a mouse, a keyboard, a touch panel, a switch, a joystick, etc. In addition, the input device 915 may also include a microphone that detects the user's voice. For example, the input device 915 may be a remote control device that uses infrared light or some other radio waves, or may be an externally connected device 929, such as a mobile phone compatible with the operation of the information processing device 900. The input device 915 includes an input control circuit that generates an input signal based on the information input by the user and outputs the input signal to the CPU 901. By operating the input device 915, the user inputs various data to the information processing device 900 or gives an instruction to execute a processing operation. In addition, as will be described later, the imaging device 933 can function as an input device by capturing an image of the movement of the user's hand, fingers, etc. At this time, the pointing position can be determined based on the movement of the hand and the orientation of the fingers.

[0233] The output device 917 is formed by a device that can visually or audibly notify the user of the acquired information. For example, the output device 917 may be a display device such as a liquid crystal display (LCD) or an organic electroluminescence (EL) display, a sound output device such as a speaker or headphones, etc. In addition, the output device 917 may include a plasma display panel (PDP), a projector, a hologram, a printer device, etc. The output device 917 outputs the result obtained by the processing executed by the information processing device 900 as a video such as text or an image, or outputs the result as an audio such as voice or sound. In addition, the output device 917 may include a lamp for brightening the environment.

[0234] The storage device 919 is a data storage device designed as an embodiment of the storage unit of the information processing device 900. For example, the storage device 919 is formed by a magnetic storage device such as a hard disk drive (HDD), a semiconductor storage device, an optical storage device, a magneto-optical storage device, etc. The storage device 919 stores programs and various data executed by the CPU 901, as well as various data acquired from the outside.

[0235] The drive 921 is a reader / writer for a removable recording medium 927 (such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory), and is built in or externally connected to the information processing device 900. The drive 921 reads the information recorded in the attached removable recording medium 927 and outputs the information to the RAM 903. In addition, the drive 921 writes a record to the attached removable recording medium 927.

[0236] The connection port 923 is a port for directly connecting the device to the information processing device 900. For example, the connection port 923 can be a Universal Serial Bus (USB) port, an IEEE 1394 port, a Small Computer System Interface (SCSI) port, etc. In addition, the connection port 923 can be an RS-232C port, an optical audio terminal, a High-Definition Multimedia Interface (HDMI (registered trademark)) port, etc. When the external connection device 929 is connected to the connection port 923, various types of data can be exchanged between the information processing device 900 and the external connection device 929.

[0237] For example, the communication device 925 is a communication interface formed by a communication device for connecting to the network 931, etc. For example, the communication device 925 can be a communication card for a wired or wireless Local Area Network (LAN), Bluetooth (registered trademark), or Wireless USB (WUSB), etc. In addition, the communication device 925 can be a router for optical communication, a router for Asymmetric Digital Subscriber Line (ADSL), a modem for various communications, etc. For example, the communication device 925 uses a predetermined protocol such as TCP / IP to send signals, etc. to the Internet and other communication devices, and receives signals, etc. from the Internet and other communication devices. In addition, the network 931 connected to the communication device 925 is a network connected in a wired or wireless manner, and is, for example, the Internet, a home LAN, infrared communication, radio wave communication, satellite communication, etc.

[0238] <4. Conclusion>

[0239] According to an embodiment of the present disclosure, it is feasible to extract image feature amounts from images with higher accuracy using a model obtained through learning. Moreover, it is expected that the image retrieval performance based on the image feature amounts extracted by the model will be improved. In addition, with the improvement of the image retrieval performance, it is expected that the accuracy of the device position / pose information regarding the imaging device will also be improved when capturing images. Furthermore, as the accuracy of the device position / pose information increases, it is expected that the accuracy of the superimposed display of the AR object will also increase, and as the image retrieval performance improves, it is expected that the accuracy of the superimposed display of image retrieval failures will also increase.

[0240] Moreover, according to an embodiment of the present technology, information indicating which region in the learning DB image is the overlapping region (which is the ground truth label) is automatically attached. Therefore, according to an embodiment of the present disclosure, the cost for manually attaching the ground truth label is reduced.

[0241] The preferred embodiments of the present disclosure have been described in detail with reference to the accompanying drawings so far, but the technical scope of the present disclosure is not limited to such embodiments. Obviously, those skilled in the art in the technical field of the present disclosure can design various changes or modifications within the scope of the technical idea disclosed in the claims, and it will be naturally understood that these changes or variations also belong to the technical scope of the present disclosure.

[0242] In addition, the effects disclosed in this specification are merely illustrative or exemplary and not restrictive. That is, in addition to or instead of the above effects, other effects that are obvious to those skilled in the art from the description of this specification can be provided according to the technology of the present disclosure.

[0243] Those skilled in the art should understand that various modifications, combinations, sub - combinations, and changes can occur according to design requirements and other factors, as long as they are within the scope of the appended claims or their equivalents.

[0244] Note that the following configurations also belong to the technical scope of the present disclosure.

[0245] (1) An information processing device, comprising:

[0246] A circuit system configured to:

[0247] Receive a first image,

[0248] Receive a model from a learning device, and

[0249] Output a first position and a first pose based on a first image feature amount extracted from the first image and the model,

[0250] wherein the model is obtained by the following operations:

[0251] Determine at least one overlapping point between a second image and a third image;

[0252] In response to determining at least one overlapping point, divide the second image into an overlapping region and a non - overlapping region, the overlapping region including at least one overlapping point; and

[0253] Perform training based on the second image feature amount corresponding to the overlapping region and the third image feature amount corresponding to the non - overlapping region.

[0254] (2) The information processing device according to (1), wherein the received first image is captured by a terminal device.

[0255] (3) The information processing device according to (1) or (2), wherein the model is a three - dimensional model.

[0256] (4) The information processing apparatus according to any one of (1) to (3), wherein the first position and the first pose are determined based on the first image feature amount and at least one fourth image feature amount extracted from at least one high-order inference database image.

[0257] (5) The information processing apparatus according to any one of (1) to (4), wherein the first position and the first pose are further determined based on the vector indicated by the first image feature amount with respect to a part of the model trained based on at least one high-order inference database image.

[0258] (6) The information processing apparatus according to any one of (1) to (5), wherein the circuit system receives the model from the learning device based on the difference between the vector indicated by the first image feature amount and the vector indicated by the second image feature amount.

[0259] (7) The information processing apparatus according to any one of (1) to (6), wherein the circuit system receives the model from the learning device based on the ranking of a plurality of database images and the difference between the first image feature amount and the corresponding database image feature amount of each corresponding database image in the plurality of database images.

[0260] (8) The information processing apparatus according to any one of (1) to (7), wherein the circuit system outputs the first position and the first pose based on the difference between the vector indicated by the first image feature amount and the vector indicated by at least one fourth image feature amount extracted from at least one high-order inference database image.

[0261] (9) The information processing apparatus according to any one of (1) to (8), wherein the circuit system outputs the first position and the first pose based on the ranking of a plurality of database images and the difference between the first image feature amount and the corresponding database image feature amount of each corresponding database image in the plurality of database images.

[0262] (10) The information processing apparatus according to any one of (1) to (9), wherein the difference between the first image feature amount and the corresponding image feature amount of each corresponding database image indicates whether the pixels of the first image correspond to the pixels of each corresponding database image.

[0263] (11) The information processing apparatus according to any one of (1) to (10), wherein the circuit system is further configured to determine the correspondence between the pixels of the first image and the pixels of each corresponding database image based on the depth of the pixels of the first image and the depth of the pixels of each corresponding database image.

[0264] (12) The information processing apparatus according to any one of (1) to (11), wherein the circuit system is further configured to estimate the depth of the pixels of the first image.

[0265] (13) The information processing apparatus according to any one of (1) to (12), wherein the training based on the second image feature amount corresponding to the overlapping region and the third image feature amount corresponding to the non-overlapping region includes training using a convolutional neural network.

[0266] (14) The information processing apparatus according to any one of (1) to (13), wherein at least one overlapping point between the second image and the third image is determined based on the overlapping density between the pixel set of the second image and the pixel set of the third image.

[0267] (15) The information processing apparatus according to any one of (1) to (14), wherein the first image feature amount is extracted from the first image based on the sum of the feature amounts of the pixels of the first image.

[0268] (16) The information processing apparatus according to any one of (1) to (15), wherein the circuit system outputs the first position and the first pose based on the relationship between the sum of the feature amounts of the pixels of the first image and the sum of the feature amounts of the pixels of the model.

[0269] (17) An information processing method, comprising:

[0270] Receiving a first image;

[0271] Receiving a model from a learning device; and

[0272] Outputting a first position and a first pose based on the first image feature amount extracted from the first image and the model,

[0273] wherein the model is obtained by:

[0274] Determining at least one overlapping point between a second image and a third image;

[0275] Dividing the second image into an overlapping region and a non-overlapping region in response to determining the at least one overlapping point, the overlapping region including the at least one overlapping point; and

[0276] Performing training based on a second image feature amount corresponding to the overlapping region and a third image feature amount corresponding to the non-overlapping region.

[0277] (18) A non-transitory computer-readable medium having a program recorded thereon, which when executed by a computer causes the computer to execute an information processing method, the method including:

[0278] Receiving a first image;

[0279] Receiving a model from a learning device; and

[0280] Output a first position and a first pose based on a first image feature quantity extracted from a first image and a model.

[0281] Wherein, the model is obtained through the following operations:

[0282] Determine at least one overlapping point between a second image and a third image;

[0283] In response to determining at least one overlapping point, divide the second image into an overlapping region and a non-overlapping region, where the overlapping region includes at least one overlapping point; and

[0284] Perform training based on the second image feature quantity and the third image feature quantity corresponding to the overlapping region.

[0285] (B1)

[0286] An information processing method implemented by a processor,

[0287] The method includes:

[0288] Determine whether there is an overlapping position between a first image and a second image;

[0289] When it is determined that there is an overlapping position, perform learning, and the learning is based on

[0290] The overlapping region feature quantity corresponding to the overlapping region corresponding to the overlapping position in the first image feature quantity extracted from the first image by the extraction unit, and

[0291] The non-overlapping region feature quantity corresponding to the non-overlapping region which is the region in the first image other than the overlapping region in the first image feature quantity; and

[0292] Cause the model to extract a third image feature quantity from a third image, where the model is obtained by learning to update the extraction unit.

[0293] (B2)

[0294] According to the information processing method in (1), wherein,

[0295] Perform learning based on the overlapping region feature quantity, the non-overlapping region feature quantity, and the second image feature quantity extracted from the second image by the extraction unit.

[0296] (B3)

[0297] According to the information processing method in (2), wherein,

[0298] The learning includes updating the extraction unit such that the overlapping region feature quantity and the second image feature quantity are close to each other, and the non-overlapping region feature quantity and the second image feature quantity move away from each other.

[0299] (B4)

[0300] The information processing method according to (1), wherein,

[0301] Based on the three-dimensional information related to the first image and the second image, the existence of an overlapping position is determined.

[0302] (B5)

[0303] The information processing method according to (4), wherein,

[0304] The three-dimensional information includes a group of three-dimensional feature points calculated based on corresponding point pairs between the first image and the second image.

[0305] (B6)

[0306] The information processing method according to (4), wherein,

[0307] The three-dimensional information includes information based on the first position / pose information of the first imaging device when capturing the first image and the second position / pose information of the second imaging device when capturing the second image.

[0308] (B7)

[0309] The information processing method according to (6), wherein,

[0310] Based on the first image and the second image, the first position / pose information and the second position / pose information are estimated.

[0311] (B8)

[0312] The information processing method according to (6), wherein,

[0313] The first position / pose information and the second position / pose information are estimated by a self-positioning device.

[0314] (B9)

[0315] The information processing method according to (6), wherein,

[0316] The first position / pose information is the position / pose information of the first imaging device when capturing the first image in a virtual space, in which a three-dimensional model generated by computer graphics is arranged, and

[0317] The second position / pose information is the position / pose information of the second imaging device when capturing the second image in the virtual space.

[0318] (B10)

[0319] The information processing method according to (6), wherein,

[0320] The information based on the first position / pose information and the second position / pose information includes the coordinates of the first three-dimensional point in the real space that appears in the first image and the coordinates of the second three-dimensional point in the real space that appears in the second image.

[0321] (B11)

[0322] The information processing method according to (10), wherein,

[0323] The three-dimensional information includes the normal direction with respect to the object surface at the first three-dimensional point and the normal direction with respect to the object surface at the second three-dimensional point.

[0324] (B12)

[0325] The information processing method according to (10), wherein,

[0326] The coordinates of the first three-dimensional point and the coordinates of the second three-dimensional point are calculated based on the depth from a predetermined origin.

[0327] (B13)

[0328] The information processing method according to (12), wherein,

[0329] The depth is calculated based on the first image, the first position / pose information, the second image, and the second position / pose information.

[0330] (B14)

[0331] The information processing method according to (12), wherein,

[0332] The depth is measured by a distance measuring device.

[0333] (B15)

[0334] The information processing method according to (10), wherein,

[0335] The first image is an image based on the first viewpoint in the virtual space, in which a three-dimensional model generated by computer graphics is arranged,

[0336] The second image is an image based on the second viewpoint in the virtual space, and

[0337] The coordinates of the first three-dimensional point and the coordinates of the second three-dimensional point are obtained from the three-dimensional model.

[0338] (B16)

[0339] The information processing method according to (1), wherein,

[0340] Perform learning based on a feature quantity corresponding to a non-object region and a non-overlapping region feature quantity, where the non-object region is obtained by excluding an object region including a region in which a predetermined object is detected from an overlapping region.

[0341] (B17)

[0342] The information processing method according to any one of (1) to (16) further includes:

[0343] Estimate third position / pose information regarding a third imaging device when capturing a third image based on a third image feature quantity, and the processor performs this estimation.

[0344] (B18)

[0345] The information processing method according to (17), wherein

[0346] The processor specifies a predetermined number of image feature quantities from the image feature quantities of the corresponding images in a plurality of images in ascending order of the difference from the third image feature quantity, and estimates the third position / pose information based on a fourth image and the third image, where the fourth image corresponds to each of the predetermined number of image feature quantities.

[0347] (B19)

[0348] An information processing device includes:

[0349] A model obtained by updating an extraction unit through learning,

[0350] Wherein

[0351] Check to determine whether there is an overlapping position between a first image and a second image. When it is determined that there is an overlapping position,

[0352] Perform learning based on the following:

[0353] An overlapping region feature quantity corresponding to an overlapping region corresponding to the overlapping position in the first image features extracted from the first image by the extraction unit, and

[0354] A non-overlapping region feature quantity corresponding to a non-overlapping region that is a region other than the overlapping region in the first image among the first image features, and

[0355] The model extracts a third image feature quantity from a third image.

[0356] (B20)

[0357] A program for causing a computer to:

[0358] Determine whether there is an overlapping position between a first image and a second image;

[0359] When it is determined that there is an overlapping position, learning is performed. The learning is based on

[0360] the overlapping region feature amount, which corresponds to the overlapping region corresponding to the overlapping position in the first image feature amount extracted from the first image by the extraction unit, and

[0361] the non-overlapping region feature amount, which corresponds to the non-overlapping region that is the region other than the overlapping region in the first image among the first image feature amounts; and

[0362] the model is made to extract a third image feature amount from a third image. The model is obtained by updating the extraction unit through learning.

[0363] [List of reference numerals]

[0364] 1 Information processing system

[0365] 10 Terminal device

[0366] 110 Imaging device

[0367] 120 Operation unit

[0368] 150 Storage unit

[0369] 160 Presentation unit

[0370] 20 Learning device

[0371] 200 Control unit

[0372] 210 3D restoration unit

[0373] 212 Position / orientation estimation unit

[0374] 214 Depth estimation unit

[0375] 216 Point cloud generation unit

[0376] 217 Mesh generation unit

[0377] 220 Overlapping point extraction unit

[0378] 230 Feature amount extraction unit

[0379] 231 Learning query image feature amount extraction unit

[0380] 232 Pixel feature amount extraction unit

[0381] 233 Summation processing unit

[0382] 234 Feature amount extraction unit

[0383] 235 Pixel feature amount extraction unit

[0384] 236 Region division unit

[0385] 237 Summation processing unit

[0386] 238 Summation processing unit

[0387] 240 Learning loss calculation unit

[0388] 250 Region determination unit

[0389] 260 Update unit

[0390] 270 Object detection unit

[0391] 290 Memory

[0392] 30 Inference device

[0393] 300 Control unit

[0394] 310 Image retrieval unit

[0395] 312 Image feature quantity extraction unit

[0396] 314 Image feature quantity matching unit

[0397] 320 Feature point matching unit

[0398] 322 Pixel feature quantity extraction unit

[0399] 324 Pixel feature quantity matching unit

[0400] 330 Relative position / pose estimation unit

[0401] 340 Device position / pose estimation unit

[0402] 390 Memory

[0403] 40 Network

[0404] 610 Distance measurement device

[0405] 620 SLAM device

[0406] 710 CG

[0407] 811 Imaging device

[0408] 812 Imaging device

[0409] 814 Imaging device.

Claims

1. An information processing device, comprising: A circuit system configured to: Receive a first image, Receive a model from a learning device, and Output a first position and a first pose based on a first image feature amount extracted from the first image and the model, Wherein the model is obtained by the following operations: Determine at least one overlapping point between a second image and a third image; In response to determining the at least one overlapping point, divide the second image into an overlapping region and a non-overlapping region, the overlapping region including the at least one overlapping point; and Perform training based on a second image feature amount corresponding to the overlapping region and a third image feature amount corresponding to the non-overlapping region.

2. The information processing device according to claim 1, Among them, The received first image is captured by a terminal device.

3. The information processing device according to claim 1, Among them, The model is a three-dimensional model.

4. The information processing device according to claim 1, Among them, The first position and the first pose are determined based on the first image feature amount and at least one fourth image feature amount extracted from at least one high-order inference database image.

5. The information processing device according to claim 4, Among them, The first position and the first pose are further determined based on a vector indicated by the first image feature amount relative to a part of a model trained based on the at least one high-order inference database image.

6. The information processing device according to claim 1, Among them, The circuit system receives the model from the learning device based on a difference between a vector indicated by the first image feature amount and a vector indicated by the second image feature amount.

7. The information processing device according to claim 1, Among them, The circuit system receives the model from the learning device based on a ranking of a plurality of database images and a difference between the first image feature amount and a corresponding database image feature amount of each corresponding database image in the plurality of database images.

8. The information processing device according to claim 1, Among them, The circuit system outputs the first position and the first pose based on a difference between a vector indicated by the first image feature amount and a vector indicated by at least one fourth image feature amount extracted from at least one high-order inference database image.

9. The information processing device according to claim 1, Among them, The circuit system outputs the first position and the first pose based on a ranking of a plurality of database images and a difference between the first image feature amount and a corresponding database image feature amount of each corresponding database image in the plurality of database images.

10. The information processing device according to claim 9, Among them, The difference between the first image feature amount and the corresponding image feature amount of each corresponding database image indicates whether the pixels of the first image correspond to the pixels of each corresponding database image.

11. The information processing device according to claim 10, Among them, The circuit system is further configured to determine a correspondence between pixels of the first image and pixels of each respective database image based on the depth of the pixels of the first image and the depth of the pixels of each respective database image.

12. The information processing device according to claim 11, Among them, The circuit system is further configured to estimate the depth of the pixels of the first image.

13. The information processing device according to claim 1, Among them, Training based on the second image feature amount corresponding to the overlapping region and the third image feature amount corresponding to the non-overlapping region includes training using a convolutional neural network.

14. The information processing device according to claim 1, Among them, The at least one overlapping point between the second image and the third image is determined based on an overlapping density between a pixel set of the second image and a pixel set of the third image.

15. The information processing device according to claim 1, Among them, The first image feature amount is extracted from the first image based on a sum of feature amounts of the pixels of the first image.

16. The information processing device according to claim 15, Among them, The circuit system outputs the first position and the first pose based on a relationship between a sum of feature amounts of the pixels of the first image and a sum of feature amounts of the pixels of the model.

17. An information processing method, comprising: Receiving a first image; Receiving a model from a learning device; and Outputting a first position and a first pose based on a first image feature amount extracted from the first image and the model, wherein the model is obtained by: Determining at least one overlapping point between a second image and a third image; Dividing the second image into an overlapping region and a non-overlapping region in response to determining the at least one overlapping point, the overlapping region including the at least one overlapping point; and Performing training based on a second image feature amount corresponding to the overlapping region and a third image feature amount corresponding to the non-overlapping region.

18. A non-transitory computer-readable medium having a program stored thereon, which when executed by a computer causes the computer to execute an information processing method, the method comprising: Receiving a first image; Receiving a model from a learning device; and Outputting a first position and a first pose based on a first image feature amount extracted from the first image and the model, wherein the model is obtained by: Determining at least one overlapping point between a second image and a third image; Dividing the second image into an overlapping region and a non-overlapping region in response to determining the at least one overlapping point, the overlapping region including the at least one overlapping point; and Performing training based on a second image feature amount corresponding to the overlapping region and a third image feature amount corresponding to the non-overlapping region.

Citation Information

Patent Citations

  • gaming machines

    JP2022189993A