Visual Positioning System
The visual positioning system addresses the challenge of integrating real and virtual environments by generating scene graphs to extract features, ensuring accurate location mapping through consistent feature extraction.
Patent Information
- Application Number
- JP2023083329
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-05-19
- Publication Date
- 2026-03-04
- Estimated Expiration
- 2043-05-19
AI Technical Summary
Conventional visual positioning technologies are limited to associating locations in the real world and have not considered integrating with virtual worlds, leading to potential inaccuracies due to structural differences between real and virtual environments.
A visual positioning system that extracts features from images by generating a scene graph and then extracting features from that graph, ensuring consistency between real and virtual worlds by abstracting away subtle structural differences.
Ensures high accuracy in associating camera viewpoints between real and virtual worlds by maintaining feature level consistency, enabling precise location mapping across both domains.
Smart Images

Figure 0007823626000001 
Figure 0007823626000002 
Figure 0007823626000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to visual positioning that identifies a location based on an image captured by a camera. [Background technology]
[0002] Patent Document 1 discloses an image learning device. The image learning device uses three-dimensional computer graphics images acquired in a virtual space to learn a machine learning model for object recognition in a real space. A camera is installed at a fixed position in the real space, and camera viewpoint information is obtained based on the position and orientation of the camera. A virtual camera is placed in the virtual space based on the camera viewpoint information. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-081081 Summary of the Invention [Problem to be solved by the invention]
[0004] Visual positioning is a technology that identifies locations based on images captured by a camera. However, using visual positioning to associate locations in the real world with locations in the virtual world has not been considered until now.
[0005] One object of the present disclosure is to provide a technology that can associate a position in the real world with a position in a virtual world using visual positioning. [Means for solving the problem]
[0006] One aspect of the present disclosure relates to a visual positioning system that determines a location based on an image captured by a camera. World 1 is either the real world or a virtual world that simulates the real world. The second world is the other half of the real world and the virtual world. The first image is an image taken by a first camera in the first world. The second image is an image taken by a second camera in the second world. The visual positioning system includes one or more processors. One or more processors execute common processing to generate a scene graph representing the positional relationships between objects included in the image and extract feature quantities from the scene graph. The one or more processors perform matching between a first feature extracted as a result of the common processing on the first image and a second feature extracted as a result of the common processing on the second image. The one or more processors correlate the position of the first camera in the first world and the position of the second camera in the second world based on the results of the matching. [Effects of the Invention]
[0007] According to the present disclosure, features are extracted from an image by common processing. The common processing does not extract features from the image itself, but first abstracts the image to generate a scene graph and then extracts features from that scene graph. Minor structural differences between the real world and the virtual world are absorbed by the image abstraction. Therefore, consistency between the real world and the virtual world is ensured at the feature level. Therefore, features extracted from real images and virtual images captured from the same camera viewpoint are sufficiently consistent. Because features for the same camera viewpoint are sufficiently consistent, visual positioning can be performed with high accuracy. As a result, it is possible to accurately associate the camera viewpoint in the real world with the camera viewpoint in the virtual world. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a conceptual diagram for explaining an overview of a visual positioning system according to an embodiment. [Figure 2] FIG. 10 is a conceptual diagram for explaining common processing according to the embodiment. [Figure 3] FIG. 10 is a conceptual diagram for explaining a modified example of the common processing according to the embodiment. [Figure 4] 1 is a block diagram showing an example of the configuration of a visual positioning system according to an embodiment; [Figure 5] FIG. 2 is a block diagram for explaining a first example of visual positioning according to an embodiment. [Figure 6] FIG. 10 is a block diagram for explaining a second example of visual positioning according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0009] 1. Overview Visual positioning is a technology for identifying a position based on an image captured by a camera. First, a database (gallery) is generated in advance based on a first image captured by a first camera. More specifically, a first feature, which is a feature of the first image, is extracted by a predetermined feature extraction process. Then, a correspondence between the first feature and the camera viewpoint of the first camera when the first image was captured is stored in the database. The camera viewpoint is defined by a combination of the location and orientation of the camera in an absolute coordinate system.
[0010] The process for identifying the position of the second camera is as follows. First, a second image captured by the second camera is acquired as a query. Then, a second feature, which is a feature of the second image, is extracted by the same feature extraction process as for the first feature of the first image. Then, a matching process is performed to search the database for a matching entry having a first feature that matches the second feature. The camera viewpoint included in the obtained matching entry is determined to be the camera viewpoint (position and orientation) of the second camera.
[0011] This type of visual positioning makes it possible to determine a location without using GNSS (Global Navigation Satellite System) or the like.
[0012] However, conventional visual positioning has only been applied to cameras in the real world, and has not taken the virtual world into consideration. Therefore, this embodiment proposes a technology that can extend visual positioning to the virtual world. In particular, this embodiment proposes a technology that can use visual positioning to associate a real position in the real world with a virtual position in the virtual world.
[0013] 1 is a conceptual diagram for explaining an overview of a visual positioning system 100 according to this embodiment. In the following description, the visual positioning system 100 will be referred to as a "VPS 100."
[0014] First, consider a predetermined area in the real world WO-R. For example, the predetermined area may be a city, a building, etc. One or more real cameras 10-R exist in the predetermined area of the real world WO-R. The real camera 10-R may be a stationary camera (fixed camera), or a mobile camera mounted on a mobile object. Examples of mobile objects include a vehicle, a robot, etc. A real image IMG-R of the predetermined area in the real world WO-R is captured by the real camera 10-R.
[0015] The virtual world WO-V is a world that simulates the real world WO-R. The virtual world WO-V is reproduced so as to be as similar to the real world WO-R as possible. For example, the virtual world WO-V is reproduced on a computer using DigitalTwin technology. One or more virtual cameras 10-V exist in a predetermined area of the virtual world WO-V. The virtual camera 10-V may be a stationary camera (fixed camera) or a mobile camera mounted on a mobile object. Examples of mobile objects include vehicles and robots. The virtual camera 10-V captures a virtual image IMG-V of a predetermined area of the virtual world WO-V.
[0016] The VPS 100 acquires a real image IMG-R captured by a real camera 10-R and a virtual image IMG-V captured by a virtual camera 10-V, and then performs visual positioning based on the real image IMG-R and the virtual image IMG-V.
[0017] For example, the above-mentioned database is generated based on a virtual image IMG-V (first image) captured by a virtual camera 10-V (first camera). In this case, the VPS 100 acquires a real image IMG-R (second image) captured by a real camera 10-R (second camera) as a query. The VPS 100 performs visual positioning based on the database and acquires a camera viewpoint in the virtual world WO-V corresponding to the real image IMG-R. The acquired camera viewpoint in the virtual world WO-V can also be said to be the camera viewpoint of the real camera 10-R in the real world WO-R. In other words, the camera viewpoint in the real world WO-R and the camera viewpoint in the virtual world WO-V are associated with each other.
[0018] As an example, the above-mentioned database may be generated based on a real image IMG-R (first image) captured by a real camera 10-R (first camera). In this case, the VPS 100 acquires a virtual image IMG-V (second image) captured by a virtual camera 10-V (second camera) as a query. The VPS 100 performs visual positioning based on the database and acquires a camera viewpoint in the real world WO-R corresponding to the virtual image IMG-V. The acquired camera viewpoint in the real world WO-R can also be said to be the camera viewpoint of the virtual camera 10-V in the virtual world WO-V. In other words, the camera viewpoint in the real world WO-R and the camera viewpoint in the virtual world WO-V are associated with each other.
[0019] The inventors of the present application have recognized the following point: Even for the same object, the granularity (fineness) of information may differ between the real world WO-R and the virtual world WO-V. For example, the structure of a building in the virtual world WO-V is represented by CAD data. This CAD data does not necessarily represent the detailed structure of the actual building in the real world WO-R. In other words, although the virtual world WO-V simulates the real world WO-R as closely as possible, there may be errors in the detailed structure. A real image IMG-R and a virtual image IMG-V captured from the same camera viewpoint are generally the same, but there may be differences in the detailed structure. The feature amounts extracted from such real image IMG-R and virtual image IMG-V do not necessarily match. If the feature amounts do not match despite the same camera viewpoint, the accuracy of localization may decrease, or visual positioning may not function well.
[0020] For accurate visual positioning, it is desirable to match the features for the same camera viewpoint as much as possible. Even if there are differences in the detailed structure, it is desirable to ensure consistency between the real world WO-R and the virtual world WO-V at least at the feature level.
[0021] From the above perspective, the VPS 100 according to this embodiment is configured to ensure consistency between the real world WO-R and the virtual world WO-V at the feature level. To achieve this, the VPS 100 does not extract features directly from the original image IMG, but rather extracts features after applying some additional processing to the original image IMG. This processing will be referred to as "common processing" hereinafter.
[0022] 2 is a conceptual diagram for explaining common processing in the VPS 100. The VPS 100 includes a common processing unit 110 that performs common processing. The common processing unit 110 receives an image IMG, performs common processing on the image IMG, and extracts feature values FE. The common processing unit 110 includes a semantic segmentation processing unit 111, a scene graph generation unit 112, and a feature value extraction unit 113.
[0023] The semantic segmentation processing unit 111 applies well-known semantic segmentation to the image IMG. Semantic segmentation divides the image IMG into multiple regions by classifying each pixel of the image IMG into a category and grouping pixels of the same category together. Regions of the same category are hereinafter referred to as "objects." Objects correspond to various landmarks such as buildings, trees, roads, etc. In the example shown in FIG. 2, the image IMG includes multiple objects S1 to S6. By performing semantic segmentation in this manner, it is possible to detect and identify objects included in the image IMG.
[0024] The scene graph generation unit 112 receives the results of the semantic segmentation. Then, the scene graph generation unit 112 generates a scene graph that represents the positional relationships between multiple objects included in the image IMG. The scene graph has a graph structure, and each node corresponds to each object included in the image IMG. Note that scene graph generation (SGG) is a well-known technique.
[0025] The scene graph obtained in this way accurately represents the characteristics of the scene shown in the original image IMG. On the other hand, the very fine structural information shown in the image IMG is lost in the scene graph. In other words, the scene graph retains the characteristics of the original image IMG but eliminates the fine structural information. In other words, the scene graph is a moderately abstracted version of the original image IMG. Such a scene graph can also be called a mid-level representation of the image IMG.
[0026] The feature extraction unit 113 receives the scene graph from the scene graph generation unit 112. Then, the feature extraction unit 113 extracts a feature FE of the scene graph. For example, the feature FE of the scene graph is extracted by using a graph neural network (GNN).
[0027] The common process described above extracts feature quantities FE from the image IMG. Instead of extracting feature quantities FE from the image IMG itself, the common process first abstracts the image IMG to generate a scene graph and then extracts feature quantities FE from that scene graph. The subtle structural differences between the real world WO-R and the virtual world WO-V are absorbed by the abstraction of the image IMG. Therefore, even if there are subtle structural differences between the real world WO-R and the virtual world WO-V, those differences disappear at the feature quantity FE level. In other words, consistency between the real world WO-R and the virtual world WO-V is ensured at the feature quantity FE level. Therefore, the feature quantities FE extracted from the real image IMG-R and the virtual image IMG-V, which were captured from the same camera viewpoint, are sufficiently consistent. Because the feature quantities FE for the same camera viewpoint are sufficiently consistent, accurate visual positioning is possible. As a result, it is possible to accurately associate the camera viewpoint in the real world WO-R with the camera viewpoint in the virtual world WO-V.
[0028] 2. Variations of common processing 2-1. First modified example Instead of semantic segmentation, a well-known object detection process using an object detection model such as YOLOX may be performed. The object detection process makes it possible to detect objects appearing in the image IMG.
[0029] 2-2. Second modified example An image IMG may contain a moving object. Examples of moving objects include vehicles, pedestrians, and animals. A moving object is not always visible in the image IMG. Therefore, the moving object becomes noise for the feature FE. Therefore, in the second modified example, the objects visible in the image IMG are classified into static objects and dynamic objects. Then, a scene graph is generated based only on the static objects, without using the dynamic objects. This further improves the accuracy of visual positioning.
[0030] 2-3.Third modified example As a result of abstracting the image IMG, there is a possibility that the feature values FE for different camera positions may coincide by chance. In order to suppress such coincidence of the feature values FE for different camera positions, the third modification takes into consideration the depth information of the image IMG.
[0031] 3 is a conceptual diagram for explaining a third modified example. In addition to the above-mentioned functional blocks, the common processing unit 110 further includes a depth estimation unit 114. The depth estimation unit 114 estimates the depth of the image IMG and obtains a depth map. For example, the image IMG is an RGB image. Techniques for estimating depth from an RGB image are well known.
[0032] The scene graph generation unit 112 receives a depth map of the image IMG from the depth estimation unit 114. Based on the depth map, the scene graph generation unit 112 generates a scene graph that represents the three-dimensional positional relationships between multiple objects included in the image IMG. In other words, the scene graph generation unit 112 converts a two-dimensional scene graph into a three-dimensional scene graph by combining the depth map with the two-dimensional scene graph. This reduces the coincidence of feature quantities FE for different camera positions. As a result, the accuracy of visual positioning is further improved.
[0033] 3. Example of a visual positioning system FIG. 4 is a block diagram showing an example configuration of a VPS 100 according to this embodiment. The VPS 100 includes one or more processors 101 (hereinafter simply referred to as "processors 101"), one or more storage devices 102 (hereinafter simply referred to as "storage devices 102"), and an interface 103. The processor 101 executes various processes. For example, the processor 101 includes a CPU (Central Processing Unit). The storage device 102 stores various information required for the processes. Examples of the storage device 102 include a hard disk drive (HDD), a solid state drive (SSD), a volatile memory, and a non-volatile memory. The interface 103 includes a network interface and a user interface.
[0034] The computer program PROG is a computer program for visual positioning. The computer program PROG is stored in the storage device 102. The computer program PROG may be recorded on a computer-readable recording medium. The computer program PROG is executed by the processor 101. The functions of the VPS 100 are realized by cooperation between the processor 101, which executes the computer program PROG, and the storage device 102.
[0035] The virtual world configuration information CONF indicates the configuration of the virtual world WO-V. For example, the virtual world configuration information CONF indicates the three-dimensional arrangement of structures (e.g., roads, road structures, buildings, etc.) within the virtual world WO-V. For example, the three-dimensional arrangement of structures is expressed using CAD data. The virtual world configuration information CONF is stored in the storage device 102.
[0036] The processor 101 recreates the virtual world WO-V on a computer using DigitalTwin technology. At this time, the processor 101 places structures in the virtual world WO-V based on the virtual world configuration information CONF. The processor 101 also places a virtual camera 10-V in the virtual world WO-V. The processor 101 acquires a virtual image IMG-V captured by the virtual camera 10-V based on the virtual world configuration information CONF.
[0037] The processor 101 also communicates with a real camera 10-R present in the real world WO-R via the interface 103. The processor 101 acquires a real image IMG-R captured by the real camera 10-R.
[0038] Furthermore, the processor 101 generates in advance a database 200 (gallery) to be used in visual positioning. The database 200 has a sufficient number of entries. Each entry indicates a correspondence between a camera viewpoint CV and a feature value FE. The camera viewpoint CV is defined by a combination of the position and orientation of the camera 10 in the absolute coordinate system. The feature value FE is extracted by performing the above-mentioned common processing on the image IMG captured by the camera 10. The processor 101 generates the database 200 by performing the common processing on at least one of the real image IMG-R and the virtual image IMG-V. The database 200 is stored in the storage device 102.
[0039] The processor 101 performs visual positioning based on the real image IMG-R, the virtual image IMG-V, and the database 200. An example of visual positioning will be described below.
[0040] 3-1. First example 5 is a block diagram illustrating a first example of visual positioning. In the first example, the VPS 100 generates a database 200-V using a virtual image IMG-V captured by a virtual camera 10-V. Each entry in the database 200-V indicates a correspondence between a camera viewpoint CV-V of the virtual camera 10-V and a feature value FE-V extracted as a result of common processing of the virtual image IMG-V.
[0041] The VPS 100 includes a matching unit 120 in addition to the above-mentioned common processing unit 110. The common processing unit 110 acquires a real image IMG-R captured by a real camera 10-R as a query. The common processing unit 110 extracts a feature value FE-R by performing common processing on the real image IMG-R.
[0042] The matching unit 120 matches the feature FE-R with the feature FE-V. Specifically, the matching unit 120 searches for a matching entry having a feature FE-V that matches the feature FE-R from among multiple entries in the database 200-V. The feature FE-V that matches the feature FE-R is the feature FE-V that is closest to the feature FE-R. The matching unit 120 then determines that the camera viewpoint CV-V included in the matching entry corresponds to the camera viewpoint CV-R of the real camera 10-R in the real world WO-R. In other words, the matching unit 120 determines that the camera viewpoint CV-R of the real camera 10-R that captured the real image IMG-R is the camera viewpoint CV-V included in the matching entry. In this way, the camera viewpoint CV-R in the real world WO-R and the camera viewpoint CV-V in the virtual world WO-V are associated with each other.
[0043] According to the first example, the database 200-V is generated based on the virtual image IMG-V. Therefore, it is possible to easily expand the coverage of the database 200-V. This also contributes to improving the accuracy of visual positioning.
[0044] According to the first example, a camera viewpoint CV-V in the virtual world WO-V corresponding to the camera viewpoint CV-R of the real camera 10-R that captured the real image IMG-R is obtained. This makes it possible to project, for example, an instance (e.g., a person, a vehicle, or an object) captured in the real image IMG-R into the virtual world WO-V.
[0045] 3-2. Second example 6 is a block diagram illustrating a second example of visual positioning. In the second example, the VPS 100 generates a database 200-R using real images IMG-R captured by a real camera 10-R. Each entry in the database 200-R indicates a correspondence between a camera viewpoint CV-R of the real camera 10-R and a feature value FE-R extracted as a result of common processing of the real image IMG-R.
[0046] The common processing unit 110 acquires a virtual image IMG-V captured by a virtual camera 10-V as a query. The common processing unit 110 extracts a feature amount FE-V by performing common processing on the virtual image IMG-V.
[0047] The matching unit 120 matches the feature FE-V with the feature FE-R. Specifically, the matching unit 120 searches for a matching entry having a feature FE-R that matches the feature FE-V from among multiple entries in the database 200-R. The feature FE-R that matches the feature FE-V is the feature FE-R that is closest to the feature FE-V. The matching unit 120 then determines that the camera viewpoint CV-R included in the matching entry corresponds to the camera viewpoint CV-V of the virtual camera 10-V in the virtual world WO-V. In other words, the matching unit 120 determines that the camera viewpoint CV-V of the virtual camera 10-V that captured the virtual image IMG-V is the camera viewpoint CV-R included in the matching entry. In this way, the camera viewpoint CV-V in the virtual world WO-V and the camera viewpoint CV-R in the real world WO-R are associated with each other.
[0048] According to the second example, a camera viewpoint CV-R in the real world WO-R corresponding to the camera viewpoint CV-V of the virtual camera 10-V that captured the virtual image IMG-V is obtained. For example, the future of the virtual world WO-V is predicted by a simulation in DigitalTwin. If a future event is detected based on the virtual image IMG-V, the future event can be overlaid on the real image IMG-R captured by the real camera 10-R.
[0049] 3-3.Generalization The first and second examples described above can be generalized as follows: The first world WO-1 is one of the real world WO-R and the virtual world WO-V. The second world WO-2 is the other of the real world WO-R and the virtual world WO-V. The first camera 10-1 is the camera 10 in the first world WO-1, and is one of the real camera 10-R and the virtual camera 10-V. The second camera 10-2 is the camera 10 in the second world WO-2, and is the other of the real camera 10-R and the virtual camera 10-V. The first image IMG-1 is the image IMG captured by the first camera 10-1. The second image IMG-2 is the image IMG captured by the second camera 10-2.
[0050] The VPS100 executes common processing. The first feature FE-1 is a feature FE extracted as a result of the common processing on the first image IMG-1. The second feature FE-2 is a feature FE extracted as a result of the common processing on the second image IMG-2. The VPS100 matches the first feature FE-1 with the second feature FE-2. Then, based on the matching result, the VPS100 associates the camera viewpoint of the first camera 10-1 in the first world WO-1 with the camera viewpoint of the second camera 10-2 in the second world WO-2.
[0051] The database 200 is generated based on a first image IMG-1 captured by the first camera 10-1. Each entry in the database 200 indicates a correspondence between a camera viewpoint CV-1 of the first camera 10-1 and a first feature FE-1. The VPS 100 extracts the second feature FE-2 by performing common processing on the second image IMG-2 captured by the second camera 10-2. The VPS 100 searches the multiple entries in the database 200 for a matching entry having a first feature FE-1 that matches the second feature FE-2. The first feature FE-1 that matches the second feature FE-2 is the first feature FE-1 that is closest to the second feature FE-2. The VPS 100 then determines that the camera viewpoint CV-2 of the second camera 10-2 in the second world WO-2 is the camera viewpoint CV-1 included in the matching entry. [Explanation of symbols]
[0052] 10 Camera 100 Visual Positioning System (VPS) 110 Common Processing Section 111 Semantic segmentation processing unit 112 Scene Graph Generation Unit 113 Feature Extraction Unit 120 Matching Department CV camera view FE features IMG image
Claims
1. A visual positioning system that identifies a location based on an image captured by a camera, The first world is one of the real world and a virtual world that simulates the real world, a second world being the other of the real world and the virtual world; the first image is an image taken by a first camera in the first world; the second image is an image taken by a second camera in the second world; the visual positioning system comprises one or more processors; the one or more processors: classifying objects included in the image into static objects and dynamic objects; generating a scene graph representing a positional relationship between the static objects without using the dynamic objects, and executing a common process of extracting a feature amount of the scene graph; performing matching between a first feature extracted as a result of the common processing on the first image and a second feature extracted as a result of the common processing on the second image; Based on the result of the matching, the position of the first camera in the first world and the position of the second camera in the second world are associated with each other. It was configured as Visual positioning system.
2. 2. The visual positioning system of claim 1, the visual positioning system further comprising one or more storage devices storing a database having a plurality of entries; each of the plurality of entries indicates a correspondence relationship between a camera viewpoint including a position and an orientation of the first camera and the first feature amount; The one or more processors further extracting the second feature amount by performing the common processing on the second image; searching the plurality of entries in the database for a matching entry having the first feature amount that matches the second feature amount; determining that the camera viewpoint of the second camera in the second world is the camera viewpoint included in the matching entry; It was configured as Visual positioning system.
3. 3. The visual positioning system of claim 2, the first world is the real world, the second world is the virtual world, the first image is an image captured by the first camera in the real world; The second image is an image taken by the second camera in the virtual world. Visual positioning system.
4. 2. The visual positioning system of claim 1, The common process includes detecting the object in the image by performing semantic segmentation. Visual positioning system.
5. 5. A visual positioning system according to any one of claims 1 to 4, comprising: The common processing includes generating a two-dimensional scene graph representing a positional relationship between the static objects in the image, acquiring a depth map of the image, generating a three-dimensional scene graph representing a three-dimensional positional relationship between the static objects by combining the depth map and the two-dimensional scene graph, and extracting a feature amount of the three-dimensional scene graph. Visual positioning system.
Citation Information
Patent Citations
Image processing device, image processing system, and image processing method
JP2013182523A
Image learning device and image learning method
JP2022081081A
Semantic graph embedding lifted for all azimuth direction location recognition
JP2023059794A