Method and system for absolute pose estimation of street view images using automatically generated control features
The method and system automatically generate control features for street view images, addressing the lack of accurate location and orientation by using feature matching, enhancing precision and reducing costs.
Patent Information
- Application Number
- JP2025534951
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-01-30
- Filing Date
- 2023-11-07
- Publication Date
- 2026-01-06
AI Technical Summary
Street View images lack accurate absolute coordinate location and orientation information, requiring costly and error-prone manual placement of markers for control points, leading to mismatches and inaccurate estimation.
A method and system for estimating absolute pose using automatically generated control features, including ground and building control points, through feature matching between street view images and aerial data, enabling precise three-dimensional coordinate and orientation determination.
Enables quick and accurate estimation of three-dimensional absolute coordinate and orientation information for street view images, improving data quality without manual intervention.
Smart Images

Figure 2026500328000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a method and system for estimating absolute pose of street view images, and more particularly to a method and system for estimating absolute pose of street view images using a plurality of automatically generated control features containing three-dimensional absolute coordinate position information. [Background technology]
[0002] With the development of information technology, map information services have been commercialized, and street view images are provided as one type of map information service. For example, a map information service provider may acquire images of an actual space by using a ground mobile object, and then provide the images taken at a specific point on an electronic map as a street view image of the point.
[0003] However, Street View images do not contain accurate absolute coordinate location information or orientation information. To obtain absolute 3D coordinate location information for Street View images taken using a ground-based mobile device, a large number of markers must be placed and surveyed over a wide area to collect control points / control lines, which requires significant cost and effort. Even if control points / control lines are automatically collected using image processing technology, there are frequent mismatches between the collected control points / control lines and feature points in the Street View images, making it difficult to estimate accurate absolute coordinate location information and orientation information. Summary of the Invention [Problem to be solved by the invention]
[0004] The present disclosure provides a method, a non-transitory computer-readable recording medium having instructions recorded thereon, and an apparatus (system) for solving such problems. [Means for solving the problem]
[0005] The present disclosure can be implemented in numerous ways, including as a method, an apparatus (system), or a non-transitory computer-readable storage medium having instructions recorded thereon.
[0006] According to one embodiment of the present disclosure, a method for estimating an absolute pose of street view images, performed by at least one processor, includes: receiving a plurality of control features associated with a specific area; receiving a plurality of matching points obtained through feature matching between a plurality of street view images associated with the specific area; and estimating three-dimensional absolute coordinate position information and orientation information of at least one of the plurality of street view images based on the plurality of control features and the plurality of matching points, wherein each of the plurality of control features includes three-dimensional absolute coordinate position information, and the plurality of control features includes at least one of ground control points, building control points, or ground control lines.
[0007] A non-transitory computer-readable storage medium having instructions recorded thereon for a computer to perform a method according to one embodiment of the present disclosure is provided.
[0008] According to one embodiment of the present disclosure, an information processing system includes a communications module, a memory, and at least one processor connected to the memory and configured to execute at least one computer-readable program contained in the memory, the at least one program including instructions for receiving a plurality of control features associated with a specific area, receiving a plurality of matching points obtained through feature matching between a plurality of street view images associated with the specific area, and estimating three-dimensional absolute coordinate position information and orientation information of at least one of the plurality of street view images based on the plurality of control features and the plurality of matching points, wherein each of the plurality of control features includes three-dimensional absolute coordinate position information, and the plurality of control features includes at least one of ground control points, building control points, or ground control lines. [Effects of the Invention]
[0009] According to one embodiment of the present disclosure, three-dimensional absolute coordinate position information and direction information of a street view image can be quickly and accurately obtained automatically using automatically obtained control features through feature matching between an aerial image containing absolute coordinate information and a street view image.
[0010] According to one embodiment of the present disclosure, it is possible to obtain high-quality three-dimensional absolute coordinate position information and direction information of a street view image by using different weights assigned to different types of control features.
[0011] The effects of the present disclosure are not limited to the effects described above, and other effects not described above will be clearly understood by a person having ordinary knowledge in the technical field to which the present disclosure pertains (referred to as an "ordinary engineer") from the description of the claims. [Brief explanation of the drawings]
[0012] Embodiments of the present disclosure will be described with reference to the accompanying drawings, as described below, in which like reference numerals indicate like elements, but are not limited to the drawings. [Figure 1] FIG. 1 illustrates an example method for aligning a three-dimensional model with street view data according to one embodiment of the present disclosure. [Figure 2] 1 is a schematic diagram illustrating a configuration in which an information processing system according to an embodiment of the present disclosure is communicably connected to a plurality of user terminals. [Figure 3] 1 is a block diagram illustrating an internal configuration of a user terminal and an information processing system according to an embodiment of the present disclosure. [Figure 4] FIG. 1 is a block diagram illustrating a method for estimating absolute three-dimensional coordinate position information and orientation information for street view images associated with a particular area according to one embodiment of the present disclosure. [Figure 5] FIG. 10 is a diagram illustrating a method for determining the weight of a control feature according to an embodiment of the present disclosure. [Figure 6] 10A and 10B are diagrams illustrating a method for estimating three-dimensional absolute coordinate position information and direction information of a specific street view image using ground control points according to an embodiment of the present disclosure. [Figure 7] FIG. 10 is a diagram illustrating a method for estimating three-dimensional absolute coordinate position information and direction information of a street view image using building control points according to an embodiment of the present disclosure. [Figure 8] FIG. 10 is a diagram illustrating a method for estimating three-dimensional absolute coordinate position information and direction information of a street view image using a ground control line according to an embodiment of the present disclosure. [Figure 9] FIG. 1 is a diagram illustrating a method for estimating camera parameters for a street view image according to an embodiment of the present disclosure. [Figure 10] FIG. 10 is a diagram illustrating a specific example of a method for estimating camera parameters for a street view image according to an embodiment of the present disclosure. [Figure 11]FIG. 10 is a diagram illustrating an example of a loss function used when performing filtering according to an embodiment of the present disclosure. [Figure 12] 1 is a flowchart illustrating an example method for absolute pose estimation of street view images according to one embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0013] Hereinafter, specific details for implementing the present disclosure will be described in detail with reference to the accompanying drawings. However, in the following description, detailed descriptions of well-known functions or configurations will be omitted if they may unnecessarily obscure the gist of the present disclosure.
[0014] In the accompanying drawings, identical or corresponding components are denoted by the same reference numerals. In addition, in the following description of the embodiments, duplicated descriptions of identical or corresponding components may be omitted. However, even if technology related to a component is omitted, it is not intended that such a component is not included in a certain embodiment.
[0015] The advantages and features of the disclosed embodiments, as well as methods for achieving them, will become clearer with reference to the following embodiments described in conjunction with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below, and can be realized in various different forms. However, the present embodiments are provided merely to complete the disclosure and to fully convey the scope of the invention to those skilled in the art.
[0016] The terms used in this specification will be briefly explained, and the disclosed embodiments will be specifically described. The terms used in this specification are currently commonly used and general terms that have been selected as much as possible while taking into consideration the functions of the present disclosure. However, these terms may change depending on the intentions or precedents of engineers in the relevant field, the emergence of new technologies, etc. In addition, in certain cases, the applicant may arbitrarily select terms, and in such cases, the meanings thereof will be described in detail in the relevant description of the invention. Therefore, the terms used in this disclosure should be defined based on the meanings of the terms and the overall content of this disclosure, rather than simply by their names.
[0017] In this specification, the singular expression includes the plural expression unless the context clearly dictates otherwise. Furthermore, the plural expression includes the singular expression unless the context clearly dictates otherwise. When a part in the entire specification includes a certain element, this does not mean that other elements are excluded, but that other elements may also be included, unless otherwise specified.
[0018] Additionally, the terms "module" and "unit" as used herein refer to software or hardware components, and a "module" or "unit" performs a certain function. However, the term "module" or "unit" is not limited to software or hardware. A "module" or "unit" may be configured to reside on an addressable storage medium or to execute one or more processors. Thus, by way of example, a "module" or "unit" may include components such as software components, object-oriented software components, class components, and task components, as well as at least one of processes, functions, attributes, procedures, subroutines, program code segments, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The components and "modules" or "units" may be combined into fewer components and "modules" or "units," or the functionality provided therein may be further separated into additional components and "modules" or "units."
[0019] According to one embodiment of the present disclosure, a "module" or "unit" may be implemented with a processor and memory. "Processor" should be broadly interpreted to include a general-purpose processor, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a controller, a microcontroller, a state machine, etc. In some environments, "processor" may refer to an application-specific semiconductor (ASIC), a programmable logic device (PLD), a field-programmable gate array (FPGA), etc. "Processor" may also refer to a combination of processing devices, such as a combination of a DSP and a microprocessor, a combination of multiple microprocessors, a combination of one or more microprocessors coupled with a DSP core, or any other such configuration. "Memory" should also be broadly interpreted to include any electronic component capable of storing electronic information. "Memory" may also refer to various types of processor-readable media, such as random access memory (RAM), read-only memory (ROM), nonvolatile random access memory (NVRAM), programmable ROM (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, magnetic or optical data storage, registers, etc. Memory is said to be in electronic communication with a processor if the processor can read information from and / or write information to the memory. Memory that is integrated into a processor is in electronic communication with the processor.
[0020] In the present disclosure, a "system" may include, but is not limited to, at least one of a server device and a cloud device. For example, a system may be configured with one or more server devices. As another example, a system may be configured with one or more cloud devices. As another example, a system may be configured and operated by a server device and a cloud device together.
[0021] In this disclosure, "display" may refer to any display device associated with a computing device, for example, any display device capable of displaying any information / data controlled by or provided by the computing device.
[0022] In the present disclosure, "each of a plurality of A's" or "each of a plurality of A's" may refer to each of all the components included in the plurality of A's, or may refer to each of some of the components included in the plurality of A's.
[0023] In this disclosure, "street view data" may refer to data including not only road view data including images and location information taken on roadways, but also pedestrian view data including images and location information taken on sidewalks. Furthermore, "street view data" may further include not only roadways and sidewalks, but also images and location information taken at any point outdoors (or indoors with a view of the outdoors).
[0024] 1 illustrates an example method for aligning a 3D model 110 and street view data 120 according to one embodiment of the present disclosure. An information processing system may obtain / receive the 3D model 110 and street view data 120 for a particular area.
[0025] The 3D model 110 may include 3D geometric information expressed in absolute coordinate positions and corresponding texture information. Here, the position information included in the 3D model 110 may be information with higher accuracy than the position information included in the street view data 120. Furthermore, the texture information included in the 3D model 110 may be information with lower quality (e.g., lower resolution) than the texture information included in the street view data 120. According to one embodiment, the 3D geometric information expressed in absolute coordinate positions may be generated based on an aerial photograph taken above the specific area.
[0026] The 3D model 110 for a specific area may include a 3D building model 112, a digital elevation model (DEM) 114, a true ortho image 116 for the specific area, a digital surface model (DSM), a road layout, a road DEM, etc. As a specific example, the 3D model 110 for a specific area may be, but is not limited to, a model generated based on a digital surface model (DSM) including geometric information about the ground of the specific area and the corresponding ortho image 116 for the specific area. In one embodiment, a precise ortho image 116 for a specific area may be generated based on multiple aerial photographs and absolute coordinate position information and direction information for each aerial photograph.
[0027] The street view data 120 may include a plurality of street view images captured at a plurality of nodes within a specific area, and absolute coordinate position information and direction information (i.e., image capture direction information) for each of the plurality of street view images. Here, the position information and direction information included in the street view data 120 may be less accurate than the position information and direction information included in the 3D model 110, and the texture information included in the street view images may be higher quality (e.g., higher resolution) than the texture information included in the 3D model 110. For example, the position information and direction information included in the street view data 120 may be position information and direction information obtained using a GPS device when capturing street view images at the nodes. Position information obtained using a vehicle's GPS device may have an error of approximately 5 to 10 meters.
[0028] The information processing system may perform map matching 130 between the 3D model 110 and the street view data 120. Specifically, the information processing system may perform feature matching between texture information included in the 3D model 110 and a plurality of street view images included in the street view data 120. To perform map matching 130, the information processing system may convert at least some of the plurality of street view images included in the street view data 120 into top view images. As a result of map matching 130, a plurality of map matching points / map matching lines 132 may be automatically extracted without user input. The map matching points / map matching lines 132 are also referred to as control features. Each control feature may include three-dimensional absolute coordinate position information and a visual descriptor.
[0029] A map matching point may represent a correspondence pair between one point in a street view image and one point in the 3D model 110. The types of map matching points may vary depending on the type of 3D model 110 used in map matching 130, the location of the points, and other factors. For example, a map matching point may include at least one of a ground control point (GCP), which is a correspondence pair of points on the ground within a specific area, a building control point (BCP), which is a correspondence pair of points on a building within a specific area, or a structure control point, which is a correspondence pair of points on a structure within a specific area. Map matching points can be extracted from any area of the street view image and the 3D model 110, in addition to the above-mentioned ground, buildings, and structures.
[0030] A map matching line may represent a corresponding pair of a line in a street view image and a line in the 3D model 110. The map matching line may be of various types depending on the type of 3D model 110 used in map matching 130, the position of the line, and the like. For example, the map matching line may include at least one of a ground control line (GCL), which is a corresponding pair of lines on the ground within a specific area; a building control line (BCL), which is a corresponding pair of lines on a building within a specific area; a structure control line, which is a corresponding pair of lines on a structure within a specific area; or a lane control line, which is a corresponding pair of lines on a lane within a specific area. Map matching lines can be extracted from any area of the street view image and the 3D model 110, in addition to the above-mentioned ground, buildings, structures, and lanes.
[0031] The information processing system may also perform feature matching 150 between multiple street view images to extract multiple feature point correspondence sets 152. According to one embodiment, for robust feature matching, feature matching 150 between multiple street view images may be performed using at least a portion of the 3D model 110. For example, feature matching 150 between street view images may be performed using 3D building models 112 included in the 3D model 110.
[0032] The information processing system may then estimate (160) absolute coordinate position information and direction information for the plurality of street view images based on at least one of the plurality of map matching points / map matching lines 132 and at least a portion of the plurality of feature point correspondence sets 152. For example, the processor may estimate (160) the absolute coordinate position information and direction information for the plurality of street view images using a bundle adjustment technique. According to one embodiment, the estimated absolute coordinate position information and direction information 162 is information in an absolute coordinate system representing the three-dimensional model 110 and may be six degrees of freedom (DoF) parameters. The absolute coordinate position information and direction information 162 estimated through this process may be data with higher accuracy than the absolute coordinate position information and direction information included in the street view data 120. A specific method for estimating (160) three-dimensional absolute coordinate position information and direction information for the plurality of street view images will be described in detail below with reference to FIGS. 4 to 12.
[0033] According to one embodiment, an information processing system can estimate 3D absolute coordinate position information and orientation information (i.e., 6-DoF pose) of a Street View image based on matching points obtained by feature matching between a control feature associated with a specific area and multiple Street View images associated with the specific area. Here, the control feature 3D model 110 can be automatically obtained by performing map matching 130 between the control feature 3D model and Street View data 120 without user input, and each control feature can include 3D absolute coordinate position information and a visual descriptor. This configuration allows the 3D absolute coordinate position information and orientation information of the Street View image to be estimated with high accuracy.
[0034] 2 is a schematic diagram illustrating a configuration in which an information processing system 230 according to an embodiment of the present disclosure is communicatively connected to multiple user terminals 210_1, 210_2, and 210_3. As illustrated, the multiple user terminals 210_1, 210_2, and 210_3 can be connected to the information processing system 230, which can provide a map information service, via a network 220. Here, the multiple user terminals 210_1, 210_2, and 210_3 may include terminals of users who receive the map information service. Furthermore, the multiple user terminals 210_1, 210_2, and 210_3 may be automobiles that capture street view images at nodes. In one embodiment, the information processing system 230 may include one or more server devices and / or databases, or one or more cloud computing service-based distributed computing devices and / or distributed databases, that can store, provide, and execute computer-executable programs (e.g., downloadable applications) and data related to the provision of the map information service.
[0035] The map information service provided by the information processing system 230 may be provided to users via an application, a web browser, or the like installed on each of the multiple user terminals 210_1, 210_2, and 210_3. For example, the information processing system 230 may provide information corresponding to a street view image request, an image-based location recognition request, or the like received from the user terminals 210_1, 210_2, and 210_3 via an application, or may perform corresponding processing.
[0036] A plurality of user terminals 210_1, 210_2, and 210_3 can communicate with the information processing system 230 via a network 220. The network 220 can be configured to enable communication between the plurality of user terminals 210_1, 210_2, and 210_3 and the information processing system 230. Depending on the installation environment, the network 220 can be configured by, for example, a wired network such as Ethernet, a wired home network (Power Line Communication), a telephone line communication device, and RS-serial communication, a mobile communication network, a wireless network such as WLAN (Wireless LAN), Wi-Fi, Bluetooth, and ZigBee, or a combination thereof. The communication method is not limited, and can include not only a communication method utilizing a communication network that the network 220 can include (for example, a mobile communication network, a wired Internet, a wireless Internet, a broadcast network, a satellite network, etc.), but also short-range wireless communication between the user terminals 210_1, 210_2, and 210_3.
[0037] 2, a mobile terminal 210_1, a tablet terminal 210_2, and a PC terminal 210_3 are shown as examples of user terminals, but are not limited thereto, and the user terminals 210_1, 210_2, and 210_3 may be any computing devices capable of wired and / or wireless communication and capable of installing and executing an application or a web browser, etc. For example, the user terminals may include AI speakers, smartphones, mobile phones, navigation systems, computers, laptops, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), tablet PCs, game consoles, wearable devices, internet of things (IoT) devices, virtual reality (VR) devices, augmented reality (AR) devices, set-top boxes, etc. Also, while FIG. 2 shows three user terminals 210_1, 210_2, and 210_3 communicating with the information processing system 230 via the network 220, this is not limited thereto, and a different number of user terminals may be configured to communicate with the information processing system 230 via the network 220.
[0038] According to one embodiment, the information processing system 230 may receive, from the user terminals 210_1, 210_2, and 210_3, a plurality of matching points obtained through feature matching between a plurality of control features associated with a specific area and a plurality of street view images associated with the specific area. The information processing system 230 may then estimate three-dimensional absolute coordinate position information and direction information of at least one of the plurality of street view images based on the plurality of control features and the plurality of matching points. The information processing system 230 may then transmit the three-dimensional absolute coordinate position information and direction information of the estimated street image to the user terminals 210_1, 210_2, and 210_3. Additionally, the information processing system 230 may transmit various service-related data based on data generated using the three-dimensional absolute coordinate position information and direction information of the estimated street view image to the user terminals 210_1, 210_2, and 210_3.
[0039] 3 is a block diagram illustrating the internal configuration of a user terminal 210 and an information processing system 230 according to an embodiment of the present disclosure. The user terminal 210 may refer to any computing device capable of executing an application, a web browser, or the like, and capable of wired / wireless communication, and may include, for example, the mobile phone terminal 210_1, the tablet terminal 210_2, and the PC terminal 210_3 of FIG. 2. As illustrated, the user terminal 210 may include a memory 312, a processor 314, a communication module 316, and an input / output interface 318. Similarly, the information processing system 230 may include a memory 332, a processor 334, a communication module 336, and an input / output interface 338. As illustrated in FIG. 3, the user terminal 210 and the information processing system 230 may be configured to communicate information and / or data over the network 220 using their respective communication modules 316 and 336. Additionally, the input / output device 320 may be configured to input information and / or data to the user terminal 210 or output information and / or data generated from the user terminal 210 via the input / output interface 318 .
[0040] The memories 312 and 332 may include any non-transitory computer-readable recording medium. According to one embodiment, the memories 312 and 332 may include a permanent mass storage device such as a read-only memory (ROM), a disk drive, a solid-state drive (SSD), or a flash memory. As another example, a non-transitory mass storage device such as a ROM, an SSD, a flash memory, or a disk drive may be a separate permanent storage device distinct from the memory and may be included in the user terminal 210 or the information processing system 230. The memories 312 and 332 may also store an operating system and at least one program code (e.g., code for an application installed and run on the user terminal 210).
[0041] These software components may be loaded from a computer-readable recording medium separate from the memories 312 and 332. Such separate computer-readable recording media may include recording media directly connectable to the user terminal 210 and the information processing system 230, such as a floppy drive, a disk, a tape, a DVD / CD-ROM drive, or a memory card. As another example, the software components may be loaded into the memories 312 and 332 via a communication module rather than a computer-readable recording medium. For example, at least one program may be loaded into the memories 312 and 332 based on a computer program to be installed by a developer or a file provided via the network 220 by a file distribution system that distributes application installation files.
[0042] The processors 314, 334 may be configured to process computer program instructions by performing basic arithmetic, logic, and input / output operations. The instructions may be provided to the processors 314, 334 by the memories 312, 332 or the communications modules 316, 336. For example, the processors 314, 334 may be configured to execute received instructions according to program code stored in a storage device, such as the memories 312, 332.
[0043] The communication modules 316, 336 may provide a configuration or function for the user terminal 210 and the information processing system 230 to communicate with each other via the network 220, and may provide a configuration or function for the user terminal 210 and / or the information processing system 230 to communicate with other user terminals or other systems (e.g., a separate cloud system). For example, a request or data (e.g., data related to a request for a control feature related to a specific area and a plurality of street view images related to the specific area) generated by the processor 314 of the user terminal 210 in accordance with program code stored in a storage device such as the memory 312 may be transmitted to the information processing system 230 via the network 220 under the control of the communication module 316. Conversely, a control signal or command provided under the control of the processor 334 of the information processing system 230 may be received by the user terminal 210 via the communication module 316 of the user terminal 210 via the communication module 336 and the network 220. For example, the user terminal 210 may receive data related to street view images, including three-dimensional absolute coordinate position information and direction information, from the information processing system 230.
[0044] The input / output interface 318 may be a means for interfacing with the input / output device 320. As an example, the input device may include a camera including an audio sensor and / or an image sensor, a keyboard, a microphone, a mouse, etc., and the output device may include a display, a speaker, a haptic feedback device, etc. As another example, the input / output interface 318 may be a means for interfacing with a device that integrates input and output configurations or functions, such as a touchscreen. For example, when the processor 314 of the user terminal 210 processes instructions of a computer program loaded into the memory 312, a service screen configured using information and / or data provided by the information processing system 230 or another user terminal may be displayed on the display via the input / output interface 318. Although FIG. 3 illustrates the input / output device 320 as not being included in the user terminal 210, this is not limiting and the input / output device 320 may be configured as a single device together with the user terminal 210. Furthermore, the input / output interface 338 of the information processing system 230 may be a means for interfacing with an input or output device (not shown) that is connected to the information processing system 230 or that may be included in the information processing system 230. In Fig. 3, the input / output interfaces 318, 338 are shown as elements configured separately from the processors 314, 334, but are not limited to this, and the input / output interfaces 318, 338 may be configured to be included in the processors 314, 334.
[0045] The user terminal 210 and the information processing system 230 may include more components than those shown in FIG. 3 . However, it is not necessary to explicitly illustrate most of the conventional components. According to one embodiment, the user terminal 210 may be implemented to include at least a portion of the input / output device 320 described above. The user terminal 210 may also include other components such as a transceiver, a Global Positioning System (GPS) module, a camera, various sensors, and a database. For example, if the user terminal 210 is a smartphone, it may include components typically included in a smartphone, such as an acceleration sensor, a gyro sensor, an image sensor, a proximity sensor, a touch sensor, an illuminance sensor, a camera module, various physical buttons, buttons using a touch panel, input / output ports, and a vibrator for vibration. According to one embodiment, the processor 314 of the user terminal 210 may be configured to run an application that provides a map information service. In this case, code related to the application and / or program may be loaded into the memory 312 of the user terminal 210.
[0046] During operation of a program for an application that provides a map information service, the processor 314 may receive text, images, pictures, sounds, and / or actions entered or selected via an input device, such as a touch screen, keyboard, camera including an audio sensor and / or image sensor, or microphone, connected to the input / output interface 318, and may store the received text, images, pictures, sounds, and / or actions in the memory 312 or provide them to the information processing system 230 via the communication module 316 and the network 220. For example, the processor 314 may receive a user input requesting multiple matching points between multiple control features for a specific area and multiple street view images of the specific area, and provide the received input to the information processing system 230 via the communication module 316 and the network 220.
[0047] The processor 314 of the user terminal 210 can be configured to manage, process, and / or store information and / or data received from the input / output device 320, other user terminals, the information processing system 230, and / or multiple external systems. The information and / or data processed by the processor 314 can be provided to the information processing system 230 via the communication module 316 and the network 220. The processor 314 of the user terminal 210 can transmit and output information and / or data to the input / output device 320 via the input / output interface 318. For example, the processor 314 can display the received information and / or data on a screen of the user terminal.
[0048] The processor 334 of the information processing system 230 may be configured to manage, process, and / or store information and / or data received from multiple user terminals 210 and / or multiple external systems. The information and / or data processed by the processor 334 may be provided to the user terminals 210 via the communication module 336 and the network 220.
[0049] 4 is a block diagram illustrating a method for estimating three-dimensional absolute coordinate position information and orientation information 490 of a street view image associated with a specific area according to one embodiment of the present disclosure. According to one embodiment, at least one processor (e.g., at least one processor of an information processing system and / or a user terminal) can receive input data 410 associated with a specific area to estimate three-dimensional absolute coordinate position information and orientation information 490 of a street view image associated with the specific area. Here, the input data 410 can indicate a first set of matching points associated with the specific street view image and a corresponding first set of control features. The control features can include three-dimensional absolute coordinate position information and visual descriptors. For example, the control features can include at least one of a ground control point associated with a specific point on the ground, a building control point associated with a specific point on a building, and a ground control line associated with a specific line on the ground.
[0050] According to one embodiment, the processor may determine (420) weights of a first set of control features to estimate absolute coordinate position information and direction information of a particular street view image. Here, the weights of the first set of control features may be determined differently depending on whether the control feature is a ground control point, a building control point, or a ground control line. A specific method for determining the weights of the control features will be described in detail below with reference to FIG. 5.
[0051] According to one embodiment, the processor may estimate (430) three-dimensional absolute coordinate position information and orientation information to a first precision for the specific street view image based on the input data 410 and the weights of the control features. Here, the three-dimensional absolute coordinate position information and orientation information to a first precision may refer to absolute coordinate position information and orientation information with a precision lower than the precision of the three-dimensional absolute coordinate position information and orientation information to a second precision and the precision of the final three-dimensional absolute coordinate position information and orientation information 490. Specifically, the processor may project a first set of control features onto the specific street view image to generate a first set of projection data, which may be used to estimate the three-dimensional absolute coordinate position information and orientation information to the first precision for the specific street view image. For example, the processor may estimate (430) three-dimensional absolute coordinate position information and orientation information to the first precision for the specific street view image based on the first set of matching points, the first set of projection data, and the respective weights of the first set of control features, such that an error between the first set of matching points associated with the specific street view image and the first set of projection data is minimized. Here, the processor may not estimate camera parameters when estimating the initial position information and orientation information of a particular street view image.
[0052] According to one embodiment, the processor may perform a first filtering (440) based on the estimated three-dimensional absolute coordinate position information and orientation information with a first accuracy. Specifically, the processor may project the first set of control features onto the specific street view image based on the estimated three-dimensional absolute coordinate position information and orientation information with the first accuracy to generate a second set of projection data. The processor may then determine a second set of control features by removing some of the first set of control features based on an error between the first set of matching points associated with the specific street view image and the second set of projection data. The filtering criterion (loss function) may vary depending on the type of control feature (GCP, BCP, GCL). The filtering criterion will be described in detail later with reference to FIG. 11 . In this case, the processor may determine a second set of matching points by removing matching points corresponding to the removed control features from the first set of matching points. That is, the processor extracts data that exceeds a predetermined threshold range from the first set of matching points and the second set of projection data, and based on this, performs first filtering to remove outliers from the first set of control features and the first set of matching points (440), thereby obtaining a second set of matching points and a second set of control features associated with the specific street view image.
[0053] According to one embodiment, the processor may determine (450) individual weights for the input data 410 based on the estimated three-dimensional absolute coordinate position information and orientation information with a first accuracy. Here, the input data 410 to which the individual weights are applied may represent a second set of matching points and a second set of control features from which outliers have been removed by a first filtering. To this end, the processor may calculate the covariance of the estimated three-dimensional absolute coordinate position information and orientation information with a first accuracy. Here, the covariance indicates the confidence level for the initial values of the estimated three-dimensional absolute coordinate position information and orientation information with a first accuracy. Reverse calculation of the confidence level using an error propagation formula reveals the confidence levels for each of the second set of matching points and the second set of control features. The individual weights for each of the second set of matching points and the second set of control features may be determined using the finally calculated confidence levels for each of the second set of matching points and the second set of control features. In other words, the processor can calculate individual confidence levels for each of the second set of matching points and the second set of control features based on the confidence levels of the estimated three-dimensional absolute coordinate position information and orientation information with the first accuracy, and can determine individual weights for each of the second set of matching points and the second set of control features based on the calculated individual confidence levels.
[0054] According to one embodiment, the processor may estimate (460) three-dimensional absolute coordinate position information and orientation information of the specific Street View image to a second precision based on, for example, the second set of matching points, the second set of control features, and the calculated individual weights. To this end, the processor may estimate camera parameters for the specific Street View image, and use these, along with the second set of matching points, the second set of control features, and the calculated individual weights, to estimate (460) three-dimensional absolute coordinate position information and orientation information of the specific Street View image to a second precision. Using the estimated camera parameters in conjunction with the geometric information allows for more accurate estimation of the three-dimensional absolute coordinate position information and orientation information of the Street View image. A specific method for estimating the camera parameters of the specific Street View image will be described in detail below with reference to Figures 9 and 10.
[0055] According to one embodiment, the processor may perform second filtering (470) based on the estimated three-dimensional absolute coordinate position information and orientation information with a second accuracy. The second filtering is performed according to a procedure identical or similar to that of the first filtering. Specifically, the processor may project the second set of control features onto the specific Street View image based on the estimated three-dimensional absolute coordinate position information and orientation information with the second accuracy to generate a third set of projection data. The processor may then determine a third set of control features by removing some of the second set of control features based on an error between the second set of matching points associated with the specific Street View image and the third set of projection data. In this case, the third set of matching points may be determined by removing matching points corresponding to the removed control features from the second set of matching points.
[0056] According to one embodiment, the processor may determine (480) whether the estimated three-dimensional absolute coordinate position information and direction information of the specific Street View image are equal to or less than a predetermined threshold. For example, if the error between the second set of matching points associated with the specific Street View image and the third set of projection data is equal to or less than the predetermined threshold, the processor may determine the three-dimensional absolute coordinate position information and direction information with second precision as final three-dimensional absolute coordinate position information and direction information (490). As a specific example, using the predetermined threshold, the angle between the lines connecting the matching points and the projection data with the image capture location information as the origin may be set to 6 degrees, but is not limited to this. In another example, if the error between the second set of matching points associated with the specific Street View image and the third set of projection data exceeds the predetermined threshold, the processor may repeat steps 450 to 480 until the error is equal to or less than the predetermined threshold.
[0057] FIG. 5 is a diagram illustrating a method for determining the weight of a control feature according to an embodiment of the present disclosure. According to an embodiment, the weight of a control feature may be determined differently depending on the distance between the reference point of the control feature and a specific Street View image. For example, the weight may be determined as a value between 0 and 1 depending on the distance between the reference point of the control feature and a specific Street View image. Here, the control feature may include three-dimensional absolute coordinate position information. Furthermore, the control feature may include at least one of a ground control point, a building control point, or a ground control line.
[0058] In one embodiment, the ground control points 520, 530 may be assigned a lower weight the greater the distance between the street view image 510 and the ground control points 520, 530. This is because, in the case of a region of the street view image corresponding to an object that is far away from the shooting position, the image quality of the region is low, increasing the probability of mismatching between the matching point (feature point) and the ground control point. For example, with reference to the planar image 512 of the street view image shown in FIG. 5, it can be seen that the image quality of a first region 522 that is far away from the shooting position (origin) is lower than the image quality of a second region 532 that corresponds to a relatively closer object. The first region 522 may correspond to the first ground control point 520, and the second region 532 may correspond to the second ground control point 530. In this case, the first ground control point 520 that is far away from the second ground control point 530 from the shooting position (origin) of the street view image 510 may have a lower weight than the second ground control point 530.
[0059] In one embodiment, the relative positions between the shooting position (origin) and the ground control points 520, 530 may be determined as y-axis values within the planar image 512 of the street view image. For example, referring to FIG. 5, the y-axis value may increase toward the bottom of the image, which may indicate closer to the shooting position (origin). Therefore, the larger the y-axis value, the higher the weighting of the ground control point.
[0060] In one embodiment, a building control point may be weighted higher the greater the distance between the Street View image and the building control point. This is because if a building is located close to the shooting location in the Street View image, there is a high probability that the shooting location is a narrow road such as an alley, and in an alley, a building control point itself is generated, which increases the probability that a matching point (feature point) and the building control point are erroneously mismatched.
[0061] In one embodiment, the weight of a building control point may be calculated by dividing the distance value of the target building relative to the shooting position by the maximum distance value that can be set relative to the shooting position. Here, the maximum distance that can be set relative to the shooting position may be calculated to be 100 m to 200 m, but is not limited thereto. If the maximum distance value that can be set is 100 m and the distance between the shooting position and the target building is 50 m, the weight of the building control point may be determined to be 0.5.
[0062] In one embodiment, the ground control lines may not be assigned a separate weight, for example, the ground control lines may all be assigned a weight of 1.
[0063] Through this configuration, mismatches that occur between the automatically collected control features and the matching points included in the Street View image can be effectively eliminated, thereby enabling high-quality 3D absolute coordinate position information and direction information for a specific Street View image to be quickly and accurately obtained.
[0064] 6 is a diagram illustrating a method for estimating three-dimensional absolute coordinate position information and orientation information of a specific street-view image 630 using ground control points according to one embodiment of the present disclosure. According to one embodiment, the control features may include ground control points 622. The ground control points may be obtained based on feature matching between the street-view image and an aerial image 620 including three-dimensional absolute coordinate position information, and the ground control points may include three-dimensional absolute coordinate position information and visual descriptors associated with specific points on the ground of the street-view image. Here, the specific points on the ground of the street-view image may be referred to as matching points (feature points) obtained through feature matching between multiple street-view images.
[0065] According to one embodiment, the information processing system can estimate three-dimensional absolute coordinate position information and orientation information of the specific street view image 630 based on the ground control points 622 and matching points 612 in the specific street view image 630. For example, the information processing system can project the ground control points 622 onto the specific street view image 630 to generate projection data 624. The information processing system can then estimate the three-dimensional absolute coordinate position information and orientation information such that the error between the matching points 612 in the specific street view image and the projection data 624 is minimized. The information processing system can obtain final three-dimensional absolute coordinate position information and orientation information for the specific street view image 630 by iteratively estimating the three-dimensional absolute coordinate position information and orientation information based on the estimation results. A specific method for this has been described above with reference to FIG. 4.
[0066] While Fig. 6 illustrates estimating the three-dimensional absolute coordinate position information and direction information of a particular street view image 630 using one pair of ground control point 622 and matching point 612, the present invention is not limited to this, and multiple pairs of control features and matching points can be used to estimate the three-dimensional absolute coordinate position information and direction information of a particular street view image 630. In this case, the weights described in Fig. 5 are applied to each control feature, thereby estimating the three-dimensional absolute coordinate position information and direction information of a particular street view image 630.
[0067] FIG. 7 is a diagram illustrating a method for estimating three-dimensional absolute coordinate position information and direction information of a street view image 730 using building control points according to one embodiment of the present disclosure. According to one embodiment, the control feature may include a building control point 722. The building control point may be obtained based on feature matching between a street view image and an aerial image including three-dimensional absolute coordinate position information, and the building control point may include three-dimensional absolute coordinate position information and a visual descriptor associated with a specific point on a building in the street view image. Here, a specific point on a building in a street view image may be referred to as a matching point (feature point) obtained through feature matching between multiple street view images. For example, the building control point 722 may be obtained based on feature matching between a street view image 710 and an aerial image 720 including three-dimensional absolute coordinate position information.
[0068] According to one embodiment, the information processing system may estimate three-dimensional absolute coordinate position information and direction information of the specific street view image 730 based on the building control points 722 and matching points 712 in the specific street view image 730. For example, the information processing system may project the building control points 722 onto the specific street view image 730 to generate projection data 724. The information processing system may then estimate the three-dimensional absolute coordinate position information and direction information such that the error between the matching points 712 in the specific street view image and the projection data 724 is minimized. The information processing system may iteratively estimate the three-dimensional absolute coordinate position information and direction information based on the estimation results to obtain final three-dimensional absolute coordinate position information and direction information for the specific street view image 730. A specific method for this has been described above with reference to FIG. 4.
[0069] 7 illustrates estimating the three-dimensional absolute coordinate position information and direction information of a particular street view image 730 using one pair of a building control point 722 and a matching point 712, but is not limited to this, and multiple pairs of control features and matching points can be used to estimate the three-dimensional absolute coordinate position information and direction information of a particular street view image 730. In this case, the weights described in FIG. 5 are applied to each control feature, thereby estimating the three-dimensional absolute coordinate position information and direction information of a particular street view image 730.
[0070] FIG. 8 is a diagram illustrating a method for estimating three-dimensional absolute coordinate position information and direction information of a street view image 830 using ground control lines according to one embodiment of the present disclosure. According to one embodiment, the control feature may include a ground control line 822. The ground control line may be obtained based on feature matching between a street view image and an aerial image including three-dimensional absolute coordinate position information, and the ground control line may include three-dimensional absolute coordinate position information and a visual descriptor associated with a specific line on the ground of the street view image. Here, a specific line on the ground of a street view image may be referred to as a matching line (feature line) obtained by feature matching between multiple street view images. For example, the matching line may be represented by two matching points. For example, the ground control line 822 may be obtained based on feature matching between a street view image 810 and an aerial image 820 including three-dimensional absolute coordinate position information.
[0071] According to one embodiment, the information processing system can estimate three-dimensional absolute coordinate position information and direction information of the specific street view image 830 based on the ground control line 822 and matching line 812 in the specific street view image 830. For example, the information processing system can project the ground control line 822 onto the specific street view image 830 to generate projection data 824. The information processing system can then estimate the three-dimensional absolute coordinate position information and direction information such that the error between the matching line 812 in the specific street view image 830 and the projection data 824 is minimized. For example, the information processing system can estimate the three-dimensional absolute coordinate position information and direction information such that the error between the first normal vector n1 of the plane connecting the shooting position (origin) and the projection data 824 and the second normal vector n2 of the plane connecting the shooting position (origin) and the matching line 812 is minimized. The information processing system can obtain final three-dimensional absolute coordinate position information and direction information for the specific street view image 830 by iteratively estimating the three-dimensional absolute coordinate position information and direction information based on the estimation results. A specific method for this has been described above with reference to FIG.
[0072] 8 illustrates estimating the three-dimensional absolute coordinate position information and direction information of the specific street view image 830 using one pair of ground control line 822 and matching line 812, but is not limited to this, and the three-dimensional absolute coordinate position information and direction information of the specific street view image 830 can be estimated using multiple pairs of control features and matching points. In this case, the weights described in FIG. 5 are applied to each control feature, thereby estimating the three-dimensional absolute coordinate position information and direction information of the specific street view image 830.
[0073] 9 is a diagram illustrating a method for estimating camera parameters for a street view image 900 according to an embodiment of the present disclosure. According to an embodiment, an information processing system can estimate camera parameters for a particular street view image 900 based on the three-dimensional absolute coordinate position information and direction information estimation result of the particular street view image 900. Here, the particular street view image 900 may be a 360-degree panoramic image generated by equirectangular projection.
[0074] In one embodiment, the information processing system can estimate camera region parameters that define boundaries for dividing the specific street view image 900 into six images (cam1 to cam6). The boundaries can be represented by lines drawn on a sphere. The specific street view image 900 is a 360-degree panoramic image, and can be arbitrarily processed to be provided to the user without creating a sense of dissonance with the real world. Therefore, the specific street view image 900 does not need to match the geometric features of the panoramic image. The information processing system can estimate the camera region parameters to align the specific street view image 900 with the geometric features of the panoramic image and improve the quality of estimating three-dimensional absolute coordinate position information and direction information for the specific street view image 900.
[0075] Specifically, assuming that a specific street view image 900, which is a 360-degree panoramic image, is an image obtained by stitching together images captured by six cameras, the areas corresponding to each camera (e.g., cam1 to cam6) can be estimated. For example, the area corresponding to the first camera (cam1) can be defined by estimating four nodes 912, 914, 916, and 918 associated with the area corresponding to the first camera (cam1) and determining edges (lines) connecting these nodes. Similarly, boundaries for dividing the areas corresponding to the remaining five cameras (e.g., cam2 to cam5) can be defined. In this case, as shown in FIG. 9 , a boundary 920 within the panoramic image can be redefined in accordance with the geometric characteristics of the panoramic image. For example, because the boundary 920 is redefined in accordance with the geometric characteristics of the panoramic image, the boundary of the image corresponding to cam1 may not match the boundary of the adjacent images corresponding to cam2 and cam3.
[0076] According to one embodiment, the information processing system can determine camera distortion parameters for transforming six images (e.g., cam1 to cam6) into an undistorted planar image. The information processing system can use the camera distortion parameters for transforming into an undistorted planar image together with the camera region parameters to estimate three-dimensional absolute coordinate position information and orientation information of the street view image.
[0077] FIG. 10 is a diagram illustrating a specific example of a method for estimating camera parameters for a street view image according to an embodiment of the present disclosure. According to an embodiment, an information processing system may estimate camera region parameters for a street view image, which is a 360-degree panoramic image. A first example 1010 illustrates an example in which a street view image is divided into six images, each with a 90-degree field of view (FOV). For example, a front image (Front) may be defined as an area corresponding to 45 degrees left, 45 degrees right, 45 degrees up, and 45 degrees down based on the center of the street view image, which is a 360-degree panoramic image. A left image (Left) may be defined so that the FOV is 90 degrees based on a position rotated 90 degrees left from the front. Similarly, a right image (Right), a rear image (Back), an upper image (Up), and a lower image (Down) may be defined. Each image may include four nodes corresponding to a rectangular grid. For example, the front image may include four nodes 1012, 1014, 1016, 1018 that correspond to a rectangular grid.
[0078] A second example 1020 shows an example result of estimating camera region parameters to align a street view image with the geometric features of a panoramic image. For example, as described with reference to Figure 9, camera region parameters can be estimated to define boundaries for dividing a street view image into six images. For example, a region of the front image (Front) defined based on the estimated camera region parameters can include four nodes 1022, 1024, 1026, and 1028 corresponding to a distorted rectangle.
[0079] Through this configuration, the street view image is aligned with the geometric features of the panoramic image, and based on this, three-dimensional absolute coordinate position information and direction information relative to the street view image can be estimated, thereby obtaining high-quality three-dimensional absolute coordinate position information and direction information.
[0080] FIG. 11 is a diagram illustrating an example of a loss function used when performing filtering according to an embodiment of the present disclosure. According to one embodiment, the loss function can be used as a criterion for determining an outlier based on the error between a matching point in a street view image and projection data of a control feature. Here, the X-axis of the illustrated graph is associated with the magnitude of the error between the matching point in the street view image and projection data of the control feature. Furthermore, the Y-axis of the graph is associated with the weight assigned to the control feature.
[0081] According to one embodiment, the loss function used to remove outliers included in the control features can be scaled differently depending on the type of control feature, for example, a larger scale loss function can be used for ground control lines than for ground control points and building control points.
[0082] The first example 1110 shows an example of a loss function with a small scale (e.g., scale: 1). For example, the first example 1110 may show an example of a loss function for ground control points and building control points. In the first example 1110, data with an error magnitude equal to or greater than 1 may be determined as an outlier and removed.
[0083] The second example 1120 shows an example of a loss function with a large scale (e.g., scale: 2). For example, the second example 1120 may show an example of a loss function for a ground control line. In the second example 1120, data with an error magnitude of 2 or greater may be determined to be an outlier and removed. Because ground control lines have larger errors than ground control points and building control points due to their line characteristics, a loss function with a larger scale than ground control points and building control points may be used for ground control lines.
[0084] As a result, for ground control lines using a large-scale loss function, even if the error between the matching point in the Street View image and the projection data of the control feature is relatively larger than for ground control points or building control points, it can still be determined to be normal data.In contrast, for ground control points and building control points using a small-scale loss function, even if the error between the matching point in the Street View image and the projection data of the control feature is relatively smaller than for ground control lines, it can still be determined to be normal data.
[0085] FIG. 11 shows an example in which the scale value of the loss function is set to 1 or 2, but this is for convenience of explanation, and the scale value is not limited to this.
[0086] 12 is a flowchart illustrating an example method 1200 for estimating absolute poses from street view images according to one embodiment of the present disclosure. The method 1200 may begin with a processor (e.g., at least one processor of an information processing system and / or a user terminal) receiving a plurality of control features associated with a specific area (S1210). The processor may then receive a plurality of matching points obtained through feature matching between a plurality of street view images associated with the specific area (S1220). Here, each of the plurality of control features may include three-dimensional absolute coordinate position information, and the plurality of control features may include at least one of ground control points, building control points, or ground control lines.
[0087] According to one embodiment, the plurality of control features may include three-dimensional absolute coordinate location information and a visual descriptor associated with a particular point on the Street View image. In one embodiment, the plurality of control features may include a plurality of ground control points, each of which may include three-dimensional absolute coordinate location information and a visual descriptor associated with a particular point on the ground. In another embodiment, the plurality of control features may include a plurality of building control points, each of which may include three-dimensional absolute coordinate location information and a visual descriptor associated with a particular point on a building. In another embodiment, the plurality of control features may include a plurality of ground control lines, each of which may include three-dimensional absolute coordinate location information and a visual descriptor associated with a particular line on the ground.
[0088] According to one embodiment, the processor may estimate three-dimensional absolute coordinate position information and orientation information of at least one of the plurality of street view images based on a plurality of control features and the plurality of matching points (S1230).
[0089] Specifically, the processor may estimate three-dimensional absolute coordinate position information and orientation information with a first accuracy for a specific street view image included in the plurality of street view images. For example, the processor may identify a first set of matching points associated with the specific street view image from among the plurality of matching points. The processor may also identify a first set of control features corresponding to the first set of matching points from among the plurality of control features. The processor may also project the first set of control features onto the specific street view image to generate a first set of projection data. The processor may then estimate three-dimensional absolute coordinate position information and orientation information with a first accuracy for the specific street view image based on the first set of matching points and the first set of projection data. Here, the three-dimensional absolute coordinate position information and orientation information with a first accuracy may be estimated such that an error between the first set of matching points and the first set of projection data is minimized.
[0090] According to one embodiment, estimating three-dimensional absolute coordinate position information and orientation information of at least one of the plurality of Street View images may further include determining a weight for each of the first set of control features. In one embodiment, if the control features are ground control points, the processor may assign a lower weight to ground control points that are farther away from the particular Street View image. In another embodiment, if the control features are building control points, the processor may assign a higher weight to building control points that are farther away from the particular Street View image.
[0091] Furthermore, the processor may perform a first filtering process based on the estimated three-dimensional absolute coordinate position information and orientation information with a first accuracy. Specifically, the processor may generate a second set of projection data by projecting a first set of control features onto a specific Street View image based on the estimated three-dimensional absolute coordinate position information and orientation information with a first accuracy. The processor may then determine a second set of control features by removing some of the first set of control features based on the error between the first set of matching points and the second set of projection data. The processor may then determine a second set of matching points by removing matching points corresponding to the removed control features from the first set of matching points. In one embodiment, when removing some of the first set of control features, a loss function with a larger scale may be used for the ground control lines than for ground control points and building control points. Therefore, in the case of ground control lines using a larger scale loss function, there may be a higher probability that the ground control lines will be determined to be normal data even if the error between the matching points in the Street View image and the projection data of the control features is relatively larger than that for ground control points or building control points.
[0092] Furthermore, the processor may estimate three-dimensional absolute coordinate position information and direction information with a second accuracy based on the estimated three-dimensional absolute coordinate position information and direction information with a first accuracy. Here, the second accuracy may be higher than the first accuracy. Specifically, the processor may determine individual weights for each of the second set of matching points and the second set of control features. For example, the processor may calculate individual confidence levels for each of the second set of matching points and the second set of control features based on the confidence levels of the three-dimensional absolute coordinate position information and direction information with the first accuracy. The processor may then determine individual weights for each of the second set of matching points and the second set of control features based on the calculated individual confidence levels. The processor may then estimate camera parameters for a specific street view image based on the second set of matching points, the second set of control features, and the individual weights. For example, the processor may estimate camera region parameters defining boundaries for dividing the specific street view image into six images. Here, the specific street view image may be a 360-degree panoramic image generated using equirectangular projection. The processor may further determine camera distortion parameters for transforming the six images into an undistorted planar image, and the processor may estimate three-dimensional absolute coordinate position information and orientation information of a second precision for the particular street view image based on the second set of matching points, the second set of control features, the individual weights, and the camera parameters.
[0093] Furthermore, the processor may perform second filtering based on the estimated three-dimensional absolute coordinate position information and orientation information with the second accuracy. Specifically, the processor may project the second set of control features onto the specific street view image based on the estimated three-dimensional absolute coordinate position information and orientation information with the second accuracy to generate a third set of projection data. Then, the processor may determine a third set of control features by removing some of the second set of control features based on the error between the second set of matching points and the third set of projection data. Thereafter, the processor may determine a third set of matching points by removing matching points corresponding to the removed control features from the second set of matching points.
[0094] The processor may further project the third set of control features onto the specific Street View image to generate a fourth set of projection data. The processor may then determine whether errors between the third set of matching points and the fourth set of projection data are all equal to or less than a predetermined threshold. In one embodiment, if it is determined that errors between the third set of matching points and the fourth set of projection data are all equal to or less than the predetermined threshold, the processor may determine the three-dimensional absolute coordinate position information and orientation information with the second accuracy as the final three-dimensional absolute coordinate position information and orientation information for the specific Street View image. On the other hand, if it is determined that errors between the third set of matching points and the fourth set of projection data are all greater than the predetermined threshold, the processor may repeat the step of estimating the three-dimensional absolute coordinate position information and orientation information for the specific Street View image.
[0095] 12 and the above description are merely examples, and the scope of the present disclosure is not limited thereto. For example, at least one step may be added / modified / deleted, or the order of the steps may be changed.
[0096] The above-described method can be provided as a computer program stored on a computer-readable recording medium for execution by a computer. The medium may permanently store a computer-executable program or may temporarily store the program for execution or download. The medium may be various recording or storage means combined with one or more pieces of hardware, but is not limited to media directly connected to a computer system and may be distributed over a network. Examples of media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and media configured to store program instructions, including ROM, RAM, and flash memory. Other examples of media include recording or storage media managed by app stores that distribute applications, or by sites or servers that provide or distribute various other software.
[0097] The methods, operations, or techniques of the present disclosure may be implemented by various means. For example, the techniques may be implemented in hardware, firmware, software, or a combination thereof. Those of ordinary skill in the art will understand that the various illustrative logic blocks, modules, circuits, and algorithm steps described in connection with the disclosure herein may also be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate such interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design requirements for the overall system. Those of ordinary skill in the art may implement the described functions in various ways for each particular application, but such implementation should not be interpreted as causing a departure from the scope of the present disclosure.
[0098] In a hardware implementation, the processing units used to execute the techniques may be implemented within one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, electronic devices, other electronic units designed to perform the functions described in this disclosure, computers, or combinations thereof.
[0099] Accordingly, the various illustrative logic blocks, modules, and circuits described in connection with this disclosure may be implemented or performed with a general-purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, such as a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other configuration.
[0100] In a firmware and / or software implementation, the techniques may be implemented as instructions stored on a computer-readable medium, such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, compact disc (CD), magnetic or optical data storage device, etc. The instructions may be executable by one or more processors and may cause the processors to perform certain aspects of the functions described in this disclosure.
[0101] While the above-described embodiments have been described as utilizing aspects of the presently disclosed subject matter on one or more stand-alone computer systems, the present disclosure is not limited thereto and may be implemented in connection with any computing environment, such as a network or distributed computing environment. Furthermore, aspects of the subject matter in this disclosure may be implemented on multiple processing chips or devices, and storage may be affected similarly across multiple devices. These devices may include PCs, network servers, and handheld devices.
[0102] Although the present disclosure has been described herein with reference to several embodiments, various modifications and changes can be made thereto without departing from the scope of the present disclosure, as would be understood by one of ordinary skill in the art to which the presently disclosed invention pertains, and it should be understood that these modifications and changes also fall within the scope of the claims appended hereto.
Claims
1. 1. A method for absolute pose estimation of street view images performed by at least one processor, comprising: receiving a plurality of control features associated with a particular region; receiving a plurality of matching points obtained through feature matching between a plurality of street view images related to the specific area; and estimating three-dimensional absolute coordinate position information and orientation information of at least one of the plurality of street view images based on the plurality of control features and the plurality of matching points; each of the plurality of control features includes three-dimensional absolute coordinate position information; The method for absolute pose estimation of street view imagery, wherein the plurality of control features include at least one of ground control points, building control points, or ground control lines.
2. the plurality of street view images includes a specific street view image; The step of estimating three-dimensional absolute coordinate position information and direction information of at least one of the plurality of street view images includes: identifying a first set of matching points from the plurality of matching points that are associated with the particular street view image; identifying a first set of control features of the plurality of control features that correspond to the first set of matching points; projecting the first set of control features onto the particular street view image to generate a first set of projection data; and estimating three-dimensional absolute coordinate position and orientation information to a first precision for the particular street view image based on the first set of matching points and the first set of projection data.
3. 3. The method of claim 2, wherein the first precision 3D absolute coordinate position information and orientation information is estimated so as to minimize an error between the first set of matching points and the first set of projection data.
4. The step of estimating three-dimensional absolute coordinate position information and direction information of at least one of the plurality of street view images includes: determining a weight for each of the first set of control features; 3. The method of claim 2, wherein the first precision of three-dimensional absolute coordinate position information and orientation information of the particular street view image is estimated based on the first set of matching points, the first set of projection data, and respective weights of the first set of control features.
5. Determining a weight for each of the first set of control features comprises:
5. The method for estimating an absolute pose of street view images according to claim 4, further comprising the step of: if the control feature is a ground control point, assigning a lower weight to a ground control point that is farther away from the specific street view image.
6. Determining a weight for each of the first set of control features comprises: The method for estimating an absolute pose of street view images according to claim 4 , further comprising the step of: if the control feature is a building control point, assigning a higher weight to the building control point when the distance between the specific street view image and the building control point is greater.
7. projecting the first set of control features onto a specific street view image based on the estimated three-dimensional absolute coordinate position information and orientation information with the first precision to generate a second set of projection data; determining a second set of control features obtained by removing a portion of the first set of control features based on an error between the first set of matching points and the second set of projection data; The method of claim 2 , further comprising: determining a second set of matching points by removing matching points corresponding to the removed control features from the first set of matching points.
8. 8. The method of claim 7, wherein a loss function with a larger scale is used for ground control lines than for ground control points and building control points when removing some of the first set of control features.
9. determining an individual weight for each of the second set of matching points and the second set of control features; estimating camera parameters for the particular street view image based on the second set of matching points, the second set of control features, and the individual weights; and estimating three-dimensional absolute coordinate position information and orientation information of the particular street view image to a second precision based on the second set of matching points, the second set of control features, the individual weights, and the camera parameters; The method of claim 7 , wherein the second accuracy is higher than the first accuracy.
10. The method for determining the individual weights comprises: calculating individual confidence levels for the second set of matching points and the second set of control features based on the confidence levels of the three-dimensional absolute coordinate position information and orientation information with the first accuracy; and determining an individual weight for each of the second set of matching points and the second set of control features based on the calculated individual confidences.
11. The step of estimating camera parameters includes: The method of claim 9, further comprising estimating camera region parameters that define boundaries for dividing the particular street view image into six images.
12. The method for estimating an absolute pose of street view images according to claim 11 , wherein the specific street view image is a 360-degree panoramic image generated using equirectangular projection.
13. The step of estimating camera parameters includes: The method for absolute pose estimation of street view images of claim 11 , further comprising determining camera distortion parameters for transforming the six images into an undistorted planar image.
14. projecting the second set of control features onto a specific street view image based on the estimated three-dimensional absolute coordinate position information and orientation information with the second precision to generate a third set of projection data; determining a third set of control features by removing a portion of the second set of control features based on an error between the second set of matching points and the third set of projection data; 10. The method of claim 9, further comprising: determining a third set of matching points by removing matching points corresponding to the removed control features from the second set of matching points.
15. projecting the third set of control features onto a particular street view image to generate a fourth set of projection data; 15. The method of claim 14, further comprising: determining whether errors between the third set of matching points and the fourth set of projection data are all below a predetermined threshold.
16. the plurality of control features includes a plurality of ground control points; The method of claim 1 , wherein each ground control point includes three-dimensional absolute coordinate position information and a visual descriptor associated with a specific point on the ground.
17. the plurality of control features includes a plurality of building control points; The method of claim 1 , wherein each building control point includes three-dimensional absolute coordinate position information and a visual descriptor associated with a specific point on a building.
18. the plurality of control features includes a plurality of ground control lines; The method of claim 1 , wherein each ground control line includes three-dimensional absolute coordinate position information and a visual descriptor associated with a particular line on the ground.
19. A non-transitory computer-readable storage medium having recorded thereon instructions for performing the method of claim 1 on a computer.
20. An information processing system, communication module, memory, and at least one processor coupled to said memory and configured to execute at least one computer readable program contained in said memory; The at least one computer readable program receiving a plurality of control features associated with a particular region; receiving a plurality of matching points obtained through feature matching between a plurality of street view images related to the specific area; instructions for estimating three-dimensional absolute coordinate position information and orientation information of at least one of the plurality of street view images based on the plurality of control features and the plurality of matching points; each of the plurality of control features includes three-dimensional absolute coordinate position information; The information processing system, wherein the plurality of control features include at least one of ground control points, building control points, or ground control lines.
Citation Information
Patent Citations
Estimating location method and apparatus for autonomous driving
KR102219843B1
Method, apparatus and system for generating and exposing advertisement content for three-dimensional visual effect
KR102512259B1
High-precision multi-layer visual and semantic map by autonomous units
US10794710B1
Image processing apparatus, method, and program
US20180150974A1