Method and system for generating visual feature maps using 3D models and street view images

By aligning 3D models with Street View imagery to generate accurate visual feature maps, the method addresses GPS inaccuracy and cost issues, enhancing map quality for autonomous driving applications.

JP2026503824AInactive Publication Date: 2026-01-30NAVER CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025526205
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-08
Filing Date
2023-11-06
Publication Date
2026-01-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing methods for generating visual feature maps for autonomous driving using Street View footage are limited by inaccurate GPS location information and the high cost of installing precise GPS equipment, and they cannot utilize pre-captured Street View data effectively.

Method used

A method and system that aligns a 3D model with Street View imagery using absolute coordinate positions to generate a visual feature map, extracting feature points and determining their accurate coordinates, thereby reducing costs and leveraging existing Street View data.

Benefits of technology

This approach improves the quality of visual feature maps for autonomous driving by accurately aligning 3D geometric information with Street View imagery, reducing costs and effort compared to traditional multi-sensor surveying systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026503824000001_ABST
    Figure 2026503824000001_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method for generating a visual feature map using a 3D model and Street View imagery, the method being performed by at least one processor, and includes the steps of receiving a 3D model of a specific area including 3D geometric information expressed by absolute coordinate positions, receiving first Street View imagery captured at a first node within the specific area, projecting at least a portion of the 3D geometric information included in the 3D model onto the first Street View imagery based on absolute coordinate position information and direction information of the first Street View imagery to render a depth map associated with the first Street View imagery, extracting a first set of feature points from the first Street View imagery, determining absolute coordinate position information of all or some of the first set of feature points based on the depth map, and generating a first visual feature map associated with the first Street View imagery by associating and storing the absolute coordinate position information and a visual feature descriptor for each of the first set of feature points.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a method and system for generating a visual feature map using a 3D model and street view video, and more specifically, to a method and system for automatically generating a visual feature map using street view video in which absolute coordinate position information of a 3D model is aligned. [Background technology]

[0002] Autonomous driving technology refers to technology that enables a vehicle to drive autonomously with minimal or no human intervention by recognizing the surrounding environment using radar, lidar (light detection and ranging), GPS (global positioning system), cameras, etc., mounted on the vehicle. Because the driving environment is influenced by various factors, such as vehicles on the road, traffic structures, and buildings in areas outside the road, accurately recognizing the vehicle's surrounding environment using devices mounted on the vehicle is an important factor in ensuring safety toward the commercialization of autonomous driving technology.

[0003] To accurately recognize the vehicle's surroundings, it is essential to generate a visual feature map of the driving environment using video information from the autonomous vehicle's perspective, which contains highly accurate 3D geometric information and location information. However, to obtain 3D geometric information and video information, mapping data for the target area must be acquired using a vehicle equipped with a vehicle-based multi-sensor surveying system (Mobile Mapping System, MMS) that includes expensive lidar sensors, cameras, and high-precision GPS. However, this requires the vehicle to directly drive in the area where data is needed to acquire the data, which is costly and labor-intensive.

[0004] Additionally, methods for generating visual feature maps using captured Street View footage to provide Street View services have been considered, but there is a problem with the location information acquired by the vehicle's GPS equipment when capturing Street View footage, which has an error of about 5 to 10 meters. To resolve this inaccuracy in location information, expensive GPS equipment could be installed in the vehicle to capture Street View footage while acquiring highly accurate location information, but this method is expensive and has the problem of not being able to utilize Street View footage that has already been captured. Summary of the Invention [Problem to be solved by the invention]

[0005] The present disclosure provides a method for solving the above problems, a computer-readable non-transitory recording medium having instructions recorded thereon, and an apparatus (system). [Means for solving the problem]

[0006] The present disclosure may be embodied in various forms, including a method, an apparatus (system), or a computer-readable non-transitory storage medium having instructions recorded thereon.

[0007] According to one embodiment of the present disclosure, a method for generating a visual feature map using a 3D model and Street View imagery, performed by at least one processor, includes receiving a 3D model of a specific area including 3D geometric information expressed by absolute coordinate positions; receiving first Street View imagery captured at a first node within the specific area; projecting at least a portion of the 3D geometric information included in the 3D model onto the first Street View imagery based on the absolute coordinate position information and direction information of the first Street View imagery, and rendering a depth map associated with the first Street View imagery; extracting a first set of feature points from the first Street View imagery; determining absolute coordinate position information of all or some of the first set of feature points based on the depth map; and generating a first visual feature map associated with the first Street View imagery by associating and storing the absolute coordinate position information and a visual feature descriptor for each of the first set of feature points.

[0008] A non-transitory computer-readable recording medium is provided having recorded thereon instructions for causing a computer to perform a method according to an embodiment of the present disclosure.

[0009] According to an embodiment of the present disclosure, an information processing system includes: a communications module; a memory; and at least one processor connected to the memory and configured to execute at least one computer-readable program stored in the memory, the at least one program including instructions to receive a 3D model of a specific area including 3D geometric information expressed by absolute coordinate positions; receive first street view imagery captured at a first node within the specific area; project at least a portion of the 3D geometric information included in the 3D model onto the first street view imagery based on the absolute coordinate position information and direction information of the first street view imagery; render a depth map associated with the first street view imagery; extract a first set of feature points from the first street view imagery; determine absolute coordinate position information of all or a portion of the first set of feature points based on the depth map; and associate and store the absolute coordinate position information and a visual feature descriptor for each of the first set of feature points, thereby generating a first visual feature map associated with the first street view imagery. [Effects of the Invention]

[0010] According to one embodiment of the present disclosure, instead of using a vehicle equipped with a vehicle-based multi-sensor surveying system, 3D building model information generated by aerial image surveying, etc., and Street View images already captured for the Street View service are used, thereby reducing the cost and effort of obtaining a visual feature map.

[0011] According to one embodiment of the present disclosure, an object related to autonomous driving is extracted from a plurality of pieces of environmental information included in a street view image using a binary mask, and a visual feature map is generated for the extracted object, thereby improving the quality of the visual feature map.

[0012] The effects of the present disclosure are not limited to those mentioned above, and other effects not mentioned will be clearly understood by a person with ordinary skill in the art to which the present disclosure pertains (hereinafter referred to as "a person skilled in the art") from the description in the claims. [Brief explanation of the drawings]

[0013] Embodiments of the present disclosure will now be described with reference to the accompanying drawings, as described below, in which like reference numerals indicate like, but not limited to, elements and in which:

[0014] [Figure 1] 1 illustrates an example method for aligning a 3D model with street view data, according to an embodiment of the present disclosure. [Figure 2] 1 is a schematic diagram illustrating a configuration in which an information processing system is communicably connected to a plurality of user terminals according to an embodiment of the present disclosure. [Figure 3] FIG. 2 is a block diagram illustrating an internal configuration of a user terminal and an information processing system according to an embodiment of the present disclosure. [Figure 4] 1 illustrates an example process for generating a visual feature map associated with street view imagery based on a 3D model and street view imagery, according to one embodiment of the present disclosure. [Figure 5] FIG. 2 is a block diagram illustrating a specific method for generating a visual feature map associated with a first street view image based on aligned street view data and a three-dimensional model, according to one embodiment of the present disclosure. [Figure 6] FIG. 10 is a block diagram illustrating a specific method for generating a visual feature map associated with a first street view image based on aligned street view data, a 3D model, a binary mask, and 3D planar information about road traffic structures, according to another embodiment of the present disclosure. [Figure 7] 1 is a diagram illustrating an example process for generating a binary mask based on street view imagery, according to an embodiment of the present disclosure. [Figure 8]10A-10C are diagrams illustrating an example of a process for performing stereo matching on road traffic structures in a first street view image and a second street view image according to an embodiment of the present disclosure. [Figure 9] 1 is a diagram illustrating an example of a three-dimensional planar information map relating to a plurality of road traffic structures according to an embodiment of the present disclosure. [Figure 10] 1 illustrates an example process for extracting a first set of feature points from Street View footage using a plurality of planar images generated based on the Street View footage, according to an embodiment of the present disclosure. [Figure 11] 1 is a diagram illustrating an example of a process for obtaining absolute coordinate position information for a road traffic structure according to an embodiment of the present disclosure. [Figure 12] 1 is a diagram illustrating an example of a visual feature map for a specific region, according to an embodiment of the present disclosure. [Figure 13] 1 is a flowchart illustrating an example of a method for generating a visual feature map using a 3D model and street view video according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0015] Hereinafter, specific contents for implementing the present disclosure will be described in detail with reference to the accompanying drawings. However, in the following description, specific descriptions of well-known functions and configurations will be omitted if they may obscure the gist of the present disclosure.

[0016] In the accompanying drawings, identical or corresponding components are denoted by the same reference numerals. In addition, in the following description of the embodiments, duplicated descriptions of identical or corresponding components will be omitted. However, omission of a description of a component does not mean that the component is not included in any of the embodiments.

[0017] The advantages and features of the disclosed embodiments, as well as methods for achieving them, will be apparent from the following embodiments described in conjunction with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below, and may be embodied in various other forms. The present embodiments are provided merely to complete the disclosure and fully convey the scope of the invention to those skilled in the art.

[0018] The terms used in this specification will be briefly explained, and the disclosed embodiments will be specifically described. The terms used in this specification are currently commonly used and generally selected as much as possible while taking into consideration the functions of the present disclosure. However, these terms may change depending on the intentions of engineers in the relevant field, precedents, the emergence of new technology, etc. In addition, in certain cases, the applicant may arbitrarily select terms, and in such cases, the meanings of these terms will be described in detail in the relevant description of the invention. Therefore, the terms used in this disclosure should be defined based on the meanings of the terms and the overall content of the present disclosure, rather than simply by the names of the terms.

[0019] In this specification, the singular expression includes the plural expression unless the context clearly dictates otherwise. Furthermore, the plural expression includes the singular expression unless the context clearly dictates otherwise. Throughout this specification, when a part includes one element, this does not mean that it may further include other elements, but does not exclude other elements, unless otherwise specified.

[0020] Furthermore, the terms "module" and "module" used in this specification refer to software or hardware components, and the "module" or "module" may perform either function. However, the term "module" or "module" is not limited to software or hardware. A "module" or "module" may be configured to be provided on an addressable storage medium or to execute one or more processors. Thus, as an example, a "module" or "module" may include components such as software components, object-oriented software components, class components, and task components, as well as at least one of processes, functions, attributes, procedures, subroutines, program code segments, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. Components and "modules" or "modules" may be combined into fewer components and "modules" or "modules" so that the functionality provided by them can be further separated into additional components, "modules," or "modules."

[0021] According to one embodiment of the present disclosure, a "module" or "unit" may be embodied as a processor and memory. "Processor" should be broadly interpreted to include a general-purpose processor, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a controller, a microcontroller, a state machine, etc. In some environments, "processor" may refer to an application specific integrated circuit (ASIC), a programmable logic device (PLD), a field programmable gate array (FPGA), etc. "Processor" may also refer to a combination of processing devices, such as, for example, a combination of a DSP and a microprocessor, a combination of multiple microprocessors, a combination of one or more microprocessors in conjunction with a DSP core, or any other such configuration. Additionally, "memory" should be broadly interpreted to include any electronic component capable of storing electronic information. "Memory" may refer to various types of processor-readable media, such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable-programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, magnetic or optical data storage devices, registers, etc. Memory is in electronic communication with a processor if the processor can read information from and / or store information in the memory. Memory that is integrated into a processor is in electronic communication with the processor.

[0022] In the present disclosure, a "system" may include at least one of a server device and a cloud device, but is not limited thereto. For example, a system may be configured with one or more server devices. As another example, a system may be configured with one or more cloud devices. As yet another example, a system may be configured and operated by both a server device and a cloud device.

[0023] In this disclosure, "display" may refer to any display device associated with a computing device, for example, any display device capable of displaying any information / data controlled by or provided by the computing device.

[0024] In the present disclosure, "each of a plurality of A's" or "each of a plurality of A's" may refer to each of all the components included in the plurality of A's, or may refer to each of some of the components included in the plurality of A's.

[0025] In this disclosure, "street view data" may refer to data including not only road view data including video and location information taken from a roadway, but also foot view data including video and location information taken from a sidewalk. "Street view data" may further include video and location information taken from any point outdoors (or indoors looking outdoors) in addition to roadways and sidewalks.

[0026] 1 illustrates an example method for aligning a 3D model 110 with street view data 120 according to one embodiment of the present disclosure. An information processing system may acquire / receive the 3D model 110 and street view data 120 for a particular area.

[0027] The 3D model 110 may include 3D geometric information expressed in absolute coordinate positions and corresponding texture information. Here, the location information included in the 3D model 110 may be information with higher accuracy than the location information included in the street view data 120. Furthermore, the texture information included in the 3D model 110 may be information with lower quality (e.g., lower resolution) than the texture information included in the street view data 120. According to an embodiment, the 3D geometric information expressed in absolute coordinate positions may be generated based on an aerial photograph taken of a specific area from above the specific area.

[0028] The 3D model 110 of the specific area may include a 3D building model 112, a digital elevation model (DEM) 114, a true ortho image 116 of the specific area, a digital surface model (DSM), a road layout, a road DEM, etc. As a specific example, the 3D model 110 of the specific area may be, but is not limited to, a model generated based on a digital surface model (DSM) including geometric information about the ground of the specific area and the corresponding true ortho image 116 of the specific area. In one embodiment, a precise ortho image 116 of the specific area can be generated based on multiple aerial photographs and absolute coordinate position information and direction information of each aerial photograph.

[0029] The street view data 120 may include a plurality of street view images captured at a plurality of nodes within a specific area, and absolute coordinate position information for each of the plurality of street view images. Here, the position information included in the street view data 120 may be less accurate than the position information included in the 3D model 110, and the texture information included in the street view images may be higher quality (e.g., higher resolution) than the texture information included in the 3D model 110. For example, the position information included in the street view data 120 may be position information acquired using a GPS device when capturing street view images at the node. Position information acquired using a vehicle's GPS device may have an error of approximately 5 to 10 meters. Furthermore, the street view data may include direction information for each of the plurality of street view images (i.e., image capture direction information).

[0030] The information processing system may perform map matching 130 between the 3D model 110 and the street view data 120. Specifically, the information processing system may perform feature matching between texture information included in the 3D model 110 and a plurality of street view images included in the street view data 120. To perform map matching 130, the information processing system may convert at least some of the plurality of street view images included in the street view data 120 into top view images. As a result of map matching 130, a plurality of map matching points / map matching lines 132 may be extracted.

[0031] A map matching point may indicate a corresponding pair between one point in the Street View imagery and one point in the 3D model 110. The map matching points may be categorized into various types depending on the type of 3D model 110 used in map matching 130, the location of the point, etc. For example, the map matching points may include at least one of a ground control point (GCP), which is a corresponding pair of points on the ground in a specific area, a building control point (BCP), which is a corresponding pair of points on a building in a specific area, or a structure control point, which is a corresponding pair of points on a structure in a specific area. In addition to the above-mentioned ground, buildings, and structures, map matching points may be extracted from any region of the Street View imagery and the 3D model 110.

[0032] A map matching line may indicate a corresponding pair of one line in the street view image and one line in the 3D model 110. The map matching line may be of various types depending on the type of 3D model 110 used in map matching 130, the position of the line, etc. For example, the map matching line may include at least one of a ground control line (GCL), which is a corresponding pair of lines on the ground in a specific area, a building control line (BCL), which is a corresponding pair of lines on a building in a specific area, a structure control line, which is a corresponding pair of lines on a structure in a specific area, or a lane control line, which is a corresponding pair of lines on a lane in a specific area. In addition to the ground, buildings, structures, and lanes described above, the map matching line may be extracted from any region of the street view image and the 3D model 110.

[0033] The information processing system may also perform feature matching 150 between a plurality of street view images to extract a plurality of feature point correspondence sets 152. According to an embodiment, for robust feature matching, feature matching 150 between a plurality of street view images may be performed using at least a portion of the 3D model 110. For example, feature matching 150 between street view images may be performed using the 3D building model 112 included in the 3D model 110.

[0034] The information processing system may then estimate absolute coordinate position information and direction information 160 for the plurality of street view images based on at least one of the plurality of map matching points / map matching lines 132 and at least a portion of the plurality of feature point correspondence sets 152. For example, the processor may use a bundle adjustment technique to estimate absolute coordinate position information and direction information 160 for the plurality of street view images. According to one embodiment, the estimated absolute coordinate position information and direction information 162 is information in an absolute coordinate system representing the 3D model 110 and may be six-degree-of-freedom (DoF) parameters. The absolute coordinate position information and direction information 162 estimated by this process may be more accurate than the absolute coordinate position information and direction information included in the street view data 120.

[0035] According to one embodiment of the present disclosure, at least one of the predefined map matching points / map matching lines 132 and the 3D model 110 can be used to automatically generate a visual feature map for each street view image in the street view data that is matched with the absolute coordinate position information of the 3D model. A method for generating a visual feature map for street view imagery will be described later with reference to FIGS. 4 to 13. This configuration can reduce the cost and effort required to generate a visual feature map for map-based services such as autonomous driving. Furthermore, the automatically generated visual feature map can be used for various map-based services such as autonomous driving services.

[0036] FIG. 2 is a schematic diagram illustrating a configuration in which an information processing system 230 according to an embodiment of the present disclosure is communicatively connected to multiple user terminals 210_1, 210_2, and 210_3. As illustrated, the multiple user terminals 210_1, 210_2, and 210_3 may be connected to the information processing system 230 via a network 220, which can provide map information services. Here, the multiple user terminals 210_1, 210_2, and 210_3 may include terminals of users receiving map information services, autonomous driving services, and the like. Furthermore, the multiple user terminals 210_1, 210_2, and 210_3 may be automobiles capturing street view images at nodes. In one embodiment, the information processing system 230 may include one or more server devices and / or databases, or one or more distributed computing devices and / or distributed databases based on a cloud computing service, which can store, provide, and execute computer-executable programs (e.g., downloadable applications) and data related to providing the map information service.

[0037] The map information service provided by the information processing system 230 may be provided to a user through an application or a web browser installed in each of the user terminals 210_1, 210_2, and 210_3. For example, the information processing system 230 may provide information corresponding to a street view image request, an image-based location recognition request, or the like received from the user terminals 210_1, 210_2, and 210_3 through an application or the like, or perform corresponding processing.

[0038] A plurality of user terminals 210_1, 210_2, and 210_3 can communicate with the information processing system 230 through the network 220. The network 220 may be configured to enable communication between the plurality of user terminals 210_1, 210_2, and 210_3 and the information processing system 230. Depending on the installation environment, the network 220 may be configured as a wired network such as Ethernet, a wired home network (Power Line Communication), a telephone line communication device, and RS-serial communication, a mobile communication network, a wireless network such as WLAN (Wireless LAN), Wi-Fi, Bluetooth, and ZigBee, or a combination thereof. The communication method is not limited, and may include a communication method utilizing a communication network that the network 220 may include (e.g., a mobile communication network, a wired Internet, a wireless Internet, a broadcast network, a satellite network, etc.), as well as short-range wireless communication between the user terminals 210_1, 210_2, and 210_3.

[0039] 2, a mobile phone terminal 210_1, a tablet terminal 210_2, and a PC terminal 210_3 are shown as examples of user terminals, but are not limited thereto, and the user terminals 210_1, 210_2, and 210_3 may be any computing device capable of wired and / or wireless communication and on which an application or a web browser, etc., can be installed and executed. For example, the user terminal may include an AI speaker, a smartphone, a mobile phone, a navigation system, a computer, a laptop, a digital broadcasting terminal, a PDA (Personal Digital Assistant), a PMP (Portable Multimedia Player), a tablet PC, a game console, a wearable device, an IoT (Internet of Things) device, a VR (Virtual Reality) device, an AR (Augmented Reality) device, a set-top box, etc. Also, although FIG. 2 shows three user terminals 210_1, 210_2, and 210_3 communicating with the information processing system 230 through the network 220, this is not limited thereto, and other numbers of user terminals may be configured to communicate with the information processing system 230 through the network 220.

[0040] FIG. 3 is a block diagram illustrating the internal configuration of a user terminal 210 and an information processing system 230 according to an embodiment of the present disclosure. The user terminal 210 may refer to any computing device capable of executing an application, a web browser, or the like, and capable of wired / wireless communication, and may include, for example, the mobile phone terminal 210_1, the tablet terminal 210_2, and the PC terminal 210_3 of FIG. 2 . As illustrated, the user terminal 210 may include a memory 312, a processor 314, a communication module 316, and an input / output interface 318. Similarly, the information processing system 230 may include a memory 332, a processor 334, a communication module 336, and an input / output interface 338. As illustrated in FIG. 3 , the user terminal 210 and the information processing system 230 may be configured to communicate information and / or data over the network 220 using their respective communication modules 316 and 336. Additionally, the input / output device 320 may be configured to input information and / or data to the user terminal 210 and / or output information and / or data generated from the user terminal 210 via the input / output interface 318 .

[0041] The memories 312 and 332 may include any non-transitory computer-readable recording medium. According to one embodiment, the memories 312 and 332 may include a permanent mass storage device such as a read only memory (ROM), a disk drive, a solid state drive (SSD), or a flash memory. As another example, a non-transitory mass storage device such as a ROM, an SSD, a flash memory, or a disk drive may be included in the user terminal 210 or the information processing system 230 as a permanent storage device separate from the memory. The memories 312 and 332 may also store an operating system and at least one program code (e.g., code for an application installed and run on the user terminal 210).

[0042] Such software components may be loaded from a computer-readable recording medium separate from the memory 312, 332. Such separate computer-readable recording medium may include a recording medium directly connectable to the user terminal 210 and the information processing system 230, but may also include a computer-readable recording medium such as a floppy drive, a disk, a tape, a DVD / CD-ROM drive, or a memory card. As another example, the software components may be loaded into the memory 312, 332 through a communication module rather than a computer-readable recording medium. For example, at least one program may be loaded into the memory 312, 332 based on a computer program installed by a file provided over the network 220 by a developer or a file distribution system that distributes application installation files.

[0043] The processors 314, 334 may be configured to process computer program instructions by performing basic arithmetic, logic, and input / output operations. The instructions may be provided to the processors 314, 334 by the memory 312, 332 or the communication modules 316, 336. For example, the processors 314, 334 may be configured to execute instructions received by program code stored in a storage device, such as the memory 312, 332.

[0044] The communication modules 316 and 336 may provide a configuration or function for the user terminal 210 and the information processing system 230 to communicate with each other via the network 220, and may provide a configuration or function for the user terminal 210 and / or the information processing system 230 to communicate with other user terminals or other systems (e.g., a separate cloud system, etc.). For example, a request or data (e.g., data related to street view images captured on the ground) generated by the processor 314 of the user terminal 210 via program code stored in a storage device such as the memory 312 may be transmitted to the information processing system 230 via the network 220 under the control of the communication module 316. Conversely, a control signal or command provided under the control of the processor 334 of the information processing system 230 may be received by the user terminal 210 via the communication module 316 of the user terminal 210 via the communication module 336 and the network 220. For example, the user terminal 210 may receive data related to street view images of a specific area from the information processing system 230.

[0045] The input / output interface 318 may be a means for interfacing with the input / output device 320. For example, the input device may include a camera including an audio sensor and / or an image sensor, a keyboard, a microphone, a mouse, etc., and the output device may include a display, a speaker, a haptic feedback device, etc. As another example, the input / output interface 318 may be a means for interfacing with a device that integrates input and output configurations or functions, such as a touchscreen. For example, when the processor 314 of the user terminal 210 processes instructions of a computer program loaded in the memory 312, a service screen configured using information and / or data provided by the information processing system 230 or another user terminal may be displayed on the display through the input / output interface 318. Although FIG. 3 illustrates the input / output device 320 as not being included in the user terminal 210, this is not limiting and the input / output device 320 may be configured as a single device together with the user terminal 210. Furthermore, the input / output interface 338 of the information processing system 230 may be a means for interfacing with an input or output device (not shown) that may be connected to or included in the information processing system 230. While the input / output interfaces 318 and 338 are shown in FIG. 3 as elements configured separately from the processors 314 and 334, the present invention is not limited thereto, and the input / output interfaces 318 and 338 may be configured to be included in the processors 314 and 334.

[0046] The user terminal 210 and the information processing system 230 may include more components than those shown in FIG. 3 . However, it is not necessary to explicitly show most of the conventional components. According to one embodiment, the user terminal 210 may be embodied to include at least some of the input / output devices 320 described above. The user terminal 210 may also include other components such as a transceiver, a global positioning system (GPS) module, a camera, various sensors, a database, etc. For example, if the user terminal 210 is a smartphone, it may include components typically included in a smartphone, such as an acceleration sensor, a gyro sensor, an image sensor, a proximity sensor, a touch sensor, an illuminance sensor, a camera module, various physical buttons, buttons using a touch panel, input / output ports, and a vibrator for vibration. According to one embodiment, the processor 314 of the user terminal 210 may be configured to run an application that provides a map information service. In this case, code related to the application and / or program may be loaded into the memory 312 of the user terminal 210.

[0047] While a program for an application that provides a map information service is running, the processor 314 may receive text, images, videos, sounds, and / or actions entered or selected through an input device, such as a touch screen, keyboard, camera including an audio sensor and / or image sensor, or microphone, connected to the input / output interface 318, and may store the received text, images, videos, sounds, and / or actions in the memory 312 or provide them to the information processing system 230 via the communication module 316 and the network 220. For example, the processor 314 may receive a user input requesting a street view image for a specific area and provide it to the information processing system 230 via the communication module 316 and the network 220.

[0048] The processor 314 of the user terminal 210 may be configured to manage, process, and / or store information and / or data received from the input / output device 320, other user terminals, the information processing system 230, and / or multiple external systems. The information and / or data processed by the processor 314 may be provided to the information processing system 230 via the communication module 316 and the network 220. The processor 314 of the user terminal 210 may transmit and output information and / or data to the input / output device 320 via the input / output interface 318. For example, the processor 314 may display the received information and / or data on a screen of the user terminal.

[0049] The processor 334 of the information processing system 230 may be configured to manage, process, and / or store information and / or data received from multiple user terminals 210 and / or multiple external systems. The information and / or data processed by the processor 334 may be provided to the user terminal 210 via the communication module 336 and the network 220.

[0050] 4 illustrates an example process for generating a visual feature map 450 associated with street view imagery 420 based on a 3D model 410 and street view imagery 420, according to one embodiment of the present disclosure. Here, the 3D model 410 may represent a 3D model of a specific area including 3D geometric information expressed in absolute coordinate positions. For example, the 3D model 410 may include 3D building models and 3D road models of buildings within the specific area.

[0051] In one embodiment, the street view image 420 may be a 360° panoramic image generated using equirectangular projection and may be street view image captured at a specific node in a specific area. The street view image 420 may include absolute coordinate position information and direction information matched with absolute coordinate position information of the 3D model. For example, the absolute coordinate position information and direction information included in the street view image 420 may be information matched with the absolute coordinate position information of the 3D model 410 using predefined map matching points or map matching lines.

[0052] According to one embodiment, the information processing system may render a depth map 430 associated with the street view image 420 using the 3D model 410. For example, the information processing system may project 3D geometric information included in the 3D model 410 onto the street view image 420 based on absolute coordinate position information and direction information of the street view image 420, thereby rendering the depth map 430 associated with the street view image 420.

[0053] According to one embodiment, the information processing system may extract the plurality of feature points 440 from the Street View image 420. For example, the information processing system may convert the Street View image into a plurality of planar images using a perspective projection method, and then extract the plurality of feature points 440 from the plurality of planar images. In another example, the information processing system may extract the plurality of feature points 440 using a binary mask generated based on the Street View image 420.

[0054] According to one embodiment, the information processing system may generate a visual feature map 450 associated with the street view image 420 based on the depth map 430 and a plurality of feature points 440 in the street view image. For example, the information processing system may determine absolute coordinate position information of the plurality of feature points 440 based on the depth map 430. The information processing system may then generate the visual feature map 450 associated with the street view image 420 by associating and storing the absolute coordinate position information and a visual feature descriptor for each of the plurality of feature points 440.

[0055] Thus, according to one embodiment of the present disclosure, instead of using a vehicle equipped with a vehicle-based multi-sensor surveying system, the cost and effort of obtaining a visual feature map can be reduced by using 3D building model information generated by aerial image surveying, etc., and Street View images already captured for the Street View service.

[0056] 5 is a block diagram illustrating a specific method for generating a visual feature map 590 associated with a first street view image 512 based on aligned street view data 510 and a 3D model 520 according to one embodiment of the present disclosure. Here, the aligned street view data 510 may be a plurality of street view images captured at multiple nodes within a specific area, and may be a panoramic image generated using a square projection. In one embodiment, the aligned street view data 510 may include highly accurate (high-precision) absolute coordinate position information and direction information aligned with the absolute coordinate position information of the 3D model 410 using predefined map matching points or map matching lines. In another embodiment, the aligned street view data 510 may include absolute coordinate position information and direction information obtained from high-precision location estimation using GPS / INS (Global Positioning System / Inertial Navigation System) sensor fusion.

[0057] The 3D model 520 is a model of a specific area generated based on aerial photography, and may include a 3D road model, a 3D building model, etc. The 3D model 520 may also include 3D geometric information expressed in absolute coordinate positions. For example, the 3D model 520 may be configured as a triangle mesh consisting of vertices and edges, but is not limited thereto. The 3D model 520 may be a point cloud reconstructed in 3D by aerial photography or a digital elevation model (DSM) determined by aerial surveying. Alternatively, the 3D model 520 may be configured as a combination of three types of data structures, such as a triangle mesh, a point cloud, or a digital elevation model.

[0058] According to one embodiment, an information processing system may receive a plurality of 3D building and road models 522 within a specific region included in a 3D model 520 for the specific region. The information processing system may also receive a first street view image 512 captured at a first node within the specific region included in the aligned street view data 510. The information processing system may then project at least a portion of the 3D geometric information included in the 3D building and road models 522 onto the first street view image 512 based on absolute coordinate position information and direction information of the first street view image 512, thereby rendering / generating a depth map 540 associated with the first street view image 512. Here, the depth map 540 may include depth information of the buildings and roads.

[0059] According to one embodiment, the information processing system may extract a first set of feature points 530 from the first street view image 512. Specifically, the information processing system may detect the first set of feature points 530 from the first street view image 512 and extract visual feature descriptors 580 for each of the first set of feature points. To address errors in feature detection or feature matching due to geometric distortion of the panoramic street view image, the information processing system may extract the first set of feature points 530 from multiple planar images generated based on the first street view image 512. For example, the information processing system may horizontally rotate the first street view image 512 at a height that views the horizon, project the image as 12 planar images, detect feature points from each planar image, and extract visual feature descriptors. The first set of feature points 530 may be extracted using, but is not limited to, a Scale-Invariant Feature Transform (SIFT), SuperPoint, or Repeatable and Reliable Detector and Descriptor (R2D2) technique.

[0060] According to one embodiment, the information processing system may determine absolute coordinate position information 550 for each of the first set of feature points based on the depth map 540. The information processing system may then generate a visual feature map 590 associated with the first street view image 512 by associating and storing the absolute coordinate position information 550 for each of the first set of feature points and the visual feature descriptors 580 for each of the first set of feature points.

[0061] 6 is a block diagram illustrating a specific method for generating a visual feature map 690 associated with a first street view image 612 based on aligned street view data 610, a 3D model 620, a binary mask 616, and 3D planar information about road traffic structures 660, according to another embodiment of the present disclosure. In FIG. 6, configurations that overlap with those in FIG. 5 will be briefly described or omitted based on the embodiment shown in FIG. 6.

[0062] According to one embodiment, the information processing system may extract a first set of feature points 630 associated with buildings and roads using a binary mask 616 based on the first street view image 612. Specifically, the information processing system may perform semantic segmentation based on the first street view image 612 to generate a binary mask 616 indicating road regions and building regions included in the first street view image. In another example, the binary mask 616 may indicate road regions, building regions, and road traffic structure regions included in the first street view image.

[0063] In one embodiment, the information processing system can convert the first street view image 612 into a plurality of undistorted planar images. For example, the information processing system can convert the first street view image 612 into six cube images, including top, bottom, left, right, front, and back, using cube mapping. The information processing system can then perform semantic segmentation on the plurality of undistorted planar images to detect road areas, building areas, vehicles, lanes, road traffic structures (e.g., traffic lights, signboards, etc.), etc. The information processing system can then combine the semantic segmentation results for the plurality of undistorted planar images into a panoramic image again to generate a binary mask 616 that indicates the road areas and building areas in the first street view image 612.

[0064] In one embodiment, the information processing system may use the binary mask 616 to detect a first set of feature points 630 from the first street view image 612 and extract visual feature descriptors 632 for each of the first set of feature points. For example, the information processing system may use the binary mask 616 to extract the first set of feature points 630 from a portion of the first street view image 612. In another example, the information processing system may extract a plurality of feature points from the first street view image 612 and then filter the extracted plurality of feature points using the binary mask 616 to extract the first set of feature points 630. Here, the first set of feature points may be feature points extracted from road regions and building regions of the first street view image 612, or feature points extracted from road regions, building regions, and road traffic structures.

[0065] According to one embodiment, the information processing system may determine absolute coordinate position information 670 of feature points associated with road traffic structures using a plurality of street view images 612, 614. Here, the road traffic structures may refer to structures and / or facilities associated with road travel, such as signboards, traffic lights, medians, etc. Specifically, the information processing system may generate 3D planar information 660 for road traffic structures commonly included in the first street view image 612 and the second street view image 614, based on a first street view image 612 captured at a first node and a second street view image 614 captured at a second node within a specific area.

[0066] The information processing system can then determine absolute coordinate position information 670 of the feature points associated with the road traffic structures by projecting the feature points associated with the road traffic structures from the first set of feature points 630 onto a three-dimensional plane. With this configuration, the information processing system can obtain absolute coordinate position information of the road traffic structures included in both the first street view image 612 and the second street view image 614, even if the 3D model 620 does not include position information for the road traffic structures.

[0067] According to one embodiment, the information processing system may determine absolute coordinate position information 650 of feature points associated with buildings and roads based on the depth map 640. Here, the depth map 640 may represent the depth map 640 generated by the method described with reference to Fig. 5. For example, the information processing system may determine absolute coordinate position information 650 of feature points associated with buildings and roads among the first set of feature points 630 using the depth map 640 including depth information of buildings and roads.

[0068] According to one embodiment, the information processing system can generate a visual feature map 690 associated with the first street view image 612 by correlating and storing absolute coordinate position information 670 of feature points associated with road traffic structures among the first set of feature points 630, absolute coordinate position information 650 of feature points associated with buildings and roads among the first set of feature points 630, and respective visual feature descriptors 632 of the first set of feature points.

[0069] 6 illustrates, but is not limited to, an example in which absolute coordinate position information 650 of feature points associated with buildings and roads is determined using a depth map 640 including depth information of buildings and roads generated based on a 3D building and road model 622. For example, the 3D building and road model 622 may further include some or all of road traffic structures associated with road driving, in which case the depth map 640 may include depth information for some or all of the road traffic structures in addition to the roads and buildings. In this case, the absolute coordinate position information 650 of feature points associated with buildings and roads may also include absolute coordinate position information 650 of feature points associated with some or all of the road traffic structures in addition to the roads and buildings.

[0070] 6 illustrates a method for generating a visual feature map 690 for buildings, roads, and road traffic structures by using the binary mask 616 and the three-dimensional planar information 660 for road traffic structures, but some of these may be omitted. For example, the visual feature map 690 for buildings, roads, and road traffic structures may be generated without using the binary mask 616. In another example, the visual feature map 690 for buildings and roads may be generated without generating the three-dimensional planar information 660 for road traffic structures.

[0071] 7 illustrates an example process for generating a binary mask based on street view imagery according to an embodiment of the present disclosure. The information processing system may generate a binary mask based on street view imagery through a first state 710, a second state 720, and a third state 730. The first state 710 illustrates an example of street view imagery captured at a specific node in a specific region. Here, the street view imagery may be a 360° panoramic imagery generated using a square projection.

[0072] The second state 720 illustrates an example of a semantic segmentation result based on the Street View image. In one embodiment, the information processing system may convert the Street View image into multiple undistorted planar images and then perform semantic segmentation. For example, the information processing system may use a perspective projection method to convert the Street View image into six cube images and then perform semantic segmentation. In another example, the information processing system may perform semantic segmentation on the Street View image in one go. Semantic segmentation may be performed using techniques such as DeepLab v3 and Mask-RCNN, but is not limited to these.

[0073] The third state 730 is an example of a binary mask obtained as a result of performing semantic segmentation based on the Street View image. In one embodiment, the binary mask may indicate road and building regions within the Street View image.

[0074] According to one embodiment, the information processing system may extract a first set of feature points from the Street View image using a binary mask. In one example, the information processing system may extract the first set of feature points from a partial region of the Street View image using the binary mask. Here, the partial region of the Street View image may be a region corresponding to the binary mask. In another example, the information processing system may extract the first set of feature points in the region corresponding to the binary mask by filtering a plurality of feature points extracted from the entire region of the Street View image using the binary mask.

[0075] 7, the binary mask represents road and building regions in the Street View image, but is not limited to this. For example, the binary mask may represent road regions, building regions, and road traffic structure regions in the Street View image.

[0076] 8 illustrates an example process for performing stereo matching on road traffic structures in a first street view image 812 and a second street view image 814, according to an embodiment of the present disclosure. In one embodiment, when a visual feature map associated with a street view is generated using a depth map including depth information for buildings and roads generated based on a 3D building and road model, absolute coordinate position information for road traffic structures associated with the driving environment may be missing from the street view. To address this, the information processing system may generate 3D planar information for the road traffic structures based on the multiple street view images 812 and 814, and acquire absolute coordinate position information associated with the road traffic structures.

[0077] According to one embodiment, in order to obtain absolute coordinate position information associated with a road traffic structure, the information processing system may generate 3D planar information for a specific road traffic structure included in each of the plurality of street view images 812, 814 based on the plurality of street view images 812, 814. For example, the information processing system may generate 3D planar information for a specific road traffic structure included in the first street view image 812 and the second street view image 814 by passing through a first state 810 and a second state 820. Here, the first street view image 812 and the second street view image 814 may be images captured at adjacent nodes.

[0078] First state 810 illustrates an example of the result of detecting areas including road traffic structures in first street view image 812 and second street view image 814 and determining a connection relationship between the same road traffic structures. In one embodiment, the information processing system may detect a first area including a first road traffic structure (e.g., a traffic light) in first street view image 812. Similarly, the information processing system may detect a second area including a second road traffic structure (e.g., the same traffic light) in second street view image 814. For example, as shown, the information processing system may detect a four-color traffic light in first street view image 812 as a first area and a four-color traffic light in second street view image 814 as a second area. The information processing system may then determine that the first area and the second area include the same road traffic structure and determine a connection relationship between them. For example, the information processing system may determine that the first road traffic structure and the second road traffic structure are the same road traffic structure (e.g., the same traffic light) based on the visual similarity between the first region and the second region, where the visual similarity between the first region and the second region may be determined using at least one of hue similarity, visual feature descriptor similarity, or a deep learning-based matching model.

[0079] A second state 820 illustrates exemplary results 826 and 828 of stereo matching a first region and a second region to generate 3D planar information about a specific road traffic structure. Road traffic structures have similar structures, forms, and shapes, and it is necessary to identify road traffic structures that match with each other. To do this, the information processing system can divide the first region in the first street view image 812 into first region images 822. Similarly, the information processing system can divide the second region in the second street view image 814 into second region images 824. Then, dense stereo matching and triangulation can be performed on the first region image 822 and the second region image 824 to obtain depth information about the matched road traffic structure. As a result, the information processing system can generate 3D planar information about the same road traffic structure (e.g., a four-color traffic light) that is commonly included in the first street view image 812 and the second street view image 814.

[0080] FIG. 9 illustrates an example of a 3D planar information map 900 relating to a plurality of road traffic structures according to an embodiment of the present disclosure. Here, the 3D planar information map 900 relating to a plurality of road traffic structures may represent visualized information obtained by merging 3D planar information relating to each of a plurality of road traffic structures acquired from a plurality of street view images captured at different nodes. As illustrated in FIG. 9 , road traffic structures may include, but are not limited to, road traffic signs, safety signs, and traffic lights. For example, road traffic structures are structures or facilities that may affect the road driving environment and may further include medians, curbs, roadside trees, bus stops, etc. Such road traffic structures are objects that do not change visually much over time and are suitable for use in camera-based localization after a visual feature map is generated.

[0081] According to one embodiment, the 3D plane information about the specific road traffic structure may be used to determine absolute coordinate position information of a feature point associated with the specific road traffic structure among a plurality of feature points in the Street View image. In one embodiment, the information processing system may determine absolute coordinate position information of the feature point associated with the specific road traffic structure by projecting the feature point associated with the specific road traffic structure among a plurality of feature points in the Street View image onto a 3D plane.

[0082] 10 illustrates an example process for extracting a first set of feature points from Street View imagery using a plurality of planar images generated based on the Street View imagery, according to an embodiment of the present disclosure. Here, the Street View imagery may be a 360° panoramic image generated using a square projection and captured at a specific node in a specific region. To address errors in feature detection or feature matching due to geometric distortion of the panoramic Street View imagery, the information processing system may extract the first set of feature points from the Street View imagery via a first state 1010 and a second state 1020.

[0083] The first state 1010 illustrates an example of multiple planar images generated based on the Street View image. According to one embodiment, the information processing system may convert the first Street View image into multiple planar images using a perspective projection method. The information processing system may then extract multiple feature points from each of the multiple planar images. For example, as shown, the information processing system may horizontally rotate the Street View image at a height that views the horizon, project the images into 12 planar images, detect feature points from each planar image, and extract visual feature descriptors. In the first state 1010, one red dot may represent one feature point.

[0084] The second state 1020 illustrates an example of the result of projecting feature points extracted from the plurality of planar images onto the Street View imagery. According to one embodiment, the information processing system may acquire the first set of feature points by projecting coordinate information associated with the plurality of feature points in each of the plurality of planar images onto the Street View imagery. For example, the information processing system may use an inverse perspective projection method to project the coordinate information associated with the plurality of feature points extracted from each of the plurality of planar images in the first state 1010 onto the Street View imagery to acquire the first set of feature points in the Street View imagery. In the second state 1020, one red dot may represent one feature point.

[0085] 11 illustrates an example of a process for acquiring absolute coordinate position information related to a road traffic structure according to an embodiment of the present disclosure. According to an embodiment, an information processing system may determine absolute coordinate position information of a feature point 1114 associated with a road traffic structure using three-dimensional planar information 1116 related to a specific road traffic structure acquired by the method described with reference to FIG. 8. Feature points displayed as green dots in the street view image 1112 may represent feature points associated with buildings and roads, and feature points displayed as red dots may represent feature points associated with road traffic structures.

[0086] In one embodiment, the information processing system can project feature points 1114 associated with road traffic structures from among a plurality of feature points extracted from street view imagery 1112 onto a three-dimensional plane 1116 related to the road traffic structures. The information processing system can then determine three-dimensional coordinate information of feature points 1118 projected onto the three-dimensional plane 1116 based on the information on the three-dimensional plane 1116.

[0087] FIG. 12 illustrates an example of a visual feature map 1200 for a specific region according to an embodiment of the present disclosure. According to an embodiment, an information processing system may generate the visual feature map 1200 for the specific region by merging visual feature maps associated with multiple street view images captured at different nodes in the specific region. For example, the information processing system may generate a first visual feature map associated with a first street view image based on a first street view image captured at a first node in the specific region. Similarly, the information processing system may generate a second visual feature map associated with a second street view image based on a second street view image captured at a second node in the specific region. The information processing system may then generate the visual feature map 1200 for the specific region by merging the first visual feature map and the second visual feature map based on absolute coordinate position information and direction information of the first street view image and absolute coordinate position information and direction information of the second street view image. Here, the absolute coordinate position information and direction information of the street view image and the absolute coordinate position information and direction information of the second street view image may be information that is consistent with the absolute coordinate position information of the 3D model.

[0088] In the visual feature map 1200, feature points displayed as green dots may represent feature points related to buildings and roads, and feature points displayed as red dots may represent feature points related to road traffic structures. With this configuration, instead of using a vehicle equipped with a vehicle-based multi-sensor surveying system, an entire visual feature map for a specific area can be obtained using 3D building model information generated by aerial image surveying or the like and Street View images already captured for the Street View service, thereby reducing costs and labor.

[0089] 13 is a flowchart illustrating an example of a method 1300 for generating a visual feature map using a 3D model and street view imagery according to an embodiment of the present disclosure. The method 1300 may begin with a processor (e.g., at least one processor of an information processing system) receiving a 3D model of a specific area including 3D geometric information expressed in absolute coordinate positions (S1310). Here, the 3D model may include multiple 3D building models and road models within the specific area.

[0090] The information processing system may then receive a first street view image captured by a first node within a specific area (S1320). Here, the first street view image may be a panoramic image generated using equirectangular projection. Furthermore, absolute coordinate position information and direction information of the first street view image may be position information matched with absolute coordinate position information of the 3D model. For example, the absolute coordinate position information and direction information of the first street view image may be matched with absolute coordinate position information of the 3D model using predefined map matching points or map matching lines. Here, the map matching points may include ground control points (GCPs) and building control points (BCPs). Each ground control point may be a corresponding pair of a point on the ground and three-dimensional absolute coordinate position information. Each building control point may be a corresponding pair of a point on a building and three-dimensional absolute coordinate position information. The map matching lines include ground control lines (GCLs), and each ground control line may be a corresponding pair of a line on the ground and at least one piece of three-dimensional absolute coordinate position information.

[0091] In one embodiment, the information processing system may project at least a portion of the 3D geometric information included in the 3D model onto the first street view image based on the absolute coordinate position information and direction information of the first street view image, and render a depth map associated with the first street view image (S1330). Here, the depth map may include depth information of buildings and roads.

[0092] In one embodiment, the information processing system may extract a first set of feature points from the first street view image (S1340). For example, the information processing system may use a perspective projection method to convert the first street view image into a plurality of planar images, extract a plurality of feature points from each of the plurality of planar images, and project coordinate information associated with the plurality of feature points in each of the plurality of planar images onto the first street view image to obtain the first set of feature points.

[0093] In one embodiment, the information processing system may perform semantic segmentation on the first Street view image to generate a binary mask indicating road areas and building areas included in the first Street view image, and extract a first set of feature points from the first Street view image using the binary mask. In this case, generating the binary mask may include converting the first Street view image into a plurality of undistorted planar images and performing semantic segmentation on the plurality of undistorted planar images to detect road areas and building areas. Here, the plurality of undistorted planar images may be generated by converting the first Street view image into six cube images using a perspective projection method. In addition, extracting the first set of feature points from the first Street view image using the binary mask may include extracting the first set of feature points from a portion of the first Street view image using the binary mask. Alternatively, extracting the first set of feature points from the first Street View image using the binary mask may include filtering the plurality of feature points extracted from the first Street View image using the binary mask to extract the first set of feature points.

[0094] In one embodiment, the information processing system may determine absolute coordinate position information of all or some of the first set of feature points based on the depth map (S1350). Specifically, absolute coordinate position information of feature points associated with buildings and roads among the first set of feature points may be determined based on the depth map.

[0095] In one embodiment, the information processing system may receive a second street view image captured at a second node within a specific region. In this case, the information processing system may generate 3D planar information about a specific road traffic structure included in the first street view image and the second street view image based on the first street view image and the second street view image. Here, generating the 3D planar information may include detecting a first region including the first road traffic structure within the first street view image, detecting a second region including the second road traffic structure within the second street view image, determining that the first road traffic structure and the second road traffic structure are the same specific road traffic structure based on a visual similarity between the first region and the second region, and performing stereo matching and triangulation on the first region and the second region to generate the 3D planar information about the specific road traffic structure. Here, the visual similarity between the first region and the second region may be determined using at least one of hue similarity, visual feature descriptor similarity, or a deep learning-based matching model. The information processing system can then determine absolute coordinate position information of the feature points associated with the specific road traffic structure among the first set of feature points based on the three-dimensional plane information. For example, determining the absolute coordinate position information of the feature points associated with the specific road traffic structure among the first set of feature points may include projecting the feature points associated with the specific road traffic structure among the first set of feature points onto the three-dimensional plane.

[0096] The information processing system can then generate a first visual feature map associated with the first street view image by associating and storing absolute coordinate position information and a visual feature descriptor for each of the first set of feature points (S1360).

[0097] In one embodiment, the information processing system may receive second street view images captured at a second node within a specific region and generate a second visual feature map associated with the second street view images. The information processing system may then merge the first visual feature map and the second visual feature map based on absolute coordinate position information and direction information of the first street view images and absolute coordinate position information and direction information of the second street view images. Here, the absolute coordinate position information and direction information of the first street view images and the absolute coordinate position information and direction information of the second street view images may be position information that is consistent with the absolute coordinate position information of the 3D model.

[0098] 13 and the above description are merely examples, and the scope of the present disclosure is not limited thereto. For example, at least one step may be added / modified / deleted, or the order of the steps may be changed.

[0099] The above-described method may be provided as a computer program stored on a computer-readable recording medium for execution by a computer. The medium may be a medium that permanently stores a computer-executable program or a medium that temporarily stores the program for execution or download. The medium may be various recording or storage means in the form of a single or multiple pieces of hardware, and is not limited to media directly connected to a computer system but may also be distributed over a network. Examples of media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and media configured to store program instructions, including ROM, RAM, and flash memory. Other examples of media include recording or storage media managed by app stores that distribute applications or by websites or servers that provide or distribute various software.

[0100] The methods, operations, or techniques of the present disclosure may be embodied in various means. For example, such techniques may be embodied as hardware, firmware, software, or a combination thereof. Those skilled in the art will understand that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the present disclosure may be embodied as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability between hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is embodied as hardware or software depends on the particular application and design requirements imposed on the overall system. Those skilled in the art may also implement the described functionality in various ways for each particular application, but such implementations should not be interpreted as departing from the scope of the present disclosure.

[0101] In a hardware implementation, the processing units used in performing the techniques may be embodied within one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, electronic devices, other electronic units designed to have the functionality described in this disclosure, computers, or combinations thereof.

[0102] Accordingly, the various illustrative logic blocks, modules, and circuits described in connection with this disclosure may be embodied or implemented as a general-purpose processor, a DSP, an ASIC, an FPGA or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination designed to have the functionality described herein. A general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be embodied as a combination of computing devices, such as a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other configuration.

[0103] In a firmware and / or software implementation, the techniques may be embodied as instructions stored on a computer-readable medium such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, compact disc (CD), magnetic or optical data storage device, etc. The instructions may be executable by one or more processors and may cause the processors to perform certain aspects of the functions described in this disclosure.

[0104] While the embodiments described above utilize aspects of the presently disclosed subject matter on one or more stand-alone computer systems, the present disclosure is not limited thereto and may be implemented in connection with any computing environment, such as a network or distributed computing environment. Furthermore, aspects of the subject matter in this disclosure may be implemented on multiple processing chips or devices, and storage may be similarly affected across multiple devices. Such devices may include PCs, network servers, and handheld devices.

[0105] Although the present disclosure has been described herein with reference to some embodiments, various modifications and variations that are understandable to those skilled in the art to which the present disclosure pertains can be made without departing from the scope of the present disclosure, and such modifications and variations should be considered to fall within the scope of the claims appended hereto.

Claims

1. 1. A method for generating a visual feature map using a 3D model and street view imagery, performed by at least one processor, comprising: receiving a three-dimensional model of a particular region including three-dimensional geometric information expressed by absolute coordinate positions; receiving a first street view image captured at a first node within the specific area; projecting at least a portion of 3D geometric information included in the 3D model onto the first street view image based on absolute coordinate position information and direction information of the first street view image, and rendering a depth map associated with the first street view image; extracting a first set of feature points from the first Street View image; determining absolute coordinate position information of all or a portion of the first set of feature points based on the depth map; and generating a first visual feature map associated with the first street view image by associating and storing absolute coordinate position information and a visual feature descriptor for each of the first set of feature points.

2. the first street view image is a panoramic image generated using square projection; The visual feature map generating method of claim 1 , wherein absolute coordinate position information and direction information of the first street view image are aligned with absolute coordinate position information of the three-dimensional model.

3. The step of extracting the first set of feature points comprises: generating a binary mask representing road areas and building areas included in the first street view image by performing semantic segmentation based on the first street view image; and extracting the first set of feature points from the first street view image using the binary mask.

4. generating the binary mask comprises: converting the first street view image into a plurality of undistorted planar images; The method of claim 3 , further comprising: detecting road regions and building regions by performing semantic segmentation on the plurality of undistorted planar images.

5. The method of claim 4 , wherein the plurality of undistorted planar images are generated by converting the first street view image into six cube images using a perspective projection method.

6. the three-dimensional model includes a plurality of three-dimensional building models and road models within the specific area; the depth map includes depth information of buildings and roads; The method of claim 1 , further comprising determining absolute coordinate position information of feature points associated with buildings and roads among the first set of feature points based on the depth map.

7. extracting the first set of feature points from the first Street View image using the binary mask includes: The method of claim 3 , further comprising: extracting the first set of feature points from a portion of the first street view image using the binary mask.

8. extracting the first set of feature points from the first Street View image using the binary mask includes:

4. The method of claim 3, further comprising: extracting the first set of feature points by filtering the plurality of feature points extracted from the first street view image using the binary mask.

9. receiving a second street view image captured at a second node within the specific area; generating three-dimensional planar information regarding a specific road traffic structure included in the first street view image and the second street view image based on the first street view image and the second street view image; 2. The visual feature map generating method of claim 1, further comprising: determining absolute coordinate position information of feature points associated with the specific road traffic structure among the first set of feature points based on the three-dimensional plane information.

10. The step of determining absolute coordinate position information of a feature point associated with the specific road traffic structure includes: The method of claim 9 , further comprising projecting the feature points associated with the specific road traffic structure from the first set of feature points onto a three-dimensional plane.

11. The step of generating three-dimensional plane information includes: detecting a first area including a first road traffic structure within the first street view image; detecting a second area including a second road traffic structure within the second street view image; determining that the first road traffic structure and the second road traffic structure are the same road traffic structure, based on the visual similarity between the first area and the second area; and generating three-dimensional planar information about the specific road traffic structure by performing stereo matching and triangulation on the first and second regions.

12. 12. The method of claim 11, wherein the visual similarity between the first region and the second region is determined using at least one of hue similarity, visual feature descriptor similarity, or a deep learning-based matching model.

13. The step of extracting the first set of feature points comprises: converting the first street view image into a plurality of planar images using a perspective projection method; extracting a plurality of feature points from each of the plurality of planar images; and acquiring the first set of feature points by projecting coordinate information associated with a plurality of feature points in each of the plurality of planar images onto the first street view image.

14. 2. The visual feature map generating method of claim 1, wherein the absolute coordinate position information and direction information of the first street view image are matched with the absolute coordinate position information of the 3D model using predefined map matching points or map matching lines.

15. the map matching points include ground control points and building control points; Each ground control point is a pair of a point on the ground and its corresponding 3D absolute coordinate position information. The visual feature map generating method of claim 14, wherein each building control point is a corresponding pair of a point on the building and three-dimensional absolute coordinate position information.

16. the map matching lines include ground control lines; The visual feature map generating method of claim 14, wherein each ground control line is a corresponding pair of a line on the ground and at least one piece of three-dimensional absolute coordinate position information.

17. receiving a second street view image captured at a second node within the specific area; generating a second visual feature map associated with the second street view image; merging the first visual feature map and the second visual feature map based on absolute coordinate position information and direction information of the first street view image and absolute coordinate position information and direction information of the second street view image, 2. The visual feature map generation method of claim 1, wherein the absolute coordinate position information and direction information of the first street view image and the absolute coordinate position information and direction information of the second street view image are aligned with the absolute coordinate position information of the three-dimensional model.

18. A computer-readable non-transitory recording medium having recorded thereon instructions for causing a computer to execute the visual feature map generating method of claim 1.

19. An information processing system, a communication module; Memory and at least one processor coupled to the memory and configured to execute at least one computer-readable program contained in the memory; The at least one program receiving a three-dimensional model of a specific region including three-dimensional geometric information expressed by absolute coordinate positions; receiving a first street view image captured at a first node within the specific area; projecting at least a portion of 3D geometric information included in the 3D model onto the first street view image based on absolute coordinate position information and direction information of the first street view image, and rendering a depth map associated with the first street view image; extracting a first set of feature points from the first Street View image; determining absolute coordinate position information of all or a portion of the first set of feature points based on the depth map; an information processing system including instructions for generating a first visual feature map associated with the first street view image by associating and storing absolute coordinate position information and a visual feature descriptor for each of the first set of feature points.

Citation Information

Patent Citations

  • Image-based positioning method and system

    JP2022077976A

  • Method and system for generating depth information of street view image using 2d map

    KR102234461B1

  • Method and system for generating visual feature map

    KR102383499B1