Depth information-based visual localization method and apparatus for performing same

WO2026205990A1PCT designated stage Publication Date: 2026-10-01INDUSTRY UNIVERSITY COOPERATION FOUNDATION HANYANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2026/004770
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-26
Filing Date
2026-03-25
Publication Date
2026-10-01

Smart Images

  • Figure KR2026004770_01102026_PF_FP_ABST
    Figure KR2026004770_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed, according to one embodiment, is a visual localization method comprising: generating a fixed point from three-dimensional point cloud data; generating three-dimensional line cloud data by connecting a plurality of three-dimensional points included in the three-dimensional point cloud data to the generated fixed point; receiving an input image captured via a camera and depth information captured via a depth camera; and extracting two-dimensional feature point information included in the input image, and estimating a pose of the camera by which the input image was captured on the basis of the extracted two-dimensional feature point information and the depth information matched to two-dimensional feature points.
Need to check novelty before this filing date? Find Prior Art

Description

Depth information-based visual localization method and device for executing the same

[0001] The present invention relates to a visual localization method capable of real-time camera pose estimation based on input images and depth information while preventing personal information leakage, and an apparatus for executing the same.

[0002] With the increasing usage and demand for Augmented Reality (AR), Virtual Reality (VR), Mixed Reality (MR), autonomous vehicles, and autonomous guidance robots that emerged with the Fourth Industrial Revolution, there is a growing need for more precise user location estimation technology. Existing Global Navigation Satellite Systems (GNSS), such as GPS, have limitations that make them unsuitable for application in industrial environments, as they are difficult to apply indoors and possess margins of error. Visual localization technology is garnering attention as a potential replacement.

[0003] Visual localization technology is a technique that identifies the precise location and orientation of a user device on a spatial map based on images captured by a camera. With the widespread adoption of camera sensors, it can achieve high location identification accuracy at a relatively low cost. Products equipped with visual localization technology transmit query images to a cloud server, compare them with a stored 3D point cloud spatial map, and estimate the user's camera position and orientation. However, since feature information is stored within the 3D point cloud map on the cloud server, if the spatial map is leaked, the point cloud data can be processed through deep learning models (e.g., Inverse Structure-from-Motion, InvSfM) to synthesize a realistic image (reverse reconstruction). Consequently, if the 3D map stored on the cloud server is leaked, the user's sensitive personal information may be compromised.

[0004] To address this problem, a technique for estimating the position and attitude of a user device was proposed (hereinafter referred to as Prior Art 1) by using a geometrically hidden 3D spatial map, which is a 3D line passing through a point in a random direction, instead of using a 3D point cloud spatial map. However, this technique also raised concerns about privacy infringement as it was found that it is possible to reconstruct a point cloud from random line cloud data (hereinafter referred to as Prior Art 2).

[0005] In Prior Art 3 published in 2024, a method for reconstructing point cloud data was proposed using an optimization method that utilizes the distance between neighboring lines within the line cloud data. This method demonstrated that all previously proposed geometric point cloud hiding methods could be reverse-reconstructed, pointing out the limitations of the hiding performance of existing methods.

[0006] [Previous Paper 1] P. Speciale et al., “Privacy preserving image-based localization”, CVPR 2019.

[0007] [Previous Paper 2] K. Chelani et al., "How Privacy-Preserving Are Line Clouds? Recovering Scene Details From 3D Lines", CVPR, 2021.

[0008] [Previous Paper 3] K. Chelani and A. Benbihi et al., Obfuscation Based Privacy Preserving Representations are Recoverable Using Neighborhood Information, 3DV 2025.

[0009] One embodiment disclosed to solve these problems is a depth information-based visual localization method and an apparatus for executing the same, which prevents the possibility of reverse reconstruction by generating line group data based on one fixed point and enables real-time visual localization through depth information even for a single image.

[0010] A visual localization method according to a disclosed embodiment comprises: generating a fixed point from three-dimensional point cloud data; generating three-dimensional line cloud data by connecting a plurality of three-dimensional points included in the three-dimensional point cloud data with the generated fixed point; receiving an input image captured through a camera and depth information captured through a depth camera; extracting two-dimensional feature point information included in the input image, and estimating the attitude of the camera that captured the input image based on the extracted two-dimensional feature point information and depth information matching the two-dimensional feature points.

[0011] Generating the above 3D line group data includes projecting a 3D point onto a unit sphere.

[0012] Generating the above 3D line group data may include: randomly removing 3D points from among 3D points projected onto a unit sphere; selecting a preset number of feature point descriptors of the removed 3D points; and generating a plurality of fake feature points based on the selected feature point descriptors.

[0013] Generating the above 3D line group data may include pairing the generated spurious feature points and unremoved 3D points with the fixed points; and generating the 3D line group data by generating lines connecting the paired points to each other.

[0014] Generating the above three-dimensional line cloud data may include generating a plurality of unit spheres based on a subspace divided from the above three-dimensional point cloud data.

[0015] Generating the fixed point may include setting a three-dimensional spatial region where the fixed point exists from the three-dimensional point cloud data; sampling candidate fixed points; and generating the fixed point based on whether the candidate fixed point exists within the three-dimensional spatial region.

[0016] Setting the above three-dimensional spatial region may include calculating three basis axes and variances for the basis axes based on the principal component analysis of the above three-dimensional point cloud data; and setting the above three-dimensional spatial region based on the calculated variances.

[0017] Setting the above three-dimensional spatial region may include calculating the centroid of a three-dimensional point included in the above three-dimensional point cloud data; and setting the three-dimensional spatial region by setting the distance between the centroid and the three-dimensional point as a radius.

[0018] Sampling the candidate fixed points may include selecting any 3D point among the 3D point cloud data or sampling the candidate fixed points based on 3D coordinate values ​​generated through random number generation.

[0019] Estimating the pose of the camera may include generating three-dimensional feature points based on the two-dimensional feature point information and the depth information; matching the three-dimensional line group data with the three-dimensional feature points; and estimating the pose of the camera based on the intersection points of the matched three-dimensional lines and an absolute pose estimation algorithm.

[0020] A server according to another disclosed embodiment includes: a processor; a program for operating the processor and a memory for storing three-dimensional point cloud data received from the outside; wherein the processor generates a fixed point from the three-dimensional point cloud data and generates three-dimensional line cloud data by connecting a plurality of three-dimensional points included in the three-dimensional point cloud data and the generated fixed point, receives an input image captured through a camera and depth information captured through a depth camera, extracts two-dimensional feature point information included in the input image, and estimates the attitude of the camera that captured the input image based on the extracted two-dimensional feature point information and depth information matching the two-dimensional feature point.

[0021] The processor can generate the 3D line group data by removing a randomly selected 3D point among the plurality of 3D points; selecting a preset number of feature point descriptors of the removed 3D points; generating a plurality of fake feature points by assigning feature point coordinates based on Gaussian noise to the selected feature point descriptors; pairing the generated fake feature points and the unremoved 3D point with the fixed point; and generating a line connecting the paired points.

[0022] The above processor can generate multiple fake feature points by assigning feature point coordinates based on random numbers or specified coordinates received from a user to the selected feature point descriptor.

[0023] The processor can project the generated fake feature points and the unremoved 3D points onto a unit sphere based on the map scale of the 3D point cloud data.

[0024] The processor can generate three-dimensional feature points based on the two-dimensional feature point information and the depth information; match the three-dimensional line group data with the three-dimensional feature points; and estimate the pose of the camera based on the intersection points of the matched three-dimensional lines and an absolute pose estimation algorithm.

[0025] A system according to another disclosed embodiment includes a user terminal comprising a camera and a depth camera; and a server that communicates with the user terminal; wherein the server generates a fixed point from three-dimensional point cloud data and generates three-dimensional line cloud data by connecting a plurality of three-dimensional points included in the three-dimensional point cloud data with the generated fixed point, and the user terminal extracts two-dimensional feature point information included in an input image captured through the camera and can estimate the attitude of the camera that captured the input image based on the extracted two-dimensional feature point information and depth information matching the two-dimensional feature point.

[0026] The user terminal can generate a 3D feature point based on the 2D feature point and the depth information; match the 3D line group data with the 3D feature point; and estimate the pose of the camera based on the intersection point of the matched 3D lines and an absolute pose estimation algorithm.

[0027] The depth information-based visual localization method and the device implementing it disclosed herein can prevent damage caused by the leakage of map data in fields where security for spatial facilities is emphasized, such as defense and industry, and enable privacy protection in preparation for the proliferation of high-definition video devices and spatial restoration technologies.

[0028] In particular, the disclosed embodiment is applicable to products where real-time visual localization for immediate interaction with the surrounding environment is essential, such as in the fields of autonomous driving and robotics, and can maintain high spatial information security by hindering attempts to reconstruct from a 3D line cloud map to 3D point cloud data.

[0029] FIG. 1 is a diagram schematically illustrating a system for implementing the disclosed visual localization method.

[0030] FIG. 2a is a control block diagram of a system according to a disclosed embodiment.

[0031] FIG. 2b is a control block diagram of a system according to another disclosed embodiment.

[0032] Figure 3 is an overall flowchart of the disclosed visual localization method.

[0033] FIG. 4 is a flowchart for specifically explaining the first embodiment of the fixed point generation step of FIG. 3.

[0034] FIG. 5 is a flowchart for specifically explaining a second embodiment of the fixed point generation step of FIG. 3.

[0035] FIG. 6 is a drawing for explaining a first embodiment of setting a three-dimensional spatial region in FIG. 5.

[0036] FIG. 7 is a drawing for explaining a second embodiment of setting a three-dimensional spatial region in FIG. 5.

[0037] FIG. 8 is a flowchart for specifically explaining the step of sampling candidate fixed points of FIG. 5.

[0038] Figure 9 is a diagram illustrating an example of generating line group data based on point group data.

[0039] Figure 10 is a flowchart specifically explaining the steps for generating military data.

[0040] Figure 11 is a diagram for visually explaining the generation steps of the line data of Figure 10.

[0041] FIG. 12 is a flowchart illustrating various embodiments for generating fake feature points.

[0042] Figure 13 is a flowchart illustrating a specific method for performing camera pose estimation.

[0043] Figure 14 is an example illustrating the Songun map.

[0044] Fig. 15 is a diagram for specifically explaining the camera attitude estimation method of Fig. 13.

[0045] Throughout the specification, the same reference numerals refer to the same components. This specification does not describe all elements of the embodiments, and general content in the art to which the invention pertains or content that overlaps between embodiments is omitted.

[0046] Throughout the specification, when a part is described as being "connected" to another part, this includes not only cases where they are directly connected but also cases where they are indirectly connected, and indirect connections include connections made via a wireless communication network.

[0047] Furthermore, when it is stated that a part "includes" a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components.

[0048] Singular expressions include plural expressions unless there is an obvious exception in the context.

[0049] In addition, terms such as "~part," "~unit," "~block," "~part," and "~module" may refer to a unit that processes at least one function or operation. For example, the above terms may refer to at least one piece of hardware such as an FPGA (field-programmable gate array) or ASIC (application specific integrated circuit), at least one piece of software stored in memory, or at least one process processed by a processor.

[0050] The symbols attached to each step are used to identify each step and do not indicate the order of the steps relative to one another; the steps may be performed differently from the specified order unless a specific order is clearly indicated in the context.

[0051] Hereinafter, an embodiment relating to an electronic device and a control method according to one aspect will be described in detail with reference to the attached drawings.

[0052] FIG. 1 is a diagram schematically illustrating a system for implementing the disclosed visual localization method.

[0053] Referring to FIG. 1, a system (1) executing a visual localization method according to one disclosed embodiment may include a user terminal (10), a server (20), and a communication network (2) connecting the user terminal (10) and the server (20).

[0054] A user terminal (10) acquires an image (4, hereinafter referred to as an input image or query image) captured by a camera (RGB Camera, 11-1 in FIG. 2a and FIG. 2b). The user terminal (10) can transmit the input image or feature points detected in the input image to a server (20).

[0055] According to the embodiment of FIG. 2a, the server (20) can perform pose estimation of the camera (11-1) based on the input image (4) transmitted by the user terminal (10). Specifically, the server (20) stores 3D point cloud data in advance and converts the 3D point cloud data into 3D line cloud data. The server (20) can estimate the pose of the camera (11-1) that captured the input image (4) by matching the feature points of the input image transmitted by the user terminal (10), the depth information collected by the depth camera (11-2 in FIG. 2a and FIG. 2b), and the 3D line cloud data. The server (20) transmits the pose estimation result (5) back to the user terminal (10).

[0056] Meanwhile, according to the embodiment of FIG. 2b, the user terminal (10) may estimate the attitude of the camera (11-1) by using the depth information of the depth camera (11-2), the 3D line group data, and the input image after receiving 3D line group data from the server (20).

[0057] A specific method for estimating the pose of the camera (11-1) will be described later through other drawings below.

[0058] The user terminal (10) can be implemented as a computer or portable terminal that can connect to a communication network (2). Here, the computer includes, for example, a laptop, desktop, laptop, tablet PC, slate PC, etc. equipped with a web browser, and the portable terminal is, for example, a wireless communication device that ensures portability and mobility, and may include all kinds of handheld-based wireless communication devices such as PCS (Personal Communication System), GSM (Global System for Mobile communications), PDC (Personal Digital Cellular), PHS (Personal Handyphone System), PDA (Personal Digital Assistant), IMT (International Mobile Telecommunication)-2000, CDMA (Code Division Multiple Access)-2000, W-CDMA (W-Code Division Multiple Access), WiBro (Wireless Broadband Internet) terminal, smartphone, etc., and wearable devices such as glasses, contact lenses, or head-mounted devices (HMD).

[0059] The communication network (2) is a channel for transmitting line data or the attitude estimation result (5) of the camera (11-1) between the user terminal (10) and the server (20). The attitude estimation result (5) is map coordinates including latitude and longitude, and may be expressed in a format such as decimal degrees (DD), degrees, minutes, seconds (DMS), degrees and decimal minutes (DMM), etc.

[0060] The server (20) is configured to store large-capacity map data capable of performing pose estimation of the camera (11-1). The server (20) can be implemented in the form of a cloud server in which an unlimited number of users can request access in a cloud computing environment. The server (20) can be composed of a processor (25, see FIG. 2a and FIG. 2b) that converts large-capacity 3D point cloud data into 3D line cloud data and performs pose estimation of the camera (11-1) through the 3D line cloud data, and a memory (23, see FIG. 2a and FIG. 2b) that stores the aforementioned large-capacity data, and can also be implemented in a form that can be directly connected to the 3D point cloud data through an external processor (20-1), etc.

[0061] FIG. 2a is a control block diagram of a system according to a disclosed embodiment.

[0062] Referring to FIG. 2a, a user terminal (10) according to a disclosed embodiment may include, in hardware, a camera (11-1) for capturing an input image, a depth camera (11-2) for collecting depth information, a communication unit (12) for communicating with a server (20) through a communication network (2), and an output unit (14) for displaying an input image captured by the camera (11-1) or outputting a pose estimation result of the camera (11-1) transmitted by the server (20).

[0063] Specifically, the camera (11-1) may include various shooting means such as a CMOS (Complementary Metal-Oxide Semiconductor) image sensor and a CCD (Charge-Coupled Device) image sensor. The camera (11-1) may include all various devices for capturing two-dimensional images, such as RGB images.

[0064] The depth camera (11-2) is a camera that measures the distance to an object in 3D space and, unlike an RGB camera, generates a map containing depth information of the scene. Specifically, the depth camera (11-2) generates a depth map containing depth information based on data, i.e., depth information, where each pixel represents a distance value from the depth camera (11-2) to a specific object.

[0065] The communication unit (12) may include one or more components that enable connection to a communication network (2), and may include, for example, at least one of a short-range communication module, a wired communication module, and a wireless communication module.

[0066] The short-range communication module may include various short-range communication modules that transmit and receive signals using a wireless communication network at short range, such as a Bluetooth module, an infrared communication module, an RFID (Radio Frequency Identification) communication module, a WLAN (Wireless Local Access Network) communication module, an NFC communication module, and a Zigbee communication module.

[0067] Wired communication modules may include various wired communication modules such as Local Area Network (LAN) modules, Wide Area Network (WAN) modules, or Value Added Network (VAN) modules, as well as various cable communication modules such as USB (Universal Serial Bus), HDMI (High Definition Multimedia Interface), DVI (Digital Visual Interface), RS-232 (recommended standard 232), power line communication, or POTS (plain old telephone service).

[0068] In addition to Wi-Fi modules and WiBro (Wireless broadband) modules, the wireless communication module may include wireless communication modules that support various wireless communication methods such as GSM (global System for Mobile Communication), CDMA (Code Division Multiple Access), WCDMA (Wideband Code Division Multiple Access), UMTS (universal mobile telecommunications system), TDMA (Time Division Multiple Access), and LTE (Long Term Evolution).

[0069] The output unit (14) outputs a display that displays an input image and a pose estimation result (5) of a camera (11-1) transmitted by a server (20). To this end, the output unit (14) may be provided with a Digital Light Processing (DLP) panel, a Plasma Display Panel, a Liquid Crystal Display (LCD) panel, an Electro Luminescence (EL) panel, an Electrophoretic Display (EPD) panel, an Electrochromic Display (ECD) panel, a Light Emitting Diode (LED) panel, or an Organic Light Emitting Diode (OLED) panel, but is not limited thereto.

[0070] The user terminal (10) may include various additional configurations in addition to the configuration shown in FIG. 2a, and is not limited to the names used to refer to the configurations.

[0071] The server (20) includes an input unit (21) for receiving input commands from a user, a communication unit (22) for communicating with a user terminal (10), a memory (23) for storing input images received by the communication unit (22) or for storing large amounts of data and algorithms necessary for executing a disclosed visual localization method, and a processor (25) for controlling each component of the server (20).

[0072] Specifically, the input unit (21) receives various input commands necessary to generate three-dimensional line group data. To this end, the input unit (11) may include hardware devices such as various buttons, switches, keyboards, mice, trackballs, various levers, handles, or sticks that are connected to the server (20) and can receive setting commands from the user. Additionally, the input unit (11) may include a GUI (Graphical User Interface), i.e., a software device such as a touch pad. The touch pad may be implemented as a touch screen panel (TSP) and form a layered structure with the display.

[0073] The memory (23) may be implemented as at least one of a non-volatile memory device such as a cache, ROM (Read Only Memory), PROM (Programmable ROM), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), and flash memory, a volatile memory device such as RAM (Random Access Memory), or a storage medium such as a hard disk drive (HDD) and CD-ROM, but is not limited thereto. The memory (23) may be a memory implemented as a separate chip from the processor (25) described below, or it may be implemented as a single chip with the processor (25).

[0074] The processor (25) is a configuration that performs the disclosed visual localization method while controlling the hardware included in the server (20). To this end, the processor (25) may refer to a data processing device embedded in hardware having a physically structured circuit to perform a function expressed by code or instructions included in a program. As an example of a data processing device embedded in hardware, the processor (25) may include, but is not limited to, processing devices such as a microprocessor, a central processing unit (CPU), a processor core, a multiprocessor, an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), and a Graphics Processing Unit (GPU). The processor (25) may include one or more processors.

[0075] The processor (25) can be software-wise divided into a fixed point generation unit (26) that generates at least two fixed points from three-dimensional point cloud data, a three-dimensional line cloud data generation unit (27) that generates three-dimensional line cloud data by connecting one of the generated fixed points with a three-dimensional point included in the three-dimensional point cloud data, and a visual localization unit (28) that estimates the pose of a camera (11-1) that captured an input image based on the generated three-dimensional line cloud data. The processor (25) transmits the pose estimation result generated by the visual localization unit (28) to a user terminal (10) through a communication unit (22).

[0076] A detailed description of the operation of the software-separated processor (25) will be provided later through other drawings below.

[0077] Meanwhile, the server (20) may omit some of the configurations described in FIG. 2a. For example, if the server (20) is configured as a cloud server, the server (20) may be configured only with a processor (25) and memory (23), or multiple chips implementing the processor (25) and memory (23) may be provided.

[0078] FIG. 2b is a control block diagram of a system according to another disclosed embodiment. Descriptions of configurations that overlap with those shown in FIG. 2a are omitted.

[0079] Referring to FIG. 2b, a user terminal (10) according to another disclosed embodiment may include, in hardware, a camera (11-1) that captures an input image, a communication unit (12) that communicates with a communication unit (22) of a server (20) through a communication network (2), a processor (15) that controls the entire user terminal (10), performs image processing of the input image captured by the camera (11-1), receives three-dimensional line data transmitted by the server (20), and performs pose estimation of the camera (11-1), an output unit (14) that displays the input image captured by the processor (15) or outputs the pose estimation result performed by the processor (15), and a memory (13) that stores various data necessary for the operation of the processor (15).

[0080] Unlike the embodiment described in FIG. 2a, the system (1) according to the embodiment of FIG. 2B allows the processor (15) of the user terminal (10) to perform a visual localization process, that is, pose estimation of the camera (11-1). To do this, the processor (15) can receive three-dimensional line group data through the server (20). The processor (15) can perform pose estimation of the camera by collecting the input image captured by the camera (11-1) and depth information generated from the depth camera (11-2), and matching the three-dimensional line group data received from the server (20) with the extracted feature points and depth information. The method of performing pose estimation of the camera is the same as the method of operation of the visual localization unit (28) of the server (20) described above in FIG. 2a, and is merely a difference in the entity performing the visual localization.

[0081] Meanwhile, the user terminal (10) according to the embodiment of FIG. 2b may have the processor (15) and memory (13) provided as the same chip, may include additional configurations other than those shown in FIG. 2b, or some configurations may be omitted.

[0082] Below, a visual localization method that can be implemented in both the embodiments of FIG. 2a and FIG. 2b will be described in detail.

[0083] Figure 3 is an overall flowchart of the disclosed visual localization method.

[0084] Referring to FIG. 3, a system (1, hereinafter the system) that performs the disclosed visual localization method loads three-dimensional point cloud data (100).

[0085] Here, the 3D point cloud data is a large amount of map data prepared in advance to determine where the input image captured by the camera (11-1) that is trying to estimate the pose was taken.

[0086] Three-dimensional point cloud data can be generated in various ways. For example, three-dimensional point cloud data can be generated through a Structure-from-Motion (SfM) pipeline. The three-dimensional point cloud data can be stored in advance in the memory (23) of the server (20), and when a request is received from the user terminal (10), it can be loaded from the memory (23).

[0087] The system (1) generates a fixed point from three-dimensional point cloud data (200).

[0088] Unlike conventional technology, the disclosed visual localization method generates a single fixed point, and the specific method for generating the fixed point will be described in detail later through the drawings from Fig. 4 onwards.

[0089] Conventionally, to remove personal or sensitive information contained in 3D point cloud data, a method was proposed to generate 3D line cloud data by randomly selecting two points from the 3D point cloud data. Additionally, to enhance the cryptographic strength of such 3D line cloud data, a new method for generating 3D line cloud data was disclosed in which 3D straight lines intersect at two or more predefined fixed points. However, it has been argued that all of these conventional methods are capable of reverse reconstruction, which could reveal personal or sensitive information. The disclosed visual localization method generates 3D line cloud data based on a single pre-set fixed point, but by generating a fake feature point described below, it proposes a visual localization method that makes reverse reconstruction impossible.

[0090] The system (1) generates three-dimensional line group data by connecting the generated fixed points and three-dimensional points included in the three-dimensional point group data (300).

[0091] Specifically, the system (1) randomly selects some of the three-dimensional points included in the point cloud data. Afterward, the system (1) removes the selected three-dimensional points and selects a preset number of feature point descriptors from the removed three-dimensional points. The system (1) generates multiple fake feature points by assigning feature point coordinates based on Gaussian noise, random numbers, or user-specified coordinates to the selected feature point descriptors. The system (1) pairs the generated fake feature points with fixed points and simultaneously pairs the unremoved three-dimensional points with fixed points. That is, the system (1) generates line cloud data in the form of a unit sphere by connecting the three-dimensional points and fixed points of the point cloud data, and the generated fake feature points and fixed points. The specific method by which the system (1) generates three-dimensional line cloud data is explained in more detail through the drawings from FIG. 10 onwards.

[0092] The system (1) receives input images and depth information (400).

[0093] Specifically, the system (1) estimates the attitude of the camera (11-1) through generated line data stored in the server (20). To do this, the system (1) uses the RGB image of the camera (11-1), i.e., the input image. However, unlike the prior art, the disclosed system (1) estimates the attitude of the camera by using not only the input image but also the depth information of the depth camera (11-2).

[0094] According to the embodiment of FIG. 2a, the server (20) receives input images and depth information from the user terminal (10). According to the embodiment of FIG. 2b, the user terminal (10) receives three-dimensional line group data from the server (20) and collects input images from the camera (11-1) and depth information from the depth camera (11-2).

[0095] The system (1) estimates the orientation of the camera (11-1) (500).

[0096] The system (1) estimates the pose of the camera (11-1) through three-dimensional line group data generated based on a single fixed point. Specifically, the system (1) generates three-dimensional feature points based on two-dimensional feature points and depth information of the input image. The system (1) matches the generated three-dimensional feature points with the generated three-dimensional line group data. The system (1) estimates the pose of the camera (11-1) based on the intersection points of the matched three-dimensional lines and an absolute pose estimation algorithm. A detailed explanation of the pose of the camera (11-1) is provided later through the drawings from FIG. 14 onwards.

[0097] FIG. 4 is a flowchart for specifically explaining the first embodiment of the fixed point generation step of FIG. 3.

[0098] Referring to FIG. 4, the system (1) receives an input command instructing the user to execute the creation of a fixed point (210).

[0099] The input command may be received by the user terminal (10), or the server (20) may receive it from the user terminal (10) through the communication unit (22).

[0100] Based on the user's input command, the system (1) generates N clusters based on a clustering algorithm (211).

[0101] N clusters can be generated through a general clustering algorithm (e.g., K-means clustering algorithm).

[0102] The system (1) extracts the center point of N clusters (212) and determines the extracted center point as a fixed point (213).

[0103] FIG. 5 is a flowchart for specifically explaining a second embodiment of the fixed point generation step of FIG. 3. FIG. 6 is a diagram for explaining a first embodiment of setting a three-dimensional spatial region in FIG. 5, and FIG. 7 is a diagram for explaining a second embodiment of setting a three-dimensional spatial region in FIG. 5. To avoid redundant descriptions, they are described together below.

[0104] Referring to FIG. 5, the system (1) sets a three-dimensional spatial region where a fixed point exists (220).

[0105] The step (200) of generating a fixed point, unlike the first embodiment (Fig. 4), pre-sets a candidate region where a candidate fixed point can be generated, i.e., a three-dimensional space region, and generates a candidate fixed point that is a candidate for a fixed point within the three-dimensional space region. In setting the three-dimensional space region, the system (1) may include the first embodiment (see Fig. 6) and the second embodiment (see Fig. 7).

[0106] Referring to FIG. 6, the system (1) can perform Principal Component Analysis (PCA) (221-2) on three-dimensional point cloud data (221-1). Here, Principal Component Analysis (PCA) is a technique used for dimensionality reduction and is used in machine learning, data mining, statistical analysis, and noise removal. That is, the system (1) converts high-dimensional data contained in the three-dimensional point cloud data into low-dimensional data.

[0107] The system (1) calculates three basis axes (v1, v2, v3) and variances (σ1, σ2, σ3) for each axis through PCA analysis. The system (1) can set a region that is 2 sigma (2xσ1, 2xσ2, 2xσ3) away from each basis axis relative to the centroid as a three-dimensional spatial region (221-3).

[0108] Meanwhile, in the embodiment of FIG. 6, the setting standard for the three-dimensional space area is shown as 2 sigma, but it is not necessarily limited to this and can be changed to various values ​​such as 1.96 sigma.

[0109] Referring to FIG. 7, the system (1) can calculate the centroid and radius (r) of three-dimensional points included in three-dimensional point cloud data (221-1) (222-2). Specifically, the system (1) can calculate the distance between the centroid and the three-dimensional points, set the distance farthest from the centroid as the radius (r), and set the area within three times the radius (3r) from the centroid as the three-dimensional space area.

[0110] Meanwhile, in the embodiment of FIG. 7, the setting standard for the three-dimensional spatial area is shown as three times the radius, but it is not necessarily limited to this and can be changed to various values.

[0111] Referring again to FIG. 5, the system (1) samples candidate fixed points (230).

[0112] The method for generating candidate fixed points is specifically explained through Figs. 8 and 9.

[0113] The system (1) determines whether a sampled candidate fixed point exists within a set three-dimensional space region (240).

[0114] If the candidate fixed point is not included within the set three-dimensional space area (No. 240), the system (1) discards the previous candidate fixed point and creates a new candidate fixed point (230).

[0115] If a candidate fixed point is included within a set three-dimensional space area (Yes of 240), the system (1) determines the candidate fixed point as a fixed point for generating line group data (250).

[0116] FIG. 8 is a flowchart for specifically explaining the step of sampling candidate fixed points of FIG. 5.

[0117] Referring to FIG. 8, the system (1) can select any 3D point among the point clouds of 3D point cloud data included in a set 3D spatial area (231). Additionally, the system (1) can generate a virtual point based on 3D coordinate values ​​through random number generation among the point clouds of 3D point cloud data included in the set 3D spatial area (232).

[0118] The 3D point generated in step 231 or step 232 is sampled as a candidate fixed point (233).

[0119] Here, the coordinates of the virtual point in step 232 are arbitrary points generated as a result of random number generation. The system (1) determines whether the generated virtual point exists within a three-dimensional space area (step 240 of FIG. 5). If the virtual point does not exist within the three-dimensional space area, the virtual point is discarded, and a new virtual point is generated again through random number generation. If the virtual point is included within the three-dimensional space area, the system (1) determines it as a candidate fixed point.

[0120] Figure 9 is a diagram illustrating an example of generating line group data based on point group data.

[0121] The system (1) may set a plurality of virtual points included in the 3D point cloud data or within the 3D spatial region as candidate fixed points, and may generate a centroid within the cluster as a fixed point through the K-means clustering algorithm. The system (1) may generate unit sphere-shaped line cloud data (233-3) based on one fixed point.

[0122] Figure 10 is a flowchart specifically explaining the steps for generating Seon-gun data. Figure 11 is a diagram for visually explaining the steps for generating Seon-gun data of Figure 10. To avoid redundant explanations, they are described together below.

[0123] Referring first to FIG. 10, the system (1) removes a randomly selected 3D point among a plurality of 3D points (310).

[0124] Referring to FIG. 11, the system (1) has generated a centroid from three-dimensional point cloud data (A). The system (1) projects the three-dimensional points (True Points) of the three-dimensional point cloud data (A) onto a unit sphere (B) centered on the centroid. Here, the unit sphere is based on the map scale of the three-dimensional point cloud data, and the radius of the unit sphere may vary depending on the map scale.

[0125] The system (1) removes a randomly selected 3D point among a plurality of 3D points (320).

[0126] In the example of FIG. 11, the system (1) selects 6 three-dimensional points to be removed (Rejected Points) out of 12 three-dimensional points projected onto a unit sphere. However, the embodiment of FIG. 11 is just one example, and the number of three-dimensional points to be removed may vary. The system (1) removes (C) the selected three-dimensional points from the unit sphere.

[0127] The system (1) selects a preset number of feature point descriptors of the removed 3D points (330).

[0128] Each 3D point included in the 3D point cloud data may include various feature point descriptors. The system (1) selects one feature point descriptor from among the feature point descriptors of the removed feature points to generate a fake feature point. However, the number of selected feature point descriptors may vary.

[0129] The system (1) generates multiple fake feature points (340).

[0130] Specifically, the system (1) generates a fake feature point by assigning feature point coordinates to the feature point descriptor selected in step 330. That is, in the example of FIG. 11, the system (1) generates a fake feature point (D) having coordinates different from the coordinates of the removed 3D point.

[0131] Meanwhile, to prevent the generation of fake feature points with coordinates that are too different from the existing coordinates, the system (1) may generate fake feature points based on Gaussian noise, random numbers, or specified coordinates entered by the user. Various embodiments for generating fake feature points in addition to the example in FIG. 11 are described later in FIG. 12.

[0132] When a fake feature point is generated, the system (1) pairs the generated fake feature point and the unremoved 3D point with the fixed point (350).

[0133] After a fake feature point is generated, the fixed point can be matched with an existing 3D point, that is, an unremoved 3D point (True Point). Additionally, the fixed point can be matched with the generated fake feature point. The system (1) pairs the unremoved 3D point with the fixed point (Centroid) and pairs the fake feature point with the fixed point.

[0134] The system (1) generates line data by connecting paired points to each other (360).

[0135] Specifically, the system (1) generates a line in the direction of the line generation direction as the direction of a 3D point paired from a fixed point, and the reference point of the line becomes the fixed point. After generating the line, the system (1) deletes the location of the 3D point (including fake feature points). In addition, the system (1) generates 3D line group data by generating a 3D line based on preset color information and line feature information. The unit sphere generated through this, i.e., the line group data, can have high security and encryption power.

[0136] FIG. 12 is a flowchart illustrating various embodiments for generating fake feature points.

[0137] The system (1) selects a preset number (N) of feature point descriptors (341).

[0138] For example, the system (1) can generate a three-dimensional random number (342) and generate N fake feature points based on the generated random number (345).

[0139] As another example, the system (1) can generate new coordinates by adding N random numbers to the feature point coordinates (343).

[0140] Here, feature point coordinates may be generated based on the coordinates of feature points included in a 3D spatial map. The system (1) can generate new coordinates by adding N random numbers to N feature point coordinates of the 3D spatial map. The system (1) generates N fake feature points based on the generated new coordinates (345).

[0141] As another example, the system (1) generates coordinates where a fake feature point is located based on coordinate values ​​specified by the user (344).

[0142] Here, the system (1) can receive coordinate values ​​through a user terminal (10) or receive coordinate values ​​from a user who can operate the server (20). The system (1) generates N new coordinates through the coordinate values ​​specified by the user and generates N fake feature points based on the generated new coordinates (345).

[0143] FIG. 13 is a flowchart illustrating a specific method for performing camera attitude estimation. FIG. 14 is an example illustrating a line map. FIG. 15 is a diagram illustrating the camera attitude estimation method of FIG. 13 in detail. To avoid redundancy, they are described together below.

[0144] Referring first to FIG. 13, the system (1) loads three-dimensional line group data (510).

[0145] Here, the 3D line group data is line group data generated through a single fixed point (step 300 of FIG. 3). The generated line group data is stored in memory (23).

[0146] In accordance with one disclosed embodiment (Fig. 2a), when the server (20) performs pose estimation of the camera (11-1), the server (20) uses previously stored line group data. However, in accordance with another embodiment (Fig. 2b), when the user terminal (10) performs pose estimation of the camera (11-1), the server (20) transmits three-dimensional line group data stored through memory (23), etc., via the communication unit (22), and the user terminal (10) uses the line group data received from the server (20).

[0147] The system (1) extracts feature point information and depth information of the input image (520).

[0148] Here, the feature point information may include various information of feature points included in the input image, that is, the two-dimensional image captured by the camera (11-1). The two-dimensional feature information may include the position of the point, pixel information, and color information. When the server (20) performs pose estimation of the camera (11-1) according to one disclosed embodiment (Fig. 2a), the user terminal (10) extracts two-dimensional feature point information from the input image and transmits it to the server (20). When the user terminal (10) performs pose estimation of the camera (11-1) according to another disclosed embodiment (Fig. 2b), the user terminal (10) extracts two-dimensional feature point information and performs the following steps.

[0149] Meanwhile, the system (1) extracts not only two-dimensional feature point information of the input image but also depth information obtained from the depth camera (11-2). When the server (20) performs pose estimation of the camera (11-1) according to one disclosed embodiment (Fig. 2a), the user terminal (10) transmits the depth information extracted from the depth camera (11-2) to the server (20). When the user terminal (10) performs pose estimation of the camera (11-1) according to another disclosed embodiment (Fig. 2b), the user terminal (10) extracts depth information from the depth camera (11-2) and performs the following steps.

[0150] The system (1) matches the 2D feature point information of the input image with the depth information of the depth camera (11-2) and the loaded line group data (530).

[0151] Specifically, the system (1) first matches depth information corresponding to 2D feature points extracted from an input image. Then, the system (1) performs one-to-one matching between the 2D feature point information, which matches the depth information based on the feature descriptor, and the lines (lines) included in the line group data.

[0152] The system (1) clusters lines based on fixed points (540), and the server (20) samples lines within the line cluster (550).

[0153] The line group data stored in advance in the server (20) may be in the form of multiple unit sphere-shaped line group data generated from a single fixed point merged together. That is, the system (1) uses a line group map containing a line group distribution in which there is at least one point where at least three straight lines intersect.

[0154] As in the embodiment of FIG. 14, the loaded line group data may be a line group map (275) comprising unit spheres (271, 272, 273, 274) generated based on one fixed point for each subspace. The system (1) clusters units among the unit spheres where at least three straight lines intersect. After clustering multiple line group data in the line group map, the system (1) may sample straight lines (lines) included within the line group clusters. The sampling method may vary.

[0155] Based on the sampled line cluster, the system (1) performs pose estimation of the camera (560).

[0156] Specifically, the system (1) uses an absolute pose estimation algorithm that considers the intersection point (fixed point) of three-dimensional lines as a virtual camera center. Here, the system (1) generates three-dimensional feature points obtained by extracting depth information that matches two-dimensional feature points from a depth map and then back-projecting it.

[0157] Referring to FIG. 15, the system (1) can generate virtual points, namely three-dimensional feature points (L1, L2, L3), that can connect feature points of the input image (111) matched from a map containing depth information (hereinafter, depth information map, 112) with unit spheres. In order to correspond the three-dimensional feature points (L1, L2, L3) with a three-dimensional line, the system (1) performs the matching by using feature descriptors included in the three-dimensional feature points and the three-dimensional line.

[0158] Generally, visual localization methods using a map containing depth information (hereinafter referred to as a depth information map) have a problem in that the localization accuracy is lowered due to sensor noise inherent in the depth information map (image). The disclosed system (1) can improve localization accuracy by introducing a depth map regularization loss function. The total loss function used in the disclosed visual localization method can be expressed as Equation 1 below.

[0159]

[0160] Here, is the projection-based loss function (Reprojection error), and is the loss function of the depth information map, and λ represents the weight of the loss function of the depth information map. The disclosed system (1), based on verification results using an actual dataset, has a depth map regulation with appropriate weights (λ=10) compared to when no depth map regulation is applied (λ=0). -4 Localization accuracy improved when ) was applied.

[0161] As described above, the system (1) can estimate the absolute pose by using a perspective camera model for the localization technique of the line map, by considering the lines among the lines matched with 3D feature points in the line map as camera rays coming from each virtual camera (fixed point).

[0162] Through this, the system (1) can prevent damage caused by the leakage of map data in fields where security for spatial facilities is emphasized, such as defense and industrial fields, and can protect privacy in preparation for the distribution of high-definition video devices and spatial restoration technology. In particular, the system (1) can be applied to products where real-time visual localization for immediate interaction with the surrounding environment is essential, such as in the fields of autonomous driving and robotics, and can maintain high spatial information security by hindering attempts to restore from a 3D line cloud map to 3D point cloud data.

[0163] Those skilled in the art to which the present invention pertains should understand that the embodiments described above are illustrative in all respects and not restrictive, as the present invention may be implemented in other specific forms without altering its technical concept or essential features. The scope of the present invention is defined by the claims set forth below rather than by the detailed description above, and all modifications or variations derived from the meaning and scope of the claims and equivalent concepts thereof should be interpreted as being included within the scope of the present invention.

Claims

1. Generate a single fixed point from 3D point cloud data; By connecting a plurality of 3D points included in the above 3D point cloud data with the generated fixed point, 3D line cloud data is generated; Receive input video captured through a camera and depth information captured through a depth camera; A visual localization method comprising: extracting 2D feature point information included in the input image; and estimating the pose of the camera that captured the input image based on the extracted 2D feature point information and depth information matching the 2D feature points.

2. In Paragraph 1, Generating the above 3D line group data is, A visual localization method including projecting a three-dimensional point onto a unit sphere.

3. In Paragraph 2, Generating the above 3D line group data involves randomly removing 3D points from among the 3D points projected onto a unit sphere; Select a preset number of feature point descriptors of the above-mentioned removed 3D points; A visual localization method comprising generating a plurality of fake feature points based on the selected feature point descriptors.

4. In Paragraph 3, Generating the above 3D line group data is, Pairing the above-mentioned generated sham feature points and unremoved 3D points with the above-mentioned fixed points; A visual localization method comprising generating three-dimensional line group data by generating lines connecting paired points to each other.

5. In Paragraph 1, Generating the above 3D line group data is, A visual localization method comprising: generating a plurality of unit spheres based on a subspace divided from the above three-dimensional point cloud data.

6. In Paragraph 1, Creating the above fixed point is, A 3D spatial region where a fixed point exists is set from the above 3D point cloud data; Sample candidate fixed points; A visual localization method comprising generating the fixed point based on whether the candidate fixed point exists within the above three-dimensional spatial region.

7. In Paragraph 6, Establishing the above three-dimensional spatial region is, Based on the principal component analysis of the above 3D point cloud data, three basis axes and the variance for said basis axes are calculated; A visual localization method comprising establishing the three-dimensional spatial region based on the variance calculated above.

8. In Paragraph 6, Establishing the above three-dimensional spatial region is, Calculate the centroid of the 3D points included in the above 3D point cloud data; A visual localization method comprising: establishing a three-dimensional spatial region by setting the distance between the center of gravity and the three-dimensional point as a radius.

9. In Paragraph 6, Sampling the above candidate fixed points is, Select any 3D point from the above 3D point cloud data, or A visual localization method comprising sampling the candidate fixed points based on 3D coordinate values ​​generated through random number generation.

10. In Paragraph 1, Estimating the attitude of the above camera is, Generate 3D feature points based on the above 2D feature point information and the above depth information; Matching the above 3D line group data with the above 3D feature points; A visual localization method comprising: estimating the pose of the camera based on the intersection points of matched 3D lines and an absolute pose estimation algorithm.

11. Processor; It includes a program for operating the above processor and a memory for storing three-dimensional point cloud data received from the outside; and The above processor is, A fixed point is generated from the above 3D point cloud data, and By connecting a plurality of 3D points included in the above 3D point cloud data with the generated fixed point, 3D line cloud data is generated, and Receives input video captured through a camera and depth information captured through a depth camera, and A server that extracts 2D feature point information included in the input image and estimates the pose of the camera that captured the input image based on the extracted 2D feature point information and depth information matching the 2D feature points.

12. In Paragraph 11, The above processor is, Remove a randomly selected 3D point among the above plurality of 3D points; Select a preset number of feature point descriptors of the above-mentioned removed 3D points; By assigning feature point coordinates based on Gaussian noise to the selected feature point descriptors above, multiple fake feature points are generated; Pairing the above-mentioned generated sham feature points and unremoved 3D points with the above-mentioned fixed points; A server that generates the three-dimensional line group data by generating lines that connect paired points to each other.

13. In Paragraph 12, The above processor is, A server that generates multiple fake feature points by assigning feature point coordinates based on random numbers or specified coordinates received from a user to the selected feature point descriptor.

14. In Paragraph 12, The above processor is, A server that projects the generated fake feature points and the unremoved 3D points onto a unit sphere based on the map scale of the 3D point cloud data.

15. In Paragraph 11, The above processor is, Generate 3D feature points based on the above 2D feature point information and the above depth information; Matching the above 3D line group data with the above 3D feature points; A server that estimates the pose of the camera based on the intersection points of matched 3D lines and an absolute pose estimation algorithm.

16. A user terminal including a camera and a depth camera; and A server that communicates with the above user terminal; including The above server is, Generate a single fixed point from 3D point cloud data, and By connecting a plurality of 3D points included in the above 3D point cloud data with the generated fixed point, 3D line cloud data is generated, and The above user terminal is, A system for extracting 2D feature point information included in an input image captured through the camera, and estimating the attitude of the camera that captured the input image based on the extracted 2D feature point information and depth information matching the 2D feature points.

17. In Paragraph 16, The above user terminal is, Generate 3D feature points based on the above 2D feature points and the above depth information; Matching the above 3D line group data with the above 3D feature points; A system for estimating the pose of the camera based on the intersection points of matched 3D lines and an absolute pose estimation algorithm.