Indoor positioning correction system and method using deep learning model

The deep learning-based indoor positioning correction system addresses the challenges of error accumulation in SLAM and high costs of lidar-based methods by using low-resolution images and a deep learning model to accurately determine a user's location in indoor spaces.

WO2025121993A1PCT designated stage expired Publication Date: 2025-06-123I INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/096408
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-07
Filing Date
2024-10-29
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Existing indoor positioning systems face challenges in accurately determining a user's location in indoor spaces due to error accumulation in SLAM technology and the high cost and data processing time required by lidar-based methods.

Method used

A deep learning-based indoor positioning correction system that uses low-resolution images acquired at regular intervals to calculate a user's location, correct node positions, and apply these corrections as learning data to improve the accuracy of the deep learning model.

Benefits of technology

The system reduces data processing time and error in determining a user's location by using a deep learning model that can perform candidate corrections independently on the user terminal and continuously improve its accuracy through recalibration and retraining.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024096408_12062025_PF_FP_ABST
    Figure KR2024096408_12062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided is an indoor positioning correction system using a deep learning model, according to an embodiment of the present invention. The indoor positioning correction system using a deep learning model comprises: a camera that acquires a plurality of images according to a change in a location of a user; a node generation unit that generates nodes on the basis of first feature points extracted from the images acquired according to the change in the location of the user; and a deep learning model that outputs result data obtained by correcting locations of the nodes by using the nodes and the plurality of images as input data, wherein the deep learning model is trained and established by using, as training data, data obtained by correcting the locations of the nodes on the basis of the location of the camera or the user derived from the result of analyzing second feature points extracted from the plurality of images, respectively.
Need to check novelty before this filing date? Find Prior Art

Description

Indoor positioning correction system and method using a deep learning model

[0001] The present invention relates to an indoor positioning correction system and method using a deep learning model that corrects data on a user's path generated based on Visual SLAM (Simultaneous Localization And Mapping) based on a user's location derived using SFM (structure from Motion).

[0002] With the advancement of electronic technology, technologies are being developed to implement various environments online. One such technology, virtualization, involves creating a virtual environment by virtualizing a real-world environment online. To virtualize a real-world environment, it's crucial to create a virtual environment using images captured from various angles of the real-world environment and to associate the created virtual environment with the location where the images were captured.

[0003] In outdoor environments, information about the image capture location can be obtained using GPS (Global Positioning System), making it easy to obtain location information. However, in indoor environments, it is difficult to confirm the indoor image capture location and its movement path. SLAM (Simultaneous Localization And Mapping) technology is used to determine location in indoor spaces. SLAM is a technology that generates location information (movement path information) simultaneously with movement based on analysis of sensing data / video data regarding the movement process. However, this SLAM technology has a problem in that errors accumulate when movement path information is continuously generated.

[0004] To compensate for this accumulated SLAM error, a method is being used to configure a loop along the path, setting the first and last points to be identical to compensate for the error. However, while this method may be useful for estimating the path of unmanned vehicles that circulate along a fixed path, such as robot vacuum cleaners, it may not be a useful method for building a digital twin while circulating along an irregular path. In other words, configuring a scan path as a loop to compensate for SLAM errors is inefficient and practically impossible.

[0005] In addition, a SLAM method that applies lidar to indoor positioning has been developed, but there is a problem of significant cost increase due to the mandatory use of lidar, and there is a problem of excessive time required to process data acquired by lidar.

[0006] The technical problem of the present invention is to provide an indoor positioning correction system that uses a minimum amount of data to facilitate data processing in the process of determining a user's location.

[0007] The technical task of the present invention is to reduce the error of path data for the user's location generated through Visual SLAM by post-correcting the path data using the user's location information calculated through a plurality of images, and to apply a deep learning model constructed using the post-corrected result data as learning data to an indoor positioning correction system.

[0008] The technical task of the present invention is to provide an indoor positioning correction system capable of locating a user's location in an indoor space in real time by performing a time-consuming task in a data calculation process in advance on a server side and operating a deep learning model built based on the result data performed on the server side on a user terminal.

[0009] A server according to an embodiment of the present invention can reduce data processing time by calculating the location of a moving user using low-resolution images acquired at regular intervals. Furthermore, by correcting nodes derived from user terminals using the user locations derived from the images, errors occurring during node generation can be reduced.

[0010] According to an embodiment of the present invention, since a deep learning model that uses data on the locations of nodes corrected by a server as learning data is driven by a user terminal, the user terminal can derive data similar to the result of performing SFM on its own without data processing by the server as result data.

[0011] According to an embodiment of the present invention, by constructing a deep learning model that utilizes the candidate correction data derived from the time-consuming SFM task performed on a server as training data, the user terminal can independently perform candidate correction of nodes. Furthermore, the resulting data from the deep learning model can be recalibrated on the server, and the recalibrated data can then be used as training data for the deep learning model, thereby continuously improving the accuracy of the deep learning model.

[0012] FIG. 1 is a drawing showing devices for implementing an indoor positioning correction system using deep learning according to an embodiment of the present invention.

[0013] FIG. 2 is a diagram showing an indoor positioning correction system using deep learning according to an embodiment of the present invention.

[0014] FIG. 3 is a diagram showing initial nodes generated in a user terminal according to an embodiment of the present invention.

[0015] FIG. 4 is a diagram showing changes in the position of a user who captured multiple images generated by a server according to an embodiment of the present invention.

[0016] FIG. 5 is a diagram illustrating the correction of initial nodes in a server according to an embodiment of the present invention.

[0017] Figure 6 is a flowchart illustrating an indoor positioning correction method using deep learning according to an embodiment of the present invention.

[0018] An indoor positioning correction system using a deep learning model according to an embodiment of the present invention is provided. The indoor positioning correction system using a deep learning model includes a camera that acquires a plurality of images according to a change in a user's position, a node generation unit that generates nodes based on first feature points extracted from images acquired according to a change in the user's position, and a deep learning model that outputs data obtained by correcting the positions of the nodes using the nodes and the plurality of images as input data, and the deep learning model is constructed by learning data obtained by correcting the positions of the nodes based on the positions of the camera or the user, which are derived as a result of analyzing second feature points extracted from each of the plurality of images, as learning data.

[0019] By way of example, the plurality of images includes a plurality of reference images and a plurality of sub-images captured between the reference images.

[0020] For example, the node generation unit matches the information about the locations where the plurality of images were taken with the nodes, and the data correcting the locations of the nodes is data correcting the locations of the reference nodes matching the plurality of reference images based on reference points representing the user's location derived as a result of analyzing the second feature points.

[0021] For example, the node generation unit generates the nodes through images acquired in real time based on Visual SLAM (Simultaneous Localization And Mapping), and the images are not stored in the user terminal on which the node generation unit is mounted.

[0022] For example, the deep learning model uses data for a plurality of points representing the user's location, which are derived by determining common features from the second features extracted from the plurality of images using SFM (structure from Motion), and correcting the locations of the nodes as learning data.

[0023] For example, the result data of the deep learning model is recalibrated based on data for a plurality of points representing the user's location derived by determining common features from the second features extracted from the plurality of images using SFM (structure from Motion), and the recalibrated data is used as training data of the deep learning model.

[0024] An indoor positioning correction system using a deep learning model according to an embodiment of the present invention is provided. The indoor positioning correction system using a deep learning model includes a user terminal that generates nodes based on first feature points extracted from an image acquired according to a change in the user's position, a server that corrects the positions of the nodes based on a plurality of positions of the user terminal derived from analyzing second feature points extracted from each of the plurality of images, and a user terminal that acquires a plurality of images according to a change in the user's position, and a server that analyzes second feature points extracted from each of the plurality of images, and the user terminal includes a deep learning model constructed by learning data obtained by correcting the positions of the nodes corrected by the server as training data.

[0025] By way of example, the plurality of images includes a plurality of reference images and a plurality of sub-images captured between capturing the reference images.

[0026] For example, the user terminal matches the information about the location where the plurality of images were taken with the nodes, and the server corrects the positions of the reference nodes matching the plurality of reference images based on reference points indicating the position of the user terminal derived from the analysis of second feature points extracted from each of the plurality of images.

[0027] For example, the user terminal generates the nodes through images acquired in real time based on Visual SLAM (Simultaneous Localization And Mapping), and the images are not stored in the user terminal.

[0028] For example, the server uses SFM (structure from Motion) to determine common feature points from the second feature points extracted from the plurality of consecutive images to inversely calculate the location of the user terminal, and the data correcting the locations of the nodes is data correcting the nodes based on a plurality of points representing the user's location.

[0029] For example, the deep learning model outputs data correcting the positions of the nodes as result data using the nodes generated in real time by the user terminal and the plurality of images acquired by the user terminal as input data.

[0030] For example, the user terminal transmits the result data to the server, and the server re-corrects the result data, which is data corrected by the deep learning model, using the input data of the deep learning model.

[0031] For example, the server transmits the recalibrated data to the user terminal, and the user terminal uses the recalibrated data as training data for the deep learning model.

[0032] For example, in a storage medium storing computer-readable instructions, the instructions are executed by a processor, and the processor performs an operation of generating nodes based on first feature points extracted from an image acquired according to a change in the user's position, an operation of inputting a plurality of images captured according to a change in the user's position and the nodes as input data to a deep learning model, and an operation of outputting data obtained by correcting the positions of the nodes, wherein the deep learning model is constructed by learning data obtained by correcting the nodes based on the user's position derived as a result of analyzing second feature points extracted from each of the plurality of images as learning data.

[0033] The advantages and features of the present invention, and the methods for achieving them, will become clearer with reference to the embodiments described in detail below together with the accompanying drawings. However, the present invention is not limited to the embodiments disclosed below and may be implemented in various different forms. The present embodiments are provided only to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the scope of the invention, and the present invention is defined only by the scope of the claims. Like reference numerals refer to like elements throughout the specification.

[0034] The terms “... part,” “... unit,” “... module,” etc., described in the specification mean a unit that processes at least one function or operation, which may be implemented by hardware, software, or a combination of hardware and software.

[0035] The detailed description exemplifies the present invention. Furthermore, the foregoing description illustrates and describes preferred embodiments of the present invention, and the present invention can be used in various other combinations, modifications, and environments. In other words, changes or modifications may be made within the scope of the inventive concept disclosed herein, the scope equivalent to the disclosed disclosure, and / or the scope of technology or knowledge in the art. The described embodiments illustrate the best possible state for implementing the technical idea of ​​the present invention, and various modifications required for specific applications and uses of the present invention are also possible. Therefore, the detailed description of the invention above is not intended to limit the present invention to the disclosed embodiments. Furthermore, the appended claims should be construed to include other embodiments.

[0036] FIG. 1 is a drawing showing devices for implementing an indoor positioning correction system using deep learning according to an embodiment of the present invention.

[0037] Referring to Fig. 1, an indoor positioning correction system using deep learning can be implemented by a user terminal (100), a 360 camera (200), and a server (300). The indoor positioning correction system is intended to precisely determine the position of a user moving in an indoor space, and an embodiment of the present invention describes a system capable of precisely implementing indoor positioning without utilizing a lidar. The position of the user measured by the indoor positioning correction system can be applied to a three-dimensional virtual space virtualized using a digital twin. A digital twin is a virtual model that digitally virtualizes a physical object, and a digital twin can be constructed by combining the position and path of the user measured by the indoor positioning correction system and a captured image according to the present invention.

[0038] Each of the user terminal (100) and the server (300) may include at least one processor and memory. The processor may be composed of one or more cores and may include a processor for data analysis and deep learning, such as a central processing unit (CPU), a general purpose graphics processing unit (GPGPU), a tensor processing unit (TPU), and an application central processing unit (AP) of a computing device. One or more processors may be controlled to process input data according to predefined operation rules or artificial intelligence models stored in memory. If one or more processors are artificial intelligence-dedicated processors, the artificial intelligence-dedicated processors may be designed with a hardware structure specialized for processing a specific artificial intelligence model.

[0039] The processor can read a computer program or command stored in the memory and perform data processing for the indoor positioning correction system according to the present embodiment. The memory can store various information required for the indoor positioning correction system according to the present embodiment. The memory can include at least one type of storage medium among a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, and an optical disk. In addition, the memory can include any type of computer-readable recording medium well known in the art to which the present invention pertains. The description of the above-described memory is merely an example, and the present disclosure is not limited thereto.

[0040] The 360 ​​camera (200) can capture a 360-degree panoramic image based on the user's location. The image captured by the 360 ​​camera (200) can be transmitted to the user terminal (100) and / or server (300) via a wired / wireless communication network.

[0041] The user terminal (100) can acquire real-time video and multiple images while the user moves through a location. The user terminal (100) is a portable electronic device, and may include, for example, a mobile phone, a smart phone, a laptop computer, a digital broadcasting terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device, a smartwatch, a smart glass, a head mounted display (HMD), etc. The user terminal (100) can generate nodes indicating the user's location based on Visual SLAM (Simultaneous Localization And Mapping), which will be described later. The nodes and multiple images generated by the user terminal (100) can be transmitted to the server (300) through a wired / wireless communication network.

[0042] The server (300) can correct the nodes generated by the user terminal (100) based on the received nodes and multiple images. The server (300) can inversely calculate the location of the user terminal (100) or the user who captured the received images using SFM (structure from motion), which will be described later. The server (300) can correct the location of the nodes generated by the user terminal (100) based on the estimated location of the user terminal (100) or the user. For example, the server (300) can be implemented as any type of computing device, such as a notebook, desktop, or laptop.

[0043] The communication network described in the present invention may include a wired / wireless communication network, and examples of such wireless communication networks include Wireless LAN (WLAN), Digital Living Network Alliance (DLNA), Wireless Broadband (Wibro), World Interoperability for Microwave Access (Wimax), Global System for Mobile communication (GSM), Code Division Multi Access (CDMA), Code Division Multi Access 2000 (CDMA2000), Enhanced Voice-Data Optimized or Enhanced Voice-Data Only (EV-DO), Wideband CDMA (WCDMA), High Speed ​​Downlink Packet Access (HSDPA), High Speed ​​Uplink Packet Access (HSUPA), IEEE 802.16, Long Term Evolution (LTE), Long Term Evolution-Advanced (LTE-A), Wireless Mobile Broadband Service (WMBS), 5G mobile communication service, This may include Bluetooth, LoRa (Long Range), RFID (Radio Frequency Identification), Infrared Data Association (IrDA), Ultra Wideband (UWB), ZigBee, Near Field Communication (NFC), Ultra Sound Communication (USC), Visible Light Communication (VLC), Wi-Fi, and Wi-Fi Direct.Additionally, wired communication networks may include wired Local Area Networks (LANs), wired Wide Area Networks (WANs), Power Line Communications (PLCs), USB communications, Ethernet, serial communications, and optical / coaxial cables.

[0044] FIG. 2 is a diagram showing an indoor positioning correction system using deep learning according to an embodiment of the present invention.

[0045] Referring to FIGS. 1 and 2, the indoor positioning correction system using deep learning may be implemented as a software, computer program, or smartphone-driven application, or may be a combination of these and hardware. The indoor positioning correction system using deep learning may be implemented by a single processor, but is preferably implemented by at least one processor included in a user terminal (100) and at least one processor included in a server (300).

[0046] The user terminal (100) may include a camera (110), a sensor (130), a node generation unit (150), and a deep learning model (170). Since the user terminal (100) performs shooting in an indoor space, it cannot obtain location information based on absolute coordinates, such as a Global Positioning System (GPS). Therefore, the user terminal (200) may generate relative location information. The relative location information may include a relative movement path from an indoor point where a previous omnidirectional image was acquired to an indoor point where a current omnidirectional image was acquired.

[0047] The camera (110) can capture images for Visual SLAM (Simultaneous Localization And Mapping) and capture multiple images for SFM (structure from Motion). The images may not be stored in the user terminal (100), and the multiple images may be transmitted to the server (300) after being stored in the user terminal (100). The multiple images may include multiple reference images and multiple sub-images captured between the reference images. That is, multiple sub-images may be captured between consecutive reference images.

[0048] For example, the reference images may be 360-degree panoramic images, while the sub-images may be lower-resolution images compared to the reference images. The reference images may serve as reference points for future node position correction. The 360-degree panoramic images may be used in a digital twin to be built in the future. For example, when a user clicks on a specific location in the digital twin, a 360-degree panoramic image or a post-processed version of the 360-degree panoramic image may be displayed.

[0049] For example, both the reference images and sub-images may be low-resolution images. Reference images may refer to images that serve as reference points for future position correction of nodes.

[0050] Multiple images can be captured whenever the user's position changes by a preset distance or at regular time intervals. For example, sub-images captured between two consecutive images can be captured whenever the user moves by 60 cm to 1 m. However, the preset distance may not be particularly limited. Since multiple images are acquired whenever the user's position changes by a preset distance or at regular time intervals, the time required for future data processing using the multiple images can be reduced.

[0051] The sensor (130) may include an inertial measurement unit (IMU). The sensor (130) may detect inertial characteristics of the user terminal (100) and generate electrical signals or data values ​​corresponding to the detected state. For example, the sensor (130) may include a gyro sensor and an acceleration sensor. The data measured by the sensor (130) may be defined as inertial sensing data, and the inertial sensing data may be data including changes in the user's position, movement direction, and position.

[0052] The node generation unit (150) can generate nodes based on first feature points extracted from an image acquired according to changes in the user's location. The nodes can express a path according to changes in the user's location. The node generation unit (150) can include a processor, program, or software that drives deep learning-based Visual SLAM (Simultaneous Localization And Mapping). The node generation unit (150) can extract first feature points from an image acquired in real time by the camera (110), and extract first common feature points that match the first feature points according to changes in the frame. The node generation unit (150) can generate nodes by identifying the direction in which the user moves, the distance moved, etc. based on the first common feature points. At this time, the image acquired by the camera (110) may not be stored in the user terminal (130).

[0053] The node generation unit (150) may consider changes in inertial sensing data in the process of generating nodes according to the user's movement direction and movement distance. The node generation unit (150) may reduce errors occurring in the generation of nodes based on changes in the user's position, movement direction, and position acquired through the inertial sensing data. Specifically, if outlier data exceeding a preset threshold value occurs in nodes generated based on first feature points acquired through an image, the node generation unit (150) may compare the outlier data with the inertial sensing data to determine whether to apply the outlier data. For example, the threshold value may include at least one of a threshold value for a change in direction and a threshold value for a movement distance, and the node generation unit (150) may not apply the outlier data if a change exceeding the threshold value occurs in data (outlier data) for nodes generated based on the first feature points, but no corresponding change occurs in the inertial sensing data. The nodes generated by the node generation unit (150) may be defined as basic nodes.

[0054] The node generation unit (150) can match the generated basic nodes with a plurality of images corresponding to the positions of the basic nodes. The node generation unit (150) can be linked with the camera (110) and / or the 360 ​​camera (200) to match the basic nodes indicating the user's position at the time when the plurality of images are captured. Specifically, the node generation unit (150) can match the reference images with the reference nodes indicating the user's position at the time when the reference images are captured. Since the initial location setting of the user is important in the process of determining the user's path, the node generation unit (150) can match the initial reference image captured at the user's initial position with the user's initial position on a map of the indoor space stored in advance.

[0055] A deep learning model (170) can be constructed by learning data processed by a server (300). The deep learning model (170) may refer to an artificial neural network driven by a user terminal (100). Specific construction and learning methods for the deep learning model (170) will be described later.

[0056] The user terminal (100) can transmit data including basic nodes, multiple images, and matching relationships between them to the server (300).

[0057] The server (300) can use multiple images acquired by the user terminal (100) to correct the basic nodes generated by the user terminal (100). The server (300) can include a camera position calculation unit (310) and a correction unit (330).

[0058] The camera position calculation unit (310) can inversely calculate the position of the moving user terminal (100) by analyzing the second feature points extracted from each of the plurality of images acquired by the user terminal (100). The camera position calculation unit (310) may include a processor, program, or software that drives SFM (structure from motion). The position of the user terminal (100) may have the same meaning as the user's position or the position at which each of the plurality of images was captured. The camera position calculation unit (310) can inversely calculate the position of the user terminal (100) that captures the plurality of images based on changes in the second feature points extracted from each of the plurality of images. At this time, the plurality of images used by the camera position calculation unit (310) may be images having a relatively low resolution. By using the plurality of low-resolution images captured at regular intervals, the camera position calculation unit (310) may take a relatively short time to calculate the position of the user terminal (100).

[0059] The correction unit (330) can correct the basic nodes generated by the user terminal (100) based on the location of the user terminal (100) calculated by the camera position calculation unit (310). The location of the user terminal (100) calculated by the camera position calculation unit (310) can be expressed as a plurality of points. To simplify the calculation, the correction unit (330) can correct the locations of the reference nodes among the basic nodes to the locations of the reference points, which represent the locations where the reference image was captured. The nodes corrected by the correction unit (330) can be defined as final nodes.

[0060] In general, deriving a user's location using SFM (structure from motion) can be more accurate than deriving the user's location using Visual SLAM (Simultaneous Localization And Mapping). However, real-time image-based SFM uses excessive images for data processing, which slows down data processing. The server (300) according to an embodiment of the present invention can reduce the time required for data processing by calculating the location of a moving user using low-resolution images acquired at regular intervals. In addition, by correcting the nodes derived by the user terminal (100) using the user's location derived from the images, errors occurring during node generation can be reduced.

[0061] Data on the final nodes corrected by the correction unit (330) can be transmitted to the user terminal (100). The deep learning model (170) can use the data on the final nodes corrected by the correction unit (330) as learning data.

[0062] The deep learning model (170) can be constructed by learning data in which the positions of nodes are corrected based on the positions of the camera (110) or the user derived from the analysis of the second feature points extracted from each of a plurality of images as training data. The correction unit (330) can correct the nodes generated by the node generation unit (150) using the positions of the camera (110) or the user derived from the camera position calculation unit (310). The deep learning model (170) can use data on the final nodes corrected by the correction unit (330), data on the nodes generated by the node generation unit (150), and a plurality of images acquired by the camera (110) as training data. In other words, the deep learning model (170) can use data in which the positions of nodes are corrected based on data on a plurality of points representing the positions of the user derived by determining common feature points from the second feature points extracted from a plurality of images using SFM (structure from Motion) as training data.

[0063] An artificial neural network model can be trained using at least one of supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. The artificial neural network model can be trained in a direction that minimizes the error of the output. The training process of the artificial neural network model may include a process of repeatedly inputting training data into the artificial neural network model, calculating the output of the artificial neural network model and the target error for the training data, and backpropagating the error of the artificial neural network model from the output layer of the artificial neural network model toward the input layer in a direction that reduces the error, thereby updating the weights of each node of the artificial neural network model. In the case of supervised learning, training data in which the correct answer is labeled for each training data is used (i.e., labeled training data), and in the case of unsupervised learning, the correct answer may not be labeled for each training data.

[0064] For example, the training data for the final nodes corrected by the correction unit (330) may be labeled data. Data for the nodes generated by the node generation unit (150) and a plurality of images may be input as training data to the deep learning model (170), and an error may be calculated by comparing the output data of the deep learning model (170) with the labels of the training data (data for the corrected final nodes).

[0065] The deep learning model (170) can perform candidate correction of basic nodes generated by the node generation unit (150) without going through the server (300). That is, the deep learning model (170) can perform a function similar to the SFM process performed by the server (300). Accordingly, the user terminal (100) can perform Visual SLAM based on real-time images to generate basic nodes, and input data including a plurality of images acquired by the user terminal (100) and basic nodes can be input to the deep learning model (170). As a result, the deep learning model (170) can output data correcting the positions of nodes in real time.

[0066] The result data corrected by the deep learning model (170) can be transmitted back to the server (300). The server (300) can recalibrate the result data output by the deep learning model (170) based on data on nodes generated by the node generation unit (150) input as input data to the deep learning model (170) and a plurality of images. The recalibration of the result data by the server (300) may be the same as the process in which the correction unit (330) corrects the basic nodes generated by the node generation unit (150). Specifically, the result data of the deep learning model (170) can be recalibrated by the correction unit (330) based on a plurality of points representing the user's location derived as a result of analyzing the second feature points by the camera position calculation unit (310). The result data recalibrated by the correction unit (330) can be transmitted to the user terminal (100). As the recalibrated result data is used as training data for the deep learning model (170), the accuracy with which the deep learning model (170) selects nodes can be continuously improved.

[0067] According to an embodiment of the present invention, as the data processing process becomes longer when determining a user's location using Visual SLAM, an error occurs. Therefore, by correcting nodes generated by Visual SLAM using SFM, which inversely calculates the user's location through analysis of images, the precision of determining a user's location in an indoor space can be improved.

[0068] According to an embodiment of the present invention, by generating nodes indicating the location of a user using Visual SLAM on the user terminal (100) side based on real-time images that are not stored, and by driving SFM that inversely calculates the location of the user terminal (100) or the user using a plurality of images on the server (300), the number of data processed on the user terminal (100) can be minimized. In addition, by performing SFM based on images with a relatively large amount of data processing on the server (300) side, a fast indoor positioning and data correction process can be performed according to the classification of data to be processed between the user terminal (100) and the server (300).

[0069] According to an embodiment of the present invention, as the deep learning model (170) that uses data on the locations of nodes corrected by the server (300) as learning data is driven by the user terminal (100), the user terminal (100) can derive data similar to the result of performing SFM on its own without a data processing process by the server (300) as result data.

[0070] FIG. 3 is a diagram showing initial nodes generated in a user terminal according to an embodiment of the present invention.

[0071] Referring to FIGS. 1 to 3, the user terminal (100) can perform Visual SLAM based on real-time images according to changes in the user's location. As a result of performing Visual SLAM, a plurality of basic nodes indicating the real-time user's location can be generated. The camera (110) of the user terminal (100) can capture images at regular intervals according to changes in the user's location. For example, the image captured by the 360 ​​camera (200) can be a reference image, and the image captured by the camera (110) can be a sub-image. The location where the 360 ​​camera (200) captures an image can be defined as a reference node (P0, P1, P2, P3, P4, P5, P6, P7). The camera (110) can also capture sub-images at the reference nodes (P0, P1, P2, P3, P4, P5, P6, P7). The node generation unit (150) can match reference images and / or sub-images captured at reference nodes (P0, P1, P2, P3, P4, P5, P6, P7) with each other, and the reference image and / or sub-image captured at the first reference node (P0) can be defined as the first image. The first image can be stored by being matched with the reference node (P0) indicating the first location in a map of a pre-stored indoor space. However, as the execution time of Visual SLAM gets longer, an error may occur between the actual user's location and the generated basic nodes.

[0072] The user terminal (100) can transmit basic nodes generated by the node generation unit (150), sub-images acquired by the camera (110), and reference images acquired by the 360 ​​camera (200) to the server (300). However, the 360 ​​camera (200) can also directly transmit reference images to the server (300).

[0073] FIG. 4 is a diagram showing changes in the position of a user who captured multiple images generated by a server according to an embodiment of the present invention.

[0074] Referring to FIGS. 1, 2, and 4, the server (300) can receive information on basic nodes, sub-images, reference images, and matching basic nodes and images from the user terminal (100).

[0075] The camera position calculation unit (310) can reverse-calculate the position of the user terminal (100) or the user capturing the sub-images by performing SFM based on the sub-images. Since the sub-images are captured at regular intervals, the position of the user terminal (100) or the user derived as a result of performing SFM using the sub-images may be less accurate than the position of the user terminal (100) or the user derived as a result of performing SFM based on real-time video. However, in the embodiment of the present invention, the position of the user terminal (100) or the user derived as a result of performing SFM is used to correct the basic nodes generated by the user terminal (100), so that the position of the user can be precisely estimated even if sub-images acquired at regular intervals are used.

[0076] The user terminal (100) or the user's location derived as a result of performing SFM can be divided into points where sub-images were captured and points where reference images were captured. The points where reference images were captured can be defined as reference points (C0, C1, C2, C3, C4, C5, C6, C7). That is, the camera position calculation unit (310) can generate a plurality of reference points (C0, C1, C2, C3, C4, C5, C6, C7) and points between two consecutive reference points (C0, C1, C2, C3, C4, C5, C6, C7).

[0077] FIG. 5 is a diagram illustrating the correction of initial nodes in a server according to an embodiment of the present invention.

[0078] Referring to FIGS. 1 to 5, the correction unit (330) can correct the basic nodes generated by the node generation unit (150) of the user terminal (100) based on the locations of a plurality of points.

[0079] For example, the correction unit (330) can determine whether the positions of the basic nodes and points match by matching the positions of the basic nodes and points. The correction unit (330) can correct the positions of the basic nodes generated by the node generation unit (150) with the positions of the points generated by the camera position calculation unit (310). Through the correction process as described above, the precision of the user position positioning can be improved. In addition, since the process of estimating the user position based on a plurality of images and the process of correcting the basic nodes based on the points representing the user position derived from the plurality of images are performed on the server (300) with a relatively fast data processing speed rather than on the user terminal (100), the time required for the entire data processing process can be reduced.

[0080] As another example, the correction unit (330) can correct the positions of the reference nodes (P0, P1, P2, P3, P4, P5, P6, P7) among the basic nodes to the positions of a plurality of reference points (C0, C1, C2, C3, C4, C5, C6, C7). In order to further reduce the overall data processing speed, the correction unit (330) can correct the positions of the reference nodes (P0, P1, P2, P3, P4, P5, P6, P7) based on the positions of the reference points (C0, C1, C2, C3, C4, C5, C6, C7) where the reference images were captured. If the data for some of the reference nodes (P0, P1, P2, P3, P4, P5, P6, P7) is the same as or within a preset range of the data for the reference points (C0, C1, C2, C3, C4, C5, C6, C7), the correction unit (330) may not perform position correction for the some of the nodes. In the present embodiment, the positions of the three reference nodes (P0, P1, P2) are matched with the three reference points (C0, C1, C2) based on the initial positions, and therefore the correction unit (330) may not perform position correction for the positions of the reference nodes (P0, P1, P2). The reference node corrected by the correction unit (330) may be defined as the final reference node, and the positions of the remaining basic nodes may be corrected based on a plurality of final reference nodes. Specifically, the correction unit (330) can correct existing nodes that have not been corrected based on the positions of the final reference nodes and changes in the basic nodes, and thus, final nodes that have been corrected can be derived.

[0081] Figure 6 is a flowchart illustrating an indoor positioning correction method using deep learning according to an embodiment of the present invention. For simplicity, redundant descriptions are omitted.

[0082] Referring to FIG. 6, basic nodes can be created using Visual SLAM at the user terminal. While the user terminal creates basic nodes in real time, a camera mounted on the user terminal can capture multiple images (S100).

[0083] The basic nodes and multiple images generated at the user terminal can be matched with each other. That is, the basic nodes and the images measured at the locations where the basic nodes are generated can be matched with each other (S200).

[0084] Basic nodes, images, and information matching basic nodes and images can be transmitted from the user terminal to the server (S300).

[0085] The server can receive basic nodes, images, and information matching the basic nodes and images transmitted by the user terminal. A process of inversely calculating the location of the user terminal or user can be performed on the server side using multiple images. The calculated location of the user terminal or user can be defined by multiple points (S400).

[0086] The server side can correct the positions of the basic nodes created by the user terminal based on the positions of multiple points. If the positions of the multiple points created by the server and the basic nodes created by the user terminal do not match, the server can determine that the user's position indicated by the multiple points created based on the images is more accurate (S500).

[0087] The candidate-corrected basic nodes can be defined as final nodes. Information about the candidate-corrected final nodes can be transmitted from the server side to the user terminal (S600).

[0088] Information about the final nodes selected by the server can be used as training data. A deep learning algorithm can be installed on the user terminal, and deep learning of the algorithm can be performed using the selected final nodes as training data (S700).

[0089] As a result of deep learning, a deep learning model can be built that independently performs candidate correction of basic nodes on the user terminal. The deep learning model can perform functions similar to the SFM process performed by the server (S800).

[0090] A user terminal can perform Visual SLAM based on real-time images to generate basic nodes, and input data including a plurality of images and basic nodes acquired by the user terminal can be input to a deep learning model. As a result, the deep learning model can generate final nodes by correcting the positions of the basic nodes based on changes in feature points extracted from the plurality of images. In other words, the deep learning model can derive result data by post-correcting the positions of the basic nodes using the basic nodes and a plurality of images generated by the user terminal as input data (S900).

[0091] The resulting data corrected by the deep learning model can be transmitted to a server. The server can derive multiple points regarding the user terminal's location using multiple images input as input data to the deep learning model. The server can recalibrate the resulting data of the deep learning model based on the multiple points (S1000).

[0092] The recalibrated result data can be transmitted from the server to the user terminal (S1100).

[0093] The recalibrated result data can be used as training data for a deep learning model. That is, by using the recalibrated result data of a deep learning model as training data for the deep learning model, the accuracy of the deep learning model's result data extraction can be enhanced (S1200).

[0094] According to an embodiment of the present invention, by constructing a deep learning model that utilizes the candidate correction data derived from the time-consuming SFM task performed on a server as training data, the user terminal can independently perform candidate correction of nodes. Furthermore, the resulting data from the deep learning model can be recalibrated on the server, and the recalibrated data can then be used as training data for the deep learning model, thereby continuously improving the accuracy of the deep learning model.

[0095] While the embodiments of the present invention have been described above with reference to the attached drawings, those skilled in the art will appreciate that the present invention can be implemented in other specific forms without altering the technical spirit or essential characteristics thereof. Therefore, the embodiments described above should be understood to be illustrative in all respects and not restrictive.

[0096] - Assignment information

[0097] Ministry Name: Ministry of SMEs and Startups

[0098] Research Management Specialist Organization: Small and Medium Business Technology Information Promotion Agency

[0099] Research Project Name: SME Technology Innovation Development Project 'Export-Oriented Project'

[0100] Research Project Title: Development of a Platform for Creating and Sharing Ultra-Realistic Content Based on AI Motion Capture

[0101] Host organization: 3RII Co., Ltd.

[0102] Research period: May 1, 2022 - December 31, 2025

Claims

1. A camera that acquires multiple images according to changes in the user's position; A node generation unit that generates nodes based on first feature points extracted from an image acquired according to changes in the user's location; and Includes a deep learning model that outputs data that corrects the positions of the nodes using the above nodes and the plurality of images as input data, The above deep learning model is, Based on the location of the camera or user derived from the analysis of the second feature points extracted from each of the above multiple images, the data is learned as learning data and the locations of the nodes are corrected. Indoor positioning correction system using deep learning model.

2. In paragraph 1, The above plurality of images include a plurality of reference images and a plurality of sub-images captured during the shooting of the reference images. Indoor positioning correction system using deep learning model.

3. In paragraph 2, The above node generation unit matches the information about the locations where the multiple images were taken with the above nodes, The data correcting the positions of the above nodes is data correcting the positions of the reference nodes matching the plurality of reference images based on the reference points indicating the user's position derived as a result of analyzing the second feature points. Indoor positioning correction system using deep learning model.

4. In paragraph 1, The above node generation unit generates the nodes through images acquired in real time based on Visual SLAM (Simultaneous Localization And Mapping). The above video is not stored in the user terminal in which the above node generation unit is installed. Indoor positioning correction system using deep learning model.

5. In paragraph 4, The above deep learning model uses data for correcting the positions of the nodes as learning data for multiple points representing the user's position by determining common features from the second feature points extracted from the multiple images using SFM (structure from Motion). Indoor positioning correction system using deep learning model.

6. In paragraph 1, The result data of the above deep learning model is recalibrated based on data for multiple points representing the user's location derived by determining common feature points from the second feature points extracted from the multiple images using SFM (structure from Motion). The recalibrated data is used as training data for the deep learning model. Indoor positioning correction system using deep learning model.

7. A user terminal that generates nodes based on first feature points extracted from images acquired according to changes in the user's location and acquires multiple images according to changes in the user's location; and A server is included that corrects the positions of the nodes based on the plurality of positions of the user terminal derived from the results of analyzing the second feature points extracted from each of the plurality of images. The above user terminal includes a deep learning model built by learning data that corrects the positions of the nodes corrected by the server as learning data. Indoor positioning correction system using deep learning model.

8. In paragraph 7, The above plurality of images include a plurality of reference images and a plurality of sub-images captured during the shooting of the reference images. Indoor positioning correction system using deep learning model.

9. In paragraph 8, The above user terminal matches the information about the location where the multiple images were taken with the above nodes, The server corrects the positions of the reference nodes matching the plurality of reference images based on the reference points indicating the positions of the user terminals derived from the analysis of the second feature points extracted from each of the plurality of images. Indoor positioning correction system using deep learning model.

10. In paragraph 7, The above user terminal generates the nodes through images acquired in real time based on Visual SLAM (Simultaneous Localization And Mapping). The above video is not stored within the user terminal. Indoor positioning correction system using deep learning model.

11. In clause 10, The above server uses SFM (structure from Motion) to determine common feature points from the second feature points extracted from the plurality of consecutive images and inversely calculates the location of the user terminal. The data that corrects the positions of the above nodes is data that corrects the above nodes based on multiple points that indicate the user's position. Indoor positioning correction system using deep learning model.

12. In paragraph 7, The above deep learning model outputs data that corrects the positions of the nodes as result data using the nodes generated in real time by the user terminal and the input data of the multiple images acquired by the user terminal. Indoor positioning correction system using deep learning model.

13. In paragraph 12, The above user terminal transmits the above result data to the server, The above server re-corrects the result data, which is data corrected by the deep learning model, using the input data of the deep learning model. Indoor positioning correction system using deep learning model.

14. In paragraph 13, The above server transmits the above recalibrated data to the user terminal, The above user terminal uses the recalibrated data as training data for the deep learning model. Indoor positioning correction system using deep learning model.

15. In a storage medium storing computer-readable instructions, the instructions are executed by a processor, and the processor, An operation of generating nodes based on first feature points extracted from an image acquired according to changes in the user's position; An operation of inputting multiple images and nodes captured according to changes in the user's location into a deep learning model as input data; and An operation is performed to output data as a result of correcting the positions of the above nodes, The above deep learning model is constructed by learning data that corrects the nodes based on the user's location derived from analyzing the second feature points extracted from each of the above multiple images. Storage medium.

Citation Information

Patent Citations

  • Method and apparatus for predicting dementia based on Activity of daily living

    KR1020210155335A

  • Method for map matching of user terminal

    KR102054349B1

  • Mobile Augmented Reality Service Apparatus and Method Using Deep Learning Based Positioning Technology

    KR102166586B1

  • LED Display board for CCTV

    KR102238700B1

  • Pre-slit holder

    KR102749487B1