Method and computer device for location recognition based on wayfinding map
By employing a particle filter algorithm and integrating FPV and BEV image modules, the technology effectively estimates a mobile robot's location on a distorted pathfinding map within indoor environments, overcoming the limitations of existing technologies.
Patent Information
- Application Number
- PCT/KR2024/013681
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-20
- Filing Date
- 2024-09-10
- Publication Date
- 2025-06-26
AI Technical Summary
Existing technologies struggle to accurately estimate a mobile robot's location in indoor environments without prior information or GPS signals, especially when using distorted pathfinding maps that do not accurately represent the real environment.
A location recognition technology that applies a particle filter algorithm and uses a combination of first-person view (FPV) and bird's eye view (BEV) image modules to estimate a location on a pathfinding map by understanding the semantic structure according to the spatial structure and layout of the road.
This solution enables accurate location estimation in indoor environments without prior information or GPS, improving position accuracy by probabilistically using continuous position estimation results and effectively handling distorted maps.
Smart Images

Figure KR2024013681_26062025_PF_FP_ABST
Abstract
Description
Method and computer device for map-based location recognition for pathfinding
[0001] The description below is about a technique for estimating location based on images.
[0002] A mobile robot must not only be able to determine its own position within a given environment, but also be able to create its own map of its surroundings when placed in a new, previously unfamiliar environment.
[0003] Mapping a mobile robot means identifying the locations of obstacles and objects in the surrounding area, as well as open spaces where it can move freely, and memorizing them in an appropriate way.
[0004] As an example of a map-making technology for a mobile robot, Korean Patent Publication No. 10-2010-0070922 (published on June 28, 2010) discloses a technology that can create a final grid map for location recognition of a mobile robot by creating a grid map using distance information to surrounding objects and then linking it with location information of landmarks.
[0005] Meanwhile, public spaces such as large shopping malls, theme parks, train stations, and airports provide wayfinding maps. Wayfinding maps are designed to guide visitors to their locations or destinations. These maps typically display various landmarks or stores in polygonal shapes. These polygons may not correspond to real-world structural features (such as walls or doors) but may instead outline semantic areas, such as food courts or amusement parks. Because the primary purpose of wayfinding maps is to convey information, not to accurately represent the real-world environment, they are often designed to emphasize certain areas and simplify less important ones.
[0006] Given such a pathfinding map, a technology is needed that allows a robot or device to estimate its own location.
[0007] It provides a visual localization technology that recognizes and understands the semantic structure according to the spatial structure and layout of the road on a distorted map that is different from the actual environment.
[0008] It provides a technology that can estimate a location on a pathfinding map in an indoor environment where there is no prior information about the space and global positioning system (GPS) signals are not available.
[0009] A position recognition technology is provided that applies a particle filter algorithm to probabilistically use continuous position estimation results.
[0010] A method for localization executed on a computer device, the computer device including at least one processor configured to execute computer-readable instructions contained in a memory, the method comprising: receiving, by the at least one processor, a first-person view image as visual information for localization in an indoor environment; and estimating, by the at least one processor, a location matching the first-person view image in a given wayfinding map for the indoor environment using the first-person view image.
[0011] According to one aspect, the above-mentioned pathfinding map corresponds to a guide map or a map of the indoor environment, and the above-mentioned estimating step can estimate a location on the pathfinding map using the first-person image in an environment where no prior information about the space is given and a global positioning system (GPS) signal cannot be used.
[0012] According to another aspect, the estimating step may include a step of converting the first-person image into a BEV (bird-eye view) image through a location recognition module and then comparing the BEV image with the pathfinding map through a cross-correlation operation.
[0013] According to another aspect, the location recognition module can be trained through a loss function between the heatmap, which is the output of the cross-correlation operation, and the ground truth (GT) heatmap.
[0014] According to another aspect, the estimating step may include: converting the first-person image into a BEV image, which is an egocentric bird-eye view representation; converting the pathfinding map into a BEV pathfinding map by converting it into a latent space identical to that of the BEV image; and comparing the BEV image with the BEV pathfinding map to search for a location on the pathfinding map.
[0015] According to another aspect, the step of converting the first-person image into a BEV image may include the steps of: processing the first-person image as a first latent feature of a first image encoder to embed a first BEV representation; processing the first-person image as a second latent feature of a second image encoder to embed a second BEV representation; and concatenating the first BEV representation and the second BEV representation to create the BEV image as a final BEV representation.
[0016] According to another aspect, the step of converting the first-person image into a BEV image includes the step of embedding a BEV representation by processing the first-person image into a latent feature of an image encoder, and the latent feature can be learned through two types of image modules, an FPV image module (Imagine-FPV) and a BEV image module (Imagine-BEV).
[0017] According to another aspect, the FPV image module can be trained through a loss function between an estimated map mask output through the FPV latent feature for the first-person image and a ground truth (GT) map mask that labels the observable area and masks the remaining area.
[0018] According to another aspect, the BEV image module can be trained through a loss function between the egocentric map predicted from the BEV latent representation for the first-person image and the GT egocentric map according to each camera pose.
[0019] According to another aspect, the estimating step may include a step of performing sequential inference for location recognition by assigning weights to the searched location through a particle filter algorithm.
[0020] According to another aspect, the step of performing the above may include a step of grouping the estimated particles on the pathfinding map using K-means clustering and estimating the center of the cluster with the highest sum of weights of particles in each cluster as a representative pose.
[0021] According to another aspect, the performing step may include a step of updating the pose of the estimated particle on the pathfinding map using wheel odometry data of the robot or device providing the first-person image.
[0022] A non-transitory computer-readable recording medium storing a computer program for executing the above location recognition method on the computer device is provided.
[0023] In a computer device, a computer device is provided, comprising at least one processor configured to execute computer-readable instructions contained in a memory, wherein the at least one processor processes a process of receiving a first-person image as visual information for location recognition in an indoor environment; and a process of estimating a location matching the first-person image in a given pathfinding map for the indoor environment using the first-person image.
[0024] According to embodiments of the present invention, a current location can be recognized on a distorted map that is different from the actual environment through a location recognition technology that recognizes and understands a semantic structure according to the spatial structure and layout of a road.
[0025] According to embodiments of the present invention, a location can be estimated in an indoor environment where there is no prior information about the space and GPS signals cannot be used, through a location recognition technology that recognizes and understands a semantic structure according to the spatial structure and layout of a road.
[0026] According to embodiments of the present invention, by applying a particle filter algorithm for probabilistically using continuous position estimation results, position accuracy can be improved by weighting positions according to probability.
[0027] FIG. 1 is a block diagram illustrating an example of the internal configuration of a computer device according to one embodiment of the present invention.
[0028] FIG. 2 is a flowchart illustrating an example of a location recognition method that a computer device according to one embodiment of the present invention can perform.
[0029] FIG. 3 illustrates an outline of a pathfinding map-based location recognition system according to one embodiment of the present invention.
[0030] FIG. 4 illustrates a detailed configuration of a pathfinding map-based location recognition framework according to one embodiment of the present invention.
[0031] FIG. 5 illustrates an example of a GT (ground-truth) label generation process in one embodiment of the present invention.
[0032] FIG. 6 illustrates an example flowchart of a particle filter algorithm used for location recognition in one embodiment of the present invention.
[0033] FIG. 7 illustrates an example of a visualization of a location recognition episode in one embodiment of the present invention.
[0034] Hereinafter, embodiments of the present invention will be described in detail with reference to the attached drawings.
[0035]
[0036] Embodiments of the present invention relate to a technique for estimating a location based on an image.
[0037] Embodiments including those specifically disclosed in this specification can estimate a location on a pathfinding map through location recognition that understands the semantic structure of a path when a pathfinding map created in a distorted form different from an actual environment for guiding people is given.
[0038] FIG. 1 is a block diagram illustrating an example of a computer device according to an embodiment of the present invention. For example, a location recognition system according to embodiments of the present invention may be implemented by the computer device (100) illustrated in FIG. 1.
[0039] As illustrated in FIG. 1, a computer device (100) may include a memory (110), a processor (120), a communication interface (130), and an input / output interface (140) as components for executing a location recognition method according to embodiments of the present invention.
[0040] The memory (110) is a computer-readable recording medium, and may include a random access memory (RAM), a read only memory (ROM), and a permanent mass storage device such as a disk drive. Here, the ROM and the permanent mass storage device such as the disk drive may be included in the computer device (100) as a separate permanent storage device distinct from the memory (110). In addition, the memory (110) may store an operating system and at least one program code. These software components may be loaded into the memory (110) from a computer-readable recording medium separate from the memory (110). This separate computer-readable recording medium may include a computer-readable recording medium such as a floppy drive, a disk, a tape, a DVD / CD-ROM drive, a memory card, etc. In another embodiment, the software components may be loaded into the memory (110) through a communication interface (130) rather than a computer-readable recording medium. For example, software components may be loaded into the memory (110) of a computer device (100) based on a computer program that is installed by files received over a network (160).
[0041] The processor (120) may be configured to process instructions of a computer program by performing basic arithmetic, logic, and input / output operations. Instructions may be provided to the processor (120) via the memory (110) or the communication interface (130). For example, the processor (120) may be configured to execute instructions received according to program code stored in a storage device such as the memory (110).
[0042] The communication interface (130) may provide a function for the computer device (100) to communicate with other devices via a network (160). For example, requests, commands, data, files, etc. generated by the processor (120) of the computer device (100) according to program codes stored in a recording device such as a memory (110) may be transmitted to other devices via the network (160) under the control of the communication interface (130). Conversely, signals, commands, data, files, etc. from other devices may be received by the computer device (100) via the communication interface (130) of the computer device (100) via the network (160). Signals, commands, data, etc. received via the communication interface (130) may be transmitted to the processor (120) or the memory (110), and files, etc. may be stored in a storage medium (the aforementioned permanent storage device) that the computer device (100) may further include.
[0043] The communication method is not limited, and may include not only a communication method that utilizes a communication network (e.g., a mobile communication network, a wired Internet, a wireless Internet, a broadcasting network) that the network (160) may include, but also short-range wired / wireless communication between devices. For example, the network (160) may include any one or more of a personal area network (PAN), a local area network (LAN), a campus area network (CAN), a metropolitan area network (MAN), a wide area network (WAN), a broadband network (BBN), the Internet, etc. In addition, the network (160) may include any one or more of a network topology including, but not limited to, a bus network, a star network, a ring network, a mesh network, a star-bus network, a tree, or a hierarchical network.
[0044] The input / output interface (140) may be a means for interfacing with an input / output device (150). For example, the input device may include a device such as a microphone, a keyboard, a camera, or a mouse, and the output device may include a device such as a display or a speaker. As another example, the input / output interface (140) may be a means for interfacing with a device that integrates input and output functions, such as a touchscreen. The input / output device (150) may also be configured as a single device with the computer device (100).
[0045] Furthermore, in other embodiments, the computer device (100) may include fewer or more components than those illustrated in FIG. 1. However, it is not necessary to explicitly illustrate most conventional components. For example, the computer device (100) may be implemented to include at least some of the input / output devices (150) described above, or may further include other components such as a transceiver, a camera, various sensors, a database, and the like.
[0046] Below, a specific embodiment of a pathfinding map-based location recognition technology is described.
[0047] In this specification, a route finding map is a map image created for the purpose of guiding people, and may refer to a map drawing that briefly depicts only some of the main information in an actual environment using various shapes or drawing tools, such as a guide map or a map.
[0048] Humans inherently have the ability to match maps to the real world through their ability to perceive the layout and structure of paths and to identify features like corridors and junctions. Therefore, even if a route-finding map is distorted or inaccurately constructed due to abstraction or simplification of the real world, humans can easily understand it.
[0049] The present invention can provide a location recognition system that recognizes and understands a semantic structure including a spatial structure and layout of a road, thereby estimating a location on a route finding map using only an image without using prior information about the space, a precise map, a positioning sensor, or the like.
[0050] FIG. 2 is a flowchart illustrating an example of a location recognition method that a computer device according to one embodiment of the present invention can perform.
[0051] The computer device (100) according to the present embodiment can provide a desired service to a client through a dedicated application installed on the client or access to a web / mobile site related to the computer device (100). The computer device (100) may be configured with a computer-implemented location recognition system. For example, the location recognition system may be implemented in the form of an independently operating program, or may be configured in the form of an in-app of a specific application so as to be operable on the specific application.
[0052] The processor (120) of the computer device (100) may be implemented as a component for performing the following location recognition method. Depending on the embodiment, the components of the processor (120) may be selectively included or excluded from the processor (120). Furthermore, depending on the embodiment, the components of the processor (120) may be separated or combined to express the functions of the processor (120).
[0053] These processors (120) and components of the processor (120) can control the computer device (100) to perform steps included in the following location recognition method. For example, the processor (120) and components of the processor (120) can be implemented to execute instructions according to the code of the operating system included in the memory (110) and the code of at least one program.
[0054] Here, the components of the processor (120) may be representations of different functions performed by the processor (120) according to instructions provided by the program code stored in the computer device (100).
[0055] The processor (120) can read necessary commands from the memory (110) loaded with commands related to the control of the computer device (100). In this case, the read commands may include commands for controlling the processor (120) to execute steps to be described later.
[0056] The steps included in the location recognition method described below may be performed in a different order than the illustrated order, and some of the steps may be omitted or additional processes may be included.
[0057] Referring to FIG. 2, a location recognition method according to the present embodiment may include a step (S210) of receiving a first-person view (FPV) image, which is a first-person image, from a robot or device as visual information for location recognition for a given pathfinding map, a step (S220) of converting the FPV image into a bird-eye view (BEV) image, which is an image observation representation, and then comparing the BEV image with the given pathfinding map to estimate a location indicated by the FPV image on the pathfinding map, and a step (S230) of displaying the estimated location on the pathfinding map.
[0058] The processor (120) can compare the BEV image with the pathfinding map through a cross-correlation operation. The location recognition framework according to the present embodiment can be trained to compare the BEV image with the pathfinding map in the same latent space. The processor (120) can create an egocentric map from the first-person image and perform a cross-correlation operation using the egocentric map as the query image to search for a location matching the query image on the pathfinding map. In this case, the more information included in the BEV egocentric map, the better the location recognition performance.
[0059] FIG. 2 is a flowchart illustrating an example of a location recognition method that a computer device according to one embodiment of the present invention can perform.
[0060] FIG. 3 illustrates an outline of a pathfinding map-based location recognition system according to one embodiment of the present invention.
[0061] Referring to FIG. 3, the position recognition system (300) according to the present embodiment may include two types of image modules, an FPV image module (Imagine-FPV) (310) and a BEV image module (Imagine-BEV) (320), to help understand the surrounding environment.
[0062] The FPV image module (310) represents how the pathfinding map will appear in the current RGB image in FPV (first person view), and learns which parts of the FPV image will be drawn on the BEV map and which parts will not be drawn.
[0063] The BEV image module (320) expresses how the current surrounding environment will be drawn as a pathfinding map in BEV (bird's eye view).
[0064] In this embodiment, the FPV image module (310) and the BEV image module (320) are attached to a localizer module (330).
[0065] The location recognition module (330) converts the FPV image into an egocentric BEV representation and compares it with a given BEV map, which is a pathfinding map. At this time, the location recognition module (330) learns latent features useful for location recognition with the help of two image modules (310, 320).
[0066] In this embodiment, attaching two image modules (310, 320) to the location recognition module (330) can significantly improve location recognition accuracy, especially in indoor environments. Furthermore, indoor environments often consist of simple, repetitive structures, making it difficult to accurately identify a specific location from a single image. Therefore, the present invention can extend the location recognition method using a particle filter algorithm, which effectively resolves location ambiguity by resampling and weighting locations based on likelihood.
[0067] The present invention provides a novel position recognition system (300) based on an inaccurate pathfinding map that relies on images and wheel odometry data given from a robot or device, and the position recognition system (300) according to the present invention can accurately estimate a position in a large-scale indoor environment without additional knowledge such as GPS signals or satellite images. In addition, the position recognition system (300) according to the present invention can apply two image modules (310, 320) designed to learn effective latent interpretations by expressing from both a first-person view (FPV) and a bird's eye view (BEV). The FPV image module (310) and the BEV image module (320) play an essential role in improving the position recognition accuracy. Furthermore, the position recognition system (300) according to the present invention can efficiently handle position ambiguity by using a particle filter algorithm and support consistent tracking in a dynamic environment.
[0068] Details of the pathfinding map-based location awareness framework are as follows.
[0069] This example addresses the problem of location recognition in unfamiliar, large-scale indoor spaces. A pathfinding map Assume that no photographic knowledge is given. Here, H map The height of the pathfinding map M, Wmap represents the width of the pathfinding map M. At this time, the pathfinding map M can be composed of multiple polygons representing stores or arbitrary areas. In addition, we target a pathfinding map M that has no specific semantic label other than a free space area or a polygon area. This allows for general implementation in various environments without relying on a specific area. Each pixel of the pathfinding map M can be one of three classes: a blank area outside the map, a free space area, or a polygon area inside a polygon.
[0070] In the present invention, a robot exploring its surroundings is considered, and an RGB image I observed at time step t t The purpose is to perform position recognition of a 3DOF pose (x, y, θ) on a pathfinding map M based on . At this time, the robot includes three perspective cameras that can cover a 180° field of view. The robot configuration is exemplary and is not limited thereto, and the number of cameras is not limited and may vary depending on the user. Depending on the embodiment, a robot configuration that uses LiDAR instead of a camera for position recognition or that uses both a camera and LiDAR may be considered. In addition to a camera or LiDAR, any sensor that can be replaced for position recognition can be applied.
[0071] The position recognition system (300) according to the present embodiment can access wheel odometry information of the robot and may include a position recognition module (330), an FPV image module (310), and a BEV image module (320).
[0072] Figure 4 illustrates the detailed configuration of the position recognition module (330), the FPV image module (310), and the BEV image module (320).
[0073] Location recognition module (330)
[0074] The location recognition module (330) converts the FPV image into a self-BEV representation and then compares it with the BEV pathfinding map M. The location recognition module (330) may utilize a convolution-based localization method. Image encoder F img1 Silver FPV Image I t ∈R 3ХHХWХC to the latent feature f t ∈R 3ХHХWХD where H and W represent the height and width of the image, and C and D represent the channel sizes. Then, each pixel feature is embedded into the BEV representation according to the estimated 3D position. Various techniques are already known to estimate the 3D position of each pixel, and for example, an embedding method that utilizes depth information and known camera parameters can be used. Since only RGB information is given, the depth information is estimated from RGB using a known model (e.g., Omnidata). Since Omnidata outputs normalized depth and surface normal, a conversion process to a metric scale is required to generate a consistent BEV representation. At this time, the floor area can be estimated using the estimated surface normal, and the depth scale can be adjusted so that the floor height matches the known camera height. Using the estimated depth information and known camera parameters, f t Each FPV pixel feature can be mapped to a 3D position in . Features corresponding to the same xy location are averaged, and the averaged features are expressed in the BEV representation m t ∈R HegoХWegoХD is recorded in the corresponding grid cell. The above FPV-to-BEV conversion process is shown in F of Fig. 4. embed It corresponds to the process.
[0075] The pathfinding map M is a convolutional neural network F mapProcessed by self-BEV expression m t It is converted to the same potential space as . At this time, the pathfinding map k t ∈R HmapХWmapХD becomes. Then, self-BEV expression m t is a pathfinding map k through a rotated cross-correlation operation (or convolution). t is compared to . The output of the cross-correlation is a heatmap H t ∈R HmapХWmapХR , where R represents the number of rotations, and 360° is divided into R discrete intervals. The position recognition module (330) calculates H, which is the possibility of the camera pose in each grid cell. t are trained to generate .
[0076] The pathfinding map M is provided after being cropped to a fixed size (HmapХWmap), and the robot pose can be any pixel of the cropped map M. In this embodiment, the cropped map M represents a real area of, for example, 64mХ64m. The loss for location recognition is the cross entropy loss L loc =CrossEntropy(H t , H gt ) is. Here, H gt refers to a one-hot distribution that indicates in which grid cell the robot exists.
[0077] The present invention's localization approach involves the robot gradually learning where to look for localization and gaining insight into how the observed area is represented on the BEV map. Furthermore, this understanding is a crucial element in allowing localization from two different perspectives. To enhance this ability, the present invention includes two image modules (310, 320) designed to explicitly guide this process.
[0078] The location recognition system (300) according to the present invention is different from other BEV potential expressions n tIntroduce a separate stream to build F loc m that implicitly learns latent features through t In contrast, n t is explicitly guided to learn features useful for location recognition by additional loss using image modules (310, 320). n t is an image encoder F img Another FPV latent feature g encoded by 2 t ∈R 3ХHХWХD is generated from n t is m t Connected with μ t becomes the final BEV representation μ t is used for convolution-based location recognition in a location recognition system (300).
[0079] Image module (310, 320) BEV latent representation n t Let's explain what effect it has.
[0080] FPV Image Module (310)
[0081] U-Net encoder F in FPV image module (310) FPV is a FPV potential feature g t Take and estimate the map mask v for each image t ∈R 3ХHХW Outputs the FPV image module (310) which is trained to estimate which part of the observed image will be drawn on the pathfinding map. GT (ground truth) virtual map mask v gtThe process of creating is as shown in Fig. 5. First, a local egocentric map is cropped for each camera in the pathfinding map. Then, starting from the GT camera pose, a ray is cast on each map to obtain only the observable part. As the ray is cast, the area before the line is hit is labeled as 'Free Space'. When the line of a polygon is met, the meeting point is labeled as 'Polygon'. The subsequent part of the ray beyond this point is labeled as 'Unknown'. In this embodiment, the remaining unobservable area is masked as 'Unknown'. The labeled egocentric map Then, using the known camera height and unique parameters, the self-map We render the 'free space' of each FPV image and assume that they are on the ground. The rendered portion is v gt , and the value of the FPV image area displayed on the map is 1, and the rest is 0. In addition, obstacles in the rendered free space area are filtered using the estimated surface normal. Pixels that do not share a similar orientation to the floor are filtered by v gt is represented as 0 in the convolutional neural network F FPV is trained by L1 loss v gt : L FPV = L1(v t , v gt ) and v t Make v gt The area considered is different from simply classifying the floor in the image. In Fig. 5, v gt It can be seen that the interior area of the store is not highlighted. The estimated vt is the building n t It is used to mask unimportant areas in FPV latent feature g t is v t Only n free space pixels are multiplied and estimated t =Fembed (g t ×v t ) is embedded in.
[0082] BEV image module (320)
[0083] Neural network F in BEV image module (320) BEV is BEV latent representation n t From your local map e t is trained to predict. During this process, the network can understand its surroundings, in particular by representing how the observed area appears from the BEV perspective within the context of the pathfinding map M. For each camera pose, a local self-map Combined to form a single 180° magnetic map for the central camera e gt becomes e gt has three classes corresponding to 'Polygon', 'Free space', and 'Unknown'. e t The loss for is cross entropy loss L BEV =CrossEntropy(e t , e gt ) is formulated as the estimated magnetic map e t is not used during inference time, but serves as an auxiliary task in the learning process of the BEV image module (320), encouraging better feature extraction for scene understanding and location recognition.
[0084] learning
[0085] The pathfinding map-based location recognition framework according to the present invention is trained using three loss functions corresponding to the FPV image module (310), the BEV image module (320), and the location recognition module (330), and the total loss is as shown in mathematical expression 1.
[0086] [Mathematical Formula 1]
[0087]
[0088] Using the above loss, the neural network (F) of the pathfinding map-based location recognition framework img1 , F img2 , F FPV , F BEV , F map ) can be studied jointly.
[0089] The FPV image module (310) and the BEV image module (320) are the final BEV representation μ used in the position recognition module (330). t contributes to learning. As a result, L FPV Wow L BEV If you learn L loc It helps to improve the performance of overall location awareness by affecting the .
[0090] Particle filter algorithm
[0091] This embodiment focuses on large-scale indoor environments, considering repetitive patterns and numerous dynamic obstacles. Maps for these environments may not always be accurate, resulting in the positioning system (300) sometimes selecting incorrect locations that are similar to the correct answer. To address this issue, a particle filter algorithm can be applied to combine temporal information to improve position estimation results.
[0092] Figure 6 illustrates an example flowchart of a particle filter algorithm.
[0093] The processor (120) can perform sequential inference for location recognition through a particle filter algorithm.
[0094] At the beginning of the episode, particles are scattered on the map and wheel odometry information is used to update each particle pose, each particle has a pose (h,ω,r), which is H t ∈R Hmap×Wmap×R corresponds to the grid of H tThe values from become the particle weights. The weights are continuously multiplied as the particles move, and the particles are resampled according to their weights. In most situations, particles often cluster in similar areas, which is common in particle filter-based localization. These clusters maintain similar scores until they receive important information that can clarify the pose. Therefore, selecting particles based on the highest weights can lead to inaccurate estimates due to possible outliers or inaccurate particle movements. To achieve consistent localization over time, we use K-means clustering to group particles and treat these clusters as potential pose candidates. For example, we use five clusters, and the sum of the particle weights in each cluster determines the score. The center of the cluster with the highest sum of weights is the representative estimated pose.
[0095] Figure 7 illustrates an example of a visualization of a location awareness episode. As shown in Figure 7, the estimated e of the robot's location over time t Wow v t can be displayed on the route finding map.
[0096] The present embodiments can estimate a location on a pathfinding map based on a pathfinding map and a first-person image, and at this time, the location recognition performance can be improved through a framework including an image module that predicts where an important part on the image is and how the surrounding map will be drawn based on the first-person image. Furthermore, the location estimation result can be improved by applying a particle filter algorithm to combine information over time, and the location recognition performance can be improved by additionally utilizing semantic information included in the map through map parsing. The processor (120) can parse the pathfinding map and the first-person image given as input into 2D, obtain information on landmarks (e.g., store names, stairs, trees, landmarks, etc.) through the parsing result, and then search for a location on the pathfinding map by comparing the information between the pathfinding map and the first-person image.
[0097] According to embodiments of the present invention, the current location can be recognized on a distorted map that differs from the actual environment through a location recognition technology that recognizes and understands the semantic structure according to the spatial structure and layout of the road, and effective location estimation is possible, especially in indoor environments where there is no prior spatial information and GPS signals are unavailable. Furthermore, according to embodiments of the present invention, by applying a particle filter algorithm for probabilistically utilizing continuous location estimation results, weighting of locations according to probability can be improved, thereby improving location accuracy.
[0098] The devices described above may be implemented as hardware components, software components, and / or a combination of hardware components and software components. For example, the devices and components described in the embodiments may be implemented using one or more general-purpose computers or special-purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding to them. The processing device may execute an operating system (OS) and one or more software applications running on the operating system. The processing device may also access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing device is sometimes described as being used alone; however, one of ordinary skill in the art will recognize that the processing device may include multiple processing elements and / or multiple types of processing elements. For example, a processing unit may include multiple processors, or a processor and a controller. Other processing configurations, such as parallel processors, are also possible.
[0099] Software may include a computer program, code, instructions, or a combination of one or more of these, which may configure a processing device to perform a desired operation or may independently or collectively command the processing device. The software and / or data may be embodied in any type of machine, component, physical device, computer storage medium, or device for interpretation by the processing device or for providing instructions or data to the processing device. The software may also be distributed over networked computer systems and stored or executed in a distributed manner. The software and data may be stored on one or more computer-readable recording media.
[0100] The method according to the embodiment may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. In this case, the medium may be one that continuously stores a computer-executable program or one that temporarily stores it for execution or download. In addition, the medium may be various recording means or storage means in the form of a single or multiple hardware combinations, and is not limited to a medium directly connected to a computer system, but may also be distributed over a network. Examples of the medium may include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and those configured to store program commands, including ROM, RAM, and flash memory. In addition, examples of other media may include recording media or storage media managed by app stores that distribute applications, sites that supply or distribute various software, servers, etc.
[0101] Although the embodiments described above have been described by way of limited examples and drawings, those skilled in the art will appreciate that various modifications and variations can be made based on the above teachings. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components of the described systems, structures, devices, circuits, etc. are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.
[0102] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims described below.
Claims
1. In a localization method executed on a computer device, The computer device comprises at least one processor configured to execute computer-readable instructions contained in a memory, The above location recognition method is, A step of receiving a first-person view image as visual information for location recognition in an indoor environment by at least one processor; and A step of estimating a location matching the first-person image in a given wayfinding map for the indoor environment by using the first-person image, by the at least one processor. A location recognition method comprising:
2. In paragraph 1, The above route finding map is a guide map or map for the above indoor environment. The above estimating steps are: Estimating a location on the pathfinding map using the first-person image in an environment where no prior information about the space is given and the use of a global positioning system (GPS) signal is not possible. A location recognition method characterized by:
3. In paragraph 1, The above estimating steps are: A step of converting the first-person image into a BEV (bird-eye view) image through a location recognition module and then comparing the BEV image with the pathfinding map through a cross-correlation operation. A location recognition method comprising:
4. In paragraph 3, The above location recognition module is learned through a loss function between the heatmap, which is the output of the cross-correlation operation, and the GT (ground truth) heatmap. A location recognition method characterized by:
5. In paragraph 1, The above estimating steps are: A step of converting the above first-person image into a BEV image, which is an egocentric bird-eye view representation; A step of converting the above pathfinding map into a BEV pathfinding map by converting it into the same latent space as the BEV image; and A step of comparing the BEV image with the BEV pathfinding map and searching for a location on the pathfinding map. A location recognition method comprising:
6. In paragraph 5, The steps for converting to the above BEV image are: A step of embedding a first BEV representation by processing the first-person image as a first latent feature of a first image encoder; A step of processing the first-person image as a second latent feature of a second image encoder to embed a second BEV representation; and A step of creating the BEV image, which is the final BEV expression, by connecting the first BEV expression and the second BEV expression. A location recognition method comprising:
7. In paragraph 5, The steps for converting to the above BEV image are: Step of embedding the BEV representation by processing the above first-person image as a latent feature of the image encoder. Including, The above latent features are learned through two types of image modules: the FPV image module (Imagine-FPV) and the BEV image module (Imagine-BEV). A location recognition method characterized by:
8. In paragraph 7, The above FPV image module is trained through a loss function between the estimated map mask output through the FPV latent feature for the first-person image and the ground truth (GT) map mask that labels the observable area and masks the remaining area. A location recognition method characterized by:
9. In paragraph 7, The above BEV image module is learned through a loss function between the egocentric map predicted from the BEV latent representation for the first-person image and the GT egocentric map according to each camera pose. A location recognition method characterized by:
10. In paragraph 5, The above estimating steps are: A step of performing sequential inference for location recognition by assigning weights to the searched locations through a particle filter algorithm. A location recognition method comprising:
11. In paragraph 10, The steps performed above are: A step of grouping the estimated particles on the pathfinding map using K-means clustering and estimating the center of the cluster with the highest weight sum of particles in each cluster as the representative pose. A location recognition method comprising:
12. In paragraph 10, The steps performed above are: A step of updating the pose of the estimated particle on the pathfinding map using wheel odometry data of the robot or device providing the first-person image. A location recognition method comprising:
13. A non-transitory computer-readable recording medium storing a computer program for executing the location recognition method of paragraph 1 on the computer device.
14. In computer devices, At least one processor configured to execute computer-readable instructions contained in memory Including, At least one processor of the above, A process of receiving a first-person image as visual information for location recognition in an indoor environment; and A process of estimating a location matching the first-person image on a given pathfinding map for the indoor environment using the first-person image. A computer device that processes data.
15. In paragraph 14, The above route finding map is a guide map or map for the above indoor environment. At least one processor of the above, Estimating a location on the pathfinding map using the first-person image in an environment where no prior information about the space is given and GPS signals are unavailable. A computer device characterized by:
16. In paragraph 14, At least one processor of the above, After converting the first-person image into a BEV image through a location recognition module, the BEV image is compared with the pathfinding map through a cross-correlation operation. The above location recognition module is learned through a loss function between the heatmap, which is the output of the cross-correlation operation, and the GT heatmap. A location recognition method characterized by:
17. In paragraph 14, At least one processor of the above, Convert the above first-person image into a BEV image, which is a self-BEV representation, Convert the above pathfinding map into a BEV pathfinding map by converting it into the same latent space as the above BEV image, Comparing the above BEV image with the above BEV pathfinding map to search for a location on the above pathfinding map. A computer device characterized by:
18. In paragraph 17, At least one processor of the above, Embed the BEV representation by processing the above first-person image as a latent feature of the image encoder, The above latent features are learned through two types of image modules: the FPV image module and the BEV image module. A computer device characterized by:
19. In Article 18, The above FPV image module is trained through a loss function between the estimated map mask output through the FPV latent feature for the first-person image and the GT map mask that labels the observable area and masks the remaining area. The above BEV image module is learned through a loss function between the self-map predicted from the BEV latent representation for the first-person image and the GT self-map according to each camera pose. A computer device characterized by:
20. In paragraph 17, At least one processor of the above, By performing sequential inference for location recognition by giving weight to the searched location through particle filter algorithm, Using K-means clustering to group the estimated particles on the pathfinding map, and estimating the center of the cluster with the highest weighted sum of particles in each cluster as the representative pose. A computer device characterized by:
Citation Information
Patent Citations
Method and apparatus of pose estimation in a mobile robot based on particle filter
KR100877071B1
Navigation method and target position estimation apparatus for mobile robot
KR101722339B1
Method for detecting indoor position based on images and mobile terminal using same
KR1020140058861A
Burner apparatus for pulverized coal combustion
KR102622029B1
Cited By
Cognition-enhanced urban space unmanned aerial vehicle visual target searching method
CN122192337A