A method and system for characterizing visual identities used for location identification to assist in precision and long-term map management

The method enhances SLAM algorithms by generating a data field with observation comparison scales and reliability checks to address computational and environmental challenges, improving navigation reliability and accuracy in dynamic environments.

JP2025522884APending Publication Date: 2025-07-17OPTERAN TECH LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025500136
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-07-04
Filing Date
2023-07-04
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing SLAM algorithms face challenges such as high computational complexity, susceptibility to sensor noise, environmental changes, and memory consumption, leading to unreliable map generation and navigation in dynamic environments.

Method used

A method involving generating a data field with observation comparison scale data, performing reliability checks, and managing digital maps by maintaining, updating, or deleting information based on reliability scales, using reduced-size sensor data and incorporating action vectors to enhance navigation accuracy.

Benefits of technology

Improves navigation reliability and accuracy by reducing computational burden and adapting to environmental changes, while minimizing memory usage and maintaining map integrity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025522884000001_ABST
    Figure 2025522884000001_ABST
Patent Text Reader

Abstract

A method for navigating an agent in a multi-dimensional space and a computer program are provided. The method includes generating a data field that spatially corresponds to the multi-dimensional space in which the agent exists. The data field incorporates observation comparison scale data regarding saved observations, and the saved observations include second sensor data that describes the environment of a first location within the multi-dimensional space. The step of generating the data field includes obtaining a plurality of current observations made by the agent at different corresponding positions within the multi-dimensional space. The current observations include first sensor data that describes the environment of each of the different corresponding positions within the multi-dimensional space. Each of the current observations is compared with the saved observations to obtain corresponding observation comparison scale data regarding different corresponding positions within the multi-dimensional space. Then, the observation comparison scale data is incorporated into the data field according to different corresponding positions within the multi-dimensional space. The method further includes determining a reliability measurement associated with the saved observations by performing a reliability check on the data device field, and managing a digital map of the multi-dimensional space by maintaining, updating, or deleting information of the digital map regarding the saved observations based on the reliability scale.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a computer-implemented method, apparatus, system, and computer program for map management, and more particularly to improving the reliability and accuracy of maps for the purpose of navigating in a two-dimensional or three-dimensional space.

Background Art

[0002] Simultaneous Localization and Mapping (SLAM) is the problem and process of computationally constructing or updating a map of an unknown environment while simultaneously tracking the position of an agent within that environment. SLAM algorithms are based on concepts in computational geometry and computer vision and are used in robotic navigation, robotic mapping, and odometry measurements in virtual or augmented reality.

[0003] SLAM algorithms are also used in the control of autonomous vehicles to map the unknown environment around the vehicle. The resulting map information can be used to perform tasks such as path planning and obstacle avoidance.

[0004] Visual SLAM uses an approach consisting of a front end that processes images from a camera to perform "feature extraction". In this feature extraction process, small regions of an image that meet certain criteria that are visually distinguishable, such as high local spatial frequency or contrast, are detected. These regions can then be saved for the purpose of finding the same regions in subsequent camera images taken at a slight interval with the same camera. Next, a camera model that takes into account the distortion of the camera lens is applied to map the positions of these features in the direction of the real world. These are provided to the back end of an algorithm called a bundle adjuster. The bundle adjuster compares the predicted positions of features in 3D space based on previous camera images with the new directions from the current camera image and simultaneously determines the relative positions of the features with respect to each other in 3D space and the position of the camera with respect to the features. By repeating this process, the bundle adjuster tracks the position of the moving camera and the position of the feature map over time.

[0005] This approach has several drawbacks that limit its effectiveness in real-world robot applications. First, both feature extraction and bundle adjustment are computationally very expensive processes. Feature extraction can be easily parallelized for efficiency improvements, but bundle adjustment is a strictly sequential process.

[0006] Second, in bundle adjustment, it is necessary for the new feature directions and predicted feature positions to converge to a stable solution, which is greatly affected by sensor noise and environmental changes such as wind blowing the leaves of a thicket. If a feature is erroneously detected at the wrong location on the camera image, the bundle adjustment cannot converge to a stable solution, and the system will get "lost" and be unable to identify its position.

[0007] To address the convergence obstacles of the bundle adjuster, two approaches are adopted in the state-of-the-art. One is to mitigate the obstacles, and the other is to enable the system to continue by degrading the performance when convergence fails.

[0008] The first approach is outlier removal. A set of constraints is used to evaluate the likelihood that a feature, when found, is the same as one previously found. For example, if a feature moves from one side of the camera to the opposite side, it is unlikely to be the same, barring very high camera movement speed, and is rejected as a match. If it is close to the previous location, it is likely to be the same feature and is accepted. The accepted features are passed to the bundle adjuster, but the rejected features are not. Outlier removal is another very computationally expensive process and can be the largest computational component in some SLMA systems.

[0009] The second approach is visual-inertial odometry. This is a system that uses the features identified by the feature extractor in a different way. Using a second motion information source, the measured linear acceleration from an inertial sensor, and a tracking function, it predicts the velocity rather than the position of the camera. This velocity measurement can be integrated over time and used in combination with a bundle adjuster to correct for convergence failures. However, when integrating VIO, position errors accumulate over time, and since this approach to SLMA relies on metric measurements of position, problems occur as the integration time lengthens. Since the same features are used in VIO and bundle adjustment, feature extraction becomes a single point of failure for the entire system.

[0010] Another weakness of the feature extraction front-end is that due to the size of the features, the dimensionality of the data for a single feature is relatively low. Therefore, especially in aliased environments such as hotel corridors where identical doors are evenly placed, multiple parts of the image are likely to match the features. In such an environment, the bundle adjuster may converge to the wrong position.

[0011] Loop closure is another weakness of the state-of-the-art. When a moving camera crosses a complete loop and returns to the same location from different directions, the system needs to recognize that the camera has returned to the same location. Since this system is metric, it also needs to confirm that the map matches the same position at the same location. Due to the errors accumulated in the bundle adjustment process, this rarely holds true, and there is usually a metric error between the representations of the positions at both ends of the loop. Next, the system needs to go back and adjust all the features stored throughout the loop to correct and remove this error. This loop closure process is computationally expensive and prone to errors.

[0012] The computational complexity of algorithms that can address such environments and problems is extremely high, and the platforms on which they can be introduced are limited. Furthermore, the memory consumption of the generated maps is very large, ranging from several hundred megabytes to several gigabytes. Additionally, due to small changes in the environment, such as displacements of the characteristic parts of the environment, the reliability of the generated maps may decrease. These changes need to be considered by steps of adding or deleting points to the map, which is a cumbersome process when using SLAM algorithms.

[0013] It is understood that there is a need for a method and system for managing maps and improving their reliability by using alternative means to SLAM.

Summary of the Invention

[0014] This summary is provided to introduce, in a simplified form, some concepts that will be described in more detail in the detailed description. This summary is not intended to identify the main or essential features of the subject matter of the claims, nor is it intended to be used to determine the scope of the subject matter of the claims. Modifications and alternative features that facilitate the implementation of the invention and / or achieve substantially similar technical effects are considered to fall within the scope of the invention disclosed herein.

[0015] In a first aspect, the present disclosure provides a computer-implemented method for navigating an agent in a multi-dimensional space. The method includes generating a data field that spatially corresponds to the multi-dimensional space in which the agent exists. The data field incorporates observation comparison scale data regarding saved observations, and the saved observations include second sensor data that describes the environment of a first location in the multi-dimensional space. The step of generating the data field includes obtaining a plurality of current observations made by the agent at different corresponding positions in the multi-dimensional space, where the current observations include first sensor data that describes the environment of each of the different corresponding positions in the multi-dimensional space. The step further includes comparing each of the current observations with the saved observations to obtain respective observation comparison scale data for different corresponding positions in the multi-dimensional space, and incorporating the observation comparison scale data into the data field according to different corresponding positions in the multi-dimensional space. The method further includes determining a reliability scale associated with the saved observations by performing a reliability check on the data field, and managing a digital map of the multi-dimensional space by maintaining, updating, or deleting information of the digital map regarding the saved observations based on the reliability scale.

[0016] The agent may be a real device or system such as a vehicle or a robot, or may be virtual.

[0017] The first sensor data and the second sensor data may be first and second image data.

[0018] Each of the first and second sensor data is arranged in a first vector and a second vector respectively extracted from the first and second image data.

[0019] The method further includes steps of obtaining first sensor data and second sensor data. The step of obtaining first sensor data includes obtaining original first sensor data of the agent's environment at respective different corresponding positions in a multi-dimensional space, and processing the original first sensor data to reduce the size of the original first sensor data, obtaining the first sensor data such that the first sensor data is in a reduced-size format relative to the original first sensor data. The step of obtaining second sensor data includes obtaining original second sensor data of the agent's environment at a first location, processing the original second sensor data to reduce the size of the original second sensor data, obtaining the second sensor data such that the second sensor data is in a reduced-size format relative to the original second sensor data, and storing the second sensor data in the reduced-size format.

[0020] Processing the original first sensor data and the original second sensor data may include applying one or more filters and / or masks to the original first sensor data and the original second sensor data to reduce the size of each dimension of the original first sensor data and the original second sensor data respectively.

[0021] The method may further include the step of navigating the agent according to a digital map to a saved observation.

[0022] The digital map may include a series of saved observations and information regarding displacement vectors linking the series of observations. The step of navigating includes navigating the agent to one of the series of saved observations using the digital map information.

[0023] The data field includes an action vector field, the action vector field includes a plurality of action vectors, the observation comparison scale data are the plurality of action vectors, and each of the plurality of action vectors starts from one of different corresponding first locations within the vector field and is directed in the estimated direction of the first location within the vector field.

[0024] The step of obtaining corresponding observation comparison scale data for different corresponding positions in a multi-dimensional space by comparing each current observation with a saved observation includes, for each current observation, obtaining a plurality of first sub-regions of first sensor data, where each first sub-region describes a corresponding first part of the environment around the agent at one of the different corresponding positions, and each first part is associated with a corresponding first direction from one of the different corresponding positions, and obtaining a plurality of second sub-regions of second sensor data, where each second sub-region describes a respective second part of the environment around the first location, and each second part is associated with a respective second direction from the first location, and for each second sub-region, comparing the second sub-region with each first sub-region using a similarity comparison scale to determine the first sub-region that is most similar to the second sub-region, and determining the relative rotation between the second direction associated with the second sub-region and the first direction associated with the most similar first sub-region. The method further includes aggregating the relative rotations of the plurality of second sub-regions to obtain an action vector, the action vector indicating the estimated direction from one of the different corresponding positions to the first location, and the observation comparison scale includes the action vector.

[0025] The step of determining a reliability measure associated with a saved observation by performing a reliability check on a data field is a step of obtaining a model action vector field that spatially corresponds to an action vector field, the model action vector field including a plurality of model action vectors, each model action vector being directly induced to a first location and the model vector field being adapted to radially converge to the first location, the step of obtaining the model action vector field, the step of comparing a plurality of action vectors of the action vector field with spatially corresponding model action vectors of the model action vector field to determine an angular deviation between each action vector and the corresponding model action vector, and the step of aggregating or averaging the angular deviations between the plurality of action vectors and the plurality of model action vectors to determine a first scalar measure, the reliability measure including the first scalar measure.

[0026] The step of determining a reliability measure associated with a saved observation by performing a reliability check on a data field includes the step of determining a first number of different respective positions where an action vector cannot be obtained, the step of determining a second number of different respective positions where an action vector can be obtained, and the step of dividing the first number by the second number to determine a second scalar measure, the reliability measure including the second scalar measure.

[0027] Based on a reliability measure, the step of managing a digital map of multi-dimensional space by maintaining, updating, or deleting information of the digital map regarding stored observations may include a step of comparing the reliability measure with one or more thresholds to determine whether to update, delete, or maintain the information of the digital map. The step of managing the map is important to enable more accurate or better navigation within the space. The step of deleting information of the digital map may include a step of deleting information regarding stored observations in the digital map and clarifying the deleted information, and / or the step of updating information in the digital map may include a step of obtaining replacement sensor data at a first location in the multi-dimensional space to form replacement stored observations and a step of regenerating a data field according to the replacement stored observations.

[0028] The step of adjusting the map to clarify the deleted information may include, for example, when the deleted information corresponds to intermediate nodes, a step of adjusting displacements and other connections such as odometric data between adjacent nodes (stored observations), as these adjacent nodes may need to be connected to each other.

[0029] The method may further include a step of aggregating a first scalar measure and a second scalar measure such that the reliability measure includes both the first and second scalar measures.

[0030] The data field includes a similarity scalar field, which includes a plurality of similarity magnitude values indicating the similarity between stored observations and current corresponding observations at respective different corresponding positions, and the observation comparison scale data are the plurality of similarity magnitude values.

[0031] The step of obtaining corresponding observation comparison scale data for different corresponding positions in a multi-dimensional space by comparing each of the current observations with the stored observations may include the step of comparing each of the current observations with the stored observations using a similarity comparison scale and determining a similarity value for each of the different corresponding positions, and the observation comparison scale data includes a similarity magnitude value.

[0032] The current observation and the stored observation may be arranged as vectors, and the similarity comparison scale is the inner product of these vectors.

[0033] The step of determining a reliability scale associated with the stored observation by performing a reliability check on the data field may include the step of determining the maximum similarity magnitude value among a plurality of similarity magnitude values of the similarity scalar field, and the step of dividing the maximum similarity magnitude value by the average of the plurality of similarity magnitude values to determine a third scalar value, and the reliability scale includes the third scalar scale.

[0034] The method may further include generating a parameterized model of the similarity scalar field that approximates the similarity scalar field.

[0035] The step of determining the reliability scale may be performed with respect to the parameterized model.

[0036] The step of determining a reliability scale associated with the stored observation by performing a reliability check on the data field may include the step of determining the variability between the similarity magnitude value and the model value of the parameterized model, and the step of obtaining a fourth scalar value based on the variability, and the reliability scale includes the fourth scalar scale.

[0037] The step of managing a digital map by maintaining, updating, or deleting information related to the stored observation based on the reliability scale may include the step of comparing the reliability scale with one or more thresholds to determine whether to update, delete, or maintain the information in the digital map.

[0038] The method may include the step of aggregating a third scalar measure and a fourth scalar measure such that the reliability measure includes both the first and second scalar measures.

[0039] The method may include the step of aggregating the first scalar measure, the second scalar measure, the third scalar measure, and the fourth scalar measure such that the reliability measure includes all of the first, second, third, and fourth scalar measures.

[0040] According to a second aspect of the present disclosure, a computing device or system including a processor and a memory is provided, and instructions are stored in the memory, and when executed by the processor, the instructions enable the computing device to execute the method of the first aspect.

[0041] According to a third aspect of the present disclosure, a computer program is provided that, when executed by a processor, causes the processor to execute the method of the first aspect.

[0042] According to a fourth aspect of the present disclosure, there is provided a computer-implemented method for determining the position of an agent in a multi-dimensional space. The method includes obtaining a similarity field that spatially corresponds to the multi-dimensional space in which the agent exists, the similarity field including a plurality of similarity magnitude values indicating the similarity between stored observations corresponding to a first location in the multi-dimensional space and current corresponding observations respectively corresponding to one of a plurality of different corresponding locations in the multi-dimensional space; obtaining, by the agent, a new observation at the current position of the agent in the multi-dimensional space; comparing the new observation with the stored observations to obtain a new similarity magnitude value; comparing a new location scale based on the new similarity magnitude value with a field location scale based on the similarity magnitude values of the similarity field to identify the most likely or matching field location scale of the new location scale; and identifying one or more possible positions of the agent in the multi-dimensional space based on one or more field locations in the similarity field corresponding to the most likely or matching field location scale.

[0043] The new location scale and the field location scale can combine one or more different scales or functions based on the similarity field and / or other modalities. In the following description, the new location scale and the field location scale may be based on the first to fourth "modes".

[0044] The new location scale can include a new similarity magnitude value, the field location scale is each similarity magnitude value of the similarity magnitude value field, and the most likely or matching field location scale is the matching similarity magnitude value of the similarity magnitude value field, or a range of similarity magnitude values of the similarity magnitude value field, which range includes the new similarity magnitude value. Also, the step of identifying one or more possible positions of an agent in a multi-dimensional space can include obtaining one or more field locations of the matching similarity magnitude value of the similarity magnitude value field, or a range of similarity magnitude values of the similarity magnitude value field, and determining one or more positions of the agent in the multi-dimensional space according to the field locations.

[0045] The method can further include obtaining saved observations, where the saved observations include second sensor data describing the environment of a first location in a multi-dimensional space, obtaining a plurality of current observations made by an agent at different corresponding positions in the multi-dimensional space, where the current observations include first sensor data describing the environment of each of the different corresponding positions in the multi-dimensional space, generating a similarity field, comparing each of the current observations with the saved observations to obtain a similarity comparison measure using a similarity comparison scale, determining a similarity magnitude value for each of the different corresponding positions in the multi-dimensional space, and incorporating the similarity magnitude values into the similarity field according to the different corresponding positions in the multi-dimensional space.

[0046] The current observations and the saved observations may be arranged as vectors, and the similarity comparison scale is the inner product of these vectors.

[0047] This method includes the step of obtaining a second similarity field that spatially corresponds to the multi-dimensional space where the agent exists, the second similarity field including second stored observations corresponding to second locations in the multi-dimensional space and second plurality of similarity magnitude values indicating similarities between each current observation respectively corresponding to one of a plurality of different corresponding positions within the multi-dimensional space; and the step of comparing a new observation with the second stored observations to obtain a second new similarity magnitude value, the new location scale including a ratio between the new similarity magnitude value and the second new similarity magnitude value, and the field location scale including a plurality of field ratios between the plurality of similarity magnitude values and the second plurality of similarity magnitude values, each field ratio of the plurality of field ratios being determined for different positions within the similarity field, whereby the most likely or matching field location scale for the new location scale includes one or more field ratios; and the step of identifying one or more positions of the agent within the multi-dimensional space based on one or more field locations within the similarity field corresponding to the most likely or matching field location scale includes the step of obtaining one or more field locations of one or more matching field ratios within the similarity field and the step of determining one or more positions of the agent within the multi-dimensional space according to the field locations.

[0048] The second stored observations and the second similarity field may be used without a ratio to narrow down one or more possible positions of the agent within the space. This is because the position needs to match both the stored observations and the second stored observations. This may also apply to two or more stored observations. Using three stored observations improves the accuracy of the method and reduces the number of possible positions of the agent. This is because three matches are required to register the position as a possible position of the agent.

[0049] The new location metric can include both the ratio of the new similarity magnitude value and the second new similarity magnitude value, the field location metric includes the similarity magnitude value of the similarity magnitude value field and each of a plurality of field ratios, and the most likely or matching field location metric is the matching similarity magnitude value of the similarity magnitude value field, or a range of the similarity magnitude values of the similarity magnitude value field, which range includes the new similarity magnitude value, and both of one or more matching field ratios, whereby the matching similarity magnitude value and the matching field ratio are associated with the same one or more positions in the multi-dimensional space.

[0050] Using both the ratio and the similarity magnitude value improves accuracy and narrows down the possible positions of the agent in the space. This is because matching requires not just any one of these matchings, but both ratio matching and similarity magnitude value matching.

[0051] The method may further include generating a parameterized model of the similarity field before comparing the new location metric based on the new similarity magnitude value with the field location metric based on the similarity magnitude value of the similarity field, whereby the comparing step is performed with respect to a model field location metric based on the model similarity magnitude value of the model similarity field.

[0052] The method can include obtaining an action vector field that spatially corresponds to the multi-dimensional space in which the agent exists, the action vector field including a plurality of action vectors indicating directions from a plurality of different corresponding positions in the multi-dimensional space toward the first location of the saved observation, and the method can further include discounting one or more of the one or more possible positions of the agent based on the directions of one or more action vectors associated with the field location.

[0053] The action vector field can be obtained in a manner similar to the first aspect described above.

[0054] The method may further include a step of generating an output field, where the output field is based on the one or more identified positions of the agent in the multi-dimensional space, the output field has the same dimension as the similarity field, and represents the probability distribution of the one or more identified positions of the agent in the multi-dimensional space. The output field may sometimes be referred to as an array or field of possible positions.

[0055] The step of discounting one or more of the possible positions of the agent based on the directions of one or more action vectors associated with one or more field locations may include the step of adjusting the probability distribution of the one or more identified positions of the agent in the output field based on the directions of one or more action vectors associated with one or more field locations.

[0056] The probability distribution of the output field may be based on one or more possible positions of the agent determined according to the similarity field values, ratios, and action vectors described above. Each of the possible positions of the agent determined from each of these modalities or modes may be mapped to the probability distribution of the possible positions of the agent according to each of the similarity field values, ratios, and action vectors. These arrays are summed element-wise and optionally normalized to form the output field.

[0057] This method includes a step of obtaining a plurality of spatially corresponding similarity fields within a multi-dimensional space in which an agent exists, where the plurality of similarity fields include respective sets of similarity magnitude values indicating the similarity between a plurality of respective saved observations corresponding to a plurality of respective first locations within the multi-dimensional space and each current observation corresponding to one of a plurality of different corresponding positions within the multi-dimensional space; a step of comparing a new observation with each of the plurality of saved observations to obtain a plurality of new similarity magnitude values for each of the plurality of saved observations; a step of comparing a plurality of new location scales based on the plurality of new similarity magnitude values with respective field location scales based on the sets of similarity magnitude values of the plurality of similarity fields to identify the most likely or matching field location scale for each of the plurality of new location scales; and a step of identifying one or more possible positions of the agent within the multi-dimensional space based on one or more field locations within a similarity field corresponding to the most likely or matching field location scale for each of the plurality of new position measurements.

[0058] Accordingly, one or each of a similarity magnitude value, a ratio, an action vector, and odometric information, or a combination thereof, can be used with respect to a plurality of saved observations, and the outputs can be combined to determine the possible position of the agent from a plurality of different fields and related saved observations.

[0059] Taking the similarity field as an example, the possible positions of the agent can be determined using three saved observations and three related similarity fields. These possible positions of the agent can be combined based on a probability distribution for each similarity field.

[0060] The method may further include receiving odometry data regarding the movement of an agent, determining an estimated distance the agent has moved from a previous position in a multi-dimensional space, and modifying the probability of one or more identified positions of the agent in the multi-dimensional space based on the estimated distance. The modified output field may be referred to as an odom update array.

[0061] The method may further include modifying the probability of one or more identified positions of the agent in the multi-dimensional space based on a reliability measure according to a first aspect.

[0062] The method may further include navigating the agent based on one or more identified positions of the agent.

[0063] According to a fifth aspect of the present disclosure, a computer-implemented method is provided for navigating an agent and determining the reliability of its navigation path. The method may be executed independently or in combination with the first aspect and / or the fourth aspect.

[0064] The method includes navigating from one or more specified positions of the agent in a multi-dimensional space along a first path from the specified one or more positions to a first location to a first location corresponding to a saved observation, determining the length of the first path and calculating a metric of edge traversal reliability, where the length of the path is inversely proportional to the metric of edge traversal reliability, comparing the metric of edge traversal reliability with a historical metric of edge traversal reliability calculated from the average length of the previously traveled path from one or more positions of the agent in the multi-dimensional space to the first location, and determining that the first path is not reliable if the metric of edge traversal reliability is lower than a threshold value than the historical metric.

[0065] When it is determined that the first path is not reliable, the method may include the steps of performing at least one of the following, recording new saved observations along the first path between one or more positions of the agent and the first location in the multi-dimensional space, excluding the first path from future navigation, optimizing the first path by retrying the first path one or more times with different path parameters to maximize the metric of the reliability of the edge traversal, and warning the operator of the agent.

[0066] According to a sixth aspect of the present disclosure, a computing device or system including a processor and a memory is provided, instructions are stored in the memory, and when executed by the processor, the computing device executes the method of the fourth or fifth aspect above.

[0067] According to a seventh aspect of the present disclosure, a computer program is provided that, when executed by a processor, causes the processor to execute the method of the fourth or fifth aspect above.

[0068] The methods described herein may be in the form of a computer program comprising computer program code means adapted to perform all the steps of any of the methods described herein when the program is run on a computer, and when the computer program may be embodied on a computer-readable medium, it may be executed by software in machine-readable form on a tangible storage medium. Examples of tangible (or non-transitory) storage media include disks, thumb drives, memory cards, etc., and do not include propagated signals. The software is suitable for execution on a parallel or serial processor and the steps of the method may be performed in any suitable order or simultaneously.

[0069] In this use, it is recognized that firmware and software are goods of individually tradable value. This is intended to include software that is computed or controlled on "dams" or standard hardware to perform desired functions. It is also intended to include software that "describes" or defines the configuration of hardware, such as HDL (Hardware Description Language) software used to design silicon chips or configure general-purpose programmable chips to perform desired functions.

[0070] The preferred features are appropriately combined as will be apparent to those skilled in the art and can be combined with any aspect of the present invention.

[0071] Embodiments of the present invention will be described as examples with reference to the following drawings.

Brief Description of the Drawings

[0072]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13A

Figure 13B

Figure 14

Figure 15A

Figure 15B

Figure 15C

Figure 15D

Figure 16A

Figure 16B

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

[0073] Throughout the drawings, like reference numerals are used to represent like features.

[0074] The embodiments described herein relate to computer-implemented methods and systems for analyzing a multi-dimensional space, determining the structure of that space, and managing a map generated as a result of the analysis of that space. The approach employed in the examples described here treats the problem of understanding the space as a closed-loop problem, where sensor data from sensors or virtual sensors that describe the agent's environment directly drives changes in the agent's behavior.

[0075] The term "agent" is used to refer to either, or both, physical agents that exist in the physical space of the real world, such as vehicles or robots, and non-physical agents that exist in a logical space, such as points or pixels in an image or map.

[0076] Agents can move within their respective spaces. In the physical space, an agent moves by activating movement modules such as actuators and motors. In the logical space, an agent moves from a first point or pixel to a second point or pixel. An agent can be either real or virtual. An actual agent such as a robot occupies a region within the space. A virtual agent is a projection into the space and does not actually exist within the space. For example, a virtual agent may be a projection of the center point of the field of view of a camera observing a two-dimensional scene. The camera can move and / or rotate, changing the center point of the field of view on the scene and thus changing the position of the virtual agent. This will be explained in more detail below.

[0077] The physical space refers to a three-dimensional space that can be physically reached or traversed by real-world objects, vehicles, robots, or other agents. The three-dimensional space may be composed of one or more other physical or material objects. The three-dimensional space may also include the distance (or time) interval between two points, objects, or events that an agent may interact with when moving through the physical space.

[0078] The logical space refers to a space with a set of arbitrary points, such as a mathematical space or a topological space. Thus, the logical space may be, for example, an image, a map, a concept map, etc. A series of points may satisfy a series of axioms or constraints. The logical space may correspond to the physical space as described above.

[0079] Data is captured regarding the position of an agent within the space, and the captured data describes the environment of the space in the local vicinity of the agent. The capture of data is performed by sensors within the physical space, such as cameras, radars, tactile sensors, LIDAR sensors, infrared sensors, etc. The capture of data is also performed by virtual sensors within the logical space that acquire data of the logical space in the vicinity of the agent.

[0080] In the physical space, there is an observability problem that an agent cannot observe all of the physical space at a specific point in time. There are cases where the physical space is too large for the sensor to observe, or where the physical space includes objects or terrain that block the area of the physical space.

[0081] In the logical space, it is possible that all of the logical space can be observed at once. For example, when the logical space is an image, the entire image may be stored in the memory accessible to the agent.

[0082] However, in the logical space, it is not necessary for the agent to observe the entire logical space at once, which is more efficient. Therefore, the virtual sensor simulates the sensor in the physical space in that it only acquires a part of the logical space, and only the vicinity of the agent in the logical space is captured and / or acquired by the virtual sensor.

[0083] Therefore, in both scenarios of the agent in the logical space and the agent in the physical space, it is not necessary for the agent to observe the entire space, and it is computationally advantageous for the agent to observe only a part of the vicinity of the agent in each space.

[0084] When the agent is a non-physical agent in the logical space, it is not necessary for the vicinity of the agent to be directly observed by the sensor or the virtual sensor. For example, the vicinity of the agent may be a part or area of a pre-recorded image that defines the entire logical space.

[0085] Define the data captured by the sensor or virtual sensor that observes the space in the vicinity of the agent as observation data. This observation data is used to perform "observation". Therefore, observation is the result of processing the observation data in a way that concisely describes the environment in the vicinity of the agent.

[0086] In the physical space, an observation may be represented by a matrix or a vector that provides spatial information obtained from observation data. The observation data may be, for example, an image around the agent or point cloud data.

[0087] In the logical space, an observation may be represented by a matrix or a vector that provides spatial information in the local neighborhood of the agent. The observation data from which this observation is derived may be a set of pixels or points from the logical space around the agent. For example, the observation data may be a matrix of pixels within a threshold distance observable from the agent.

[0088] An agent may make multiple different observations at different locations in the space. The agent moves between these locations to make observations. The vector between two observations is defined as a "displacement" having both magnitude and direction. Thus, the agent moves between two specific locations and observations are made by moving along the displacement between those two specific locations.

[0089] Each of the observations and displacements made by the agent in the space may be stored in a memory associated with the agent. If the agent makes a series or sequence of observations and displacements in the space, that series or sequence of observations and displacements may be stored together as a set of observations and displacements. A map of the space may be generated using the sequence of observations and displacements. The map may include both observations and displacements, either observations or displacements, or characteristics associated with the observations and / or displacements such as action vectors or similarity measures as described later.

[0090] A series of observations and displacements may form an "identity". The identity may describe a specific feature or attribute of the space in which it was found. The terms attribute and feature are used interchangeably hereinafter.

[0091] For example, in the physical space, an identity may describe attributes such as a specific object or a path to a specific destination within a space. The observations within the set that forms the identity indicate how an agent perceives the space and thus how specific features are perceived. For example, when there is a specific object such as a chair in the space, the identity of the chair may include multiple observations obtained from various locations on the parts or sides of the chair. Similarly, when there is a specific destination in the space, the identity may include multiple observations obtained from locations on the path to the destination. The displacements within the set that forms the identity indicate how an agent moves between the observation locations. For example, the first observation of a chair in the space may be made from a position near the rear leg of the chair. The second observation of the chair in the space may be made from a position near the front leg of the chair. The first displacement of the identity is the displacement required to move from a position near the rear leg of the chair to a position near the front leg of the chair.

[0092] An exemplary identity is shown in FIG. 1. FIG. 1 shows a series of observations 101 connected by a series of displacements 102 in an exemplary indoor space that includes a door and a central obstacle. The space of this exemplary room can be a physical space such as a real-world apartment. The observations 101 may be captured using a camera on a robot. The displacements 102 may be made by moving the robot. In this embodiment, the identity represents the path to the door, which is considered the destination. Each observation 101 provides data that describes the environment from the location where the observation was made.

[0093] Although the displacements are shown as one-directional in FIG. 1, it is understood that a corresponding set of negative displacements is also recorded and that the identity is composed of the observations 101, the set of displacements 102, and the negative displacements that link the observations.

[0094] The attributes of the identity space shown in FIG. 1 may be a destination such as a door. Therefore, this identity can be used for the purpose of moving around in a room, avoiding a central obstacle, or moving to a door.

[0095] A second embodiment of the identity is shown in FIG. 2. FIG. 2 shows a series of observations 201 connected by a series of displacements 202 on a chair. The chair may be three-dimensional in a physical space such as the real world or two-dimensional in a logical space such as an image. Similar to FIG. 1, the displacements 202 may also include opposite negative displacements.

[0096] The camera, for example in a physical space, plays a role of capturing the observations 201 and performing the displacements 202 in order to obtain an identity as shown in FIG. 2. In a logical space, this may be performed by a virtual agent.

[0097] In the case of a camera in a physical space, the camera may be in a fixed position or not. When it is not fixed with respect to the space, the camera may move within the space. For example, the camera may be attached to a robot. In this case, the agent may be the robot (or the camera), and the agent's position is where the observation data is recorded by simply recording the images around the robot or the camera. When in a fixed position, the camera can rotate or zoom to focus on various parts of a chair. In this case, from the perspective of the camera's field of view, the camera's focus and the point of fixation (i.e., the direction of the line of sight) can represent the agent. In this case, the agent may be defined, for example, as the center point or projection of another specific pixel within the camera's field of view. The camera does not move translationally, but when the camera zooms or rotates, the agent "moves". The reason for this is that such actions change the position of the center point or a specific pixel within the camera's field of view with respect to the observed scene / image. Therefore, the agent's position does not need to match the camera's position. When fixed, the camera may not move between observations to obtain displacement 202. Instead, the displacement may be created and saved as the distance between two fixed points within the camera's field of view over two observations. The displacement can be executed by rotating the camera to "move" the agent and moving the center point of the camera's field of view.

[0098] For example, the camera may first capture an image centered on the legs of the chair. This image is then processed in a first observation. The position of the agent is considered to be the center point of the camera's field of view with respect to the first image. Next, the camera is rotated and panned upward to capture and process a second image of the chair backrest. The second observation is processed from the image of the backrest. The position of the agent in this second observation is considered to be the center point of the camera's field of view with respect to the second image, and the displacement between the first observation and the second observation is related to the pixel unit distance between the center point of the field of view with respect to the first image and the center point of the field of view with respect to the second image and / or the distance related to the size of the rotation. This can be estimated, for example, using visual inertial odometry.

[0099] If the camera is not fixed, the agent is the camera itself (or the device to which the camera is attached). In this embodiment, the camera / device moves to a specific location, makes observations at that location, and executes the movement. The displacement is recorded, for example, using odometry.

[0100] In the case of the logical space, FIG. 2 can represent an image of the chair. In this embodiment, since the entire image is saved, there is no problem of observability, and it is possible to observe the entire image at once. However, in order to improve the computational efficiency, the logical space is treated in the same way as the physical space, and instead of processing the entire image at once, each part of the image is observed in detail. Specifically, a part of the image plane, for example, the part including the legs of the chair, is observed. The size of the part may be fixed or may vary according to the size of the image. Similar to the example of the fixed camera in the physical space, the position of the agent is considered to be a specific point in a part of the image, for example, a specific pixel such as the central pixel of the image part.

[0101] In this case, the "agent" or virtual agent can perform a displacement by "moving" by the number of pixels on the image plane before making the second observation. In practice, this simply involves retrieving a second portion of the image plane from memory, whereby the second portion is centered around the new position of the agent (the pixels where the virtual agent "moved"). Thus, in the logical space, the agent is a tool for selecting various portions of the image based on fixed points within the image. The movement of the virtual agent involves a translation from the pixel where the virtual agent was previously located to another pixel.

[0102] Therefore, when referring to an "agent", it is necessary to understand that the agent can be one of three things. First, the agent can be a movable physical device, whereby the agent physically moves to a location within the physical space and makes observations at the location of the device. In this case, the agent performs the movement between observations by moving the physical device, and the observation data corresponds to the data captured by the agent at its current location. Second, the agent can be a projection of a fixed physical device. In this case, observations are made in various directions of the fixed physical device, and the agent becomes the projection of a point within the field of view of the physical device related to the orientation of the physical device. For example, the agent is located at the center point of the field of view of the fixed physical device. The displacement in this case is measured based on the distance between the projections of points within the field of view of the physical device caused by a change in the orientation of the physical device. Third, the agent can be a virtual agent within the logical space. In this case, the agent is located at a point or pixel within the logical space. The point or pixel designates the portion of the surrounding space corresponding to the observation data. For example, the agent may be located at the central pixel of a portion of the space corresponding to the observation data of the observation. The displacement is measured as the distance between pixels or points within the logical space.

[0103] Therefore, it should be noted that when referring to an agent, particularly "the position of the agent" in the foregoing description, any of the above definitions of "agent" may apply. The location of the agent does not necessarily have to be the location of a physical device. Similarly, when referring to an agent that performs actions such as moving or observing, the above definition applies. That is, the physical device does not necessarily have to move, and the observation may be performed by a physical device or a computing device that is not necessarily the agent itself.

[0104] Returning to FIG. 2, the attributes of the space defined by the identity formed by a series of displacements 202 and observations 201 may be the identification of a feature part (in this case, a chair).

[0105] Therefore, FIGS. 1 and 2 show that an identity can be formed in both two-dimensional or three-dimensional physical space and logical space for the purpose of identifying an object such as a three-dimensional chair, an image or a feature part of an image such as a two-dimensional chair, or for the purpose of navigating a three-dimensional space to reach a destination. It should be understood that identities have other uses, and object / image recognition and navigation are merely examples.

[0106] Multiple identities such as those in FIGS. 1 and 2 may be stored in a set of saved observations and displacements from previous experiences in one or more spaces. It is not necessary to label the identities. That is, in the above example, it is not necessary to label the identity corresponding to the chair with the label "chair", nor is it necessary to know that it corresponds to a chair. Rather, when a set of observations and displacements is stored as an identity, the agent can use this information to locate and identify what matches the identity in other regions of the space or in other spaces as a whole. In a system or method, it is not necessary to understand the real-world features to which the identity is related.

[0107] In some of the examples described below, an identity may be labeled according to the characteristics or attributes of the space it represents. In the above example, the identity corresponding to a chair may be labeled as a chair. The process of labeling identities within a set of stored observations and displacements may be performed as part of a general training process, which is carried out according to existing techniques that can be understood by a skilled person. For example, various machine learning techniques can be used, such as supervised learning processes, unsupervised learning processes, the use of neural networks and classifiers.

[0108] For example, the set of stored observations and displacements may be obtained from a verification space, such as a verification image that includes or indicates one or more known features. For example, a verification image of a chair is presented to an agent, and it is found that the identity formed from that image and stored in the set of stored observations and displacements corresponds to a chair and is labeled as such. The process of obtaining the set of stored observations will be described in detail later.

[0109] When using a set of observations and displacements to effectively describe the features of a space, rather than the entire map or image of the space, since each part of the feature only is included in the observations and displacements, less memory and computational effort are required to identify it in a part of the space, resulting in high computational efficiency. Furthermore, the observations and displacements can be calculated using various data, such as vectors encoded from the outputs of a series of spatial filters applied to the visual data of the observations, or in the case of displacements, the local temporal correlation of low-level features. This eliminates common failure modes, enables two parts of the system (observations and displacements) to compensate for the weaknesses of the other, and provides diversity. In the application area of 3D space SLMA, since the system essentially performs not only position determination and mapping but also navigation and path planning, no additional components are needed to achieve a complete implementation in the real world.

[0110] Compared to standard SLAM, the methods and systems described herein are more robust to environmental factors and occlusion. The methods and systems described herein also have the advantage of being able to store the minimum set of measurements necessary to navigate and observe a particular space. Observations are made within the space in response to the complexity and / or changes within the space, thereby adapting the methods and systems to optimally function in any space. For example, in vast fields or areas with few information and features, fewer observations are made and stored compared to environments rich in features. In this method, it is not involved in tracking features or objects that occupy a small portion of the field of view. Instead, since most of the characteristics of the entire solid angle are fully utilized, compared to conventional systems and approaches, the impact of occlusion and changes on the part of the visual scene on the system's ability to complete the task is reduced.

[0111] Destinations within a space, such as stores or landmarks, may be navigated from a location initially known in the world or a previously visited location based on a stored set of observations and displacements that connect the starting point and the destination. Additionally, destinations within a space may be navigated from an unknown region of the world or from an unknown starting point from a set of observations and displacements.

[0112] To identify the attributes of the space in this way, it relies on comparing the stored set of observations and displacements with the current observations and displacements made in a particular space. In the next part of the description, this process will be explained in more detail. In the next part of the description, the stored observations and displacements are referred to as predicted observations and displacements, and target observations and displacements. It is important to understand that these terms define the same technical characteristics and simply differ in order to indicate the context in which they are used and when. Similarly, the current observations from the space are referred to as first observations, second observations, and transition observations. Again, these terms define the same technical characteristics and are the observations made by an agent within the space where the agent exists, but are labeled differently to provide context regarding the reason or time of creation or use. The same terms apply to displacements.

[0113] A method for performing this process of identifying / determining the attributes of a space using existing knowledge in the form of a set of saved observations and displacements will be described in detail below with reference to FIG. 1. The advantage of this method is that only a part of the space is required from the perspective of current observations and displacements to identify the attributes.

[0114] FIG. 3 shows a flowchart of a method 300 for determining the attributes of a space in which an agent can explore or move according to various embodiments.

[0115] In the initial condition of method 300, the agent is present in the space and can observe a part of the space through sensors, virtual sensors, etc. This space may have been previously explored by the agent or may be completely unknown to the agent. The starting position may be known to the agent or may be completely unknown to the agent.

[0116] In the first step 301, the agent makes observations in the space from the current position of the agent in the space. This observation can be considered as the current observation or the first observation from the current location or the first location, but it should be understood that this observation does not have to be the exactly first observation in that space. Rather, here "first" is simply used to distinguish the current observation from subsequent observations.

[0117] The first observation includes data that describes a first portion of space in the local neighborhood of the agent. For example, if the agent is a robot or a vehicle, the first observation is made using sensor data from sensors such as cameras, radars, LIDARs, etc., and the sensor data shows a view of the world from the current position of the robot. The local neighborhood of the agent is the local neighborhood of the current position of the robot. In physical space, the range of the local neighborhood is determined by the sensor constraints and the environment of the robot. For example, if the range of the sensor is limited, the data of the observed sensor data is only acquired for the environment of the space within that range. Similarly, the environment of the robot may include one or more features that limit the observable surroundings, such as walls or other obstacles that block the field of view of the sensor. Therefore, the neighborhood of the agent is not fixed, and there may be a maximum distance from the position of the agent based on the sensor constraints. In logical space, the local neighborhood of the agent is the portion around the position of the agent and is smaller than the entire space. Since the entire space can be saved, for example, if the space is an image, it is not necessary to acquire data from the sensor to make an observation. Rather, the portion of the space around the position of the agent is obtained from the entire saved space. This "virtual sensor" simulates the same effect as using a sensor in physical space, but only considers a part of the space rather than the entire space in the observation. Therefore, the neighborhood of the agent in logical space may be set based on the distance from the position of the agent. In an example where the space is an image and the agent is placed at a pixel within the image, the neighborhood may be set, for example, by the number of pixels away from the pixel of the agent.

[0118] When the first observation is made in the first step 301, in the second step 302, it is compared with the saved observations from the set of saved observations and displacements.

[0119] By comparing the first observation with the saved observations, an observation comparison measure is generated that indicates how similar each of the saved observations is to the first observation. The observation comparison measures for each saved observation are ranked, the highest observation comparison measure is identified, and the corresponding saved observation is retrieved. In this way, the saved observation most similar to the first observation is determined and selected. The process of comparing the observation results will be described in detail later.

[0120] In the third step 303, a hypothesis regarding the saved observation most similar to the first observation is determined. In particular, the set of saved observations and displacements includes one or more identities that form a subset of the set of saved observations and displacements. Each subset includes one or more observations and one or more displacements each associated with a specific attribute or feature to which the identity is related. In the third step 303, the attribute or feature associated with the selected saved observation is determined. This attribute forms the basis of the hypothesis and, in effect, will predict what the agent is observing within the space. The prediction can be, for example, a two-dimensional or three-dimensional object or image, or a navigable destination.

[0121] The set of saved observations and displacements may include multiple subsets of observations and displacements, each subset being associated with a specific attribute and thus a specific hypothesis. In one example, there are subsets associated with the attribute "chair", subsets associated with the attribute "table", and subsets associated with the attribute "door". The comparison in the second step 302 is performed for all observations within the set of saved observations and displacements, and for a particular observation, the observation comparison measure may be the highest or strongest, and that particular observation is part of the subset associated with the attribute "door". Thus, in the third step 303, a hypothesis is established that the first observation is part of a door, i.e., the attribute of the space observed by the agent is a door within the space. The subset of saved observations and displacements associated with an attribute is called a hypothesis subset.

[0122] In the fourth step 304, observations and displacements of the hypothesis subset are obtained. In particular, the saved observation that is most similar to the first observation is, as explained above, part of the hypothesis subset of saved observations and displacements. Further, the observations and displacements within the hypothesis subset are linked and form a network of observations separated by displacements. The links between the observations and displacements within the hypothesis subset are saved in the set of saved observations and displacements, and the network configuration is saved in memory. When obtaining the hypothesis subset of observations and displacements, the method includes obtaining the observations and displacements associated with the saved observation that is most similar to the first observation. In other words, the adjacent observations and displacements necessary to reach that observation from the saved observation that is most similar to the first observation are obtained from the hypothesis subset. These observations and displacements become predicted observations and predicted displacements, and according to the hypothesis, if the agent moves by only the predicted displacement from the current position, it is expected that the agent will arrive at a position within the space where the predicted observations can be observed.

[0123] It should be understood that the terms "predicted observations" and "predicted displacements" simply refer to subsets of observations and displacements related to the determined hypothesis. According to the hypothesis, each of these observations and displacements is predicted to exist within the space currently occupied by the agent.

[0124] In the fifth step 305, the hypothesis is verified. To verify the hypothesis, the agent is sequentially moved along the network of saved observations and displacements that form the subset of the hypothesis, i.e., the predicted observations and predicted displacements. In each iteration of this sequential process of moving the agent along the network, the agent is configured to repeatedly perform new observations, compare the new observations with the predicted observations, and / or compare the predicted displacements with the actual displacements, and determine whether to maintain the hypothesis, reject the hypothesis, or confirm the hypothesis. This hypothesis verification will be further explained in more detail below with reference to steps 305-1 to 305-5.

[0125] In the first step 305-1 of the hypothesis verification process, the agent moves from the current location where the first observation was made to a second location in the space based on the movement function. The movement function depends on the predicted displacements of a subset of hypotheses. In particular, the movement function depends on the predicted displacement connected to the saved observation that is most similar to the first observation, identified from the linked network of observations and displacements forming the subset. Thus, the agent cannot recognize the second location until it executes the movement based on the movement function.

[0126] In the second step 305-2 of the hypothesis verification process, the agent makes a second observation from the second location of the agent in the space, and the second observation includes data describing a second portion of the space in the local neighborhood of the second location. For example, if the agent is a robot or a vehicle, the second observation is made in the same way as the first observation using sensor data from sensors such as cameras, radars, LIDARs, etc., and the sensor data is processed to be an observation showing a view of the world from the position of the robot at the second location.

[0127] In the third step 305-3 of the hypothesis verification process, the actual displacement of the agent from the first location to the second location is determined. In the physical space, the actual displacement can be estimated using, for example, odometry. In the logical space, the actual displacement can be measured according to known techniques using vector calculations or matrix calculations. For example, if the logical space is an image, the number of pixels between the first location and the second location can be determined as the distance. In the physical space, displacement is measured using an odometry system that combines visual inertial odometry and, if applicable, motion odometry such as wheel odometry. This system provides an interface for the input of visual inertial odometry and motion odometry. Displacement indicates the physical or logical distance between the observation pairs. Odometry is stored in a coordinate system that references the rotation of the starting observation (the predicted observation most similar to the first observation). Additional data can be stored, but is not limited to, characterizing other aspects of the displacement, such as vibrations that occur during the transition of the displacement or the energy expended during the transition of the displacement. The similarity of the displacements is measured by the endpoint error between the displacement pairs. In this case, the displacements start from the same location, and the endpoint error is the physical or logical Euclidean distance between the points identified by transitioning along the two displacements.

[0128] In the fourth step 305-4 of the hypothesis verification process, a comparison is performed. In the comparison, observations, displacements, or both are compared. The second observation is compared with the predicted observation from the hypothesis subset, and in particular, may be compared with the predicted observation linked to the predicted displacement on which the movement function was based. The comparison of the second observation and the predicted observation provides an observation comparison scale. Similarly, a displacement comparison scale can be obtained by comparing the predicted displacement based on the movement function when the agent moves from the first location to the second location in the space with the actual displacement between the first location and the second location. Therefore, by performing the comparison, both an observation comparison scale and a displacement comparison scale can be generated. As will be explained in more detail later, it is advantageous to use both of these scales. It should be understood that the order of the above steps after the agent moves is compatible.

[0129] In the fifth step 305-5 of the hypothesis verification process, the hypothesis is adjusted, maintained, or confirmed based on the observation comparison scale and / or displacement comparison scale from the fourth step of the hypothesis verification process 305-4.

[0130] In this step, the observation comparison scale and / or displacement comparison scale are evaluated to effectively determine whether the recognition of the agent's occupied space from the perspective of the second observation and actual displacement matches the recognition of the attributes according to the hypothesis from the perspective of the predicted observation and predicted displacement. Although the observation comparison scale and displacement comparison scale will be explained in detail later, generally, the stronger these scales are, the more likely the hypothesis is correct and the agent is actually observing the attributes according to the hypothesis in the space where the agent currently exists.

[0131] The observation comparison scale and / or displacement comparison scale are compared with one or more respective threshold values to determine whether to confirm, maintain, or reject the hypothesis.

[0132] Adjusting the hypothesis includes rejecting the hypothesis. This occurs when it is determined that the space observed by the agent is unlikely to contain or represent the attributes according to the hypothesis. This conclusion can be reached in several different ways. First, if the observation comparison scale and / or displacement comparison scale fall below the first observation threshold level and / or the first displacement threshold level, the hypothesis may be rejected. The first observation threshold level and the first displacement threshold level may be globally set as the minimum possibility required to maintain the hypothesis and for the agent to further verify the hypothesis. Alternatively, these thresholds may be adaptable and changeable based on environmental conditions of the space, such as factors like brightness / darkness. In particular, when the observation by the agent is performed in an environment with a brightness level different from the saved brightness level where a subset of the saved observations related to the hypothesis is captured, the lighting of the space may affect the observation comparison scale. The thresholds may be adjusted based on color, contrast, and other characteristics of the image / sensor data.

[0133] Furthermore, to address additional aspects of space that differ, different observational comparison scales can be used, and by using them in combination, it may be possible to identify specific dimensions of differences with respect to the hypothesis. For example, a change in lighting has a greater impact on the measurement of vector comparison than on the measurement of the spatial distribution of filtered elements around the position of the agent. A change in the color of an object affects the filtered elements of the color of the observation more than the filtered elements of luminance or direction. By considering various observational comparison scales, it becomes possible to observe the nature of the differences. When the lighting conditions change, the threshold can be adapted using the scale of the spatial distribution of the filtered elements of the observation around the location of the agent, and a strong matching against that scale can be used for pre-updating the hypothesis, and the acceptance criterion (second observation threshold level) of the "similarity" observational comparison scale can be lowered. The various types of observational comparison scales described here arise from different ways of comparing observations, which will be explained in detail later.

[0134] The first observation threshold level may be set according to the initial ranking of the observational comparison scale of each saved observation with respect to the first observation made by the agent. In particular, the first observation threshold level may be set equal to the next highest-ranked observational comparison scale corresponding to the saved observations of the set of saved observations and displacements that are not included in the subset related to the current hypothesis. Thus, this particular saved observation is related to a different hypothesis. In this way, the first observation threshold level is set according to the similarity between the first observation and the saved observations related to a hypothesis different from the hypothesis currently being verified. This enables the agent to effectively consider a hypothesis different from the verified hypothesis when the verified hypothesis generates an observational comparison scale that is weaker than the one first found for a different hypothesis. Thus, the verified hypothesis is rejected and ultimately replaced by another hypothesis associated with the first observation threshold level.

[0135] Finally, adjusting the hypothesis involves rejecting the current hypothesis and replacing it with another. This may involve, as shown in FIG. 3, the method returning to the second step 302, selecting the next best observation comparison measure from the ranked observation comparison measures, and determining the next best observation and the associated hypothesis. The agent can also return to the position of the first observation in the space via a movement opposite to the movement actually made before selecting a new hypothesis. The replacement of the hypothesis will be described in detail later.

[0136] The maintenance of the hypothesis occurs when the hypothesis is neither adjusted nor confirmed. Thus, when the observation comparison measure and / or the displacement comparison measure respectively match or exceed the first observation threshold level and / or the first displacement threshold level, the hypothesis is maintained. This means that there is no need to adjust the hypothesis. According to the example introduced above, the first observation threshold level is set according to the "next highest observation comparison measure", which means that the current hypothesis still provides the highest observation comparison measure regarding the predicted observation from any subset of the saved set of observations and displacements, that is, the attributes according to the current hypothesis are still the attributes most likely to be found in the space where the agent exists. Another condition for the hypothesis to be maintained is that the hypothesis confirmation condition is not met. The hypothesis confirmation condition may be a second higher observation threshold level and / or a second higher displacement threshold level. The hypothesis confirmation condition indicates the minimum reliability that the hypothesis is correct and the attributes of the space are the attributes according to the hypothesis.

[0137] In other words, when it is determined that the observation comparison measure and / or the displacement comparison measure for the predicted observation and the predicted displacement respectively match or exceed the first observation threshold level and / or the first displacement threshold level, but do not match or exceed the second observation threshold level and / or the first displacement threshold level respectively, the hypothesis is maintained.

[0138] If the hypothesis is maintained, the next predicted observation and the next predicted displacement that link the current predicted observation to the next predicted observation are obtained from a subset of the stored observations and displacements related to the hypothesis. Next, the process of verifying the above hypothesis is repeated for the next predicted observation and the next predicted displacement, as shown in FIG. 3. Thus, in method 300, it is involved in the sequential verification of the hypothesis subset of predicted observations and displacements until the hypothesis is confirmed or adjusted.

[0139] In each observation made by the agent while maintaining the hypothesis, the past observation comparison scale and / or displacement comparison scale are updated, and then the past observation scale and / or displacement scale are incorporated into future observation comparisons and / or displacement comparisons. The past observation scale and / or displacement scale are stored in memory and function to increase the future observation comparison scale and / or displacement comparison scale based on the number of consecutive times the hypothesis has been maintained. This effectively increases the reliability of being able to confirm the hypothesis based on the fact that the hypothesis has not been rejected for a plurality of consecutive observations and displacements. It also helps to prevent the agent from falling into a local minimum where the hypothesis is never rejected or confirmed. This process of using past observation scale data and displacement scale data may include a sequential probability ratio test of a single hypothesis against the null hypothesis (the data does not match the identity), or may be included as a sequential probability ratio test of multiple hypotheses between competing hypotheses including the null hypothesis (none of the hypotheses match).

[0140] Confirming the hypothesis includes determining that the hypothesis confirmation conditions are met and determining the attributes of the space based on the hypothesis confirmation conditions being met. As described above, the hypothesis confirmation conditions may be a second, higher observation threshold level and / or a second, higher displacement threshold level that the observation comparison scale and / or displacement comparison scale must match or not exceed.

[0141] Similar to the first observation threshold level, the second observation threshold level may be changed or adapted based on the lighting conditions in the space using the light intensity coefficient.

[0142] When the hypothesis is confirmed, since the agent identifies the attributes of the space, method 300 may stop. The identity of the attributes is saved in memory, and the observations and displacements executed by the agent to identify the attributes are saved in memory and may be linked to the space in which the agent exists.

[0143] Method 300 has the advantage of determining the attributes of the space using observations and displacements without the need to consider the completeness of the entire space. This realizes a spatial analysis method with much higher computational efficiency. Method 300 is an example of a method for traversing a space. However, it is not necessary to use both observations and displacements to navigate the space. In some embodiments, the use of displacements is optional and observations are mainly used.

[0144] The method 400 for performing and comparing observations will be described with reference to FIG. 4. The comparison of observations is used when navigating the space and can be performed without using displacements. The performing and comparing of observations are carried out at least at two points in method 300. First, in the first step 301 and the second step 302, the agent performs an observation, compares it with the set of saved observations, and finally determines a hypothesis. Next, in the fifth step 105, the hypothesis is verified. Therefore, in this process, at least two observations and two comparisons are required. However, it should be understood that more observations can be made, and not only the above two specific examples, but also in the transition observations between the first location and the second location, it is advantageous to sequentially perform observations while the agent moves from the first observation location to the second observation location. These transition observations will be described in detail later. Here, the process of making and comparing observations will be described.

[0145] In the first step 401, observation sensor data is captured. In the physical space, the observation data is input data captured using sensors such as cameras. In this example, the observation data is a captured image. The captured image and its content vary depending on the orientation of the camera. In particular, the camera may be attached to a robot at the same position in the world, for example, and may be at the same position. However, when the orientation of the robot is corrected, the content of the observation data may change even if the position of the robot has not changed.

[0146] Therefore, it is advantageous to perform observations in the physical space using a camera system or other sensors capable of capturing a three-axis stabilized cylindrical projection of a three-dimensional space, where the main axis of the cylinder is oriented perpendicular to the direction of gravity. By fixing it with respect to gravity in this way, it becomes possible to perform observations with the same roll and pitch (since these can also be fixed according to the direction of gravity). The capture resolution of the image may be, for example, 256 columns x 64 rows of pixels, but can also be set to any resolution. This 256x64 pixel image represents the input data for the observation and is then processed in the second step 402. In the non-physical or logical space, the observation data may be data obtained from the logical space around the position of the agent. It is also possible to save the entire logical space in advance. The observation data may be provided from a portion within a predetermined distance from the location of the agent within the logical space. For example, if the logical space is an image or a map, the observation data may be pixels or points included within a distance of 5, 10, 50, 100, or any number of pixels / points from the position of the agent.

[0147] In the second step 402, observations are generated from the input data captured by sensors, virtual sensors, or obtained from memory. The input data is encoded into a set of vector / matrix representations through a step-by-step pipeline.

[0148] The process of generating observations to be saved to form part of a set of saved observations and displacements is very similar to the process of generating current observations from input data for comparison with the saved observations. However, when comparing current observations, there may be a possibility that the relative orientation when the agent is at the current position and the orientation when the agent made the saved observations are not the same and may not be easily distinguishable. To solve this problem, when making current observations from the input data, the process includes making observations for each possible permutation of the agent's orientation. In practice, this means that the current observations are associated with multiple rotations. Each rotation then represents input data that has been shifted or adjusted to represent a different orientation of the agent. How this affects the process compared to making "saved" observations is described below.

[0149] The general process of generating observations (current or saved) includes the following steps.

[0150] First, low-level features are extracted from the input data. This may include, for example, the use of one or more filters or masks for preprocessing the input data. Examples of low-level features include edges, colors, shapes, etc. Preprocessing may also include color conversions to provide multiple channels of the input data. The result of this first step is processed data that may potentially include multiple channels.

[0151] Second, the resolution of the input data is reduced by extracting or removing parts of the input data, such as rows of pixels, using filters, etc., and / or by summing and averaging the data over large regions of the input data.

[0152] Thirdly, the input data may be convolved using one or more kernels to obtain useful results for identifying, for example, regions of interest, edges, regions of greatest change, etc. If there are multiple channels, this process is performed for all channels.

[0153] Fourthly, the resulting data forms a vector or matrix of processed and reduced input data. This is then concatenated and normalized across dimensions (columns, rows, channels) to generate an observation vector o.

[0154] As described above, when storing observations, only one observation vector o is required, for example, for the purpose of storing a set of observations such as identities or maps. When making current observations for comparison purposes, this general process is performed for several rotational permutations of the input data. If the resolution of the data is XxY, this may mean including X permutations, each representing an X+1 shift from the previous permutation. As a result, X observation vectors are generated for the current observation, and each vector can then be compared with the single stored observation vector.

[0155] To encode the input data to form one or more observation vectors o, a computer device such as a central processing unit (CPU), a graphical processing unit (GPU), a field programmable gate array (FPGA), or an application specific integrated circuit (ASIC) is used.

[0156] A detailed example of this process is shown below for the above example of an image (cylindrical projection image) captured from a camera. In this example, since the resolution of the image is 256x64 pixels, there are 256 possible rotations of the projection for the current observation, and the columns are rotated (shifted) from the original capture position. As described above, these rotations represent various permutations of the original observation data. It should be understood that there may be any number of permutations based on the input resolution of the image. The detailed process is as follows.

[0157] In the first stage, color conversion processing is performed on each pixel, whereby the color channel data from the captured image is converted. For example, the data of the red, green, and blue channels is converted to the red, green, and luminance channels, where the luminance is 1 / 3 (red + green + blue). It should be understood that these colors and conversions are exemplary, and that there are different color channels in the image and that they may be converted to different color channels.

[0158] In the second stage, the luminance channel result of the first stage is convolved with a filter using the convolution kernel [1, -1; 1, -1; 1, -1] or a similar vertical edge filter (e.g., Sobel filter or other filter) to create a luminance vertical (luma_v) channel.

[0159] In the third stage, the luminance channel result of the first stage is also filtered by convolving it with a filter using the convolution kernel [1, 1, 1; -1, -1, -1] or a similar horizontal edge filter (e.g., Sobel filter or other filter) to create a luminance horizontal (luma_h) channel.

[0160] As a result of this process, a new image is generated that includes pixels of four channels: red, green, luma_h, and luma_v. It should be understood that other pixel filter options can also be used. With the selected filter set, all sub-pixels of all channels (red, green, luma_h, or luma_v) are routed to the next processing pipeline.

[0161] In the fourth stage, a series of box filters are calculated and arranged to surround the rows around the set number of equally spaced rows of the image for each channel. It should be understood that the number of filters can be more or less than this, and the size of 19x19 is exemplary. The box filters smooth the image for each channel. Each pixel for each channel is changed by the box filter from the sum of the pixels contained within each box filter, thereby providing a smoothed (or blurred) image. The box filters may overlap each other.

[0162] In the fifth stage, the blurred image is reduced to form a slice image. In particular, the resolution of the blurred image is significantly reduced. This process is shown in FIG. 9. Specifically, four equally spaced rows 902 are extracted from the blurred image (not shown) processed from the original image 901 (or other observation dataset), and a sliced image 904 is generated.

[0163] As shown in FIG. 9, the sliced image contains 256 columns and 4 rows.

[0164] In the sixth stage, the columns of the slice image formed in the fifth stage are divided into a set of equally spaced columns. As an example, 256 columns are divided into 256 sets, each set containing a 16-column x 4-row image, and the first column of each set is different from the complete set of 256 columns. As can be seen from FIG. 9, there are 16 sets of 16x4 images (a total of 256 permutations), the columns are shown in different colors, and it shows how the data is divided from the 256-column sliced image 904 into 16x16x16x4 sets. These sets represent all possible rotational positions of the input data (an increment of one column rotation for all possible 256 rotations). FIG. 9 is an example for illustration, and the number of rows, columns, and permutations can be more or less than this, that is, the number of rotations can be more or less than 256.

[0165] When storing the observations in memory and forming part of a set of stored observations and displacements, only one of the 16x4 images of the set formed in the fifth stage needs to be stored. For example, the first permutation 906 of the first set may be selected to be stored as an observation. Rotation (i.e., 256 permutations of the 16x4 image) only needs to be considered when creating an observation for comparison with previously stored observations. Therefore, the permutations only need to be considered for the current observation. When comparing, it is not necessary to store each permutation of the 16x4 image in memory. Instead, each permutation is repeatedly compared with the stored 16x4 observation vector, and the optimal matching indicated by one or more observation comparison metrics is identified. It is not necessary to store the permutations. The comparison of observations will be described in detail later.

[0166] Returning to the detailed example of creating an observation, in the seventh stage, the 16×4 image is periodically convolved with a center-on, periphery-off (or vice versa) kernel. For example, the element with a value of 1 is at the center of the kernel. The kernel for the first row is, for example, [-1 / 5,1,-1 / 5;-1 / 5,-1 / 5,-1 / 5], the second and third rows are [-1 / 8,-1 / 8,-1 / 8;-1 / 8,1,-1 / 8;-1 / 8,-1 / 8,-1 / 8], and the fourth row is [-1 / 5,-1 / 5-1 / 5;-1 / 5,1,-1 / 5], etc. This example is shown in FIG. 10. FIG. 10 shows the center-on, periphery-off horizontal lap filters 1001 to 1003 of the 16x4 image. This kernel identifies the regions in the image with the greatest variation. As described above, this is only necessary for one 16x4 image when storing the observations. When creating the current observation for comparison with the stored observations, this is performed for each separate set (i.e., the 16x16 sets from the sixth stage).

[0167] Following this pipeline, a 16x4x4 image is generated, and in the case of comparison, the 16x4x4 image is rotated 256 times. There are 16 columns, 4 rows, and 4 channels. The channels of the 16x4 image may be divided into one data subset composed of red and green, and two data subsets composed of luma_h and luma_v. Each of these data subsets is normalized by vector normalization, where each element describes a value on a set of orthogonal values and the result is a unit vector. These normalized vectors are each scaled by the square root of 2 and concatenated into a single unit vector. This provides the observation vector o. The observation vector for this example is concatenated from 16 columns, 4 rows, and 4 channels, and a 256-element vector is generated and saved. When comparing, 256 permutations of the vector o are compared with one saved vector.

[0168] In addition to the observation vector o, the final vector can store additional data related to the location of the observation (such as user-defined labels, angular location and identification of objects, data from additional sensory modalities such as auditory or olfactory information, traversability in different directions in physical or logical space, and previously explored directions, etc.).

[0169] The observation vector o is a rough representation of the original observation data captured by the sensor and provides a statistical summary of the environment in the vicinity of the agent within the physical or logical space. The coarseness of the vector compared to the original data allows the method to be executed more efficiently. Also, the observation data is filtered at high resolution before coarsening, creating various channels. Coarsening has the advantage of providing a representation of areas that are less sensitive to changes in lighting or objects in the environment than state-of-the-art techniques, as information is integrated over most of the field of view.

[0170] It should be understood that the observation vector o can be either a vector or a matrix. The vector o may be composed of a plurality of sub-regions, each corresponding to a part of the entire vector. The sub-regions may be content / color sub-regions divided based on the channels used to form the vector (in the example, red, green, lumam_h, luma_v), or may also be spatial sub-regions corresponding to various directions around the agent. The sub-regions can be represented by sub-vectors, and the entire observation vector o is formed by a combination of sub-vectors. As described above, the vector representing the rotation permutation is necessary when comparing observations. When the current observation is created and needs to be compared with the saved observations, it is not necessary to save all permutations.

[0171] It should be understood that the numerical values given above regarding the dimensions, filter and kernel sizes, row selection, and rotation permutations may be more or less than the exemplary values.

[0172] Returning to FIG. 4, in the third step 403, the current observation vector o is compared with the saved observations from the saved observation set. That is, the current observation is compared with the target observation.

[0173] The observations may be compared using several functions. By comparing the observations, an observation comparison scale is obtained. The observation comparison scale may vary depending on the function used to compare the observations. In some examples, multiple observation comparison scales may be generated by comparing the observations using multiple comparison functions. These comparison functions may be used alone or in combination in method 300 for verifying hypotheses. They may also be used in other related methods such as the map traversal described later. Environmental characteristics such as the brightness level of the scene affect one observation comparison scale but not another. Therefore, by using a combination of the observation comparison function and the resulting observation comparison scale, method 300 can be executed more accurately under various environmental conditions.

[0174] Functions for comparing observations include an observation similarity function 403-1, an observation rotation function 403-2, an observation action function 403-3, and an observation variance function 403-4. As shown in FIG. 4, one or more of these functions can be executed on the current observation and the saved target observation. Therefore, the initial conditions and systems required for each method of comparing observations are the same. This system is schematically shown in FIG. 5.

[0175] FIG. 5 shows a schematic functional diagram of a system or apparatus 500 that includes various modules involved in the comparison of observations. These modules may be implemented by a computer, a computing device, or a computer system. It should be understood that the functions of these modules may be further divided into other modules and separated among components, or their functions may be executed by one device or component. The first module of the apparatus or system is a memory module 501. The memory module 501 stores a set of saved observations and displacements. These observations and displacements may be stored as vectors. A set of observations and displacements including one or more hypothesis subsets can be stored in a cyclic graph or an acyclic graph, and / or a directed graph. Alternatively, the observations and displacements are stored with metadata that identifies the displacements and observations to which each is connected, or organized in a reference table that holds the links between each observation and displacement.

[0176] The memory module 501 is connected to or otherwise communicable with the saved observation processing module 502. The saved observation processing module 502 is configured to select saved observations from the memory module 501 to be compared. The saved observation processing module 502 can also select one or more masks for the purpose of reducing the saved observations to a part of the saved observations for input to a comparison function. The masks include observation masks, masks selected by the user / system, or combinations thereof. The mask masks a specific area of the selected saved observation vector and retains other areas called regions of interest (ROIs). The mask is a binary mask and may mask or retain using values of 0 and 1, or vice versa.

[0177] The saved observation processing module 502 is connected to or communicable with a comparison engine 503 that performs a comparison according to one or more of the comparison functions 403-1, 403-2, 403-3, and 403-4.

[0178] A second aspect of the device system or the like is provided with a sensor module 504. The sensor module 504 includes sensors capable of inferring spatial information from the environment, and may include cameras, LIDAR sensors, tactile sensors, and the like. The sensor module 504 also includes control circuits and memory necessary for recording sensor data. The sensor module 504 is configured to capture current observation data.

[0179] The sensor module 504 is connected to or communicable with the current observation processing module 505. The current observation processing module 505 processes sensor data from the sensor module 504 and obtains current observations such as a current observation vector. This includes all possible rotations on the sensor data according to the processes described above with respect to FIGS. 9 and 10. The current observation processing module 505 can also select a mask for masking the current observation vector.

[0180] The current observation processing module 505 is connected to or communicable with a comparison engine 503 that performs comparisons according to one or more of comparison methods 403-1, 403-2, 403-3, and 403-4.

[0181] The comparison engine 503 is configured to obtain or receive saved observations from the saved observation processing module 502, current observations from the current observation processing module 505, and respective masks for the purpose of performing comparison observations according to any of comparison methods 403-1, 403-2, 403-3, and 403-4. The comparison engine 503 may be implemented by a computing device, a server, a distributed computing system, etc.

[0182] The system 500 further includes a movement module (not shown) configured to move an agent within a space. In the physical space, this is a motor, an engine, etc., and is configured to move the agent by physically moving the position of the agent or changing the orientation of the agent. In the logical space, the movement module may be implemented by a computer or the like.

[0183] It should be understood that the comparison methods are described individually below, but they may be executed simultaneously or within the same overall process executed by the comparison engine 503. Since these methods each depend on the same basic data points and operations, they can be executed in the same overall process. In particular, each of the following methods compares current observations with saved target (predicted) observations. For example, if the observations are represented as vectors, each method performs this by taking the inner product between the vectors or sub-vectors thereof. The comparison methods differ by additional preprocessing / postprocessing steps, as shown below.

[0184] In the first method 403-1, observations are compared based on a similarity function to obtain a first observation comparison scale component. The first observation comparison scale component is the "similarity" observation comparison scale described above. In this similarity function, the vectors of the observations to be compared need to be vector-normalized. The similarity function includes calculating the vector inner product of the observation vectors to be compared. The inner product is composed of the cosine of the angle between two vectors. Therefore, the similarity function returns a single value or score indicating the similarity between the two vectors corresponding to the observations being compared. When the current observation includes multiple possible rotations (in the above example, when there are 256 possible rotations), the vectors corresponding to each of the rotations are compared with the predicted observation. In this case, the inner product that provides the highest value is saved as the first observation comparison scale component. This single value is used in the second step 302 of method 300, for example, to generate a set of ranked observation comparison scales corresponding to the set of saved observations, from which the saved observation most similar to the first observation is determined and selected to determine the hypothesis. In other words, this similarity function may be used to determine the similarity between the first observation and the set of saved observations, identify which of the saved observations is most similar to the first observation, and determine which hypothesis to verify.

[0185] This similarity function may be similarly used in the fifth step 305 when comparing the second observation with the predicted observation.

[0186] When the current observation is associated with multiple rotational positions and each rotational position corresponds to an observation vector, the similarity function is configured to find the maximum inner product between different observation vectors and the target observation. To do this, the similarity function tracks the "maximum inner product value" while repeatedly calculating the inner product for each of the observation vectors rotated with respect to the target observation. When all rotations are evaluated, the maximum inner product value is obtained as the first observation comparison scale component.

[0187] The advantage of using this similarity function is that the computational efficiency during execution is relatively high, and the resulting first observation comparison metric component decreases very smoothly as a function of the distance from the original observation location (or the target observation location where the current observation is being compared).

[0188] However, this function may be affected by environmental conditions such as the brightness or darkness of the environment. For example, if the saved observations are related to observation data captured under low-level light conditions and the current observation is related to observation data captured under direct sunlight, even if the features captured in each observation are spatially the same, the current observation data may be different from the saved observation data, resulting in different vectors being generated. To cancel out these effects, the first observation comparison component can be combined with or used in conjunction with a fourth observation comparison component formed from a fourth observation comparison function 403-4 shown below.

[0189] In the second observation comparison function, an observation rotation function is used to compare observations, and a second observation comparison metric component is obtained. This rotation function obtains the rotation o of the current observation and, similar to the case of the similarity function, compares each of these with the target observation. For example, during the fifth step 305 of method 300 where the hypothesis is being verified, the observation rotation function may be used to compare the second observation (the current observation) with the predicted observation (the target observation). In the above example, there are 256 rotations corresponding to 256 different column positions from the original cylindrical observation data. While the similarity function is configured to obtain the inner product indicating the optimal matching between the rotated observation vector and the target observation, the observation rotation function is configured to identify the specific rotation that causes this optimal matching. The observation rotation function outputs the rotation direction of this optimal matching as an offset value. This corresponds to the index of the rotational position of the observation. For example, the optimal matching between the observation vector o and the predicted observation may be 120 番目 out of 256. This is 120 out of 256 columns 番目The column corresponding to the vector formed from the observed data rotated such that the [[ID=]] column is at the center of the data, i.e., at any other defined rotational position. Thus, the rotational observation function may output an offset value corresponding to the index of the rotational position (which may be, for example, 120). 番目 In the physical space, when an inertial measurement unit (IMU) is used as part of the sensor or is included in the sensor, the offset value indicates how much the yaw (roll and pitch can be fixed by gravity) has drifted since the target observation was first made (assuming that the saved observations forming the target observation have occurred previously in the current space). Thus, the offset value of the observation rotation function provides an output indicating the IMU yaw drift. This IMU yaw drift can be identified from the output and used to update the set of saved observations and displacements, since they are all relative to the original yaw at which they were created and saved. Thus, if the reliability from the first observation comparison scale component and / or the second observation comparison scale component is such that what is indicated as the current observation is the target observation, the offset value can be used to adjust the set of saved observations and displacements accordingly.

[0190] If the physical space in which the agent exists is not the same space as the space in which the saved observations were captured, the offset value indicates how the space in which the agent exists is oriented compared to the space in which the saved observations were captured.

[0191]

[0192] ​In summary, the second observed comparison scale component generated by the observed rotation function provides an offset result that shows the strongest similarity between the target observation and multiple rotations (e.g., 256 rotations) of the current observation. Further, the index of the rotation showing the strongest similarity indicates the alignment between the current orientation of the agent (i.e., the x-axis and y-axis that the agent currently recognizes) and its orientation when the target observation was created and saved. Therefore, since the direction information can be recovered by the offset, the target (predicted) observation and displacement can be remapped accordingly.

[0193] This rotation function may be used when comparing the first observation with the set of saved observations in the second step 302 of method 300, and may also be used when comparing the second observation and the predicted observation with the predicted displacement and observation recalculated based on the relative orientation (reference frame) of the agent in the fifth step 305.

[0194] The above-described observed similarity function 403-1 and observed rotation function 403-2 can be executed either in the same process or independently. An example of a single process capable of executing these functions simultaneously is shown below in the form of an example of pseudo-code for executing both the observed similarity function 403-1 and the observed rotation function 403-2. The pseudo-code is as follows. combined_function(cam_observation[16x4x4x256],cam_mask[16x4x4x256],memory_observation[16x4x4],memory_mask[16x4x4],user_mask[16x4x4]) · Initialize max_similarity to -1 · Initialize offset to 0 · For each of the 256 rotations: o Initialize cam_sum, mem_sum, match_sum to 0 o Combine cam_mask, memory_mask, user_mask with logical binary OR (use the element if 0, do not use if 1) o For each of the 16x4x4 (total 256) vector elements: ■ If the mask is 0: · Add the square of the elements of cam_observation to cam_sum · Add the square of the elements of memory_observation to mem_sum · Add the product of the elements of cam_observation and memory_observation to match_sum ■ Otherwise: · Do nothing o Normalize match_sum by dividing it by the product of the square roots of cam_sum and mem_sum o This is the curr_similarity score for this rotation o If the similarity score is greater than max_similiarity: ■ Set max_similarity to curr_similarity ■ Set the offset to the rotation index o Otherwise: ■ Do nothing · After considering all rotations, max_similarity is the output similarity, the first observation comparison metric component. · The offset after considering all rotations is the output rotation, i.e., the second observation comparison metric component.

[0195] In the above pseudo-code, "cam" refers to a sensor (e.g., a camera) and thus the current observation. Similarly, "mem" and "memory" refer to memory, i.e., the saved target observation to be compared.

[0196] The cam, memory, and user mask serve the function of including and / or ignoring specific parts of the current observation vector and the target observation vector in the comparison function. The cam mask and the memory mask are composed of information regarding the reliability of information from parts of the visual scene within the observation vector. In the case of the memory mask, this refers to a single observation vector, and in the case of the cam mask, this refers to a set of observation vectors for various rotations of the field of view. This can be used, for example, to exclude parts of a camera image where the input data is saturated due to excessive contrast and the data quality has deteriorated. The user mask is composed of information provided by an internal function of the system attempting to select a subset of the complete observation vector elements, or an external function performing the same. This can be used, for example, for specific input functions or to determine the matching of observations for various parts of the visual scene encoded in the observations.

[0197] The user mask is configured to remove outliers / incorrect data from the observation data.

[0198] The above pseudo-code is iterated for each possible rotation of the current observation, identifying the maximum normalized similarity measure (the first observation comparison measure component) across all rotations and the rotation offset (the second observation comparison measure component) corresponding to this maximum value.

[0199] In the above first and second observation comparison functions 403-1 and 403-2, and in the third comparison observation 403-3 described below, the observation vector o may be created from an image in physical space. In this image, the x-axis is the angle around the agent, and the y-axis has the property of being some (radial distance or vertical altitude) property perpendicular to the x-axis. In the above example, the vector is composed of four channels with sub-pixels (red, green, luma_h, luma_v). The number of channels may increase or decrease, and accordingly, the observation vector may become larger or smaller. Further, the image and each of its channels are composed of a series of rows and columns providing spatial information about a part of the space where the agent is present.

[0200] Therefore, the observation vector o includes spatial information in multiple directions and color / edge information from different channels, and may be relatively large. This relatively large vector may be manipulated or sliced for various comparisons. For example, using only the edge channels luma_h and luma_v makes the comparison less susceptible to the influence of color changes. Using only the green channel can, for example, determine changes in the distribution of plants. If all channels are used but only for a spatial sub-region of the input image, only the corresponding part of the visual world is compared. This enables partial comparison and allows checking how well specific parts / features of the world match.

[0201] To perform slicing of the observation vector o, one or more masks are used. Each mask is a vector having the same number of elements as the observation vector o and is configured to mask specific regions of the observation vector o while retaining other regions. For example, a binary mask of 0s and 1s may be used for such purposes. If the observation vector o has 256 elements, the mask consists of 256 elements. Here, the value 1 indicates that the corresponding index of the observation vector o needs to be included / ignored in the subset of the observation vector o being compared. On the other hand, when the value of the mask is 0, it indicates that the corresponding element of vector o is not to be included / ignored in the subset of the observation vector. The mask can be used, as described above, to filter out regions of the data with errors.

[0202] In a third method of comparing observations, an observation action function 403-3 is used to obtain a third observation comparison metric component. The observation action function can be used to compare observations and can also be used as part of a movement function to move the agent to a second location while verifying a hypothesis according to the fifth step 305 of method 300. The observation action function can also be used to create an action vector field. The action vector target field effectively maps an estimated direction in space towards the estimated location of a particular observation.

[0203] The observation action function relies on the same basic principle as the observation similarity function and the observation rotation function. Observations differ from the previous two action functions in that multiple sub-regions of the observation vector o are evaluated and an offset is determined for each sub-region. Thus, instead of the entire observation vector o, a subset of the observation vector corresponding to the spatial sub-regions is evaluated by the observation action function, from which it is possible to obtain the offset value of each sub-region corresponding to the subset of vectors relative to the target observation. A mask is used to obtain the relevant sub-regions. Each sub-region may overlap with one or more adjacent sub-regions.

[0204] The use of sub-regions for this purpose stems from the concept that when the agent moves away from or towards the predicted observation location, not all parts of the field of view change equally. This differs from conventional triangulation in that it considers the distortion of the entire visual scene rather than the movement of the identified object or signal source. Therefore, it is robust even if a part of the visual scene is occluded and more resistant to environmental changes. Furthermore, since it does not require metric calculations to function, it becomes more robust to measurement noise. Using this property, not only can the agent determine how far it is from the target observation location, but it is also possible to calculate an additional vector to move to that location. This additional vector calculated by the observation action function uses the same principle as the observation similarity function, but the inner product is obtained for N spatial sub-regions, and these are compared with the corresponding sub-regions of the target observation, which is different. For example, when N = 8, the eight spatial sub-regions have half of the x-axis range of the full field of view of the entire observation vector o and are centered on different directions. In this example, the observation action function combines the observation similarity function and the observation rotation function and is executed eight times, once for each sub-region, and eight similarity points and eight offsets are detected. The eight similarity points and eight offsets differ only with respect to a specific sub-region of the observation vector o and are substantially the same as the first observation comparison scale component and the second observation comparison scale component, respectively.

[0205] Therefore, for the actual comparison performed, the saved target observation vector is divided into N sub-regions, and for each of the N sub-regions, the corresponding sub-region of the current observation vector is compared for all possible rotation permutations, and it is identified which permutation is most similar to the saved target sub-region. For the most similar specific permutation, an index (e.g., 120) is used to identify the offset 番目There is (rotation). Offsets are calculated for each sub-region of the saved observation vector, and these offsets are used to obtain the action vector required to reach the position of the target observation. Similarity is used to evaluate and weight the reliability of the movement vector when this movement vector is used in the movement function described later. Therefore, the overall output of the observation action function 403-3 and the third observation comparison scale component become this movement vector.

[0206] FIG. 6 shows a schematic diagram of observation data corresponding to sub-regions of an observation vector generated for the purpose of executing an observation action function. FIG. 6 also shows the process in which sub-regions are executed in the observation action function (with respect to observation data for visualization purposes).

[0207] As shown in FIG. 6, the observation data is represented by a two-dimensional circle 601 of spatial data around the agent 602. This is a representative example, and it should be understood that the observation data may be three-dimensional, such as inside a cylinder. The observation action function processes by comparing the observation vector o with the target observation. According to the observation action function, a plurality of sub-vectors of the observation vector are processed, and each sub-vector is focused in a different direction or placed at the center. In FIG. 6, these directions are shown as a plurality of directions 604-1 to 604-n in the corresponding observation data 603. There are eight directions and eight sub-regions in FIG. 6. However, it should be understood that in the process of creating sub-regions, there may be more or fewer selected directions than this.

[0208] Each of the eight directions corresponds to eight sub-regions 605-1 to 605-n of the observation data (not all are shown). These sub-regions correspond to each part of the visual field (sensor field if not a camera), and each part is centered around one of the eight directions. The specific direction associated with a sub-region is called the sub-region direction.

[0209] In the observation action function, as will be described in more detail below, pairs of sub-regions 605-1 to 605-n are compared with the corresponding parts of the saved observations. This is shown in the observation data 606 in which the first sub-region 606a and the second sub-region 606b are compared with the corresponding parts (not shown) of the target observation.

[0210] The observation action function calculates the offsets between each of the sub-regions 606a and 606b and the corresponding parts of the target observation. Thereby, as shown in the observation data 607, a pair of offsets 607a and 607b is generated. From these offsets 607a and 607b, as shown in the observation data 608, a resulting offset vector 609 is determined. The sub-regions 606a and 606b used in this process are different sub-regions and do not necessarily have to be so, but may also be opposite sub-regions.

[0211] Here, the observation action function 403-3 will be described in more detail with reference to the sub-regions of the observation data described above and shown in FIG. 6 and the sub-vectors of the corresponding observation vectors.

[0212] As described above, the observation action function uses a set of masked sub-vectors. Each sub-vector is a subset of the saved observation vector o and corresponds to a continuous sub-region of the field of view. The sub-region of the saved observation vector is compared with the corresponding sub-region of the current observation vectors arranged at equal angles around the agent. Following the previous example, the sub-region of the saved vector may cover a set of consecutive columns within the saved 16x4 observation vector. An even number of overlapping sub-regions are extracted. There are N pairs of sub-regions, and the central column of each sub-region is projected into the real world with opposite vectors as shown in FIG. 6. FIG. 6 contains eight sub-regions. These opposite sub-regions match 256 rotations of the current observation, and the most matching rotation is extracted using the observation rotation function 403-2. In particular, an offset corresponding to the rotation index that best matches between the sub-vector corresponding to the sub-region and the rotated saved observation is determined for each opposing sub-region. The difference in the direction of the offsets of the opposite pairs of subsets of the observation vector is compared. This provides a vector indicating the direction and magnitude of the displacement required to reach the saved observation (in FIG. 6, along the resulting offset vector 609 perpendicular to the line connecting the centers of the two opposing sub-regions 606a and 606b). This direction and magnitude are the direction where the target observation being compared was captured. The weighted average of the offset vectors 609 obtained from all opposing sub-regions provides a single action vector indicating the direction to approach the target observation location being compared from the current observation location. This vector is rotated from the saved observation rotation to the current observation rotation using a second observation comparison scale component or another rotation measurement to align the movement direction with the agent's current world frame. It should be understood that alternative implementations can be performed where the sub-regions are evaluated in the current observation world frame and such rotation is not necessary, but this would be computationally expensive as the sub-region mask would need to be evaluated for each current observation rotation, and it would be impossible to accurately evaluate the sub-regions in the saved observation, resulting in low accuracy.

[0213] Therefore, it can be considered that the observation action function 403-3 is a process of repeatedly using the observation similarity function 403-1 and the observation rotation function 403-2 to obtain the vector of the direction of the position of the target observation. Therefore, the above observation action function 403-3, the observation similarity function 403-1 and the observation rotation function 403-2 can be executed by the same process, by the same comparison engine 503, or independently.

[0214] An exemplary process for executing the observation action function 403-3 using the comparison engine 503 is shown below in the form of exemplary pseudo-code.

[0215] Action_function: · For the current observation vector: o Generate a mask for selecting the spatial sub-region of the observation in the following way: ■ Given an observation: observation^6x4x4], where the first two dimensions are spatial (X, Y) and the third is a different channel. ■ For the spatial elements of each sub-pixel: · Select a periodically continuous subset of the columns within the spatial element. By periodically continuous, in the case of a 16-column dimension, it represents an example of a continuously subset of columns 2, 3, 4 and columns 15, 0, 1, 2, but note that columns 1, 3, 4 are not continuous because there is an intervening column 2. These subsets should have the same width (e.g., 8 columns) and be evenly spaced between columns. For example, in the case of 4 subsets with a width of 8 columns, indices 0, 4, 8, 12 constitute an example of equal spacing when referring to the first column, the last column, or other determined columns within each subset. · The determined subset is applied as user_mask, and the elements of each column within the subset are set to 0 (used), and the elements of the outer columns become 1 (not used). · For each sub-region: o Execute combined_function (the same as above) · The max_similarity and offset of the sub-region are saved · A set of max_similarity and offset is given, and there is one max_similarity and offset for each sub-region o For pairs of offsets corresponding to opposite directions within the agent's field of view for sub-regions (e.g., for 8 sub-regions, the pairs composed of indices 0 and 4, 1 and 5, 2 and 6, 3 and 7), calculate by subtracting 128 from the circular distance between the first index and the second index of the offsets within the pair. This corresponds to the angle between the offsets and how the two sub-regions have moved within the field of view. The direction perpendicular to the pair of sub-region directions is the action direction. The offset difference is used to calculate the magnitude of the action. The directions and magnitudes of the actions for all pairs are vector-averaged to form the output of the final action function, i.e., the action vector. o The magnitude of the action is the magnitude of the result of the vector described by the offset (the second offset increases by 128 each time and corresponds to a rotation of PI radians).

[0216] Therefore, in the observation action function, N pairs of opposing offsets (N = 8) are taken, and each pair represents a direction described by the perpendicular to the direction of the mask center of the opposing offsets. This direction is combined with the magnitude calculated from the difference between the two offsets, and when both offsets are zero (i.e., when the center of the sub-region of the current observation vector is the same as the sub-region of the saved observation), the magnitude becomes zero. When the difference between the first offset and the second offset is periodic, for example, when the first offset is 200 and the second offset is 10 (256 rotations), the difference is -66, but when the first offset is 10 and the second offset is 200, it is +190. This difference can be converted to an angle and used to calculate the exact magnitude, or, for example, reduced by 16 times to provide a rough magnitude for the purpose of the action vector. Next, this is used in the movement function to rotate using the second comparison measurement and combined as the vector average of all pairs of action vectors to relocate the agent to the predicted observation. Based on quality criteria such as the minimum similarity score of the sub-region, vectors can be excluded from the average. Since the vectors generated by the pairs are redundant, it is understood that parts of the visual scene with a large environmental difference will have a lower weight compared to parts with a small environmental difference.

[0217] In a fourth method of comparing observations, the observation dispersion function 403-4 is used to determine a fourth observation comparison scale component. The observation dispersion function uses characteristics similar to the similarity function and the action function. The observation dispersion function 403-4 includes calculating the circular dispersion or circular standard deviation of the offsets of the sub-region around the agent. In effect, this observation dispersion function identifies a measure of the spread of the offsets on the circle around the agent.

[0218] The observation dispersion function 403-4 uses the same sub-region of the storage vector as that used by the observation action function 403-3 to determine the offset in the same way with respect to the rotation permutation of the sub-region of the current observation vector. For example, assume there are eight sub-regions providing eight offsets. The eight offsets represent how different parts of the field of view move differently when the agent moves towards or away from the target observation location. When the agent is at the target observation location, they are all the same and equal to the offset of a perfect vector match. As the agent moves away, they disperse and spread. By using circular statistics, such as the observation dispersion function 403-4, this spread can be measured, and the spread increases with the distance from the observation location. The circular statistics can be the circular dispersion or circular standard deviation of the eight offsets and provide a scalar value that forms the fourth observation comparison scale component. The circular statistics can be any measure of circular dispersion.

[0219] To execute the observation dispersion function 403-4, the observation action function 403-3 is first executed to obtain N offsets. At this point, the periodic circular dispersion of the offsets is calculated after first converting from the offset index range (0->255) to the appropriate angular range (e.g., radians (0->2*PI)). At this point, sub-region offsets that exceed a specified angle from the circular mean angle (e.g., PI / 2 radians) are identified, and as long as there are more than a specified number (e.g., 4) of sub-regions remaining, these sub-regions are excluded and the circular dispersion is recalculated.

[0220] The observation variance function 403-4 is another way to determine the similarity between the current observation and the saved observations. Unlike the observation similarity function 403-1, the observation action function is very robust to changes in the lighting environment because it depends on the offset of the sub-region rather than the entire observation vector. However, since the sub-vector used has a lower dimension than the complete observation vector o, there is a tendency for discontinuities to occur, and incorrect matches may be generated over long distances, which can have a significant impact on the comparison. For this reason, the observation similarity function and the first observation comparison scale component provided by the observation similarity function result in a better distance measurement of similarity.

[0221] The circle variance function 403-4 can be used to measure similarity in the same way as the similarity observation function 403-1, as it provides a scalar output indicating similarity. Additionally, the circle variance function 403-4 can be used to determine an environmental adjustment coefficient, which can be used to weight or modify other comparison scale components such as the first observation comparison scale component from the similarity function 403-1. In particular, when there is a high level of confidence from another observation comparison function such as the observation action function 403-3 that the agent is at the target observation, each offset between the saved vector and the current observation vector should be zero, so the circle variance function 403-4 should provide a high similarity measure in the complete space. Therefore, the output of other functions can be weighted based on the deviation provided by the fourth observation comparison scale component of the circle variance function 403-4. For example, combining the similarity function 403-1 with the circle variance function 403-4 means that the first observation comparison scale component from the similarity function 403-1 can be modified according to the fourth observation comparison scale component from the circle variance function 403-4 to reduce and quantify the influence of environmental conditions.

[0222] Examples of the observation comparison functions 403-1, 403-2, 403-3, and 403-4, including numerical examples, are provided above, but it should be understood that these functions can be executed with arbitrarily selected numbers of sub-regions, masks, channels, rotations, and array sizes.

[0223] The observation comparison scale may include one or more of the first, second, third, and fourth observation comparison scale components respectively obtained from the observation similarity function, the observation rotation function, the observation action function, and the observation dispersion function. For example, when the first observation is compared with the set of saved observations in the second step 302 of method 300, when the predicted observation is compared with the second observation in the fourth step 305-4 of hypothesis verification, and when the transition observation is compared with the predicted observation, the observation comparison scale is generated each time an observation is compared, which will be described in detail later.

[0224] It should be understood that the observation similarity function, the observation rotation function, the observation action function, and the observation dispersion function described above are examples of functions that can be used to compare the current observation with the target observation. Other functions or adaptations of the above functions can also be used. Generally, such functions are used to compare observations to determine how similar the current observation is to the saved observations, what rotation / direction the current observation has compared to the saved observations, and in what direction the location of the saved observations is from the current observation when the agent is not at the location of the saved observations. These characteristics can be mapped with respect to the saved observations or the target observations, as will be described later.

[0225] When navigating a map using observations or displacements, the movement of the agent is executed according to a movement function. In particular, the movement function depends on the predicted displacement required to reach the location of the predicted observation to be compared. Therefore, it is necessary to select the correct predicted displacement from a subset of hypotheses that may include multiple predicted displacements. In order to ensure that the correct predicted displacement is selected for the movement function, the predicted displacements and predicted observations stored in the hypothesis subset are linked to each other in a network of predicted observations connected by the predicted displacements, as described above. The link of an observation to one or more displacements within this network can be stored and maintained in any suitable way. For example, the set of stored observations and displacements, including the hypothesis subset, may be stored in a cyclic graph, an acyclic graph, or a directed graph. Alternatively, the observations and displacements are stored with metadata that identifies the displacements and observations to which each is connected, or are organized in a reference table that holds the links between each observation and displacement.

[0226] As described with respect to FIG. 3, the determination of the hypothesis is made based on the stored observation that is most similar to the first observation made by the agent. Therefore, the predicted displacement selected for the purpose of the movement function is the displacement linked to the observation that is most similar to the first observation within the set of stored observations and displacements.

[0227] When the predicted displacement from the hypothesis subset that is expected to lead from the most similar observation to the next predicted observation in the hypothesis subset is identified, the movement function is updated to be a movement based on the specific predicted displacement.

[0228] It is understood that the predicted displacement has both magnitude and direction.

[0229] The movement function may include one or more movement components. The first movement component of the movement function depends on the predicted displacement. In particular, the first movement component is a contribution to the movement function that decreases as the agent moves according to the predicted displacement. Thus, the first movement component is effectively weighted based on the magnitude of the predicted displacement. At the first observation location, since the magnitude of the predicted displacement is relatively large, the first movement component is weighted more strongly accordingly. When the agent starts moving according to the movement function, the magnitude of the predicted displacement is updated based on the actual displacement that the agent has already moved. Thus, as the agent moves towards the end point of the predicted displacement, the magnitude of the predicted displacement decreases, and the first movement component is weighted to gradually become weaker. In one example, the first movement component is a function of the predicted displacement and is expressed as predicted displacement - movement displacement. The moved displacement is the actual displacement estimated using odometry, visual inertial odometry, etc. In this example, the weighting described above may be a coefficient provided simply by subtracting the travel distance. As the travel distance increases, the first movement component decreases from an initial value equal to the predicted displacement. In other words, this can be considered as the predicted displacement being "resolved" as the agent moves along the predicted displacement. Thus, the value of the first movement component at any given time becomes a function of the "remaining displacement" of the predicted displacement. This provides the first value contributed to the movement function. This first value is scaled according to a scaling factor, and as a result, it is understood that the contribution of the first movement component to the movement function is appropriate considering additional movement components.

[0230] As the agent approaches the end point of the predicted displacement, the contribution of the predicted displacement to the movement function tends to become zero. In one example, when the measured actual displacement is equal to the predicted displacement and the entire predicted displacement is moved, the first movement component based on the predicted displacement becomes zero and no longer contributes to the agent's movement function. If the first movement component is the only movement component of the movement function and the movement function depends only on the predicted displacement, the position of the agent after moving the predicted displacement becomes a second location, and a second observation is recorded there.

[0231] Alternatively, the first movement component may stop contributing to the movement function before the entire predicted displacement is moved. In particular, the movement function may be configured to ignore the first movement component when the magnitude of the first movement component, in other words, the remaining distance of the predicted displacement, is within a displacement threshold distance from the end point of the predicted displacement. This is advantageous for reducing the effect of drift on the agent's movement. In particular, the estimation of the actual displacement using odometry is susceptible to the influence of drift-related errors, and the estimated value of the actual displacement (the actual distance traveled) may be smaller than the true value. Therefore, even when the estimated actual displacement is smaller than the predicted displacement, the agent may already have moved the entire predicted displacement. To avoid overshooting the end point of the predicted displacement, the first movement component may be stopped before the end point of the predicted displacement based on the displacement threshold distance. The displacement threshold distance can be set based on the specific environment or scenario in which the agent is used. In physical space, this can be an appropriate value such as 5 cm, 10 cm, 1 m, 5 m, 10 m, etc. before the end point of the predicted displacement, depending on the scenario.

[0232] The second movement component that may be included in the movement function depends on the observation action function described above with respect to observation comparison. When using this second movement component in the movement function, it is necessary to continuously, repeatedly, or sequentially perform transition observations between the first location in the space where the first observation is made and the second location in the space where the second observation is made. In other words, as the agent moves away from the first location, transition observations are made.

[0233] In this way, one or more transition observations are made along the movement path of the agent. Each transition observation is a type of "current" observation that is compared with a predicted observation that is expected to be derived from the predicted displacement according to the hypothesis. From each comparison, an action vector is generated that defines both the expected magnitude and direction to a location where the predicted observation might be found, using the observation action function described above. The action vector forms the second movement component of the movement function. Similar to the first movement component of the movement function, the contribution of the second movement component is scaled by a scaling factor and may also be weighted. In particular, the second movement component may be weighted such that the contribution of the second movement component to the movement function becomes stronger as the action vector of the second movement component indicates that the agent has approached the location of the predicted observation. For example, the magnitude of the second vector resulting from the observation action function is configured to increase as the agent approaches the location of the predicted observation. This increase in magnitude may directly provide a weight that strengthens the second movement component as the agent approaches the location of the predicted observation.

[0234] The second movement component may be used alone as the only contribution to the movement function, or it may be used in combination with the first movement component. Advantageously, both the first and second movement components are included, and when both of these components contribute to the movement function, the movement of the agent depends on both the predicted displacement and the transition comparison between the predicted observation and the transition observation. This means that there are two separate data sources used to move to and position the predicted observation within the space. This is more reliable and accurate than using only one set of data, effectively providing redundancy. Further, the first movement component and the second movement component are complementary to each other in that the relative weighting of the first movement component decreases the contribution of the first movement component to the movement function as the agent approaches the end of the predicted movement, whereas the relative weighting of the second movement component increases the contribution of the second movement component to the movement function as it approaches the location of the predicted observation. Thus, the first movement component and the second movement component operate in opposite ways.

[0235] If the hypothesis is correct, at least partially correct, or similar, and the attributes related to the hypothesis are at least similar to the attributes of the space occupied by the agent, the end point of the predicted displacement should be at least close to the location where the predicted observation was found. If the hypothesis is correct, at least partially correct, or similar, and the attributes related to the hypothesis are at least similar to the attributes of the space occupied by the agent, the end point of the predicted displacement should be at least close to the location where the predicted observation was found.

[0236] The relationship between the first movement component and the second movement component regarding their contributions to the agent's movement function can take any suitable form. An example of the relationship between these two movement components will be described with reference to FIG. 7. FIG. 7 shows a schematic map of an agent moving between a first location and a second location within a space.

[0237] In a first example of the relationship between the first and second movement components of the movement function, as shown in the first graph 701, the relationship is simple and depends only on the relative weighting of the first and second movement components. As the agent moves away from the first location 701-1, the relative weightings of the first and second movement components decrease and increase respectively, so that as the agent approaches 701-1, the movement component of the first location becomes naturally dominant with respect to its contribution to the movement function. This is because the transition observations made for the purpose of the observation action function are usually very different from the predicted observations because the distance between the capture points of these observations is likely to be large and thus the observations are in different regions of space. This is because the transition observations made for the purpose of the observation action function are usually very different from the predicted observations because the distance between the capture points of these observations is likely to be large and thus the observations are in different regions of space. Since the observation action function depends on the comparison of observations and the identification of similarities between observations, it is less reliable when the observations are very different. Therefore, the fact that the first movement component becomes dominant when the agent is close to the first location 701-1 of the agent is advantageous because the predicted displacement can be used to approach the location where the predicted observations should be. This is more reliable than using only the second movement component.

[0238] In the first graph 701, as the agent moves, a series of transition observations 701-2 are made, and the second movement component, particularly the observation action function responsible for the second movement component, is updated periodically, continuously, or iteratively. The agent moves along path 701-3 according to the movement function and moves to the second location 701-6 in the space. As the agent approaches the end point of the displacement from the first movement component, the second movement component from the observation action function becomes dominant and relocates the agent to the second location 701-6. The second location 701-6 was not recognized before the agent moved. Rather, it should be understood that the second location 701-6 represents the position that is the same or most consistent in the space with the predicted observation (target observation) linked to the predicted displacement. The second location 701-6 is determined by the position where the second from the observation action function settles at a point in the space, and the magnitude and direction of the observation action function vector indicate that the second location 701-6 is the location of the predicted observation or its optimal matching.

[0239] In a second example of the relationship between the first movement component and the second movement component of the movement function, as shown in the second graph 702, the relationship is sequential, and the movement function initially depends only on the first movement component. When the expiration date of the first movement component expires, it depends only on the second movement component. Since this relationship depends on the first movement component not being a predictive observation, there is an advantage that it does not require transition observations for most of the agent's movement between the first location 702-1 and the predicted displacement of the second location 702-6. According to this relationship, the agent is configured to move from the first location 702-1 via the first path 702-3 corresponding to movement following only the first movement component based on the predicted displacement. When the predicted displacement is depleted and the agent reaches the end point of the predicted displacement at the intermediate location 702-5 or reaches a predetermined threshold from the end point of the displacement, the movement function switches from using the first movement component to using the second movement component. At this point, the agent performs sequential or iterative transition observations 702-2 and executes an observation action function to obtain a vector of magnitude and direction indicating the location of the predictive observation. Using this information, the agent follows the second path 702-4 and heads towards the second location 702-6. Thus, in the second relationship, the agent actually moves by the first displacement before repositioning using the observation action function.

[0240] In a third example of the relationship between the first movement component and the second movement component of the movement function, as shown in the third graph 703, this relationship is a winner-takes-all type of relationship, and the movement function depends only on the component with the highest weight among the first movement component and the second movement component at a specific time and position within the space along the path to the second location. As described above, the first movement component that depends on the predicted displacement naturally has a higher weighting than the second movement component closer to the first location 703-1. Therefore, the agent moves along the first path 703-3 based on the first movement component and the predicted displacement. During this movement along the first path 703-3, the agent performs a series of transition observations 703-2 for the purpose of updating the second movement component. This purpose is to determine the time when the second movement component becomes dominant over the first movement component. The frequency of the transition observations made along the first path 703-3 is not as important when compared to the first relationship shown in the first graph 701. This is because the transition observations 703-2 along the first path 703-3 are made only to compare the first movement component with the second movement component, not to guide the agent itself. Increasing the frequency only increases the accuracy of determining that the "winner" in the winner-takes-all relationship becomes the second movement component. At the point 703-5 along the path, this determination is made when it is determined that the second movement component is dominant over the first movement component. At this point, according to the observation action function, the second movement component is used in the movement function and the agent is repositioned to the second location 703-6. In this phase, the agent follows the second path 703-4 and performs further transition observations 703-2. This is necessary to obtain a vector from the observation action function and provide the direction and magnitude for the agent to move to the second location 703-6.

[0241] Therefore, the movement function uses two components to relocate to the location of the predicted observation. The three relationships between these two components are exemplary, and any suitable method or combination thereof may also be appropriate.

[0242] The ability of the movement function to relocate the agent to the location of the predicted observation depends on the observation-action function, which depends on whether similarities are found between the current observation and the predicted observation. Therefore, if the similarity between observations falls below a certain level, the action function is not relocated. In this case, the agent may move according to the first movement component of the movement function based on the predicted displacement. If it is determined that the agent cannot be relocated because there is no similarity between the current observation (transition observation) and the predicted observation, and the predicted observation is not found at or near the end point of the predicted displacement, the hypothesis is rejected. In this case, the agent moves in the reverse direction along the predicted displacement and returns to the position within the space of the previous observation where the hypothesis was determined. Next, the hypothesis is adjusted to another hypothesis related to a different subset of the saved observations and displacement hypotheses.

[0243] Therefore, the movement function provides the agent with the ability to move to the location where the predicted observation was found, the "optimal matching" location in the space for the predicted observation (e.g., if the agent relocates to a position that does not exactly match the predicted observation but is still similar), or the location where the predicted observation was not found. Since the movement function may consist of two elements, there may also be a difference between the predicted displacement and the actual displacement. For example, if the optimal or exact matching of the predicted and transition observations is at the end point of the predicted displacement, the agent may not need to use the second movement component to relocate. Alternatively, it may be necessary to relocate the agent slightly to make the actual movement distance equivalent to the predicted movement distance. Finally, the second movement component may relocate the agent over a wider range compared to the predicted displacement. In this case, the predicted displacement is not similar to the actual displacement.

[0244] These possibilities ultimately provide this method with a way to determine in two separate dimensions how closely a series of observations and displacements executed by the agent resemble the attributes of the hypothesis. This ultimately helps determine whether to adjust, maintain, or confirm the hypothesis as described above. In particular, the first dimension is the similarity between the current observation (or an observation for which method 300 is continued or repeated for two or more observations) and a hypothesized subset of the saved observations. This similarity is determined by an observation comparison metric and is referred to as "style". The second dimension is the similarity between the current displacement (or a displacement for which method 300 is continued or repeated for two or more displacements) and a hypothesized subset of the saved displacements. This similarity is determined by a displacement comparison metric and is referred to as "composition". The displacement comparison metric is determined by comparing the actual displacement and the predicted displacement.

[0245] Generally, there are four combinations of style and composition. These are high style and high composition, high style and low composition, low style and high composition, and low style and low composition. What each of these possibilities means for the hypothesis will be explained in detail below.

[0246] First, the fact that the style and composition are high indicates that the observations made by the agent are very similar / identical to the predicted observations of the hypothesis subset, and the actual displacement made by the agent is very similar / identical to the predicted displacement of the hypothesis subset. "Very similar" means that, as explained above with respect to the fifth step of hypothesis verification 305-5, both the observation comparison scale and the displacement comparison scale exceed a second, higher observation threshold level and a second, higher displacement threshold level. If the displacement and the observation comparison scale match or exceed both of these levels, it is determined that the space in which the agent exists contains attributes that match the identity defined by the hypothesis subset. As described above, there is no need to label the identity, but a label can be associated that defines the attributes associated with the identity. When the attributes of the space match a plurality of consecutive predicted observations and hypothesis subsets of displacements, the reliability that the attributes within the space match the identity increases. Therefore, when determining whether the attributes of the space match the identity, it may be necessary to compare a plurality of consecutive predicted observations and displacements and match them to a sufficient level.

[0247] Second, the fact that the style is high and the configuration is low indicates that the observations made by the agent are very similar / identical to the predicted observations of the hypothesis subset, but the actual displacements made by the agent are not similar / identical to the predicted displacements of the hypothesis subset. "Very similar" means that, with respect to the fifth step of hypothesis verification 305-5, the observation comparison scale exceeds the second high observation threshold level, but the displacement comparison scale does not exceed the second high displacement threshold. In terms of the space in which the agent exists, this means that although the space appears to be the same as or very similar to the attributes defined by the hypothesis subset from the perspective of the sensor (observation) point data, at least one of the size, position, and / or distance between the attributes is not as predicted according to the hypothesis. This relationship is called "context". In the above examples shown in FIGS. 1 and 2, the context is represented by a room with a similar appearance equipped with furniture similar to that in FIG. 1, but perhaps the room is larger and the furniture arrangement has been changed (e.g., the central element has been moved to one side). In FIG. 2, the context may be found when the chair is photographed from the back or when the ratio of the chair is changed (e.g., the legs are made shorter or the backrest is made longer).

[0248] Thirdly, the fact that the style is low and the configuration is high indicates that the observations made by the agent are not similar / identical to the predicted observations of the hypothesis subset, but the actual displacements made by the agent are very similar / identical to the predicted displacements of the hypothesis subset. "Very similar" means that the displacement comparison measure exceeds a second, higher displacement threshold level, but for the fifth step of hypothesis verification 305-5, the observation comparison measure does not exceed a second, higher observation threshold. In terms of the space in which the agent exists, this means that, according to the hypothesis, the size, location, and / or distance between points of that space appear to be very similar, but do not appear to be / are not the same attributes defined by the hypothesis. This relationship is called a "class". In the examples of FIGS. 1 and 2 above, the class is represented by rooms of the same size and structure, but the color and furniture may be different. In FIG. 2, a class may be found when the color of the chairs is different, as in some types of chairs designed for dining at a table, or when the detailed style is different while maintaining the same ratio. In terms of navigation, it may be different floors of a building, and even though the floor plan is the same, different companies may provide the decoration. Multiple observation comparison measures may be used to determine the class relationship between the attributes within the space in which the agent exists and the saved observation data. While the similarity indicated by the first observation comparison measure component may be low, the second and / or third observation comparison measure components can be used from the observation similarity function 403-1. If the observation action function does not provide convergence to a point, there is no match, the basis for comparison is invalid, and in that case, the hypothesis is rejected.

[0249] Finally, a low style and configuration indicates that the observations made by the agent are not similar / identical to the predicted observations of the hypothesis subset, and that the actual displacement made by the agent is not similar / identical to the predicted displacement of the hypothesis subset. In this case, with respect to the fifth step of hypothesis verification 305-5, the displacement comparison measure does not exceed the second higher displacement threshold level, and the observation comparison measure does not exceed the second higher observation threshold. In terms of the space in which the agent exists, this means that the space appears to have no relation to the structure and appearance of the attributes related to the hypothesis. The hypothesis may be maintained for further comparison or rejected and adjusted as described above.

[0250] It is possible to identify the characteristics of space from the perspectives of two dimensions of composition and style (structure and appearance) by comparing the current observations and displacements with the saved observations and displacements for one or more hypothesis subsets. When the space is repeatedly or periodically evaluated by an agent, these dimensions can also be used to determine changes in the space. In this scenario, the saved hypothesis subset may include observations and displacements directly obtained from the same space in previous operations by the agent. Using this information, the agent can detect what has changed in the space over time in terms of structure and appearance. The step of identifying differences can be useful, for example, in route planning applications or in searching for alternative routes. The agent can save the average displacement up to a specific saved observation, and this average displacement is updated each time the agent displaces to the observation. This average displacement provides a historical average of past displacements made for a specific observation. Additionally, the standard deviation of the previous displacements made by the agent for a specific observation may also be saved. Statistical analysis can be performed using the average displacement and the standard deviation to determine the likelihood that the space has changed. For example, if the new displacement for an observation is significantly different from the average displacement and / or different from the standard deviation of the previous displacements, it is judged that the space is likely to have changed. This can occur, for example, when the attributes of the space have been moved. To perform the statistical analysis, historical records of each observation and displacement, such as the average displacement, observation, standard deviation, etc., can be saved. This historical record may be updated with each exploration by the agent.

[0251] The second history record can also be saved and updated in a similar way. The second history record is related to the N most recent observations and displacements made by the agent when the agent moves to a specific observation. In particular, it is related to the similarity or overall similarity score between these N most recent observations and displacements and the corresponding saved observations and displacements that they are compared to. Thus, the second history record represents the localization reliability indicating to what extent the system is confident that the agent is localized to the saved observations and displacements while moving towards a specific observation. If the second history record indicates that the N most recent observations made by the agent match very well with the N most recent observations that are saved, from the perspective of observations and displacements, the reliability that the agent is appropriately localized to the subset or path it is following is obtained. If the displacement made for a specific observation, or the specific observation itself, is significantly different from the corresponding saved observation and displacement, the second history record can be used to reliably indicate and determine that the space has changed in some way with respect to that specific observation and displacement. For example, when N = 5, the last 5 displacements made by the agent may exactly match the saved displacements. If the 6 displacement made is different from the 6 displacement that is saved, the fact that the second history record shows a high level of match (in this case, identity) for the previous 5 displacements means that 6 it can be easily determined that the space has changed with respect to the

[0252] Furthermore, as described above, hypotheses can be separately confirmed, maintained, or rejected based on observations or displacements, but when determining the outcome of a hypothesis, it is advantageous to use both observations and displacements, i.e., composition and style. Adjusting the hypothesis by forming a two-dimensional similarity measure by combining an observation comparison measure and a displacement comparison measure may include adjusting the hypothesis when the two-dimensional similarity measure falls below a first two-dimensional similarity measure threshold level. Maintaining the hypothesis may include maintaining the hypothesis when the two-dimensional similarity measure exceeds the first two-dimensional similarity measure threshold level but does not exceed a higher second two-dimensional similarity measure threshold level. When exceeding this second two-dimensional similarity measure threshold level (also referred to as the hypothesis confirmation two-dimensional similarity measure threshold level), the step of confirming the hypothesis may be performed. If there is a change in the visual identity of observations encountered at different times, the agent may update the saved observations using the current observation information if it can be confident that the current and past observations are captured with the same identity. Since both observations are represented by normalized vectors, the final observation can be made to approximately match the past or current observation vector by applying optional weighting to each and averaging the two vectors. For example, multiplying the past observation by 5 and the current observation by 1 and re-normalizing the total vector will make the final observation vector closer to the past vector.

[0253] Adjusting a hypothesis involves rejecting the hypothesis and going back to previous observations to identify a second hypothesis. In particular, in any observation, if the current observation does not match the predicted observation with appropriate reliability and / or if the actual displacement does not match the predicted displacement with appropriate reliability, the hypothesis is rejected. Next, the agent is configured to return to a previous observation by making a movement opposite to the movement made to reach the current observation. This may include returning to the first observation, but this is not essential. Next, method 300 repeats to verify different hypotheses selected based on a comparison of each of a plurality of hypothesis subsets of previous observations and predicted observations. Optionally, the agent may save the observations and displacements even if the observations and displacements do not match the hypothesis, and the hypothesis is then rejected. Saving these observations and displacements is useful when generating a map of the space for future navigation and attribute identification purposes. Thus, the agent may make and save one or more observations and displacements before returning to a previous observation even if the hypothesis is rejected.

[0254] Confirming an observation includes, but is not limited to, identifying or exploring an object, image, or destination within a space. When an attribute is found, a reference to that attribute (including observations and displacements within the space by the agent) is saved in memory. These may form part of a map of the space in which the agent exists. If the hypothesis is associated with a labeled identity or attribute, the agent can also save a reference to the label in memory along with the observations and displacements made within the space. If the hypothesis is based on a previous visit to the same space, the hypothesis subset effectively defines the attributes previously observed in the space, so the hypothesis subset may be updated with the set of displacements and observations most recently made by the agent when verifying the hypothesis. In this way, the hypothesis subset is kept up-to-date with respect to a particular space.

[0255] It should be understood that the above method may be executed by one agent or may be executed by one or more agents cooperating. If one or more agents are included in the space, the one or more agents may communicate with each other or access the same memory. For example, multiple agents may each be linked to the same server or computer system. Multiple agents may update the map of the space simultaneously or verify one or more hypotheses simultaneously. The generation of the map will be described in detail below.

[0256] In the above description, method 300 for analyzing a space and identifying the attributes of the space has been described. The attributes may be objects or images, in which case method 300 is used for the purpose of object or image recognition. Alternatively, they may be destinations, in which case method 300 is used for the purpose of navigation. Method 300 depends on the agent making observations and displacements within the space and comparing them with the saved observations and displacements. Therefore, one principle of method 300 is that the agent can compare observations with displacements. This is shown in FIGS. 4 to 6 and the above description. Another principle is the agent's ability to move, which is shown in FIG. 7 and the above description.

[0257] Next, further aspects for complementing method 300 will be described. These aspects may be incorporated into method 300 or may be executed or implemented separately.

[0258] A first further aspect is the generation and storage of linked observations and displacements that form an identity and a hypothesis subset. This may be considered, as briefly touched on before, a "training process". The process of forming an identity may be performed either specifically with respect to the space in which the agent is expected to exist, or in another space. There are advantages to both of these implementations. If an identity is formed in the same space that the agent is configured to later investigate, after further exploring that space, it becomes possible to identify how the attributes of that space change over time. If an identity is formed in a separate space, it becomes possible to identify that the attributes of one space repeatedly appear in another space, or to identify similarities and differences between spaces.

[0259] The process of forming and storing an identity is described here with reference to FIG. 8. FIG. 8 shows a schematic diagram 800 of an identity. It is necessary to consider that observations and displacements are not stored. The agent is "dropped" or activated into a space that has hitherto been unknown and unexplored. An initial observation o1 is made within the space. The initial observation o1 may be made based on low-level features detected within the space, but this is not essential. The use of low-level features will be described separately later as a further complementary aspect. The initial observation can be selected either randomly or based on a learned process. From the initial observation o1, an initial displacement is made from the location of o1 and a new observation o2 is made, and the displacement between o1 and o2 is stored in memory as d12, and the reverse displacement from o2 to o1 (negative value of d12) is stored in memory as d21.

[0260] Optionally, an additional effort value, which is a better measure of the characteristics (such as energy usage and / or vibrations encountered) optimized in routing, is stored.

[0261] Next, a continuous set of observations Oi {oi,1...oi,n} and its corresponding set of displacements Di {di,1...di,n-1} are moved by the agent and stored in the same way. This forms an identity. The set of observations and displacements that form the identity describe the predicted path through space and the expected environment at points along the path. The steps of moving along the path allow the agent to verify the prediction (i.e., the hypothesis) by comparing the measured displacements and observations one by one with the prediction. The continuous set of observations and displacements that form the identity is stored as described above. In some examples, the identity can be stored as a subset of a larger stored set, and the larger stored set contains multiple subsets, each subset defining a path in the same or different space. Thus, identities can be linked to other identities and may share one or more common observations and / or displacements.

[0262] As shown in FIG. 8, a series of observations o1 to o4 and corresponding displacements are obtained in a closed loop to form an identity, and then, using the styles and composition dimensions described above, for successive subsets or the entire set, future observations and displacements can be verified in sequential order. This identity can be used for purposes such as identifying or locating a specific object, extracting general repeating classes such as corridors and rooms in a building, objects and navigation terrain (not limited to these), and extracting contexts of areas or objects that have similar components or characteristics (e.g., forests, offices, roads, or wheeled vehicles, flying objects).

[0263] While or after an identity is stored as a subset of observations and displacements, further optional processes may involve the use of another system to determine attributes associated with the identity. For example, for the purpose of navigation to a destination in space, the identity can be associated with a route. Alternatively, the identity may be associated with an object such as a chair, or an image such as a cartoon chair or a painting of a chair. A training process may be used to correctly associate the identity with these attributes. The training process may involve, for example, identifying the attributes via established training algorithms such as the use of training data to train a classifier. It should be understood that any machine learning process may be used and may include the use of a neural network or the like.

[0264] The methods and systems described above show how an agent can use observations and / or displacements to identify, determine, or navigate a space. Once an observation has been made and the space has been at least partially explored, it is convenient to save a record of the observation by generating a map. It should be understood that map generation can be performed independently or in combination with the above identity generation and attribute determination / hypothesis verification.

[0265] An agent can generate a new map if the space is completely unknown from the agent's previous explorations, and can update an existing map in memory if the space has been previously visited by the agent.

[0266] To generate a map, the agent makes a first observation in the space. This may be the same observation as the first observation made in step 301 of method 300 for verifying the hypothesis described above. The first observation is stored in memory as indicating the corresponding "first location" of the map.

[0267] Next, the agent is configured to move away from the first location and make one or more additional observations.

[0268] If the map is being created or updated in combination with the agent performing hypothesis verification, the agent moves according to the movement function described above. In this case, the agent is configured to make one or more new observations when moving away from the first location. The new observations may be made either while the agent is moving according to the movement function (during the transition) or afterwards. In this regard, the agent can wait until it arrives at the second location of the second observation of method 300 described above and then make a "new" observation, in which case the second observation becomes the new observation.

[0269] The new observations are compared with the first observations using one or more of the observation comparison functions described above. This provides one or more observation comparison metrics. The observation comparison metrics are compared with a mapping threshold that defines the maximum probability that consecutive observations must not fall below in order to be recorded on the map. In other words, in the method of generating the map, it is necessary that the observations stored in the map are not too similar. This prevents unnecessary points from being added to the map, simplifies the map, and minimizes the memory required for storing and accessing the map. If, as a result of the observation comparison between the first observation and the new observation, the observation comparison metric is lower than the mapping threshold, New observations are stored in the memory of "new locations" on the map. The displacement between the first observation and the new observation is measured and / or obtained from the movement function and is also stored on the map together with the inverse displacement from the new observation to the first observation. When the map is generated while hypothesis verification is also being performed, the agent performs two comparisons. The first is a comparison between the current observation and the predicted observation (to verify the hypothesis), and the second is a comparison between the current observation and previous observations (such as the first observation), to determine whether these observations are not similar enough to be included in the map. Thus, even if the hypothesis is rejected and the current observation does not sufficiently match the predicted observation, the agent can make observations for the purpose of adding or updating to the map or use the current observation by comparing it with previous observations.

[0270] If the mapping threshold is exceeded, the new observation is not stored on the map. Instead, the agent is configured to move again with additional displacement before making further observations until the mapping threshold is no longer exceeded.

[0271] When the new observation is stored on the map together with the displacement between the first observation and the new observation, the new observation effectively becomes the first observation (or previous observation), and this process is repeated for further displacements and further observations. This process iteratively constructs a connected network of observations that forms a map describing the space in which the agent exists.

[0272] The map of the space can also be generated separately from the identification of the attributes of the space. The map generation method in this case is almost the same. The agent makes a first observation and stores it in memory as the first location of the map. Next, the agent moves from the first location of the first observation in an unexplored direction that is physically or logically passable. The movement of the agent is recorded as the first displacement. As described above, the agent makes one or more new observations at a location away from the first location, compares the first observation with the new observations, and determines whether the mapping threshold is exceeded. If the mapping threshold is not exceeded, the new observations are stored in the map in a state where the first observation and the new observations are linked by the first displacement and the reciprocal of the first displacement. The relative direction between the first observation and the new observations is marked as "explored". Thereafter, this process is repeated with further displacements and observations in unexplored directions from the new observation locations. The use of the mapping threshold means that the number and density of observations made in a particular space depend on the complexity and characteristics of that space. In an environment rich in features, the distance at which the similarity between observations falls below the mapping threshold is shorter and decreases more rapidly, so more observations are made. In an environment with few features, the space is not so complex, so the number of observations required is reduced. Therefore, the method of mapping and / or identifying the attributes is performed in such a way that it automatically adapts to the constraints of the space in which the agent observes or moves. Therefore, these methods are sensitive to the complexity of the environment.

[0273] Multiple agents may contribute to the same map. In particular, multiple agents are each configured to contribute to and access a shared database or memory, and update the shared database with a set of observations and displacements to generate a map. When a hypothesis is verified and an attribute is identified, an ID associated with the attribute is stored on the map and labeled as such if possible. Thus, the map may include multiple identified attributes and the stored labels of those attributes. Agents can contribute to map generation in real-time or near real-time, each communicatively coupled to the shared database via appropriate means. For example, in a physical space, multiple agents may be a fleet of drones, robots, vehicles, etc., each of which can communicate with a computer system or a distributed computer system such as a plurality of servers.

[0274] Subsequently, this map can be used for navigation of the space. In particular, a destination on the map can be selected as the destination for moving an agent. This destination is represented by observations that form part of the map.

[0275] Navigation of the map functions in a similar manner to method 300 for verifying a hypothesis and determining the attributes of a space. When navigating the map, the attributes serve substantially as destinations. The process of navigation will be described in detail here.

[0276] The map is generated as described above, stored in memory, forms a set of stored observations and displacements, and at least a subset of the stored observations and displacements is connected to observations associated with a destination or forms a network. Thus, this map is similar to the hypothesis subset described above.

[0277] In some embodiments, the map does not include displacement data. In some embodiments, the map may include data regarding characteristics associated with the observations, such as action vectors or similarities. These characteristics may form a vector field or a scalar field that functions as a map for navigation purposes.

[0278] The agent is configured to make a first observation in space and identify which of the stored observations in the stored map the agent is closest to. In particular, the first observation is compared to each observation connected to the destination within the stored map. From these comparisons, the observation closest to the first observation is identified. Next, the agent is moved to relocate to the location of the closest matching observation.

[0279] From the location of the closest matching observation on the map, a route is planned using a routing algorithm. Any suitable routing algorithm can be used, such as Dijkstra's algorithm, which determines a route that minimizes distance, effort, vibration, or other measurements obtainable from the displacements stored in the map.

[0280] Once the route is planned, the agent sequentially moves along the route, and the process is repeated as the agent moves from observation point to observation point on the map. The agent starts this process from the "closest matching observation", and the route may include, for example, multiple observations between the closest matching observation and the destination. The observation where the agent currently is is the current observation, and the next observation on the route to the destination on the map is called the target observation.

[0281] The iterative process includes moving the agent in the direction of the target observation using the saved displacement vector between the current observation and the target observation. This is similar to moving the agent according to the movement function according to the predicted displacement, as described above with reference to the hypothesis verification method 300. Similar to hypothesis verification, this process also includes performing one or more observation comparisons using one or more of an observation similarity function, an observation rotation function, and an observation action function. For example, the saved displacement to reach the target observation can be used in combination with the observation action function to move the agent from the current observation to the target observation. Thus, the agent can make multiple transition observations while moving between the current observation and the target observation and re-localize to the location of the target observation using the observation action function. This is similar to the movement function including the first movement component and the second movement component described above, and when changing the movement direction and reproducing the displacement between the current observation and the target observation, odometry drift and uncertainty can be considered.

[0282] When the observation comparison metric (or metrics) exceeds a predetermined threshold or a continuously calculated threshold, it is determined that the agent has reached the target observation. This is similar to determining the match between the second observation and the predicted observation in the hypothesis verification method 300. When it is determined that the agent has reached the target observation, this process is repeated for the next observation on the route to the destination until the destination is reached. Thus, the target observation becomes the current observation, the next observation becomes the target observation, and the displacement to move is the displacement between the current observation and the target observation. When the agent reaches the observation associated with the destination, the agent is configured to wait at that location unless otherwise instructed.

[0283] In this way, by using observation, displacement, and movement functions, the attributes and features of a space can be identified, a map of the space can be generated, and the space can be navigated. It should be noted that one or more, or all, of these processes may be executed simultaneously. For example, an agent can be instructed to follow an existing route to a destination, identify attributes and features along the way to the destination, and use transition observations to add them to an existing map. Therefore, the agent can perform learning and environmental navigation simultaneously.

[0284] In the above method and system for analyzing a logical or physical space, which includes identifying the attributes / features of a space, generating a map of the space, and navigating the map of the space, the agent is configured to make observations, compare them, and move between these observations. As described above, the logical space or physical space does not need to be recognized by the agent, or the agent does not need to have visited it before. Therefore, there are several sets of initial conditions that the agent may encounter when placed in a space.

[0285] In the first example, the agent is placed in a space that has been visited before, and the agent can access a saved map or saved subset that includes the observations and displacements made within that space. In this case, when making the first observation, the first observation may be identical to, or match, a saved observation (such as a predicted observation of a hypothesis subset or an observation along the route to a destination). In this case, there is no need to relocate the agent or navigate the saved route to the destination before executing method 300 to verify the hypothesis.

[0286] In a second example, the agent is placed in a previously visited space, and the agent can access a stored map or a stored subset that includes observations and displacements made within that space. In this case, when making a first observation, the first observation may not be identical to any of the stored observations in the map or the stored subset. Since the location of the agent is unknown, it is not possible to use the stored displacements to relocate the agent on the stored observations. Rather, the agent first needs to relocate to the stored observations without using predicted displacements in order to understand the location on the map or the position relative to a subset of stored observations and displacements. To do this, the first observation is compared with each of the stored observations, and an observation action function is used to determine a vector to one of the stored observations (here called the target observation). Next, the agent moves along this vector and repeatedly or continuously updates further transition observations to relocate to the target observation.

[0287] The selection of which stored observation becomes the target observation can be determined in any reasonable way. In one example, the stored observations may be compared with the first rotation for each rotation using an observation similarity function. Next, the most similar stored observation may be selected for the purpose of comparing with the first observation to determine an action vector from the observation action function.

[0288] Alternatively, an action vector can be determined for each of the stored observations, and the strongest one among them can be selected to relocate the agent to the stored observations. The action vector is determined from the observation action function as described above and is weighted based on the similarity between the sub-region of the stored observation vector and the corresponding double region of the current observation vector, or based on an overall similarity measure, and thus becomes stronger.

[0289] In the third example, the agent can be placed in a previously unvisited space, in which case the agent can access a saved subset that includes observations and displacements from different spaces or multiple spaces. In this case, when making the first observation, the first observation may not be identical to any of the saved observations in the saved subset. Since the location of the agent is unknown, it is not possible to use the saved displacements to relocate the agent onto the saved observations. Rather, the agent first needs to relocate to the saved observations without using predicted displacements in order to understand its position relative to the subset of saved observations and displacements. To do this, the first observation is compared to each saved observation, and an observation action function is used to determine a vector to one of the saved observations (referred to here as the target observation). Next, the agent moves along this vector, repeating or continuously updating further transition observations to relocate to the target observation.

[0290] The selection of which saved observation becomes the target observation can be determined in any reasonable way. In one example, the saved observations may be compared to the first rotation for each rotation using an observation similarity function. Next, the most similar saved observation may be selected for the purpose of comparing it to the first observation to determine the action vector from the observation action function.

[0291] Alternatively, an action vector can be determined for each saved observation, and the strongest one among them can be selected to relocate the agent to the saved observation.

[0292] However, unlike the first and second examples above, in this third example, the subset of saved observations and displacements is from a space different from the space where the agent currently exists. Therefore, there may be no saved observations in the space where the agent exists. In this case, the agent may not be able to re-localize to the saved observations. When the agent is verifying a hypothesis, if the subset of saved observations and displacements is the hypothesis subset, the agent may reject the hypothesis related to the hypothesis subset if it cannot position itself to the saved observations of the hypothesis subset. Thereafter, the agent can return to a location within the space of the first observation by a negative displacement. From here, the agent can select another hypothesis to verify and the corresponding hypothesis subset. The agent can repeat this process as needed for all saved hypotheses and subsets of hypotheses. If all hypotheses are rejected, i.e., if the space does not contain recognizable features or attributes corresponding to the subset of saved observations and displacements, the agent enters the exploration mode and, as described above, makes observations and displacements to generate a map of the space.

[0293] In the fourth example, the agent is placed in a space that has not been visited before, and the agent does not have access to the saved subset including observations and displacements. In this case, the agent enters the exploration mode as described above, maps the unknown space where the agent exists, and finally saves a series of observations and displacements to describe the space.

[0294] The above description focuses on an agent that makes sequential observations and displacements to perform one or more of map generation, navigating routes within a space, or identifying features of a space. It is possible to perform these methods using only observations and displacements, but additional elements can also be incorporated to gain further efficiency and accuracy advantages.

[0295] In particular, low-level features of the space may be identified or detected and used to inform the process of observations within the space and movement between observations. In particular, observations by an agent include spatial information indicating the environment of the space in the agent's local vicinity. Therefore, it is possible to extract features from the observations or observation data to obtain more information about the agent's vicinity. For example, clues in the environment such as high-contrast lines, colors, edges, shapes, etc. can be extracted to determine detailed information about the environment and used as a reference for decision-making and actions. For example, in map generation where an agent is configured to sequentially perform observations and movements to explore a space, the selection of the direction to move the agent after each observation may be controlled by low-level features. Specifically, after each observation, the selection of the agent's next movement direction may be weighted based on low-level features and / or the previously navigated direction. For example, a first weighting can be applied in a direction that brings the agent closer to a previous observation when the agent has moved. Since the purpose of map generation is for the agent to explore and map the space, it is beneficial to encourage the agent to move away from previous observations, and this is achieved by this first weighting. A second weighting may be applied to the direction based on low-level features. For example, if an edge or a high-contrast line appears to be proceeding in a particular direction, that direction may be weighted higher than a direction without low-level features. In the physical space, the agent may be a vehicle. The low-level features to be extracted may include roads (from edge detection, etc.). The direction in which a road is indicated from the current observation or the direction leading to the road may be weighted higher than the direction from the current observation without a road. This second weighting based on low-level features encourages the agent to follow or approach potential features of the environment. For example, by following a road, the likelihood of being able to explore the space increases, and the likelihood of the agent making observations that can be easily traversed and connected also increases.Roads are artificial structures added by humans to the environment, but it has been found that there are also things in the natural environment that bring similar benefits, such as the banks of waterways, the edges of forests with grasslands, and the paths of animals searching for food carved by plants. By following such characteristics in both the exploration stage and the navigation stage, agents can confirm that there are additional guides that can ensure reliable movement even when there are large differences in the environment, such as thick fog. Also, it is necessary to note that such a function is also useful when an agent recognizes the ID of an object such as a chair. Guide the movement of the observation points along the edges of the strings formed by the legs, seat surface, and backrest of the chair, and ensure high robustness when reproducing the identity even with 3D rotation.

[0296] It should be understood that the first direction weighting and the second direction weighting described here can be used separately (without the other) or in combination.

[0297] Low-level features identified or detected by observations or observation data can also be linked to a series of actions performed by the agent. These actions allow the agent to approach the features in a similar way when encountering a new part of a space or object for the first time and when encountering the same part of the space or object again. For example, when a high-contrast line is identified or detected, the agent may be configured to move along the high-contrast line. Since the agent can move along roads, sidewalks, or paths, it may be configured to move to the vanishing point along a consistent line in the environment. Such structured actions simplify the problem of robust navigation. The agent can move in another predetermined direction relative to the low-level feature, such as perpendicular, parallel, clockwise, or counterclockwise to the low-level feature.

[0298] Observations can also be made in response to the presence of low-level features in the environment. Using characteristic low-level visual statistics, including but not limited to spatial frequency distribution, luminance distribution, color distribution, etc., an agent can determine when and where to make a first observation.

[0299] The use of low-level features is not limited to map generation. In particular, low-level features can be used in any of the above methods. For example, in hypothesis verification, according to method 300 described above, when determining how to move an agent from a current observation to a predicted observation, low-level features or cues can be used in the movement function. The movement function depends on the predicted distance of movement required to move the agent to the predicted distance of movement, but may also depend on low-level features such as high-contrast lines and edges. In physical space, an agent may be configured to follow low-level features such as roads even if the road does not directly coincide with the direction of the predicted displacement. A road may be identified, for example, by being displayed as a single color that extends upward from the lower part of the field of view where the agent is located toward the upper part of the field of view that mainly includes more distant locations. Alternatively, a road may be identified as an area that has been extended as described above but has no strong horizontal edges and includes strong vertical edges. For example, if a road is within a threshold angle from the direction of the predicted displacement, the agent may be configured to follow the road rather than the predicted displacement. If the predicted displacement is inaccurate for some reason, this may improve the system. The agent makes periodic transition observations during movement to confirm that the agent's movement is appropriately along the predicted observations. If the predicted observations are not reached, or if, for example, from the observation action function, the agent appears to be moving further away, the agent may leave the road. The agent can also move to the end point of the low-level feature (e.g., the end of the road or the next intersection of the road). Even if the agent does not find the predicted observations at this point, before rejecting the hypothesis and returning to the previous observations, observations can be made and saved, and maps can be generated or the map can be further embedded with new knowledge about the space.

[0300] Tracking or otherwise using low-level features simplifies the process of hypothesis verification and space navigation, facilitates backtracking from a point in space to a previous observation point, and is advantageous because it only requires tracking the low-level features in the reverse direction. Further, since the agent can track features without distant descriptive information, its robustness to environmental changes that have a significant visual impact, such as thick fog, is improved.

[0301] In the above description, various concepts are introduced. One such concept is to use an action function to compare saved observations with current or new observations, effectively determine the similarity between the saved observations and the new observations, and provide a vector from the new observations to the saved observations. Thereby, the agent can use the vector when determining how to move to reach the location corresponding to the saved observations. As briefly described above, the action vector can be used to generate a vector field map of the space with respect to the saved observations. Here, with reference to FIGS. 11 to 15, and FIGS. 6 and 9 above, the process of generating the action vector will be described in more detail.

[0302] As described above, the observation action function compares a sub-region of the current observation with a sub-region of the saved observation. Each sub-region of the current observation and the saved observation is a part of the entire current observation and the entire saved observation, respectively.

[0303] The use of sub-regions for this purpose stems from the concept that when an agent moves away from or towards the location where a saved observation was acquired, not all parts of the field of view change evenly. This differs from conventional triangulation in that it takes into account the distortion of the entire visual scene rather than the movement of identified objects or signal sources. As a result, it is robust even when a part of the visual scene is occluded and more resistant to changes in the environment. Furthermore, since it does not require metric calculations to function, it is more robust to measurement noise. Using this property, not only can an agent understand how far it is from the location corresponding to a saved observation, but it can also calculate an additional vector to move to that location. This additional vector calculated by the observation action function uses the same principle as the observation similarity function, but the dot product is obtained for N spatial sub-regions of the current observation and N corresponding sub-regions of the saved observation at different points.

[0304] To obtain the sub-regions, the processes described above with respect to FIGS. 6, 9, and 10 are executed. This process will be described in detail here with reference to FIG. 11.

[0305] FIG. 11 shows a schematic diagram of the results of several processes performed on the original observation data to obtain an observation vector or matrix o from which a plurality of sub-regions are obtained. In FIG. 11, in the first step, the original observation data is acquired, and the original image 1101 is acquired. The original image 1101 is a 360-degree view of the agent's environment acquired from the first location, and the first location is the position of the agent. In this case, the original image 1101 forms the first or current observation data, also called the first sensor data. Similarly, the original image 1101 may be historical data or stored data. The process of forming sub-regions is the same for both saved observation data and current observation data. However, for saved observation data, also called second sensor data, it is not necessary to save the original image, but only a single observation vector or matrix corresponding to the original image (an image in the form of a slice or reduced size) needs to be saved. This will be explained in more detail below. To obtain the original observation data such as the original image 1101, a sensor data acquisition process by the agent is executed. It will be understood that any sensor or virtual sensor capable of recording spatial information can be used. In the case of the original image 1101, a camera may be used.

[0306] In the second step, the original image 1101 is divided into individual feature channels, and for example, four feature channels are provided. These feature channels include, for example, red, green, luminance vertical edge, luminance horizontal edge, etc. The resolution may be the same as that of the original image 1101, such as 256x64 pixels, for example. The red feature channel 1102 is shown in FIG. 11. Other types of channels are also possible. The edges can be obtained by convolution with a 1x2 kernel.

[0307] In the third step, for each feature channel of the original image 1101 such as the red channel 1102, the image is processed with a box filter of a 17×17 square (or a square of any other size such as 19×19), and the information of each channel is blurred to generate a filtered image 1103 of each channel. It should be noted that other types of image processing are also possible as is well understood.

[0308] In the fourth step, a reduction process is performed on the filtered image 1103. This reduction process includes performing a slice operation to reduce the dimensions of the filtered image 1103. In particular, the filtered image 1103 is sliced, and a part of the rows and a part of the columns of the filtered image 1103 are deleted. The first reduced format image 1104 is shown in FIG. 11. This first reduced format image 1104 is the result of performing a slice of a predetermined number of rows from the filtered image 1103. In FIG. 11, 4 rows are "sliced" from the filtered image 1103 (the remaining rows are discarded), and a first reduced format image 1104 of 256x4 is obtained. The selection of the rows to be cut and used is arbitrary, but it is advantageous to use non-adjacent rows such as rows 8, 24, 40, 56 or rows 6, 22, 38, 54. FIG. 11 also shows a second reduced format image 1105. This second reduced format image 1105 is the result of slicing a predetermined number of columns from the filtered image 1103 (in addition to the slice operation performed to obtain the first reduced format image 1104). In FIG. 11, 11, 16 columns are "sliced" from the filtered image 1103 (the remaining columns are discarded), and a reduced format image 1105 of 16x4 seconds is obtained. The selection of the columns to be cut out and used is arbitrary, but it is advantageous to use non-adjacent columns, for example, columns 1, 17, 33, 49, 65, 81, 97, 113, 129, 145, 161, 177, 193, 209, 225, 241. To evenly disperse the data from around the agent, it is even more advantageous to select evenly spaced columns.

[0309] When both slice operations are performed (to obtain the second reduced format image 1105 as shown in FIG. 11), an observation vector o is obtained. The observation vector or matrix roughly represents the agent's 360-degree environment according to the original sensor data 1101. In the example of FIG. 11, this observation vector is formed from the second reduced format image 1105. This image itself is a 16x4 matrix representing a reduced and simplified version of the original data 1101. Other dimensions are possible and understood. Selecting rows and columns to hold evenly distributed from the image 1103 filtered by the slice operation ensures that the observation vector or matrix appropriately represents the entire field of view associated with the original image 1101. In other words, it is advantageous if the columns and rows of the filtered image 1103 that are sliced to form the observation vector or matrix are evenly distributed throughout the filtered image 1103.

[0310] Additional processing can be performed to obtain an observation vector in a form suitable for performing the comparison. In particular, the 16x4 second reduced format image 1105 is periodically convolved with a kernel that is either centered or off-centered (or vice versa). The channels of the 16x4 second reduced format image 1105 can be split into two data subsets. One is composed of red and green, or in other cases, a "color" vector, and the other is composed of luma_h and luma_v, or in other cases, an "edge" vector. These data subsets are each normalized by vector normalization (division by their respective Euclidean norms), so that the elements of each color vector and edge vector represent values on a set of orthogonal values, and the result is a unit vector. These normalized vectors are each scaled by the square root of 2 and concatenated into a single unit vector. This provides an observation vector o associated with the second reduced format image 1105 and the original image 1101. The observation vector in this example is concatenated from 16 columns, 4 rows, and 4 channels, resulting in a 256-element vector. These processes are performed on the 16x4 second reduced format image 1105 to generate a one-dimensional vector suitable for comparison calculations, but it should be understood that the following reference to the second reduced format image 1105 (16x4 matrix) is synonymous with the observation vector. The observation vector is substantially the same information as the second reduced format image 1105, but in a different form and in a normalized form. Therefore, the second reduced format image 1105 may be referred to as the observation vector.

[0311] The above first to fourth steps are applicable to either the first (current) sensor data or the second (stored) sensor data, meaning that the process of forming an observation vector or matrix for each of them is exactly the same.

[0312] In some embodiments where the observation action function is executed, an additional fifth step is performed to obtain a sub-region from the observation vector.

[0313] In this fifth aspect, the observation vector or matrix 1105 is subsampled and sub-regions 1106a through 1106n of the sensor data are obtained. FIG. 11 shows sub-regions 1106a - 1106n of the red feature channel. Sub-regions 1106a - 1106n are substantially portions of the field of view roughly represented by a 16x4 observation vector or matrix. In this fifth step, a mask is applied to the observation vector or matrix 1105 to obtain each sub-region. The mask is overlaid on the observation vector or matrix 1105, and the output of this operation is a reproduction of the cells or pixels of the observation vector or matrix 1105 overlaid by the mask. Since the mask is smaller than the dimensions of the observation vector or matrix 1105, only a portion of the pixels of the observation vector or matrix 1105 are reproduced by the mask. This reproduced portion forms the sub-region. In the example, the mask is 8x4 pixels and is overlaid on a 16x4 observation vector or matrix. This means that 8 columns of pixels are discarded by the mask, occupying half of the field of view of the original image 1101. To obtain multiple sub-regions 1106a - 1106n, the mask is repeatedly shuffled across different columns of the observation vector or matrix 1105, or alternatively, the observation vector or matrix 1105 is shuffled column by column each time with the entire mask. These two operations are substantially the same. Since the observation vector or matrix 1105 is a rough representation of the full view (e.g., 360 - degree view) corresponding to the original image 1101, it is necessary to understand that the last column of the observation vector or matrix 1105 is substantially next to the first column, and the mask can be cyclically shuffled past the last column and back to the first column.

[0314] The size of the mask and the number of columns by which the mask moves with respect to the observation vector or matrix 1105 are changeable, and changing these variables can result in various effects. In the above example, the width of the mask is 8 columns, which is half the width of the observation vector or matrix 1105. As a result of this arrangement, the sub-regions 1106a - 1106n generated by the mask substantially correspond to half of the field of view of the observation vector or matrix 1105. Increasing the size of the mask increases the portion of the field of view held in the sub-regions 1106a - 1106n, and decreasing the size of the mask decreases the portion of the field of view held in the sub-regions 1106a - 1106n. Since the number of pixels in each sub-region serving as a comparison basis is large, comparisons between larger sub-regions are more reliable. However, reducing the sub-region has the advantage that relatively limited details of the scene acquired in the original image 1101 can be compared. In other words, smaller sub-regions may be less affected by more extensive changes in the environment. Therefore, smaller sub-regions may be used to improve the effective range within which the comparison between the stored data and the current data can be made. In particular, when the agent moves away from an object or other attribute, the object is displayed smaller and will ultimately be found in a smaller region of the pixels within the original image 1101 compared to when the object is close to the agent. Similarly, when the agent moves farther away, the area around the object may appear to change relatively dramatically in the original image 1101. Using smaller sub-regions (e.g., those adjusted to the expected size of such an object at a specific distance) allows for comparison between sub-regions and finding the object even when the adjacent pixels of the observation vector or matrix 1105 vary significantly depending on the distance from the object. This example shows a balance between selecting a large sub-region or a small sub-region according to the size of the observation vector or matrix 1105. Both of these options have advantages. Even when larger sub-regions are used, comparison methods such as the observation action function remain robust with respect to changes in the distance from the target / specific object.

[0315] As described above, the movement of the mask in each iteration to obtain a sub-region for the observation vector or matrix 1105 can also be changed. This is equivalent to determining the number of sub-regions to be obtained. In one example, the mask may move one column at a time for each iteration. That is, each obtained sub-region is reordered by one column compared to the previous sub-region. This is the finest difference possible between sub-regions. In the above example of the 16x4 observation matrix, 16 sub-regions can be obtained in this way (the maximum number of sub-regions possible with a 16x4 matrix). However, it is also possible to reorder the mask by 8 columns at a time to obtain 2 sub-regions, or reorder by 2 columns at a time in each iteration to obtain 8 sub-regions. The number of sub-regions to be obtained is arbitrary, but as will be described later, it is advantageous to obtain four or more in order to enhance the reliability of the action vector formed from the comparison of sub-regions. With this variable as well, a balance is achieved between reliability (the number of sub-regions to be compared increases) and computational load (the computational amount increases as the number of sub-regions to be compared increases). Therefore, the choice of the number of sub-regions depends on the needs and requirements of a specific use case.

[0316] When a normalized observation vector is masked to obtain a sub-region (substantially a sub-vector), the normalization of the vector is lost. Therefore, in addition to the above process, the sub-regions (especially the sub-vectors corresponding to these sub-regions) are renormalized.

[0317] Before explaining in more detail the comparison of the sub-region of the current (first) sensor data and the sub-region of the stored (second) sensor data with respect to the action function, first, the recording and storage of the observation vector or matrix 1105 will be described.

[0318] It should be understood that the sub-regions do not need to be stored and held by the agent or the memory associated with the agent, and are only required, for example, when the comparison is calculated using the observation action function. Therefore, the sub-regions can be generated from the observation vector or matrix 1105 as needed.

[0319] Regarding the recording and storage of observation vectors or matrices, several options are possible. As described in the fourth step above, the observation vector or matrix 1105 is a rough representation of the original image 1101 and provides information from the entire field of view represented in the original image 1101 in a reduced size. As part of the process of forming the observation vector or matrix 1105, a slice is performed that effectively discards the rows and columns of the filtered image 1103 to reduce the size. However, for the purpose of performing a comparison between the current observation vector and the saved observation vector, it is advantageous that the information deleted by these slice operations is not lost. Various options for saving this data are possible for the saved (second) sensor data and the current (first) sensor data.

[0320] In the first option, when saving observations from the original image 1101 using the above process, multiple observation vectors or matrices 1105 are saved for each original image 1101. The multiple observation vectors or matrices 1105 are obtained by cutting out different columns in the fourth step above. In the above example, the original image 1101 is a 256x64 pixel image. This image is filtered and sliced to provide a 16x4 observation vector. Assuming that the columns that survived the slice operation are evenly distributed in the original image 1101, there will be 16 16x4 observation vectors for each channel from which the original image 1101 is obtained (256 / 16 = 16). These different observation vectors are obtained by rearranging the columns of the slice operation one column at a time. 17 番目の In the permutation, the observation vector is 1 番目のIt becomes the same as the permutation. For example, when the first observation vector is generated from columns 1, 17, 33, 49, 65, 81, 97, 113, 129, 145, 161, 177, 193, 209, 225, and 241 of the original image 1101, the second observation vector is generated from columns 2, 18, 34, 50, 66, 82, 98, 114, 130, 146, 162, 178, 194, 210, 226, and 242. This process can be repeated until the first column of the observation vector corresponds to column 17 of the original image 1101. In this case, this observation vector is just a rotation of the first observation vector. Therefore, in this example, to ensure that all columns of the original image 1101 are retained, a set of 16 observation vectors is required.

[0321] Furthermore, each of the 16 observation vectors (obtained by slicing from different columns) may be rearranged up to 16 times each to move the data within the observation vector to different columns. This is effectively a rotation of the observation vector (for example, the data originally in column 1 is rearranged to column 2, and the data in column 16 is rearranged to column 1, etc.). Therefore, there are 16 possible rotations for each observation vector, and 16 different observation vectors can be formed by changing the slicing operation. Therefore, in total, 256 observation vectors can be considered from the 256x64 original image 1101.

[0322] In the above first option, a set of 256 observation vectors or matrices is stored for the saved sensor data. When comparing with the current observation (from the current sensor data), a corresponding set of 256 current observation vectors can be compared with the set of saved observation vectors. Using the same process of searching for the set of observation vectors of the current original image and the saved original image, for example, a set of 256 16x4 observation vectors that are saved will be compared with a set of 256 current 16x4 observation vectors. This improves the comparison accuracy, but since it is necessary to calculate a large number of vector inner products, the calculation efficiency deteriorates. Saving and comparing the complete set of saved observation vectors also means that when performing the comparison with the current observation vectors, an additional step of selecting the best comparison result from the complete set of saved observation vectors is required. Therefore, comparing with the complete set of saved observation vectors introduces an additional layer or step in the comparison process, further increasing the computational load. However, obtaining an optimal comparison for one of the complete sets of saved observation vectors may potentially result in more accurate and useful results.

[0323] In the second option, when saving the observation vector, only one observation vector or matrix needs to be saved. This saved observation vector may be called the zero-th observation vector. In the above example, for instance, it could be the observation vector obtained from columns 1, 17, 33, 49, 65, 81, 97, 113, 129, 145, 161, 177, 193, 209, 225, and 241 of the original image. The image itself and other possible observation vectors are not saved. This means that the details of the original image 1101 are substantially lost, and the only data available to the agent regarding the original image 1101 is the rough representation provided by the zero-th 16x4x4 observation vector. The reason this is allowed is that when compared to the current observation (which may not be saved in persistent memory and instead may be recorded on-the-fly or in volatile memory etc.), the saved zero-th observation vector is compared with each possible permutation of the current observation vector (for example, a set of 256). Thus, the saved observation vector is compared with a set of 256 current 16x4x4 observation vectors. That is, the saved zero-th observation vector is saved against all the data obtained in the current observation (all the data provided by the 256x64 current original image). This way, if any of the current observation data is similar to the zero-th saved observation vector, such similarity should be detected by comparing the zero-th saved observation vector with the complete set of current observation vectors.

[0324] The second option is computationally efficient compared to the first option above. This is because only one observation vector needs to be saved for a given original image for comparison with new current data. Further, in the second option, there is no need to perform an additional procedure to determine which saved observation vector to use (from the complete set) when performing the comparison with the current observation vector. This is particularly advantageous when an agent or a computer associated with the agent needs to save hundreds / thousands of observation vectors related to various locations, saved objects, or other attributes.

[0325] According to the second option, when comparing observations, the saved observation vector is compared with a set of current observation vectors that are each possible permutation of the current original image. When using the observation action function, as will be described in detail below, a sub-region of the saved observation vector (i.e., the saved sub-vector) is compared with a sub-region of the set of current observation vectors (i.e., the current sub-vector). As shown in FIG. 11, in an example where the saved observation vector and the current observation vector are formed from a 16x4 second reduced format image 1105 (for each of the four channels), there are 256 16x4 current observation vectors, and each current observation vector may be used to form a sub-region (or current sub-vector). This is done by processing each current observation vector through a mask used to generate the sub-region.

[0326] This means that there are a total of 256 possible sub-regions for the current observation vector to be compared against 16 possible sub-regions of the saved observation vectors. Figure 12 shows this possibility, where the spatial graph 1200 shows the number and arrangement of sub-regions 1201 around the agent 1202. It should be understood that Figure 12 is not to scale and is for illustrative purposes only. Each arm 1203 of the graph 1200 represents a sub-region obtained from the same mask position with respect to the current observation vector. Each sub-region 1201 moving along each arm 1203 is obtained from a different one of the set of current observations (split into different slices). Each concentric set 1204 of sub-regions on the graph 1200 is from the same current observation vector, but corresponds to different directions in space because the position of the mask used to generate the sub-regions is different (or permuted) for each sub-region within the concentric set 1204, or because the current observation vector itself is rotated before extraction of the sub-regions. Although shown as rectangles in Figure 12, it should be understood that each sub-region contains a part of the 360-degree view around the agent, and that part depends on the size of the mask as explained above. In one example, each sub-region corresponds to half of the field of view corresponding to the current original image to which this data relates. The sub-regions overlap each other.

[0327] In the observation action function, the current sub-region (from the current first sensor data) is compared against a set of sub-regions derived from the saved observation vectors. From the set of current sub-regions, an optimal match is found for each saved sub-region, and an offset angle and magnitude are determined for each optimal match. This process will be explained in more detail here with respect to the observation action function.

[0328] As described above, the saved observation vector is divided into N sub-regions (also called "saved" sub-regions), and for each of the N saved sub-regions, a comparison is performed with the complete set of sub-regions of the current observation vector (also called "current" sub-regions). All possible rotation slice permutations (e.g., 256 permutations for a 16x4 observation vector) are compared, and the permutation most similar to the saved sub-region (also called the most matching current sub-region) is identified. The most similar specific permutation has an index such as the 120th rotation, which is used to identify the offset with respect to the saved sub-region. An offset is calculated for each saved sub-region, and optionally the magnitude of similarity is calculated. This magnitude indicates the similarity between the saved sub-region and the most matching current sub-region. The magnitude of similarity is determined by performing the comparison itself. The comparison is a combination of the observation similarity function and the rotation function described above. In particular, the scalar product of each of the saved sub-region and the set of current sub-regions is calculated, and the most matching current sub-region is determined. By normalizing the sub-regions, the scalar product becomes maximum (1) when the vectors are aligned (i.e., when the relative difference between pixels is the same), zero (0) when the vectors are orthogonal, and minimum (-1) when the vectors are exactly opposite. Therefore, this comparison returns a value in the range [-1,1] and the index of the position of the most matching current sub-region. This associates the offset and magnitude with the current sub-region and / or the saved sub-region.

[0329] Figure 13A shows a schematic diagram of the comparison of sub-regions. Figure 13A shows agent 1301. Agent 1301 is located at the center of the "current observation" cylindrical projection 1302, which consists of a green feature channel projection 1302a and a red feature channel projection 1302b. The current cylindrical projection 1302 is an exemplary spatial projection of the current observation vector, indicating how the current observation vector describes the environment around the first location of agent 1301. In particular, the current observation vector describes how the actual scene 1303 (i.e., the world) around agent 1301 at the first location is represented by sensor data. Also shown in Figure 13A is a similar depiction of the "stored observation" cylindrical projection 1302' around the agent, which includes a green feature channel projection 1302a' and a red feature channel projection 1302b'. The cylindrical projection 1302' is an exemplary spatial projection of the stored observation vector, indicating how the stored observation vector describes the environment around the second location of agent 1301. In Figure 13A, the projections of the stored observation vector and the current observation vector are divided into sub-regions. In particular, Figure 13A shows a first (front) sub-region 1304a for each of the "current observation" cylindrical projection 1302 and the "stored observation" cylindrical projection 1302, and also shows a second (rear) sub-region 1304b for each of the "current observation" cylindrical projection 1302 and the "stored observation" cylindrical projection 1302. In other words, Figure 13A shows the front stored and current sub-regions 1304a and the rear stored and current sub-regions 1304b. This is exemplary, and it is understood that in reality, the current sub-regions 1304a and 1304b are samples of the set of all possible permutations of the current sub-regions used for comparison.

[0330] The set of the saved front and the current sub-region 1304a are compared with each other using an observation similarity function and a rotation function, and a scalar product and an offset are obtained. The set of the saved back and the current sub-region 1304a are also compared with each other using an observation similarity function and a rotation function, and a scalar product and an offset are obtained. These operations respectively provide offset angles 1305a and 1305b of the sub-region. These operations also provide a magnitude MF indicating the similarity between the saved front sub-region 1304a and the current sub-region 1304a, and a magnitude MB indicating the similarity between the saved back sub-region 1304b and the current sub-region 1304b.

[0331] FIG. 13B shows the same arrangement of the agent 1301. The only difference in FIG. 13B is that the saved sub-regions and the current sub-regions to be compared are the current sub-regions on the right and the saved sub-region 1304c and the current sub-regions on the left and the saved sub-region 1304d, instead of the front sub-region and the back sub-region. The comparison process is otherwise the same, and the complete sets of the saved right sub-region 1304c and the current sub-region 1304c are compared with each other using an observation similarity function and a rotation function, and a scalar product and an offset are obtained. The complete sets of the saved left sub-region and the current sub-region 1304d are also compared with each other using an observation similarity function and a rotation function, and a scalar product and an offset are obtained. These operations respectively provide offset angles 1305c and 1305d of the sub-region. These operations also provide a magnitude MR indicating the similarity between the saved right sub-region 1304c and the current optimally matching sub-region, and a magnitude ML indicating the similarity between the saved left sub-region 1304d and the most matching current sub-region.

[0332] Figures 13A and 13B show the saved sub-regions and current sub-regions of the front, back, right side, and left side, respectively, and the offset angles determined from their comparison. It should be understood that any part of the entire saved observation vector, and thus the number of saved sub-regions corresponding to the directions, can be arbitrary. Figures 13A and 13B show sub-regions that are substantially half the size of the observation vector, and thus contain sensor data representing half of the field of view associated with the observation vector. Thus, in the example of the original 16x4x4 observation vector above, these sub-regions correspond to 8x4x4 fields of view centered on different directions (front, back, right, left). The sub-regions overlap each other. In alternative embodiments, the sub-regions can represent smaller or larger portions of the observation vector.

[0333] The sub-region offset angles 1304a, 1304b, 1304c, and 1304d are measured with respect to the axis corresponding to the currently observed vector that most closely matches overall. It is understood that the determination of the currently observed vector that most closely matches overall can be made by a similar comparison process performed before or after the comparison of the sub-regions described above.

[0334] The currently observed vector with the overall best matching is determined in the same way as the comparison between the saved sub-regions and the set of current sub-regions. The only difference is that instead of performing a comparison of the sub-regions, the entire saved observation vector is compared against a set of permutations of the current observation vector (including rotational permutations and slice permutations). The complete saved observation vector (e.g., a 16x4x4 vector) is compared against each permutation of the current observation vector using an observation similarity function and an observation rotation function, and the currently observed vector that most closely matches is determined. These functions provide an index (e.g., 120 番目の rotation permutation) corresponding to the offset angle of the most closely matching observed vector, and a related magnitude indicating the similarity between the saved observation vector and the most closely matching observed vector.

[0335] FIG. 14 shows a schematic diagram comparing each of a complete stored observation vector and a set of current observation vectors for the purpose of determining the most matching current observation vector. FIG. 14 shows an agent 1401 at a first location. The agent 1401 is disposed at the center of a cylindrical projection 1402 including a green feature channel projection 1402a and a red feature channel projection 1402b. The cylindrical projection 1402 is an exemplary spatial projection of the current observation vector and shows how the current observation vector describes the environment around the first location of the agent 1401. In particular, the current observation vector represents how the actual scene 1403 (i.e., the world) around the agent 1401 at the first location is displayed. Although one current observation vector projection is shown, it should be understood that the comparison is actually performed on the entire set of current observation vectors. FIG. 14 also shows a similar depiction of an agent 1401' at a second location, along with a cylindrical projection 1402' including a green feature channel projection 1402a' and a red feature channel projection 1402b'. The cylindrical projection 1402' is an exemplary spatial projection of the stored observation vector and shows how the stored observation vector describes the environment around the second location of the agent 1401'. In this example, it is shown that both the current observation vector and the stored observation vector are composed of only two channels, but it should be understood that there may be more channels, such as four channels as described above.

[0336] The stored observation vectors shown by the stored cylindrical projection 1402’ are compared with each possible rotation and slice permutation of the current observation vector, and an optimal matching is searched for using the observation similarity function and the observation rotation function described above. The observation similarity function returns a magnitude indicating the similarity between a particular current observation vector and a stored observation vector. The magnitude indicating the highest similarity between a particular current observation vector and a stored observation vector (among the set of current observation vectors) is recorded, and the index of the particular current observation vector associated therewith is also recorded. This index provides the direction 1405 of the optimal matching. Thereby, for the current observation vector that most matches the stored observation vector, the most matching observation vector offset angle 1404 can be determined. In other words, the optimal matching of the stored observation vector in the set of current observation vectors is determined by comparing the stored observation vector with each current observation vector using the scalar product. The optimal matching index is recorded, and from this, the optimal matching observation offset angle 1404 is inferred.

[0337] This optimal matching observation vector offset angle 1404 indicates the rotation that the agent 1401 can perform at the first location to align itself with the optimal matching to the stored observation vector. In the case of FIG. 14, assuming that the stored observation vector is recorded at coordinates (2, 0) and the agent 1401 has moved forward from the second (original) position when it was at 1401’, the current observation vector represented by the projection 1402 rotates to the left, and as a result, the current observation vector at the optimal matching observation vector offset angle 1404 (measured as a positive rotation counterclockwise from the forward orientation) most matches the stored observation vector.

[0338] It should be understood that since agent 1401 can record observations with a wide field of view (e.g., 360 degrees), it is not essential to perform a rotation at this point. Instead, a simple readjustment with respect to the axis (reference frame) of the agent is performed by the optimal matching observation vector offset angle 1404, and for the purpose of comparing sub-regions, the optimal matching direction 1405 is aligned with the axis of the agent. This can be performed before, during, or after the sub-region comparison operation. Since the sub-regions are generated based on a specific direction (not representing the entire field of view), aligning the axis of the agent is important to ensure that the sub-regions of the saved observation vectors (e.g., forward direction, backward direction) are properly oriented (by adjusting the axis with the offset angle 1404 of the most matching observation vector), such that when the action vector is determined, it is based on the offset of the sub-regions with respect to the adjusted axis.

[0339] Figures 15A - 15D show the offset angles of each sub - region with respect to the axis of the agent 1301. In each of Figures 15A - 15B, the axis 1310 with respect to the agent 1301 is shown. The axis 1310 represents the re - adjusted axis, and the axis 1310 is re - adjusted based on the offset angle 1404 of the most - matching observation vector obtained from the comparison between the saved observation vector and the current observation vector. Figure 15A also shows the sub - region offset angles 1304a and 1304b related to the front - face sub - region comparison and the back - face sub - region comparison respectively. Similarly, Figure 15B shows the sub - region offset angles 1304c and 1304d related to the comparison of the right - hand side and left - hand side sub - regions respectively. Once the offsets of each sub - region are determined in this way, the direction of the action vector is determined. To determine the direction of the action vector, a result vector is calculated. This process is shown in Figures 15A and 15B. In Figure 15A, the front / back composite vector 1306a is calculated using the pair on the opposite sides of the front - face sub - region offset 1304a and the back - face sub - region offset 1304b. In Figure 15B, the right / left composite vector 1306b is calculated using the pair on the opposite sides of the right - hand sub - region offset 1304c and the left - hand sub - region offset 1304d.

[0340] Next, the front - back composite vector 1306a and the right - left composite vector 1306b are aggregated to form the final action vector 1308a as shown in Figure 15C.

[0341] Figures 15A and 15B are intended to show a process of defining front / back and right / left result vectors using opposite pairs of sub-region comparisons. However, in reality, it is not essential to form the result vectors. Instead, it should be understood that the action vector 1308a may be calculated directly from the aggregation of all sub-region offset angles 1304a to 1304d. Furthermore, additional sub-regions centered on different directions with respect to the alignment axis of the agent 1301 can also be used. For example, Figure 15D shows a figure similar to Figure 15C, but with respect to the sub-regions, the comparison is made at an angle of 45 degrees with respect to the front-back and left-right directions. Aggregating the sub-region offset angles of these different sub-regions provides the action vector 1308b. It should be understood that these processes provide the direction components of the action vectors 1308a and 1308b.

[0342] The direction of the action vector 1308a determined from the front-back and left-right sub-region comparisons may be aggregated with the direction of the action vector 1308b determined from additional sub-region comparisons. In the example of a 16x4x4 observation vector, there may be up to 16 sub-regions contributing to the determination of the action vector in this way. The direction of the action vector 1308a and the direction of the action vector 1308b are aggregated in any suitable way to provide the final action vector. In one example, the vector average is calculated to determine the final action vector. Alternatively, the final action vector may be selected as either the action vector 1308a or 1308b depending on one or more parameters.

[0343] The magnitude of the final action vector is determined separately from the magnitude calculated in the process of executing the observation similarity function for the purpose of identifying the most matching observation vector and the most matching current sub-region. The magnitude of the final action vector is based on the size and direction of the offset angles 1304a to 1304d. As explained above, the offset angles 1304a to 1304d are determined based on the axis 1310 that coincides with the direction 1405 of the optimal matching. The magnitude of the action vector is determined by components. In particular, the first component is the front-back composite vector 1306a. This composite vector 1306a is a component of the action vector in a direction orthogonal to the front-back direction of the axis 1310, as shown in FIG. 15A. The magnitude of this composite vector 1306a is determined based on the magnitudes of the offset angles 1304a and 3014b. The offset angles 1304a and 1304b may be measured in individual steps based on a rotation index (for example, index 4 of 256). The magnitude can be determined by averaging the sizes of the offset angles 1304a and 1304b (obtaining the average offset angle in discrete steps of 1 / 256). This magnitude is multiplied by a constant k. The constant k may be modified based on specific scenario requirements and may be a positive number such as 2, 4, 6, 8, 10, or any other number.

[0344] A further aspect of the magnitude determination may include using an additional constant to define an upper limit on the use of specific offset angle pairs. In some cases, if the offset angle is large, it may indicate an incorrect match or a match between unreliable sub-regions. To prevent such a match from being used in the determination, the constant j can be set such that the magnitude is set to 0 when the average offset angle between opposing pairs is greater than j. The constant j can be changed for the action vector and can take any value such as 2, 5, 10, 20, 50, etc.

[0345] The magnitude of the composite vector 1306b is calculated in a similar manner with respect to the offset angles 1304c and 1304d. This provides additional orthogonal components as shown in the Figure 15B action vector. The resulting vectors 1306a and 1306b are added together to determine the action vector.

[0346] When the offset angles in both the front - rear direction and the left - right direction are both zero (that is, when the center of the sub - region that best matches the current observation vector is the same as the sub - region of the saved observation), the magnitude of the related orthogonal component becomes zero. The magnitude follows the relationship that the larger the offset from the saved sub - region, the larger the magnitude. Such magnitudes are determined for each pair of opposing sub - region offsets, such as a pair of offset angles including the forward offset angle 1304a and the rearward offset angle 1304b, or a pair including the right - hand offset angle 1304c and the left - hand offset angle 1304d. Next, when forming the action vectors 1308a and 1308b and the final action vector, the magnitudes are appropriately combined as part of the vector averaging determination. Thus, the magnitude of each pair of opposing sub - region offset pairs provides a weighting to the final action vector based on the size of the offset angle compared to the saved sub - region. This essentially means that the action vectors are more weighted from the result vectors 1306a and 1306b generated from larger offsets (offset angles away from the aligned axis). Thereby, a final action vector is provided that points in the direction where there is a high probability of finding the location within the space where the saved observation was made, along with a magnitude indicating the distance or the effort required to reach this location. This magnitude is also inversely proportional to the reliability indicating that the location associated with the saved observation is in the direction indicated by the final action vector. Thus, the larger the value of the final action vector, the more effort is required to reach that location, and it also indicates a lower certainty that the location is in the indicated direction. Conversely, the lower the magnitude, the less effort is required to reach that location, and it indicates a higher certainty that the location is in the indicated direction.

[0347] The magnitude of the action vector is determined based on the offset of the opposing pair of the best-match sub-regions, but the magnitudes of similarity ML, MR, MF, and MB obtained from the comparison between the saved sub-region and the current sub-region may still be used in the process. In particular, based on the magnitude of similarity of the opposing sub-region pairs used to calculate the action vector, weighting can be applied to the magnitude of the action vector. Opposing sub-regions with high similarity are weighted such that the contribution of the magnitude of the action vector from a specific opposing sub-region increases, while sub-regions with low similarity may be weighted such that the contribution of the magnitude of the action vector decreases.

[0348] Furthermore, the magnitudes of similarity ML, MR, MF, and MB are used in threshold processing before determining the action vector, and sub-regions can also be removed or excluded from the determination of the action vector. In particular, the magnitude of similarity of the comparison of each sub-region may be compared with a definable reliability threshold (e.g., 0, 0.1, 0.5, or any other arbitrary numerical value). If the magnitude of similarity does not match or exceed the reliability threshold, the comparison of the corresponding sub-region is not used in the determination of the action vector. As described above, it should be understood that similarities and sub-region directions other than ML, MR, MF, and MB may be used in the determination of the action vector.

[0349] The final action vector can be used alone or in combination with the other functions and methods described above to determine a way to move the agent towards the saved observation corresponding to the saved observation vector. For example, the action vector may be used in a movement function to relocate the agent to the predicted observation, and as shown in Figure 7, it may also be used in addition to displacement data or relocation techniques.

[0350] The observation action function is a tool useful for navigation as it provides the action vector by which an agent can move to reach the destination indicated by the saved observation vector. By saving multiple observations of a particular space, the agent can repeatedly use the action function on the multiple saved observations associated with that space to move through the space. The action function is relatively resistant to small environmental changes as multiple sub-regions are used in determining the action vector and the sub-regions themselves are pre-processed to reduce sensor data information to its most basic components (edges, distinct features, etc.).

[0351] As briefly outlined above, the action vector for the saved observations, determined from comparison with the current observation, can be used to generate a map or a part of a map. The map may include data regarding the action vector and data regarding the magnitude of similarity of the saved observations. The map is generated as a connected series of observations (nodes) and may include the displacements linking the observations. The map is stored as a spatially corresponding array or series of arrays, with each array being composed of a different dataset. Each array can either form its own map or contribute as part of the overall map.

[0352] Figures 16A and 16B respectively show a first and a second vector field 1600 and 1610 formed from action vectors calculated for observations stored at locations 1602 and 1604. Figure 16A represents a theoretical vector field of action vectors for observations stored at the central location 1602. The vectors represented by the arrows respectively point in the direction of the stored observations, and the magnitude is proportional to the distance from the stored observation location 1602. Figure 16B is a realistic vector field of action vectors based on the stored observations at the central location 1604. In this example, many of the action vectors still point in the direction of the stored observations, but some point in other directions and have different magnitudes. The realistic vector field 1610 forms the first part of the map and may be stored in an array, for example.

[0353] As shown in Figures 16A and 16B, each action vector is determined from a comparison of the current observations at locations in space with the stored observations at specific positions 1602 and 1604. To set up the vector field, the current observations are compared with the stored observations (to obtain the action vectors), and the action vectors are added to the vector field at the current location (the location corresponding to the current observations being compared). Although the global location corresponding to the current observations may be unknown, the local locations of each current observation and the corresponding action vectors can potentially be determined based on other current observations and action vectors already existing on the vector field 1610. In particular, as explained above with regard to displacement, odometry such as visual or inertial odometry can be used to determine the relative positions of the continuously made current observations. By mapping the continuously made current observations over a short distance compared to the accuracy / precision of the odometry (travel measurement), the inherent drift of the odometry is not a problem. This means that the map (vector field) can be created using the action vectors, thereby making the distances between the vectors substantially accurate. To increase the accuracy of the map with respect to the locations of the action vectors, it is necessary to shorten the distance that the agent moves between the current observation points.

[0354] Since each action vector corresponds to a specific saved observation, the resulting vector field map is effectively created within the coordinate frame of the specific saved observation. In this method, it is not necessary to know the absolute location (within the world reference frame) of the saved observation or the current observation.

[0355] When sufficient comparison with the saved observations at locations 1602, 1604 is made, the agent can use the vector field to move towards the saved observations. However, as shown by vector field 1610, the action vectors may not be uniform or accurate, and if the vector field is composed of vectors of random lengths and orientations, the agent may have to move randomly around location 1604 to find the action vector that points to that location. Instead, if the vector field is composed of vectors that precisely point directly to the target location 1604 with high precision, the agent can move along the shortest path. Therefore, when using the action vectors as a characteristic for generating the map, it is beneficial to have a uniform and reliable vector field. Since the vector field depends on the comparison of the saved observations with some current observations, the reliability of the vector field is synonymous with the reliability of the saved observations.

[0356] A highly reliable saved observation is considered to be an observation that has an effective action vector field whose vectors always guide the agent along the path towards the target.

[0357] The map can be used to verify the reliability of the observations stored therein. If the map contains data regarding the action vectors (such as a map including the action vector field), the reliability of the saved observations can be verified according to several methods.

[0358] In the first method, the angular deviations between each action vector in the action vector field and the vectors in a field that converges radially (such as that shown in FIG. 16A) are summed and averaged to obtain a scalar measure that is inversely proportional to the reliability of the stored observations. FIG. 17 shows a visual representation 1700 of the vector angular deviations of each action vector. The black hexagons correspond to low deviations from the ideal radial field, and the white hexagons correspond to the maximum possible deviation (π radians / 180 degrees). The hexagons marked with a "star" correspond to locations where the action vector is zero. In the example of FIG. 17, the average deviation is 51.7 degrees. The larger the average deviation, the lower the reliability of the stored observations, and the smaller the average deviation, the higher the reliability of the stored observations.

[0359] In the second method, a calculation is performed to obtain a value by dividing the number of locations in the motion vector field where the action vector cannot be calculated by the number of locations where it can be calculated. As an example, the action vector may not be calculable if the offsets used in the calculation of the action vector are all canceled out during the calculation of the determination of the action vector. The value obtained by dividing the number of locations in the action vector field where the action vector cannot be calculated by the number of locations where the action vector can be calculated is a measure that is inversely proportional to the reliability of the stored observations. In the example of FIG. 16B, 10% of the locations in the vector field 1610 correspond to locations where the action vector could not be calculated. In this example, the measure is equal to 1 / 9 (10% / 90%), indicating that the stored observations may be fairly reliable. A larger numerical value means that the reliability of the stored observations may be lower.

[0360] The map may alternatively or additionally include a similarity magnitude field obtained using an observation similarity function as described above, which includes data regarding the magnitude of similarity to stored observation vectors. This data may be stored in an additional array. FIG. 18 shows an example of a realistic similarity magnitude field 1800. This is a scalar field, and each data point represents a magnitude returned from applying an observation similarity function to find the similarity between a particular current observation vector and a stored observation vector. Similar to the vector field described above, the location corresponding to each current observation is determined using odometry. The value of the similarity magnitude between the current observation and the stored observation may be determined simultaneously with, or during the same process as, the determination of the action vector so that the current observations are the same.

[0361] FIG. 18 shows a scalar field around the location 1602 of the stored observations (the same stored observations as in FIG. 16B). The bright-colored hexagons within this scalar field indicate higher magnitude values returned as a result of executing the observation similarity function for current observation values that are close to, and thus more similar to, the observation values stored at location 1602. The dark regions indicate low magnitude. As can be seen from FIG. 18, generally, the value of the magnitude of the scalar within the scalar field increases in the vicinity of the location 1602 of the stored observations. Since the magnitude similarity scalar field spatially matches the action vector field, both of these data sets or arrays can be combined to create an overall map.

[0362] The similarity magnitude scalar field 1800 may be used to form a parameterized model 1900 of the similarity magnitude scalar field. The parameterized model 1900 is a smooth approximation of the actual field 1800 and is adapted according to one or more characteristics of the actual field 1800. An example of a parameterized model of the similarity magnitude scalar field 1800 is shown in FIG. 19. In particular, FIG. 19 shows the model field 1900 of the similarity magnitude, showing a smooth and consistent change as the field approaches the saved observation location 1604 which is at the center of the field. Similar to FIG. 18, white indicates a high similarity magnitude and dark regions indicate a low similarity magnitude. Using the parameterized field 1900, data from the actual field 1800 can be extended without the need for additional measurements or additional current observations.

[0363] The parameters of the model of the similarity magnitude field may include the maximum similarity, the shape or form of the slope measured with respect to the distance from the location, and whether the field is circular or elliptical. An example of a parameterizable model is a two-dimensional elliptical Gaussian distribution g(x,y) centered and rotated. TIFF2025522884000002.tif16150 TIFF2025522884000003.tif17150 TIFF2025522884000004.tif14150 TIFF2025522884000005.tif14150

[0364] Here, the five parameters are O, the scalar offset, M, the multiplier of the Gaussian hill with the standard deviation, o x 、o y 、and the rotation angle θ.

[0365] Using any standard optimization method, the model g(x,y) can be fitted to the data f(x,y) to find the parameters that minimize the objective function o of the following form. TIFF2025522884000006.tif11150

[0366] In the example shown in FIG. 19, the Nelder-Mead simplex method was used.

[0367] In the case of saved observations, the map may include one or both of the action vector field and the scalar similarity magnitude field.

[0368] From the perspective of the similarity magnitude field, a highly reliable saved observation is considered to have a consistent scalar field and be able to identify when an agent arrived when approaching repeatedly (e.g., by verifying whether the magnitude exceeded a threshold).

[0369] The reliability of the saved observations on the map including the similarity magnitude field can be verified according to several additional methods.

[0370] A third method for verifying the reliability of the map and the observations saved therein includes calculating the modeled maximum similarity magnitude divided by the average similarity magnitude of the parameterized model of the similarity magnitude field. A large calculated value indicates that the saved observations can be clearly recognized (because it is different from the average similarity magnitude of the modeled field). Therefore, the measure provided by this calculation is proportional to the reliability of the saved observations. In the example of FIG. 18, this measure is calculated to be 1.32. Therefore, this third method can also be applied to the actual field 1800 or the parameterized field 1900.

[0371] A fourth method for verifying the reliability of saved observations involves determining the cumulative deviation from the similarity scalar field using a parameterized model of the similarity field. This requires determining the variability of the similarity magnitude and comparing this variability with the parameterized model that best fits it. Surfaces that are more variable and less smooth increase the probability of misidentifying the wrong location as the location of the saved observation. This value may be used as an objective function for optimization to find the best model fit when improving or updating the parameterized model. Similar to the third method above, this fourth method can be performed either with the parameterized model 1900 or the actual field 1800.

[0372] The four methods described above for verifying the reliability of observations saved in a map can each be used individually or in combination. In some examples, one or more of the first or second methods (with respect to the action vector field) are used in combination with one or more of the third and fourth methods (with respect to the similarity magnitude field). By combining the metrics determined from these two fields, it becomes possible to more reliably judge the reliability of the saved observations. The metrics from the methods for determining reliability can be combined in any suitable way, and the contribution from each method to the overall metric may be weighted according to application-specific requirements.

[0373] The reliability metric for the saved observations may be recorded or associated with the saved observations in the map. The map may include one or more of the action vector field and the similarity magnitude field, or data from them. For example, a map of multiple saved observations may include the vector field within a limited range around each saved observation, and optionally, data points or contours near each saved observation corresponding to the similarity magnitude data or its parameterized model.

[0374] The reliability metric can be used to determine whether a particular saved observation needs to be excluded from the map (if the saved observation is not reliable), or whether the saved observation needs to be repeatedly measured / recorded to improve reliability. This can be done using one or more reliability thresholds. The reliability thresholds can be compared to one or more metrics determined from the first through fourth methods of verifying reliability described above. For example, the second method above returned a value of 1 / 9, which can be compared to the reliability threshold. The reliability threshold is adjustable and could be 1 in this example. That is, a saved observation that contains less than or equal to 1 / 2 of the total possible action vectors in the vector field is considered unreliable.

[0375] Unreliable saved observations may be removed from the map or the agent may be instructed to re-capture and save the observations to obtain more reliable observations.

[0376] When the reliability of a saved observation is confirmed from the perspective of the similarity magnitude field and / or vector field associated with that saved observation, this saved observation and the associated fields can be used to improve the determination of the agent's location. Also, it should be understood that the vector field and similarity magnitude field can be used to determine the agent's location even if the reliability has not been pre-checked.

[0377] The action vector field and the similarity magnitude field form two of several modalities that can be used to determine the location of an agent. Specifically, these modalities may include one or more of an action vector or action vector field for a saved observation, a similarity magnitude or similarity magnitude field for a saved observation, and a displacement between data regarding the saved observation and the saved observation itself. On a map, these different data sets can be combined. For example, a map can consist of data regarding a series of observations and displacements as shown in identity 800 of FIG. 8, a vector field do for each saved observation such as vector field 1610 shown in FIG. 16B, and a scalar similarity magnitude field for each saved observation such as field 1800 of FIG. 18 or parameterized field 1900 of FIG. 19. This example is shown in FIG. 20. FIG. 20 shows a map 2000 that includes the displacement data and observation data of FIG. 8 and the similarity magnitude field data of FIG. 19. It should be understood that the action vector field data may also be present on the map near saved observations o1 to o4. Saved observations within the map may sometimes be referred to as "nodes". For each node, the surrounding similarity magnitude field is shown. These similarity magnitude fields overlap each other. It should be understood that the similarity magnitude fields may be extrapolated according to a parameterized model to be shorter or longer from the nodes than shown in FIG. 20. Map 2000 is constructed to be a topographically correct map, meaning that the relationships between the nodes are correct. The exact angles and distances between the nodes do not have to be perfectly accurate. Map 2000 can be constructed using one or more modalities / data sets by generating a probability distribution of each data set evaluated at individual locations on a grid representing the agent's local area. The local area is defined according to the scenario and may, for example, range from 1 square meter, 10 square meters, or 100 square meters. Thus, map 2000 can be thought of as an array or grid, and each cell or grid square is composed of data corresponding to each data set used to form the map.

[0378] In the following description for improving the determination of the agent's location, reference is made to "modes". These are synonymous with methods or processes, but are named as such to avoid confusion with the various methods described above. In the following description, several modalities are also mentioned, such as similarity magnitude values, ratios, action vectors, odometry values, etc. These may each be referred to as a "location scale".

[0379] In the following description, reference is also made to location, which is synonymous with position. Similarly, the locations within each of the following arrays may also be referred to as "field locations".

[0380] Furthermore, it should be further understood that each of the following modes can be applied to multiple saved observations. In this case, the agent's location can be determined from multiple saved observations, and each saved observation is associated with the performance of one or more modes (and one or more of their associated arrays). Thus, each saved observation corresponds to a respective set of arrays, and a set of arrays from multiple saved observations can be combined to determine the agent's location.

[0381] In a first mode of improving the determination of the location of an agent in space according to the generated spatial map, the similarity magnitude for one or more saved observations, and thus a scalar field of similarity magnitudes, is used. In this first method, an initial array is generated from a parameterized model of the similarity magnitude field. The similarity magnitude measured for a particular saved observation can be used to narrow down the distribution of the possible locations of the agent by using the distance from the saved observation and a model of how the similarity magnitude changes as a function of angle, albeit to a lesser extent. For example, consider the case where the similarity magnitude decreases linearly from 1 to 0.5 over a distance of 1 meter from the saved observation location, the similarity magnitude is radially symmetric, and the agent observes a similarity magnitude of, for example, 0.75 with respect to that particular saved observation. In this case, the agent may be determined to be at a radial distance of 0.5 m from the location of the saved observation. Thus, the similarity magnitude with respect to the saved observation can be used to place the agent on a contour (a circle in this example) around the saved observation. Thus, the initial array can be used in a manner similar to an elevation model. Connected cells or entries within the array are considered to be on the same "contour" if they are within a certain range of values with respect to each other. Specific ranges include, for example, values of similarity magnitude such as 0.01, 0.05, 0.1, etc. Alternatively, only cells with exactly the same similarity magnitude value are considered to be part of the same contour line. Each contour line is determined using any suitable algorithm and may be labeled as such. For example, a clustering algorithm may be used. In this way, the first array can represent a series of contours, each corresponding to an equally spaced range of similarity magnitude values, and the contours surround the saved observations. When a new current observation is compared to the saved observations and a new similarity magnitude value is determined, the initial array can be used to narrow down the possible locations of the agent to somewhere on the contour of the initial array that has a range that includes the new similarity magnitude value.

[0382] If the agent is also close enough to other nodes (other saved observations), i.e., if it is within the radial scalar field of other saved observations, the similarity magnitude can be calculated, and the same operation can be performed on those nodes. Based on the determination for two nodes (the first and second saved observations), a new similarity magnitude value can be used to determine possible locations for the first contour of the first of the two nodes and possible locations for the second contour of the second of the two nodes. In this case, the possible locations can be further narrowed down to the region or location where the first and second contours intersect. If the model of the similarity magnitude field is complete, the two contours representing possible locations for each of the two nodes may intersect at one point, and thus accurate localization can be achieved using only the two nodes. However, this is likely not the case. If the model of the similarity magnitude field is inaccurate, the two contours may not intersect at all or may overlap, resulting in two possible locations.

[0383] If the agent is within the similarity magnitude fields of three nodes, since the intersection of the three rings generates one point in space, the agent can effectively use the three contours to more accurately localize itself. In one example, this can be done using a parameterized conical model of the similarity magnitude field for each of the three nodes. However, it should be understood that any complex parameterized model of the actual similarity field can be used as needed. For example, a 2D Gaussian distribution with covariance, or a discrete sampling approach combined with interpolation if the field is more diverse. Additionally, the actual similarity magnitude field can also be used.

[0384] The first mode involves using the similarity magnitude as the input dataset to construct a first array that can be used in Map 2000. To construct the first array, first, a landscape / actual field of the predicted similarity magnitude is created for each node.

[0385] The first array that may contain the actual fields of multiple nodes can be used to obtain further estimates of the agent's location. In particular, at each discretized location within the local area, a parameterized model of the node's similarity magnitude field can be used to predict what the contribution of the similarity magnitude at that location will be. Using this estimate, for example, when forming the similarity magnitude scalar field of each node, it may be extended to locations farther from the directly sampled locations. In this method, at all discrete locations, the absolute value of the difference between the observed similarity magnitude and the estimated value of the similarity magnitude at that location is determined. This difference is recorded, and the process is repeated at each individual location by the agent. Each of the recorded differences is recorded in an array, and a landscape is generated that indicates the most likely location of the agent where the difference is closest to zero. Total: 1 - The difference can be inverted by performing an addition: the most likely location of the agent will be the location where the difference is closest to 1. As explained above, this method can be repeated for multiple nodes (the nodes visible to the agent). Therefore, multiple first arrays corresponding to multiple saved observations may be generated and then combined.

[0386] This localization method depends on the accuracy of the agent model of the similarity magnitude field of nearby nodes. Since this model is unlikely to be exactly accurate, a probabilistic approach is adopted, where each value within the generated array is passed through a sigmoid function, which asymptotes to 0 for low values and stabilizes at 1 for high values (values close to 1). This has the effect of expanding the region containing the most likely location of the agent based on the model.

[0387] This array modified by the sigmoid function is finally converted into a probability distribution by normalizing the area under the landscape to 1 by dividing each element of the array by the sum of all elements.

[0388] A second mode that improves the determination of the location of agents in space according to the generated spatial map also uses a scalar field of similarity magnitude, and thus the similarity magnitude of one or more saved observations. In this second mode, a second array is generated from a parameterized model of the similarity magnitude field, although it may also be generated from the actual field itself.

[0389] Depending on changes in the environment, the similarity magnitude around a particular node may change. For example, a change in overall brightness may cause the maximum value of the observed similarity magnitude of all nodes to decrease. If the observed similarity magnitude is low due to a change in brightness, in the location identification procedure according to the above first mode and the first array, the agent is predicted to be farther from the node than it actually is.

[0390] To partially mitigate the risk of overall changes in similarity, the ratio between the observed similarity magnitudes for pairs of nodes can be used. For a particular observation of the ratio of two similarity magnitudes corresponding to two nodes, the agent may be somewhere on the front between the two nodes. In the case of a pair of nodes with the same radially symmetric similarity magnitude field, this front will be a straight line perpendicular to the straight line connecting the two nodes. However, if the similarity magnitude field of the nodes is more complex, the front may also become more complex.

[0391] To construct the second array, first, a single estimated similarity magnitude ratio landscape is created for each node pair. At each discretized location within a local area of the map, the estimated similarity magnitude of the first node (obtained from the above first mode) is divided by the estimated similarity magnitude of the second node. Next, at all discretized locations, the absolute value of the difference between the ratio of the observed similarity magnitude and the ratio of the estimated similarity magnitude at that location is calculated. This generates a landscape where the most likely location of the agent is the location closest to where this calculation is zero. Similar to the above first mode, when this array is subtracted from 1 element by element, the array is modified so that the most likely location is the location closest to 1.

[0392] Similar to the first mode described above, this localization method also depends on the accuracy of the model of the similarity magnitude field of the nodes near the agent. Since this model is unlikely to be exactly accurate, a probabilistic approach is adopted in the second mode, where each value of the input array is passed through a sigmoid function, which asymptotes to 0 for low values and stabilizes at 1 for high values (values close to 1). This has the effect of spreading the region of the most likely location of the agent when the model is given.

[0393] Finally, this is converted into a probability distribution by dividing each element of the array by the sum of all elements to normalize the area under the landscape to 1.

[0394] This second mode is similar to the first mode, but since the ratio between nodes should remain almost constant regardless of changes in global luminance or other global environmental changes, the ratio of similarity magnitude values rather than the values themselves is used.

[0395] The third mode that improves the determination of the location of the agent in the space according to the generated spatial map uses an action vector field. Observing the action vector field shows that, with high probability, it is directed inward towards the nodes. In other words, the angle between the radius emitted from the node and the action vector is usually within the interval [-90, 90]. This fact, while not imposing strict constraints on the location by itself, can be used to substantially halve the number of possible locations identified by the similarity magnitude field. The similarity magnitude field only constrains the location to a ring surrounding the nodes. By using the action vector and the similarity magnitude field in combination, the possible locations of the agent can be reduced. The process by this third mode is shown in Figure 21.

[0396] Figure 21 shows a first visual representation 2100 of an action vector field overlaid with rings corresponding to the contours of the observed similarity magnitudes. The similarity magnitude contour lines are obtained by performing the first mode described above on a single node. In the absence of an action vector, the agent could be located anywhere on the contour ring. A second visual representation 2110 shows the result of combining the observed similarity magnitude and the observed action vector. The observed action vector indicates direction A. Combining direction A with the rings of the observed similarity magnitude yields a location probability distribution indicating that, as shown in the figure, the agent's location is most likely to be in the lower right of the contour ring due to the direction A of the action vector. Further, since the action vector typically points to a node (in this case the center of the ring), as shown in the figure, the likelihood of the location being in the upper left of the ring is low.

[0397] By combining the first, second, and third modes described above, the most likely location of the agent can be determined. This location is represented by the sum of the first array, second array, and third array described above. That is, the difference between the observed and predicted similarity scores for each node within the agent's current view, the ratio of the similarity scores between each node pair within the agent's current view, and the action vector for each node. The most likely location of the agent is obtained by summing the contributions from each of the above modes at each discrete location. This array is normalized by dividing each element by the sum of all elements. This is hereinafter referred to as the "plausible location array".

[0398] A fourth mode for improving the determination of the location of an agent within a space according to the generated spatial map uses odometric information. This can be in the form of physical travel distance recorded by an odometer (distance traveled meter), virtual odometry by pixels or data odometry cells, or visual odometry. The odometric information may be determined from or be the same as the displacement determined according to the methods described above. The odometric information is stored together with other data sets within the map and, as described above with reference to FIG. 7, can be used to locate the agent at the nearest observation point. When the odometric information becomes available to the agent, the likely location of the agent may be updated taking this new information into account. This is done as follows.

[0399] The use of odometric information may be used to form a fourth array. In the absence of displacement data (as described above), the fourth array is generated using odometric data. This fourth array is generated by extending the first, second, third, or "likely location" arrays described above. The use of odometric information effectively updates the estimate of the location of the agent.

[0400] The odometric information is obtained over time from a previously determined position using any suitable technique. In particular, the velocity components over time are integrated to obtain the travel distance in two or three dimensions as required. Next, the first, second, third, or "likely location" arrays described above (e.g., the second visual representation 2110 providing a probability distribution of locations, or the probability distribution of the first or second arrays, or a combination thereof) are modified according to the odometric information (travel distance) to model the uncertainty of the likely location of the agent.

[0401] The fourth array is an extended array created by inputting values calculated from a two-dimensional Gaussian distribution into the cells of each non-zero element of a specific array to be extended, where the mean is the value obtained by adding the change in location calculated from the previous position and the odometry data, and the covariance is the value given by the odometry data. The entire array is multiplied element-wise by the value of the original element of the array to be extended. As a result, each element of the original first, second, third, or "plausible location" array is converted into a complete array of values representing a 2D Gaussian hill based on the odometry values. Alternatively, the contribution of each element of the first, second, or third array (i.e., the "plausible location" array) is summed with the corresponding odometry data to create a new array updated according to the odometry data. A single bump is shifted two-dimensionally according to the travel distance and distorted according to the covariance. Thereby, a fourth array is provided. This can be any combination of the first to third arrays, modified to take into account the probability of the agent's location based on odometry. Although the Gaussian distribution was exemplified above, any distribution suitable for the odometry data can be used.

[0402] The fourth array is called the "odom update array" and represents the best estimate of the agent's current location. When the agent later receives new information regarding the similarity magnitudes for all nodes within the view and the action vectors, it recalculates the possible location arrays in the manner described above. Next, this is multiplied by the odom update array to generate an updated estimate of the agent's location based on a combination of the new observations (represented by the plausible location array) and the previous location estimate (represented by the odom update array).

[0403] The above-described first to fourth arrays can be used in any combination. However, it should be understood that since the third array (using the action vector) effectively adjusts one of the first or second arrays, the third array needs to be used together with at least one of the first or second arrays. Similarly, the fourth array is an extension of the first, second, third, or fourth arrays, and thus, the fourth array can be used together with at least one of the first to third arrays.

[0404] In an exemplary combinatorial method for improving the determination of the location of agents in a space according to the generated spatial map, each of the first, second, third, and fourth arrays outlined above can be used.

[0405] In the first thing of the combinatorial method, the agent measures the similarity magnitude values and identifies all nearby nodes by identifying the nodes as "visible" if these values exceed the baseline value. If the list of visible nodes has changed since the previous iteration, the agent calculates the similarity magnitude field of the newly visible nodes and the similarity magnitude ratio field of the new pairs of visible nodes. According to the first mode above, an array is generated for each displayable node. According to the second mode above, one more array is generated for each pair of displayable nodes. According to the third mode, a further set of arrays is generated.

[0406] This may include generating a plurality of current observation points and using odometry to record the location of the current observation points, and the associated action vectors and / or similarity magnitude values.

[0407] In the second thing, all the arrays are summed element by element and the resulting array is normalized. Next, each element is compared with a low threshold. All elements below the low threshold are set to zero. This reduces the number of calculation times in terms of odometry as described in the above section, since only non-zero elements are considered. When the location is known with high precision, the number of non-zero elements is reduced, so the number of elements for which the location needs to be updated is also reduced, and the calculation efficiency is increased.

[0408] This generates a "synthetic" array, which is renormalized by dividing each element by the sum of the array as a third step. As a result, a discretized landscape (a plausible location array representing the probability of an agent existing at each location within the discretized field) is generated. This landscape can arbitrarily become complex. The advantage of this method is that it is not necessary to make assumptions about the statistical structure of the input.

[0409] As a third step, using the method described in the fourth mode above, the plausible location array is transformed according to the new odometry data (odom update array). Therefore, the reliability regarding the agent's location is updated iteratively and continuously. This is done by multiplying the previous reliability regarding the location by the new measurement when the location is observed.

[0410] The plausible location array or the odom update array can be used directly for localization in certain situations, but it is beneficial to calculate summary statistics of the agent's most plausible location along with a measure of the uncertainty of the location. This can be achieved by fitting a 2D Gaussian distribution to the output array. The peak of the Gaussian distribution becomes the location estimate, and its covariance gives the uncertainty of the estimate.

[0411] Another approach is to take the center of mass of the output array, where the weight contributed by each location is simply the value of the output array at that location. The measure of reliability at a location is given by the value of the output array at the center of mass location. A wide and flat output array results in a low peak at the center of mass, while a high certainty at a location results in an output array with a high and sharp maximum value. This approach works well when the output array is a single continuous peak, but the accuracy decreases in more complex cases (for example, the center of mass of a donut-shaped landscape is at the center of the donut, and the probability there is very low).

[0412] In the sec...

Claims

1. A computer-implemented method for navigating an agent in a multi-dimensional space, generating a data field that spatially corresponds to the multi-dimensional space in which the agent resides, wherein the data field incorporates observation comparison scale data for stored observations, and the stored observations include second sensor data that describes the environment of a first location in the multi-dimensional space, the generating including, the generating of the data field includes, acquiring a plurality of current observations generated by the agent at different corresponding positions in the multi-dimensional space, wherein the current observations include first sensor data that describes the environment of each of the different corresponding positions in the multi-dimensional space, the acquiring, comparing each of the current observations with the stored observations to obtain corresponding observation comparison scale data for the different corresponding positions in the multi-dimensional space, and incorporating the observation comparison scale data into the data field according to the different corresponding positions in the multi-dimensional space, the method further includes, determining a reliability scale associated with the stored observations by performing a reliability check on the data field, and managing a digital map of the multi-dimensional space by maintaining, updating, or deleting information in the digital map regarding the stored observations based on the reliability scale. A computer-implemented method.

2. The first sensor data and the second sensor data are first image data and second image data respectively, and the first sensor data and the second sensor data are arranged in a first vector and a second vector respectively extracted from the first image data and the second image data. The method according to claim 1.

3. The method further includes acquiring the first sensor data and the second sensor data, and acquiring the first sensor data includes, acquiring original first sensor data of the environment of the agent at each of the different corresponding positions in the multi-dimensional space, and processing the original first sensor data to reduce the size of the original first sensor data, so that the first sensor data is in a form with a reduced size relative to the original first sensor data. Obtaining the second sensor data includes: obtaining the original second sensor data of the environment of the agent at the first location; processing the original second sensor data to reduce the size of the original second sensor data and obtain the second sensor data, such that the second sensor data is in a reduced-size format with respect to the original second sensor data; and storing the second sensor data in the reduced-size format, the method according to claim 2. **Claim 4** Processing the original first sensor data and the original second sensor data includes applying one or more filters and / or masks to the original first sensor data and the original second sensor data to reduce the size of each dimension of the original first sensor data and the original second sensor data, respectively, the method according to claim 3. **Claim 5** The method according to any one of the preceding claims, further comprising navigating the agent to the saved observation point according to the digital map. **Claim 6** The digital map includes information regarding a displacement vector linking a series of saved observations, and navigating includes using the digital map information to navigate the agent to one of the series of saved observations, the method according to claim 5. **Claim 7** The data field includes an action vector field, the action vector field includes a plurality of action vectors, the observation comparison scale data is the plurality of action vectors, and each of the plurality of action vectors originates from one of the different corresponding first locations in the vector field and is directed in an estimated direction of the first location in the vector field, the method according to any one of the preceding claims. **Claim 8** Comparing each of the current observations with the saved observations to obtain corresponding observation comparison scale data for the different corresponding positions in the multi-dimensional space includes, for each current observation: Obtaining a plurality of first sub-regions of the first sensor data, each first sub-region describing a corresponding first portion of the environment around the agent at one of the different corresponding positions, each first portion being associated with a corresponding first direction from one of the different corresponding positions, the obtaining, Obtaining a plurality of second sub-regions of the second sensor data, each second sub-region describing a corresponding second portion of the environment around the first location, each second portion being associated with a corresponding second direction from the first location, the obtaining, and for each second sub-region, Comparing the second sub-region with each first sub-region using a similarity comparison measure to determine the first sub-region that is most similar to the second sub-region, Determining a relative rotation between the second direction associated with the second sub-region and the first direction associated with the most similar first sub-region, and including, The method is, Aggregating the relative rotations for the plurality of second sub-regions to obtain an action vector, the action vector indicating an estimated direction from one of the different corresponding positions to the first location, the observation comparison measure including the action vector, the aggregating, further including the method according to claim 7.

9. The determining the reliability measure associated with the saved observation by performing a reliability check on the data field is, Obtaining a model action vector field that spatially corresponds to the action vector field, the model action vector field including a plurality of model action vectors, each model action vector being directly directed to the first location such that the model action vector field converges radially to the first location, the obtaining, Comparing the plurality of action vectors of the action vector field with the spatially corresponding model action vectors of the model action vector field to determine an angular deviation between each action vector and the corresponding model action vector, To determine the first scalar measure, aggregating or averaging the angular deviation between the plurality of action vectors and the plurality of model action vectors, wherein the reliability measure includes the first scalar measure, the aggregating or averaging, the method according to claim 7 or 8.

10. By performing a reliability check on the data field, the determining the reliability measure associated with the stored observation is Determining a first number of the different corresponding positions where it is not possible to obtain an action vector; Determining a second number of the different corresponding positions where it is not possible to obtain an action vector; Dividing the first numerical value by the second numerical value to determine a second scalar measure, wherein the reliability measure includes the second scalar measure, the determining, the method according to claim 7 or 8.

11. Managing the digital map of the multi-dimensional space by maintaining, updating, or deleting the information in the digital map regarding the stored observation based on the reliability measure is Including comparing the reliability measure with one or more thresholds to determine whether to update, delete, or maintain the information in the digital map; Deleting the information in the digital map includes deleting the information regarding the stored observation in the digital map and adjusting the digital map to account for the deleted information, and / or Updating the information in the digital map is Obtaining second sensor data for replacement at the first location in the multi-dimensional space to form a stored observation for replacement; Regenerating the data stored observation field according to the replacement value, the method according to claim 9 or 10.

12. Further including aggregating the first scalar measure and the second scalar measure such that the reliability measure includes both the first scalar measure and the second scalar measure, the method according to claim 11.

13. The data field includes a similarity magnitude scalar field, and the similarity magnitude scalar field includes a plurality of similarity magnitude values indicating the similarity between the stored observations and the corresponding current observation values at each of the different corresponding positions. The observation comparison scale data is the plurality of similarity magnitude values, according to the method of any one of the preceding claims.

14. In order to obtain the corresponding observation comparison scale data for the different corresponding positions in the multi-dimensional space, comparing each of the current observations with the stored observations is to compare each current observation with the stored observations using a similarity comparison scale to determine a similarity magnitude value for each of the different corresponding positions, and the observation comparison scale data includes the similarity magnitude value, including the determining, according to the method of claim 13.

15. The method according to claim 14, wherein the current observation and the stored observation are arranged as vectors, and the similarity comparison scale is the inner product of these vectors.

16. Determining a reliability scale associated with the stored observations by performing a reliability check on the data field is to determine the maximum similarity magnitude value among the plurality of similarity magnitude values of the similarity scalar field and to divide the maximum similarity magnitude value by the average of the plurality of similarity magnitude values to determine a third scalar value, and the reliability scale includes the third scalar scale, including the dividing, according to the method of any one of claims 13 to 15.

17. The method according to any one of claims 13 to 17, further including generating a parameterized model of the similarity magnitude scalar field that approximates the similarity magnitude scalar field.

18. Determining the reliability scale according to claim 16 is performed with respect to the parameterized model, according to the method of claim 17.

19. Determining the reliability scale associated with the stored observations by performing a reliability check on the data field is to determine the variability between the similarity magnitude value and the model value of the parameterized model and to obtain a fourth scalar value based on the variability, and the reliability scale includes the fourth scalar scale, according to the method of claim 17 or 18.

20. The method according to any one of claims 16 to 19, further comprising aggregating the scalar measures of the third scalar measure and the fourth reliability measure such that both the first and second scalar measures are included in the foregoing.

21. The method according to claim 20 when dependent on claim 12, further comprising aggregating the first scalar measure, the second scalar measure, the third scalar measure, and the fourth scalar measure such that the reliability measure includes all of the first, second, third, and fourth scalar measures.

22. A computing device comprising a processor and a memory, wherein instructions are stored in the memory and, when executed by the processor, cause the computing device to perform the method according to any one of claims 1 to 21.

23. A computer program which, when executed by a processor, causes the processor to perform the method according to any one of claims 1 to 21.

24. A computer-implemented method for determining the position of an agent in a multi-dimensional space, obtaining a similarity field that spatially corresponds to the multi-dimensional space in which the agent resides, the similarity magnitude field including a plurality of similarity magnitude values indicating the similarity between a stored observation corresponding to a first location in the multi-dimensional space and corresponding current observations each corresponding to one of a plurality of different corresponding locations in the multi-dimensional space, the obtaining; obtaining, by the agent, a new observation at the current position of the agent in the multi-dimensional space; comparing the new observation with the stored observation to obtain a new similarity magnitude value; comparing a new location scale based on the new similarity magnitude value with a field location scale based on the similarity magnitude values of the similarity magnitude field to identify the most likely or matching field location scale for the new location scale, and identifying one or more possible positions of the agent in the multi-dimensional space based on one or more field locations in the similarity magnitude field corresponding to the most likely or matching field location scale.

25. The new location scale includes the new similarity magnitude value, and the field location scale is each of the similarity magnitude values of the similarity field, whereby the most plausible or matching field location scale is the matching similarity magnitude value of the similarity magnitude field or the range of similarity magnitude values of the similarity magnitude value field, and the range includes the new similarity magnitude value, identifying the one or more possible positions of the agent in the multi-dimensional space, retrieving the one or more field locations of the matching similarity magnitude value of the similarity magnitude field, or the range of similarity magnitude values of the similarity magnitude field, and determining the one or more positions of the agent in the multi-dimensional space according to the one or more field locations, the method according to claim 24. **Claim 26** The method is acquiring the stored observation, the stored observation including second sensor data describing the environment of the first location in the multi-dimensional space, by the acquiring, acquiring the plurality of current observation points generated by the agent at different corresponding positions in the multi-dimensional space, the current observation including first sensor data describing the environment of each of the different corresponding positions in the multi-dimensional space, by the acquiring, comparing each of the current observations with the stored observation to obtain, using a similarity comparison measure, a value of similarity magnitude for each of the different corresponding positions in the multi-dimensional space, and further including generating the similarity magnitude field by incorporating the similarity magnitude value into the similarity magnitude field according to the different corresponding positions in the multi-dimensional space, the method according to any one of claims 24 to 25. **Claim 27** The method according to claim 26, wherein the current observation and the stored observation are arranged as vectors, and the similarity comparison measure is the inner product of these vectors. **Claim 28** Obtaining a second similarity magnitude field that spatially corresponds to the multi-dimensional space in which the agent resides, where the second similarity magnitude field includes second stored observations corresponding to second positions within the multi-dimensional space and second plural similarity magnitude values indicating similarities between the second stored observations and corresponding current observations each corresponding to one of a plurality of different corresponding positions within the multi-dimensional space; said obtaining; Comparing the new observation with the second stored observation to obtain a second new similarity magnitude value; The new location scale includes a ratio between the new similarity magnitude value and the second new similarity magnitude value, the field location scale includes a plurality of field ratios between the plurality of similarity magnitude values and the second plurality of similarity magnitude values, and each field ratio of the plurality of field ratios is determined for different positions within the similarity magnitude field; The most likely or matching field location scale of the new location scale includes one or more field ratios; Identifying the one or more positions of the agent within the multi-dimensional space based on the one or more field locations within the similarity magnitude field corresponding to the most likely or matching field location scale; Extracting one or more field locations of one or more matching field ratios within the similarity field; Determining the one or more positions of the agent within the multi-dimensional space according to the one or more field locations; further including said comparing, the method according to any one of claims 24 to 27; [

29. ] The new location scale includes a ratio between the new similarity magnitude value and the second similarity magnitude value, the field location scale includes each similarity magnitude value of the similarity magnitude field and the plurality of field ratios, and the most likely or matching field location scale is both of the following: A matching similarity magnitude value of the similarity magnitude field, or a range of similarity magnitude values of the similarity magnitude field, the range including the new similarity magnitude value; and One or more matching field ratios; The method according to claim 28, when dependent on claim 25, wherein the matching similarity magnitude value and the matching field ratio are associated with the same one or more positions in the multi-dimensional space.

30. Before comparing the new location scale based on the new similarity magnitude value with the field location scale based on the similarity magnitude value of the similarity magnitude field, generating a parameterized model of the similarity magnitude field, such that the comparison is performed with respect to the model field location scale based on the model similarity magnitude value of the model similarity magnitude field, the generating further comprising the method according to any one of claims 24 to 29.

31. obtaining an action vector field that spatially corresponds to the multi-dimensional space in which the agent resides, the action vector field including a plurality of action vectors indicating directions from the plurality of different corresponding positions in the multi-dimensional space towards the first location of the stored observation, the obtaining further comprising, The method is, excluding one or more of the one or more possible positions of the agent based on the directions of the one or more action vectors associated with the one or more field locations, the method according to any one of claims 24 to 29.

32. generating an output field, the output field being based on the identified one or more positions of the agent in the multi-dimensional space, the output field having the same dimension as the similarity magnitude field and indicating a probability distribution of the identified one or more positions of the agent in the multi-dimensional space, the generating further comprising the method according to any one of claims 24 to 31.

33. Excluding one or more of the one or more possible positions of the agent based on the directions of the one or more action vectors associated with the one or more field locations is, adjusting the probability distribution of the identified one or more positions of the agent in the output field based on the directions of the one or more activity vectors associated with the one or more field locations, the method according to claim 32 when dependent on claim 31.

34. The method according to claim 32 or 33, wherein the probability distribution of the output field is based on the one or more possible positions of the agent determined according to each of claims 25, 28 and 31.

35. Receiving odometric data related to the movement of the agent; Determining an estimated distance by which the agent has moved from a previous position in the multi-dimensional space; Further comprising modifying the probability of the identified one or more positions of the agent in the multi-dimensional space based on the estimated distance, the method according to any one of claims 32 to 34.

36. The method according to claim 34 or 35, further comprising modifying the probability of the identified one or more positions of the agent in the multi-dimensional space based on the reliability measure according to any one of claims 1 to 21.

37. Obtaining a plurality of similarity magnitude fields that spatially correspond to the multi-dimensional space in which the agent resides, the plurality of similarity magnitude fields including a corresponding set of similarity magnitude values indicating the similarity between a corresponding plurality of stored observations corresponding to a plurality of corresponding first locations in the multi-dimensional space and corresponding current observations each corresponding to one of a plurality of different corresponding positions in the multi-dimensional space, the obtaining; Comparing the new observations with the stored observations for each of the plurality of stored observations to obtain a corresponding plurality of new similarity magnitude values; Comparing a plurality of new location scales based on the plurality of new similarity magnitude values with corresponding field location scales based on the set of similarity magnitude values of the plurality of similarity fields to identify the most likely or matching field location scale for each of the plurality of new location scales; Further comprising identifying one or more possible positions of the agent in the multi-dimensional space based on one or more field locations in the similarity magnitude field corresponding to the most likely or matching field location scale for each of the plurality of new location scales, the method according to any one of claims 24 to 36.

38. The method according to any one of claims 24 to 37, further comprising navigating the agent based on the one or more identified positions of the agent.

39. Navigate from the identified one or more positions of the agent in the multi-dimensional space to the first position along a first path from the identified one or more positions to the first position, to the first position corresponding to the stored observation. Determining the length of the first path to calculate a metric of edge traversal reliability, wherein the length of the path is inversely proportional to the metric of edge traversal reliability, said calculating. Comparing the metric of edge traversal reliability with a historical metric of edge traversal reliability, wherein the historical metric is calculated from an average length of a path previously traveled from the one or more positions of the agent to the first position in the multi-dimensional space, said comparing. When the metric of the reliability of the edge traversal is lower than a threshold amount than the historical metric. Determining that the first path is not reliable, further comprising the method according to claim 38 or claim 5 or 8.

40. When it is determined that the first path is not reliable. Recording new stored observations along the first path between the one or more positions of the agent and the first position in the multi-dimensional space. Excluding the first path from future navigation. Retrying the first path one or more times using different path parameters to optimize the first path by maximizing the metric of the edge traversal reliability. Warning the operator of the agent, performing at least one of the methods according to claim 39.

41. A computing device comprising a processor and a memory, wherein instructions are stored in the memory, and when executed by the processor, it causes the computing device to execute the method according to any one of claims 24 to 40.

42. Instructions are stored in the memory, and when executed by the processor, the system executes the method according to any one of claims 24 to 40.