Fusion of Vision and RF Sensors for Multi-Agent Tracking

By fusing wireless-based ranging information with visual odometry information, the system effectively addresses the challenge of accurate location identification in GPS-denied environments, achieving a high tracking accuracy of approximately 15 cm.

JP7698066B2Active Publication Date: 2025-06-24NEC CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023572190
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-05-26
Filing Date
2022-05-27
Publication Date
2025-06-24
Estimated Expiration
2042-05-27

AI Technical Summary

Technical Problem

Existing location identification and tracking technologies face challenges in environments where GPS signals are unavailable, particularly in providing accurate and robust location determination using multiple data sources.

Method used

A method and system that fuse wireless-based ranging information with visual odometry information to determine a device's location, allocating resources based on the final location estimate. This approach combines the strengths of infrastructure-assisted wireless localization and visual tracking to achieve robust and accurate location identification.

Benefits of technology

The solution achieves high tracking accuracy, with results showing a final location estimate accuracy of around 15 cm, significantly improving upon standalone wireless-based and visual tracking methods, which achieved accuracies of 40 cm and 32 cm respectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007698066000010
    Figure 0007698066000010
  • Figure 0007698066000011
    Figure 0007698066000011
  • Figure 0007698066000012
    Figure 0007698066000012
Patent Text Reader

Abstract

A method and system for determining a location of a device includes determining (704) a first location estimate using radio-based ranging information. A second location estimate is determined (708) using visual odometry information. The first and second location estimates are fused (710) based on radio and visual environmental conditions to determine a final location estimate. Resources are committed (606) based on the final location estimate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Related Application Information This application claims priority to U.S. Patent Application No. 63 / 196,387, filed on June 3, 2021, and U.S. Patent Application No. 63 / 194,262, filed on May 28, 2021, the entire contents of which are incorporated herein by reference.

Background Art

[0002] Technical Field The present invention relates to location identification and tracking, and more particularly to identifying the location of a device using multiple data sources.

[0003] Description of Related Art The ability to identify and track both people and assets in real time is useful for a variety of applications, particularly in environments where Global Positioning System (GPS) signals are not available. For example, such location identification may be used to facilitate collaboration between humans and robots.

Summary of the Invention

[0004] A method of determining the location of a device includes determining a first location estimate using wireless-based ranging information. A second location estimate is determined using visual odometry information. The first location estimate and the second location estimate are fused based on wireless environmental conditions and visual environmental conditions to determine a final location estimate. Resources are allocated based on the final location estimate.

[0005] A system for determining a device position includes a hardware processor and a memory. When executed by the hardware processor, the memory causes the hardware processor to determine a first position estimate using radio-based ranging information, determine a second position estimate using visual odometry information, fuse the first position estimate and the second position estimate based on radio environment conditions and visual environment conditions to determine a final position estimate, and store a computer program that causes resources to be allocated based on the final position estimate.

[0006] These and other features and advantages will become apparent from the following detailed description of its exemplary embodiments, read in conjunction with the accompanying drawings.

Brief Description of the Drawings

[0007] The present disclosure provides details in the following description of preferred embodiments with reference to the following figures.

[0008]

Figure 1

[0009]

Figure 2

[0010]

Figure 3

[0011]

Figure 4

[0012]

Figure 5

[0013]

Figure 6

[0014]

Figure 7

[0015]

Figure 8

[0016]

Figure 9

[0017]

Figure 10

DETAILED DESCRIPTION OF THE INVENTION

[0018] Dual-layer diversity can be used to improve the localization of multiple agents in space. For example, complementary tracking modalities such as passive / relative modalities (e.g., visual odometry) and active / absolute modalities (e.g., infrastructure-assisted wireless localization) are fused. To exhibit robustness even in unfamiliar environments without sacrificing the accuracy of tracking, various techniques that fuse the complementary strengths of algorithms and data-driven approaches are also adopted. In this way, for example, the complementary advantages of wireless location sensing and visual tracking can be combined to track agents with algorithms and data-driven technologies that share the burden of maintaining accuracy.

[0019] Passive tracking may include odometry-based techniques such as visual-inertial odometry that combines visual information from a camera and motion information from an inertial sensor. Passive tracking can provide position information within tens of centimeters under good visual conditions, but is weak under general environmental conditions such as dim lighting and featureless surfaces. Furthermore, since passive tracking provides relative position information, it is difficult to recover from adverse events and also has limitations in its ability to provide localization within a global reference frame.

[0020] Active tracking involves the use of fixed anchor nodes and provides absolute localization within a reference frame defined by the known positions of the anchor nodes. By using anchor tracking, the errors accumulated in passive tracking can be eliminated. However, there is a trade-off between the operating range and accuracy in active systems. High-resolution active tracking systems such as infrared, millimeter-wave, and acoustic systems have high accuracy but are limited to use within the line-of-sight range. Low-resolution active tracking systems that use general wireless network technologies can handle non-line-of-sight positioning and have a longer operating range, but their tracking accuracy is lower compared to high-resolution systems.

[0021] A hybrid approach that uses both passive and active systems enables scalable and accurate multi-agent tracking. Thus, the tracking device can include both a passive tracking device such as a stereo camera and a wireless interface. While the camera provides relative tracking based on visual odometry, the wireless interface provides absolute tracking by estimating the range and position of one or more anchor nodes in the environment.

[0022] The algorithmic solution can estimate the absolute position from wireless information and the relative translational information from visual data. On the other hand, the data-driven model can provide data filtering, feature synthesis, and fusion of different data modalities. Data filtering helps to separate the ranging estimates affected by non-line-of-sight propagation that degrade accuracy, and feature synthesis estimates the certainty of the absolute and relative position information by considering the environment and sensor artifacts. To provide robustness and maintain high tracking accuracy, fusion jointly considers the sensor streams and automatically adapts to appropriate sensor estimates based on the features and relative importance at each time. The fusion model does not need to capture the complex problem structure purely from data and can rely on the physics and geometry involved in each localization modality. This reduces the latency and computational requirements for real-time operation on resource-constrained devices.

[0023] This combined approach provides excellent accuracy in a variety of environments. For example, in some tests of the hybrid approach, a tracking accuracy within about 15 cm was achieved. In contrast, an accuracy of about 40 cm was achieved in equivalent wireless-based tests and about 32 cm in visual tracking. The combination of the algorithmic approach and the data-driven approach was equally effective in unknown environments, achieving a tracking accuracy of about 30 cm, compared to about 60 cm for the algorithmic approach alone and about 80 cm for the data-driven approach alone.

[0024] Referring now to FIG. 1, an exemplary multi-agent tracking environment 100 is shown. Environment 100 includes obstacles 102 and a plurality of agents 104 that can move freely around the obstacles. Obstacles 102 can include, for example, walls, furniture, doors, and other physical objects that may affect the freedom of movement of agents and / or the propagation of wireless signals. For example, even if a window obstructs a user's passage, the wireless signal can propagate freely.

[0025] Also shown is an anchor node 106. Anchor node 106 provides wireless ranging and positioning information to agents 104. The ranging and positioning information can include, for example, signal strength or time of flight indicating the distance between anchor node 106 and agent 104, and may further include direction information indicating the angle at which the signal is received from anchor node 106 by agent 104.

[0026] Anchor node 106 and agent 104 can use appropriate wireless location technologies. Representative technologies include WIFI (registered trademark) and ultra-wideband radio (UWB). These technologies have a balance between tracking resolution and non-line-of-sight operation, but other technologies such as millimeter-wave technology are also being considered.

[0027] In wireless-based location determination, the distance between agent 104 and the anchor node is estimated using wireless ranging. The estimated distances to three anchors 106, or one anchor 106 with multiple antennas, can provide direction information. Combining the estimated distances and directions with the known positions of the anchors 106 yields the absolute position within environment 100. The accuracy of position estimation depends on the accuracy of the individual distance and angle estimations. By combining two-way ranging of UWB with its wide bandwidth, high ranging accuracy that is robust to indoor multipath, for example, on the order of several tens of centimeters, can be obtained.

[0028] Visual odometry uses a stream of camera images to track the movement of agent 104. By utilizing changes in texture, color, and shape in consecutive camera images of a stationary environment, movement is tracked with high precision on the order of centimeters, for example. However, in environments with poor lighting, or few textures, or environments containing perceptual aliasing, the accuracy of visual odometry decreases. Therefore, inertial odometry can be used to improve the tracking accuracy from visual information sources.

[0029] Visual odometry provides high-precision relative positioning, but errors can accumulate over time. Once an error occurs, the relative measurements of visual odometry can propagate that error forward in time and cause a large drift in the estimated position. Wireless-based active tracking is less susceptible to such drift because each position estimate is independent of previous estimates. For this reason, errors caused by incorrect ranging estimates do not propagate, and UWB can improve the accuracy of absolute tracking over a long period.

[0030] Wireless-based position estimation provides a coarser resolution than vision-odometry-based estimation. Therefore, the fusion of the two types of sensor information brings their respective complementary advantages, and wireless-based estimation is used to remove the errors of more accurate visual odometry measurements that would otherwise occur.

[0031] Accurate multi-agent tracking may be used in a variety of different applications. For example, augmented reality and virtual reality games and collaborative work environments are becoming increasingly popular with the improvement of related technologies. In such virtual or augmented environments, by tracking users with respect to each other, it becomes possible to track their physical relationships and provide functions sensitive to their positions. For example, by recognizing that two users are close to each other in space, it becomes possible to provide an augmented reality function that enables interaction within the virtual environment.

[0032] Tracking can also be used in situations where accurate real-time position data is used to navigate indoor spaces. In emergencies such as fires, visibility may be blocked by smoke and debris. Agent tracking may be used to identify the location of emergency service personnel and building occupants to facilitate search and rescue.

[0033] By fusing visual / inertial data with wireless-based data, it becomes possible to connect high-resolution relative position estimates from visual odometry to a global reference frame. By measuring the position relative to known anchor positions, for example, by associating the coordinates of an agent with a position on a map, it becomes possible to localize the agent within a known space.

[0034] When fusing different methodologies, various models of algorithm-driven and data-driven types can be considered. Algorithmic solutions include filter-based data fusion that aims to minimize statistical noise using time-series data from individual sensors, such as Kalman filters and Bayesian filters. Data-driven solutions, including black-box systems, may focus on passive tracking. The black-box model may be trained to interpret the input sequence of raw sensor data and convert it into relative position and pose estimation. Using these solutions to fuse wireless-based ranging estimation and visual odometry camera images to predict absolute position estimation may be difficult to implement in a single model. Such models may be too dependent on the distribution of input data and may result in insufficient results in unfamiliar environments. However, by combining these two methodological approaches, the pitfalls of each can be avoided. The algorithmic model exhibits robustness even in untrained environments, and the data-driven model achieves high accuracy through sensor fusion.

[0035] Next, referring to FIG. 2, a block diagram of one of the agent devices 104 is shown. The agent device 104 can be any suitable mobile device such as a mobile phone, a headset, or an autonomous unit (e.g., a robot). The agent 104 can include a hardware processor and a memory 204, as well as appropriate software necessary to operate the device. The agent 104 can include an ultra-wideband transceiver 206 configured to communicate with the anchor node 106 in the environment 100 and other agent devices 104 on the floor in the environment 100.

[0036] Although UWB communication is particularly assumed, it should be understood that other radio frequency technologies may be used instead. As used herein, UWB may refer to signals in the frequency range of 3 - 6 GHz with a relatively wide bandwidth. It is also possible to use millimeter wave radio frequencies, for example, in the range between 24 GHz and 30 GHz or between 57 GHz and 66 GHz, or WIFI (registered trademark) signals at approximately 2.4 GHz and 5 GHz. However, UWB signals have a good balance between spatial resolution and the ability to locate non-line-of-sight objects.

[0037] An inertial sensor 208 and a camera 210 can be used to provide visual / inertial odometry information. The camera can be a monocular or stereo camera, and can also provide additional views and any appropriate number of video streams. The inertial sensor 208 can include an acceleration sensor capable of capturing six degrees of freedom, and provides information about how the device moves in space. The localization 212 integrates all this information. In the case of visual / inertial odometry, the localization 212 combines visual information indicating the movement of the agent 104 with directly measured acceleration information to estimate how the agent 104 moves in space. The localization 212 can select between the visual / inertial odometry estimate and the estimate generated from the UWB transceiver by determining which is more appropriate for a given environment and conditions.

[0038] In the case of UWB, the distance R to a known anchor node i i can be used to measure the time of flight. Distance estimation can be performed between the device 104 and a plurality of known anchor nodes 106 (e.g., at least three), or a single anchor node 106 if angle information (e.g., from a plurality of antennas) is available. To solve for the absolute position estimation of the agent 104, multilateration may be used based on the known positions of the anchor nodes 106. If there are n anchor nodes 106, each having a fixed position (x i , y i ), the absolute two-dimensional position (x, y) of the device can be estimated by minimizing the error

Number

[0039] Factors affecting the accuracy of location determination include multilateration optimization errors and environmental conditions. These may appear as inaccurate ranging estimations. The error of multilateration optimization can be estimated as the output of the optimization solution, but it is more difficult to quantify the error due to environmental conditions. Furthermore, errors resulting from the variation of wireless-based ranging can have a significant impact on the accuracy of location determination.

[0040] Considering an indoor environment with five exemplary anchor nodes 106, agent 104 can move along a trajectory within the environment that exposes it to both line-of-sight and non-line-of-sight paths to various anchors. Various scenarios are possible, where different numbers of anchor nodes 106 are exposed to agent 104 via line-of-sight paths and the remainder are blocked via non-line-of-sight paths, for example by obstacles 102. Having line-of-sight paths to all available anchors 106 produces superior accuracy from wireless-based localization compared to scenarios where one or more anchors 106 are blocked via non-line-of-sight paths. In a mixed scenario where some anchors are in line-of-sight to agent 104 and some are in non-line-of-sight, an optimal number of anchor nodes 106 is selected. Filtering out non-line-of-sight anchor nodes can reduce errors, but without knowledge of the environmental conditions, it is not always straightforward to determine the optimal set of anchors.

[0041] Non-line-of-sight distance estimation can be inaccurate even at short distances. To identify the quality of the estimate, measurements of received signal power exhibit highly discriminative behavior. The received signal power can be estimated as follows:

Equation

[0042] For example, as the non-line-of-sight path distance increases, the received signal strength may vary significantly, while along a line-of-sight path, the signal strength may remain above a high threshold even as the position changes significantly. Thus, the received power Pi between anchor i and agent 104 i can function as an effective discriminative feature for capturing the impact of non-line-of-sight paths on the accuracy of position estimation, along with the distance Ri. i

[0043] The extracted features (P, R) help to identify accurate distance estimates when a sufficient number of anchors (e.g., three or more) are available, indirectly capturing the certainty of position estimation for subsequent fusion.

[0044] The number of anchor nodes 106 visible from an agent 104 may be relatively small. For optimal anchor selection, machine learning models such as support vector machines and logistic regression can be used. Separate anchor classification datasets can be derived to train a classifier that selects the optimal K anchors for location identification from the collected data. The model can be fitted with the distances from all the anchors 106 and their corresponding received powers as inputs. The best set of anchors that provides the minimum error compared to known ground truth can be set as a binary output vector.

[0045] Multi-output classification using a classifier chain utilizes the correlation between the anchors 106. Multiple different models (e.g., support vector machines, logistic regression, random forests) can be optimized using grid search to adjust their parameters. After grid search, the model with the best performance can be selected. The model helps to filter the inputs, while multi-positioning is responsible for position estimation.

[0046] Anchor selection filters non-line-of-sight anchors to improve position estimation, but the device may not be able to utilize three available line-of-sight anchors, resulting in a decrease in location identification accuracy. Signal strength information and distance can be combined with the absolute position estimation from multi-positioning to form a wireless-based input to sensor fusion. Thus, the input from the wireless-based path to location identification 212 is represented as follows.

Number

Number

[0047] Visual odometry can perform tracking, local mapping, pose optimization, and loop closing. It tracks using a stream of image frames (e.g., stereo camera images) and incrementally and relatively locates the device in frame units. This can be done by extracting features from the images and establishing correspondences of matching keypoints between frames. At a given time instance t, two or more stereo frames

Number

[0048] The final features can be counted by collating the features across all stereo images. The matching features are used to find correspondences with a previous reference frame

Number

[0049] Since visual odometry provides relative tracking, temporary environmental artifacts such as limited visual features or dynamic scenes can degrade a small number of displacement estimates and lead to the accumulation of continuous errors over time. Instead of using the final relative estimate of the position, the relative displacement r v of the translation and the orientation θv can be directly used for sensor fusion. Even if visual odometry generates a temporary displacement error, the resulting error propagation is only in displacement, and the orientation continues to track the absolute trajectory direction. Therefore, the error is transient and does not propagate. The relative estimated values (r v , θ v ) can be fused with the radio-based absolute position estimated values (x u , y u ) to eliminate even transient errors.

[0050] Even in the absence of error drift, there may be errors in the relative estimated values themselves. Short-term environmental artifacts such as dynamic lighting occlusion can cause significant position inaccuracies even in a visually characteristic environment. To compensate for this, additional functions may be used to grasp the certainty of tracking estimation.

[0051] The features extracted from the image determine the accuracy and robustness of the tracking. The number of keypoints in the image can vary based on environmental factors such as different lighting conditions. If the number of keypoints exceeds a threshold (e.g., about 500), the error rate is relatively constant, but if the number of keypoints is below the threshold, the error rate increases. When the number of keypoints is particularly small (e.g., less than about 100), the tracking may fail completely. The keypoint match grasps the certainty of the estimation provided by the tracking. This confidence feature M is combined with the relative position estimated values (r, θ) generated by visual odometry and provided as a synthetic odometry feature input to sensor fusion. The input from the visual odometry path can be represented as follows.

Equation

[0052] Next, referring to FIG. 3, an example of sensor fusion is shown. In the wireless-based branch, wireless data 310 is collected from the transceiver 206 of agent 104 and may also be collected from anchor node 106. Block 312 performs anchor selection as described above and identifies a set of anchor nodes that provide the most reliable ranging and angle information. For example, this is done by observing the signal strength and can indicate whether the anchor has a line-of-sight path to the agent. Block 314 performs multi-positioning to generate a wireless-based position estimate. These estimates can be used to generate wireless-based features 316.

[0053] In the odometry path, visual / inertial data 320 is used to generate a set of features. For example, visual simultaneous localization and mapping can be used for feature detection 322. Feature matching 324 can establish corresponding keypoint matches between images. Mapping 326 maps the position of the image to coordinates in the environment, and pose graph optimization uses this information to identify the relative position information of agent 104. A pose graph 328 may be generated from the relative position information. Odometry features 329 are generated based on the relative position information.

[0054] Feature fusion 330 combines the wireless-based features and the odometry features 329 to generate an absolute position estimate 332. This fusion can adopt a cross-attention model to combine the features, which are then processed by a long short-term memory (LSTM) layer and a fully connected layer to output the absolute position estimate value 332.

[0055] Next, referring to FIG. 4, an exemplary neural network model for generating the wireless-based feature 316 is shown. This model includes a set of two-dimensional convolutional sections, and each section includes a convolutional layer with a leaky rectified linear unit (ReLU) activation function, a batch normalization layer, and a dropout layer. Such a first section 402 has 16 units, and the subsequent two sections 404 and 406 each have 64 units. A flattened dense dropout layer 408 of 128 units follows, and a dense dropout layer 410 of 64 units and a dense dropout layer 412 of 32 units further process the output to generate the wireless-based feature 316.

[0056] Next, referring to FIG. 5, an exemplary neural network model for generating the odometry-based feature 329 is shown. This model starts with a set of one-dimensional convolutional sections that include a convolutional layer with a leaky ReLU activation function, a batch normalization layer, and a dropout layer. Such a first section 502 has 16 units, and the subsequent two sections 504 and 506 each have 64 units. A flattened dense dropout layer 508 of 64 units follows, and a dense dropout layer 510 of 64 units and a dense dropout layer 512 of 32 units further process the output to generate the odometry-based feature 329.

[0057] When processing the wireless-based feature 316 and the odometry feature 329, the feature fusion model 330 can first prepare the features by passing them through a simple convolutional neural network, and embed the features related to position and certainty into more representative features that capture both position and certainty.

[0058] The fusion model 330 can use attention to adaptively weight the features of each sensor path and leverage their complementary properties. In particular, cross-attention can be used to weight each sensor relative to the others in order to extract the correlation between sensors. The model can weight the wireless-based estimate higher when the odometry-based estimate is troubled by unfavorable environmental conditions, and can weight the odometry-based estimate higher when the wireless-based estimate is troubled by non-line-of-sight paths to the anchor 106. Cross-attention can further incorporate the advantages of self-attention, which takes into account features correlated with the tracking error.

[0059] Cross-attention mask A rv and A vr refers to the wireless-based feature 316 (F r ) and the odometry-based feature 329 (F v ) can be jointly learned using. The mask is defined as follows:

Number

[0060] After the mask is learned, each mask is applied to the respective sensor features element-wise, and then the masked features are merged by concatenation, and the fused features

Number

[0061] When agent 104 is in a dimly lit place, the attention to wireless-based features may increase and the attention to odometry-based features may decrease. When agent 104 does not have a sufficient line-of-sight path to anchor 106, the attention to wireless-based features may decrease and the attention to odometry-based features may increase. Considering the independence and complementarity of the two sensor paths and their respective environmental artifacts, cross-attention helps provide robustness against various adverse conditions.

[0062] The model can be trained on a dataset collected in various indoor environments (such as office buildings, homes, conference centers, etc.) with different landscapes, textures, and lighting conditions. This diverse dataset allows the model to avoid overfitting. During learning, the input data is normalized by subtracting the mean of the dataset. An appropriate loss function can be used during learning, and the loss function can be minimized to adjust the model parameters. Once the model is trained, it can be deployed in any environment and generalize to environments not included in the training dataset.

[0063] Next, referring to FIG. 6, a method for identifying the positions of agents within an environment and interacting with the agents is shown. Block 602 identifies the positions of agents within an environment, such as inside a building. The identification of the positions of agents can employ the fusion of wireless signal ranging, inertial sensor information, and visual odometry. Each individual agent 104 can communicate its respective data to a central system that calculates the positions of all agents 104.

[0064] The central system uses this data to generate a map of the environment at block 604. For example, the positions of the agents can be overlaid on an existing map of the building interior to identify the positions of each agent. Block 604 may further include generating a map of the building interior itself by tracking the movement of devices within the building. For example, by tracking the movement of a mobile phone, the passageways within the building can be identified.

[0065] Block 606 deploys resources based on the map. This map can be used for various purposes. For example, asset tracking is used to identify the inventory and inventory levels within a store, and deploying resources may include replenishing items that are in short supply. Additionally, in the event of a fire or natural disaster, device tracking may be used for emergency response purposes to identify the positions of people within the building for rescue. Mapping may also be used to assist responders in moving through the building by following the paths generated by tracking devices within the building. Other applications include tracking workers and assets at construction sites where it is not possible to introduce an infrastructure-based positioning system, and tracking workers in large factories where introducing a dedicated positioning system is costly. In the case of an augmented reality system, the map helps identify when users are approaching each other, and deploying resources also includes electronically displaying augmented reality elements to reflect their positions.

[0066] Next, referring to FIG. 7, further details regarding the determination 602 of the device position are shown. Block 702 collects wireless data of the agent 104 using, for example, the UWB transceiver 206. Next, block 704 uses the collected wireless data to determine the wireless-based feature 316 as described above. Block 706 collects visual information of the agent using, for example, the camera 210. Block 708 uses the collected visual information to determine the odometry-based feature 329. Block 710 fuses the wireless-based feature and the odometry-based feature to provide a position estimation that is sensitive to the quality of the sensor data obtained by the agent 104.

[0067] Embodiments described herein may be entirely hardware, may be entirely software, or may include both hardware and software elements. In a preferred embodiment, the invention is implemented in software, including but not limited to firmware, resident software, microcode, etc.

[0068] Embodiments can include a computer program product accessible from a computer-usable or computer-readable medium providing program code for use by or in connection with a computer or any instruction execution system. A computer-usable medium or computer-readable medium can include any device that stores, communicates, propagates, or transports a program for use by or in connection with an instruction execution system, apparatus, or device. The medium can be magnetic, optical, electronic, electromagnetic, infrared, or semiconductor system (or apparatus or device) or a propagation medium. The medium can include computer-readable storage media such as semiconductor or solid state memory, magnetic tape, removable computer diskette, random access memory (RAM), read only memory (ROM), rigid magnetic disk, and optical disk.

[0069] Each computer program can be tangibly stored on a machine-readable storage medium or device (such as a program memory or magnetic disk) readable by a general-purpose or special-purpose programmable computer to configure and control the operation of a computer when the storage medium or device is read by the computer for executing the procedures described herein. The system of the present invention can also be considered to be implemented on a computer-readable storage medium constituted by a computer program, in which case the configured storage medium causes the computer to operate in a specific predetermined manner to execute the functions described herein.

[0070] A data processing system suitable for storing and / or executing program code may include at least one processor directly or indirectly coupled to memory elements via a system bus. The memory elements can include local memory employed during actual execution of the program code, bulk storage, and cache memory that provides at least some temporary storage of at least some program code to reduce the number of times code is retrieved from bulk storage during execution. Input / output or I / O devices (including, but not limited to, keyboards, displays, pointing devices, etc.) can be coupled to the system directly or via intervening I / O controllers.

[0071] A network adapter can also be coupled to the system to enable the data processing system to become coupled to other data processing systems or remote printers or storage devices via intervening private or public networks. Modems, cable modems, Ethernet cards are but a small part of the types of network adapters currently available.

[0072] As used herein, the term "hardware processor subsystem" or "hardware processor" can refer to a processor, memory, software, or a combination thereof that cooperate to perform one or more specific tasks. In useful embodiments, the hardware processor subsystem can include one or more data processing elements (e.g., logic circuits, processing circuits, instruction execution devices, etc.). The one or more data processing elements can include a central processing unit, an image processing unit, and / or a controller based on a separate processor or computing element (e.g., logic gates, etc.). The hardware processor subsystem can include one or more on-board memories (e.g., cache, dedicated memory arrays, read-only memory, etc.). In some embodiments, the hardware processor subsystem can include one or more memories that can be on-board or off-board, or dedicated for use by the hardware processor subsystem (e.g., ROM, RAM, basic input / output system (BIOS), etc.).

[0073] In certain embodiments, the hardware processor subsystem can include one or more software elements and can execute them. The one or more software elements can include an operating system and / or one or more applications and / or specific code to achieve a particular result.

[0074] In other embodiments, the hardware processor subsystem can include dedicated circuits that execute one or more electronic processing functions to achieve a specified result. Such circuits can include one or more application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), and / or programmable logic arrays (PLAs).

[0075] These and other variations of the hardware processor subsystem are also contemplated in accordance with embodiments of the present invention.

[0076] Referring now to FIG. 5, an exemplary arithmetic unit 500 according to an embodiment of the present invention is shown. The arithmetic unit 500 is configured to perform classifier enhancement.

[0077] The arithmetic unit 500 can be embodied as any type of computing or computer device capable of performing the functions described herein, including but not limited to computers, servers, rack-based servers, blade servers, workstations, desktop computers, laptop computers, notebook computers, tablet computers, mobile arithmetic units, wearable arithmetic units, network devices, web devices, distributed arithmetic systems, processor-based systems, and / or user electronic devices. Additionally or alternatively, the arithmetic unit 500 may be embodied as one or more compute threads, memory threads, or other components of a rack, thread, arithmetic chassis, or physically decomposed arithmetic unit.

[0078] As shown in FIG. 6, the arithmetic unit 600 illustratively includes a processor 610, an input / output subsystem 620, a memory 630, a data storage device 640, and a communication subsystem 650, and / or other components and devices commonly found in a server or similar arithmetic unit. The arithmetic unit 600 may include other or additional components (e.g., various input / output devices) commonly found in a server computer in other embodiments. Further, in some embodiments, one or more of the exemplary components may be incorporated into or otherwise form part of another component. For example, the memory 630, or a portion thereof, may be incorporated into the processor 610 in some embodiments.

[0079] Processor 610 can be embodied as any type of processor capable of executing the functions described herein. Processor 610 may be embodied as a single processor, a multiprocessor, a central processing unit (CPU), a graphics processing unit (GPU), a single or multi-core processor, a digital signal processor, a microcontroller, or other processor or processing / sing / control circuitry.

[0080] Memory 630 can be embodied as any type of volatile or non-volatile memory or data storage capable of executing the functions described herein. During operation, memory 630 can store various data and software used during the operation of arithmetic unit 600, such as an operating system, applications, programs, libraries, and drivers. Memory 630 is communicatively coupled to processor 610 via I / O subsystem 620 and can be embodied as circuitry and / or components for facilitating input / output operations with processor 610, memory 630, and other components of arithmetic unit 600. For example, I / O subsystem 620 may be embodied as, or otherwise include, a memory controller hub, an input / output control hub, a platform controller hub, an integrated control circuit, a firmware device, a communication link (e.g., a point-to-point link, a bus link, a wire, a cable, a light guide, a printed circuit board trace, etc.) and / or other components and subsystems for facilitating input / output operations. In some embodiments, I / O subsystem 620 forms part of a system-on-chip (SOC) and may be incorporated into a single integrated circuit chip together with processor 610, memory 630, and other components of arithmetic unit 600.

[0081] The data storage device 640 can be embodied as any type of device or apparatus configured for short - term or long - term storage of data, such as, for example, a memory device and circuitry, a memory card, a hard disk drive, a solid - state drive, or other data storage devices. The data storage device 640 can store program code 640A for identifying the position of devices in the environment based on wireless ranging information, inertial sensor information, and pressure sensor information, and program code 640B for mapping the interior of a building and responding to the positioning of the devices. The communication subsystem 650 of the computing device 600 can be embodied as any network interface controller or other communication circuitry, device, or collection thereof that can enable communication between the computing device 600 and other remote devices via a network. The communication subsystem 650 can be configured to implement such communication using any one or more communication technologies (e.g., wired or wireless communication) and associated protocols (e.g., Ethernet, InfiniBand®, Bluetooth®, Wi - Fi®, WiMAX® etc.).

[0082] As shown in the figure, the computing device 600 can also include one or more peripheral devices 660. The peripheral devices 660 may include any number of additional input / output devices, interface devices, and / or other peripheral devices. For example, in some embodiments, the peripheral devices 660 can include a display, a touch screen, a graphics circuit, a keyboard, a mouse, a speaker system, a microphone, a network interface, and / or other input / output devices, interface devices, and / or peripheral devices.

[0083] Of course, the computing device 600 can also include other elements (not shown), and certain elements can also be omitted, as would be readily envisioned by those skilled in the art. For example, various other sensors, input devices, and / or output devices can be included in the computing device 600 depending on the particular implementation of the same, as would be readily understood by those skilled in the art. For example, various types of wireless and / or wired input and / or output devices can be used. Further, processors, controllers, memories, etc. can be added and utilized in various configurations. These and other variations of the processing system 600 are readily contemplated by those skilled in the art in view of the teachings of the present invention provided herein.

[0084] Next, refer to FIGS. 9 and 10. Referring to FIGS. 7 and 8, exemplary neural network architectures are shown, which can be used to implement a part of this model. A neural network is a generalized system, and its function and accuracy are improved by being exposed to additional empirical data. A neural network is learned by being exposed to empirical data. During training, the neural network stores and adjusts a plurality of weights applied to the input empirical data. By applying the adjusted weights to the data, it is possible to identify that the data belongs to a specific class predefined from a set of classes, or to output the probability that the input data belongs to each class.

[0085] Empirical data (also called training data) obtained from a series of examples is formatted as a string of values and supplied as input to a neural network. Each example is associated with a known result or output. Each column is represented as a pair (x, y), where x represents the input data and y represents the known output. The input data can have various data types and can contain multiple different values. The network can have one input node for each value that makes up the input data of an example, and separate weights can be applied to each input value. The input data can be formatted as a vector, an array, or a string, for example, depending on the architecture of the neural network being constructed and trained.

[0086] A neural network "learns" by comparing the neural network output generated from the input data with the known values of the examples and adjusting the stored weights to minimize the difference between the output value and the known value. The adjustment can be done to the stored weights through backpropagation, and the influence of the weights on the output value is determined by calculating the mathematical gradient and adjusting the weights in a way that shifts the output to the minimum difference. This optimization, called gradient descent, is a non-limiting example of how training is performed. A subset of examples with known values that were not used in training can be used to test and validate the accuracy of the neural network.

[0087] During operation, a trained neural network can be used for new data that was not previously used for training or validation through generalization. The adjusted weights of the neural network can be applied to the new data, and the weights estimate a function developed from the training examples. The parameters of the estimated function captured by the weights are based on statistical inference.

[0088] In a layered neural network, the nodes are arranged in layers. An exemplary simple neural network has an input layer 920 of source nodes 922 and a single computational layer 930 having one or more computational nodes 932 that also function as output nodes, with a single computational node 932 for each possible category into which an input example can be classified. The input layer 920 can have a number of source nodes 922 equal to the number of data values 712 of the input data 910. The data values 912 of the input data 910 can be represented as a column vector. Each computational node 932 of the computational layer 930 generates a weighted linear combination of the values from the input data 910 supplied to the input nodes 920 and applies a non - linear activation function that is differentiable with respect to the sum. An exemplary simple neural network can perform classification on linearly separable examples (e.g., patterns).

[0089] Deep neural networks, such as multi - layer perceptrons, can have an input layer 920 of source nodes 922, one or more computational layers 930 having one or more computational nodes 932, and an output layer 940 having one output node 942 for each category into which an input example might be classified. The input layer 920 can have a number of source nodes 922 equal to the number of data values 912 of the input data 910. The computational nodes 932 of the computational layer 930 are between the source nodes 922 and the output nodes 942 and are not directly observable, so they are also called hidden layers. Each node 932, 942 of the computational layer generates a weighted linear combination of the values output from the nodes of the previous layer and applies a non - linear activation function that is differentiable over the range of the linear combination. The weights applied to the values from each previous node can be represented, for example, as w1, w2,... w n-i , w n and so on. The output layer provides the overall response of the network to the input data. Deep neural networks can be fully connected, where each node of a computational layer is connected to all nodes of the previous layer, or the connections between layers can be in other configurations. If links between nodes are missing, the network is called partially connected.

[0090] For the training of a deep neural network, there are two phases: a forward phase in which the weights of each node are fixed and the input is propagated through the network, and a backward phase in which the error value is propagated backward through the network to update the weight values.

[0091] The computing nodes 932 of one or more computing (hidden) layers 930 perform a non-linear transformation on the input data 912 that generates a feature space. Classes or categories may be more easily separable in the feature space than in the original data space.

[0092] In the specification, references to "an embodiment" or "an embodiment" of the present invention, and other variations, mean that the specific features, structures, characteristics, etc. described in connection with the embodiment are included in at least one embodiment of the present invention. Accordingly, the expressions "in one embodiment" or "in an embodiment" that appear throughout this specification, and any other variations, do not necessarily all refer to the same embodiment. However, it should be understood that the features of one or more embodiments can be combined in view of the teachings of the present invention provided herein.

[0093] For example, in the case of "A / B", the use of any of the following, such as " / ", "and / or", "at least one of", like "A and / or B", "at least one of A and B", is intended to include the selection of only the first-listed option (A), or only the second-listed option (B), or the selection of both options (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C", such expressions are intended to include the selection of only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or the selection of only the first and second-listed options (A and B), the selection of only the first and third-listed options (A and C), the selection of only the second and third-listed options (B and C), or the selection of all three options (A and B and C). This can be extended by the number of listed items.

[0094] The above is to be understood as illustrative and exemplary in all respects and not restrictive. The scope of the invention disclosed herein is determined from the claims construed in accordance with the full breadth permitted by the patent law, rather than from the detailed description. The embodiments shown and described herein are merely illustrative of the invention, and it should be understood by those skilled in the art that various modifications can be made without departing from the scope and spirit of the invention. Those skilled in the art can implement various other combinations of features without departing from the scope and spirit of the invention. Thus, while the aspects of the invention have been described with the detail and particularity required by the patent law, what is desired to be claimed and protected by patent is as set forth in the appended claims.

Claims

1. A computer-implemented method for determining a device position, comprising: Determining radio-based location features using collected radio data (704); Determining odometry-based location features using collected visual information (708); Adopting a cross-attention model to combine the radio-based location features and the odometry-based location features based on radio environmental conditions indicating the propagation environment of radio signals and visual environmental conditions indicating visually recognizable physical factors, and processing using a long short-term memory (LSTM) layer and a fully connected layer (710) to determine a final position estimate; Arranging the device (606) based on the final position estimate.

2. The method according to claim 1, further comprising determining the radio signal strength of a plurality of anchor devices in the radio-based location features.

3. The method according to claim 2, wherein determining the radio signal strength of the plurality of anchor devices includes determining that the number of visually recognizable anchor devices is less than a threshold.

4. The method according to claim 1, further comprising determining a keypoint match between image frames in the odometry-based location features.

5. The method according to claim 4, wherein determining the keypoint match between image frames includes determining that the number of keypoint matches is less than a threshold.

6. The method according to claim 1, further comprising determining radio-based ranging information using an ultra-wideband transceiver and determining the visual information using a stereo camera.

7. A system for determining a device position, comprising: A hardware processor (810); and A memory (840) storing a computer program, the computer program causing the hardware processor to: Determine radio-based location features using collected radio data (704); Determine odometry-based location features using collected visual information (708); Based on the wireless environmental conditions indicating the propagation environment of wireless signals and the visual environmental conditions indicating physical factors that can be visually recognized, a cross-attention model is adopted to combine the wireless-based position features and the odometry-based position features, and processed using a long short-term memory (LSTM) layer and a fully connected layer (710), and a procedure for determining the final position estimation, A system, which is a computer program for causing the device to execute a procedure of arranging the device (606) based on the final position estimation.

8. The computer program according to claim 7, which further causes the hardware processor to execute a procedure for determining the wireless signal strength of a plurality of anchor devices in the wireless-based position features.

9. The system according to claim 8, wherein the computer program further causes the hardware processor to execute a procedure for determining that the number of visually recognizable anchor devices is less than a threshold value.

10. The system according to claim 7, wherein the computer program further causes the hardware processor to execute a procedure for determining the matching of keypoints between image frames in the odometry-based position features.

11. The system according to claim 10, wherein the computer program further causes the hardware processor to execute a procedure for determining that the number of matched keypoints is less than a threshold value.

12. The system according to claim 7, further comprising an ultra-wideband transceiver configured to capture the wireless-based ranging information and a stereo camera configured to capture the visual information.

Citation Information

Patent Citations

  • Distributed localization system and method and self-localizing device

    JP2018510366A

  • Moving object, control method for moving object, and program

    JP2020095339A

  • Information terminal device, method, and program

    JP2021050969A

  • Robot movement control method, apparatus and robot using the same

    US20200206921A1