Temporal separation in large-scale image-based positioning

An automated system for robot localization addresses the challenge of scaling robot fleets by implementing image-based positioning models with efficient data strategies and asynchronous deployment, ensuring accurate navigation and adaptability across diverse environments.

JP7860254B2Active Publication Date: 2026-05-15BEAR ROBOTICS INC
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
BEAR ROBOTICS INC
Filing Date
2023-03-02
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

The challenge of scaling robot localization technology from a single instance deployment to a production scale across diverse environments is significant, especially with large fleets of robots, requiring manual adjustments and substantial technical effort.

Method used

An automated robot training and deployment system for image-based positioning models, utilizing data collection and storage strategies, model training, and asynchronous deployment to manage large robot swarms efficiently.

Benefits of technology

Facilitates reliable and scalable robot localization by generating and maintaining image-based positioning models across multiple environments, ensuring accurate robot navigation and adaptability to environmental changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007860254000004
    Figure 0007860254000004
  • Figure 0007860254000005
    Figure 0007860254000005
  • Figure 0007860254000006
    Figure 0007860254000006
Patent Text Reader

Abstract

A computer-implemented method and apparatus for generating an image-based positioning model for mobile robot navigation. The method includes the steps of: generating collected data by performing data collection at a plurality of different service locations where a mobile robot swarm can be deployed; dividing the collected data into a plurality of blocks of continuous portions of the collected data; and generating a first image-based positioning model and a second image-based positioning model for a first service location and a second service location, respectively, of the plurality of different service locations, using the collected data. The method further includes the steps of distributing the first image-based positioning model and the second image-based positioning model on a first mobile robot and a second mobile robot, respectively, of the mobile robot swarm, and the first image-based positioning model and the second image-based positioning model are used to navigate the first service location and the second service location, respectively, of the plurality of different service locations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application claims the benefit of priority to U.S. Patent Application No. 63 / 268,792, filed on Mar. 2, 2022, which is hereby incorporated by reference in its entirety.

[0002] This disclosure generally relates to robot navigation and robot localization.

Background Art

[0003] Mobile robot localization technology seeks to provide a reliable solution for determining the position that a mobile robot may be at any given time. The estimation of the robot pose (e.g., position and orientation) with respect to a coordinate system poses various technical challenges and can be achieved using a robot navigation model.

Summary of the Invention

Problems to be Solved by the Invention

[0004] To design and deploy a user-defined model for robot operation in a particular environment, many components of the navigation model may have to be manually adjusted or “tweaked” until the user-defined model provides satisfactory performance. Considerable technical problems arise when scaling the robot deployment from a single instance deployment to a production scale across many diverse environments. Such technical challenges can be substantial when the robot fleet size includes a large number of robots (e.g., dozens, hundreds, and ultimately thousands of robots), and they multiply as the robot fleet size increases.

Brief Description of the Drawings

[0005] [Figure 1]As illustrated by several examples, this is a schematic representation of an environment in which a large number of mobile robots (e.g., a swarm of service robots) are deployed in individual locations or environments such as cafeterias, hospitals, or elderly care facilities. [Figure 2] This is a diagram of an exemplary mobile robot that could be deployed in a place such as a cafeteria. [Figure 3] This is a block diagram illustrating one perspective of the components and modules of a mobile robot relating to several examples. [Figure 4] This is a block diagram illustrating other aspects of the components and modules of several example mobile robots. [Figure 5] The output of several example image-based positioning models is shown. [Figure 6] This block diagram shows a model system that operates to generate and maintain image-based positioning models deployed on various mobile robots in one or more locations, according to several examples. [Figure 7] This flowchart shows the workflow for generating, maintaining, and deploying image-based positioning models, as illustrated in several examples. [Figure 8] This flowchart illustrates several example data preparation flows that can be performed by data collection and preparation modules to generate data stored in a cloud storage facility. [Figure 9] As illustrated by several examples, we present a random sample of raw data corresponding to several trips collected by a mobile robot, within the context of a map divided into grid cells of configurable size. [Figure 10] The diagram illustrates several exemplary block-based data splitting processes that can be performed by data collection and preparation modules during the splitting operation. [Figure 11] This flowchart illustrates the image-based positioning model development process, including several examples, which can be realized through model training and evaluation modules. [Figure 12] This flowchart illustrates several example image-based positioning model deployment processes that can be implemented using the model deployment module. [Figure 13] This flowchart illustrates several example image-based positioning model refresh processes that can be implemented by the model refresh module. [Figure 14] This flowchart illustrates several example model versioning processes that can be implemented by a model system. [Figure 15] This block diagram illustrates a software architecture that may be installed in any one or more of the devices described herein, according to several examples. [Figure 16] A schematic representation of a machine in which instructions (e.g., software, programs, applications, applets, apps, or other executable code) can be executed to cause the machine to perform any one or more of the methodologies discussed herein, as illustrated by several examples. [Figure 17] A directory structure for machine learning programs, illustrating several examples. [Figure 18] This is a schematic representation of the processing environment for several examples. [Modes for carrying out the invention]

[0006] To facilitate identification of discussions concerning any particular element or action, the most important digit or number in the reference number indicates the drawing number in which the element is first introduced.

[0007] The examples disclosed herein embody an automated robot training and deployment system for an image-based positioning model that can be scaled to large robot swarms. The examples disclosed herein also embody numerous data collection and data storage strategies or policies for data used in generating image-based positioning models.

[0008] A process and system for generating a single model for a given location are described for illustrative purposes. However, an example of a model generation and deployment pipeline that can be extended to multiple locations is also described.

[0009] Figure 1 is a schematic representation of an environment in which numerous mobile robots 104 (for example, a group of service robots) are deployed in individual locations 102 or environments such as a cafeteria, hospital, or elderly care facility. Depending on the location, the mobile robots 104 can perform a variety of functions within location 102. For example, if location 102 is a service location such as a cafeteria, the mobile robots 104 can operate to deliver food from the kitchen to the table and assist in transporting dishes, trash, etc., from the table to the kitchen.

[0010] Each mobile robot 104 is connected via a network 106 or multiple networks 106 to a cloud service 110 that resides on one or more server systems 108.

[0011] Figure 2 is a drawing of an exemplary mobile robot 104 that may be placed in a location 102 such as a cafeteria. The mobile robot 104 has a chassis or housing 202 that houses various components and modules, including a transport system having wheels or tracks that allow the mobile robot 104 to propel itself within the service location. The power source for the mobile robot 104 may be a rechargeable battery that provides the energy necessary to operate such components and modules. The operating and recognition systems are also housed within the housing 202. The housing 202 supports a number of trays 204 that can hold plates and other tableware that are transported from the kitchen to the table and from the table to the kitchen within location 102. The mobile robot 104 may also have manipulators (not shown) such as a robotic arm that allow the robot to perform tasks such as manipulating or lifting objects or opening doors.

[0012] The mobile robot 104 may include a number of sensors, including exeroceptive sensors for capturing information about the environment or location in which the mobile robot 104 can operate, and proprioceptive sensors for capturing information related to the mobile robot 104 itself. The operating system enables the mobile robot 104 to map its environment and plan a path to reach its destination. The recognition system processes the sensor data to recognize objects, detect obstacles, and interpret the operating environment of the mobile robot 104.

[0013] Examples of external sensors can include visual sensors (e.g., 2D, 3D, depth, and RGB cameras), optical sensors, acoustic sensors (e.g., microphones or ultrasonic sensors), proximity sensors (e.g., infrared (IR) transceivers, ultrasonic sensors, photoresistors), tactile sensors, temperature sensors, exploration and position sensors (e.g., GPS (Global Positioning System) sensors). Visual distance measurement and visual SLAM (Simultaneous Localization And Mapping) can be useful for the mobile robot 104 to operate in all indoor and outdoor environments where lighting conditions are reasonable and can be maintained in some examples. 3D cameras, depth, and stereo vision cameras provide pose (e.g., position and orientation) information.

[0014] Examples of proprioceptive sensors include inertial sensors (e.g., tilt and acceleration), accelerometers, gyroscopes, magnetometers, compasses, wheel encoders, and temperature sensors. An inertial measurement unit (IMU) within the mobile robot 104 can include multiple accelerometers, gyroscopes, as well as magnetometers and barometers. The instantaneous pose (e.g., position and orientation), velocity (linear velocity, angular velocity), acceleration (linear acceleration, angular acceleration), and other parameters of the mobile robot 104 can be obtained through the IMU.

[0015] Figure 3 is a block diagram illustrating one perspective of the components and modules of a mobile robot 104 according to some examples. The mobile robot 104 includes a robotics open platform 302, an operating stack 304, and a robotics controller 330.

[0016] The robotics open platform 302 provides various application program interfaces (APIs), and such application program interfaces (APIs) include ● Device API 306 ●Diagnostic API 308 ●Robotics API310 ●Data API 312 and ● Swarm API 314 Includes.

[0017] Operation stack 304 is, ●Recognition 316 ● Peer-to-peer (P2P) service 318 ● Semantic operation 320 ● Sensor correction 322 ●Sensor processing 324 ● Obstacle Avoidance 326 Includes components that support it.

[0018] The ROS operation stack 328 also forms part of the operation stack 304.

[0019] The robotics controller 330 is ●Power management 332 ●Wireless charging 334 ●Device Interface 336 ●Motor control 338 Includes components that support it.

[0020] Figure 4 is a block diagram illustrating other aspects of the components and modules of a mobile robot 104 relating to several examples. The mobile robot 104 includes a robotics stack 402 and an application stack 406. The robotics stack 402 further includes a recognition stack 404 and a navigation stack 304. The application stack 406 provides telemetry 408 and login 410 services for the mobile robot 104.

[0021] Figure 5 shows the output of the image-based positioning model 604 for several examples. The generation and maintenance of the image-based positioning model will be discussed further in Figures 6 and 7 below.

[0022] Image-based positioning models (or image positioning models) can be used for mobile robots (on which the model is deployed) to navigate areas or environments such as service locations. An example of an image-based positioning model can take an image (e.g., an image captured in a specific environment or location) as input and output a set of potential poses or pose estimates. Poses can include candidate locations (e.g., grid cells of a configurable size corresponding to a map of a traversable area). Poses can include a set of (x, y, yaw) values, where x and y are map or world coordinates and yaw is the yaw angle (corresponding to orientation information). In some examples, poses can include pitch and roll values ​​(corresponding to orientation information). In some examples, the positioning model can output a set of potential locations such as grid cells and predicted x-offset and y-offset values ​​within each grid cell, along with one or more of the yaw, pitch, and roll values. Each type of predicted value (e.g., grid cell, x-offset, and y-offset) may be accompanied by a confidence score or probability value representing the model's confidence in the prediction.

[0023] Figure 5 shows an exemplary output of an image-based positioning model 604 that takes an image (e.g., an image from the front camera of the mobile robot 104) as input and generates a probability distribution for a set of potential locations represented by grid cells in a grid corresponding to a map of the traversable area. In this example, the model returns a list of (grid cell, probability) pairs such as [((e, 3), 0.5);((b, 3), 0.2);((b, 4), 0.1);((d, 2), 0.1);((d, 3), 0.1)]. Thus, in this example, the model can accurately identify the grid cell corresponding to the location where the mobile robot 104 is most likely to be in a given environment.

[0024] Figure 6 is a block diagram showing a model system 602 that operates to generate and maintain image-based positioning models 604 to be deployed on various mobile robots 104 at one or more locations 102, according to several examples. The model system 602 may include the following components or modules: ●Data acquisition and preparation module 606; ●Model training and evaluation module 608; ●Model placement module 610; and ●Model Refresh Module 612

[0025] Further details regarding the operation of such example modules are provided below.

[0026] Figure 7 is a flowchart of the model workflow 702 for generating, maintaining, and deploying an image-based positioning model in several examples. The model workflow 702 can be performed by the model system 602 shown in Figure 6 within the context of the environment shown in Figure 1. The workflow 702 begins with data collection (e.g., map metadata, images, pause data, or pause messages, as illustrated in more detail in Figure 8) by the data collection and preparation module 606 at a location 102, such as a cafeteria. The data collection and preparation module 606 employs a data collection strategy that can use sensors on a mobile robot 104 located at location 102 for data aggregation (see below for more details on the data collection strategy). In some examples, the data collection strategy may also include the use of sensors installed at location 102 or movably positioned within location 102 (e.g., manually captured by user-carried sensors) separately from the mobile robot 104. A sufficient amount of data for location 102 is collected and stored in a cloud storage (e.g., as part of cloud service 110). The sufficiency of the collected data volume can be objectively determined based on a number of criteria, including location area coverage ratio, volume-based criteria, or data resolution criteria. Data can be stored in a variety of formats (for example, images may be stored in formats such as JPG, GIF, and PNG, while map metadata and / or image or pause messages may be stored in formats such as YAML, XML, and JSON).

[0027] Subsequently, the collected data is preprocessed (or prepared) by converting the collected raw data into a dataset, which is a form of prepared data that can be consumed by the model training and evaluation module 608. Data preprocessing or preparation involves subsampling and splitting the data into training / development / test sets (or training / validation / test sets), as further explained in the descriptions of Figures 8 and 10. The preprocessed data is stored (for example, in a cloud storage) using an appropriate format (e.g., CSV, Parquet, etc.). Subsequently, the image-based positioning model 604 is trained using the preprocessed data and cloud machine learning capabilities (see the "Model Training and Evaluation" section below). The trained model is stored / storaged or selectively compressed in an appropriate format in a remote storage location such as a cloud storage (e.g., part of cloud service 110) (for example, the "model.tar.gz" file corresponds to a model stored in a compressed .tar file or tarball, and compression is performed by the gzip utility). Other exemplary compression utilities include bzip2, zstd, etc.

[0028] Subsequently, the trained image-based positioning model 604 is deployed to the mobile robot 104 at a given location 102. In some examples, the trained image-based positioning model is used by an image localizer, which is a Robot Operating System (ROS) node that provides image-based positioning responses to client requests (see Figure 12 for details).

[0029] The following subsections provide additional details for steps 702 of the workflow.

[0030] Data collection strategy

[0031] The data acquisition and preparation module 606 employs a data acquisition strategy that utilizes the sensors of the mobile robot 104 within location 102 to collect data including map metadata, images (e.g., image messages), poses (e.g., pose messages including pose estimates), etc. In some examples, the image and pose messages may be Robot Operating System (ROS) messages, and the pose messages may include all position and orientation data.

[0032] The service mobile robot 104 can use a probabilistic positioning method to calculate a pose estimate for the service robot in the context of a pre-calculated map of the service environment. For example, the service mobile robot 104 can use an Adaptive Monte Carlo Localization (AMCL) method with automatically acquired laser scanning data (e.g., as embodied by the AMCL node in the ROS operation stack, as shown in Figure 7). In some examples, the service mobile robot 104 can use other positioning algorithms (e.g., General Monte Carlo Localization (GMCL)) and identical or additional types of automatically acquired service environment data (e.g., mileage data, sonar data).

[0033] Image and pause messages may also be further associated with a timestamp identifying the day (e.g., date) and time information when the payload information was captured or generated, and a mission identifier identifying the specific mission or movement in which the payload information was captured. The timestamp and mission identifier information can be used, for example, to evaluate, verify, and visualize the data acquisition process by the mobile robot 104. Sensors used for data acquisition include those mentioned above in relation to the mobile robot 104, and include, for example, numerous RGB cameras, LiDAR sensors, and radar sensors. While the primary example describes data acquisition by a mobile robot, alternative or additional acquisition processes may include fixed sensors such as cameras, LiDARs, or radars placed at different points in the environment. Such sensors can capture information over time to generate a more complete picture of the environment.

[0034] In some examples, the data collection strategy employed by the data collection and preparation module 606 at location 102 may be intentional or passive.

[0035] Intentional data acquisition can refer to the process of sending a mobile robot 104 to perform a specific task: to collect data (e.g., image messages, pause messages, map metadata, etc.) using a variety of sensors. The advantage of intentional data acquisition is that it can provide adequate (e.g., more comprehensive and uniform) data coverage for many accessible points on a map.

[0036] Passive data collection can refer to the process by which a mobile robot 104 samples data and uploads it to a cloud service 110 while performing normal delivery tasks (e.g., at a configurable speed). In this case, data collection can be turned on or off via a config flag in some examples. Passive data collection does not require the movement of any additional service robots at a location and therefore does not interfere with the regular operation of the cafeteria / client.

[0037] The examples can support both passive and intentional data collection. However, for explanatory purposes, passive data collection will be discussed later in relation to the example mobile robot 104.

[0038] Data preservation policy

[0039] When data collection is generally activated, the data collection and preparation module 606 accumulates a large amount of data. Insufficient data can cause bottlenecks in many machine learning applications, but excessively large amounts of data are costly, difficult to handle, and sometimes unnecessary. To address this, a data organization strategy is implemented in the data collection and preparation module 606, which prevents the amount of data from becoming arbitrarily large while ensuring that there is enough data to train new models.

[0040] Exemplary data preservation strategies or policies are age-based and volume-based.

[0041] Age-based data retention policy: This exemplary policy retains data for a specified period, such as one month or six months, after which it is deleted or stored. For example, a data retention policy could specify that data older than 30 days should be deleted from the system.

[0042] Volume-based data storage policy: This exemplary policy stores only a specified amount of data, expressed by the number of images, the amount of data, or both (e.g., 5,000 to 10,000 images or 500MB to 1GB of data).

[0043] Since mobile robots 104 are used at varying frequencies in different locations 102 (for example, some cafeterias may be busier than others), volume-based preservation policies and related preservation work can be implemented to ensure that a uniform amount of data is preserved across multiple locations 102.

[0044] Other exemplary conservation strategies or policies that may be used include:

[0045] Rolling window retention policy: In this exemplary policy, the retention period is not fixed; instead, data is retained for a specific period, such as the past 30 or 90 days. At the end of each window, the oldest data is deleted or stored to make space for new data. For example, a rolling window retention policy can specify that only data from the past 30 days be retained at any given time. This ensures that only the most recent data is retained and is particularly useful for applications where data changes continuously.

[0046] Quality-Based Data Retention Policy: This exemplary policy relies on data quality rather than volume for data retention. For example, data with high accuracy, completeness, or relevance is prioritized for retention, while lower-quality data can be discarded. This policy can be implemented when the quality of collected data is variable, and data retention policies can be adjusted according to data quality metrics.

[0047] Hybrid preservation policies: Under such exemplary policies, volume, age, and / or quality considerations are taken into account when deciding which data to preserve. For example, a certain amount of data may be preserved, but high-quality data may be given priority over low-quality data. Such a hybrid approach can strike a balance between having enough data for model training and ensuring high data quality.

[0048] Adaptive preservation policies: Under such exemplary policies, preservation can be automatically adjusted based on changes in data volume, age, quality, and / or other relevant factors. For example, if data volume exceeds a certain critical value, the preservation policy can be adjusted to reduce the amount of data preserved. Similarly, if data quality deteriorates, the preservation policy can be adjusted to preserve higher quality data. Such an approach can help ensure that data preservation policies remain effective over time, even as conditions change.

[0049] Data preparation

[0050] Figure 8 is a flowchart illustrating a data preparation flow 802 for several examples, which may be performed by the data collection and preparation module 606 to generate data stored in the cloud storage 804. The data preparation or preprocessing in the data preparation flow 802 includes the task of converting the collected raw data into a dataset, which is in the form of prepared data 820 that can be directly consumed by the model training and evaluation module 608 in several examples.

[0051] The collected raw data includes map metadata 818, pause data, or pause messages 806 and images (or image messages) 808. As illustrated in Figure 8, since images 808 and pause messages 806 can be collected independently, the matching operation 810 matches images 808 with their corresponding pause messages 806 to generate matching data 814 (for example, based on their timestamps). Images can be matched with pause messages if the individual timestamps are identical, or if the difference between individual capture or collection / generation times, as calculated based on the timestamps, is less than a predetermined time difference (e.g., a maximum of 0.2 or 0.3 seconds).

[0052] In assignment task 812, the preprocessing logic of the data acquisition and preparation module 606 divides the map of location 102 (e.g., represented in map metadata 818) into a grid of cells (see Figure 9 for an example of map 902). Each robot location (e.g., position information or coordinates captured by pause message 806) is assigned to a specific grid cell (of a configurable size) on the map. The output of assignment task 812 is stored as grid assignment data 816. Matching task 810 and assignment task 812 may be performed at the level of a robot task and, in some examples, may be performed in batches (e.g., daily).

[0053] When raw data (e.g., pause messages 806 and images 808) is collected (e.g., passively) in a natural working environment, a particular area may be traversed by the robot more frequently than other areas, and therefore that particular location may be represented more prominently in the collected dataset than other locations (see Figure 9). For example, in a cafeteria environment, a narrow corridor between the kitchen and dining area may be included in the frequently traversed areas. To ensure adequate coverage of the traversed environment, the preprocessing logic of the data acquisition and preparation module 606 may subsample (or resample) the raw data to readjust the data distribution so that the data distribution is more uniform across the traversable space, such as location 102. In some examples, given a set of pairs (e.g., image, pause) augmented with pause-level grid cell assignment information, the data acquisition and preparation module 606 samples individually from each subset corresponding to a particular grid cell (e.g., more frequently occurring grid cells may be subsampled).

[0054] The splitting operation 822 splits or divides the grid-assigned data 816 for training purposes. For example, the grid-assigned data 816 may be split into a training / validation set 824 (or training / development set 824) and a test set 826 (or test / evaluation set 826). As previously mentioned, the data splitting logic of the data acquisition and preparation module 606 (e.g., in splitting operation 822) may take grid cell assignments into consideration so that the observed locations appear in a balanced manner in the training dataset (e.g., prepared data 820). Splitting operation 822 further pursues ensuring that the training / validation set 824 and the test set 826 are not excessively similar (e.g., sufficiently independent), as will be discussed further in relation to Figure 10.

[0055] In some examples, data augmentation techniques can be used to enhance data used for training purposes. For example, photometric augmentation techniques (e.g., adjusting the brightness, hue, or contrast of images in a dataset to obtain additional examples corresponding to a given grid cell / pause message) can be used.

[0056] Figure 9 shows a random sample of raw data corresponding to several trips collected by the mobile robot 104, in the context of a map divided into grid cells of configurable size. The percentage labels displayed with the subset of cells indicate the percentage of data points corresponding to each particular cell. Figure 9 shows that the raw data is not uniformly distributed; for example, the four cells with the highest data point ratios cumulatively account for 48% of the collected data in the random sample. Such a non-uniform distribution of data by the mobile robot 104 may occur at location 102. To ensure that the observed locations appear in a balanced manner in the training dataset (e.g., prepared data 820), the preprocessing logic of the data acquisition and preparation module 606 (e.g., in the splitting operation 822) can assign grid cells as described above.

[0057] Temporal data joining

[0058] Figure 10 illustrates several exemplary block-based data partitioning processes 1002 that can be performed by the data acquisition and preparation module 606 in partitioning operation 822.

[0059] A challenge in developing machine learning (ML) models is to eliminate or minimize systematic dependencies between the data used to train the model and the data used to evaluate the model. In the case of a sequence of images (or pose pairs to which images are matched) collected by a mobile robot (e.g., mobile robot 104), there can be strong dependencies in the form of temporal coupling between images 808 collected at different time intervals (e.g., any given image may often appear similar to an image observed around 1 second earlier or later). It is desirable to resolve such temporal couplings to avoid distorted evaluation results.

[0060] When data is abundant, an exemplary strategy to eliminate or reduce such temporal coupling between training / validation (or training / development) data and test / evaluation data may be time-based data partitioning (e.g., date-based time partitioning). Time-based data partitioning, such as date-based data partitioning, refers to using data collected during different missions or different capture periods for training / validation and test / evaluation datasets. Missions are indicated by mission identifiers associated with the payload (e.g., images, pause messages, etc.). The capture period (e.g., seconds / minutes / hours / days / weeks / months of capture) may be calculated by the data collection and preparation module 606 based on timestamps associated with the payload (e.g., images, pause messages, etc.). For example, the data collection and preparation module 606 may partition the captured data (e.g., images, matched image-pause data, etc.) so that data captured from a particular period is used in the training / validation set but not in the test / evaluation set (or vice versa).

[0061] In addition to such absolute capture periods calculated based on payload timestamps, the data acquisition and preparation module 606 can use relative capture periods (e.g., 30 minutes to 1 month) to embody time-based data partitioning. Such relative capture periods can capture the time difference between the first capture time of a selected first image (or a separately selected first capture time) and the second capture times of subsequent images. For example, the data acquisition and preparation module 606 can ensure that images in a training / validation set are captured within a relative period of 1 hour with respect to a predetermined capture time (e.g., the capture time of the earliest captured image included in the training set).

[0062] However, in some examples, time-based data partitioning can be difficult to implement, for instance, when 1,000 robots operate in diverse environments, because the amount of data collected in one location for a day may be much smaller or larger than in some other locations. Therefore, time-based data partitioning may not be optimal in some use cases.

[0063] Further exemplary strategies to prevent temporal coupling could be random shuffling (e.g., random rearrangement) of collected payload data such as images (or images and matching pause messages) and their assignment to training / validation sets 824 and test sets 826. However, random shuffling, in some situations and when used alone, can still result in a large number of identical or similar images that are unacceptable in training / validation and test / evaluation datasets.

[0064] For example, in the case of sparser data, yet another exemplary strategy for dealing with temporal coupling is to divide the stream of images into continuous blocks of images (e.g., matched with corresponding pause messages), where each block has a fixed and configurable size. Finally, dealing with temporal coupling can be accomplished by combining such exemplary strategies serially or in other ways.

[0065] Figure 10 shows an exemplary block-based data partitioning process 1002. First, the collected data stream is partitioned into blocks 1004 of the same size, consisting of continuously captured images (e.g., image-pose pairs). In some examples, the images are not captured continuously but may be captured within a determinable period or interval. Thereafter, the blocks 1004 are randomly shuffled and assigned to the training / validation set 824 (or training / development set 824) and the test set 826.

[0066] If the block size is sufficiently large and two consecutive blocks are assigned to the training / validation set 824 and the test set 826, temporal coupling may be reduced. Larger blocks improve temporal disconnection but can worsen the imbalance in data distribution with respect to location (as reflected in the distribution across grid cells corresponding to map 902). For example, if a robot is sent to one destination for the first 10 tasks and to another destination for the next 10 tasks, and the collected data is divided into two blocks of the same size (the block size corresponds to the number of tasks), then the first block will not contain data collected from the second destination, and vice versa. Therefore, in some examples, the block size may be empirically selected by varying the block size so that the data distribution is sufficiently balanced with respect to observed locations (as reflected in the distribution across grid cells of the grid corresponding to the exemplary map 902, for example), and then selecting the largest block size.

[0067] Finally, in some examples, the data collection and preparation module 606 may prepare a separate dataset for evaluation or testing purposes, where the separate dataset is completely independent of the data used during training. In some examples, such a separate dataset may include images and matching pose messages from different (e.g., more recent) times and / or dates. Such a dataset may not contain every location that appears in the training dataset, but it may be more up-to-date, accurate, and independent, and therefore provide a more useful evaluation metric.

[0068] Data Quality Considerations

[0069] Generally, supervised machine learning models depend on the quality of the data used for training (the rule "put in garbage and get garbage out" applies). In some examples, model system 602 can use uncurated data for model training, so data quality issues can be addressed in various ways.

[0070] Firstly, unless the robot gets lost in a certain regularity, we can assume that the robot's pose (or the generated pose estimate) is reasonably accurate in most cases.

[0071] Secondly, if there are particularly difficult locations in the cafeteria where the pause estimate is likely to consistently make errors, such unreliable data points can be filtered using the pause covariance values ​​reported by a probabilistic positioning method (e.g., AMCL).

[0072] Finally, since positioning methods like AMCL provide pose estimates using "ground truth," these pose estimates will inevitably contain some noise. However, the presence of such noise is mitigated by the ML model training settings. Furthermore, some noise in the training data can actually make the model training more robust, as it reduces the likelihood of overfitting by making it less likely for the model to remember location values ​​corresponding to specific images in the training set.

[0073] Data from numerous robot configurations

[0074] In some examples, such as a multi-robot setup, data may be collected from multiple robots operating in a single location. A particular robot may not cover all parts of the site (for example, one robot might be sent to only one part of a building, while others spend most of the time in other parts). Therefore, in some examples, the model system 602 can combine data from different robots and use this for training and / or evaluation.

[0075] Data collected by and received from multiple robots can be combined in some examples to generate a single training / validation set 824. A test / evaluation set 826 can be generated using aggregated data from multiple robots, or in some examples using separate data sourced from a single robot.

[0076] Maintaining data from a single robot exclusively to generate the test / evaluation set 826 can offer several advantages, such as a better understanding of whether any hardware-related differences in the data negatively impact the model's performance. For example, this can be considered if a robot needs to be replaced in a cafeteria, but its model was trained on data collected before the replacement.

[0077] Model training and evaluation

[0078] Figure 11 is a flowchart illustrating the image-based positioning model development process 1102 for several examples, which can be embodied by the model training and evaluation module 608. For simplicity, this diagram illustrates the flow of model training and evaluation for a model for one location, but numerous models 604 for numerous locations are generated and / or updated at various times (see the "model refresh" section below).

[0079] Once the collected data has been preprocessed by the data acquisition and preparation module 606 and divided into training / validation sets 824 and test / evaluation sets 826, the model training and evaluation module 608 is deployed to train the model 604.

[0080] In the model training task 1104, the image-based positioning model 604 is trained using the training / validation set 824. The image-based positioning model 604 can use various model architectures and any combination thereof in various training scenarios. For example, the model training and evaluation module 608 can use a transfer learning technique, but the training module starts with a backbone neural network (NN) pre-trained on a large public dataset (e.g., ImageNet), replaces the final K layers of the backbone neural network (where K is a predetermined constant, e.g., 1 or 2) with a set of customized layers, and trains the customized layers using the training / validation set 824 specified for the task. Thereafter, the trained customized layers are used to generate location predictions for input images. In some examples, the backbone neural network could be a Vision Transformer or EfficientNet. In some examples, a set of custom layers may include one or more fully concatenated layers along with a softmax layer as the output layer, where the size of the softmax layer corresponds to the number of different grid cell IDs for the grid cells of a grid map at a particular location.

[0081] Subsequently, the trained model is evaluated using the test set 826 in the model evaluation task 1106. The evaluation results are reviewed by the evaluation-related decision node 1108, which outputs an index (such as a binary value, actual confidence score, or satisfaction score) indicating whether the model's performance (for example, relative to the test set 826) is considered satisfactory with respect to one or more evaluation metrics (e.g., accuracy, precision / recall, AUC, etc.).

[0082] If the model evaluation is satisfactory, model 604 may be deployed as a production model on the mobile service robot 104. If the model performance is unsatisfactory, update operation 1110 may allow the model training and evaluation module 608 to retrain the model by augmenting or replacing the training / validation set 824.

[0083] Model Placement

[0084] Figure 12 is a flowchart illustrating several example image-based positioning model placement processes 1202 that can be implemented by the model placement module 610.

[0085] As discussed in the following section (see Model Refresh), models 604 for different locations will be generated / updated at different times. Also, deployed robots (e.g., mobile robot 104) are not necessarily restarted regularly, and in the case of many robots, not all robots are updated at once. Therefore, (although model training can be done at the single-location level) in some examples, model deployment is done at the robot instance level. Specifically, the model system 602 can be trained / prepared in the cloud (e.g., stored in cloud storage 804 using cloud service 110) regardless of robot activity. To achieve asynchronous model deployment, a particular robot can check for the existence of a new model upon restart. Figure 12 illustrates an example of such a process.

[0086] The image localizer 1204 is a Robot Operating System (ROS) node that provides image-based positioning responses to client requests using image-based positioning models loaded into the mobile robot 104's memory. The image localizer 1204 is responsible for loading the latest available image-based positioning models from the mobile robot 104's disk 1208 into the mobile robot's memory (see below).

[0087] The model fetcher 1206 is an independent ROS node that checks (for example, through decision task 1214 during reboot 1212) whether the latest model 604 (e.g., model XYZ) for the current location (e.g., location XYZ) in the cloud storage 804 matches the current model stored on the robot disk 1208. If the latest model 604 for the current location does not match the model stored on the disk, the model fetcher 1206 downloads the latest model for the current location from the cloud storage in download task 1210 and updates model 604 stored on disk 1208. Once model 604 is available on the mobile robot 104's local robot disk (as shown in task 1216), the image localizer 1204 can upload the available model to the mobile robot 104's memory.

[0088] In some examples, ModelFetcher 1206 can use other policies to trigger a check to see if a new model for the current location can be downloaded from Cloud Storage 804. For example, ModelFetcher 1206 could periodically check for new location-specific models (e.g., based on a set time schedule). Alternatively, ModelFetcher 1206 could use a conservative policy to check for the availability of a new location-specific model only if no new model has been downloaded within a predetermined period. Finally, ModelFetcher can use such policies in combination with other policies.

[0089] Model refresh

[0090] Figure 13 is a flowchart illustrating several example image-based positioning model refresh processes 1302 that can be implemented by the model refresh module 612.

[0091] The layout of location 102 (for example, a dining room) can be changed and evolved over time. In some examples, the changes can be seasonal (e.g., holiday decorations or more outdoor seating in warmer months) and permanent (e.g., interior renovations). In either case, such changes will affect the performance of model 604, which is trained on previous image data. Therefore, model system 602 provides a function to automatically update model 604 to prevent the model from becoming outdated and its performance from degrading.

[0092] For example, model refreshes using the Model Refresh Module 612 can be performed proactively based on a schedule (e.g., monthly / weekly) or reactively based on metrics (e.g., when the repositioning convergence rate falls below a critical value). Proactive model refreshes attempt to maintain nearly identical up-to-dateness of models at all locations, but in some cases, this can be wasteful (e.g., if some locations change their layout / decoration more frequently than others). Reactive model refreshes may involve more engineering / design work (e.g., the need to collect the correct metrics that accurately reflect model performance), but they can be less resource-intensive (e.g., there is no need to update models at locations with a somewhat static layout).

[0093] The model refresh process 1302 illustrated in Figure 13 is a lifecycle of reactive model refresh related to several examples. The operation stack 304 (e.g., part of the robotics stack 402 of a mobile robot 104) reports one or more online model metrics 1304 of the model performance of an image-based positioning model 604 used by one or more mobile robots 104. As previously mentioned, the online model metrics may include the repositioning convergence rate (e.g., the rate at which a probabilistic positioning method such as AMCL converges after repositioning of the mobile robot 104). The model metrics 1304 are stored in a database (e.g., a cloud storage 804). When the performance of the image-based positioning model (e.g., determined by the model metrics 1304) deteriorates, the model 604 is automatically retrained (e.g., in task 1306) to generate a new production model 604. The model retraining task uses the most recently updated training and test data available. The evaluation scale may include the scales used for initial model training and subsequent evaluation (e.g., accuracy, precision, recall, etc.).

[0094] In addition to retraining Model 604 using online metrics (e.g., triggering model retraining based on monitoring metrics reported by the operation stack 304), in some examples Model 604 can be evaluated offline (in task 1308) using the most recently collected data uploaded by the mobile robot 104's data uploader 1310. The collected data is prepared through the data preparation flow 802, which generates new training / validation and test / evaluation sets. The offline evaluation of Model 604's performance (task 1308) can generate offline metric values ​​using a newly generated test set containing recently collected data. Such offline metrics can be used by the model refresh module 612 to trigger retraining of Model 604. Offline metrics may include metrics used in the initial model training / evaluation (e.g., accuracy / precision / recall to random test samples). As described above, the model retraining task can use the most recently collected data from the operational environment, ensuring that the retrained model does not become outdated and is better adapted to the current environment configuration and characteristics.

[0095] Model Version Management

[0096] Figure 14 is a flowchart illustrating several example model version control processes 1402 that can be implemented by the model system 602.

[0097] During the model improvement process, it may become necessary to change the model architecture and / or serving code. Model system 602 periodically refreshes models (not necessarily requiring any changes to the model serving code), and in doing so, can support numerous different model types in addition to supporting multiple versions of the same model. However, simultaneously updating models for numerous robots presents numerous technical challenges in terms of feasibility and reasonableness.

[0098] If a new type of model 604 is generated in the cloud (for example, using the cloud service 110), the new type of model 604 does not need to be used by the mobile robot 104 until a new version of robotics software (for example, one that embodies support for the model type) is deployed. Furthermore, even if the mobile robot 104 is equipped with a new version of robot software, it does not need to be able to use the new model type at that location.

[0099] Therefore, a mechanism is provided for updating model types and / or model versions, which includes support for other different models deployed in different locations 102 and mobile robots 104. For a single robot, new model deployment can be managed manually. However, when there are many robots (e.g., 1000) with many different model types, automating this process offers technical and operational advantages.

[0100] Some examples handle asynchronous updates through pointers to different locations. For example, in some examples, the following path structure could be used to save a model to the cloud.

[0101]

number

[0102] For example, the listed paths indicate that in location 1, versions vl, v2, and v3 of model 1 can be used (for example, in location 1, version v1 of model 1 is stored in the v1 / subdirectory of the location1 / directory). The listed paths may also indicate that, for example, at location 1, at least two types of models exist, namely model 2 and model 1.

[0103] On the robot side, the updated code can include fallback logic that allows the robot to handle the existence of multiple model types and / or multiple model versions of multiple model types for a given location. An example of such fallback logic is shown below.

[0104]

number

[0105] For example, given a model for a specific location, the fallback logic checks if a subdirectory corresponding to the latest version (e.g., v3) exists in the directory corresponding to that location. If it does, the v3 version of the corresponding model is used as the current model version. If no such path exists, the fallback logic checks if subdirectories corresponding to other model versions (e.g., v2 or vl) exist to identify the corresponding model version (e.g., v2 or vl).

[0106] Figure 14 illustrates locations 1, 2, and 3, where different versions of location-related models are stored for different locations (e.g., in cloud storage 804) (e.g., using some of the storage routes listed above). Location 1 has model versions 1, 2, and 3; location 2 has model versions 1 and 2; and location 3 has model versions 1 and 3. By examining the directory structure (e.g., directory file structure) and storage routes described above, an exemplary image localizer node for a mobile robot in location 1, 2, or 3 can retrieve the latest version of a particular type of location-specific model (e.g., through a download operation 1210 not shown). For example, an image localizer 1204 for a mobile robot 104 in location 1 can retrieve model version 3 for location 1. On the other hand, an image localizer 1204 for a mobile robot in location 2 can retrieve model version 2 for location 2, and an image localizer 1204 for a mobile robot in location 3 can retrieve model version 3 for location 3.

[0107] Figure 15 is a block diagram 1500 illustrating a software architecture 1504 that may be installed in any one or more of the devices described herein in several examples. The software architecture 1504 is supported by hardware such as a machine 1502, which includes a processor 1520, memory 1526, and / or I / O components 1538. In this example, the software architecture 1504 may be conceptualized as a stack of layers, each layer providing a specific function in several examples. The software architecture 1504 includes layers such as an operational structure 1512, a library 1510, a framework 1508, and an application 1506. Operationally, the application 1506 invokes an API call 1550 through the software stack and receives a message 1552 in response to the API call 1550.

[0108] The operational structure 1512 manages hardware resources and provides common services. The operational structure 1512 includes, for example, the kernel 1514, services 1516, and drivers 1522. The kernel 1514 acts as an abstraction layer between hardware and other software layers. For example, the kernel 1514 provides memory management, processor management (e.g., scheduling), component management, networking, and security configuration, among other functions. Services 1516 can provide other common services for other software layers. Drivers 1522 are responsible for controlling or interfacing the basic hardware. For example, drivers 1522 may include display drivers, camera drivers, BLUETOOTH® or BLUETOOTH® Low Energy drivers, flash memory drivers, serial communication drivers (e.g., General Purpose Serial Bus (USB) drivers), Wi-Fi® drivers, audio drivers, and power management drivers.

[0109] Library 1510 provides a low-level common infrastructure used by application 1506. Library 1510 may include system libraries 1518 (for example, the C standard library) that provide functions such as memory allocation functions, string manipulation functions, and mathematical functions. Furthermore, Library 1510 may include API libraries 1524 such as a media library (e.g., a library to assist in the display and manipulation of various media formats such as MPEG4 (Moving Picture Experts Group-4), AVC (Advanced Video Coding) or H.264, MP3 (Moving Picture Experts Group Layer-3), AAC (Advanced Audio Coding), AMR (Adaptive Multi-Rate) audio codecs, JPEG (Joint Photographic Experts Group) or JPG, or PNG (Portable Network Graphics)), a graphics library (e.g., the OpenGL framework used to render graphic content on a display in two dimensions (2D) and three dimensions (3D)), a database library (e.g., SQLite to provide various relational database functions), and a web library (e.g., Web Kit to provide web browsing functions). Library 1510 may also include various other libraries 1528 to provide many other APIs to Application 1506.

[0110] Framework 1508 provides a high-level common infrastructure used by Application 1506. For example, Framework 1508 provides diverse graphical user interface (GUI) functions, high-level resource management, and high-level location services. Framework 1508 can provide a wide range of other APIs that can be used by Application 1506 in several examples, some of which may be limited to specific operational structures or platforms.

[0111] In some examples, application 1506 may include a wide range of other applications such as a home application 1536, a contacts application 1530, a browser application 1532, an e-book reader application 1534, a location application 1542, a media application 1544, a messaging application 1546, a game application 1548, and a third-party application 1540. Application 1406 is a program that performs programmatically defined functions. In some examples, a variety of programming languages ​​structured in various ways, such as object-oriented programming languages ​​(e.g., Objective-C, Java, or C++) or procedural programming languages ​​(e.g., C or assembly language), may be used to generate one or more of applications 1506. In certain examples, a third-party application 1540 (e.g., an application developed by an entity other than the vendor of a particular platform using the ANDROID™ or IOS™ Software Development Kit (SDK)) may be mobile software that runs on a mobile operating system such as IOS™, ANDROID™, WINDOWS® Phone, or other mobile operating systems. In this example, the third-party application 1540 can invoke the API call 1550 provided by the operational structure 1512 to implement the functions described herein.

[0112] Figure 16 is a schematic representation of machine 1600 in which instructions 1610 (e.g., software, programs, applications, applets, apps, or other executable code) can be executed to cause machine 1600 to perform any one or more of the methodologies discussed herein, according to several examples. For example, instructions 1610 can cause machine 1600 to perform any one or more of the methods described herein. Instructions 1610 can transform an unprogrammed general machine 1600 into a specific machine 1600 programmed to perform the functions described and illustrated in the described manner. Machine 1600 can operate as a standalone device or be coupled to other machines (e.g., connected to a network). In a network configuration, machine 1600 can operate as a server or client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. Machine 1600 may include (but is not limited to) any machine capable of sequentially or differently executing instruction words 1610 that specify the actions performed by Machine 1600, such as a server computer, client computer, personal computer (PC), tablet computer, notebook computer, netbook, set-top box (STB), entertainment media system, cellular phone, smartphone, mobile device, wearable device (e.g., smartwatch), smart home device (e.g., smart appliance), other smart device, web appliance, network router, network switch, network bridge, or any other machine. Although a single Machine 1600 is illustrated, the term “machine” may also include a collection of machines that individually or collectively execute instruction words 1610 to perform any one or more of the methodologies discussed herein.

[0113] Machine 1600 may include a processor 1604, memory 1606, and I / O components 1602, which may be configured to communicate via bus 1640. In some examples, processor 1604 (e.g., a CPU (Central Processing Unit), a RISC (Reduced Instruction Set Computing) processor, a CISC (Complex Instruction Set Computing) processor, a GPU (Graphics Processing Unit), a DSP (Digital Signal Processor), an ASIC (Application-Specific Integrated Circuit), an RFIC (Radio-Frequency Integrated Circuit), other processors, or any appropriate combination thereof) may include processors 1608 and 1612 that execute, for example, instruction 1610. The term “processor” is intended to include multicore processors, which may include two or more independent processors (sometimes referred to as “cores”) capable of executing instruction words simultaneously. Figure 16 illustrates a plurality of processors 1604, but machine 1600 may include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), a plurality of processors with a single core, a plurality of processors with multiple cores, or any combination thereof.

[0114] Memory 1606 includes main memory 1614, static memory 1616, and storage unit 1618, all of which are accessible to processor 1604 via bus 1640. Main memory 1606, static memory 1616, and storage unit 1618 store instruction words 1610 that embody any one or more of the methodologies or functions described herein. Instruction words 1610 may also reside, in whole or in part, in main memory 1614, in static memory 1616, in machine-readable medium 1620 in storage unit 1618, in processor 1604 (e.g., in the processor's cache memory), or in an appropriate combination thereof, while being executed by machine 1600.

[0115] The I / O component 1602 can include a variety of components for input reception, output provision, output generation, information transmission, information exchange, or measurement capture. The specific I / O component 1602 included in a particular machine will vary depending on the type of machine. For example, a portable machine such as a mobile phone may include a touch input device or other such input mechanism, while a headless server machine is unlikely to include such a touch input device. The I / O component 1602 can include many other components not shown in Figure 16. In various examples, the I / O component 1602 can include an output component 1626 and an input component 1628. The output component 1626 can include a visual component (e.g., a display such as a PDP (Plasma Display Panel), LED (Light-Emitting Diode) display, LCD (Liquid Crystal Display), projector, or CRT (Cathode Ray Tube)), an acoustic component (e.g., a speaker), a tactile component (e.g., a vibration motor, resistance mechanism), or other signal generators. The input component 1628 may include alphanumeric input components (e.g., keyboard, touchscreen configured to receive alphanumeric input, photo-optical keyboard, or other alphanumeric input components), point-based input components (e.g., mouse, touchpad, trackball, joystick, motion sensor, or other pointing device), haptic input components (e.g., physical buttons, touchscreen providing position and / or force for touch or touch gestures, or other haptic input components), audio input components (e.g., microphone), etc.

[0116] In additional examples, the I / O component 1602 may include a variety of other components, such as a biometric component 1630, a motion component 1632, an environmental component 1634, or a position component 1636. For example, the biometric component 1630 includes components that detect facial expressions (e.g., hand expressions, facial expressions, vocal expressions, body movements, or eye tracking), measure biosignals (e.g., blood pressure, heart rate, body temperature, sweat, or electroencephalogram), or identify people (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or electroencephalogram-based recognition). The motion component 1632 includes acceleration sensor components (e.g., accelerometers), gravity sensor components, and rotation sensor components (e.g., gyroscopes). The environmental components 1634 include, for example, one or more cameras, illuminance sensor components (e.g., photometers), temperature sensor components (e.g., one or more thermometers for detecting ambient temperature), humidity sensor components, pressure sensor components (e.g., barometers), acoustic sensor components (e.g., one or more microphones for detecting background noise), proximity sensor components (e.g., infrared sensors for detecting surrounding objects), gas sensors (e.g., gas detection sensors for detecting the concentration of harmful gases for safety or measuring airborne pollutants), or other components that can provide indicators, measurements, or signals corresponding to the surrounding physical environment. The location components 1636 include, for example, location sensor components (e.g., GPS receiver components), altitude sensor components (e.g., altimeters or barometers for detecting atmospheric pressure from which altitude can be derived), orientation sensor components (e.g., magnetometers), etc.

[0117] Communication can be implemented using a variety of technologies. The I / O component 1602 further includes a communication component 1638 that can be operated to connect machine 1600 to network 1622 or device 1624 through individual couplings or connections. For example, the communication component 1638 may include a network interface component or other suitable device for interfacing with network 1622. In additional examples, the communication component 1638 may include a wired communication component, a wireless communication component, a cellular communication component, an NFC (Near Field Communication) component, a BLUETOOTH® component (e.g., BLUETOOTH® Low Energy), a flash memory driver, a serial communication driver (e.g., a general-purpose serial bus (USB) driver), a WI-FI® component, and other communication components for providing communication in other forms. Device 1624 may be any other machine or a variety of peripheral devices (e.g., peripheral devices coupled via USB).

[0118] Furthermore, the communication component 1638 may include components that detect identifiers or components that can be operated to detect identifiers. For example, the communication component 1638 may include an RFID (Radio Frequency Identification) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor that detects one-dimensional barcodes such as UPC (Universal Product Code) barcodes, QR (Quick Response) codes, Aztec codes, data matrices, data glyphs, Maxi Codes, PDF417, Ultra Codes, UCC RSS-2D barcodes, and other optical codes), or an acoustic detection component (e.g., a microphone for identifying tagged audio signals). In addition, a variety of information can be derived through the communication component 1638, such as location via Internet Protocol (IP) geolocation, location via Wi-Fi® signal triangulation, or location via detection of NFC beacon signals that can point to a specific location.

[0119] Various memories (e.g., main memory 1614, static memory 1616, and / or memory of processor 1604) and / or storage unit 1618 can store one or more sets of instruction words and data structures (e.g., software) that are embodied or used in one or more of the methodologies or functions described herein. When executed by processor 1604, such instruction words (e.g., instruction word 1610) trigger various tasks to embody the disclosed examples.

[0120] Instruction 1610 can be transmitted or received over network 1622 using a transmission medium through a network interface device (e.g., a network interface component included in communication component 1638) and using any one of a variety of widely known transmission protocols (e.g., HTTP (Hypertext Transfer Protocol)). Similarly, instruction 1610 can be transmitted or received using a transmission medium through a connection to device 1624 (e.g., a peer-to-peer connection).

[0121] Figure 17 is a block diagram showing several examples of machine learning programs 1700. The machine learning programs 1700, also referred to as machine learning algorithms or tools, are used as part of the systems described herein to perform tasks related to retrieval and query response.

[0122] Machine learning is a field of study that grants computers the ability to learn without explicit programming. Machine learning explores the study and construction of algorithms, also referred to herein as tools, that can learn from or be trained on existing data and make predictions about or based on new data. Such machine learning tools operate by building models from exemplary training data 1708 to make data-based predictions or decisions, which are expressed as outputs or evaluations (e.g., evaluations 1716). Although examples are presented in relation to some machine learning tools, the principles presented herein can be applied to other machine learning tools.

[0123] In some examples, different machine learning tools may be used. For example, logistic regression (LR), naive-bayes, random forest (RF), neural network (NN), matrix factorization, and support vector machine (SVM) tools may be used.

[0124] Two common types of problems in machine learning are classification problems and regression problems. Classification problems, also known as categorization problems, aim to classify an item into one of many categories (for example, is this object an apple or an orange?). Regression algorithms aim to quantify some items (for example, by providing real numerical values).

[0125] The machine learning program 1700 supports two types of stages: the training stage 1702 and the prediction stage 1704. In the training stage 1702, supervised learning, unsupervised learning, or reinforcement learning may be used. For example, the machine learning program 1700 (1) receives / receives features 1706 (e.g., as structured or labeled data for supervised learning) or (2) identifies features 1706 (e.g., unstructured or unlabeled data for unsupervised learning) in the training data 1708. In the prediction stage 1704, the machine learning program 1700 uses features 1706 to analyze query data 1712 and generates results or predictions as an example of evaluation 1716.

[0126] In training phase 1702, feature engineering is used to identify features 1706, which may include identifying useful, distinctive, and independent features for the effective operation of the machine learning program 1700 in pattern recognition, classification, and regression. In some examples, the training data 1708 includes pre-identified features 1706 and labeled data which are known data for one or more outcomes. Each feature 1706 may be a variable or attribute, such as an individually measurable characteristic of a process, article, system, or phenomenon represented in the dataset (e.g., training data 1708). Feature 1706 may also be of different types, such as digit features, strings, and graphs, and may include one or more of, for example, content 1718, concepts 1720, attributes 1722, historical data 1724, and / or user data 1726.

[0127] In training phase 1702, the machine learning program 1700 uses the training data 1708 to look for correlations between features 1706 that influence the predicted outcome or evaluation 1716.

[0128] Using the training data 1708 and the identified features 1706, the machine learning program 1700 is trained in machine learning program training 1710 during training phase 1702. The machine learning program 1700 evaluates the values ​​of features 1706 by their correlation with the training data 1708. The result of the training is the trained machine learning program 1714 (e.g., the trained or learned model).

[0129] In another example, training stage 1702 may involve machine learning where the training data 1708 is structured (e.g., labeled during preprocessing) and the trained machine learning program 1714 embodies a relatively simple neural network 1728 capable of performing tasks such as classification and clustering.

[0130] The neural network 1728 generated during training phase 1702 and embodied within the trained machine learning program 1714 may include a hierarchical (e.g., layered) organization of neurons. For example, neurons (or nodes) may be arranged hierarchically in many layers, including an input layer, an output layer, and many hidden layers. A layer in the neural network 1728 may have one or many neurons, which work-wise compute a small function (e.g., an activation function). For example, if the activation function produces a result that exceeds a certain critical value, the output may be transmitted from the neuron in question (e.g., a transmitting neuron) to connected neurons in a continuous layer (e.g., a receiving neuron). The connections between neurons also have associated weights, which determine the influence of the input from the transmitting neuron to the receiving neuron.

[0131] In some examples, neural network 1728 could also be one of many different types of neural networks, including, for example, single-layer feed-forward networks, artificial neural networks (ANNs), recurrent neural networks (RNNs), symmetrically concatenated neural networks and unsupervised pre-trained convolutional neural networks (CNNs), or recursive neural networks (RNNs).

[0132] During the prediction phase 1704, the trained machine learning program 1714 is used to perform evaluations. Query data 1712 is provided as input to the trained machine learning program 1714, and in response to receiving the query data 1712, the trained machine learning program 1714 generates an evaluation 1716 as output.

[0133] Returning to Figure 18, a schematic representation of the processing environment (1800), including processors 1802, 1806, and 1808 (for example, a GPU, CPU, or a combination thereof), is illustrated.

[0134] The processor 1802 is coupled to the power supply 1804 and is illustrated as including modules (permanently configured or temporarily instantiated), namely a data acquisition and preparation module 606, a model training and evaluation module 608, and a model deployment module 610.

[0135] Example

[0136] 1. A method for generating an image-based localization model for the navigation of a mobile robot, comprising the steps of: generating collected data by performing data collection at a plurality of different service locations where a swarm of mobile robots may be deployed; dividing the collected data into a plurality of blocks of continuous portions of the collected data; generating a first image-based localization model for a first service location among a plurality of different service locations using the collected data; generating a second image-based localization model for a second service location among a plurality of different service locations using the collected data; deploying the first image-based localization model to a first mobile robot in a swarm of mobile robots - the first mobile robot is deployed at the first service location among a plurality of different service locations, and the first mobile robot navigates the first service location using the first image-based localization model -; and deploying a second image-based localization model to a second mobile robot in a swarm of mobile robots - the second mobile robot is deployed at the second service location among a plurality of different service locations, and the second mobile robot navigates the second service location using the second image-based localization model -.

[0137] 2. A method in which one or more of the above examples include the acquisition data including image data, and the segmentation step includes the step of dividing the image data into blocks of the same size of a continuous image.

[0138] 3. A method in which one or more of the above examples include the step of generating a first image-based positioning model, wherein the step of training, developing, and testing the first image-based positioning model is performed using different blocks of the same size from a continuous image.

[0139] 4. A method comprising shuffling different blocks of the same size from a continuous image before assigning them to training, development, and testing, respectively, of an image-based positioning model, in one or more of the methods described above.

[0140] 5. A method comprising one or more of the above examples, wherein different blocks of the same size from a continuous image are randomly assigned to the training, development, and testing of the first image-based positioning model, respectively.

[0141] 6. A method, one or more of the above examples, that includes a step of automatically determining the size by balancing the size of each block of the same size in a continuous image based on a balanced distribution across grid cells of a map grid of a first service location.

[0142] 7. A method in which one or more of the above examples include image data including an image timestamp, each timestamp representing a capture period for the corresponding image, and the image timestamp for each image in each block of the same size in a sequence of images represents the same capture period for all images in the corresponding block, and the method further includes the step of assigning different blocks of the same size in a sequence of images to training, development and testing, respectively, of a first image-based positioning model based on the capture periods for the images in the blocks of the same size.

[0143] 8. A method in which, in one or more of the above examples, the step of assigning different blocks of the same size from a continuous image to training, development and testing of the first image-based positioning model, respectively, includes the step of assigning one or more blocks of the continuous image for a first capture period to only one of training, development and testing of the first image-based positioning model.

[0144] 9. A method in which, in one or more of the above examples, the image data includes a mission identifier, each image in the image data is associated with a mission identifier, each image in each block of the same size in a sequence of images is associated with a unique mission identifier from among a plurality of mission identifiers, and the method further includes the step of assigning different blocks of the same size in a sequence of images to the training, development, and testing of a first image-based positioning model, respectively, based on the mission identifier for the images in the blocks of the same size.

[0145] 10. A method in which, in one or more of the above examples, the step of assigning different blocks of the same size from a continuous image to training, development and testing of the first image-based positioning model, respectively, includes the step of assigning one or more blocks of the continuous image for a first mission identifier to only one of training, development and testing of the first image-based positioning model.

[0146] 11. A method in which, in one or more of the above examples, the step of generating a first image-based positioning model for a first service location among a plurality of different service locations includes: extracting first location data specific to the first service location from collected data; generating a plurality of online model performance measures based on the first location data associated with the current version of the first image-based positioning model; performing an offline evaluation of the current version of the first image-based positioning model using at least a portion of the first location data; and automatically generating a new version of the first image-based positioning model for the first service location based on the offline evaluation. Other technical features may be obvious to those skilled in the art from the drawings, description and claims below.

[0147] 12. A method in which, in one or more of the above examples, multiple online model performance metrics are reported by the operation stack of the first mobile robot.

[0148] 13. A method further comprising a step of retraining an image-based positioning model using multiple online model performance measures, in one or more of the methods described above.

[0149] 14. A method comprising, in one or more of the above examples, the steps of: the first mobile robot performing a restart operation; the first mobile robot automatically checking a remote storage to determine whether a new image-based positioning model has been generated and stored in the remote storage in response to the restart operation; the first mobile robot storing the new image-based positioning model in local memory in response to determining whether a new image-based positioning model has been generated and stored in the remote storage; and the first mobile robot providing an image-based positioning response to a positioning request.

[0150] 15. A method in which one or more of the above examples include a step of automatically checking a remote storage to determine whether a new image-based positioning model has been generated and stored in the remote storage, and the step of checking whether a retrained version of the image-based positioning model has now been generated.

[0151] 16. A method in which one or more of the above examples include a step of automatically checking a remote storage facility to determine whether a new image-based positioning model has been generated and stored in the remote storage facility, and further includes a step of checking whether a new image-based positioning model type has been generated.

[0152] 17. A method comprising one or more of the above examples, further comprising the steps of maintaining a plurality of image-based positioning model types and multiple versions of each of the plurality of image-based positioning model types in a cloud storage; and a first mobile robot, which implements fallback logic that enables the first mobile robot to use the plurality of image-based positioning model types and multiple versions of each of the plurality of image-based positioning model types.

[0153] 18. A method in which one or more of the above examples include a maintenance step of maintaining a file structure for storing multiple image-based positioning model types and multiple versions of each of the multiple image-based positioning model types in a cloud storage.

[0154] 19. A method in which, in one or more of the above examples, the fallback logic is included in the robotics stack of a first mobile robot and accesses a file structure to access at least one of multiple image-based positioning model types or multiple versions of each of multiple image-based positioning model types in a cloud storage.

[0155] 20. A computing device comprising at least one processor and memory, the memory storing instruction words which, when executed by the at least one processor, cause the device to perform any one of the methods described above.

[0156] 21. A non-temporary computer-readable storage medium comprising instructions that cause a computer to perform one of the methods described above when executed by the computer.

[0157] Glossary

[0158] "Carrier signal" refers to any intangible medium on which machine-executed instruction words can be stored, encoded, or transmitted, including digital or analog communication signals or other intangible media that facilitate the communication of such instruction words. Instruction words may be transmitted or received over a network using a transmission medium through a network interface device.

[0159] "Communication network" refers to one or more parts of a network that may include an ad hoc network, intranet, extranet, virtual private network (VPN), local area network (LAN), wireless LAN (WLAN), wide area network (WAN), wireless WAN (WWAN), metropolitan area network (MAN), the Internet, part of the Internet, part of a public switched telephone network (PSTN), plain old telephone service (POTS) network, mobile phone network, wireless network, Wi-Fi® network, other types of networks, or combinations of two or more such networks. For example, a network or part of a network may include a wireless or cellular network, and the coupling may be a Code Division Multiple Access (CDMA) coupling, a Global System for Mobile communication (GSM) coupling, or other types of cellular or wireless coupling.In this example, the coupling can embody any of the diverse types of data transmission technologies, such as 1xRTT (Single Carrier Radio Transmission Technology), EVDO (Evolution-Data Optimized) technology, GPRS (General Packet Radio Service) technology, EDGE (Enhanced Data rates for GSM Evolution) technology, 3GPP (third Generation Partnership Project) including 3G, 4G wireless networks, UMTS (Universal Mobile Telecommunications System), HSPA (High-Speed ​​Packet Access), WiMAX (Worldwide Interoperability for Microwave Access), LTE (Long Term Evolution) standards, other technologies defined by various standards-setting organizations, other long-range protocols, or other data transmission technologies.

[0160] "Component" refers to a device, physical entity, or logic having boundaries defined by a function or subroutine call, branch point, API, or other technique that provides a division or modularization of a particular processing or control function. Components can be coupled with other components through interfaces to perform a machine process. A component is a packaged functional hardware unit designed to be used with other components and may generally be part of a program that performs a particular function of the associated functionalities. Components may consist of software components (e.g., code embodied on a machine-readable medium) or hardware components. A "hardware component" is a tangible unit capable of performing a particular task and may be configured or arranged in a particular physical manner. For example, one or more computer systems (e.g., standalone computer systems, client computer systems, or server computer systems) or one or more hardware components of a computer system (e.g., processors or processor groups) may be configured as hardware components that operate by software (e.g., applications or application parts) to perform the particular tasks described herein. Hardware components may also be embodied mechanically, electronically, or in any appropriate combination thereof. For example, a hardware component may include a dedicated circuit or logic permanently configured to perform a specific task. A hardware component may be a special-purpose processor such as an FPGA (Field-Programmable Gate Array) or ASIC (Application-Specific Integrated Circuit). A hardware component may also include programmable logic or circuit temporarily configured by software to perform a specific task. For example, a hardware component may include software that runs on a general-purpose processor or other programmable processor.When configured by such software, a hardware component becomes a specific machine (or a specific component of a machine) customized to perform the configured function, and is no longer a general-purpose processor. Whether a hardware component is mechanically embodied, embodied by a permanently configured dedicated circuit, or embodied by a temporarily configured circuit (e.g., configured by software) can be determined by considering costs and time. Therefore, the phrase “hardware component” (or “embodied hardware component” should be understood to include tangible entities that are physically configured, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate in a particular manner or perform a particular task described herein. When considering examples where a hardware component is temporarily configured (e.g., programmed), the hardware component does not need to be configured or instantiated at any single point in time. For example, if a hardware component includes a general-purpose processor configured by software to become a special-purpose processor, the general-purpose processor may be configured as different special-purpose processors at different times (e.g., including different hardware components). Accordingly, the software configures a particular processor or processor to, for example, constitute a specific hardware component in one instance and another hardware component in another instance. Hardware components can provide information to and receive information from other hardware components. Therefore, the described hardware components can be considered to be communicatively coupled. When many hardware components exist simultaneously, communication can occur between two or more hardware components through signal transmission (e.g., through appropriate circuits and buses). In an example where many hardware components are configured or instantiated at different times, communication between such hardware components can occur, for example, through the storage and retrieval of information in a memory structure to which the many hardware components have access rights.For example, one hardware component may perform a task and save the output of that task to a memory device to which it is communicatively coupled. Subsequently, additional hardware components can access the memory device and retrieve and process the saved output. Hardware components may initiate communication with input or output devices and may operate with respect to resources (e.g., information gathering). The various tasks of the exemplary methods described herein may be performed by one or more processors that are temporarily or permanently configured (e.g., by software) to perform the relevant tasks. Whether temporary or permanent, such processors may constitute an embodiment of a processor that operates to perform one or more tasks or functions described herein. As used herein, “embodiment of a processor” refers to a hardware component embodied using one or more processors. Similarly, the methods described herein may be embodied at least partially by processors, and a particular processor or processor may be an example of hardware. For example, at least some of the tasks of the methods described herein may be performed by one or more processors or embodiment of a processor. Furthermore, one or more processors may operate to support the performance of related tasks as a “cloud computing” environment or “Software as a Service (SaaS)”. For example, at least part of the work may be performed by a group of computers (as an example of machines containing processors), and such work may be accessed through a network (e.g., the Internet) and one or more appropriate interfaces (e.g., APIs). The performance of a particular task may be distributed among processors and may reside not only within a single machine but also across multiple machines. In some examples, a processor or an embodiment of a processor may be located in a single geographical location (e.g., a home environment, an office environment, or within a server farm). In some examples, a processor or an embodiment of a processor may be distributed across multiple geographical locations.

[0161] "Computer-readable medium" refers to all machine storage and transmission media. Therefore, this term includes all storage devices / mediums and carrier / modulated data signals. The terms "machine-readable medium," "computer-readable medium," and "device-readable medium" mean the same thing and may be used interchangeably in this disclosure.

[0162] "Machine storage medium" refers to one or more storage devices and / or media (e.g., centralized or distributed databases and / or associated caches and servers) that store executable instructions, routines, and / or data. This term includes solid-state memory, including processor-internal or external memory, and optical and magnetic media. Specific examples of machine storage medium, computer storage medium, and / or device storage medium include, for example, semiconductor memory devices (e.g., EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), FPGAs, and flash memory devices), magnetic disks such as internal hard disks and portable disks, magneto-optical disks, and non-volatile memory including CD-ROMs and DVD-ROMs. The terms "machine storage medium," "device storage medium," and "computer storage medium" are synonymous and may be used interchangeably in this disclosure. The terms “machine storage medium,” “computer storage medium,” and “device storage medium” exclude carrier waves, modulated data signals, and other such media, some of which are encompassed by the term “signaling medium.”

[0163] A “module” refers to logic having boundaries defined by functions or subroutine calls, branching points, application program interfaces (APIs), or other techniques that provide the division or modularization of a particular processing or control function. Modules are typically coupled through interfaces with other modules to perform machine processes. A module is a packaged functional hardware unit designed to be used with other components and may generally be part of a program that performs a particular function of the associated function. A module may consist of a software module (e.g., code embodied on a machine-readable medium) or a hardware module. A “hardware module” is a tangible unit capable of performing a particular task and may be configured or arranged in a particular physical manner. In a variety of examples, one or more computer systems (e.g., standalone computer systems, client computer systems, or server computer systems) or one or more hardware modules of a computer system (e.g., processors or groups of processors) may be configured as hardware modules that operate by software (e.g., applications or application parts) to perform the particular tasks described herein. In some examples, hardware modules may be embodied mechanically, electronically, or in any appropriate combination thereof. For example, a hardware module may include dedicated circuitry or logic permanently configured to perform a specific task. For instance, a hardware module could be a special-purpose processor such as an FPGA or ASIC. A hardware module may also include programmable logic or circuitry ad-hocly configured by software to perform a specific task. For example, a hardware module could include software that runs on a general-purpose processor or other programmable processor. When configured with such software, the hardware module becomes a specific machine (or a specific component of a machine) uniquely customized to perform the configured function, and is no longer a general-purpose processor.It will be understood that the decision of whether to mechanically implement a hardware module, implement it with a permanently configured dedicated circuit, or implement it with a temporarily configured circuit (e.g., configured in software) can be made in consideration of cost and time. Therefore, the phrase “hardware module” (or “hardware implementation module”) should be understood to include tangible entities that are physically configured, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate in a particular manner or perform a particular task as described herein. When considering examples where a hardware module is temporarily configured (e.g., programmed), each hardware module does not need to be configured or instantiated at any single point in time. For example, if a hardware module includes a general-purpose processor configured by software to become a special-purpose processor, the general-purpose processor may be configured as different special-purpose processors at different times (e.g., composed of different hardware modules). Accordingly, the software may configure a particular processor or processor to, for example, constitute a particular hardware module in one instance and another hardware module in another instance. Hardware modules can provide information to and receive information from other hardware modules. Therefore, the described hardware modules can be considered to be communicatively coupled. When multiple hardware modules exist simultaneously, communication can occur between two or more hardware modules through signal transmission (e.g., through appropriate circuits and buses). In an example where multiple hardware modules are configured or instantiated at different times, communication between such hardware modules can occur, for example, through the storage and retrieval of information in a memory structure to which the multiple hardware modules have access. For example, one hardware module may perform a task and store the output of that task in a memory device to which it is communicatively coupled.Subsequently, additional hardware modules can access memory devices to retrieve and process the stored output. Hardware modules may initiate communication with input or output devices and may operate with respect to resources (e.g., information gathering). The diverse operations of the exemplary methods and routines described herein may be performed, at least in part, by one or more processors that are temporarily or permanently configured (e.g., by software) to perform the relevant operations. Whether temporary or permanent, such processors may constitute a processor embodiment module operating to perform one or more operations or functions described herein. As used herein, “processor embodiment module” refers to a hardware module embodied using one or more processors. Similarly, the methods described herein may be embodied, at least in part, by processors, and a particular processor or processor may be an example of hardware. For example, at least part of the operations of the method may be performed by one or more processors or processor embodiment modules. Furthermore, one or more processors may operate to support the performance of the relevant operations as a “cloud computing” environment or “software as a service (SaaS)”. For example, at least part of a task may be performed by a group of computers (as an example of a machine containing a processor), and such tasks may be accessed through a network (e.g., the Internet) and one or more appropriate interfaces (e.g., APIs). The performance of a particular task may be distributed among processors and may reside not only within a single machine but also across multiple machines. In some examples, a processor or a processor embodiment module may be located in a single geographical location (e.g., a home environment, an office environment, or within a server farm). In other examples, a processor or a processor embodiment module may be distributed across multiple geographical locations.

[0164] The term "processor" refers to any circuit or virtual circuit (a physical circuit emulated by the logic executed by the actual processor) that manipulates data values ​​using control signals (e.g., "instructions," "op codes," "machine code," etc.) and generates corresponding output signals applied to machine operation. A processor can be, for example, a CPU (Central Processing Unit), a RISC (Reduced Instruction Set Computing) processor, a CISC (Complex Instruction Set Computing) processor, a GPU (Graphics Processing Unit), a DSP (Digital Signal Processor), an ASIC (Application-Specific Integrated Circuit), an RFIC (Radio-Frequency Integrated Circuit), or any combination thereof. A processor can also be a multicore processor, having two or more independent processors (sometimes referred to as "cores") that can execute instruction words simultaneously.

[0165] "Signal medium" refers to any intangible medium on which machine executable instructions can be stored, encoded, or transmitted, and includes digital or analog communication signals or other intangible mediums that facilitate the communication of software or data. The term "signal medium" may include any form of modulated data signal, carrier wave, etc. The term "modulated data signal" means a signal in which one or more of its characteristics are set or modified in order to encode information within the signal. The terms "transmission medium" and "signal medium" mean the same thing and may be used interchangeably in this disclosure.

[0166] The various drawings of this application include block diagrams, flowcharts, and control flowcharts of methods, systems, and program products according to the present invention. It will be understood that each block or stage of a block diagram, flowchart, and control flowchart, and combinations of blocks within a block diagram, flowchart, and control flowchart, can be embodied by computer program instructions. Such computer program instructions can be loaded into a computer or other programmable device to generate a machine, and the instructions executed by the computer or other programmable device generate means for embodiing the functions specified in the block diagram, flowchart, or control flow blocks or stages. Such computer program instructions may be stored in computer-readable memory that can instruct the computer or other programmable device to function in a particular manner, and the instructions stored in computer-readable memory generate a manufactured article containing instruction means for embodiing the functions specified in the block diagram, flowchart, or control flow blocks or stages. Computer program instructions can also generate a computer implementation process, which is loaded into a computer or other programmable device and performs a series of work steps in the computer or other programmable device, so that the instructions executed in the computer or other programmable device provide steps for implementing a function specified in a block diagram, flowchart, or control flow block(s) or steps(s).

[0167] Therefore, the blocks or stages in a block diagram, flowchart, or control flowchart support combinations of means for performing a specified function, combinations of stages for performing a specified function, and program instruction means for performing a specified function. It will also be understood that each block or stage in a block diagram, flowchart, or control flowchart, and combinations of blocks or stages in a block diagram, flowchart, or control flowchart, can be embodied by a special-purpose hardware-based computer system or a combination of special-purpose hardware and computer instruction words that performs a specified function or stage.

[0168] The above description has been presented with reference to various embodiments. Those with ordinary skill in the industry and art to which this application belongs will understand that the described structures and methods of operation can be modified and altered without significant departure in principle, idea and scope. Modifications and alterations to the examples disclosed can be made without departing the scope of this disclosure. Such, and other, modifications and alterations are intended to be within the scope of this disclosure as expressed in the following claims.

Claims

1. A method for generating an image-based localization model for the navigation of a mobile robot, The stage in which a swarm of mobile robots performs data collection and generates collected data at multiple different service locations where they may be deployed; The step of dividing the collected data into a plurality of blocks, each consisting of a temporally continuous portion of the collected data; A step of generating a first image-based positioning model for a first service location among the plurality of different service locations, wherein the generation is performed using a block from the divided plurality of blocks that corresponds to the collected data collected at the first service location; A step of generating a second image-based positioning model for a second service location among the plurality of different service locations, wherein the generation is performed using a block from the plurality of divided blocks that corresponds to the collected data collected at the second service location; The first step involves deploying the first image-based positioning model to a first mobile robot among the mobile robot swarm—the first mobile robot is deployed to a first service location among the plurality of different service locations, and the first mobile robot operates at the first service location using the first image-based positioning model; The step of deploying the second image-based positioning model to the second mobile robot among the mobile robot swarm—the second mobile robot is deployed to the second service location among the plurality of different service locations, and the second mobile robot operates the second service location using the second image-based positioning model; The first image-based positioning model and the second image-based positioning model are models that receive images captured by the mobile robot as input and output pose estimates of the mobile robot. The acquired data includes image data, and the division step includes dividing the image data into blocks of the same size in a continuous image; and A method comprising the step of generating the first image-based positioning model, the step of training, developing, and testing the first image-based positioning model using different blocks of the same size from the continuous image.

2. The method according to claim 1, further comprising the step of shuffling disparate blocks of the same size from the continuous image before assigning disparate blocks of the same size from the continuous image to the training, development, and testing of the image-based positioning model, respectively.

3. The method according to claim 2, further comprising the step of randomly assigning different blocks of the same size from the continuous image to the training, development, and testing of the first image-based positioning model, respectively.

4. The method according to claim 1, further comprising the step of automatically determining the size by selecting the size of each block of the continuous image of the same size so that the distribution of the number of data points across the grid cells of the map grid of the first service location is uniform.

5. The aforementioned image data includes an image timestamp, where each timestamp represents the capture period for the corresponding image. The image timestamp for each image in the aforementioned continuous image block of the same size represents the same capture period for all images within the corresponding block. The method according to claim 1, further comprising the step of assigning different blocks of the same size in a sequence of images to the training, development, and testing of the first image-based positioning model, respectively, based on the capture period for the images in the same size blocks.

6. The method according to claim 5, wherein the step of assigning different blocks of the same size from the continuous image to training, development and testing the first image-based positioning model, respectively, includes the step of assigning one or more blocks of the continuous image for the first capture period to only one of training, development and testing the first image-based positioning model.

7. The aforementioned image data includes a mission identifier, and each image within the aforementioned image data is associated with the mission identifier. Each image in each block of the aforementioned continuous image of the same size is associated with a unique mission identifier among multiple mission identifiers. The method according to claim 1, further comprising the step of assigning different blocks of the same size in a sequence of images to the training, development, and testing of the first image-based positioning model, respectively, based on a mission identifier for an image in the same size block.

8. The method according to claim 6, wherein the step of assigning different blocks of the same size from the continuous image to training, development and testing of the first image-based positioning model, respectively, includes the step of assigning one or more blocks of the continuous image associated with a first task identifier from among a plurality of task identifiers contained in the image data to only one of training, development and testing of the first image-based positioning model.

9. The step of generating the first image-based positioning model for the first service location among the plurality of different service locations is: A step of extracting specific first location data from the collected data to the first service location; A step of generating multiple online model performance metrics based on first location data associated with the current version of the first image-based positioning model; A step of performing an offline evaluation of the current version of the first image-based positioning model using at least a portion of the first location data; and The method according to claim 1, further comprising the step of automatically generating a new version of the first image-based positioning model for the first service location based on the offline evaluation.

10. The method according to claim 9, wherein the plurality of online model performance metrics are reported by the operation stack of the first mobile robot.

11. The method according to claim 9, further comprising the step of retraining the image-based positioning model using the plurality of online model performance metrics.

12. The first mobile robot performs the restart operation; In response to the restart operation, the first mobile robot automatically checks the remote storage facility to determine whether a new image-based positioning model has been generated and stored there; In response to determining whether the new image-based positioning model has been generated and stored in the remote storage facility, the first mobile robot stores the new image-based positioning model in local memory; and The method according to claim 1, further comprising the step of providing an image-based positioning response to a positioning request using the first mobile robot.

13. The method according to claim 12, wherein the step of automatically checking the remote storage to determine whether the new image-based positioning model has been generated and stored in the remote storage includes the step of checking whether a retrained version of the image-based positioning model has now been generated.

14. The method according to claim 12, wherein the step of automatically checking the remote storage facility to determine whether the new image-based positioning model has been generated and stored in the remote storage facility includes the step of checking whether a new image-based positioning model type has been generated.

15. The steps include maintaining multiple image-based positioning model types and multiple versions of each of the said multiple image-based positioning model types in a cloud storage facility; and The method according to claim 1, further comprising the step of implementing a fallback logic in the first mobile robot that enables the first mobile robot to use the plurality of image-based positioning model types and a plurality of versions of each of the plurality of image-based positioning model types.

16. The method according to claim 15, wherein the maintenance step includes maintaining a file structure for storing the plurality of image-based positioning model types and the plurality of versions of each of the plurality of image-based positioning model types within the cloud storage.

17. The method according to claim 16, wherein the fallback logic is included in the robotics stack of the first mobile robot and accesses the file structure to access at least one of the plurality of image-based positioning model types or a plurality of versions of each of the plurality of image-based positioning model types in the cloud storage.

18. A computing device, At least one processor; and Including memory, The memory is used when the device is executed by the at least one processor. A swarm of mobile robots may be deployed to collect data at multiple different service locations and generate collected data; The collected data is divided into multiple blocks, each consisting of a temporally continuous portion of the collected data; A first image-based positioning model is generated for a first service location among the plurality of different service locations, and this generation is performed using the block corresponding to the collected data collected at the first service location from among the plurality of divided blocks; A second image-based positioning model is generated for a second service location among the multiple different service locations, and this generation is performed using the block corresponding to the collected data collected at the second service location from among the multiple divided blocks; The first image-based positioning model is placed on the first mobile robot among the mobile robot swarm—the first mobile robot is placed at the first service location among the plurality of different service locations, and the first mobile robot operates at the first service location using the first image-based positioning model; Store a command that configures the second image-based positioning model to be placed on the second mobile robot among the mobile robot swarm—the second mobile robot being placed at the second service location among the plurality of different service locations, and the second mobile robot operating at the second service location using the second image-based positioning model; The first image-based positioning model and the second image-based positioning model are models that receive images captured by the mobile robot as input and output pose estimates of the mobile robot. The collected data includes image data, and the division includes dividing the image data into blocks of the same size in a continuous image; and The generation of the first image-based positioning model includes a computing device that performs training, development, and testing of the first image-based positioning model using different blocks of the same size from the continuous image.