Distributed loading and training of machine learning models
By separating the trainer and data loader in a distributed computer system and using the service grid architecture to optimize resource utilization, the problem of high CPU resource utilization and underutilized GPU resources is solved, and the efficiency and performance of model building are improved.
Patent Information
- Application Number
- CN202480010359.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-03
- Filing Date
- 2024-02-02
- Publication Date
- 2025-09-12
AI Technical Summary
When building machine learning models, the data loading process is usually a CPU-intensive task, while the training process is a GPU-intensive task, resulting in high CPU resource utilization and underutilized GPU resources, forming a bottleneck and delaying the model building process.
By separating the trainer and data loader in a distributed computer system and using a service mesh architecture for communication management, we can separate CPU-intensive data loading from GPU-intensive training and optimize resource utilization.
It improves the efficiency of the model building process, maximizes the utilization of GPU resources, reduces the delay of the training process, and improves overall performance.
Smart Images

Figure CN120641919A_ABST
Abstract
Description
[0001] Priority claim
[0002] This application claims the benefit of priority to U.S. patent application serial number 18 / 164,230, filed on February 3, 2023, which is incorporated herein by reference in its entirety. Background Art
[0003] Machine learning is a field of study that gives computers the ability to learn without being explicitly programmed. Machine learning explores the study and construction of algorithms (also referred to as tools in this article) that can learn from existing data or be trained using existing data and make predictions on or based on new data. BRIEF DESCRIPTION OF THE DRAWINGS
[0004] In the drawings, which are not necessarily drawn to scale, like reference numerals may describe similar components in different views. To easily identify the discussion of any particular element or action, the highest-order digit or digits in a reference numeral refer to the figure in which the element is first introduced. Some non-limiting examples are shown in the figures of the accompanying drawings, in which:
[0005] Figure 1 A high-level diagrammatic representation of a distributed computer system for data loading and model training based on some examples.
[0006] Figure 2 is a diagrammatic representation of a distributed computer system for data loading and model training according to some examples.
[0007] Figure 3 is a diagrammatic representation of the control plane of a service mesh according to some examples.
[0008] Figure 4 It is a reference Figure 1 An example distributed computer system is described in accordance with some example distributed data loading and model training methods.
[0009] Figure 5 It is a reference Figure 2 An example distributed computer system is described in accordance with some examples of a method for distributed data loading and model training across a service grid for generating an object tracking model.
[0010] Figure 6 The training and use of a machine learning program according to some examples is shown.
[0011] Figure 7 is a diagrammatic representation of a networked environment in which the present disclosure may be deployed, according to some examples.
[0012] Figure 8is a diagrammatic representation of an interactive system having both client-side and server-side functionality, according to some examples.
[0013] Figure 9 is a diagrammatic representation of data structures maintained in a database according to some examples.
[0014] Figure 10 is a diagrammatic representation of messages according to some examples.
[0015] Figure 11 A system including a head wearable device according to some examples is shown.
[0016] Figure 12 is a diagrammatic representation of a machine in the form of a computer system within which a set of instructions may be executed, causing the machine to perform any one or more of the methodologies discussed herein, according to some examples.
[0017] Figure 13 is a block diagram illustrating a software architecture in which examples may be implemented. DETAILED DESCRIPTION
[0018] Examples of the present disclosure improve the speed and efficiency of the machine learning model building process through a horizontally distributed and scalable architecture.
[0019] Machine learning models are used across a variety of applications. For example, in the field of object tracking, it may be useful to create a machine learning model for hand tracking. In some applications, a machine learning model for hand tracking can be trained to detect certain gestures captured by a camera on a mobile or wearable device.
[0020] Building a model of this nature can be a time-consuming project, requiring both data loading and model training. For example, in some applications, a data loader application can be used to load and prepare training data, and then feed the training data in batches to a trainer application. The trainer provides a machine learning algorithm that iteratively learns from the training data to create a model.
[0021] In some examples, such as in the construction of an object tracking model, when comparing the data loading process to the training process, data loading is typically a more central processing unit (CPU) intensive task than training, while training is typically a graphics processing unit (GPU) intensive task. This may create one or more technical challenges. For example, when building a hand tracking model, the data loading process may create a bottleneck, resulting in high CPU resource utilization and insufficient GPU resource utilization. For example, the data loader may not be able to generate batches of training data at a speed sufficient to saturate the GPU resources available for the training process, causing the training process to sometimes be idle and ultimately delay the construction or final completion of the model.
[0022] In some examples, GPU-intensive model training is separated from CPU-intensive data loading by implementing these corresponding processes on different processors, machines, or execution units (or on different processors, machines, and execution units) in a distributed computer system. These processes can be connected via a service grid architecture. By deploying a distributed computer system according to an example of the present disclosure, the technical problem of GPU resources not being fully utilized or GPUs being idle during data loading is alleviated.
[0023] In some examples, the distributed computer system includes a trainer and multiple data loaders serving the trainer. The trainer can transmit the workload to the data loader via a remote procedure call (RPC). Traffic management and load balancing between the trainer (e.g., a model trainer client component) and the data loader (e.g., a data loader server component) can be performed, for example, by a network agent that defines the data plane of the service grid. Some examples of the present disclosure can solve the technical problem that the model building process is relatively inefficient due to the bottleneck formed at the data loading process. The architecture according to the examples of the present disclosure can be used for the construction of machine learning models such as object tracking models (e.g., machine learning models for hand tracking or human gesture recognition).
[0024] Figure 1 is a block diagram illustrating a distributed computer system 100 for building a machine learning model according to some examples. The distributed computer system 100 includes a trainer 102 and three data loaders: data loader 106a, data loader 106b, and data loader 106c.
[0025] The trainer 102 may be configured to implement a machine learning model training application. The trainer 102 receives training data in batches (referred to herein as training batches) from the data loaders 106a, 106b, and 106c and uses a machine learning algorithm to iteratively learn from the training batches to build a machine learning model, such as an object tracking model.
[0026] The data loaders 106a, 106b, 106c can be configured to implement a data loading application and communicatively coupled to a storage component 108 that stores input data. The storage component 108 can be any suitable machine storage medium (e.g., cloud storage or a persistent solid-state drive (SSD) disk) that can be accessed by the data loaders 106a, 106b, 106c to read the input data. In some examples, the input data is "raw data" that needs to be processed by the data loader before it can be fed into the training program. This can be referred to as "preprocessing," and examples of preprocessing actions are described below.
[0027] When building certain types of models (such as object tracking models), the training process can be relatively GPU-intensive, while the data loading process can be relatively CPU-intensive. Figure 1 As shown, the trainer 102 is separated from the data loaders 106a, 106b, 106c to separate CPU-intensive workloads and GPU-intensive workloads from each other. The trainer 102 and the loaders 106a, 106b, 106c (or the trainer 102 and each loader 106a, 106b, 106c) can be run, for example, on separate processors and / or machines (physical or virtual) and / or on separate computing nodes in the distributed computer system 100. For example, this can solve the technical problem that a single machine (physical or virtual) does not have enough processing power to handle both CPU-intensive workloads and GPU-intensive workloads in an efficient manner. It should be understood that in the context of this specification, the term "CPU" also extends to virtual CPUs, commonly referred to as vCPUs. Therefore, reference to CPU resources should be interpreted as extending to vCPUs to the extent applicable.
[0028] According to an example, the trainer may communicate with the data loader (and vice versa) via any suitable network, such as the Internet or a local network. Figure 1 As shown, trainer 102 is communicatively coupled to data loaders 106a, 106b, and 106c via a service grid 104. Service grid 104 allows for horizontal scaling of the data loading process to improve the efficiency of the model building process. Communication between trainer 102 and loaders 106a, 106b, and 106c is achieved via a service grid architecture, as will be described in more detail below.
[0029] In use, the trainer 102 may send a training data request to the data loaders 106a, 106b, 106c. In response, the data loaders 106a, 106b, 106c pre-process the input batches read from the input data in the storage component 108 to generate training batches. These training batches are fed to the trainer 102, allowing the trainer to train a machine learning model on the training batches (training data). Figure 2 and Figure 3 The examples shown in describe these and other aspects in more detail.
[0030] Although Figure 1 106a, 106b, 106c, and one storage component 108. The example depicts one trainer 102, three data loaders 106a, 106b, 106c, and one storage component 108. However, it should be understood that some examples may include more or fewer data loaders, more trainers, more storage components, or a combination thereof. For example, it may be desirable to horizontally scale not only data loaders but also trainers using a service mesh architecture.
[0031] Figure 2 is a block diagram illustrating further details of a distributed computer system 200 for machine learning model building according to some examples. The distributed computer system 200 is shown to include a cluster 202 and an external storage component 216 communicatively coupled to the cluster 202, the cluster 202 including a plurality of components, as described below. The external storage component 216 may be similar to that described with reference to FIG. Figure 1 The storage component 108 is described.
[0032] Cluster 202 provides a set of computing nodes that run containerized applications. Kubernetes TM is an example of an open source system for deploying clusters of this nature. Generally, clusters allow containers to run across multiple machines and environments: virtual, physical, cloud-based, or on-premises. Unlike traditional virtual machines, containers are not limited to a specific operating system. A "node" is a component that runs applications in cluster 202. They can be virtual machines or physical computers, all running as part of a single system. Therefore, clusters provide node pools, which are pools of resources that are applied as needed to run cluster components.
[0033] Each execution unit in cluster 202 can be referred to as a "pod". Figure 2 As shown, cluster 202 includes five pods: pod 208a, pod 208b, pod 208c, pod 208d, and pod 208e. A pod can typically encapsulate one or more applications. Pods can be ephemeral in nature, meaning that if a pod (or the node on which it executes) fails, a new copy of the pod can be automatically created to continue operations.
[0034] Now more specifically turning to Figure 2 pods present in , pod 208a includes a containerized application in the form of a trainer 212 and also includes a proxy container called a network proxy 210a. Pods 208b, 208c, 208d, 208e each include a containerized application in the form of a data loader 214a, 214b, 214c, 214d. Figure 2 In FIG, there is a one-to-many relationship between the trainer 212 and the data loaders 214a, 214b, 214c, 214d, where the data loaders 214a, 214b, 214c, 214d provide multiple instances of the data loader application or service. The data loaders 214a, 214b, 214c, 214d can be similar to the reference Figure 1 The data loaders 106a, 106b, 106c described above and the trainer 212 may be similar to those described with reference to Figure 1 In some examples, and as Figure 2 As shown, each data loader can be a separate instance contained in its own pod. Alternatively, a pod can include multiple data loaders operating as separate instances. Although the trainer 212 and data loaders 214a, 214b, 214c, 214d are Figure 2 2 as being deployed in the same cluster 202, but in other examples, the trainer may be deployed in a different cluster or otherwise deployed "outside" of the cluster containing one or more data loaders.
[0035] Each pod 208b, 208c, 208d, 208e also includes a network proxy 210b, 210c, 210d, 210e. The network proxy can be injected into the pod as a so-called "sidecar". The network proxies 210a, 210b, 210c, 210d, 210e define the data plane 206 in the cluster 202. The control plane 204 is provided and is communicatively coupled to each of the network proxies 210a, 210b, 210c, 210d, 210e. Figure 3 The control plane according to some examples is described in more detail.
[0036] The data plane 206 and the control plane 204 define a service mesh. The service mesh can be summarized as adding a layer to the cluster 202 at the pod level and is configured to (among other things) manage traffic between the pod 208a that houses the trainer 212 and the other pods 208b, 208c, 208d, 208e that house the data loaders 214a, 214b, 214c, 214d. Examples of platforms that can be used to define and / or deploy a service mesh include Istio TM and Linkerd TM .
[0037] Therefore, each pod in cluster 202 has a network proxy uniquely associated with it. Data plane 206 provides a network proxy that is “located” in the relevant application (in Figure 2 In the configuration shown are these network proxies between the trainer 212 and the data loaders 214a, 214b, 214c, 214d), and the control plane 204 is configured to control the operation of the network proxies and provide an interface for users operating the service mesh.
[0038] When defining a service mesh between applications or services, the applications or services do not communicate directly with each other, but rather through a network proxy. In other words, the network proxies 210a, 210b, 210c, 210d, 210e intercept and manage communications between the trainer 212, which can be considered a client application, and the data loaders 214a, 214b, 214c, 214d, which can be considered server applications. The network proxies 210a, 210b, 210c, 210d, 210e can be configured to, for example, perform request-level load balancing and implement retries. The functionality of the network proxies is further described below.
[0039] The distributed computer system 200 can be used in a model building process, for example, for building an object tracking model. When building a model of this nature, it may be desirable to maximize the utilization of the GPU resources in the distributed computer system 200. This can be facilitated by deploying a trainer 212 and allocating GPU resources in the distributed computer system 200 to the trainer 212, while horizontally scaling the CPU-intensive data loaders 214a, 214b, 214c, 214d that serve the trainer 212. Examples of the present disclosure utilize systems such as Figure 2 The architecture shown in
[15] is used to optimize or improve CPU and input / output (I / O) performance while always achieving or maintaining the desired GPU performance.
[0040] exist Figure 2 In FIG, the distributed computer system 200 includes a single cluster 202. It should be understood that in other examples, the service grid can be distributed across multiple clusters. Figure 2 One trainer, four data loaders, and one external storage device are shown in FIG, but it should be understood that some examples may include more or fewer data loaders, more trainers, more storage components, or a combination thereof.
[0041] Thus, a distributed computer system can be designed and scaled to achieve desired or optimal resource utilization or throughput. For example, to build a specific object tracking model (e.g., a hand tracking model or a gesture detection model), it may be desirable to deploy a total of six pods similar to pod 208b, each containing a data loader. Sufficient CPU resources can then be allocated to these six pods to enable, for example, a total of eight NVIDIA TM A single trainer saturates a datacenter GPU with a "V100 Tensor Core." For example, each pod can contain 76 CPUs and 600GB of memory.
[0042] Figure 32 is a block diagram illustrating control plane 204 in more detail according to some examples. As mentioned, data plane 206 is deployed by adding a network proxy to each of the pods in cluster 202. Each network proxy is uniquely associated with one of the pods, and communication to and from that pod is accomplished via the network proxy. This includes communication to and from control plane 204. By way of example only, Figure 3 The lines of communication to and from network agent 210a are shown.
[0043] The control plane 204 may provide a set of services that run in a dedicated namespace. These services may perform actions such as aggregating telemetry data, providing user-facing APIs, and providing control data to the data plane 206. Figure 3 In FIG, the control plane 204 includes a controller 302 that provides a public API 304. The public API 304 can be accessed from a client device (e.g., via the web 308 or via a command line interface (CLI 306)) or another endpoint. The deployment of the controller 302 also includes the following containers: an identity component 310, a destination component 312, a tap component 314, a proxy injector 316, and an SP (service profile) validator 318.
[0044] The identity component 310 performs certificate authority (CA) functions. For example, the identity component 310 can accept requests from a network agent and return a certificate signed with the correct identity. The destination component 312 can provide service discovery. The tap component 314 can be designed to receive requests from the CLI 306 and act on them to monitor requests and responses.
[0045] The proxy injector 316 is responsible for transforming the pod's specification to add a "sidecar" containing the relevant network proxy. The SP validator 318 is configured to validate the new service profile.
[0046] The control plane 204 may also provide access to observability components 320a and 320b. The observability component 320a may be, for example, a software application for event monitoring and alerting, such as Prometheus. TM Prometheus TM Can be used to expose data such as metrics from a service mesh. Observability component 320b can be, for example, a software application for analysis and visualization, such as Grafana TM Grafana TMProvide actionable dashboards and metrics for services running on the service mesh. One or more dashboards can be provided to allow users to be presented with a real-time view of what is happening with the services in the cluster 202. For example, for each of the pods containing a data loader, a user or client can check the success rate, request data, wait time, utilization, etc.
[0047] It should be understood that service meshes can be deployed using other techniques or patterns than those disclosed herein. For example, the service mesh pattern is not limited to a particular environment or cluster architecture, or to a particular control plane, and a service mesh can be created regardless of whether applications or services are deployed using containers, virtual machines, or other deployments, or whether the deployment is on-premises, in the cloud, or a combination thereof.
[0048] Figure 4 Shown is a reference Figure 1 The distributed computer system 100 describes a method 400 for distributed data loading and model training according to some examples.
[0049] The method 400 begins at an open-loop block 402 and proceeds to block 404, where the trainer 102 is deployed separately from the data loaders 106a, 106b, 106c in a distributed manner, as described above. Sufficient CPU resources in the distributed computer system 100 can be allocated or assigned to the data loaders 106a, 106b, 106c to enable them to perform data loading at a desired speed or throughput. Given that machine learning computations are GPU intensive in some examples (e.g., when building visual element or visual object tracking models), GPU resources can be allocated to the trainer 102. According to some examples of the present disclosure, GPU resources are allocated such that each data loader does not utilize any GPU resources.
[0050] From block 404, method 400 proceeds to block 406, where the data loaders 106a, 106b, 106c access input data from storage component 108. According to some examples, method 400 can be used to build an object tracking model. Thus, the input data can, for example, include one or more of hand detection data, hand tracking data, gesture detection data, and gesture tracking data.
[0051] Object tracking can involve landmark detection, and the input data can therefore include landmark data. For example, landmark detection using machine learning can involve identifying key points or features in an image or video frame that can be used to track an object as it moves. Examples of landmarks can include corners, edges, or other unique or identifiable features in an object. In some examples, a machine learning model is trained on a dataset of images or video frames that include objects of interest (and in some examples, annotations indicating relevant landmarks). Once the model is trained, it can be used to detect landmarks in new images or video frames and track objects based on the movement of the landmarks over time. This technology can be used as an example form of visual object tracking in hand tracking. In other words, a machine learning model can be trained to detect and track a hand (or multiple hands) based on hand landmarks. Therefore, in some examples of building a hand tracking model, the input data includes an image of the hand and annotations indicating the location of the landmarks.
[0052] Back to Figure 4 , at block 408, the trainer 102 may send a training data request for execution by one of the data loaders. From block 408, method 400 proceeds to block 410, where, in response to receiving the training data request, the relevant data loader (e.g., data loader 106a) performs a data loading task. The data loading task may include reading a batch of data from the input data, and preprocessing the batch of data to generate a training batch for sending to the trainer 102. Thus, the data loading task may include multiple steps or subtasks, such as obtaining and preparing the training data in a "raw" format, and transforming or enhancing the data as needed. The preprocessing steps may include cropping and / or rotating the image, adding noise, or adjusting aspects such as brightness, contrast, saturation, hue, etc., or a combination thereof.
[0053] Then, at block 412, the data loader can send the training batch to the trainer 102. From block 412, method 400 proceeds to block 414, where the trainer 102 receives the training batch and uses a machine learning algorithm to learn from the training batch. This can be referred to as a training task. At block 416, the trainer 102 uses what has been learned or inferred from the training task to build or adjust the parameters of the machine learning model.
[0054] In some examples, the model building process is an iterative process, in which case it is desirable to periodically feed batches of training data to the trainer 102 until the building of the model is complete (e.g., until a satisfactory number of iterations of the training task have been completed or until the model performs satisfactorily). It will be appreciated that it may be desirable to feed as many different training batches as possible to the trainer 102. In some examples, the trainer may reuse the same training batches.
[0055] Therefore, if more training is needed (see decision block 418), then at block 422, the trainer 102 sends a request for additional training data. Blocks 410, 412, 414, and 416 may then be repeated, but it should be understood that a different data loader (e.g., data loader 106b) may receive the request for additional training data and generate new training batches for the trainer 212 to learn. This process may continue until further training is no longer needed or desired (see again decision block 418), in which case the method 400 ends at closed-loop block 420.
[0056] The model can be run for several epochs over the entire training dataset, where the training dataset is repeatedly fed into the model to improve its results. In each epoch, the entire training dataset is used to train the model. Multiple epochs (e.g., iterations over the entire training dataset) can be used to train the model. In some example embodiments, the number of epochs is 10, 100, 500, or 1000. Within an epoch, one or more batches of the training dataset (training batches) are used to train the model. Thus, the batch size ranges from 1 to the size of the training dataset, and the number of batches is any positive integer value. Model parameters can be updated after each batch.
[0057] Figure 5 Shown is a reference Figure 2 The distributed computer system 200 describes a method 500 for distributed data loading and model training across a service grid for generating an object tracking model according to some examples.
[0058] Method 500 begins at open loop block 502 and proceeds to block 504 where the method 500 is implemented by deploying the Figure 2 The network proxy described above defines a service mesh. For example, and referring to Figure 2 In the cluster 202, a dedicated network proxy 210a, 210b, 210c, 210d, 210e in the form of a proxy container can be added (e.g., injected) into each pod 208a, 208b, 208c, 208d, 208e in the cluster 202 to define a service mesh data plane 206. The control plane 204 is communicatively coupled to the data plane 206 and can include reference Figure 3 Describe the element.
[0059] From block 504, method 500 proceeds to block 506, where trainer 212 sends a training data request using a remote procedure call (e.g., gRPC). The call sent from trainer 212 can be a service-to-service call. Given the service mesh architecture deployed in distributed computer system 200, the training data request is routed via network proxy 210a, and at block 508, network proxy 210a handles traffic routing and load balancing to optimize utilization of trainer 212.
[0060] In some examples, trainer 212 uses remote procedure calls (such as gRPC) to call other processes in distributed computer system 200 (via a network proxy), for example, to request that a data loader perform a data loading task and feed training batches to trainer 212. Thus, trainer 212 can be considered a client, and the data loader acting in response to the call can be considered a server.
[0061] The network proxy can be configured to automatically detect HTTP / 2 and perform load balancing. The network proxy can monitor the control plane 204 and automatically update the load balancing pool based on the performance or rescheduling of pods 208b, 208c, 208d, 208e. When using a service such as Linkerd TM When using a service mesh platform, pods can be configured to proxy Transmission Control Protocol (TCP) traffic, but automatically detect the Layer 7 protocol being used. Application code making any TCP connections goes through its local Linkerd TM The instance causes the connection to be proxied, and if the connection uses gRPC, the service mesh automatically changes its behavior to layer 7 semantics—for example, by reporting success rates, retrying idempotent requests, load balancing at the request level, etc.
[0062] Return to Figure 5 , and in particular, block 510, the service grid is used to determine which data loader in the distributed computer system 200 a particular training data request should be sent to. For example, the network agent 210a in the service grid can use an exponentially weighted moving average of response latencies to determine which of the four data loading pods (208b, 208c, 208d, 208e) to send a particular training data request to at a given point in time. In other words, the training data request can be sent to the pod determined to be the "fastest" in each case. If the service grid determines that a pod is slow or unavailable, traffic can be diverted away from it to reduce latency and improve efficiency.
[0063] In some examples, the distributed computer system 200 can monitor the efficiency of its components, such as the utilization level of the trainer 212, and initiate certain actions to improve efficiency and / or throughput. For example, if it is determined that a data loading pod (208b, 208c, 208d, 208e) is causing a bottleneck, one or more additional data loading pods can be automatically deployed to serve the trainer 212 and increase its utilization. The one or more additional data loading pods can be copies of the original pod. Alternatively, if, for example, it is determined that the trainer 212 is saturated (e.g., all available GPU resources are utilized), one or more additional trainers can be automatically deployed. If multiple trainers are deployed, they can each be served by all available data loaders, or a subset of data loaders can be assigned to each trainer, depending on the implementation.
[0064] From block 510, method 500 proceeds to block 512, where the selected data loader (e.g., data loader 214a) receives the training data request. The training data request is routed via the network proxy in its pod (e.g., network proxy 210b in pod 208b). The data loader then performs the data loading task as described above. According to some examples, method 500 includes the data loader processing the training data request performing the data loading task in parallel (or partially in parallel) with the reading task using queue processing (see block 514).
[0065] The relevant data loader can create a queue to cache the read input data so that incoming training data requests can be served from the queue. For example, the data loaders 214a, 214b, 214c, and 214d in the distributed computer system 200 can each implement a multi-producer, multi-consumer queue. This can allow the data loader to run two processes: a read task that reads or extracts input batches from the storage component and adds the input batches to the queue, and a load task that is used to preprocess (e.g., transform and enhance) the input batches as needed. Such queue processing can improve the efficiency of the data loader because each new input batch can be simply preprocessed from the preprocessing queue without having to wait for a new input batch to be read from the storage device. Queue processing can decouple data loading and data serving, at least to some extent.
[0066] At block 516 , once a data loader (eg, data loader 214 a ) has completed the data loading task to generate a training batch, the training batch is sent to the trainer 212 via a network proxy in the service mesh.
[0067] As described above, the trainer 212 may iteratively perform training tasks using different training batches received from the data loaders 214a, 214b, 214c, 214d to generate a machine learning model. Figure 5 In some examples, the machine learning model is an object tracking model. In some examples, the trainer 212 can receive a large number of training batches from each of the data loaders to allow for iterative training tasks to build and improve parameters of the object tracking model. The training batches can include hand images with annotations reflecting landmark points or landmark data.
[0068] It should be understood that the actions described with reference to blocks 506 , 508 , 510 , 512 , 514 , and / or 516 may be repeated multiple times using different data loaders, e.g., in each case based on an exponentially weighted moving average of response delays, to generate a final or near-final version of the object tracking model at block 518 , and thereafter method 500 may end at closed-loop block 520 .
[0069] Reference Figure 4 and Figure 5 In both examples, although each of the flowcharts depicts a specific order of operations, the order may be changed without departing from the scope of the present disclosure. For example, some of the operations depicted may be performed in parallel or in a different order that does not materially affect the functionality of the process. In other examples, different components of the example devices or systems that implement the process may perform functions substantially simultaneously or in a specific order.
[0070] Figure 6 6 is a block diagram illustrating a machine learning program 600 according to some examples. The machine learning program 600 (also referred to as a machine learning algorithm or machine learning tool) can be used to perform operations associated with, for example, object tracking, hand detection, landmark detection, gesture recognition, or landmark recognition. For example, the machine learning algorithm can be used as part of the systems described herein, such as in trainer 102 and / or trainer 212.
[0071] The machine learning tools can operate by building models from example training data 608 to make data-driven predictions or decisions represented as outputs or evaluations (e.g., evaluations 616). Although examples are presented with respect to several machine learning tools, the principles presented herein can be applied to other machine learning tools.
[0072] In some examples, different machine learning tools can be used, such as logistic regression (LR), naive Bayes, random forest (RF), neural network (NN), matrix factorization, and support vector machine (SVM) tools.
[0073] Two common types of problems in machine learning are classification and regression. Classification problems (also known as classification problems) aim to classify an item into one of several categorical values (e.g., is this object an apple or an orange?). Regression algorithms aim to quantify some item (e.g., by providing a value that is a real number).
[0074] The machine learning program 600 supports two types of phases, namely a training phase 602 and a prediction phase 604. In the training phase 602, supervised learning, unsupervised learning, or reinforcement learning can be used. For example, the machine learning program 600 (1) receives features 606 (e.g., as structured data or labeled data in supervised learning) and / or (2) identifies features 606 in training data 608 (e.g., unstructured data or unlabeled data for unsupervised learning). In the prediction phase 604, the machine learning program 600 uses the features 606 to analyze the query data 612 to generate a result or prediction, as an example of an evaluation 616.
[0075] In the training phase 602, feature engineering is used to identify features 606, and can include identifying informative, discriminative, and independent features for efficient operation of the machine learning program 600 in pattern recognition, classification, and regression. In some examples, the training data 608 includes labeled data, which is known data about the pre-identified features 606 and one or more outcomes. Each of the features 606 can be a variable or attribute, such as a separate measurable characteristic of a process, item, system, or phenomenon represented by a data set (e.g., training data 608). By way of example only, the features 606 can also be of different types, such as numeric features, strings, and graphs, and can include one or more of content 618, concepts 620, attributes 622, historical data 624, and / or user data 626.
[0076] During the training phase 602 , the machine learning program 600 uses training data 608 to find correlations between features 606 that influence the predicted outcome or evaluation 616 .
[0077] Using training data 608 and identified features 606, machine learning program 600 is trained at machine learning program training 610 during training phase 602. Machine learning program 600 evaluates the value of features 606 as they relate to training data 608. The result of the training is a trained machine learning program 614 (e.g., a trained or learned model).
[0078] Furthermore, the training phase 602 may involve machine learning, where the training data 608 is structured (e.g., labeled during preprocessing operations), and the trained machine learning program 614 implements a relatively simple neural network 628 capable of, for example, performing classification and clustering operations. In other examples, the training phase 602 may involve deep learning, where the training data 608 is unstructured, and the trained machine learning program 614 implements a deep neural network 628 capable of performing both feature extraction and classification / clustering operations.
[0079] The neural network 628 generated during the training phase 602 and implemented within the trained machine learning program 614 can include a hierarchical (e.g., layered) organization of neurons. For example, neurons (or nodes) can be arranged hierarchically into several layers, including an input layer, an output layer, and multiple hidden layers. Each layer in the layers within the neural network 628 can have one or more neurons, and each of these neurons can be operated to calculate a small function (e.g., an activation function). For example, if the activation function generates a result that exceeds a certain threshold, the output can be transmitted from the neuron (e.g., a sending neuron) to a connected neuron (e.g., a receiving neuron) in a consecutive layer. The connections between neurons also have associated weights that define the influence on the input from the sending neuron to the receiving neuron.
[0080] In some examples, neural network 628 can also be one of multiple different types of neural networks, including, by way of example only, a single-layer feedforward network, an artificial neural network (ANN), a recurrent neural network (RNN), a symmetrically connected neural network and an unsupervised pre-trained network, a convolutional neural network (CNN), or a recurrent neural network (RNN).
[0081] During the prediction phase 604, a trained machine learning program 614 (also referred to as a machine learning model) is used to perform an evaluation. Query data 612 is provided as input to the trained machine learning program 614, and in response to receiving the query data 612, the trained machine learning program 614 generates an evaluation 616 as output.
[0082] Networked computing environment
[0083] Figure 7is a block diagram illustrating an example interactive system 700 for facilitating interactions over a network (e.g., exchanging text messages, conducting text, audio, and video calls, or playing games). The interactive system 700 includes multiple user systems 702, each of which hosts multiple applications including an interactive client 704 and other applications 706. Each interactive client 704 is communicatively coupled to other instances of the interactive client 704 (e.g., hosted on respective other user systems 702), an interactive server system 710, and third-party servers 712 via one or more communication networks including a network 708 (e.g., the Internet). The interactive client 704 can also communicate with the locally hosted application 706 using an application programming interface (API).
[0084] Each user system 702 may include multiple user devices, such as a mobile device 714 , a head wearable device 716 , and a computer client device 718 , which are communicatively connected to exchange data and messages.
[0085] The interaction clients 704 interact with other interaction clients 704 and with the interaction server system 710 via the network 708. The data exchanged between the interaction clients 704 (e.g., interaction 720) and between the interaction clients 704 and the interaction server system 710 includes functions (e.g., commands for activating functions) and payload data (e.g., text, audio, video, or other multimedia data).
[0086] The interactive server system 710 provides server-side functionality to the interactive clients 704 via the network 708. Although certain functions of the interactive system 700 are described herein as being performed by either the interactive clients 704 or the interactive server system 710, whether certain functions are located within the interactive clients 704 or within the interactive server system 710 may be a design choice. For example, it may be technically preferable to initially deploy certain technologies and functions within the interactive server system 710, but later migrate the technologies and functions to the interactive clients 704, where the user system 702 has sufficient processing power.
[0087] The interactive server system 710 supports various services and operations provided to the interactive clients 704. Such operations include sending data to the interactive clients 704, receiving data from the interactive clients 104, and processing data generated by the interactive clients 104. The data may include message content, client device information, geographic location information, media enhancements and overlays, message content persistence conditions, social network information, and live event information. The data exchange within the interactive system 700 is activated and controlled by functions available through the user interface (UI) of the interactive client 704.
[0088] Turning now specifically to the interaction server system 710, an application program interface (API) server 722 is coupled to the interaction server 724 and provides a programming interface thereto, making the functionality of the interaction server 724 accessible to the interaction clients 704, other applications 706, and third-party servers 712. The interaction server 724 is communicatively coupled to a database server 726, thereby facilitating access to a database 728 that stores data associated with interactions processed by the interaction server 724. Similarly, a web server 730 is coupled to the interaction server 724 and provides a web-based interface to the interaction server 724. To this end, the web server 730 processes incoming network requests via the Hypertext Transfer Protocol (HTTP) and several other related protocols.
[0089] The application program interface (API) server 722 receives and sends interaction data (e.g., commands and message payloads) between the interaction server 724 and the user system 702 (and, for example, the interaction client 704 and other applications 706) and the third-party server 712. Specifically, the application program interface (API) server 722 provides a set of interfaces (e.g., routines and protocols) that the interaction client 704 and other applications 706 can call or query to activate the functions of the interaction server 724. The application program interface (API) 722 exposes various functions supported by the interaction server 724, including account registration; login functionality; sending interaction data from a particular interaction client 704 to another interaction client 704 via the interaction server 724; transferring media files (e.g., images or videos) from the interaction client 704 to the interaction server 724; setting up a collection of media data (e.g., a story); retrieving a friend list of a user of the user system 702; retrieving messages and content; adding and removing entities (e.g., friends) from an entity graph (e.g., a social graph); locating friends within a social graph; and opening application events (e.g., related to the interaction client 704).
[0090] Interactive server 724 hosts multiple systems and subsystems, see below Figure 8 Provide a description.
[0091] System Architecture
[0092] Figure 8 7 is a block diagram illustrating additional details regarding an interactive system 700 according to some examples. Specifically, the interactive system 700 is shown as including an interactive client 704 and an interactive server 724. The interactive system 700 includes a number of subsystems that are supported on the client side by the interactive client 704 and on the server side by the interactive server 724. Example subsystems are discussed below.
[0093] Image processing system 802 provides various functions that enable a user to capture and enhance (eg, annotate or otherwise modify or edit) media content associated with a message.
[0094] The camera system 804 includes control software (e.g., in a camera application) that interacts with and controls the hardware camera hardware of the user system 702 (e.g., directly or via operating system controls) to modify and enhance the real-time images captured and displayed via the interactive client 704.
[0095] The enhancement system 806 provides functionality related to the generation and publication of enhancements (e.g., media overlays) for images captured in real time by the camera of the user system 702 or retrieved from the memory of the user system 702. For example, the enhancement system 806 is operable to select, present, and display media overlays (e.g., image filters or image lenses) for the interactive client 704 to enhance the real-time imagery received via the camera system 804 or the stored imagery retrieved from the memory 1102 of the user system 702. These enhancements are selected and presented to the user of the interactive client 704 by the enhancement system 806 based on inputs and data such as:
[0096] The geographic location of the user system 702; and
[0097] ●Social network information of users of the user system 702.
[0098] The enhancement can include audio and visual content and visual effects. The example of audio and visual content includes pictures, text, logos, animations and sound effects. The example of visual effects includes color superposition. Audio and visual content or visual effects can be applied to the media content items (for example, photos or videos) at user system 702 for transmitting in a message, or are applied to video content such as video content streams or feeds sent from interactive client 704. Therefore, image processing system 802 can interact with the various subsystems of communication system 808, and supports the various subsystems of communication system 808, which are for example messaging system 810 and video communication system 812.
[0099] Media overlays can include text or image data that can be superimposed on a photo taken by user system 702 or a video stream produced by user system 702. In some examples, the media overlay can be a location overlay (e.g., Venice Beach), the name of a live event, or a business name overlay (e.g., Beach Cafe). In another example, image processing system 802 uses the geographic location of user system 702 to identify a media overlay that includes the name of a business at the geographic location of user system 702. Media overlays can include other tags associated with the business. Media overlays can be stored in database 728 and accessed by database server 726.
[0100] Image processing system 802 provides a user-based publishing platform that enables users to select a geographic location on a map and upload content associated with the selected geographic location. Users can also specify situations in which specific media overlays should be provided to other users. Image processing system 802 generates a media overlay that includes the uploaded content and associates it with the selected geographic location.
[0101] The augmented reality creation system 814 supports the augmented reality developer platform and includes applications for content creators (e.g., artists and developers) to create and publish augmentations (e.g., augmented reality experiences) for the interactive clients 704. The augmented reality creation system 814 provides content creators with a library of built-in features and tools, including, for example, custom shaders, tracking techniques, and templates.
[0102] In some examples, the enhancement creation system 814 provides a merchant-based publishing platform that enables merchants to select specific enhancements associated with a geographic location through a bidding process. For example, the enhancement creation system 814 associates the highest bidding merchant's media overlay with the corresponding geographic location for a predefined amount of time.
[0103] The communication system 808 is responsible for enabling and processing various forms of communication and interaction within the interactive system 700 and includes a messaging system 810, an audio communication system 816, and a video communication system 812. The messaging system 810 is responsible for enforcing temporary or time-limited access to content by the interactive clients 704. The messaging system 810 includes multiple timers (e.g., in a transient timer system 818) that selectively enable access (e.g., for presentation and display) of messages and associated content via the interactive clients 704 based on duration and display parameters associated with a message or a collection of messages (e.g., a story). Additional details regarding the operation of the transient timer system 818 are provided below. The audio communication system 816 enables and supports audio communication (e.g., real-time audio chat) between multiple interactive clients 704. Similarly, the video communication system 812 enables and supports video communication (e.g., real-time video chat) between multiple interactive clients 704.
[0104] The user management system 820 is operationally responsible for managing user data and profiles, and includes a social networking system 822 that maintains information about relationships between users of the interactive system 700 .
[0105] The collection management system 824 is operationally responsible for managing collections or collections of media (e.g., collections of text, images, video, and audio data). Collections of content (e.g., messages, including images, video, text, and audio) can be organized into "event libraries" or "event stories." Such collections can be made available for a specified time period (e.g., the duration of the event to which the content relates). For example, content related to a concert can be made available as a "story" for the duration of the concert. The collection management system 824 is also responsible for publishing an icon providing notification of a particular collection to the user interface of the interactive client 704. The collection management system 824 includes curation functionality that enables collection managers to manage and curate specific content collections. For example, a curation interface enables event organizers to curate collections of content related to a specific event (e.g., removing inappropriate content or redundant messages). In addition, the collection management system 824 employs machine vision (or image recognition technology) and content rules to automatically curate content collections. In some examples, users can be compensated for including user-generated content in a collection. In such cases, the collection management system 824 operates to automatically pay such users for use of their content.
[0106] The mapping system 826 provides various geolocation functions and supports the presentation of map-based media content and messages by the interactive client 704. For example, the mapping system 826 enables the display of user icons or avatars (e.g., stored in the profile data 902) on a map to indicate the current or past locations of the user's "friends" within the context of the map, as well as media content generated by these friends (e.g., a collection of messages including photos and videos). For example, on the map interface of the interactive client 704, messages posted by the user to the interactive system 700 from a particular geographic location can be displayed to the "friends" of a particular user within the context of that particular location on the map. The user can also share his or her location and status information with other users of the interactive system 700 via the interactive client 704 (e.g., using an appropriate status avatar), where the location and status information is similarly displayed to the selected user within the context of the map interface of the interactive client 704.
[0107] The gaming system 828 provides various gaming functions within the context of the interactive client 704. The interactive client 704 provides a gaming interface that provides a list of available games that can be launched by a user within the context of the interactive client 704 and played with other users of the interactive system 700. The interactive system 700 also enables a particular user to invite other users to play a particular game by sending invitations to such other users from the interactive client 704. The interactive client 704 also supports voice, video, and text messaging (e.g., chat) within the context of game play, provides leaderboards for games, and also supports the provision of in-game rewards (e.g., game coins and items).
[0108] The external resource system 830 provides an interface for the interactive client 704 to communicate with a remote server (e.g., a third-party server 712) to launch or access external resources (i.e., applications or applets). Each third-party server 712 hosts, for example, an application or a small-scale version of an application based on a markup language (e.g., HTML5) (e.g., a game application, a utility application, a payment application, or a ride-sharing application). The interactive client 704 can launch a web-based resource (e.g., an application) by accessing an HTML5 file from a third-party server 712 associated with the web-based resource. The applications hosted by the third-party server 712 are programmed in JavaScript using a software development kit (SDK) provided by the interactive server 724. The SDK includes an application programming interface (API) with functions that can be called or activated by web-based applications. The interactive server 724 hosts a JavaScript library that provides access to a given external resource for specific user data of the interactive client 704. HTML5 is an example of a technology for programming games, but applications and resources programmed based on other technologies can be used.
[0109] To integrate the SDK's functionality into a web-based resource, the third-party server 712 downloads the SDK from the interaction server 724, or the third-party server 712 receives the SDK in some other manner. Once downloaded or received, the SDK is included as part of the application code of the external web-based resource. The code of the web-based resource can then call or activate certain functions of the SDK to integrate the features of the interaction client 704 into the web-based resource.
[0110] The SDK stored on the interactive server system 710 effectively provides a bridge between external resources (e.g., applications 706 or applet) and the interactive client 704. This gives users a seamless experience of communicating with other users on the interactive client 704 while also preserving the look and feel of the interactive client 704. In order to bridge the communication between the external resources and the interactive client 704, the SDK facilitates communication between the third-party server 712 and the interactive client 704. The WebViewJavaScriptBridge running on the user system 702 establishes two one-way communication channels between the external resources and the interactive client 704. Messages are sent asynchronously between the external resources and the interactive client 704 via these communication channels. Each SDK function activation is sent as a message and a callback. Each SDK function is implemented by constructing a unique callback identifier and sending a message with the callback identifier.
[0111] By using the SDK, not all information from the interactive client 704 is shared with the third-party server 712. The SDK limits which information is shared based on the needs of the external resource. Each third-party server 712 provides an HTML5 file corresponding to the web-based external resource to the interactive server 724. The interactive server 724 can add a visual representation of the web-based external resource (e.g., a box design or other graphics) in the interactive client 704. Once the user selects the visual representation or instructs the interactive client 704 to access a feature of the web-based external resource through the GUI of the interactive client 704, the interactive client 704 obtains the HTML5 file and instantiates the resource for accessing the feature of the web-based external resource.
[0112] The interactive client 704 presents a graphical user interface (e.g., a login page or title screen) for the external resource. During, before, or after presenting the login page or title screen, the interactive client 704 determines whether the launched external resource has previously been authorized to access the user data of the interactive client 704. In response to determining that the launched external resource has previously been authorized to access the user data of the interactive client 704, the interactive client 704 presents another graphical user interface of the external resource including the functions and features of the external resource. In response to determining that the launched external resource has not previously been authorized to access the user data of the interactive client 704, after displaying the login page or title screen of the external resource for a threshold period of time (e.g., 3 seconds), the interactive client 704 slides up a menu (e.g., animating the menu to emerge from the bottom of the screen to the middle or other part of the screen) to authorize the external resource to access the user data. The menu identifies the type of user data that the external resource is authorized to use. In response to receiving a user selection of the accept option, the interactive client 704 adds the external resource to the list of authorized external resources and enables the external resource to access the user data from the interactive client 704. External resources are authorized by the interactive client 704 to access user data under the OAuth 2 framework.
[0113] The interaction client 704 controls the type of user data shared with the external resource based on the type of external resource authorized. For example, an external resource comprising a full-scale application (e.g., application 706) is provided with access to a first type of user data (e.g., a two-dimensional avatar of the user with or without different avatar characteristics). As another example, an external resource comprising a small-scale version of an application (e.g., a web-based version of the application) is provided with access to a second type of user data (e.g., payment information, a two-dimensional avatar of the user, a three-dimensional avatar of the user, and an avatar with various avatar characteristics). Avatar characteristics include different ways to customize the look and feel of an avatar (e.g., different poses, facial features, clothing, etc.).
[0114] The advertising system 832 operatively enables third parties to purchase advertisements for presentation to end users via the interactive clients 704 and also handles the delivery and presentation of these advertisements.
[0115] The machine learning model system 834 can perform functions related to the training and implementation of machine learning models, for example, for gesture recognition, hand tracking, object detection, and related functions. The machine learning model system 834 can include or be communicatively coupled to a distributed computer system for data loading and model training, such as Figure 1 or Figure 2The machine learning model built using such a distributed computer system can be deployed on the client side by the interactive client 704, or on the server side by the interactive server 724, or on both sides.
[0116] Data Architecture
[0117] Figure 9 is a diagram illustrating a data structure 900 that may be stored in a database 904 of the interactive server system 710 according to certain examples. Although the contents of the database 904 are shown as including a plurality of tables, it should be understood that data may be stored in other types of data structures (e.g., an object-oriented database).
[0118] The database 904 includes message data stored in a message table 906. For any particular message, the message data includes at least message sender data, message recipient (or receiver) data, and payload. Figure 9 Additional details regarding information that may be included in a message and within the message data stored in message table 906 are described.
[0119] The entity table 908 stores entity data and is linked (e.g., by reference) to the entity graph 910 and the profile data 902. Entities for which records are maintained within the entity table 908 may include individuals, corporate entities, organizations, objects, places, events, and the like. Regardless of the entity type, any entity for which the interactive server system 710 stores data may be an identified entity. Each entity is provided with a unique identifier and an entity type identifier (not shown).
[0120] The entity graph 910 stores information related to relationships and associations between entities. By way of example only, such relationships may be social, professional (e.g., working at a common company or organization), interest-based, or activity-based. Some relationships between entities may be unidirectional, such as a subscription by an individual user to digital content from a business or publication (e.g., a newspaper or other digital media channel or brand). Other relationships may be bidirectional, such as a "friend" relationship between various users of the interactive system 700.
[0121] Certain permissions and relationships can be attached to each relationship, and can also be attached to each direction of the relationship. For example, a two-way relationship (e.g., a friend relationship between individual users) can include authorization to publish digital content items between the individual users, but certain restrictions or filters (e.g., based on content characteristics, location data, or time of day data) can be placed on the publication of such digital content items. Similarly, a subscription relationship between an individual user and a business user can impose varying degrees of restrictions on the publication of digital content from the business user to the individual user, and can significantly limit or prevent the publication of digital content from the individual user to the business user. As an example of an entity, a particular user can record certain restrictions in the record for that entity in the entity table 908 (e.g., through privacy settings). Such privacy settings can apply to all types of relationships in the context of the interactive system 700, or can be selectively applied to certain types of relationships.
[0122] Profile data 902 stores various types of profile data about a particular entity. Based on the privacy settings specified by the particular entity, profile data 902 can be selectively used and presented to other users of the interactive system 700. In the case where the entity is a person, profile data 902 includes, for example, the user's name, phone number, address, settings (e.g., notification and privacy settings), and an avatar representation (or a collection of such avatar representations) selected by the user. The particular user can then selectively include one or more of these avatar representations within the content of messages transmitted via the interactive system 700 and on a map interface displayed to other users by the interactive client 704. The collection of avatar representations can include a "status avatar," which presents a graphical representation of the user's status or activity that they may choose to transmit at a particular time.
[0123] Where the entity is a group, the profile data 902 for the group may similarly include one or more avatar representations associated with the group, in addition to the group name, members, and various settings for the relevant group (eg, notifications).
[0124] Database 904 also stores enhancement data, such as overlays or filters, in enhancement table 912. The enhancement data is associated with and applied to videos (video data is stored in video table 914) and images (image data is stored in image table 916).
[0125] In some examples, filters are displayed as overlays on an image or video during presentation to a recipient user. Filters can be of various types, including filters selected by a user from a set of filters presented to a sending user by interactive client 704 when the sending user is composing a message. Other types of filters include geolocation filters (also referred to as geofilters), which can be presented to a sending user based on a geographic location. For example, a geolocation filter specific to a nearby or special location can be presented within a user interface by interactive client 704 based on geographic location information determined by a global positioning system (GPS) unit of user system 702.
[0126] Another type of filter is a data filter, which can be selectively presented to the sending user by the interaction client 704 based on other input or information collected during the message creation process by the user system 702. Examples of data filters include the current temperature at a particular location, the current speed the sending user is traveling, the battery life of the user system 702, or the current time.
[0127] Other augmented data that can be stored in the image table 916 include augmented reality content items (e.g., corresponding to application "lenses" or augmented reality experiences). Augmented reality content items can be real-time special effects and sounds that can be added to images or videos.
[0128] The story table 918 stores data about a collection of messages and associated images, video, or audio data that are compiled into a collection (e.g., a story or gallery). The creation of a particular collection can be initiated by a particular user (e.g., each user for whom a record is maintained in the entity table 908). A user can create a "personal story" in the form of a collection of content that has been created and sent / broadcasted by that user. To this end, the user interface of the interactive client 704 can include a user-selectable icon that enables the sending user to add specific content to his or her personal story.
[0129] The collection can also constitute a "live story", which is a collection of content from multiple users created manually, automatically, or using a combination of manual and automatic techniques. For example, a "live story" can constitute a curated stream of user-submitted content from different locations and events. Users whose client devices have location services enabled and who are at a common location event at a particular time can be presented with the option of contributing content to a particular live story, for example, via the user interface of the interactive client 704. Live stories can be identified to the user by the interactive client 704 based on his or her location. The end result is a "live story" told from a group perspective.
[0130] Another type of content collection is called a "location story," which enables users whose user systems 702 are located in a particular geographic location (e.g., on a college or university campus) to contribute to a particular collection. In some examples, contributions to location stories can employ secondary authentication to verify that the end user belongs to a particular organization or other entity (e.g., is a student on a university campus).
[0131] As mentioned above, video table 914 stores video data that, in some examples, is associated with messages whose records are maintained within message table 906. Similarly, image table 916 stores image data associated with messages whose message data is stored in entity table 908. Entity table 908 can associate various enhancements from enhancement table 912 with various images and videos stored in image table 916 and video table 914.
[0132] Database 904 also includes a training data table 920 and a model data table 922. According to some examples, training data table 920 may store input data used by a data loader for preprocessing into training batches. According to some examples, training data table 920 may also store training batches. According to some examples, model data table 922 may store data related to a machine learning model (such as an object tracking model).
[0133] Data communication architecture
[0134] Figure 10 is a schematic diagram illustrating the structure of a message 1000 according to some examples, the message 1000 being generated by an interaction client 704 for transmission to another interaction client 704 via an interaction server 724. The content of a particular message 1000 is used to populate a message table 906 stored within a database 904 accessible by the interaction server 724. Similarly, the content of the message 1000 is stored in memory as "in-transit" or "in-flight" data of the user system 702 or the interaction server 724. The message 1000 is shown as including the following example components:
[0135] ●Message identifier 1002: a unique identifier that identifies the message 1000.
[0136] • Message text payload 1004 : Text to be generated by the user via the user interface of the user system 702 and included in the message 1000 .
[0137] • Message Image Payload 1006: Image data captured by the camera component of the user system 702 or retrieved from the memory component of the user system 702 and included in the message 1000. The image data for a message 1000 sent or received may be stored in the image table 916.
[0138] • Message Video Payload 1008: Video data captured by the camera component or retrieved from the memory component of the user system 702 and included in the message 1000. The video data for a message 1000 sent or received may be stored in the image table 916.
[0139] • Message audio payload 1010 : audio data captured by a microphone or retrieved from a memory component of the user system 702 and included in the message 1000 .
[0140] Message enhancement data 1012: Enhancement data (e.g., filters, stickers, or other annotations or enhancements) representing enhancements to be applied to the message image payload 1006, message video payload 1008, or message audio payload 1010 of the message 1000. Enhancement data for a sent or received message 1000 may be stored in the enhancement table 912.
[0141] ●Message duration parameter 1014: A parameter value indicating the amount of time in seconds that the content of the message (e.g., message image payload 1006, message video payload 1008, message audio payload 1010) is to be presented to the user via the interactive client 704 or made accessible to the user.
[0142] Message geolocation parameters 1016: Geolocation data (e.g., latitude and longitude coordinates) associated with the content payload of the message. Multiple message geolocation parameter 1016 values may be included in the payload, with each of these parameter values being associated with a content item included in the content (e.g., a specific image within the message image payload 1006 or a specific video within the message video payload 1008).
[0143] Message story identifier 1018: An identifier value that identifies one or more content collections (e.g., a "story" identified in stories table 918) associated with a particular content item in the message image payload 1006 of the message 1000. For example, the identifier value can be used to associate multiple images within the message image payload 1006 with each of multiple content collections.
[0144] Message Tags 1020: Each message 1000 can be tagged with a plurality of tags, each of which indicates the subject of the content included in the message payload. For example, if a particular image included in the message image payload 1006 depicts an animal (e.g., a lion), a tag value can be included within the message tags 1020 indicating the relevant animal. Tag values can be manually generated based on user input, or can be automatically generated using, for example, image recognition.
[0145] • Message sender identifier 1022: An identifier (eg, a messaging system identifier, an email address, or a device identifier) that indicates the user of the user system 702 on which the message 1000 was generated and from which the message 1000 was sent.
[0146] • Message recipient identifier 1024: An identifier (eg, a messaging system identifier, an email address, or a device identifier) that indicates the user of user system 702 to which message 1000 is addressed.
[0147] The content (e.g., value) of each component of message 1000 may be a pointer to a location in a table where the content data value is stored. For example, the image value in message image payload 1006 may be a pointer to a location in image table 916 (or the address of a location in image table 316). Similarly, the value in message video payload 1008 may point to data stored in image table 916, the value stored in message enhancement data 1012 may point to data stored in enhancement table 912, the value stored in message story identifier 1018 may point to data stored in story table 918, and the values stored in message sender identifier 1022 and message recipient identifier 1024 may point to user records stored in entity table 908.
[0148] System with head wearable device
[0149] Figure 11 A system 1100 is shown that includes a head wearable device 716 with a selector input device according to some examples. Figure 11 is a high-level functional block diagram of an example head wearable device 716 communicatively coupled to a mobile device 714 and various server systems 1104 (e.g., interactive server system 710) via various networks 708.
[0150] The head wearable device 716 includes one or more cameras, each of which may be, for example, a visible light camera 1106 , an infrared emitter 1108 , and an infrared camera 1110 .
[0151] The mobile device 714 is connected to the head wearable device 716 using both a low power wireless connection 1112 and a high speed wireless connection 1114. The mobile device 714 is also connected to the server system 1104 and the network 1116.
[0152] The head wearable device 716 also includes two image displays of the optical assembly, an image display 1118. The two displays 1118 of the optical assembly include one image display associated with the left lateral side of the head wearable device 716 and one image display associated with the right lateral side of the head wearable device 716. The head wearable device 716 also includes an image display driver 1120, an image processor 1122, low-power circuitry 1124, and high-speed circuitry 1126. The image display 1118 of the optical assembly is used to present images and videos, including images that may include a graphical user interface, to a user of the head wearable device 716.
[0153] The image display driver 1120 commands and controls the image display 1118 of the optical assembly. The image display driver 1120 can deliver image data directly to the image display 1118 of the optical assembly for presentation or can convert the image data into a signal or data format suitable for delivery to an image display device. For example, the image data can be video data formatted according to a compression format such as H.264 (MPEG-4 Part 10), HEVC, Theora, Dirac, RealVideo RV40, VP8, VP9, etc., while the still image data can be formatted according to a compression format such as Portable Network Graphics (PNG), Joint Photographic Experts Group (JPEG), Tagged Image File Format (TIFF), or Exchangeable Image File Format (EXIF).
[0154] The head-wearable device 716 includes a frame and stems (or temples) extending from lateral sides of the frame. The head-wearable device 716 also includes a user input device 1128 (e.g., a touch sensor or push buttons), comprising an input surface on the head-wearable device 716. The user input device 1128 (e.g., a touch sensor or push buttons) is used to receive input selections from the user for manipulating the graphical user interface of the presented image.
[0155] Figure 11The components of the head wearable device 716 shown in FIG are located on one or more circuit boards (e.g., PCBs or flexible PCBs) in the frame or temples. Alternatively or additionally, the depicted components may be located in the block, frame, hinge, or nosepiece of the head wearable device 716. The left and right visible light cameras 1106 may include digital camera elements, such as complementary metal oxide semiconductor (CMOS) image sensors, charge coupled devices, camera lenses, or any other corresponding visible light or light capturing elements that can be used to capture data (including images of a scene with an unknown object).
[0156] The head wearable device 716 includes a memory 1102 that stores instructions for performing a subset or all of the functions described herein. The memory 1102 may also include a storage device.
[0157] like Figure 11 As shown, high-speed circuitry 1126 includes a high-speed processor 1130, memory 1102, and high-speed wireless circuitry 1132. In some examples, image display driver 1120 is coupled to high-speed circuitry 1126 and operated by high-speed processor 1130 to drive the left and right image displays of image display 1118 of the optical assembly. High-speed processor 1130 can be any processor capable of managing high-speed communications and operations required by any general-purpose computing system for head-wearable device 716. High-speed processor 1130 includes the processing resources required to manage high-speed data transmission over high-speed wireless connection 1114 to a wireless local area network (WLAN) using high-speed wireless circuitry 1132. In some examples, high-speed processor 1130 executes an operating system (e.g., a Linux operating system) or other such operating system for head-wearable device 716, and the operating system is stored in memory 1102 for execution. In addition to any other responsibilities, high-speed processor 1130, which executes the software architecture of head-wearable device 716, manages data transmission with high-speed wireless circuitry 1132. In some examples, high-speed wireless circuitry 1132 is configured to implement the Institute of Electrical and Electronics Engineers (IEEE) 802.11 communication standard, also referred to herein as WiFi. In some examples, high-speed wireless circuitry 1132 can implement other high-speed communication standards.
[0158] The low-power wireless circuit system 1134 and the high-speed wireless circuit system 1132 of the head wearable device 716 may include a short-range transceiver (Bluetooth TM) and a wireless wide area network transceiver, a wireless local area network transceiver, or a wide area network transceiver (e.g., cellular or WiFi). Mobile device 714—including a transceiver that communicates via low-power wireless connection 1112 and a high-speed wireless connection 1114—can be implemented using details of the architecture of head wearable device 716, as well as other elements of network 1116.
[0159] Memory 1102 comprises any storage device capable of storing various data and applications, including camera data generated by left and right visible light cameras 1106, infrared camera 1110, and image processor 1122, as well as images generated for display on an image display 1118 of the optical assembly via image display driver 1120. While memory 1102 is shown integrated with high-speed circuitry 1126, in some examples, memory 1102 may be a separate, standalone component of head-worn device 716. In some such examples, electrical wiring may provide a connection from image processor 1122 or low-power processor 1136 to memory 1102 via a chip including high-speed processor 1130. In some examples, high-speed processor 1130 may manage addressing of memory 1102, such that low-power processor 1136 will initiate high-speed processor 1130 whenever a read or write operation involving memory 1102 is required.
[0160] like Figure 11 As shown, the low-power processor 1136 or the high-speed processor 1130 of the head wearable device 716 can be coupled to a camera device (a visible light camera device 1106, an infrared emitter 1108 or an infrared camera device 1110), an image display driver 1120, a user input device 1128 (for example, a touch sensor or a push button) and a memory 1102.
[0161] The head-worn device 716 is connected to a host computer. For example, the head-worn device 716 is paired with the mobile device 714 via a high-speed wireless connection 1114 or is connected to the server system 1104 via a network 1116. The server system 1104 can be one or more computing devices as part of a service or network computing system, for example, including a processor, memory, and a network communication interface to communicate with the mobile device 714 and the head-worn device 716 via the network 1116.
[0162] The mobile device 714 includes a processor and a network communication interface coupled to the processor. The network communication interface allows communication via a network 1116, a low-power wireless connection 1112, or a high-speed wireless connection 1114. The mobile device 714 may also store at least a portion of the instructions for generating binaural audio content in a memory of the mobile device 714 to implement the functionality described herein.
[0163] The output components of the head wearable device 716 include visual components, such as a display (e.g., a liquid crystal display (LCD), a plasma display panel (PDP), a light emitting diode (LED) display, a projector, or a waveguide). The image display of the optical assembly is driven by the image display driver 1120. The output components of the head wearable device 716 also include acoustic components (e.g., a speaker), tactile components (e.g., a vibration motor), other signal generators, etc. The input components of the head wearable device 716, the mobile device 714, and the server system 1104 (e.g., user input device 1128) can include alphanumeric input components (e.g., a keyboard, a touch screen configured to receive alphanumeric input, an optical keyboard, or other alphanumeric input components), point-based input components (e.g., a mouse, touchpad, trackball, joystick, motion sensor, or other pointing instrument), tactile input components (e.g., a physical button, a touch screen or other tactile input component that provides the location and force of a touch or touch gesture), audio input components (e.g., a microphone), etc.
[0164] The head wearable device 716 may also include additional peripheral elements. Such peripheral elements may include biometric sensors, additional sensors, or display elements integrated with the head wearable device 716. For example, the peripheral elements may include any I / O components including output components, motion components, positioning components, or any other such components described herein.
[0165] For example, the biometric component includes a component for detecting expressions (e.g., hand expressions, facial expressions, voice expressions, body postures, or eye tracking), measuring biological signals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), identifying people (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or electroencephalogram-based recognition), etc. The motion component includes an acceleration sensor component (e.g., an accelerometer), a gravity sensor component, a rotation sensor component (e.g., a gyroscope), etc. The positioning component includes a position sensor component for generating position coordinates (e.g., a global positioning system (GPS) receiver component), a Wi-Fi or Bluetooth receiver for generating positioning system coordinates, etc. TMtransceiver, altitude sensor components (e.g., an altimeter or barometer that detects air pressure, from which altitude can be derived), orientation sensor components (e.g., a magnetometer), etc. Such positioning system coordinates can also be received from the mobile device 714 via the low-power wireless connection 1112 and the high-speed wireless connection 1114 via the low-power wireless circuit system 1134 or the high-speed wireless circuit system 1132.
[0166] As mentioned above, to detect or track expressions, gestures, movements, or other actions, according to some examples, the head wearable device 716 can execute a machine learning model, such as an object tracking model. The machine learning model can be executed at the head wearable device 716, at a host computer, or at an interactive server.
[0167] Typically, once a machine learning model has been built, e.g., as referenced Figure 4 or Figure 5 As described, the model may be sent to or otherwise made available to a user device (e.g., mobile device 714, head wearable device 716, or computer client device 718) to implement or facilitate such detection or tracking functionality.
[0168] Machine Architecture
[0169] Figure 121 is a diagrammatic representation of a machine 1200 within which instructions 1202 (e.g., software, programs, applications, applet, apps, or other executable code) may be executed for causing the machine 1200 to perform any one or more of the methodologies discussed herein. For example, the instructions 1202 may cause the machine 1200 to perform any one or more of the methodologies described herein. The instructions 1202 transform a general-purpose, unprogrammed machine 1200 into a specialized machine 1200 that is programmed to perform the functions described and illustrated in the manner described. The machine 1200 may operate as a standalone device or may be coupled (e.g., using a network) to other machines. In a networked deployment, the machine 1200 may operate as a server or a client machine in server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine 1200 may include, but is not limited to, a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (STB), a personal digital assistant (PDA), an entertainment media system, a cellular phone, a smartphone, a mobile device, a wearable device (e.g., a smart watch), a smart home device (e.g., a smart appliance), other smart devices, a web appliance, a network router, a network switch, a network bridge, or any machine capable of sequentially or otherwise executing instructions 1202 specifying actions to be taken by the machine 1200. Furthermore, while only a single machine 1200 is shown, the term "machine" should also be construed to include a collection of machines that individually or jointly execute instructions 1202 to perform any one or more of the methods discussed herein. For example, the machine 1200 may include the user system 702 or any of a plurality of server devices forming part of the interactive server system 710. In some examples, the machine 1200 may also include both a client system and a server system, wherein certain operations of a particular method or algorithm are performed on the server side and certain operations of the particular method or algorithm are performed on the client side.
[0170] The machine 1200 may include a processor 1204, a memory 1206, and input / output I / O components 1208 that may be configured to communicate with each other via a bus 1210. In an example, the processor 1204 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) may include, for example, a processor 1212 that executes instructions 1202 and a processor 1214. The term "processor" is intended to include multi-core processors, which may include two or more independent processors (sometimes referred to as "cores") that can execute instructions concurrently. Although Figure 12 Multiple processors 1204 are shown, but the machine 1200 may include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.
[0171] The memory 1206 includes a main memory 1216, a static memory 1218, and a storage unit 1220, all of which are accessible by the processor 1204 via the bus 1210. The main memory 1206, the static memory 1218, and the storage unit 1220 store instructions 1202 that implement any one or more of the methods or functions described herein. The instructions 1202 may also reside, completely or partially, within the main memory 1216, within the static memory 1218, within the machine-readable medium 1222 within the storage unit 1220, within at least one of the processors 1204 (e.g., within a cache memory of the processor), or any suitable combination thereof during execution thereof by the machine 1200.
[0172] The I / O components 1208 may include various components for receiving input, providing output, generating output, sending information, exchanging information, capturing measurements, etc. The specific I / O components 1208 included in a particular machine will depend on the type of machine. For example, a portable machine such as a mobile phone may include a touch input device or other such input mechanism, while a headless server machine will be less likely to include such a touch input device. It should be understood that the I / O components 1208 may include Figure 12Many other components are not shown in the drawings. In various examples, the I / O components 1208 may include user output components 1224 and user input components 1226. The user output components 1224 may include visual components (e.g., displays such as plasma display panels (PDPs), light emitting diode (LED) displays, liquid crystal displays (LCDs), projectors, or cathode ray tubes (CRTs)), acoustic components (e.g., speakers), tactile components (e.g., vibration motors, resistance mechanisms), other signal generators, etc. The user input components 1226 may include alphanumeric input components (e.g., keyboards, touch screens configured to receive alphanumeric input, optical keyboards, or other alphanumeric input components), point-based input components (e.g., mice, touch pads, trackballs, joysticks, motion sensors, or other pointing instruments), tactile input components (e.g., physical buttons, touch screens or other tactile input components that provide location and force of touches or touch gestures), audio input components (e.g., microphones), etc.
[0173] In another example, the I / O component 1208 may include a biometric component 1228, a motion component 1230, an environmental component 1232, or a positioning component 1234, as well as various other components. For example, the biometric component 1228 includes components for detecting expressions (e.g., hand expressions, facial expressions, voice expressions, body postures, or eye tracking), measuring biosignals (e.g., blood pressure, heart rate, body temperature, sweat, or brain waves), identifying people (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or electroencephalogram-based recognition), etc. The motion component 1230 includes an acceleration sensor component (e.g., an accelerometer), a gravity sensor component, and a rotation sensor component (e.g., a gyroscope).
[0174] Any biometric or other personally identifiable information (PII) collected by the biometric component or other data capture component is captured and stored only with the user's approval and is deleted upon user request. In addition, such data may be used for very limited purposes (e.g., identity verification). To ensure the limited and authorized use of biometrics and other PII, access to this data is limited to authorized personnel (if any). Data will not be shared or sold to any third party without the user's explicit consent. In addition, appropriate technical and organizational measures are implemented to ensure the security and confidentiality of this sensitive information.
[0175] Environmental components 1232 include, for example, one or more cameras (with still image / photo and video capabilities), an illumination sensor component (e.g., a photometer), a temperature sensor component (e.g., one or more thermometers that detect ambient temperature), a humidity sensor component, a pressure sensor component (e.g., a barometer), an acoustic sensor component (e.g., one or more microphones that detect background noise), a proximity sensor component (e.g., an infrared sensor that detects nearby objects), a gas sensor (e.g., a gas detection sensor for detecting concentrations of hazardous gases for safety purposes or for measuring pollutants in the atmosphere), or other components that can provide indications, measurements, or signals corresponding to the surrounding physical environment.
[0176] With respect to cameras, user system 702 can have a camera system that includes, for example, a front-facing camera on the front surface of user system 702 and a rear-facing camera on the rear surface of user system 702. The front-facing camera can, for example, be used to capture still images and videos of the user of user system 702 (e.g., “selfies”), which can then be enhanced with the enhancement data (e.g., filters) described above. For example, the rear-facing camera can be used to capture still images and videos in a more conventional camera mode, which are similarly enhanced with the enhancement data. In addition to the front-facing camera and the rear-facing camera, user system 702 can also include a 360° camera for capturing 360° photos and videos.
[0177] Furthermore, the camera system of the user system 702 may include dual rear cameras (e.g., a main camera and a depth-sensing camera), or even triple, quad, or quintuple rear camera configurations on the front and back sides of the user system 702. For example, these multi-camera systems may include a wide-angle camera, an ultra-wide-angle camera, a telephoto camera, a macro camera, and a depth sensor.
[0178] The positioning component 1234 includes a position sensor component (for example, a GPS receiver component), an altitude sensor component (for example, an altimeter or barometer that detects air pressure, and the altitude can be obtained according to the air pressure), an orientation sensor component (for example, a magnetometer), etc.
[0179] Various technologies can be used to implement communications. The I / O components 1208 also include a communications component 1236 that is operable to couple the machine 1200 to a network 1238 or device 1240 via corresponding couplings or connections. For example, the communications component 1236 may include a network interface component or another suitable device that interfaces with the network 1238. In other examples, the communications component 1236 may include a wired communications component, a wireless communications component, a cellular communications component, a near field communications (NFC) component, a Components (e.g. Low energy consumption), Device 1240 may be another machine or any of a variety of peripheral devices (eg, a peripheral device coupled via USB).
[0180] In addition, the communication component 1236 can detect an identifier or include a component operable to detect an identifier. For example, the communication component 1236 can include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting one-dimensional bar codes such as Universal Product Code (UPC) bar codes, multi-dimensional bar codes such as Quick Response (QR) codes, Aztec codes, Data Matrix, Dataglyph, MaxiCode, PDF417, Ultra Code, UCC RSS-2D bar codes, and other optical codes) or an acoustic detection component (e.g., a microphone for identifying an audio signal of a tagged item). In addition, various information can be obtained via the communication component 1236, such as location via Internet Protocol (IP), geolocation via Internet Protocol (IP), location information ... Signal triangulation to obtain location, location obtained by detecting NFC beacon signals that can indicate a specific location, etc.
[0181] Various memories (e.g., main memory 1216, static memory 1218, and memory of processor 1204) and storage unit 1220 may store one or more sets of instructions and data structures (e.g., software) implemented or used by any one or more of the methods or functions described herein. These instructions (e.g., instructions 1202) when executed by processor 1204 cause various operations to implement the disclosed examples.
[0182] Instructions 1202 may be sent or received over network 1238 using a transmission medium via a network interface device (e.g., a network interface component included in communications component 1236) and using any of several well-known transmission protocols (e.g., Hypertext Transfer Protocol (HTTP)). Similarly, instructions 1202 may be sent or received via a coupling (e.g., a peer-to-peer coupling) with device 1240 using a transmission medium.
[0183] Software Architecture
[0184] Figure 1313 is a block diagram 1300 illustrating a software architecture 1302 that can be installed on any one or more of the devices described herein. The software architecture 1302 is supported by hardware, such as a machine 1304 including a processor 1306, memory 1308, and I / O components 1310. In this example, the software architecture 1302 can be conceptualized as a stack of layers, where each layer provides specific functionality. The software architecture 1302 includes layers such as an operating system 1312, libraries 1314, frameworks 1316, and applications 1318. In operation, applications 1318 invoke API calls 1320 through the software stack and receive messages 1322 in response to API calls 1320.
[0185] The operating system 1312 manages hardware resources and provides common services. The operating system 1312 includes, for example, a kernel 1324, services 1326, and drivers 1328. The kernel 1324 serves as an abstraction layer between the hardware and other software layers. For example, the kernel 1324 provides functions such as memory management, processor management (e.g., scheduling), component management, networking, and security settings. Services 1326 can provide other common services to other software layers. Drivers 1328 are responsible for controlling or interfacing with the underlying hardware. For example, drivers 1328 may include display drivers, camera drivers, or Low-power drivers, Flash drivers, serial communication drivers (e.g., USB drivers), drivers, audio drivers, power management drivers, etc.
[0186] The libraries 1314 provide a common low-level infrastructure used by the applications 1318. The libraries 1314 may include a system library 1330 (e.g., a C standard library) that provides functions such as memory allocation functions, string manipulation functions, mathematical functions, etc. In addition, the libraries 1314 may include an API library 1332, such as a media library (e.g., a library for supporting the presentation and manipulation of various media formats, such as Moving Picture Experts Group-4 (MPEG4), Advanced Video Coding (H.264 or AVC), Moving Picture Experts Group Layer-3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) audio codec, Joint Photographic Experts Group (JPEG or JPG), or Portable Network Graphics (PNG)), a graphics library (e.g., an OpenGL framework for rendering graphical content on a display in two dimensions (2D) and three dimensions (3D), a database library (e.g., SQLite providing various relational database functions), a web library (e.g., WebKit providing web browsing functions), etc. The library 1314 may also include various other libraries 1334 to provide numerous other APIs to the application 1318 .
[0187] The framework 1316 provides a common high-level infrastructure used by applications 1318. For example, the framework 1316 provides various graphical user interface (GUI) functions, high-level resource management, and high-level location services. The framework 1316 can provide a wide range of other APIs that can be used by applications 1318, some of which may be specific to a particular operating system or platform.
[0188] In an example, applications 1318 may include a home application 1336, a contacts application 1338, a browser application 1340, a book reader application 1342, a location application 1344, a media application 1346, a messaging application 1348, a game application 1350, and a variety of other applications such as third-party applications 1352. Applications 1318 are programs that perform functions defined in the program. Various programming languages may be used to create one or more of the applications 1318 structured in various ways, such as an object-oriented programming language (e.g., Objective-C, Java, or C++) or a procedural programming language (e.g., C or assembly language). In a specific example, third-party applications 1352 (e.g., those written by an entity other than the vendor of a particular platform using ANDROID) may be used to create a third-party application 1352. TM or IOS TM Software Development Kit (SDK) can be used to develop applications on platforms such as IOS TM 、 Mobile software running on the mobile operating system of the phone or other mobile operating system. In this example, third-party applications 1352 can activate API calls 1320 provided by the operating system 1312 to facilitate the functions described in this article.
[0189] Example
[0190] In view of the above-mentioned implementation of the subject matter, the present application discloses the following list of examples, wherein one feature of an example alone or more than one feature of an example adopted in combination and optionally combined with one or more features of one or more additional examples are additional examples that also fall within the disclosure of the present application.
[0191] Example 1 is a method comprising: accessing input data from a storage component by one or more data loaders in a distributed computer system; sending a training data request by a trainer in the distributed computer system, the trainer being communicatively coupled to the one or more data loaders, and the trainer and the one or more data loaders being executed on separate processors; in response to receiving the training data request by a data loader among the one or more data loaders, performing a data loading task by the data loader, the data loading task comprising preprocessing input batches read from the input data to generate training batches; sending the training batches by the data loader to the trainer; and performing a training task by the trainer, the training task comprising executing a machine learning algorithm using the training batches in a model building process.
[0192] In Example 2, the subject matter of Example 1 includes allocating CPU resources in a distributed computer system to one or more data loaders; and allocating GPU resources in the distributed computer system to the trainers, wherein the GPU resources are allocated such that each data loader does not utilize any of the GPU resources.
[0193] In Example 3, the subject matter of any one of Examples 1 to 2 includes deploying one or more network proxies such that communication between the trainer and the one or more data loaders is implemented via a service mesh.
[0194] In Example 4, the subject matter of Example 3 includes, wherein deploying the one or more network proxies comprises defining a data plane of the service mesh by deploying each network proxy uniquely associated with a trainer or one or more data loaders, each network proxy being communicatively coupled to a control plane of the service mesh.
[0195] In Example 5, the subject matter of Example 4 includes, wherein the one or more data loaders is a plurality of data loaders, defining a one-to-many relationship between the trainer and the data loaders, each data loader being uniquely associated with one of the network agents.
[0196] In Example 6, the subject matter of Example 5 includes sending, by the trainer, a request for training data to each of the plurality of data loaders.
[0197] In Example 7, the subject matter of Example 6 includes, wherein the training data request is sent using an RPC protocol.
[0198] In Example 8, the subject matter of Example 7 includes, wherein the RPC protocol is gRPC.
[0199] In Example 9, the subject matter of any one of Examples 5 to 8 includes, wherein each data loader is a separate instance of a service in the distributed computer system.
[0200] In Example 10, the subject matter of any one of Examples 6 to 9 includes controlling, by the service mesh, traffic between the trainer and the plurality of data loaders.
[0201] In Example 11, the subject matter of any one of Examples 6 to 10 includes routing, by the network proxy, the training data request to the data loader to optimize utilization of the trainer.
[0202] In Example 12, the subject matter of Example 11 includes, wherein routing comprises using an exponentially weighted moving average of response delays to determine which of the data loaders to send each training data request to.
[0203] In Example 13, the subject matter of any one of Examples 7 to 12 includes, wherein the training data request is a service-to-service call, and wherein the network proxy associated with the trainer is configured to load balance the training data request across the plurality of data loaders.
[0204] In Example 14, the subject matter of any one of Examples 4 to 13 includes, wherein the trainer and each data loader are executed by respective pods in the distributed computer system, and each network agent is a proxy container added to the pod of the trainer or data loader associated with the network agent.
[0205] In Example 15, the subject matter of any one of Examples 1 to 14 includes, wherein the one or more data loaders and the trainer are executed on different machines in the distributed computer system.
[0206] In Example 16, the subject matter of any one of Examples 1 to 15 includes implementing, by each data loader, a multi-producer, multi-consumer queue, the implementation comprising performing preprocessing at least partially in parallel with a read task, the read task comprising obtaining, by the data loader, one or more input batches from a storage component and adding the one or more input batches to a preprocessing queue of the data loader.
[0207] In Example 17, the subject matter of any one of Examples 1 to 16 includes, wherein the input data comprises at least one of hand detection data, hand tracking data, gesture detection data, or gesture tracking data.
[0208] In Example 18, the subject matter of Example 17 includes, wherein the trainer is configured to send multiple requests for additional training data to enable the training task to be iterated using multiple training batches generated by one or more data loaders from different input batches, the method including generating, by the trainer, the object tracking model based on the results of the training task in the model building process.
[0209] Example 19 is a distributed computing system comprising: one or more processors; and a non-transitory computer-readable storage medium comprising instructions, which, when executed by the one or more processors, cause the one or more processors to perform operations, the operations comprising: accessing input data from a storage component by one or more data loaders in the distributed computer system; sending a training data request by a trainer in the distributed computer system, the trainer being communicatively coupled to the one or more data loaders, and the trainer and the one or more data loaders being executed on separate processors; in response to receiving the training data request by a data loader among the one or more data loaders, performing a data loading task by the data loader, the data loading task comprising preprocessing an input batch read from the input data to generate a training batch; sending the training batch by the data loader to the trainer; and performing a training task by the trainer, the training task comprising executing a machine learning algorithm using the training batch in a model building process.
[0210] Example 20 is a machine-readable non-transitory storage medium having instruction data that can be executed by a machine to cause the machine to perform operations, the operations including: accessing input data from a storage component by one or more data loaders in a distributed computer system; sending a training data request by a trainer in the distributed computer system, the trainer being communicatively coupled to the one or more data loaders, and the trainer and the one or more data loaders executing on separate processors; in response to receiving the training data request by a data loader among the one or more data loaders, performing a data loading task by the data loader, the data loading task including preprocessing input batches read from the input data to generate training batches; sending the training batches by the data loader to the trainer; and performing a training task by the trainer, the training task including using the training batches to execute a machine learning algorithm in a model building process.
[0211] Example 21 is at least one machine-readable medium comprising instructions that, when executed by processing circuitry, cause the processing circuitry to perform operations to implement any of Examples 1-20.
[0212] Example 22 is an apparatus comprising means for implementing any one of Examples 1 to 20.
[0213] Example 23 is a system for implementing any one of Examples 1 to 20.
[0214] Example 24 is a method of implementing any one of Examples 1 to 20.
[0215] in conclusion
[0216] According to examples of the present disclosure, the distributed framework addresses the technical challenge of resource bottlenecks by separating CPU or IO (input / output) intensive workloads (such as data reading and data enhancement / preprocessing) from GPU intensive model training. In some examples, these workloads can be separated and executed by different processors and / or machines (virtual or physical). The machines can be separate and remote from each other.
[0217] GPUs can be designed to perform the complex mathematical and geometric calculations required for graphics rendering. Therefore, in some examples, such as when creating an object tracking model, the architecture of the examples of the present disclosure improves the utilization of GPU resources while ensuring that CPU and memory resources are used as expected, thereby reducing machine learning model training or overall build time.
[0218] While the examples described herein focus on separating CPU and GPU workloads, according to other examples, other processing units such as TPUs (tensor processing units) may be deployed in place of (or in addition to) GPUs without departing from the present disclosure. (TPUs are application-specific integrated circuits designed to accelerate artificial intelligence calculations and algorithms.)
[0219] When distributing and horizontally scaling data loading tasks as described herein, a technical challenge that may arise is that requests from trainers (clients) must be routed efficiently, rather than, for example, being routed to only one data loader or to a data loader that is "busier" than other data loaders. According to examples of the present disclosure, this challenge is addressed by implementing a service mesh in which proxies handle both incoming and outgoing calls, routing traffic to optimize or improve resource utilization. The service mesh includes a control plane that is called by the data plane (defined by the proxies) to inform the data plane of its behavior, and the control plane provides an interface that allows users to modify and inspect the behavior of the data plane.
[0220] Although the examples described herein focus on object tracking models, the techniques and methods according to the examples of this disclosure can also be applied to the construction of other types of machine learning models.
[0221] As used in this disclosure, phrases of the form "at least one of A, B, or C," "at least one of A, B, and C," etc., should be interpreted as selecting at least one from the group consisting of "A, B, and C." Unless expressly stated otherwise in connection with a specific example in this disclosure, this wording does not mean "at least one of A, at least one of B, and at least one of C." As used in this disclosure, the example "at least one of A, B, or C" would encompass any of the following selections: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, and {A, B, C}.
[0222] Unless the context clearly requires otherwise, throughout the specification and claims, the words "include," "comprising," and the like should be interpreted as inclusive rather than exclusive or exhaustive; that is, as "including but not limited to." As used herein, the terms "connect," "couple," or any variation thereof mean any direct or indirect connection or coupling between two or more elements; the coupling or connection between elements may be physical, logical, or a combination thereof. In addition, when used in this application, the words "herein," "above," "below," and words of similar meaning refer to this application as a whole and not to any particular part of this application. Where the context permits, words using the singular or plural number may also include the plural or singular number, respectively. When referencing a list of two or more items, the word "or" encompasses all of the following interpretations of the word: any item in the list, all of the items in the list, and any combination of the items in the list. Similarly, when referencing a list of two or more items, the term "and / or" encompasses all of the following interpretations of the word: any item in the list, all of the items in the list, and any combination of the items in the list.
[0223] Glossary
[0224] "Carrier signal" refers to any intangible medium, such as a digital or analog communication signal, that is capable of storing, encoding, or carrying instructions for execution by a machine, or other intangible medium that facilitates communication of such instructions. Instructions may be sent or received over a network using a transmission medium via a network interface device.
[0225] "Client device" refers to any machine that interfaces with a communications network to obtain resources from one or more server systems or other client devices, for example. A client device may be, but is not limited to, a mobile phone, a desktop computer, a laptop computer, a portable digital assistant (PDA), a smartphone, a tablet computer, an ultrabook, a netbook, a laptop computer, a multiprocessor system, a microprocessor-based or programmable consumer electronics product, a game console, a set-top box, or any other communications device that a user may use to access a network.
[0226] "Communications network" refers to, for example, one or more parts of a network, which may be an ad hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless LAN (WLAN), a wide area network (WAN), a wireless WAN (WWAN), a metropolitan area network (MAN), the Internet, a part of the Internet, a part of the Public Switched Telephone Network (PSTN), a Plain Old Telephone Service (POTS) network, a cellular telephone network, a wireless network, A network, other type of network, or a combination of two or more such networks. For example, a network or a portion of a network may include a wireless network or a cellular network, and the coupling may be a code division multiple access (CDMA) connection, a global system for mobile communications (GSM) connection, or other type of cellular or wireless coupling. In this example, the coupling may implement any of various types of data transmission technologies, such as single carrier radio transmission technology (1xRTT), evolution data optimized (EVDO) technology, general packet radio service (GPRS) technology, enhanced data rates for GSM evolution (EDGE) technology, the Third Generation Partnership Project (3GPP) including 3G, fourth generation wireless (4G) networks, universal mobile telecommunications system (UMTS), high speed packet access (HSPA), worldwide interoperability for microwave access (WiMAX), long term evolution (LTE) standards, other data transmission technologies defined by various standards setting organizations, other long distance protocols, or other data transmission technologies. "Communication network" and "communication network" may be used interchangeably.
[0227] "Component" refers to a logic or device, a physical entity, for example, having the following boundaries, which are defined by function or subroutine calls, branch points, APIs, or other technical definitions that provide partitioning or modularization of specific processing or control functions. A component can be combined with other components via its interface to perform machine processing. A component can be a packaged functional hardware unit designed for use with other components, and a part of a program that generally performs a specific function of a related function. A component can constitute a software component (for example, a code implemented on a machine-readable medium) or a hardware component. A "hardware component" is a tangible unit that can perform certain operations and can be configured or arranged in a certain physical manner. In various example embodiments, one or more computer systems (for example, a stand-alone computer system, a client computer system, or a server computer system) or one or more hardware components (for example, a processor or a processor group) of a computer system can be configured by software (for example, an application or an application part) to operate to perform certain operations as described herein. Hardware components can also be implemented mechanically, electronically, or in any suitable combination thereof. For example, a hardware component can include a dedicated circuit system or logic that is permanently configured to perform certain operations. The hardware component can be a special-purpose processor, such as a field programmable gate array (FPGA) or an application-specific integrated circuit (ASIC). The hardware component can also include a programmable logic or circuit system that is temporarily configured to perform certain operations by software. For example, the hardware component can include software executed by a general-purpose processor or other programmable processor. Once configured by such software, the hardware component becomes a specific machine (or a specific component of a machine), which is uniquely customized to perform the configured function and is no longer a general-purpose processor. It will be appreciated that it can be decided whether to mechanically implement the hardware component in a dedicated and permanently configured circuit system or in a temporarily configured circuit system (e.g., configured by software) for cost and time considerations. Therefore, the phrase "hardware component" (or "hardware-implemented component") should be understood to include tangible entities, i.e., entities that are physically constructed, permanently configured (e.g., hardwired) or temporarily configured (e.g., programmed) to operate in some way or perform certain operations described herein. Considering an embodiment in which the hardware component is temporarily configured (e.g., programmed), it is not necessary to configure or instantiate each hardware component in the hardware component at any one time. For example, where a hardware component includes a general-purpose processor that is configured by software to become a special-purpose processor, the general-purpose processor can be configured as different special-purpose processors (e.g., including different hardware components) at different times. The software configures a specific processor or processors accordingly, such as to constitute a specific hardware component at one time and to constitute a different hardware component at a different time. A hardware component can provide information to other hardware components and receive information from other hardware components.Therefore, the hardware components described can be considered to be coupled in communication. In the case of multiple hardware components being present at the same time, communication can be achieved by signal transmission between or among two or more hardware components (for example, by appropriate circuits and buses). In the embodiment in which multiple hardware components are configured or instantiated at different times, communication between such hardware components can be achieved, for example, by storing information in a memory structure accessible to multiple hardware components and retrieving information in the memory structure. For example, a hardware component can perform an operation, and the output of the operation is stored in a memory device coupled in communication with it. Then, other hardware components can access the memory device at a subsequent time to retrieve the stored output and process it. The hardware component can also initiate communication with an input device or an output device, and can operate on resources (for example, a collection of information). The various operations of the example methods described herein can be performed at least in part by temporarily configuring (for example, by software) or permanently configuring one or more processors to perform related operations. Whether it is temporarily configured or permanently configured, such a processor can constitute a processor-implemented component that operates to perform one or more operations or functions described herein. As used herein, a "processor-implemented component" refers to a hardware component implemented using one or more processors. Similarly, the method described in this article can be implemented at least in part by a processor, wherein specific one or more processors are examples of hardware. For example, at least some of the various operations of the method can be performed by one or more processors or the components implemented by the processor. In addition, one or more processors can also operate to support the execution of related operations in a "cloud computing" environment or operate as "software as a service" (SaaS). For example, at least some operations in the operation can be performed by a computer group (as an example of a machine including a processor), wherein these operations can be accessed via a network (for example, the Internet) and via one or more appropriate interfaces (for example, API). The execution of some operations can be distributed between processors, not only resident in a single machine, but deployed across multiple machines. In some example embodiments, a processor or the components implemented by the processor can be located in a single geographical location (for example, in a home environment, an office environment or a server farm). In other example embodiments, a processor or the components implemented by the processor can be distributed across multiple geographical locations.
[0228] "Computer-readable storage media" refers to, for example, both machine storage media and transmission media. Thus, these terms encompass both storage devices / media and carrier waves / modulated data signals. The terms "machine-readable medium," "computer-readable medium," and "device-readable medium" mean the same thing and may be used interchangeably in this disclosure.
[0229] “Machine storage media” refers to, for example, a single or multiple storage devices and media (e.g., centralized or distributed databases, and associated caches and servers) that store executable instructions, routines, and data. Thus, the term should be taken to include, but is not limited to, solid-state memory and optical and magnetic media, including memory internal or external to the processor. Specific examples of machine storage media, computer storage media, and device storage media include: non-volatile memory, including, for example, semiconductor memory devices such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), FPGAs, and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and (and DVD-ROM disks. The terms “machine storage media,” “device storage media,” and “computer storage media” mean the same thing and are used interchangeably in this disclosure. The terms “machine storage media,” “computer storage media,” and “device storage media” expressly exclude carrier waves, modulated data signals, and other such media, at least some of which are encompassed by the term “signal media.”
[0230] “Non-transitory computer-readable storage medium” refers to a tangible medium that is capable of storing, encoding, or carrying instructions for execution by a machine, for example.
[0231] A "processor" refers to any circuit or virtual circuit (a physical circuit simulated by logic executed on an actual processor) that manipulates data values, e.g., according to control signals (e.g., "commands," "opcodes," "machine codes," etc.), and produces corresponding output signals that are applied to operate a machine. A processor may be, for example, a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), or any combination thereof. A processor may also be a multi-core processor having two or more independent processors (sometimes referred to as "cores") that can execute instructions simultaneously.
[0232] "Signal medium" refers to any intangible medium that is capable of storing, encoding, or carrying instructions for execution by a machine, and includes digital or analog communication signals or other intangible media that facilitate the communication of software or data. The term "signal medium" should be construed to include any form of modulated data signal, carrier wave, etc. The term "modulated data signal" means a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. The terms "transmission medium" and "signal medium" mean the same thing and may be used interchangeably in this disclosure.
[0233] A "user device" refers to a device that is, for example, accessed, controlled, or owned by a user and with which the user interacts to perform actions or interact with other users or computer systems.
Claims
1. A method comprising: accessing input data from the storage component by one or more data loaders in the distributed computer system; sending a training data request by a trainer in the distributed computer system, the trainer being communicatively coupled to the one or more data loaders, and the trainer and the one or more data loaders executing on separate processors; In response to receiving the training data request by a data loader among the one or more data loaders, executing a data loading task by the data loader, the data loading task including pre-processing an input batch read from the input data to generate a training batch; The data loader sends the training batch to the trainer; as well as A training task is performed by the trainer, wherein the training task includes executing a machine learning algorithm using the training batch in a model building process.
2. The method according to claim 1, comprising: allocating central processing unit (CPU) resources in the distributed computer system to the one or more data loaders; as well as Graphics processing unit (GPU) resources in the distributed computer system are allocated to the trainer, wherein the GPU resources are allocated such that each data loader does not utilize any of the GPU resources.
3. The method according to claim 1, comprising: One or more network proxies are deployed such that communication between the trainer and the one or more data loaders is achieved via a service mesh.
4. The method according to claim 3, wherein: Deploying one or more network agents includes defining a data plane of the service mesh by deploying each network agent uniquely associated with the trainer or the one or more data loaders, each network agent communicatively coupled to a control plane of the service mesh.
5. The method according to claim 4, wherein The one or more data loaders is a plurality of data loaders defining a one-to-many relationship between the trainer and the data loaders, each data loader being uniquely associated with one of the network agents.
6. The method according to claim 5, comprising: A training data request is sent by the trainer to each of the plurality of data loaders.
7. The method according to claim 6, wherein: The training data request is sent using a remote procedure call (RPC) protocol.
8. The method according to claim 7, wherein: The RPC protocol is gRPC.
9. The method according to claim 5, wherein: Each data loader is a separate instance of a service in the distributed computer system.
10. The method according to claim 6, comprising: Traffic between the trainer and the plurality of data loaders is controlled by the service grid.
11. The method according to claim 6, comprising: The training data request is routed to the data loader by the network proxy to optimize utilization of the trainer.
12. The method according to claim 11, wherein The routing includes using an exponentially weighted moving average of response delays to determine which of the data loaders to send each training data request.
13. The method according to claim 7, wherein: The training data request is a service-to-service call, and the network proxy associated with the trainer is configured to load balance the training data request across the plurality of data loaders.
14. The method according to claim 4, wherein: The trainer and each data loader are executed by corresponding pods in the distributed computer system, and each network agent is an agent container added to the pod of the trainer or the data loader associated with the network agent.
15. The method according to claim 1, wherein The one or more data loaders and the trainer execute on different machines in the distributed computer system.
16. The method according to claim 1, comprising: A multi-producer, multi-consumer queue is implemented by each data loader, the implementation comprising performing the preprocessing at least partially in parallel with a read task, the read task comprising: retrieving one or more input batches from the storage component by the data loader and adding the one or more input batches to a preprocessing queue of the data loader.
17. The method according to claim 1, wherein The input data includes at least one of hand detection data, hand tracking data, gesture detection data, or gesture tracking data.
18. The method according to claim 17, wherein The trainer is configured to send multiple requests for additional training data to allow the training task to be iterated using multiple training batches generated by the one or more data loaders from different input batches, and the method includes generating, by the trainer, an object tracking model based on the results of the training task in the model building process.
19. A distributed computing system comprising: one or more processors, and A non-transitory computer-readable storage medium comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising: accessing input data from the storage component by one or more data loaders in the distributed computer system; sending a training data request by a trainer in the distributed computer system, the trainer being communicatively coupled to the one or more data loaders, and the trainer and the one or more data loaders executing on separate processors; In response to receiving the training data request by a data loader among the one or more data loaders, executing a data loading task by the data loader, the data loading task including pre-processing an input batch read from the input data to generate a training batch; sending, by the data loader, the training batch to the trainer; and A training task is performed by the trainer, wherein the training task includes executing a machine learning algorithm using the training batch in a model building process.
20. A machine-readable non-transitory storage medium having instruction data, the instruction data being executable by a machine to cause the machine to perform operations, the operations comprising: accessing input data from the storage component by one or more data loaders in the distributed computer system; sending a training data request by a trainer in the distributed computer system, the trainer being communicatively coupled to the one or more data loaders, and the trainer and the one or more data loaders executing on separate processors; In response to receiving the training data request by a data loader among the one or more data loaders, executing a data loading task by the data loader, the data loading task including pre-processing an input batch read from the input data to generate a training batch; The data loader sends the training batch to the trainer; as well as A training task is performed by the trainer, wherein the training task includes executing a machine learning algorithm using the training batch in a model building process.