Task management using image capture monitoring of user actions
A system using video surveillance and machine learning to analyze user actions and provide real-time feedback addresses the challenge of interacting with user interfaces during tasks, enhancing efficiency and correcting errors in manufacturing environments.
Patent Information
- Application Number
- JP2025539400
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-04-26
- Filing Date
- 2023-12-28
- Publication Date
- 2026-01-16
AI Technical Summary
Users face difficulties in interacting with user interfaces while performing tasks, especially in environments requiring both hands, leading to delays and inefficiencies, particularly in manufacturing settings where users may be wearing gloves.
A system that monitors user actions through video surveillance, uses machine learning to analyze and identify task sequences, and provides real-time feedback and instructions to guide users through tasks, correcting errors and optimizing task execution.
Enhances task efficiency by providing real-time guidance, correcting errors, and optimizing task sequences, thereby reducing completion times and improving user performance.
Smart Images

Figure 2026501673000001_ABST
Abstract
Description
[Technical Field]
[0001] Incorporation by Reference, Disclaimer U.S. Patent Application No. 18 / 307,448, filed April 26, 2023, and U.S. Patent Application No. 63 / 478,326, filed January 3, 2023, are each incorporated herein by reference. Applicant hereby withdraws any disclaimer of claim scope in the parent application or its prosecution history, and reports to the USPTO that the claims in this application may be broader than any claims in the parent application.
[0002] Technical Field The present disclosure relates to managing the execution of user tasks by monitoring images of user locations to monitor user actions. In particular, the present disclosure relates to analyzing user actions recorded using a video surveillance application to determine task characteristics associated with the user actions. [Background technology]
[0003] background In many environments, a user must perform a set of steps to complete a particular task. For example, a user can perform a set of steps to assemble a component in a manufacturing environment. The set of steps can include interacting with equipment in a workspace. The system can provide feedback for completing the set of steps. For example, the system can provide an indication of what actions the user must perform to complete the operations associated with the particular task. In many cases, it is difficult for a user to interact with a user interface while completing the set of operations. For example, in a manufacturing environment, a user may be wearing gloves. Additionally, because the sequence of operations may require both hands, interrupting the sequence to interact with the user interface delays task completion.
[0004] The approaches described in this section are approaches that could be pursued, but not necessarily approaches that have been previously conceived or pursued. Thus, unless otherwise indicated, it should not be assumed that any of the approaches described in this section qualify as prior art merely by virtue of their inclusion in this section.
[0005] Embodiments are illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings. It should be noted that references to "one" or "an" embodiment in this disclosure do not necessarily refer to the same embodiment, but rather mean at least one. [Brief explanation of the drawings]
[0006] [Figure 1] FIG. 1 illustrates a system according to one or more embodiments. [Figure 2] FIG. 1 illustrates an example set of actions for presenting user task actions based on detected user actions, according to one or more embodiments. [Figure 3] FIG. 1 illustrates an example set of operations for training a machine learning model to classify recorded user actions as behaviors associated with a task, according to one or more embodiments. [Figure 4A] FIG. 1 illustrates an exemplary embodiment. [Figure 4B] FIG. 1 illustrates an exemplary embodiment. [Figure 4C] FIG. 1 illustrates an exemplary embodiment. [Figure 4D] FIG. 1 illustrates an exemplary embodiment. [Figure 4E] FIG. 1 illustrates an exemplary embodiment. [Figure 4F] FIG. 1 illustrates an exemplary embodiment. [Figure 5] FIG. 1 is a block diagram illustrating a computer system according to one or more embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0007] Detailed Description In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding. One or more embodiments may be practiced without these specific details. Features described in one embodiment may be combined with features described in a different embodiment. In some instances, well-known structures and devices are described with reference to block diagram form in order to avoid unnecessarily obscuring the present invention.
[0008] 1.Overview 2. System Architecture 3. Presenting user task actions based on detected user actions 3.1 Starting a branch task 3.2 Triggering the recording of statistics associated with a task 3.3 Monitoring task metrics based on detected user actions 4. Machine learning model training 5. Exemplary Embodiments 6. Computer Networks and Cloud Networks 7. Miscellaneous, Extensions 8. Hardware Overview 1.Overview One or more embodiments display instructions associated with performing a next action of a task in response to detecting completion of a current action of a task through real-time analysis of a video stream of a user completing the current action of the task. The system first displays instructions for performing the current action of the task in a set of actions. Simultaneously with displaying the instructions for the current action, the system continuously analyzes the user's video stream in real time while the user performs an action to complete the current action. Based on the continuous analysis of the video stream, the system detects that the user has completed the current action. The system uses the detected completion of the current action of the task as a trigger to switch from displaying instructions for the current action to displaying instructions for the next action in the set of actions.
[0009] One or more embodiments generate predictive statistics based on an analysis of a video stream of a user performing a set of actions for a task. The system analyzes the user's video stream to detect start and / or end times corresponding to the execution of the actions. The start and / or end times of one or more actions are used to generate a prediction corresponding to the set of actions. As an example, the system predicts a delay in the completion of the complete set of actions for the task based on a detected start time of a current action of the task, an estimated completion time of the current action, actions to be executed subsequent to the current action, and an estimated completion time of the task. As another example, the system predicts a delay in the completion of the complete set of actions based on a detected end time of the current action and an estimated completion time of actions to be executed subsequent to the current action. The system can store the start and / or end times for any statistical calculations associated with the execution of the actions.
[0010] One or more embodiments perform modifications to the set of tasks to be completed based on real-time analysis of a video stream of a user performing one or more of the set of tasks. In one example, the system selects a corrective action to be performed in response to determining that the user performed one or more tasks incorrectly and / or out of sequence. The system can insert and present instructions for performing the corrective action. The system can reorder the remaining tasks based on the user's performance of the incorrect and / or out-of-sequence task. The corrective action can include undoing a performed action, reperforming a performed action, and / or compensatory action that counteracts the negative impact of the incorrect and / or out-of-sequence task. Alternatively or additionally, the system can present an alert in response to determining that the user performed one or more tasks incorrectly and / or out-of-sequence.
[0011] One or more embodiments record statistics associated with a user's performance of a task based on the monitoring of user actions. For example, the system may identify a time associated with a user action corresponding to an initial operation of the task. The system may identify another time associated with a user action corresponding to completing a final operation of the task. The system may record the total time required by the user to perform the task. The system may record statistics as metrics associated with the performance of the task. For example, the system may maintain a record of how long it takes a user to perform a particular task. The system may maintain metrics that compare the user's performance to that of other users or the same user over time. One or more embodiments use the recorded statistics to generate subsequent predictions for the operations and completion times of the task.
[0012] One or more exemplary embodiments use image recognition to identify a specific user in a workspace. The system identifies a set of tasks associated with the user and equipment in the workspace. The system monitors user actions in the workspace in real time via image recording. The system identifies actions associated with the user's actions. The system presents an image to the user on a user interface display. The image is associated with additional actions of the specific task that correspond to the user's actions detected by the system.
[0013] One or more embodiments specify a task associated with a sequence of operations in a manufacturing environment. The system specifies actions associated with steps required to perform the task. The system generates instructions to direct a user to perform the actions associated with the steps. The system monitors user actions in real time via image recording to determine when actions associated with the task have started and / or completed.
[0014] One or more embodiments train a machine learning model to identify user actions corresponding to task operations. The system trains the machine learning model using a training dataset of image data corresponding to an operator's body positions and sequences of body positions. The image data can include the operator's body positions relative to a workspace. Based on identifying the user's body positions in real time via image recordings, the system identifies operations being performed and / or completed by the user.
[0015] One or more embodiments described and / or claimed herein may not be included in this General Summary section.
[0016] 2. System Architecture System 100 includes workspace monitoring platform 110, which monitors user actions in workspace 120. Workspace monitoring platform 110 acquires image data from image capture device 118, such as a video camera. Image capture device 110 captures images of user 140 in workspace 120. Workspace 120 is a designated geographic area in which user 140 performs actions to complete a task. In the example shown in FIG. 1 , workspace 120 includes resources 123, 124, and 125, as well as operational equipment 121 and 122. Resources 123, 124, and 125 can be used by user 140 during the course of performing the task. For example, resources can be part of component 126 that are added to component 126 during a manufacturing process performed by user 140. Additionally or alternatively, resources 123-125 can be provided to operational equipment 121 and 122 to perform actions. For example, the operating device 121 may be a 3D printer, and the resources 123 may include resin used as a material to perform the 3D printing operation.
[0017] The image analysis engine 111 analyzes the image data to identify user actions. According to one embodiment, the image analysis engine 111 identifies user actions by providing the image data to a machine learning model 112. The machine learning engine 115 trains the model 112 to identify user actions in the image data.
[0018] The action identification engine 113 identifies actions associated with the identified user action. For example, the system may determine that a user pressing a particular button on equipment 121 in the workspace corresponds to an action for performing a particular operation with the equipment. Additionally, the system may determine that a set of user actions, such as selecting a resource 123 and attaching the resource 123 to a component 126, corresponds to a particular action associated with assembling a manufactured component. According to one embodiment, the action identification engine 113 identifies an action 152 associated with the identified user action 153 by referencing a mapping of user actions to actions. The mapping may be stored as a table, as a directed acyclic graph (DAG) tree, or in any other format.
[0019] In some examples, one or more elements of the machine learning engine 115 can use a machine learning algorithm to identify one or both of user actions and task behaviors based on image data content. A machine learning algorithm is an algorithm that can iterate using a set of training data to learn a target model f that best maps a set of input variables to output variables. The machine learning algorithm can include supervised and / or unsupervised components. Various types of algorithms can be used, such as linear regression, logistic regression, linear discriminant analysis, classification and regression trees, naive Bayes, k-nearest neighbors, learning vector quantization, support vector machines, bagging and random forests, boosting, backpropagation, and / or clustering.
[0020] In an embodiment, the set of training data includes a dataset and associated labels. The dataset is associated with input variables for the target model f (e.g., image data including user position data in a workspace). The associated labels are associated with output variables of the target model f (e.g., specified user actions, specified task behaviors associated with user actions). The training data can be updated, for example, based on feedback regarding the accuracy of the current target model f. The updated training data is fed back to a machine learning algorithm, which then updates the target model f.
[0021] The machine learning algorithm generates a target model f such that the target model f best fits the dataset of training data to the labels in the training data. Additionally, or alternatively, the machine learning algorithm generates a target model f such that when the target model f is applied to the dataset of training data, the greatest number of results determined by the target model f match the labels in the training data.
[0022] In embodiments, the machine learning algorithm can iterate to learn user actions and / or behaviors associated with image data of the user in the workspace. In embodiments, the set of training data includes image data of user positions in the workspace. The image data is associated with labels indicative of the user actions and / or behaviors.
[0023] In an example, the system first trains a neural network using a historical data set. Training the neural network can include generating n hidden layers for the neural network and functions / weights to be applied to each hidden layer to calculate the next hidden layer. Training can further include determining functions / weights to be applied to the final nth hidden layer, thereby calculating a final label or prediction for the data points.
[0024] Training a neural network involves (a) acquiring a dataset, (b) iteratively applying the training dataset to the neural network to generate labels for data points in the training dataset, and (c) adjusting weights and offsets associated with equations constituting neurons of the neural network based on a loss function that compares values associated with the generated labels to values associated with test labels. Neurons of the neural network include activation functions that specify limits on the values output by the neurons. The activation functions may include differentiable nonlinear activation functions such as a rectified linear activation (ReLU) function, a logistic function, or a hyperbolic tangent function. Each neuron receives the values of each neuron from the previous layer, applies a weight to each value from the previous layer, and applies one or more offsets to the combined values from the previous layer. The activation functions constrain the range of possible output values from the neuron. A sigmoid activation function converts a neuron's value to a value between 0 and 1. A ReLU activation function converts a neuron's value to 0 if the neuron's value is negative and converts a neuron's value to an output value if the neuron's value is positive. The ReLU activation function can also be scaled to output values between 0 and 1. For example, after applying a weight and offset value to the value from the previous layer of one neuron, the system may scale the neuron value to a value between -1 and +1. The system may then apply a ReLU-type activation function to generate a neuron output value between 0 and 1. The system trains the neural network using training, test, and validation data sets until the labels generated by the trained neural network are within a specified accuracy level, such as 98% accuracy.
[0025] The task monitoring engine 114 identifies a task associated with the identified operation. A task 151 is a set of one or more operations. A task can specify operations that should be performed in a specific sequential order. A task can also specify operations that are not required to be performed in a specific sequential order. The task monitoring engine 114 can identify a task based on detecting a specific operation or a specific set of operations. According to one embodiment, the task monitoring engine 114 identifies a task associated with the identified operation or set of operations by referencing a mapping of tasks 151 to operations 152. The mapping can be stored as a table, as a directed acyclic graph (DAG) tree, or in any other format.
[0026] The task monitoring engine 114 determines whether a user action identified by the image analysis engine 111 corresponds to a particular operation of a task. For example, if a task includes operations A, B, and C, and the task monitoring engine 114 determines that the user performed operation A, the task monitoring engine 114 monitors the next user action to determine whether the next user action corresponds to operation B or operation C.
[0027] While user 140 is performing a particular task, task monitoring engine 114 displays subsequent actions performed by user 140 on user interface 117. For example, if task monitoring engine 114 determines that the user has completed action B in the set of actions, task monitoring engine 114 can cause user interface 117 to display an image corresponding to the next action in the set of actions. Additionally, if task monitoring engine 114 determines that a detected user action corresponds to a system action trigger, task monitoring engine 114 initiates a system action. For example, task monitoring engine 114 can determine that the detected user action does not correspond to an action in the currently executing task. Task monitoring engine 114 can perform actions such as generating an error notification, generating a prompt to add a new action to the task that corresponds to the detected user action, or initiating a branching task to correct the user error.
[0028] A user 140 can interact with the user interface 117 to perform a task. For example, the user 140 can interact with the user interface 117 to select a task to be performed. Additionally or alternatively, the user 140 can interact with the user interface 117 to review changes to a particular task. The user interface 117 includes a display device that presents instructions to the user 140 for completing an action. The instructions may include textual and / or graphical content. According to one embodiment, the task monitoring engine 114 controls the user interface 117 to present instructions to the user 140 for performing the next action in a set of actions.
[0029] The task metric recording engine 116 records metrics 154 associated with a task. Examples of task metrics include completion time, error detection rate, other anomaly detection rate, equipment utilization rate, and equipment failure detection rate. For example, the system can detect the start time of a task or operation and the end time of the task or operation. The system can store the task or operation completion time. The system can compare the completion time with a predicted completion time based on the same user or other users. The system can provide the user or other entity with efficiency metrics for a particular user, a particular operation, or a particular task based on the comparison.
[0030] The workspace metrics application platform 130 processes the metrics generated by the task metrics recording engine 116 to perform additional functions associated with the workspace 120. For example, the workspace metrics application platform 130 may schedule the operating equipment 121 for maintenance based on the detection of a particular usage level or a particular error rate level associated with operations performed with the equipment 121. The workspace metrics application platform 130 may generate operator rankings based on completion time, error rate, or productivity metrics associated with one or more users.
[0031] In one or more embodiments, system 100 may include more or fewer components than those shown in Figure 1. The components shown in Figure 1 may be local or remote from each other. The components shown in Figure 1 may be implemented in software and / or hardware. Each component may be distributed across multiple applications and / or machines. Multiple components may be combined into a single application and / or machine. Operations described with respect to one component may instead be performed by another component.
[0032] Further embodiments and / or examples relating to computer networks are described below in Section 4 entitled "Computer Networks and Cloud Networks."
[0033] In one or more embodiments, data repository 150 is any type of storage unit and / or device for storing data (e.g., a file system, a database, a collection of tables, or any other storage mechanism). Furthermore, data repository 150 may include multiple different storage units and / or devices. The multiple different storage units and / or devices may or may not be of the same type and may or may not be located in the same physical location. Furthermore, data repository 150 may be implemented or executed on the same computing system as workspace monitoring platform 110. Alternatively, or in addition, data repository 150 may be implemented or executed on a computing system separate from workspace monitoring platform 110. Data repository 150 may be communicatively coupled to workspace monitoring platform 110 via a direct connection or via a network.
[0034] The information describing the mapping of actions to operations and operations to tasks may be implemented across any of the components in system 100. However, this information is presented in data repository 150 for purposes of clarity and explanation.
[0035] In one or more embodiments, workspace monitoring platform 110 refers to hardware and / or software configured to perform the operations described herein to capture images of user actions in the workspace and perform operations such as presenting instructions to perform subsequent actions based on detected user actions. An example of an operation for presenting instructions to perform a series of actions based on detection of user actions in image data is described below with reference to FIG. 2.
[0036] In an embodiment, the workspace monitoring platform 110 is implemented in one or more digital devices. The term "digital device" generally refers to any hardware device that includes a processor. A digital device may refer to a physical device or a virtual machine that runs an application. Examples of digital devices include computers, tablets, laptops, desktops, netbooks, servers, web servers, network policy servers, proxy servers, general-purpose machines, specific-function hardware devices, hardware routers, hardware switches, hardware firewalls, hardware network address translators (NATs), hardware load balancers, mainframes, televisions, content receivers, set-top boxes, printers, mobile handsets, smartphones, personal digital assistants ("PDAs"), wireless receivers and / or transmitters, base stations, communication managers, routers, switches, controllers, access points, and / or client devices.
[0037] In one or more embodiments, interface 117 refers to hardware and / or software configured to facilitate communication between a user and workspace monitoring platform 110. Interface 117 renders user interface elements and receives input via user interface elements. Examples of interfaces include graphical user interfaces (GUIs), command line interfaces (CLIs), tactile interfaces, and voice command interfaces. Examples of user interface elements include check boxes, radio buttons, drop-down lists, list boxes, buttons, toggles, text fields, date and time selectors, command lines, sliders, pages, and forms.
[0038] In an embodiment, different components of interface 117 are specified in different languages. The behavior of user interface elements can be specified in a dynamic programming language such as JavaScript. The content of user interface elements is specified in a markup language such as HyperText Markup Language (HTML) or XML User Interface Language (XUL). The layout of user interface elements is specified in a style sheet language such as Cascading Style Sheets (CSS). Alternatively, interface 117 is specified in one or more other languages, such as Java, C, or C++.
[0039] 3. Presenting user task actions based on detected user actions 2 illustrates an example set of operations for presenting user task operations based on the detection of a user action, according to one or more embodiments. One or more of the operations illustrated in FIG. 2 may be modified, rearranged, or omitted entirely. Thus, the particular sequence of operations illustrated in FIG. 2 should not be construed as limiting the scope of one or more embodiments.
[0040] The system identifies a task comprised of a set of actions (operation 202). For example, a user may select a user interface icon associated with the task. Based on the user selection, the system identifies a set of actions to be performed by the user to complete the task. Additionally, the set of actions may further include one or more actions performed by a machine or another user. Alternatively, or in addition, the system may analyze image data of user actions, such as data recorded by a video recorder, to identify the action or set of actions performed by the user. For example, a user may interact with a machine by selecting a component from a bin. The system may recognize the action as belonging to the task "assemble a device module." The task may include six actions, including the action of selecting a component from a bin. According to one exemplary embodiment, a workspace within a manufacturing facility includes one or more pieces of equipment. The system identifies the task comprised of the set of actions based on detecting user actions in the workspace, such as the user interacting with a particular item of equipment, in the recorded video data.
[0041] According to an exemplary embodiment, the system identifies tasks based on detecting patterns of user actions. For example, one task may include cleaning subcomponents. Another task may include assembling subcomponents into a component. The system may identify which task a user is starting based on detecting different patterns of user actions. Detecting cleaning a subcomponent may include detecting the user placing an item on a surface and reaching for a cleaning agent. Detecting assembling subcomponents into a component may include detecting the user placing an item on a surface and reaching for another subcomponent that makes up the component. The system may learn patterns associated with one or more actions that correspond to a particular task.
[0042] The system presents, via a user interface, the operations in the set of operations that make up the task (operation 204). For example, in a set of 10 operations, the system identifies the next operation to be performed in the set of 10 operations. The set of operations may be arranged in a predefined sequence. Alternatively, a task may be completed by performing the operations in multiple different sequences, or in any sequence. One or more operations may be optional for completing the task. Other operations may be required to complete the task. In examples where a set of operations is required to be performed in a particular sequence, the system identifies the next operation in the sequence and presents the next operation in the sequence. In examples where multiple different operations may be performed next (e.g., as the first operation in the task or as an operation following an operation corresponding to an identified user action), the system selects the operation to present based on predetermined criteria. For example, the system may determine that performing the operations in one sequence results in a better execution time than performing the operations in another sequence. The system may present, via the user interface, the operations based on the sequence associated with the better execution time. The system may determine that resources required for one operation are available while resources required for another operation are not. Similarly, the system may predict that execution of operations in a sequence will result in delays due to waiting for resources, such as machines or subcomponents, to become available. In the absence of explicit time-based or resource-based criteria for selecting operations to present, the system may randomly or semi-randomly select operations. For example, the system may select operations to present based on the memory address of a file associated with display data for the operation, a randomly assigned number for the operation, or an alphabetical order associated with the operation name.
[0043] Presenting the actions can include displaying photographs or graphic images of the equipment or component being handled by the user and the actions being performed by the user. For example, the system can display a photo of the test equipment, a subcomponent positioned in a specific location on the test equipment, and a highlighted icon of a button to be pressed to begin testing the subcomponent. According to one exemplary embodiment, presenting the actions includes presenting a live video image of the component or equipment. As the user interacts with the component or equipment, the system can show video images of the user's hands and the component or equipment along with graphic images of the actions being performed with the component or equipment.
[0044] The system detects the performance of a user action (operation 206). According to one exemplary embodiment, the system detects the performance of the user action based on image data of the user in the workspace. The image data can be obtained from one camera or multiple cameras. The system can detect the performance of the user action based on the image data combined with additional data, such as sensor data detecting changes in weights in bins containing subcomponents, energization of equipment, or temperature changes in equipment.
[0045] In one embodiment, the system presents a video image of the equipment or subcomponent associated with the action along with a representation of the action being performed. For example, the action may include placing a subcomponent in a container. The system may provide a video image of the container along with an icon representing the placement of the subcomponent in the container. The user may observe the user performing the presented action via the video image of the container in the workspace. As the user interacts with the container corresponding to the action, the user may see this interaction in the displayed video image.
[0046] The system can detect user actions through an action recognition application. According to one or more embodiments, the system trains a machine learning model to identify user actions. The system can train the machine learning model with a dataset of images showing a user's position within a workspace. The images can include a user interacting with a particular component, such as a manufactured component, and with a particular piece of equipment. The system can train the machine learning model to identify different actions associated with the same component and the same piece of equipment. As an example, one set of images showing one user action can show a user pressing a button on an appliance to set the appliance to a desired setting. Another set of images showing another user action can show a user moving a component to the appliance. Yet another set of images showing another user action can show a user interacting with a particular actuation mechanism that causes the appliance to change the component, such as by soldering a subcomponent onto the component for a particular duration. The system can train the machine learning model to identify distinct sets of images showing distinct user actions.
[0047] The system matches the identified user actions with the actions of one or more tasks. A machine learning model can be trained to match actions with task actions. Alternatively, a machine learning model may be trained to identify actions, and a mapping engine may map actions and sets of actions to actions. For example, one action, setting equipment to a specific setting, may be part of three stored tasks. Each task may be associated with a different manufactured component. The system can identify a set of user actions as corresponding to a set of actions for a particular task by identifying (a) specific actions associated with multiple tasks and (b) another action that corresponds to a detected user action handling a particular manufactured component. The system can identify the combination of actions (a) and (b) as being associated with one particular task. If the user instead selects a different manufactured component (e.g., action (c)), the system can identify the combination of actions (a) and (c) as being associated with a different particular task.
[0048] The system determines whether the detected user action corresponds to a system action trigger (operation 208). In an exemplary embodiment in which a task consists of a set of actions to be performed in a specified sequence, the system may determine that (a) the user selected a task to perform or began performing the task, and (b) the user performed an action within the set of actions that comprise the task, but out of sequence. According to another example, the system may determine that the user performed an action incorrectly. The system may detect a variation in the user's action from the specified action corresponding to the action. Alternatively, the system may detect an abnormal condition of equipment or a manufactured component.
[0049] According to yet another example, the system may determine that the user has performed an action that is not included in the set of actions corresponding to the task. According to yet another example, the system may determine that the user has completed the last action of a particular task. According to yet another example, the system may determine that multiple possible actions can be performed following a previously performed action.
[0050] If the system determines that the user action does not correspond to a system action trigger, the system identifies one or more subsequent actions for the particular task (operation 210). For example, if the observed user action corresponds to an action that is one of a sequence of actions in the task, the system identifies the next action in the sequence of actions. According to another example, the observed user action can correspond to a task action that is one of a set of actions that are not associated with any particular sequence. For example, a task can include actions A, B, and C, which can be completed in any order. When an action is not associated with any particular sequence, the system can select the action to be the subsequent action based on one or more execution criteria.
[0051] For example, the system may calculate which sequences of actions meet certain execution thresholds, such as the likelihood of successful task execution or the shortest time to complete a task. According to another example, the system may select an action as a subsequent action based on analyzing past executions of a task to identify which actions were most frequently selected as subsequent actions by the user or other users. According to yet another example, the system may select an action as a subsequent action based on analyzing one or both of equipment status and manufacturing component status. For example, if a particular piece of equipment is being used by another user, the system may refrain from selecting an action that requires the equipment as a subsequent action. If the equipment requires warm-up time, the system may select an action that initiates warm-up of the equipment, allowing other actions to be performed while the equipment warms up. According to another example, if a particular manufacturing component is temporarily low or out of stock, the system may refrain from selecting a particular action that requires that manufacturing component as a subsequent action.
[0052] The system presents the subsequent actions to the user via the user interface (operation 212). According to one example, the system presents the subsequent actions to the user without intervening instructions from the user. For example, the user does not need to press a button or icon to advance the display of the user interface from one action to the next. Instead, the system automatically advances the display from one action to a subsequent action based on observing the user actions via the image capture device and identifying the actions and tasks associated with the user actions. According to one or more embodiments, the user actions are not user actions that interact with the user interface. Instead, they are actions that interact with equipment and / or components in the workspace.
[0053] According to one embodiment, the system presents the subsequent action via the user interface by displaying instructions for completing the next action. For example, the system may display an icon via the user interface representing the user's body performing a specific action associated with a particular piece of equipment and / or a particular manufacturing component. Additionally or alternatively, the system may display text instructions for performing a user action corresponding to the subsequent action. Additionally or alternatively, the system may present instructions for performing a user action corresponding to the subsequent action via voice or audio transmission.
[0054] According to one or more embodiments, the presentation of the subsequent operation may include a representation of the component that will be manipulated by the user during the operation. The representation may show how the component will be modified. The representation may show the final state of the component after being modified.
[0055] According to one or more embodiments, the system presents a subsequent action based on determining that a previous action has been completed by the user. For example, the system can observe that an action has been completed by the user's body position. Additionally or alternatively, the system may obtain status data from the equipment indicating that the equipment has attained a particular state (e.g., energized, test completed, manufacturing operation completed, etc.) as a result of an action performed by the user. Alternatively, the system may identify the state of a component manipulated by the user. The system can scan the component to determine that a change has been made by the user. For example, the system can identify via video analysis that component A has been attached to component B. As another example, the system can detect a change in weight on the work surface, indicating the presence of component A on the work surface.
[0056] According to an alternative embodiment, the system may suggest a subsequent action based on determining that a previous action was initiated by a user. According to one or more embodiments, the system may suggest a subsequent action based on detecting a particular state of a manufactured component, resources for manufacturing the component, or the state of equipment in the workspace.
[0057] If the system determines that the user action identified in act 206 corresponds to a system action trigger, the system performs the action associated with the trigger (act 214). For example, if the system detects that the user action identified in act 206 was the last action of a task, the system may perform one or more actions associated with completing the task, such as recording the time it took the user to complete the task, recording any errors or anomalies detected during the user's performance of the task, recording the completion of the task in a task management system (which may trigger an option to perform a dependent task in another user's interface), and identifying the next task to be performed by the user. As an example, the system may identify a pattern of tasks typically performed by the user. Upon completion of one task in the pattern, the system may display on the user interface an action to start another task in the task pattern.
[0058] Another system action trigger includes determining that a user has performed an action out of sequence. In response, the system can prompt the user to indicate whether to (a) recommend interrupting the task, (b) recommend resuming the task, (c) generate a notification on a user interface that the detected action is out of sequence, or (d) generate a new sequence of actions for the task based on the user's current sequence for performing the actions. The system can select from among the above system actions (a) through (d) based on one or both of previously generated rules and detected characteristics of the task. For example, if the system determines that a task includes sequential actions A, B, C, and D, the system may detect that the user performs action C before action B, and the user may further determine that such a change in sequence will result in an incorrectly manufactured component. Therefore, the system may determine that the sequence cannot be rearranged and that the system action must include a prompt to interrupt and resume the task. Alternatively, if the system determines that the task includes sequential operations A, B, C, and D, but that performing the operations out of order does not result in known defects in the resulting manufactured component, the system may prompt the user to indicate whether or not to store a new sequence for performing the operations.
[0059] Another system action trigger includes detecting that a user action to be performed does not match a user action associated with a particular action, such as a currently displayed action. For example, the system can detect via an imaging sensor that a user moved an appliance actuator 45 degrees, but the action specifies moving the actuator 90 degrees. The system can (a) recommend interrupting the task, (b) recommend repeating the action, (c) recommend restarting the task, (d) generate a notification on the user interface that the detected user action did not correspond to the displayed user action, or (d) prompt the user to indicate whether to modify the stored set of actions for the action to correspond to the currently detected user action. The system can select from among the above system actions (a) through (d) based on one or both of previously generated rules and detected characteristics of the task and / or action.
[0060] Another system action trigger includes detecting the status of equipment or resources associated with a task. For example, the system can detect that the equipment is in an error state. Alternatively, the system can detect that the equipment is not in an error state but is not in a state specified for a particular task. The system can perform a system action recommending pausing or resuming the operation or task, or performing one or more intervention actions to bring the equipment into a state specified for the task before continuing with the operations specified in the task. Additionally or alternatively, the system can detect an abnormal condition of a manufactured component associated with the task and / or operation. An action can include a user action of placing the component in an analysis device, such as a scanner. Based on the detection of an abnormality in the component, the system can recommend pausing or resuming the task or operation. Alternatively, the system can display one or more recommended actions to correct for the detected abnormality. For example, if the abnormality includes a tube being secured to a flange, the system can recommend an action to tighten the securing device before proceeding with operations to manufacture the component.
[0061] Another system action trigger includes determining that multiple possible actions can be performed following a previously performed action. For example, three different tasks may begin with actions A and B. The system can prompt the user to select which of the three tasks the user is performing to enable the system to recommend an appropriate follow-up action for the task. The prompt can be presented, for example, via a touch interface and / or an audio interface.
[0062] According to one or more embodiments, performing an action associated with a trigger includes changing the state of the device without user intervention in response to detecting a user action. For example, the system can detect a user action to obtain component A from bin A. The system can illuminate a light on bin B corresponding to the next action to be performed (e.g., selecting component B from bin B) without user intervention. Alternatively, the system may determine that selecting component A was an action performed out of sequence. Thus, the system can illuminate one light (such as a red light) on bin A and another light (e.g., a green light) on bin B to prompt the user to correct the sequence of operations. The system can initiate a warm-up sequence for the device associated with the next action in the sequence of actions. According to another example, the system can cause the device to begin scanning for components in response to detecting a user action to place a component on the device.
[0063] Upon executing the system action associated with the system action trigger, the system determines whether a subsequent action exists (operation 216). For example, the system action may include generating a notification to the user or a prompt for user input. Upon receiving the notification or responding to the prompt, the user still continues performing the action associated with the task. Thus, the workflow proceeds to operation 208. In contrast, the system action may include terminating the task. If the system determines that the detected action was the last action in the task, the system may refrain from displaying information about any subsequent actions. The system may, for example, generate a user interface element indicating that the task is completed. If the system determines that a detected fault or anomaly makes task execution impossible or unsatisfactory, the system may refrain from displaying additional action information. The system may generate a user interface element prompting the user to restart the task, start a new task, or obtain assistance from another operator. The additional operator may be in a different workspace from the operator being monitored by workspace monitoring platform 110. For example, the additional operator may be in an adjacent workspace or may be the supervisor of the monitored operator. According to one exemplary embodiment, the system monitors multiple different operators in different workspaces or the same workspace. For example, two operators may work to assemble a product in the same workspace. The system can monitor each operator's actions independently. As another example, the system can monitor the actions of operations in two separate workspaces. The system can generate additional operator notifications based on the operator's state, such as when the operator is resting between tasks or operations.
[0064] 3.1 Starting a branch task According to one or more embodiments, the system initiates a branching task from a primary task being performed by a user. For example, the system can detect a user action that varies from a displayed action for a particular operation. Based on the detection of the user action's variation from the displayed user action, the system can initiate the branching task. For example, a primary task may include operations A, B, and C. The system may detect that the user performs operation D instead of operation B. The system may initiate a branching task that prompts the user to perform operations D1 and D2. The system may then return to the primary task by prompting the user to perform operation C.
[0065] As an example, action D may include the user pressing one button on the machine instead of another button designated for the task. Actions D1 and D2 may correspond to user actions that interact with the machine to put it into a state ready to perform action C. According to an alternative example, actions D1 and D2 may include an instructional video showing how to perform action B. The system may then prompt the user to repeat action B or resume the task.
[0066] According to another example, a user may initiate a task including actions A, B, C, and D performed in sequence. Following the performance of action B, the system displays a graphical representation associated with action C on the user interface. The user may be assigned an urgent task by a supervisor. In response, the system detects that the user is performing action E associated with the urgent task. The system can determine whether (a) the first action must be interrupted or (b) store the user's progress in the first task for completion at a later point in time. For example, the system may determine that action E is part of a sequence of actions E, F, B, and G. If action B requires the same equipment in both tasks, the system can interrupt the first task to allow the user to access the equipment in the second task. Alternatively, the system may determine that the new task does not conflict with the initial task. In response, the system can display the actions for the new task as a branch task from the initial task. Then, upon detecting the user's completion of action G, the system can resume prompting the user to perform the actions of the initial task by returning to displaying the graphical representation associated with action C of the initial task.
[0067] 3.2 Triggering the recording of statistics associated with a task According to one or more embodiments, the execution of operations and tasks causes the system to record statistics associated with the operations and tasks. Statistics may include, for example, recording resources consumed to perform the operations or tasks. These statistics may further include, for example, recording the number of components produced as a result of completing the task. According to another example, the system records usage statistics for machines used to perform operations and tasks. For example, the system may track how long machines operate to perform operations and tasks. The usage information may feed into a maintenance log for scheduling equipment maintenance.
[0068] According to one or more embodiments, the system uses task statistics to generate and / or modify schedules for the execution of future tasks and / or operations. For example, the system may calculate future delays and their production lines, plan for resource increases at specific points in the future, predict completion times for future tasks, and use task metrics associated with one or more tasks to forecast project completion times based on one or more tasks.
[0069] 3.3 Monitoring task metrics based on detected user actions According to one or more embodiments, the system generates task metrics based on the detected user actions. Examples of task metrics include completion time, error detection rate, other anomaly detection rate, equipment utilization rate, and equipment failure detection rate. For example, the system can detect the start time of a task or operation and the end time of the task or operation. The system can store the task or operation completion time. The system can compare the completion time with a predicted completion time based on the same user or other users. The system can provide the user or other entity with efficiency metrics for a particular user, a particular operation, or a particular task based on the comparison. For example, the system can generate a task completion rate for a set of operators performing the same task.
[0070] 4. Machine learning model training 3 illustrates an example set of operations for training a machine learning model to classify patterns of user actions as corresponding to behaviors and / or tasks, according to one or more embodiments. Classifying a user action can include classification as (a) as corresponding to a particular task and / or (b) as corresponding to a particular behavior. According to one embodiment, the system also assigns a confidence level to the prediction. Based on one user action, the system can assign a low confidence level to a particular classification as corresponding to a particular task. Based on two or more user actions performed in succession, the system can assign a higher confidence level to a particular classification for a particular user action.
[0071] The method includes identifying or obtaining historical video image data (operation 302). Obtaining historical data may include obtaining data associated with user positions within the workspace. For example, a set of image data may include a user interacting with equipment, components, or another user within the workspace. The historical data may be associated with other data, such as equipment status data (e.g., whether a particular machine is on or in a particular state at the time a particular set of image data was captured). Historical image data may include both actual image data (e.g., captured by a video camera) and synthetic image data (e.g., data generated by a computer to mimic video data).
[0072] The system generates a set of training data using the historical video data (operation 304). The set of training data includes video data of recorded user actions and labels identifying the operations and / or tasks associated with the recorded user actions. The set of training data may include additional attributes associated with the video data, including equipment data such as the user's identity, the user's location within the organization, previous operations and / or tasks performed by the user before the recorded action, the location of the workspace in which the user is operating, and the status of machines in the workspace (e.g., powered, powered down, ready for operation, cooldown from operation, performing operation). The training labels may include an identifier of the operation (e.g., a subprocess of a task composed of the operation) associated with the action (e.g., "operation connect component A to component B") and / or an identifier of the task associated with the action and / or the corresponding operation (e.g., "action connect component A to component B, task assemble component XYZ").
[0073] In some embodiments, generating the training dataset includes generating a set of feature vectors for labeled examples. The feature vectors for the examples can be n-dimensional, where n represents the number of features in the vector. The number of features selected can vary depending on the particular implementation. The features can be curated in a supervised manner or automatically selected from attributes extracted during model training and / or tuning. Exemplary features include the user's identity, the user's location within the organization, previous actions and / or tasks performed by the user before the recorded action, the location of the workspace in which the user is operating, and equipment data such as the status of equipment within the workspace. In some embodiments, the features in the feature vector are represented numerically by one or more bits. The system can convert categorical attributes to numerical representations using encoding schemes such as one-hot encoding, label encoding, or binary encoding. One-hot encoding generates a unique binary feature for each possible category in the original features. In one-hot encoding, when one feature has a value of 1, the remaining features have a value of 0. For example, if the type of healthcare service has 10 different categories, the system may generate 10 different features for the input dataset. When one category is present (e.g., a value of "1"), the remaining features are assigned a value of "0." According to another example, the system may perform label encoding by assigning a unique numeric value to each category. According to yet another example, the system may perform binary encoding by converting the numeric value to binary digits and generating a new feature for each digit.
[0074] The system applies a machine learning algorithm to the training dataset to train a machine learning model (operation 306). The machine learning algorithm analyzes the training dataset to train neurons of a neural network with specific weights and offsets to associate specific recorded user actions with specific behaviors and / or tasks. According to one or more embodiments, the machine learning algorithm or post-machine learning algorithm further assigns a confidence score to the prediction. For example, as a result of training the machine learning model, a relationship between one user action and specific behaviors that are part of three different tasks may be identified. As a result of training the machine learning model, it may further determine that when a behavior is performed after the completion of task A, the behavior is most likely associated with a behavior in task B. As a result of training the machine learning model, it may further determine that when a behavior is performed after the completion of task C, the behavior is most likely associated with a behavior in task D. Thus, if the machine learning model determines that a behavior identified in the recorded user actions was performed after task C, the behavior is likely to be behavior P in task D at an 80% confidence level.
[0075] In some embodiments, the system iteratively applies a machine learning algorithm to a set of input data to generate an output set of labels, compares the generated labels to previously generated labels associated with the input data, adjusts the algorithm's weights and offsets based on the error, and applies the algorithm to another set of input data. In some cases, the system can generate and train candidate recurrent neural network models, such as long-term memory (LSTM) models. In a recurrent neural network, one or more network nodes or "cells" can include memory. The memory enables individual nodes in the neural network to capture dependencies based on the order in which feature vectors are fed through the model. The weight applied to a feature vector representing an expense or activity can depend on its position in the sequence of feature vector representations. Thus, a node can have memory for storing associated temporal dependencies between different recorded user actions. For example, a recorded user action alone can have a first set of weights applied by the node as a function of the respective feature vector for the expense. However, if a recorded user action is immediately preceded by another type of recorded user action associated with behavior in a particular task, a different set of weights may be applied by one or more nodes based on memory of the preceding recorded user action. In this case, the behavior prediction assigned to the second recorded user action may be influenced by the first recorded user action. Additionally or alternatively, the system may generate and train other candidate models, such as support vector machines, decision trees, Bayesian classifiers, and / or fuzzy logic models, as described above.
[0076] In some embodiments, the system compares the labels estimated through one or more iterations of the machine learning model algorithm with the observed labels to determine an estimation error (operation 308). The system may perform this comparison on a test set of examples, which may be a subset of examples in the training dataset that were not used to generate and fit the candidate model. The total estimation error for a particular iteration of the machine learning algorithm may be calculated as a function of the magnitude of the difference and / or the number of examples for which the estimated label was incorrectly predicted.
[0077] In some embodiments, the system determines whether to adjust weights and / or other parameters based on the estimation error (operation 310). Adjustments can be made until a candidate model is identified that minimizes the estimation error or otherwise achieves a threshold level of estimation error. The process can return to operation 308 to make adjustments and continue training the machine learning model.
[0078] In some embodiments, the system selects machine learning model parameters based on the estimation error meeting a threshold accuracy level (operation 312). For example, the system may select a set of parameter values for the machine learning model based on determining that the training model has a predictive accuracy level of at least 98% of the behaviors and / or task labels of the recorded user actions.
[0079] In some embodiments, the system trains the neural network using backpropagation. Backpropagation is the process of updating cell states in a neural network based on gradients determined as a function of the estimation error. With backpropagation, nodes are assigned a proportion of the estimation error based on their contribution to the output and are adjusted based on this proportion. In recurrent neural networks, time is also taken into account in the backpropagation process. As described above, a given example may include a sequence of related recorded user actions. Each recorded user action can be treated as a separate, discrete point in time. For example, an example may include recorded user actions c1, c2, and c3 corresponding to times t, t+1, and t+2, respectively. Diachronic error backpropagation can begin at time t+2 and move backward in time to t+1 and then to t, making adjustments via gradient descent. Furthermore, the backpropagation process can adjust the memory parameters of a cell so that the cell remembers contributions from previous recorded user actions in the sequence of recorded user actions. For example, a cell calculating the contribution of e3 may have a memory for the contribution of e2 with a memory for e1. Memory can act as feedback connections such that the output of a cell at one time (e.g., t) is used as the input to the next time in the sequence (e.g., t+1). Gradient descent techniques can account for these feedback connections so that the contribution of one recorded user action to the cell's output can influence the contribution of the next recorded user action on the cell's output. Thus, the contribution of c1 can influence the contribution of c2, and so on.
[0080] Additionally or alternatively, the system can train other types of machine learning models. For example, the system can adjust the boundaries of hyperplanes in a support vector machine or node weights in a decision tree model to minimize estimation errors. Once trained, the machine learning model can be used to estimate labels for new examples of recorded user actions.
[0081] In embodiments where the machine learning algorithm is a supervised machine learning algorithm, the system may optionally obtain feedback on various aspects of the analysis described above (operation 314). For example, the feedback may confirm or revise the labels generated by the machine learning model. The machine learning model may indicate that a particular recorded user action is associated with the label "connect component A to component B." The system may receive feedback indicating that a particular recorded user action should instead be associated with the label "disconnect component A from component B." Based on the feedback, the machine learning training set may be updated, thereby improving its analysis accuracy (operation 316). Once updated, the system may optionally further train the machine learning model by applying the model to additional training datasets.
[0082] 5. Exemplary Embodiments Detailed examples are described below for clarity. The components and / or operations described below should be understood as specific examples that may not be applicable to certain embodiments. Therefore, the components and / or operations described below should not be construed as limiting the scope of any of the claims.
[0083] Figure 4A shows a user interface display 470 that displays a set of tasks 471 that can be performed at a workstation. For example, a workstation management system can detect the presence of a particular user and display a set of tasks that can be performed by the particular user at the workstation. Based on user selection of the task "Assemble Widget," represented by a box around the task name in Figure 4A, the system displays a set of actions 472 to be performed to complete the task, as shown in Figure 4B.
[0084] 4B shows a set of actions 472 and a currently displayed action 473, "Get component A from bin 14." The system further displays a representation of bin 14 (reference numeral 474) and a representation of component A (reference numeral 475).
[0085] Referring to FIG. 4C , the system detects user actions via a video camera 428. The workspace 420 is monitored by the camera 428. The camera 428 monitors user actions of a user 440 in the workspace 420. The workspace 420 also includes equipment 421 and 422, a work surface 426, and a display device 470. The workspace monitoring platform 410 obtains camera data and status data from the equipment 421 and 422 to detect user actions. For example, the system detects a user action of moving to a bin 424 and from the bin 424 to the work surface 426. The system also detects a change in the weight of the bin 424. The system provides the video and sensor data to a user action classification machine learning model 411. The model 411 determines that the user action corresponds to a first operation of the task (e.g., “Step 1: Get component A from bin 14”) (operation 412). In response to identifying the user action as corresponding to the first action in the task, the system displays the next action in the task (eg, "Action 2") on display device 470 (act 413).
[0086] Referring to FIG. 4D , the system detects a next user action via video camera 428. Workspace monitoring platform 410 provides video and sensor data to user action classification machine learning model 411. Model 411 determines that the user action corresponds to action 3 of the task rather than the displayed action associated with second action 2 of the task (operation 414). In response to identifying the user action as corresponding to the third action in the task, the system determines whether the actions that make up the task may be performed out of sequence. Based on determining that the action of the task may be performed out of sequence (operation 415), the system identifies the next action to display. The system determines that the second action, action 2, should be performed next and re-displays the user action associated with the second action on display device 470 (operation 416).
[0087] Referring to FIG. 4E , the system detects a next user action via video camera 428. Workspace monitoring platform 410 provides video and sensor data to user action classification machine learning model 411. Model 411 determines that the user action corresponds to operation 4 of the task rather than the displayed action associated with operation 2 of the task (operation 417). In response to identifying the user action as corresponding to the fourth operation in the task, the system determines whether the fourth operation in the task may be performed before the second operation in the task. The system determines that the fourth operation should not be performed before the second operation (operation 418). For example, the system may compare the completion time associated with performing the operations in a particular sequence with a threshold completion time for the task. The system determines that performing operation 4 before operation 2 results in the task completion time exceeding the threshold. Based on the determination that operation 4 should not be performed before operation 2 in the task, the system displays a notification on display device 470 recommending that the user restart the task (operation 419). The system receives user input via a user interface instructing the system to allow the operations of the task to be performed out of sequence (operation 430). The system determines (a) whether the user has permission to reorder steps in the task and (b) whether the task can be completed in an order different from the stored sequence. For example, in a particular task being performed, some steps can be stored in a particular sequence, and other steps can be stored without any particular sequence. Thus, operation 4 can be stored with a dependency from operation 2, but operation 3 can be stored without a dependency. Thus, operation 3 can be identified by the system as being capable of being performed in any sequence.Based on a determination that (a) the user has the authority to modify dependencies in the task to change the sequence in which task operations are performed, and (b) the task can be performed in the modified sequence, the system allows the user to reorder the sequence of operations to perform the task (operation 431). The system recalculates the estimated completion time for completing the task based on the reordered sequence (operation 432). The system can notify downstream processes, such as a workflow management system, that manage the production of components using the particular component being produced by user 440 of the updated estimated completion time.
[0088] Referring to FIG. 4F , the system can detect user actions even when a task is not selected by the user. The system detects user actions via video camera 428. Workspace monitoring platform 410 provides video and sensor data to user action classification machine learning model 411. Model 411 determines that the user action corresponds to task operation 21 (“Place component in test machine 422”) (operation 451). The system identifies a set of three tasks that includes operation 21. The system further identifies tasks that are most likely to be performed by user 440. For example, the three tasks that include operation 21 include operations that can be performed in any sequence. However, through training, ML model 411 learned that the user only performed operation 21 at the beginning of the sequence of operations when performing task 3. Therefore, the system (a) determines a sequence for displaying the operations of task 3 and (b) displays the user action associated with the next operation (e.g., operation 22) to be performed in task 3 via display device 470 (operation 452).
[0089] 6. Computer Networks and Cloud Networks In one or more embodiments, the workspace monitoring system is implemented in a computer network. The computer network provides connectivity between a set of nodes. These nodes may be local and / or remote from one another. The nodes are connected by a set of links. Examples of links include coaxial cable, unshielded twisted cable, copper cable, optical fiber, and virtual links.
[0090] A subset of nodes implements computer networks. Examples of such nodes include switches, routers, firewalls, and network address translators (NATs). Another subset of nodes uses computer networks. Such nodes (also called "hosts") can run client processes and / or server processes. A client process makes a request for a computing service (such as running a particular application and / or storing a particular amount of data). A server process responds by performing the requested service and / or returning corresponding data.
[0091] A computer network may be a physical network including physical nodes connected by physical links. A physical node is any digital device. A physical node may be a function-specific hardware device such as a hardware switch, a hardware router, a hardware firewall, and a hardware NAT. Additionally or alternatively, a physical node may be a general-purpose machine configured to run various virtual machines and / or applications that perform respective functions. A physical link is a physical medium connecting two or more physical nodes. Examples of links include coaxial cable, unshielded twisted cable, copper cable, and optical fiber.
[0092] A computer network may be an overlay network. An overlay network is a logical network implemented on top of another network (e.g., a physical network). Each node in the overlay network corresponds to a respective node in the underlying network. Thus, each node in the overlay network is associated with both an overlay address (for addressing the overlay node) and an underlay address (for addressing the underlay node that implements the overlay node). An overlay node may be a digital device and / or a software process (e.g., a virtual machine, an application instance, or a thread). The links connecting overlay nodes are implemented as tunnels through the underlying network. The overlay nodes at both ends of the tunnel treat the underlying multi-hop path between these overlay nodes as a single logical link. Tunneling is performed through encapsulation and decapsulation.
[0093] In an embodiment, a user may access the workspace monitoring platform through a client. The client may be local and / or remote to the computer network. The client may access the computer network through a private network or another computer network, such as the Internet. The client may communicate requests to the computer network using a communication protocol, such as Hypertext Transfer Protocol (HTTP). The requests are communicated through an interface, such as a client interface (e.g., a web browser), a program interface, or an application programming interface (API).
[0094] In an embodiment, a computer network provides connectivity between clients and network resources. The network resources include hardware and / or software configured to run server processes. Examples of network resources include processors, data storage devices, virtual machines, containers, and / or software applications. The network resources are shared among multiple clients. The clients request computing services from the computer network independently of each other. The network resources are dynamically allocated to requests and / or clients on an on-demand basis. The network resources allocated to each request and / or client may be scaled up or down based on, for example, (a) the computing services requested by a particular client, (b) the aggregated computing services requested by a particular tenant, and / or (c) the aggregated computing services requested from the computer network. Such a computer network may also be referred to as a "cloud network."
[0095] In an embodiment, a service provider provides a cloud network to one or more end users. Various service models can be implemented by the cloud network, including, but not limited to, Software as a Service (SaaS), Platform as a Service (PaaS), and Infrastructure as a Service (IaaS). In SaaS, the service provider provides end users with the ability to use the service provider's applications running on the network resources. In PaaS, the service provider provides end users with the ability to deploy custom applications on the network resources. The custom applications can be created using programming languages, libraries, services, and tools supported by the service provider. In IaaS, the service provider provides end users with the ability to provision the processing, storage, network, and other basic computing resources provided by the network resources. Any application, including an operating system, can be deployed on the network resources.
[0096] In embodiments, various deployment models may be implemented by a computer network, including, but not limited to, private cloud, public cloud, and hybrid cloud. In a private cloud, network resources are provisioned for exclusive use by a specific group of one or more entities (the term "entity" as used herein refers to a company, organization, person, or other entity). The network resources may be local and / or remote to the premises of the specific group of entities. In a public cloud, cloud resources are provisioned for multiple entities (also referred to as "tenants" or "customers") that are independent of one another. The computer network and its network resources are accessed by clients corresponding to different tenants. Such a computer network may be referred to as a "multi-tenant computer network." Several tenants may use the same specific network resources at different times and / or at the same time. The network resources may be local and / or remote to the tenant's premises. In a hybrid cloud, the computer network comprises a private cloud and a public cloud. An interface between the private cloud and the public cloud enables data and application portability. Data stored in the private cloud and data stored in the public cloud may be exchanged through the interface. Applications implemented in a private cloud and applications implemented in a public cloud may have dependencies on each other, and calls from applications in a private cloud to applications in a public cloud (and vice versa) may be made through an interface.
[0097] In embodiments, tenants of a multi-tenant computer network are independent of one another. For example, the business or operations of one tenant may be separate from the business or operations of another tenant. Different tenants may require different network requirements from the computer network. Examples of network requirements include processing speed, data storage, security requirements, performance requirements, throughput requirements, latency requirements, resilience requirements, quality of service (QoS) requirements, tenant isolation, and / or consistency. The same computer network may be required to implement the different network requirements required by different tenants.
[0098] In one or more embodiments, tenant isolation is implemented in a multi-tenant computer network to ensure that applications and / or data of different tenants are not shared with each other. Various tenant isolation approaches may be used.
[0099] In an embodiment, each tenant is associated with a tenant ID. Each network resource in the multi-tenant computer network is tagged with a tenant ID. A tenant is granted access to a particular network resource only if the tenant and the particular network resource are associated with the same tenant ID.
[0100] In an embodiment, each tenant is associated with a tenant ID. Each application implemented by the computer network is tagged with a tenant ID. Additionally or alternatively, each data structure and / or dataset stored by the computer network is tagged with a tenant ID. A tenant is granted access to a particular application, data structure, and / or dataset only if the tenant and the particular application, data structure, and / or dataset are associated with the same tenant ID.
[0101] As one example, each database implemented by a multi-tenant computer network may be tagged with a tenant ID. Only the tenant associated with the corresponding tenant ID may access the data in a particular database. As another example, each entry in a database implemented by a multi-tenant computer network may be tagged with a tenant ID. Only the tenant associated with the corresponding tenant ID may access the data in a particular entry. However, a database may be shared by multiple tenants.
[0102] In an embodiment, the subscription list indicates which tenants have authorization to access which applications. For each application, a list of tenant IDs of tenants authorized to access the application is stored. A tenant is granted access to a particular application only if the tenant ID of the tenant is included in the subscription list corresponding to the particular application.
[0103] In an embodiment, network resources (such as digital devices, virtual machines, application instances, and threads) corresponding to different tenants are separated into tenant-specific overlay networks maintained by a multi-tenant computer network. As an example, packets from any source device in a tenant overlay network can be sent only to other devices within the same tenant overlay network. To prohibit any transmission from a source device on a tenant overlay network to a device in another tenant overlay network, an encapsulation tunnel is used. Specifically, a packet received from a source device is encapsulated within an outer packet. The outer packet is sent from a first encapsulation tunnel endpoint (communicating with the source device in the tenant overlay network) to a second encapsulation tunnel endpoint (communicating with a destination device in the tenant overlay network). The second encapsulation tunnel endpoint decapsulates the outer packet to obtain the original packet sent by the source device. The original packet is sent from the second encapsulation tunnel endpoint to a destination device in the same specific overlay network.
[0104] 7. Miscellaneous, Extensions Embodiments are directed to systems that include one or more devices that include a hardware processor and are configured to perform any of the operations described herein and / or recited in any of the appended claims.
[0105] In an embodiment, a non-transitory computer-readable storage medium includes instructions that, when executed by one or more hardware processors, cause any of the operations described herein and / or recited in any of the claims to be performed.
[0106] Any combination of the features and functions described herein may be used in accordance with one or more embodiments. In the foregoing specification, embodiments have been described with reference to numerous specific details that may vary from implementation to implementation. Accordingly, the specification and drawings should be considered in an illustrative rather than a restrictive sense. The sole and exclusive indication of the scope of the present invention and what the applicants intend to be the scope of the present invention is the literal and equivalent scope of the set of claims issuing from this application, in the specific form from such claims, including any subsequent amendments.
[0107] 8. Hardware Overview According to one embodiment, the techniques described herein are implemented by one or more special-purpose computing devices. These special-purpose computing devices may be hardwired to perform these techniques, or may include digital electronic devices such as one or more application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or network processing units (NPUs) permanently programmed to perform these techniques, or may include one or more general-purpose hardware processors programmed to perform these techniques according to program instructions in firmware, memory, other storage, or a combination thereof. Such special-purpose computing devices may also combine custom hardwired logic, ASICs, FPGAs, or NPUs with custom programming to perform these techniques. The special-purpose computing devices may be desktop computer systems, portable computer systems, handheld devices, networking devices, or any other devices incorporating hardwired logic and / or program logic to implement these techniques.
[0108] 5 is a block diagram illustrating a computer system 500 in which embodiments of the present invention may be implemented. Computer system 500 includes a bus 502 or other communication mechanism for communicating information, and a hardware processor 504 coupled with bus 502 for processing information. Hardware processor 504 may be, for example, a general-purpose microprocessor.
[0109] Computer system 500 also includes a main memory 506, such as a random access memory (RAM) or other dynamic storage device, coupled to bus 502 for storing information and instructions executed by processor 504. Main memory 506 may also be used for storing temporary variables or other intermediate information during execution of instructions executed by processor 504. Such instructions, when stored on a non-transitory storage medium accessible to processor 504, render computer system 500 a special-purpose machine customized to perform the operations specified in the instructions.
[0110] Computer system 500 further includes a read only memory (ROM) 508 or other static storage device coupled to bus 502 for storing static information and instructions for processor 504. A storage device 510, such as a magnetic disk or optical disk, is provided and coupled to bus 502 for storing information and instructions.
[0111] Computer system 500 may be coupled via bus 502 to a display 512, such as a cathode ray tube (CRT), for displaying information to a computer user. An input device 514, including alphanumeric and other keys, is coupled to bus 502 for communicating information and command selections to processor 504. Another type of user input device is a cursor control 516, such as a mouse, trackball, or cursor direction keys, for communicating directional information and command selections to processor 504 and for controlling cursor movement on display 512. This input device typically has two degrees of freedom in two axes, a first axis (e.g., x) and a second axis (e.g., y), which allows the device to specify a position in a plane.
[0112] Computer system 500 may implement the techniques described herein using customized hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic that, when combined with the computer system, makes computer system 500 a special-purpose machine or programs it to be a special-purpose machine. According to one embodiment, the techniques herein are performed by computer system 500 in response to processor 504 executing one or more sequences of one or more instructions contained in main memory 506. Such instructions may be read into main memory 506 from another storage medium, such as storage device 510. Execution of the sequences of instructions contained in main memory 506 causes processor 504 to perform the process steps described herein. In alternative embodiments, hardwired circuitry may be used in place of or in combination with software instructions.
[0113] The term "storage medium," as used herein, refers to any non-transitory medium that stores data and / or instructions that cause a machine to operate in a specific manner. Such storage media can include non-volatile media and / or volatile media. Non-volatile media include, for example, optical or magnetic disks, such as storage device 510. Volatile media include dynamic memory, such as main memory 506. Common forms of storage media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tape, or any other magnetic data storage medium, CD-ROMs, any other optical data storage medium, any physical medium with a pattern of holes, RAM, PROM, EPROM, FLASH-EPROM, NVRAM, any other memory chip or cartridge, content addressable memory (CAM), and ternary content addressable memory (TCAM).
[0114] Storage media are distinct from but may be used in conjunction with transmission media. Transmission media involves transferring information between storage media. For example, transmission media include coaxial cables, copper wire and fiber optics, including the wires that comprise bus 502. Transmission media can also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications.
[0115] Various forms of media may be involved in carrying one or more sequences of one or more instructions to processor 504 for execution. For example, the instructions may initially be carried on a magnetic disk or solid state drive of a remote computer. The remote computer may load the instructions into its dynamic memory and send the instructions over a telephone line using a modem. A modem local to computer system 500 can receive the data on the telephone line and use an infrared transmitter to convert the data to an infrared signal. An infrared detector can receive the data carried in the infrared signal and appropriate circuitry can place the data on bus 502. Bus 502 carries the data to main memory 506, from which processor 504 retrieves and executes the instructions. The instructions received by main memory 506 may optionally be stored on storage device 510 either before or after execution by processor 504.
[0116] Computer system 500 also includes a communication interface 518 coupled to bus 502. Communication interface 518 provides a two-way data communication coupling to a network link 520 that is connected to a local network 522. For example, communication interface 518 may be an Integrated Services Digital Network (ISDN) card, cable modem, satellite modem, or modem to provide a data communication connection to a corresponding type of telephone line. As another example, communication interface 518 may be a local area network (LAN) card to provide a data communication connection to a compatible LAN. A wireless link may also be implemented. In any such implementation, communication interface 518 sends and receives electrical, electromagnetic or optical signals that carry digital data streams representing various types of information.
[0117] Network link 520 typically provides data communication through one or more networks to other data devices. For example, network link 520 may provide a connection through local network 522 to a host computer 524 or to data equipment operated by an Internet Service Provider (ISP) 526. ISP 526 in turn provides data communication services through the worldwide packet data communication network now commonly referred to as the "Internet" 528. Local network 522 and Internet 528 both use electrical, electromagnetic or optical signals that carry digital data streams. The signals through the various networks and the signals on network link 520 and through communication interface 518, which carry the digital data to and from computer system 500, are exemplary forms of transmission media.
[0118] Computer system 500 can send messages and receive data, including program code, through the network(s), network link 520 and communication interface 518. In the Internet example, a server 530 might transmit a requested code for an application program through Internet 528, ISP 526, local network 522 and communication interface 518.
[0119] The received code may be executed by processor 504 as it is received, and / or stored in storage device 510, or other non-volatile storage for later execution.
[0120] In the above specification, embodiments of the present invention have been described with reference to numerous specific details that may vary from embodiment to embodiment. Accordingly, the specification and drawings should be considered in an illustrative rather than a restrictive sense. The sole and exclusive indication of the scope of the present invention and what the applicants intend to be the scope of the present invention is the literal and equivalent scope of the set of claims originating from this application, in the specific form derived from such claims, including any subsequent amendments.
Claims
1. A non-transitory computer-readable medium containing instructions that, when executed by one or more hardware processors, cause operations to be performed, the operations including: a user presenting a first instruction associated with a first action in a set of actions to perform the first action; Concurrently with presenting the first instructions, analyzing the video stream in real time as it is received to detect a first set of one or more actions performed by the user; determining whether the first set of actions corresponds to completing the first operation associated with the currently presented first instruction; In response to determining that the first set of actions corresponds to the completion of the first operation, submitting a second instruction corresponding to a second action subsequent to the first action in the set of actions to cause the user to perform the second action; 1. A non-transitory computer-readable medium comprising:
2. The operation is Concurrently with presenting the second instructions, analyzing the video stream in real time as it is received to detect a second set of one or more actions performed by the user; determining whether the second set of actions corresponds to the second behavior associated with the currently presented second instruction; In response to determining that the second set of actions corresponds to the second behavior, determining that the user has begun performing the second action; The non-transitory computer-readable medium of claim 1 , further comprising:
3. The operation is In response to determining that the user has begun performing the second action, The non-transitory computer-readable medium of claim 2 , further comprising presenting third instructions associated with a third operation in the set of operations to perform the third operation.
4. The operation is In response to determining that the second set of actions corresponds to the second behavior, The non-transitory computer-readable medium of claim 2 , further comprising detecting a start time of execution of the second operation.
5. The operation is The non-transitory computer-readable medium of claim 4 , further comprising determining scheduling-related information based on the start time of execution of the second operation.
6. determining the scheduling-related information 6. The non-transitory computer-readable medium of claim 5, comprising at least one of: calculating future delays in a workflow; planning for expansion of resources at a particular time; and predicting completion times of one or both of the second operation and a task including the second operation.
7. The operation is Detecting an end time of execution of the second operation; calculating a total time to complete the activity based on the start time and the end time; The non-transitory computer-readable medium of claim 4 further comprising:
8. The operation is Concurrently with presenting the second instructions, analyzing the video stream in real time as it is received to detect a second set of one or more actions performed by the user; determining whether the second set of actions corresponds to completing the second operation associated with the currently presented second instruction; In response to determining that the second set of actions corresponds to a third operation different from the second operation, determining that the second set of actions is out of sequence; and presenting a notification based on the second set of actions being out of sequence; and The non-transitory computer-readable medium of claim 1 , further comprising:
9. The operation is Concurrently with presenting the second instructions, analyzing the video stream in real time as it is received to detect a second set of one or more actions performed by the user; determining whether the second set of actions corresponds to completing the second operation associated with the currently presented second instruction; In response to determining that the second set of actions corresponds to a third operation different from the second operation in the set of operations, modifying the sequence of the set of actions such that an instruction corresponding to the third action is presented before a presentation of an instruction corresponding to the second action; The non-transitory computer-readable medium of claim 1 , further comprising:
10. The operation is Identifying the tasks to be completed; determining the set of actions based on the task to be completed; The non-transitory computer-readable medium of claim 1 , further comprising:
11. The operation is analyzing the video stream in real time as it is received to detect the performance of a particular action; performing an operation to capture system data in response to detecting the particular action; The non-transitory computer-readable medium of claim 1 , further comprising:
12. 10. The non-transitory computer-readable medium of claim 1, wherein a machine learning model generates a sequence for performing a set of operations that includes the first operation based on a task completion metric associated with performing the set of operations in an order corresponding to the sequence.
13. The operation is identifying a task including the first operation, the second operation, and a third operation sequentially following the second operation; Concurrently with presenting the second instructions, analyzing the video stream in real time as it is received to detect a second set of one or more actions performed by the user; determining whether the second set of actions corresponds to completing the second operation associated with the currently presented second instruction; determining that the second set of actions corresponds to a trigger criterion; based on determining that the second set of actions corresponds to the trigger criteria; presenting a modification instruction different from the third action; and presenting a prompt to receive user input regarding modifying the task to include a fourth operation that includes the second set of actions; performing a trigger action including at least one of: The non-transitory computer-readable medium of claim 1 , further comprising:
14. The operation is Concurrently with presenting the second instructions, analyzing the video stream in real time as it is received to detect a second set of one or more actions performed by the user; determining whether the second set of actions corresponds to completing the second operation associated with the currently presented second instruction; In response to determining that (a) the second set of actions corresponds to completing the second operation, and (b) the second set of actions does not correspond to actions specified in the second instructions, modifying the second instructions to include the second set of actions; The non-transitory computer-readable medium of claim 1 , further comprising:
15. The operation is In response to determining that the first set of actions corresponds to the completion of the first operation, The non-transitory computer-readable medium of claim 1 , further comprising updating a task profile associated with the set of actions to indicate that the first action has been completed.
16. The operation is calculating estimated completion times for different orders of performing the set of operations; selecting a particular order for presenting the instructions corresponding to the set of actions based on the estimated completion time; The non-transitory computer-readable medium of claim 1 , further comprising:
17. The non-transitory computer-readable medium of claim 1 , wherein the set of operations specifies a sequence for ordering operations in the set of operations, and the second operation is determined based on the sequence.
18. 10. The non-transitory computer-readable medium of claim 1, wherein presenting the second instructions is performed without user input requesting a switch from presenting the first instructions to presenting the second instructions.
19. The operation is In response to determining that the first set of actions corresponds to the completion of the first operation, The non-transitory computer-readable medium of claim 1 , further comprising updating, in a data store, a value representing at least one component consumed during execution of the first operation.
20. The operation is In response to determining that the first set of actions corresponds to the completion of the first operation, 10. The non-transitory computer-readable medium of claim 1, further comprising detecting, by at least one sensor, an amount of consumable resources available for execution of the set of operations at a subsequent instance of execution of the set of operations.
21. A method comprising the operations of any of claims 1 to 20.
22. 1. A system comprising: one or more processors; a memory having stored thereon instructions that, when executed by the one or more processors, cause the system to perform the operations of any one of claims 1 to 20; A system comprising:
23. A system comprising means for performing the operations of any of claims 1 to 20.