Automated user interface documentation platforms, systems, methods, and devices
Automated user interface documentation systems address the inefficiencies of manual documentation by recording user interactions and providing visual guidance, enhancing documentation accuracy and integration.
Patent Information
- Application Number
- PCT/US2025/030737
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-23
- Filing Date
- 2025-05-23
- Publication Date
- 2025-11-27
AI Technical Summary
Existing user interface documentation is manually generated, time-consuming, and often incomplete or incorrect, especially across different devices and after interface changes, leading to user confusion and difficulty in maintaining accurate documentation.
Automated systems and methods for generating, presenting, and maintaining user interface documentation by recording user interactions, providing visual guidance, and presenting knowledge indicators to facilitate workflow understanding and documentation integration.
Enhances the efficiency and accuracy of user interface documentation, reducing user confusion and maintenance efforts by automating the documentation process and integrating it seamlessly with the user interface.
Smart Images

Figure US2025030737_27112025_PF_FP_ABST
Abstract
Description
AUTOMATED USER INTERFACE DOCUMENTATION PLATFORMS, SYSTEMS, METHODS, AND DEVICESCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority benefit to U.S. provisional application 63 / 651,325, filed on May 23, 2024, which is hereby incorporated by reference in its entirety.FIELD
[0002] The present disclosure relates to user interface documentation, and, more specifically, to platforms, systems, methods, and devices configured to automate user interface documentation and propagation thereof.BACKGROUND
[0003] Many computing scenarios involve a presentation of a user interface, such as a collection of windows, screens, and / or views, each respectively including a series of visual content and user controls such as text, images, buttons, textboxes, lists, and the like. More specifically, many computing scenarios involve a presentation of a website by a web browser, wherein the website includes a set of web pages, each providing a layout of text, images, buttons, textboxes, lists, and the like.
[0004] The user interface may permit users to perform one or more tasks or functions, such as creating content, editing content, viewing content, searching through content, and combining content to form additional content. Such content may include text, numeric values, images, videos, sounds, records of a database, data sets, documents, the visual content of a website, or the like. In order to perform each of the one or more tasks or functions, a user may complete a workflow including a set of actions, such as locating and activating particular content or user controls of the user interface in a particular sequence. For example, in order to create a database record corresponding to an individual through a website, a user may perform a first workflow including clicking a “New User” button, typing a name of the individual into a first textbox, manipulating a calendar user control to select a birthdate of the individual, clicking an “Upload Photo” button to upload an image of the individual, and clicking a “Save Record” button. In order to delete a record for an existing individual, a user may perform a second workflow including entering the name of the individual into a search query textbox, clicking a “Search” button, finding and selecting a row of a list that matches the individual, and clicking a “Delete Record” button.
[0005] Many user interfaces may include a large amount of content and / or a large number of user controls. For example, a website may include many web pages, each of which includes many user controls and hyperlinks to other web pages. Due to the volume of content and / or user controls distributed through the user interface, a user may have difficulty determining a correct sequenceof actions to perform to complete a workflow for a selected task. Alternatively or additionally, a task that can be performed through the user interface may involve a complex workflow that includes a significant number of actions to be performed in a particular order. Due to the length and / or complexity of the workflow, a user may have difficulty determining and completing the workflow only by inspecting the user interface. For example, the user may have difficulty understanding the functionality of certain user controls within the user interface or how such functionality fits into the overall workflow. Some user interfaces may change as the user performs the workflow, such as reactive websites that respond to events occurring in the web page by dynamically altering the web page (e.g., adding user controls, relocating or modifying user controls, and / or altering a layout and / or content of the web page). Such changes may be confusing to the user and / or difficult to predict. For example, a user may believe that a particular workflow requires activating a particular button, but may have difficulty locating the button on a web page of the website, possibly because the button exists on another web page of the website or is created only by first performing other actions through the website.
[0006] In order to facilitate aid users in performing tasks and to reduce user confusion, user interfaces such as websites often include documentation that users may consult to understand commonly performed workflows. For example, a website may provide a help section featuring a list of tasks, and a user may follow a hyperlink associated with a particular task to view a step-by- step guide for performing the task. The guide may indicate the sequence of actions that the user can perform to complete the workflow of the task, optionally including a text narrative, static images that highlight a location of a particular user control on a particular web page, an animation showing a user interaction with a particular user control, or the like. Alternatively or additionally, the guide may embed documentation near or otherwise attached to a user control, such as a hover tooltip that, when activated by a cursor, displays a message indicating the functionality of the user control.
[0007] Typically, user interface documentation is manually generated by one or more individuals. For example, a software developer or documentation engineer may manually create a list of tasks that a user may wish to perform, study the user interface to understand a workflow to perform each task, manually write descriptions of the series of actions for the workflow, and manually generate and attach images to supplement the written descriptions. The software developer or documentation engineer may revise and / or update the user interface documentation if users find it to be unclear or incomplete, if users encounter problems while trying to perform the steps, if new tasks are identified to be performed through the user interface, and / or if the user interface changes, such as via upgrades to the logical functionality and / or user interfaces of the user interface.
[0008] The manual development of user interface documentation by software developers and / or documentation engineers, as well as teams thereof, may be a lengthy, complicated, and arduous task to perform adequately and completely for a given user interface. Further, manually developed user interface documentation may not match the entire range of the user interface. For instance, user interface documentation may adequately describe a user interface when a website is accessed through a desktop web browser, but may be incorrect or insufficient when the same website is accessed through a mobile browser of a mobile device such as a table or phone. The user interface documentation for a website may initially be correct and complete, but may become incorrect and / or incomplete due to changes in the content and / or functionality of the website. Users may have difficulty finding the user interface documentation, or relevant portions thereof, while accessing portions of the user interface that are not clearly associated with corresponding portions of the user interface documentation. In some cases, users may have difficulty locating the user interface documentation at all, such as when the user interface documentation is provided in a downloadable document that is not well-integrated with the workflows of a website. Users may have questions and / or may discover issues associated with portions of the user interface documentation, but may be unable to determine how to submit such questions and / or report such issues to the software developers and / or documentation engineers who manually created the user interface documentation.
[0009] Due to at least these reasons, software developers and / or documentation engineers may be required to maintain the user interface documentation, such as supplementing, correcting, clarifying, extending, and / or reorganizing the user interface documentation. Typically, such maintenance is often performed manually, which may require the software developers and / or documentation engineers to compare the documentation of a particular workflow with the current functionality of the user interface in order to identify errors, discrepancies, ambiguities, or other issues, and to perform manual editing of the workflow to complete such maintenance. In some cases, the software developers and / or documentation engineers may require clarification of a question and / or issue submitted by a user about a particular workflow, and may have to contact and engage the user via out-of-band communication mechanisms, such as telephone, email, or an in-person meeting. The manual performance of these tasks may exacerbate the length, complexity, and / or arduous nature of the maintenance of the user interface documentation.
[0010] For at least the foregoing reasons, there exists a need to aid software developers and / or documentation engineers in the generation, provision, and / or maintenance of user interface documentation. Such need includes, without limitation, a need to facilitate and expedite the creation and maintenance of user interface documentation; a need to integrate the user interface documentation with the user interface; a need to connect users with software developers and / ordocumentation engineers to foster communication about questions and / or issues with the user interface documentation; and a need to reduce the complicated and arduous nature of tasks related to the creation and maintenance of the user interface documentation.SUMMARY
[0011] The present disclosure generally includes platforms, systems, methods, and devices that include automated generation, presentation, and maintenance of documentation for a user interface. More particularly, the present disclosure generally includes platforms, systems, methods, and devices that include automated generation, presentation, and maintenance of the documentation for performing a set of workflows through one or more websites.
[0012] In some example embodiments, a method includes recording, by a processor, a workflow during a capture of a task by a user through a user interface, wherein the workflow includes a sequence of steps of the task performed through the user interface; providing, by the processor, a visual guidance to a user to perform, in the user interface, each step of the sequence of steps of the workflow; and presenting, by the processor, at least one knowledge indicator in the user interface, wherein each knowledge indicator indicates an availability of knowledge related to the user interface.
[0013] In some example embodiments, a system includes a web browser process that executes within a context of a web browser of a device during a presentation of a user interface of a web page, wherein the web browser process is configured to, record, by a processor of the device, a workflow during a capture of a task by a user through the user interface of the web page, wherein the workflow includes a sequence of steps of the task performed through the user interface of the web page; provide, by the processor of the device, a visual guidance to a user to perform, in the user interface of the web page, each step of the sequence of steps of the workflow; and present, by the processor of the device, at least one knowledge indicator in the user interface of the web page, wherein each knowledge indicator indicates an availability of knowledge related to the user interface of the web page.
[0014] In some example embodiments, a method includes detecting, by a processor, a sequence of interactions by a user while performing a task within a user interface; identifying, by a processor, at least one element of the user interface that is associated with each interaction of the sequence of interactions; determining, by the processor, a workflow of the task based on the sequence of interactions performed by the user within the user interface; wherein the workflow includes a sequence of at least one step, and each step includes at least one action; and storing, by the processor, a record of the workflow of the task, wherein the record of the workflow is usable to inform users to perform the task within the user interface.
[0015] In some example embodiments, the user interface further comprises a website including at least one web page, each interaction by the user is associated with at least one content item of the at least one web page, and each content item includes at least one of, at least one content element included in at least one web page of the website, or at least one user control included in at least one web page of the website.
[0016] In some example embodiments, identifying the at least one element of the user interface that is associated with each interaction of the sequence of interactions includes, associating, by the processor, at least one event arising within the user interface with at least one event handler, detecting, by the processor, an invocation of an event handler in response to at least one event caused by at least one interaction of the user with the user interface, and determining, by the processor, at least one event detail associated with the invocation of the at least one event, wherein each interaction based on the at least one event detail associated with the invocation of the at least one event.
[0017] In some example embodiments, identifying the at least one element of the user interface that is associated with each interaction includes, identifying, by the processor, at least one user control within the user interface that is associated with an interaction, and recording, by the processor, at least one attribute of the at least one user control to identify the at least one user control within the user interface that is associated with the interaction.
[0018] In some example embodiments, identifying the at least one element of the user interface that is associated with each interaction includes, identifying, by the processor, a first content item within the user interface that is associated with an interaction, wherein the first content item is not responsive to the interaction, determining, by the processor, a second content item within the user interface that is associated with the first content item, wherein the second content item is responsive to the interaction, and recording, by the processor, at least one attribute of the second content item to identify at least one content item within the user interface that is associated with the interaction.
[0019] In some example embodiments, identifying the at least one element of the user interface that is associated with an interaction includes, responsive to detecting an interaction by the user while performing the task, identifying, by the processor, at least one user control within the user interface that is associated with the interaction, capturing, by the processor, a screenshot of the at least one user control, and recording, by the processor, the screenshot of the at least one user control within the user interface that is associated with the interaction.
[0020] In some example embodiments, capturing the screenshot of the at least one user control includes, intercepting, by the processor, at least one event arising within the user interface with at least one event handler, wherein the at least one event is associated with the interaction by theuser, capturing, by the processor, the screenshot of the at least one user control while intercepting the at least one event associated with the interaction, and after capturing the screenshot, re-raising, by the processor, the at least one event arising within the user interface.
[0021] In some example embodiments, identifying the at least one element of the user interface that is associated with each interaction includes, displaying, by the processor, a visual highlight associated with the at least one user control, and capturing the screenshot of the at least one user control includes removing, by the processor, the visual highlight associated with the at least one user control while capturing the screenshot of the at least one user control.
[0022] In some example embodiments, identifying the at least one element of the user interface that is associated with each interaction includes, presenting, by the processor, a side panel adjacent to the user interface, responsive to detecting an interaction by the user while performing the task, presenting, by the processor, a description of the interaction in the side panel, receiving, by the processor, at least one information item provided by the user as input to the description of the interaction in the side panel, wherein the at least one information item further describes the interaction, and recording, by the processor, the at least one information item provided by the user for the interaction.
[0023] In some example embodiments, the user interface is based on a document object model, and identifying the at least one element of the user interface that is associated with each interaction includes, identifying, by the processor, a portion of the document object model that describes at least one user control within the user interface that is associated with each interaction, and recording, by the processor, at least one attribute of the portion of the document object model, wherein the at least one attribute identifies at least one element within the document object model that defines the at least one user control.
[0024] In some example embodiments, the at least one attribute that identifies at least one element of the document object model includes at least one of, a location of the at least one element within a schema of the document object model, at least one attribute of the at least one element defined by the document object model, or at least one attribute of at least one user control that is associated with the at least one element of the document object model.
[0025] In some example embodiments, generating the workflow for the task includes, determining, by the processor, at least two consecutive interactions of the sequence of interactions by the user while performing the task within the user interface, wherein at least one interaction of the at least two consecutive interactions is redundant with at least one other interaction of the at least two consecutive interactions, and removing, by the processor, the at least one interaction from the sequence of interactions.
[0026] In some example embodiments, generating the workflow for the task includes, determining, by the processor, at least two consecutive interactions of the sequence of interactions by the user while performing the task within the user interface, wherein at least one interaction of the at least two consecutive interactions renders moot at least one other interaction of the at least two consecutive interactions, and removing, by the processor, the at least two consecutive interactions from the sequence of interactions.
[0027] In some example embodiments, generating the workflow for the task includes, determining, by the processor, at least two consecutive interactions of the sequence of interactions by the user while performing the task within the user interface, wherein the at least two consecutive interactions are associated with a user control within the user interface, and substituting, by the processor, an aggregated interaction in the sequence of interactions for the at least two consecutive interactions associated with the user control within the user control.
[0028] In some example embodiments, generating the workflow for the task includes, determining, by the processor, at least one first interaction of the sequence of interactions by the user while performing the task within the user interface, wherein the at least one first interaction is associated with a first user control within the user interface, determining, by the processor, at least one second interaction of the sequence of interactions by the user while performing the task within the user interface, wherein the at least one second interaction is associated with a second user control within the user interface, and substituting, by the processor, an aggregated interaction in the sequence of interactions for the at least one first interaction and the at least one second interaction, wherein the aggregated interaction is associated with both the first user control and the second user control.
[0029] In some example embodiments, a method includes receiving, by a processor, a workflow for a task, wherein the workflow includes a sequence of steps to be performed by a user within a user interface; presenting, by the processor, the sequence of steps of the workflow to the user during a performance of the task by the user; presenting, by the processor, a current step of the workflow, wherein presenting the current step guides the user to perform at least one interaction with at least one element of the user interface, and the at least one interaction corresponds to at least one action of the current step; and advancing, by the processor, to a next step of the workflow upon detecting a completion by the user of the current step of the workflow.
[0030] In some example embodiments, the user interface further comprises a website including at least one web page, respective steps in the sequence of steps are associated with at least one content item, each content item includes at least one of, at least one content element included in at least one web page of the website, or at least one user control included in at least one web pageof the website, and presenting the sequence of steps includes displaying, by the processor, the sequence of steps of the workflow in a side panel adjacent to the user interface.
[0031] In some example embodiments, the user interface is based on a document object model, and presenting the current step of the workflow includes, identifying, by the processor, a portion of the document object model that matches a description of a current step included in the workflow, and displaying, by the processor, a visual highlight of at least a portion of the user interface indicated by the portion of the document object model.
[0032] In some example embodiments, the description of the current step includes at least one attribute of at least one element of the document object model, the at least one element is associated with a first interaction by a user while performing the task within the user interface, and the first interaction corresponds to the current step of the workflow, and identifying the portion of the document object model that matches the description of the current step included in the workflow includes determining the portion of the document object model that matches the at least one attribute of the at least one element of the document object model.
[0033] In some example embodiments, the at least one attribute of the at least one element of the document object model includes at least one of, a location of the at least one element within a schema of the document object model, at least one attribute of the at least one element defined by the document object model, or at least one attribute of at least one user control that is associated with the at least one element of the document object model.
[0034] In some example embodiments, the description of the current step includes at least two attributes of at least one element of the document object model, and identifying the portion of the document object model that matches the description of the current step included in the workflow includes, performing, by the processor, a comparison of each attribute of the at least two attributes with at least one element of the document object model, determining, a score for each element of the document object model for each attribute of the at least two attributes, wherein the score is based on the comparison, and identifying, by the processor, the portion of the document object model that matches the description of the current step included in the workflow based on the portion of the document object model including an element having a highest sum of scores for each attribute of the at least two attributes.
[0035] In some example embodiments, the comparison of each attribute of the at least two attributes with at least one element of the document object model includes at least one of, comparing, by the processor, a label associated with at least one element of the document object model during a capture of the task to generate the workflow with a label associated with at least one element of the document object model, comparing, by the processor, at least one attribute included in at least one element of the document object model during a capture of the task togenerate the workflow with at least one attribute included in at least one element of the document object model during the current step of the performance of the task by the user, comparing, by the processor, a geometry of at least one element of the document object model during a capture of the task to generate the workflow within the user interface with a geometry of at least one element of the document object model during the current step of the performance of the task by the user, comparing, by the processor, a style selector associated with at least one element of the document object model during a capture of the task to generate the workflow with a style selector associated with at least one element of the document object model during the current step of the performance of the task by the user, or comparing, by the processor, a related element that is associated with at least one element of the document object model during a capture of the task to generate the workflow with a related element that is associated with at least one element of the document object model during the current step of the performance of the task by the user.
[0036] In some example embodiments, each attribute of the at least two attributes is associated with a weight, and the score for each element of the document object model for each attribute of the at least two attributes is further based on the weight associated with each attribute.
[0037] In some example embodiments, the method further includes, responsive to failing to identify a portion of the document object model that matches the description of the current step included in the workflow, presenting, by the processor, an alternative instruction associated with the current step.
[0038] In some example embodiments, the method further includes, responsive to failing to identify a portion of the document object model that matches the description of the current step included in the workflow, determining, by the processor, an alternative description for the current step, wherein the alternative description identifies at least one alternative content item within the user interface that is associated with the current step, and substituting, by the processor, the description for the current step within the workflow with the alternative description for the current step.
[0039] In some example embodiments, presenting the current step of the workflow includes, identifying, by the processor, a first content item within the user interface that is associated with the current step, wherein the first content item is not responsive to the current step, determining, by the processor, a second content item within the user interface that is associated with the first content item, wherein the second content item is responsive to the current step, and displaying a visual highlight of the second content item within the user interface.
[0040] In some example embodiments, the method further includes determining, by the processor, a completion of the current step of the performance of the workflow by a user; responsive to the completion of the current step, removing, by the processor, a visual highlight ofat least one content item within the user interface that is associated with the current step; and advancing, by the processor, the current step of the performance of the workflow to a next step of the performance of the workflow.
[0041] In some example embodiments, the user interface further comprises a website including at least one web page, each content item includes at least one of, at least one content element included in at least one web page of the website, or at least one user control included in at least one web page of the website, and presenting the sequence of steps of the workflow includes displaying, by the processor, the sequence of steps of the workflow in a side panel adjacent to the user interface.
[0042] In some example embodiments, a device configured to automatically generate documentation for a user interface includes a processor and a memory storing instructions that, when executed by the processor, cause the device to, determine a task to be performed within the user interface, determine a sequence of interactions by a user while performing the task within the user interface, record a set of details for each interaction of the sequence of interactions, wherein the set of details identifies at least one content item within the user interface that is associated with each interaction, generate a workflow for the task, wherein the workflow includes a sequence of descriptions generated by the processor for respective actions of the workflow, and each description for each action is based on the set of details associated with each interaction, and responsive to a request to describe the workflow for the task, present the sequence of descriptions generated by the processor.
[0043] In some example embodiments, the user interface further comprises a website including at least one web page, the sequence of interactions by the user further comprises a sequence of interactions respectively associated with at least one content item, each content item includes at least one of, at least one content element included in at least one web page of the website, or at least one user control included in at least one web page of the website, and execution of instructions by the processor further causes the device to present the sequence of descriptions in a side panel adjacent to the user interface.
[0044] In some example embodiments, a system for automatically generating documentation for a user interface includes a content process configured to, determine a task to be performed within the user interface, determine a sequence of interactions by a user while performing the task within the user interface, record a set of details for each action of the sequence of interactions, wherein the set of details for each action identify at least one content item within the user interface that is associated with each interaction, and generate a workflow for the task, wherein the workflow includes a sequence of descriptions generated by the content process for each action of the workflow, and each description for each action is based on the set of details associated with eachinteraction; and a display process configured to, responsive to a request to describe the workflow for the task, present the sequence of descriptions generated by the content process.
[0045] In some example embodiments, the user interface further comprises a website including at least one web page, the sequence of interactions by the user further comprises a sequence of interactions respectively associated with at least one content item, each content item includes at least one of, at least one content element included in at least one web page of the website, or at least one user control included in at least one web page of the website, and the display process is further configured to present the sequence of descriptions generated by the content process in a side panel adjacent to the user interface.
[0046] In some example embodiments, the system further includes a background process configured to, transmit the workflow to a web application, and responsive to the request to describe the workflow for the task, retrieve the workflow from the web application.
[0047] In some example embodiments, the system further includes a web application configured to, receive the workflow from the content process, store the workflow, and display the workflow in association with the user interface.
[0048] In some example embodiments, the display process is further configured to present the sequence of descriptions generated by the content process responsive to a selection within the web application of the workflow displayed in association with the user interface.
[0049] In some example embodiments, a method includes detecting, by a processor, a presentation of at least a portion of a user interface; presenting, by the processor, a knowledge indicator in the user interface; and responsive to detecting an interaction with the knowledge indicator, presenting, by the processor, knowledge associated with the knowledge indicator.
[0050] In some example embodiments, the knowledge indicator includes at least one user comment that is associated with at least one content item included in the user interface, wherein the at least one content item is associated with least one interaction of a sequence of interactions with the user interface, and presenting the knowledge indicator includes displaying, by the processor, the knowledge indicator adjacent to the at least one content item included in the user interface.
[0051] In some example embodiments, the knowledge indicator includes at least one user comment that is associated with at least one content item included in the user interface, wherein the at least one content item is associated with at least one interaction of a sequence of interactions with the user interface, and presenting the knowledge includes displaying, by the processor, the at least one user comment associated with the at least one content item.
[0052] In some example embodiments, the method further includes, responsive to receiving, from a user, at least one user comment associated with at least one content item included in theuser interface, wherein the at least one content item is associated with at least one interaction of a sequence of interactions with the user interface; storing, by the processor, the at least one user comment in a description for each interaction of the sequence of interactions; and associating, by the processor, the at least one user comment with the at least one content item included in the user interface.
[0053] In some example embodiments, at least one content item included in the user interface is associated with at least one conversation in a conversational interface associated with the user interface, and presenting the knowledge includes displaying, by the processor, the at least one conversation in the conversational interface associated with the user interface.
[0054] In some example embodiments, the method further includes, responsive to receiving, from a user, at least one user comment associated with at least one content item included in the user interface, wherein the at least one content item is associated with at least one interaction of a sequence of interactions with the user interface, adding, by the processor, the at least one user comment to a conversation in a conversational interface associated with the user interface.BRIEF DESCRIPTION OF THE DRAWINGS
[0055] The accompanying drawings, which are incorporated into and constitute a part of this specification, illustrate one or more example aspects of the present disclosure and, together with the detailed description, serve to explain their principles and implementations.
[0056] FIG. 1 is a diagram that illustrates examples of user experience including automated capture, guidance, and distribution of knowledge associated with a task according to aspects of the present disclosure.
[0057] FIG. 2 is a diagram that illustrates an architectural view of examples of a device configured for automated capture, guidance, and distribution of knowledge associated with a task according to aspects of the present disclosure.
[0058] FIG. 3 is a diagram that illustrates an architectural view of examples of a platform configured for automated capture, guidance, and distribution of knowledge associated with a task according to aspects of the present disclosure.
[0059] FIG. 4 is a flowchart that illustrates examples of a method of automatically generating user interface documentation according to aspects of the present disclosure.
[0060] FIG. 5 is a flowchart that illustrates examples of a method of generating a record of a workflow of a task performed by a user within a user interface, according to aspects of the present disclosure.
[0061] FIG. 6 is a diagram that illustrates examples of a scenario featuring a capture of a screenshot of an event arising within a user interface of a web browser of a device, according to aspects of the present disclosure.
[0062] FIG. 7 is a diagram that illustrates examples of a scenario featuring a merging of actions within a step of a workflow, according to aspects of the present disclosure.
[0063] FIG. 8 is a flowchart that illustrates examples of a method of presenting an automatically generated user interface documentation according to aspects of the present disclosure.
[0064] FIG. 9 is a diagram that illustrates examples of a scenario featuring a presentation of a current merged step of a workflow, according to aspects of the present disclosure.
[0065] FIG. 10 is a diagram that illustrates examples of a scenario featuring a weighted comparison of scores of features of various aspects of an element between versions of a document object model, according to aspects of the present disclosure.
[0066] FIG. 11 is a flowchart that illustrates examples of another method of presenting automatically generated user interface documentation according to aspects of the present disclosure.
[0067] FIG. 12 is a diagram that illustrates examples of a scenario featuring a presentation of knowledge related to a user interface, according to aspects of the present disclosure.DETAILED DESCRIPTIONOVERVIEW
[0068] FIG. 1 illustrates an example user experience including automated capture, guidance, and distribution of knowledge associated with a task according to aspects of the present disclosure.
[0069] The example user experience of FIG. 1 begins with a capture 102 of a workflow 126. A user 104 operates a device 106 having a display 108 that displays a web browser 110 to interact with a web page 112 including a user interface 114. The device 106 may be or may include, for example, a workstation, a thin-client device, a mobile device (e.g., a mobile phone, tablet, portable game console, or the like), a vehicle such as an automobile, or the like. The web page 112 may be delivered to the device 106 by a webserver 118 that provides a set of digital resources comprising the web page 112. The set of digital resources may indicate the content and layout of the web page 112 through a document object model 120, e.g., a hierarchical arrangement of elements in a language such as Hypertext Markup Language (HTML). The document object model 120 may include or specify, and / or the webserver 118 may send, additional digital resources that are included in the web page 112, such as images, recordings of sound, music, and / or video, embedded data, executable code in compiled and / or scripted form (e.g., Java, JavaScript, TypeScript, or the like), and / or formatting documents such as cascading style sheet (CSS) documents. The web browser 110 may render the web page 112 using the document object model 120 (e.g., determining and creating elements, determining and applying a layout, integrating and / or formatting additional resources such as images, and / or executing code in response to various events) in order to present the user interface 114 to the user 104. The user interface 114 of the web page 112 may include aset of interactive elements, such as buttons that can be clicked or checked, textboxes that receive text entries, lists that allow the user to select one or more elements, or the like. The interactive elements may enable additional functionality through the device 106, such as receiving files, images captured by a camera of the device 106, coordinates received through a positioning component of the device 106, or the like.
[0070] The web page 112 may enable the user 104 to perform a task 116 by interacting with the user interface 114 of the web page 112. For example, the web page 112 may comprise a web application that allows the user 104 to perform tasks 116 such as creating a user profile, creating or editing documents or graphics, interacting with other users, purchasing or selling items or currency, or the like. The user 104 may perform the task 116 by executing a sequence 122 of interactions 144 within the user interface 114. For example, if the task 116 includes creating an article in a blog, the sequence 122 of interactions 144 may include a first interaction 144 of selecting an item in a list of the user interface 114 (e.g., a type of blog article), a second interaction 144 of entering text into a textbox (e.g., creating the body of the article), and a third interaction 144 of clicking a Publish button to publish the article on the blog. If the task 116 includes communicating with another user in a social network, the sequence 122 of interactions 144 may include a first interaction 144 of clicking a button (e.g., a button that opens a messaging interface within the user interface 114 of the social network), a second interaction 144 of selecting an item in a list (e.g., a name of another user in the social network), a third interaction 144 of clicking an item in another list (e.g., a message sent by the other user in social network), and a fourth interaction 144 of entering text into a textbox (e.g., a response to the other user).
[0071] While performing the sequence 122 of interactions 144 to perform the task 116, the user 104 may wish to generate a workflow 126 (also referred to as a “user-directed UX workflow”) that documents the sequence 122 of interactions 144 to perform the task 116. For example, the user 104 may wish to generate a workflow 126 to be shared with other users 104 to show them how to perform the task 116 through the user interface 114. The user 104 may ask the device 106 to capture 102 a workflow 126 of the task 116 based on the sequence 122 of interactions 144 performed by the user 104. In response to a request and / or instruction to generate a workflow 126 for the task 116, the device 106 may monitor and analyze the sequence 122 of interactions 144 performed by the user 104. Based on the sequence 122 of interactions 144, the device 106 may generate a record of the workflow 126 as a sequence of steps 128, wherein each step 128 includes one or more actions 124 and a description 130 of each action 124. The device 106 may record various details of each step 128 of the workflow 126, such as an identification of one or more elements of the user interface 114 associated with the step 128 (e.g., an ID, name, label, caption, and / or other HTML attributes included in the element by the document object model 120); anidentification of one or more actions 124 performed on the one or more elements of the user interface (e.g., a first action of clicking a textbox to select it, and a second step of entering text into the textbox); and / or a screenshot of one or more of the actions 124 being performed for the step 128. The device 106 may record the workflow 126 as a guidance document that guides users to perform the task 116 through the user interface 114 of the web page 112.
[0072] The example user experience of FIG. 1 includes a visual guidance 132 of the workflow 126. For example, when a user 104 (e.g., the same user or another user) seeks to perform the task 116, the device 106 may present the web page 112 including the user interface 114 and a visual depiction of the workflow 126 adjacent to the user interface 114. The device 106 may present the workflow 126 as a sequence of steps 128 that correspond to the sequence 122 of interactions 144 by which the user 104 performed the task 116 during capture 102. The device 106 may visually associate each step 128 of the visual depiction of the workflow 126 with one or more elements of the user interface 114. For example, for a first step 128 of the workflow 126, the device 106 may insert, into the user interface 114 of the web page 112, a guidance indicator 136 that highlights an element 134 of the user interface 114 that is associated with the first step 128 (e.g., a colored box that surrounds and highlights a button to be clicked). The visual depiction of the workflow 126 may include a description of one or more actions 124 to be performed to complete the first step 128 of the workflow 126 (e.g., a suggestion to the user 104 to click a particular button of the user interface 114). The device 106 may monitor the interactions 144 of the user 104 with the user interface 114 to detect a performance of the one or more actions 124 associated with the first step 128 of the workflow 126. Upon detecting the performance of the one or more actions 124 of the first step 128, the device 106 may show a guidance indicator 136 that highlights another element 134 of the user interface 114 that is associated with a second step 128 of the workflow 126 (e.g., hiding a colored box that surrounds and highlights the first button, and showing a colored box that surrounds and highlights a second button that is associated with the second step 128) and / or a description of one or more actions 124 to be performed to complete the second step 128 of the workflow 126. The device 106 may continue to monitor the interactions 144 of the user 104 to determine the completion of each step 128 of the sequence of steps 128 of the workflow 126, including iteratively showing guidance indicators 136 and / or descriptions of a current step 128 and hiding the guidance indicators 136 and / or descriptions of a previous step 128 upon completion of the previous step 128. In this manner, the device 106 may perform visual guidance 132 to guide the user 104 through a task 116 through the visual depiction of the workflow 126 associated with the task 116.
[0073] The example user experience of FIG. 1 includes distribution 138 of knowledge 142. For example, various features of the user interface 114 may raise questions, concerns, issues, or thelike, and users 104 may request knowledge about various aspects of the user interface 114. Such knowledge may have been previously created by the user 104 or by other users (e.g., users from the same workspace or team), such as the steps of a created workflow 126 and / or messages created by the users concerning the task 116 and / or user interface 114. In order to convey knowledge to a user 104 about a web page 112, the device 106 may alter the web page 112 to include a knowledge indicator 140 that the user 104 may select, activate, or the like to receive knowledge 142 about the web page 112. If the knowledge 142 is related to a particular location and / or element 134 of the user interface 114, the device 106 may position the knowledge indicator 140 in or near the location and / or element (e.g., presenting the knowledge indicator 140 adjacent to and / or overlapping a portion of the location and / or element). If the user 104 selects, activates, or otherwise performs interactions 144 with the visual indicator 140, the device 106 may present knowledge 142 related to the user interface 114 of the web page 112 (e.g., a description of how to use a button, textbox, or other HTML element to perform one or more actions 124 of one or more steps 128 of a task 116). The knowledge 142 may be shown as a message, comment, or the like that is inserted into the web page 112. Alternatively or additionally, the knowledge 142 may be shown as a visual depiction associated with the element (e.g., a depiction of the sequence of steps 128 to perform a particular task 116 that includes the element). Through the performance of the distribution 138 of knowledge, the device 106 may provide knowledge 142 about the web page in order to educate and / or assist the user 104 to understand the web page 112 and / or the user interface 114 and / or otherwise receive knowledge 142 from other users.
[0074] FIG. 2 illustrates an architectural view of an example device 202 configured for automated capture, guidance, and distribution of knowledge associated with a task according to aspects of the present disclosure. The example device of FIG. 2 may perform some or all of the capture 102, visual guidance 132, and distribution 138 of knowledge about web pages as shown in the example user experience of FIG. 1.
[0075] In some example embodiments, a web browser 110 displays a web page 112 that presents a user interface 114, such as a set of user controls that a user may manipulate to invoke various functionality, such as entering or modifying data, viewing output, and / or invoking a functionality of the web page 112. The user interface 114 may be defined by a document object model 120 (DOM) in a document that is sent by a webserver 118. The document object model 120 may be provided, for example, as an HTML document, an XML document, a combination thereof, or other variations. The device 106 may communicate with the webserver 118 via a network (e.g., the Internet) to retrieve the document containing the document object model 120 of the web page 112 and store it in a memory of the device 106. The web browser 110 may interpret the content of the document containing the document object model 120 to render the user interface 114,including content items such as information (e.g., text, images, videos, sounds, or the like) and / or user controls (e.g., buttons, textboxes, lists, or the like) that the user manipulates to interact with the user interface 114.
[0076] The memory of the device 106 may stores a set of browser processes that can execute within a sandbox of the web browser 110 (e.g. , by a virtual machine that is executed by a processor of the device 106 to emulate the virtual machine and to execute the browser processes within the emulated virtual machine). In some embodiments, the browser processes may be implemented as a web browser extension that is executed by the web browser 110. The browser processes may include a content process 210 that is configured to generate documentation of the user interface 114 of the web page 112. In some embodiments, the content process 210 may include, or may be included in, one or more scripts embedded in, linked to, and / or otherwise associated with a web page. In particular, the content process 210 may be configured to determine a task to be performed within the user interface 114, determine a set of interactions 144 by a user 104 while performing the task within the user interface 114, and record (e.g., in the memory of the device 106) a set of details for each action of the set of actions. The set of details for each action identifies at least one content item within the user interface 114 that is associated with the action. The content process 210 may generate a workflow 126 for the task 116 that includes a sequence of steps 128 and corresponding descriptions 130 generated by the content process 210 for each action 124 of the workflow 126. The description for each step 128 may be based on the set of details associated with one or more actions 124 included in the step 128. While capturing the actions 124 of the task 116, the device 106 may display a list of the currently captured actions and initial descriptions thereof (e.g., a suggested title), which the user may edit. The device 106 may also automatically merge and / or delete the actions of the workflow during the capturing of the actions. The device 106 may also permit the user to edit, rearrange, and / or delete the actions of the workflow after the capturing of the actions.
[0077] In some embodiments, the device 106 also displays, adjacent to the web page 112, a side panel 224 that presents the sequence of actions captured by the content process 210. For example, a user 104 may request help for performing the task 116. In response, the side panel 224 (optionally by a display process included in the browser processes) may cause the workflow 126 to be displayed as a sequence of steps 128. The side panel 224 may determine a current step of a current performance of the workflow 126 by the user 104, wherein the current step 128 is included in the sequence of steps 128 of the workflow 126. The side panel 224 may display a visual highlight of at least one content item within the user interface 114 (e.g., at least one user control) that is associated with the current action. For example, the side panel 224 may display a current step as a title (e.g., “Click the Search Button”), a narrative description (“Use the mouse or a pointerto click the Search button of the web page”), and / or an orange rectangle around a button of the web page 112 as a visual highlight. The browser processes may also include a background process 212 that transmits (e.g., over the network 216) the workflow 126 to a workflow server 220. The workflow server 220 may store a set of workflows 126 that have been automatically generated for one or more user interfaces 114. When the device 106 requests and receives the document containing the document object model 120 for a particular web page 112 received from a webserver, the background process 212 may query the workflow server 220 for a workflow 126 that is associated with the web page 112. If the workflow server 220 transmits at least one workflow 126 for the web page 112, the device 106 may store the workflow 126 in the memory and may display the workflow 126 in the side panel 224 as a sequence of steps 128. The workflow server 220 may also transmit a list of workflows 126 that are accessible to the device 106. Upon selection of one of the workflows 126 in the list of workflows 126, the device 106 may present the sequence of steps 128 automatically generated for the workflow 126, whereby presenting the sequence of steps 128 may include overlaying or otherwise presenting GUI elements generated by the browser processes in relation to the web page 112 to guide the user 104 through the workflow 126.
[0078] More particularly, the example device 202 includes a processor 204, a display 108 that displays a web browser 110, and a memory 206. The memory 206 stores computer-executable instructions for various processes that the example device 202 may operate. The processor 204 executes the instructions on behalf of the example device 202, including instructions of a web browser process 208 that causes a web browser 110 to be presented on the display 108. The web browser process 208 allows a user of the example device 202 to navigate to various web pages (e.g., by inputting an Internet Protocol (IP) address or a uniform resource locator (URL) address string, by following a hyperlink embedded in a web page or message, and / or by activating a quick response (QR) code). In order to present a requested web page 112 of a website, the web browser process 208 identifies a webserver 118 indicated in a request and transmits a request for web page resources 218 to the webserver 118. The request may be transmitted over a network 216, including a wide-area network (WAN) such as the Internet, a Regional Area Network (RAN) such as a cellular network for mobile devices, a local-area network (LAN) such as an intranet of an organization or a local Wi-Fi network, or a personal area network (PAN) such as a personal device mesh. The webserver 118 may respond to the request by transmitting a set of web page resources 218 that the web browser process 208 may use to render the web page 112 for presentation within the web browser 110. The web page resources 218 may include one or more HTML, XML, and CSS documents; media objects such as images, sound or video recordings, or human-readable text; executable scripts in various scripted languages such as JavaScript and TypeScript; whollyor partially compiled binary objects, such Java runtime modules or libraries; or the like. The web browser process 208 may receive the web page resources 218, render the web page 112, and display the web page 112 in the web browser 110. The web page resources 218 may be organized according to a document object model 120, such as a hierarchical arrangement of elements and content of various types and having various attributes. The rendering step of the web browser process 208 may include determining a structure of the elements of the document object model 120 (e.g., translating an HTML document into an in-memory tree structure with a root element having sets of child nodes, each child node having one or more grandchildren nodes, etc.). The rendering step of the web browser process 208 may include determining a visual layout of the elements of the document object model 120 (e.g., allocating a visible and / or virtual space of the web page 112 shown in the web browser 110 to various elements of the document object model 120). The rendering step of the web browser process 208 may include generating, retrieving, formatting, executing, or otherwise processing each element of the web page 112 to generate a visual representation of the element and / or a functional description of its properties (e.g., a set of user events that may be invoked on each element of the document object model 120). The rendering step of the web browser process 208 may include aggregating the visual representations of the elements according to the determined layout and displaying the aggregated web page 112 within the web browser 110.
[0079] After rendering and displaying the web page 112, the web browser process 208 may update the web page (e.g., with periodic updates such as animations of images or graphics objects; in response to receiving additional resources and / or information from the webserver 118, such as AJAX or server-initiated events; in response to user input, such as a mouse click, tap, drag, or text input event; and / or in response to new information received from the example device 202, such as updates to location data provided by a positioning module). For example, the rendered web page 112 presents a user interface 114 comprising a set of elements 134, such as buttons, images, textboxes, and lists. Each element 134 of the user interface 114 presents a particular visual appearance and responds to a particular set of interactions 144 by the user 104; e.g., the button element 134 may respond to click or touch interactions 144, the textbox may respond to text entry interactions 144, and the list may respond to list item selection interactions 144. Some elements 134 may not respond to any interactions 144 (e.g., an image element 134 may be static and unresponsive). In order to detect and process user interactions 144 with various elements 134, the web browser process 208 may provide an event-based mechanism, where each element 134 indicates the types of events to which it responds, and may specify logic for handling such events. For example, a button element 134 may register with the web browser process 208 to handle click events, and may indicate and / or provide a handler function to be executed when the button element134 receives a click interaction 144. The web browser process 208 may receive raw user input (e.g., a button press of a mouse at a particular location of the display 108, or a touch interaction 144 at a particular location of the display 108), translate the raw user input into one or more events (e.g., generating a “click” event for the mouse button-press or touch interaction 144, wherein the click event indicates the location on the display 108), and execute any handler functions that correspond to the event (e.g., executing an “on click” handler function of a button that is located on the web page 112 at the location indicated in the event). In addition, the web browser process 208 may execute instructions encoded in one or more scripts, compiled binary objects, libraries, extensions, or the like. The instructions may register a function of a script, object, or the like with one or more events of the web browser process 208 (e.g., a handler function for page-load or drag events). In some embodiments, the instructions may register the function by calling a function of the web browser to request execution of the function, script, object, or the like upon occurrence of the event. When the web browser process 208 generates an event (e.g., in response to user input, logic associated with one or more elements 134, and / or signals from another process of the example device 106), the web browser process 208 may execute all handler functions that are registered with the event.
[0080] As discussed in FIG. 1, a user may wish to perform a task 116 through the web page 112 by interacting with the elements 134 of the user interface 114. The user may wish to capture a representation of the task 116 as a workflow, e.g., as documentation for how to perform the task 116, and / or to teach other users how to perform the task 116. While the user performs the task by interacting with the elements 134 of the web page 112, the web browser process 208 may record the sequence of interactions 144. The web browser process 208 may translate a recorded set of interactions 144 (e.g., a movement of a pointer across the display 108 while a mouse button is pressed, or a touch event across the display 108) to one or more actions 124 (e.g., a drag action from a first location to a second location). For each interaction 144, the web browser process 208 may generate one or more events (e.g., a drag event) and may execute one or more functions that are registered with the event (e.g., on drag function handlers that are registered with the drag event by one or more elements 134 of the user interface 114 and / or by one or more scripts or binary objects included in or associated with the web page 112). The web browser process 208 may also record the set of events as a sequence 122 of interactions 144 for performing the task 116. The web browser process 208 may evaluate the sequence 122 of interactions 144 to determine a workflow 126 comprising a set of steps 128, wherein each step 128 includes one or more actions 124. For example, the user interactions 144 or raw user input may include a set of keystrokes (e.g., pressing a Shift key, pressing a character key, and releasing the Shift key and the character key). The web browser process 208 may interpret the user interactions 144 or raw user input as one ormore actions (e.g., the entry of a capital letter as text input). Each action may cause one or more events to be raised within the web browser 110. In some cases, the web browser process 208 may disregard one or more events while generating actions; e.g., if the user presses and releases the Shift key or clicks in an empty space of a web page, the web browser process 208 may not generate actions for the events arising from such interactions 144. In some example embodiments, while recording the sequence 122 of interactions 144, the web browser process 208 may intercept events raised within the web browser 110 and may record information about each event that is determined to be part of the task 116 (e.g., an identification of one or more elements 134 associated with the event; the circumstances under which the event occurred, such as a timestamp of the event or a location of the event on the web page 112; and / or a screenshot of a portion of the web page 112 in which the event occurred). The web browser process 208 may further interpret a sequence 122 of interactions 144 (e.g., a sequence of events) to determine one or more steps 128 (e.g., actions comprising a set of keystrokes may be interpreted as a text entry event, where the entered text involves entering a string into a particular textbox element 134). The web browser process 208 may record the workflow 126 as a set of steps 128 for performing the taskl l8 through the user interface 114 of the web page 112.
[0081] In some example embodiments, the web browser 110 may record and document a task 116 by interacting with a workflow server 220. The workflow server 220 may provide a set of workflow resources 222 that enable the web browser 110 to capture 102 the sequence 122 of interactions 144 performed by the user 104 as a workflow 126 for the task 116. For example, the workflow server 220 may transmit, to the example device 202 over the network 216, a set of workflow resources 222 including a web application 214. The web application 214 may include executable code that causes the web browser 110 to generate, install, configure, and / or interact with a set of additional workflow resources 222. A user may configure the example device 202 for recording workflows 126 by transmitting, to the workflow server 220, a request for the workflow resources 222. For example, the web application 214 may install a browser extension of the web browser 110 that creates various workflow-related resources within the web browser 110, including a content process 210, and a background process 212, a web application 214, and a side panel 224, which may operate as follows.
[0082] For a particular web page 112 including a user interface 114 associated with a task 116 to be recorded, the web application 214 may cause the browser process 208 to execute a content process 210 (e.g. , a script or set of functions to be inj ected into and / or combined with the document object model 120 of the web page 112). The content process 210 may modify the content of the web page 112 by inspecting, altering, and / or supplementing of the web page 112 (e.g., creating new elements 134 in the document object model 120 of the web page 112, modifying and / orsupplementing existing elements 134, and / or registering handler functions for events arising within the web page 112). During capture 102, the content process 210 may perform an overlay interrupt that intercepts each event arising within the web page 112 that is associated with the task 116, capture 102 a screenshot of a portion of the web page 112 that is associated with the task 116, and record the event within a step 128 of the workflow 126 as an action 124 including a description (e.g., an identification of the element 134 of the web page 112 that is associated with the event). During visual guidance 132, the content process 210 may inspect the document object model 120 of the web page 112 to determine one or more elements 134 with which the user may perform interactions 144 to perform a current step 128 of the workflow 126 and may inject a guidance indicator 136 that visually highlights the determined one or more elements 134. During distribution 138, the content process 210 may inject one or more knowledge indicators 140 into the web page 112 to indicate the availability of knowledge 142 related to an element 134 or portion of a web page 112, and / or may display knowledge 142 in response to the selection of a knowledge indicator 140 by the user.
[0083] The web browser 110 may include, install, and / or present (e.g., as part of the browser extension), a side panel 224 that provides a graphical user interface for the capture 102, visual guidance 132, and / or distribution 138 of knowledge related to a task 116. The side panel 224 is shown within the web browser 110 alongside a web page 112 for which capture, guidance, and / or distribution of knowledge relating to a task 116 is performed. During capture 102, the side panel 224 may show each captured step 128 of the workflow 126 as the task 116 is being performed by the user. For example, when the content process 210 performs an overlay interrupt to intercept an event and determines that the event is associated with a new step 128 of the workflow 126, the side panel 224 may update the presented workflow 126 to include a new item for the new step 128. The side panel 224 may also show the sequence of steps 128 captured so far, a screenshot and / or thumbnail of each step 128 (e.g., a visual depiction of the element 134 associated with the step 128), and / or a description of each step 128 (e.g., a determined or given name of the step 128). The side panel 224 may allow the user to modify the workflow 126 during or after the recording of the workflow 126 (e.g., renaming, reordering, merging, and / or deleting steps 128 of the workflow 126). During visual guidance 132, the side panel 224 may show the complete sequence of steps 128 included in the workflow 126, and may highlight and / or describe a current step 128 of the workflow 126 that the user is to perform and / or is currently performing. During distribution 138, the side panel 224 may show knowledge 142 that is associated with knowledge indicators 140 included in the web page 112 and selected by the user (e.g., messages exchanged between users about a particular aspect of the task that relates to a particular element 134 and / or portion of the web page 112).
[0084] The browser extension may also install, within the web browser 110, a background process 212 that performs background processing related to a workflow 126. The background processing performed by the background process 212 may include processing of data associated with the workflow 126 (e.g., recording, analysis, and / or transformation of screenshots of portions of the web page 112) and / or background communication with the workflow server 220 (e.g., the transmission of data related to capture 102, visual guidance 132, and / or distribution 138 of knowledge associated with the task 116). In some example embodiments, the background process 212 may comprise one or more web workers that persistently reside in the browser process 208 and await instructions to perform background processing associated with a task 116. The designation of a background process 212 (such as a web worker) to perform the background processing may enable the execution of functions that take a significant amount of time and / or processing without affecting the behavior the web page 112. For example, the background process 212 may execute in a separate process or thread of the browser process 208 to execute the background processing separately from the web page 112, so that the web page 112 continues to perform periodic functions (e.g., animation of images), process user input (e.g., receiving and evaluating pointer and touch movement), and the determination and execution of user interactions 144 (e.g., raising events and executing handler functions that are registered with the events). During capture 102, the background process 212 may capture images that include screenshots of portions of the web page 112 associated with various steps 128 of the workflow 126; process the images (e.g., resizing, reformatting, and / or annotating the screenshots); capture and / or evaluate snapshots of the document object model 120; generate and / or maintain the workflow 126 as a record of the task 116; and / or transmit the workflow 126, images of screenshots, and / or snapshots of the document object model 120 to the workflow server 220. During visual guidance 132, the background process 212 may request and receive the workflow 126 from the workflow server 220; request and receive images of screenshots of various steps 128 of the workflow 126; and / or transmit user interactions 144 to the workflow server 220 (e.g., as telemetry for the performance of the task 116). During distribution 138, the background process 212 may exchange messages with the workflow server 220 related to various knowledge 142 associated with the task 116 (e.g., transmitting messages created by the user that include and / or relate to knowledge 142 associated with the task 116, and / or receiving from the workflow server 220 messages from other users that that include and / or relate to knowledge 142 associated with the task 116).
[0085] In addition to installing the browser extension (e.g., the content process 210, the side panel 224, and / or the background process 212), the web application 214 may provide additional functionality to the example device 202 on behalf of the workflow server 220. For example, the web application 214 may authenticate the user with the workflow server 220 (e.g., via usercredentials such as usernames, passwords, passcodes, and / or certificates; multi-factor authentication including device authentication; and examination of biometric features of the user). The web application 214 may update and / or maintain the workflow resources 222 installed on the example device 202 (e.g., replacing versions of the browser extension with updated versions of the browser extension).
[0086] While FIG. 2 presents an architecture of an example device 202 that is configured for automated capture, guidance, and distribution of knowledge associated with a task 116, it is to be appreciated that many other architectures may be suitable for automated capture, guidance, and distribution of knowledge associated with a task 116 in accordance with the present disclosure. Such other architectures may organize the functionality of the automated capture, guidance, and distribution of knowledge in a different manner than as shown in the example of FIG. 2. As a first example, rather than working with a web page 112 provided by an external webserver 118 and / or workflow resources 222 provided by an external workflow server 220, a device may work with a local webserver 118 and / or a local workflow server 220 to receive the web page 112 and the resources for automated capture, guidance, and distribution of knowledge associated with a task 116. Alternatively or additionally, at least some functionality that is shown as being locally deployed on the example device 202 of FIG. 2 could instead be remotely executed (e.g., a localized view of a remotely rendered web page 112 and remotely executed processes for capture 102, visual guidance 132, and / or distribution 138 of knowledge 142 associated with a task 116, such as a thin- client view of remote processing that is executed on a remote device such as a remote webserver and / or in a cloud). As a second example, rather than representing the workflow resources 222 as a web application 214 that executes within a browser process 208, a device may install a native local application that directly performs interactions 144 with a web browser 110 using native processes that execute independently of any browser processes 208. As a third example, rather than organizing the workflow resources 222 as a locally executing web application 214, a content process 210, a side panel 224, and a background process 212, a device may organize the functionality of automated capture 102, visual guidance 132, and / or distribution 138 of knowledge 142 in a different manner, such as providing all functionality in one module (e.g., a specialized web browser 110) and / or further dividing the functionality into different modules (e.g., a first background process 212 that performs image processing and a second background process 212 that performs communication with the workflow server 220). These and other alterations of the organization of the resources for automated capture 102, visual guidance 132, and / or distribution 138 of knowledge 142 may be included in various embodiments of the techniques presented herein.
[0087] In some example embodiments, a platform may a suite of tools for documenting user- directed user experience (UX) workflows, guiding users through the user experience workflows, and facilitating collaboration by the users "on top of' third-party GUI's of web applications. User- directed user experience workflows may refer to an object that defines a series of steps (each of which may be comprised of one or more actions) that were performed by a user during a user- initiated "capture" session. The platform may be a web browser extension for documenting user- directed UX workflows, guiding users through the user experience workflows, and facilitating collaboration by the users "on top of' third-party web application. In some example embodiments, the web extension is configured to monitor a user-initiated capture session, capture screenshots when the user performs an action with respect to a GUI, and generates an annotated screenshot that indicates what action the user took. In some embodiments, the web extension is further configured to guide a subsequent user through a user-directed UX workflow by visually highlighting the interactive elements in the GUI that the user needs to perform interactions 144 with to complete the next step.
[0088] FIG. 3 illustrates an architectural view of an example platform 302 configured for automated capture, guidance, and distribution of knowledge associated with a task according to aspects of the present disclosure.
[0089] The example platform 302 of FIG. 3 includes a processor 204 and storage 304. The storage 304 includes a database 306 that stores information about one or more user interfaces 114 (e.g., a URL or URI of a web page 112 including a user interface 114 or of a native application including a user interface 114) associated with one or more workflows 126 for performing one or more tasks 116 through the user interface 114. The storage 304 may organize the user interfaces 114 and workflows 126 and the like in one or more databases to improve the responsiveness of the example platform 302 in fulfilling requests for the workflows 126. The storage 304 may include workflow resources 222, such as deployable components that may execute on devices 106, knowledge about the user interfaces 114 and / or workflows 126, and application programming interfaces (APIs) that may be used to access various third-party services. The storage 304 may also include service information 308 that the example platform 302 may use to provide workflows 126 as a service, such as information about users 104 (e.g., names and login credentials), devices 106 (e.g., device types and associations with users 104), and security information such as definitions of “rooms” and / or “workspaces” that particular users 104 may be permitted to access.
[0090] The example platform 320 includes a web application 214 that communicates with users 104 and / or devices 106 to perform auxiliary and / or support functions. For example, the web application 214 may allow users 104 to register with the example platform 302 by creating a user account and to be authenticated by the example platform 302. The web application 214 may beaccessed by a device 106 as a web page 112 shown in a web browser 110 of the device 106, as a native application executing on the device 106 (e.g., a desktop or mobile app), or as a combination thereof. The web application 214 may be at least partially installed on the device 106 (e.g., in a web cache of the web browser 110, or as a locally installed application) and / or may store or cache data on the device 106 (e.g., a local repository of workflows 126 for various user interfaces 114). A portion of the web application 214 executed by the platform 302 may communicate with another portion of the web application 214 executed by a device 106 (e.g., via API calls and / or HTTP / HTTPS requests) to provide the workflow service on the device 106. The web application 214 may receive one or more workflows 126 generated by a device 106 for a user interface 114 (e.g., for a user interface 114 included in a web page 112) and may store the workflow 126 and related information in the database 306. The web application 214 may receive, from the device 106, a request for one or more workflows 126 to perform a particular task 116, either in general or specifically with regard to a particular user interface 114, and may provide one or more workflows 126 retrieved from the database 306 to fulfill the request. The web application 214 may also perform security checks, e.g., verifying that a user 104 and / or device 106 is permitted to access a particular workflow 126 before transmitting the workflow 126 to the user 104 and / or device 106.
[0091] In some example embodiments, a platform includes a web browser extension that comprises a set of scripts that are executed by the web browser, including a content script that provides a set of DOM processing capabilities of the DOM processing system by a client device / web browser that executes the web browser extension. The platform may include a DOM processing system that provides a set of tools for processing a DOM. This includes monitoring the DOM for potential interactions, identifying interactions 144 with interactive elements / elements of a website, capturing target details pertaining to identified action (e.g., for guidance or knowledge distribution), identifying interactive elements based on previously captured target details, and / or the like.
[0092] The example platform 302 includes a service deployment module 310 that deploys workflow resources 222 to devices 106. For example, the workflow resources 222 may include a deployable web browser extension that a device 106 may request, receive, and install to enable the presentation of workflows 126 within user interfaces 114 shown in the web browser 110 of the device 106. The web browser extension may create a web browser process 208 that performs various workflow-related functions on the device 106, such as performing a capture 102 of a workflow 126 for a task 116 performed by a user 104; presenting a workflow 126 of a user interface 114 as visual guidance 132 to aid the user 104 in performing the task 116 through the user interface 114; and / or integrating knowledge 142 with user interfaces 114 for various tasks116 and / or workflows 126. The web browser extension may also create a background process 212 (e.g., a web worker process) that performs background processing, such as communicating with the example platform 302 to request, receive, and / or transmit workflows 126 and related resources such as screenshots and snapshots of document object models of web pages 112. The web browser extension may also create a content process 210 (e.g., as a script that is wholly or partially added to and / or imported by one or more web pages 112 presented in the web browser 110). The content process 210 may provide event handlers to detect events arising within the web browser 110 and update the document object model 120 of aweb page 112 to include additional elements 134 (e.g., a highlight of an element 134 associated with a current step 128 of a workflow 126). The service deployment module 310 may also provide at least a portion of the web application 214 to a device 106 (e.g., deploying at least a portion of the web application 214, such as a native application or a web browser extension) to provide the workflow service on the device 106.
[0093] In some example embodiments, a platform includes a Document Object Model (DOM) processing system for receiving and analyzing a DOM corresponding to a website or web application being visited by a user to detect user interactions 144 with certain objects within a graphical user interface (GUI) of the website or web application, classify the respective types of the objects being interacted with by the user, locate elements within the website or web application that were previously interacted with by another user, and / or manipulate the DOM to overlay certain types of UI display elements over the GUI of the website or web application. In some example embodiments, the DOM processing system provides a set of tools for processing a DOM. This includes monitoring the DOM for potential interactions 144, identifying interactions 144 with interactive elements / elements of a website, capturing target details pertaining to identified action (e.g., for guidance or knowledge distribution), identifying interactive objects based on previously captured target details, and / or the like. In some example embodiments, the platform may include a backend system that provides a set of services for enhancing and sharing user-directed user experience (UX) workflows users "on top of' third party web applications. The backend system may provide various services for the web extension to enhance the capture and generation of user- directed UX workflows, such as processing DOM snapshots, virtualization services, increasing resolution of screenshots used in a UX workflow, allowing users to edit the content in a particular step of a workflow, applying governance to the screenshots that were captured, editing the bounding boxes, generating descriptions / labels, processing DOM snapshots, and / or the like. The backend system may include a snapshot processing system configured to execute, analyze, and / or manipulate snapshots of a DOM captured by a respective user device to generate editable screenshots corresponding to respective actions taken by a user of third party website during a user-initiated UX capture session. In some example embodiments, a client (e.g., web extension)sends snapshots of the DOM to the backend and the backend executes a "headless" browser that renders the web-based GUI. This allows the platform more flexibility to manipulate various aspects of the platform (e.g., higher resolution, editable screenshots). This architecture may be expanded to support the virtualization of corresponding native applications, which may enable the creation of workflows that depict corresponding steps / actions of a user-defined UX workflow also for the relating of native application GUIs.
[0094] The example platform 302 also includes a workflow analysis module 312 that generates, stores, and / or analyzes workflows 126 to provide the workflow service. For example, a device 106 may perform a capture 102 of a set of actions 124, but may not have enough computational resources to generate the workflow 126, including processing a set of captured screenshots and / or videos of various actions 124 and / or translating interactions 144 performed by the user 104 into the steps 128 and actions 124 of a workflow 126. Instead, the device 106 may transmit a set of workflow resources 222 to the workflow analysis module 312, which may process the workflow resources 222 to generate a workflow 126 that may be stored in the database 306 and associated with the user interface 114. The workflow analysis module 312 may also receive and analyze document object models 120 of web pages 112 to identify elements 134 of the user interface 114, including changes to the document object model 120 over time that alter the logical representation and / or behavior of the elements 134 in the web page 112. The workflow analysis module 312 may also receive and process screenshots, e.g., to add annotations of steps 128 for presentation during visual guidance 132 and / or as knowledge 142 of the user interface 114, and / or to generate user- editable screenshots. The workflow analysis module 312 may also perform maintenance tasks of the workflows 126, such as updating a workflow 126 to maintain compatibility with a user interface 114 as the document object model 120 of the web page 112 changes. The workflow analysis module 312 may also render and / or analyze user interfaces 114 within one or more virtualization mechanisms, such as a headless web browser 110 that renders a web page 112 to determine a layout of elements 134 and / or to capture screenshots, rather than showing the rendered web page 112 on a display 108 of a device 106 for a user 104.
[0095] It is to be appreciated that FIG. 3 illustrates only one example of an architecture of a platform 302 for providing a workflow service to various devices 106. In various embodiments, a platform may be centralized (e.g., providing all functionality on one workflow server 220 or local cluster of workflow servers 220) and / or distributed (e.g., executing different portions and / or providing different modules on different workflow servers 220, which may be deployed in different geographic regions). In various embodiments, a platform 302 may organize resources in one or more databases 306 and / or in different ways, such as document-oriented stores or file systems. In various embodiments, the functions of the platform may be entirely or primarilyexecuted by a workflow server 220 (e.g., as a fully or primarily remote or cloud-based architecture) and / or may be primarily or at least partially executed on a device 106 (e.g., as a locally installed workflow service or application). In various embodiments, the platform 302 may provide the workflow service on a public basis, in which user interfaces 114 and / or workflows 126 are publicly available; on a semi-private basis, in which user interfaces 114 and / or workflows 126 are available only to users 104 and / or device 106 that meet certain criteria (e.g., subscribers to a paid service); and / or on a private basis, in which user interfaces 114 and / or workflows 126 are available only to specific users 104 and / or devices 106 (e.g., those that are associated with a private and / or governmental organization).
[0096] In some example embodiments, a platform 302 may integrate with and / or include a variety of external tools and / or information sources. Such external tools and / or information sources may be associated with a same party that provides the platform 302 and / or one of the device 106, or may be a third party. As a first such example, a platform 302 may integrate with an email provider and / or service to communicate with users 104, devices 106, and / or user interfaces 114 via email (e.g., integrating an email-based multi-factor authentication process to authenticate users 104 of the platform 302). As a second such example, a platform 302 may integrate with a third-party provider of a web page 112 to coordinate the presentation of workflows 126 related to user interfaces 114 included in the web page 112 (e.g., requesting web pages 112 directly from the third-party provider, analyzing the web page 112 to generate workflows 126, and transmitting the workflows 126 for the web page 112 to the third-party provider, so that the workflows 126 are initially included in, referenced by, and / or linked in the web page 112 when delivered by the third- party provider to users 104 and / or devices 106). As third such example, a platform 302 may integrate with a third-party information database, internet forum, chat service, or the like to receive information about a user interface 114. The platform 302 may use the information about the user interface 114 as knowledge 142 that may be presented in and / or integrated with the user interface 114 (e.g., knowledge 142 that may be presented to a user 104 upon interaction with a knowledge indicator 140 attached to the user interface 114). As a fourth example, a platform 302 may integrate with a third-party private data service, such as a team-based communication or chat tool, to include private communications by one or more users 104 of a private organization as workflows 126 and / or knowledge 142 of a user interface 114 that is accessed by members of the private organization (e.g., workflows 126 and / or knowledge 142 that is limited to members of a workspace associated with the private organization). These and other integrations may be achieved by communication between the platform 302 and third-party services (e.g., through APIs provided by the third-party services and incorporated into the platform 302, and / or by APIs provided by the platform 302 and incorporated into at least one of the third-party services and / ora device 106 of a user 104). In this manner, the platform 302 may integrate with various third- party information sources and services to extend the capabilities of the workflow service.
[0097] FIG. 4 illustrates a flowchart of an example method 400 of automatically generating user interface documentation according to aspects of the present disclosure. At least a portion of the method 400 may be performed, for example, by the example device 202 of FIG. 2.
[0098] The example method 300 includes a step 402 of recording a workflow of a sequence of steps of a task performed through a user interface during a capture of the task of the user through the user interface. The capturing may result in a user-directed workflow of the user performing the task through the user interface, wherein the workflow includes a sequence of steps of the task that are performed through the user interface. Each step of the sequence 122 may include one or more actions 124, such as entering text into a textbox, clicking on a drop down menu, selecting a button, expanding a popout, and / or the like. Each action 124 may be associated with and / or determined by one or more events raised within the user interface 114 in response to one or more interactions 144 of the user 104 with the user interface 114, such as click events, touch events, type events, or the like. That is, a set of interactions 144 may result in a set of events that together comprise an action 124 (e.g., entering information into a textbox), and a set of actions 124 may together comprise a step 128 of the workflow 126 (e.g., entering information into a set of textboxes to complete a form). The device 106 may detect the events in various ways (e.g., receiving notifications of the events from the device 106 and / or attaching event handlers to the events). The device 106 may record the sequence of actions 124 as a workflow 126.
[0099] In embodiments, the example method 300 includes a step 404 of providing visual guidance 132 to a user 104 to perform, in the user interface 114, each step 128 in the sequence of steps 128 of the workflow 126. In some scenarios, the user 104 of step 404 may be the same user 104 who performed the task 116 in the user interface 114 during the step 402 (e.g., in order to allow the user 104 to review the workflow 126 and / or to remind the user 104 how to perform the task 116). In other scenarios, the user 104 of step 404 may be different than the user 104 who performed the task 116 in the user interface 114 during the step 402 (e.g., another employee of an organization; a family member, friend, or student of the user 104 who performed the task 116 in the user interface 114 during the step 402; and / or another user 104 of the user interface 114). The visual guidance 132 of the user 104 through the workflow 126 may include presenting guidance indicators 136 for a current step 128 of the workflow 126, and, upon detecting completion of the current step 128 of the workflow 126, replacing the guidance indicators 136 for the current step 128 with guidance indicators 136 for a next step 128 of the workflow 126. The device 106 may present guidance indicators 136 until the user 104 has completed the sequence of steps 128 of the workflow 126.
[0100] The example method 400 includes a step 406 of presenting at least one knowledge indicator 140 in the user interface 114, wherein the knowledge indicators indicate an availability of knowledge 142 that is related to the user interface 114. For example, when a user 104 views a web page 112, the device 106 may insert one or more knowledge indicators 140 at various locations of the web page 112, such as locations of relevant information or of elements 134 of the user interface 114. The knowledge indicators 140 may appear as icons, pictures, buttons, visual highlights, or the like, and may indicate an availability of an interaction 144 (e.g., an indication that the knowledge indicator 140 is clickable, touchable, selectable, or the like). When the user 104 selects one of the knowledge indicators 140, the device 106 may insert, into the web page 112, additional information that presents the knowledge 142. The insertion may include inserting a window or overlay that includes text, such as messages, comments, visual indicators, videos, a dialogue between two or more users, a chat interface among a group of users, or the like, wherein the knowledge 142 conveyed by the insertion informs the user 104.
[0101] It is to be appreciated that in FIG. 4, the steps 402 of knowledge capture and the step 404 of guidance operate synergistically as documentation of tasks 116 performed in the user interface 114. For example, the step 402 of knowledge capture results in a recording of a workflow 126 of steps 128 for performing the task 116 based on the observation and recording of the interactions 144 of the user 104 while performing the task 116. The recording of the workflow 126 is then used to provide the step 404 of visual guidance 132 to the same user 104 and / or other users 104 to perform the task 116 in a similar manner as when the user 104 performed the task during the step 402 of knowledge capture. Thus, the step 402 of knowledge capture and the step 404 of guidance in the example method 400 of FIG. 4 work together to automate user interface documentation and propagation thereof of tasks 116 performed within the user interface 114 in accordance with the techniques presented herein.
[0102] Further, it is to be appreciated that the steps of FIG. 4 may be performed in various circumstances. In some example embodiments, the steps of FIG. 4 may be performed by one device 106, such as a mobile device featuring a web browser 110 that both records the workflow 126 of the task 116 performed by a user 104, provides visual guidance 132 of the workflow 126 to the same user 104 and / or other users 104, and performs distribution 138 of knowledge 142 by including knowledge indicators 140 in the user interface 114. In some example embodiments, the steps of FIG. 4 may be performed by a first device that observes interactions 144 of one or more users 104 performed on one or more other devices, such as a cloud server that receives reports of interactions 144 performed by users 104 within the web browsers 110 of one or more remote devices, records a workflow 126 based on the observed actions, and causes knowledge indicators 140 be inserted into user interfaces 114 presented on the remote devices. In some exampleembodiments, the steps of FIG. 4 may be performed by different devices, such as a first device 106 of a first user that records a workflow 126 while the first user 104 performs a task 116 on a user interface 114 presented on the first device; a second device 106 that provides visual guidance 132 to a second user 104 to perform the task on the same or similar user interface 114 presented on the second device 106; and a third device 106 that causes knowledge indicators 140 of the task 116 to be included in a user interface 114 presented on the third device 106. In some example embodiments, one or more steps may be performed by two or more devices 106 operating together, e.g., a mobile device 106 of a user 104 and a remote serve of a platform (such as a workflow server 220) may together evaluate the interactions 144 of a user 104 within the user interface 114 to determine workflows 126 of tasks 116 performed within a user interface 114 of the mobile device 106. In some example embodiments, the user interface 114 is included in a web page 112 presented by a web browser 110, such as shown in the architectural diagram of FIG. 2. In other example embodiments, the user interface 114 is included in a context other than a web page 112, such as a user interface of an application that is natively executing on a device 106 and / or that is natively executed by a first device 106 and then shown on a second device 106. In some example embodiments, a user 104 who is performing the task 116 also participates in the capture 102 and recording of a workflow 126, the visual guidance 132 provided by the device 106, and / or the inclusion of knowledge indicators 140. In some example embodiments, different users 104 are respectively involved in at least one of the performances of the interactions 144, the recording of the workflow 126, the presentation of visual guidance 132, and / or the inclusion of knowledge indicators 140 to achieve a distribution 138 of knowledge 142 associated with the task 116. For example, a first user 104 may perform interactions 144 that correspond to the actions 124 of a workflow 126 while a second user 104 indicates which interactions 144 are associated with a task 116 and / or creates, annotates, arranges, comments on, and / or otherwise participates in the generation and / or curation of the workflow 126. For instance, a teacher may cause the device 106 to generate a workflow 126 and may participate in the device 106 providing visual guidance 132 based on the recorded workflow 126. Many such variations of the context in which workflows 126 are generated (e.g., by capture 102) and / or used (e.g., by providing visual guidance 132 and / or by causing a distribution 138 of knowledge 142) are included in the techniques presented herein. KNOWLEDGE CAPTURE
[0103] In some example embodiments, a device 106 may include a UX capture system that captures a series of actions performed by a user with respect to a third-party website during a user- initiated capture session, wherein the capture system determines a series of actions taken by a user in performance of a task and highlights respective interactive elements that the user of the third- party website interacted with during the capture session. The platform includes a document objectmodel 120 processing system that provides a set of tools for processing a document object model 120. This includes monitoring the document object model 120 for potential interactions 144, identifying interactions 144 with interactive elements of a website, capturing target details pertaining to identified action (e.g., for guidance or knowledge distribution), identifying interactive elements based on previously captured target details, and / or the like. In some example embodiments, the device 106 may include an interaction monitoring component that monitors a document object model 120 of a third party website during a user-initiated capture session to identify respective interactive elements of the third-party website with which the user is engaging and identifies actions taken by the user via a GUI of the website during the capture session; and an action capture component that captures respective snapshots and / or screenshots of the GUI that collectively indicate a series of respective interactive elements that the user interacted with to undertake the actions identified by the interaction monitoring system. In embodiments, a UX capture system monitors the interactions 144 of the user 104 with the GUI to determine when the user has potentially finished a step (e.g., filling out a box, selecting a checkbox, selecting an option from a drop down menu, determine when to capture a screenshot (or DOM snapshot). When the capture session determines that a potential action has been performed, the capture system captures the step / action taken by the user. This may include capturing a screenshot and / or capturing a DOM snapshot. The capture system continues to monitor the interactions 144 of the user 104 during the capture session such that it captures a series of steps in a user-directed UX workflow. The device 106 may include a capture UI that allows a user to initiate a capture session that documents a series of actions performed by a user with respect to a third-party website and outputs visual indicia over interactive elements that the user of the third-party website performs interactions 144 with the third party -website during the capture session.KNOWLEDGE CAPTURE - OVERVIEW
[0104] FIG. 5 illustrates a flowchart of an example method 500 of generating a record of a user- directed UX workflow of a task while being performed by a user within a user interface, according to some aspects of the present disclosure.
[0105] In general, the example method 500 of FIG. 5 enables the device to perform an automated capture of a representation of the workflow for a task. In some example embodiments, during a “Capture” session, the device 106 may monitor interactions 144 of a user 104 with a user interface 114 presented by a web browser 110 to identify a sequence 122 of interactions 144 of the user 104 to complete a task through the user interface 114. The device 106 may monitor events that are raised within the web browser 110 in response to interactions 144 of the user 104 with respect to the user interface 114 (e.g., click, scroll, drag, and text entry events in response to various touch, mouse, and / or keyboard input from the user 104) to determine interactions 144 that wereperformed by the user (e.g., clicking on an icon, entering text, selecting a drop down menu, etc.). The device 106 may automatically detect, capture, describe, annotate, and optionally modify the sequence 122 of interactions 144 performed by the user 104. During subsequent “Guidance and Knowledge Distribution” phases, the device 106 may present a workflow 126 as a sequence of steps 128, each step 128 including one or more actions 124 and / or a description 130 of the step 128 and / or one or more actions 124. The description 130 may include guidance indicators 136 that guide the user 104 through the performance of the steps 128 of the workflow 126, such as visual highlights of elements 134 of the user interface 114 that are related to a current step 128 of the workflow 126. This presentation of information based on automatically captured documentation may inform, aid, expedite, and / or remind the user 104 how to perform the task 116 through the workflow 126.
[0106] The example method 500 includes a step 502 of detecting a sequence 122 of interactions 144 by a user 104 while performing a task 116 within the user interface 114. The example method 500 includes a step 504 of identifying, for each interaction 144 of the sequence 122, at least one element 134 of the user interface 114 that is associated with the interaction 144. The example method 500 includes a step 506 of determining a workflow 126 of the task 116 based on the interactions 144 performed by the user 104 within the user interface 114, wherein the workflow 126 includes a sequence of at least one step 128, and each step 128 includes at least one action 124. The example method 500 includes a step 508 of storing a record of the workflow 126, wherein the record of the workflow 126 is usable to inform users to perform the task 116 within the user interface 114.
[0107] In some example embodiments, the step 502 of detecting a sequence of interactions 144 performed by a user while performing a task within a user interface may include detecting events that are raised by the interactions 144 of the user while performing the task 116, such as click events caused by clicking on buttons, text caused by entering a character of text into a textbox, and / or drag events that are caused by dragging objects within user interface. The interactions 144 by the user 104 may be respectively associated with at least one content item. Each content item includes at least one of at least one content element included in at least one web page of the website or at least one user control included in at least one web page of the website. The device 106 may associate at least one event arising within the user interface with at least one event handler. The device 106 may detect an invocation of the event handler in response to at least one event caused by at least one interaction 144 of the user 104 with the user interface. Based on the invocation, the device 106 may determine at least one event detail associated with the invocation of the at least one event, such as a name, class, or ID of a user control associated with the event; a location within the document object model of an element associated with the event; or a geometric property of anelement of a web page associated with the event. The set of details for the action is based on the at least one event detail associated with the invocation of the at least one event.
[0108] In some example embodiments, the step 504 of identifying at least one element of the user interface that is associated with each action may include recording a set of details for each action of the set of actions, wherein the set of details for each action identifies at least one content item within the user interface that is associated with the action. In some embodiments, the identifying may include identifying at least one user control within the user interface that is associated with the action, and recording at least one attribute of the at least one user control to identify the at least one user control within the user interface that is associated with the action. The attributes may include visual attributes (e.g., a size, shape, color, and / or label of a visual element) and / or metadata attributes (e.g., tag attributes of tags within a hypertext markup language (HTML) and / or extensible markup language (XML) encoding of the document object model). The identifying of the elements of the user interface may include an analysis of the document object model (DOM) of a web page in which the elements are defined and / or described, such as the locations and / or relationships of the elements in the DOM.
[0109] In some example embodiments, the step 506 of determining a workflow 126 of the task 116 may include determining a set of steps 128 of the workflow 126, wherein each step 128 includes one or more actions 124, and each action 124 corresponds to one or more interactions 144 performed by the user 104 with the one or more determined elements of the user interface while performing the task 116. In some example embodiments, the determining of the workflow 126 includes reviewing a set of detected events generated by the interactions 144 of the user 104 with the user interface and generating a set of actions 124, grouped into steps 128, that together correspond to the set of detected events.
[0110] In some example embodiments, the step 506 of determining a workflow of the task may include modifying the determined steps 128 and / or actions 124. For example, the device 106 may identify a first content item within the user interface that is associated with an action, wherein the first content item is not responsive to the action (e.g., a click event located on an image that is unresponsive to click events). The device 106 may determine a second content item within the user interface that is associated with the first content item, wherein the second content item is responsive to the action (e.g., a button underlaying and / or adjacent to the image). Instead of recording attributes of the unresponsive first content item, the device 106 may record at least one attribute of the second content item to identify at least one content item within the user interface that is associated with the action.[OHl] In some example embodiment, when the device 106 detects an action by the user while performing the task, the device 106 may identify at least one user control within the user interfacethat is associated with the action, capture a screenshot of the at least one user control, and record the screenshot of the at least one user control in the set of details for the action. The screenshot may serve to document the appearance and / or location of the at least one user control within the web page that is to be activated to perform an action of the workflow. In some embodiments, the device 106 captures the screenshot of the user control by intercepting at least one event arising within the user interface with at least one event handler, wherein the at least one event is associated with the action by the user. The device 106 may capture the screenshot of the at least one user control while intercepting the at least one event associated with the action. After capturing the screenshot, the device 106 may re-raise the at least one event arising within the user interface. By capturing the screenshot while the event is intercepted, the device 106 may capture a visual representation of the user interface immediately prior to invoking the event. In some example embodiments, the device 106 may overlay a visual highlight with respect to the at least one user control. In this case, the device 106 may capture the screenshot of the at least one user control by interrupting presentation of the visual highlight associated with the at least one user control while capturing the screenshot of the at least one user control. The device 106 may redisplay the visual highlight associated with the at least one user control after capturing the screenshot. Additionally, in some example embodiments, the device 106 may overlay a depiction of the visual highlight over a depiction of the at least one user control in the screenshot. For example, if the action was a user entering text into a text box, the device 106 may initially overlay a bounding box over the text box while the user is entering the text. Once the user has completed the action, the device 106 may temporarily interrupt presentation of the bounding box while it captures a screenshot of the user interface. After the screenshot is taken, the device 106 (or the backend) may then modify the screenshot by overlaying the bounding box over the text box in the screenshot. In some embodiments, the bounding box in the screenshot is editable so the user can adjust the dimensions of the bounding box to ensure that it properly depicts the text box being used by the user to complete the particular action.
[0112] In some embodiments, the step 508 of storing a record of the workflow of the task may include presenting a side panel adjacent to the user interface that enables editing of the workflow. Responsive to detecting an action by the user while performing the task, the device 106 may present a description of the action in the side panel. The device 106 may receive at least one information item provided by the user as input to the description of the action in the side panel, wherein the at least one information item further describes the action. In some example embodiments, the device 106 may record the at least one information item provided by the user in the set of details for the action. The device 106 may identify a portion of the document object model that describes at least one user control within the user interface that is associated with theaction. The device 106 may record at least one attribute of the portion of the document object model, wherein the at least one attribute identifies at least one element within the document object model that defines the at least one user control. For example, the at least one attribute identifying at least one element of the document object model includes at least one of a location of the at least one element within a schema of the document object model, at least one attribute of the at least one element defined by the document object model, or at least one attribute of at least one user control that is associated with the at least one element of the document object model. In some example embodiments, the device 106 may modify the detected events while recording as the workflow. For example, the device 106 may determine at least two consecutive actions of the set of actions by the user while performing the task within the user interface, wherein at least one action of the at least two consecutive actions is redundant with at least one other action of the at least two consecutive actions. The device 106 may remove the at least one action from the set of actions to reduce the redundancy. As another example, the device 106 may determine at least two consecutive actions of the set of actions by the user while performing the task within the user interface, wherein at least one action of the at least two consecutive actions renders moot at least one other action of the at least two consecutive actions (e.g., a first click that causes a checkbox to transition from an unchecked state to a checked state, and a second click that causes the checkbox to transition from the checked state back to the unchecked state). The device 106 may remove the at least two consecutive actions from the set of actions. As yet another example, the device 106 may determine at least two consecutive actions of the set of actions by the user while performing the task within the user interface, wherein the at least two consecutive actions are associated with a user control within the user interface (e.g., a first text event that causes an entry of the letter “N” in a textbox, and a second text event that causes an entry of the additional letter “O” in the textbox). The device 106 may substitute an aggregated action in the set of actions for the at least two consecutive actions associated with the user control within the user control (e.g., an action of entering the word “NO” into the textbox). As yet another example, the device 106 may determine at least one first action of the set of actions by the user while performing the task within the user interface, wherein the at least one first action is associated with a first user control within the user interface (e.g., an entry of a name of “Joe” in a “First Name” textbox), and at least one second action of the set of actions by the user while performing the task within the user interface, wherein the at least one second action is associated with a second user control within the user interface (e.g., an entry of a name of “Smith” in a “Last Name” textbox). The device 106 may substitute an aggregated action in the set of actions for the at least one first action and the at least one second action, wherein the aggregated action is associated with both the first user control and the second user control (e.g., a multi-step action of “Enter the name into the First Name andLast Name textboxes”). The device 106 may generate a workflow for the task, wherein the workflow includes a sequence of descriptions generated by the processor for each action of the workflow, and the description for each action is based on the set of details associated with the action. The workflow may include a name of the task, an order of the actions, a name and / or screenshot for each action, and the like. The device 106 may present the sequence of descriptions responsive to a request to describe the workflow for the task. For example, the device 106 may display the actions of the workflow to teach a user how to perform the task associated with the workflow.
[0113] In some example embodiments, the example method 500 of FIG. 5 may begin with receiving an instruction to capture the interactions 144 of a user 104 while performing a task 116 within a user interface 114. As a first example, the user 104 may click a “Record” button in a side panel of a web browser. The user 104 may issue a verbal instruction to the device 106 to record the interactions 144 of the user 104 with the web page. As a second example, the device 106 may prospectively capture the interactions 144 of the user 104 with the web page prior to receiving a request or instruction to do so, in case the user 104 initiates such a request or instruction after performing at least a portion of the interactions 114 associated with the task 116. As a third example, the device 106 may retrospectively determine the interactions 114 associated with the task 116, e.g., by reviewing an event log and identifying logged events that correspond to the interactions 144 performed by the user 104 while performing the task 116. As a fourth example, the device 106 may use machine learning techniques and inference to determine which interactions 144 are associated with a task 116, wherein the input to the inference may be prospective (e.g., prior to receiving a request or instruction to record the interactions 144 associated with a task 116), responsive (e.g., in response to a request or instruction to record the interactions 144 associated with a task 116), and / or retrospective (e.g., in response to a request or instruction to review interactions 144 occurring before the request or instruction to identify the interactions 144 associated with a task 116).
[0114] The example method 500 of FIG. 5 may provide additional detail as to the step 402 of FIG. 4 involving recording a workflow 126 of a task 116 within a user interface 114. At least a portion of the example method 500 may be performed, for example, by the example device 202 of FIG. 2. While the example method 500 of FIG. 5 illustrates one way of generating a record of a workflow 126 of a task 116 performed by a user 104 within a user interface 114, it is to be appreciated that other techniques may involve different ways of generating records of workflows 126 of tasks 116. For example, other techniques may perform steps similar to those in FIG. 5 in a different order (e.g., determining the task 116 after detecting the sequence 122 of interactions 144), may perform steps concurrently (e.g., determining the task 116 performed by the user 104while generating the record of the workflow 126), may perform steps with incidental technical variations (e.g., detecting the actions 124 of the task 116 without regard to a sequential order and / or determining the workflow 126 as an unordered set of steps 128). Such variations may include and / or may be in accordance with the techniques presented herein.
[0115] The example method 500 may be performed in a variety of contexts. As a first such example, the example method 500 may be performed for a task 116 being performed through a user interface 114 of a web page 112 presented by a web browser 110. The web page 112 may be rendered by the web browser 110 based on a document object model 120 transmitted by a webserver 118, wherein the document object model 120 indicates a hierarchical arrangement of elements 134. The webserver 118 may be local to the device 106, may be remote from the device 106 but within a group that includes and / or is associated with the device 106 (e.g., an organization), or may be provided by a third party to serve third-party content to devices and users. The hierarchical arrangement of elements 134 may indicate relationships between the elements 134, e.g., a parent element 134 (such as an HTML DIV element) that includes a number of child elements 134 (e.g., HTML user controls such as labels, buttons, textboxes, images, or the like). The web browser 110 may render the web page 112 based on the document object model 120 and various properties of the web browser 110 and / or device 106, such as the dimensions of the display 108 of the device 106 or a window in which the web browser 110 presents web pages 112. For example, the document object model 120 may indicate a logical layout of the elements (e.g., a block layout or a flow layout) and relative sizes and / or locations of the elements 134 (e.g., a first DIV element that occupies a specified portion of the display 108, such as the top 50% of the window, and additional DIV elements that fill the remaining space of the display 108). The device 106 may also alter and / or reformat the user interface 114 based on changes to the device 106 (e.g., a relocation of the web browser 110 from a small display 108 to a large display 108, or a rotation of a mobile device 106 from a portrait or vertical orientation to a landscape or horizontal orientation) and / or changes to the web page 112 (e.g., an insertion of additional elements 134 in response to additional content sent by the webserver 118). Based on the document object model 120 and the properties of the display 108, the device 106 may render the web page 112 by determining specific sizes and locations of the elements 134 that fulfill the specifications of the document object model 120. The techniques presented herein may be applied to a user interface 114 included in the web page 112 that is the result of such a rendering process, and with which a user 104 may perform interactions 144 to perform a task 116. For example, the capture 102 may be applied to capture interactions 144 of the user 104 with elements according to their representations in the document object model 120 and / or based on their rendered counterparts within a resulting rendered web page 112.
[0116] As a second such example, the user interface 114 may be included in a native application executing on a device 106, and the device 106 may determine and / or control various features of the user interface 114. For example, the user interface 114 may be specified in a declarative document such as an Extensible Markup Language (XAML) document, and an operating system of the device 106 may determine the manner of presenting the user interface 114 based on various features of the device 106. The techniques presented herein may be applied to a user interface 114 of a native application that is the result of such as rendering process by an operating system of a device 106, and with which a user 104 may perform interactions 144 to perform a task 116. For example, the capture 102 may be applied to capture interactions 144 of the user 104 with elements according to their representations in the declarative document and / or based on their rendered counterparts within a resulting native application.
[0117] As a third such example, the user interface 114 may receive input and / or interactions 144 via a variety of mechanisms. For instance, the device 106 may detect gestures that a user 104 performs with a body part (e.g., a finger, hand, foot, or head) and / or with an instrument (e.g., a handheld stylus). Such gestures may be performed in contact with a surface of an input component, a non-detecting physical surface, or in the air. The device 106 may receive input via various physical features of the device 106, such as orientation, acceleration, compass heading, geolocation, temperature, or the like. The device 106 may receive input via various physiological features, such as facial expressions, posture, and / or gait. The device 106 may receive input via other means of expression, such as voice input. The device 106 may infer input from other signals, such as an inertial measurement The device 106 may detect such gestures using physical input devices such as pressable buttons or switches, cameras, range sensors (e.g., infrared sensors and / or LIDAR sensors), microphones, inertial measurement units (IMUs) that measure orientation and / or acceleration, or the like. The user interface 114 may receive such input in a variety of computing contexts (e.g., a workspace featuring input devices connected to a workstation; a mobile context via a mobile device, such as a mobile phone, tablet, palmtop computer, or a wearable device such as an earpiece, eyeglasses, or a ring); a vehicle context in which a vehicle heads-up display provides a user interface 114; a virtual and / or augmented reality context in which a user wears a full or partial headset and / or provides user input via handheld controllers or the like.KNOWLEDGE CAPTURE - DETERMINING TASK
[0118] In some example embodiments, the step 502 of determining a task 116 to be performed by a user 104 within a user interface 114 may be performed in various ways. As a first example, the user 104 may indicate the start of performing the task 116 within the user interface 114. For instance, a device 106 may present, within a web browser 110, a side panel 224 that includes controls for capturing workflows 126 of tasks 116, including a “Capture” or “Record” button thatthe user 104 may click or touch at the start of the task 116 to initiate the recording of the workflow 126. As a second example, the device 106 may store a record of interactions 144 performed by the user 104, and may receive from the user 104, retroactively, an indication that a certain set of interactions 144 were performed to complete a task 116 within the user interface 114. As a third example, the device 106 may include an artificial intelligence model that generally observes interactions 144 of a user 104 and determines that a certain sequence 122 of interactions 144 performed by the user 104 are associated with a task 116 performed within the user interface 114. The determining by the artificial intelligence model may be based, for example, on a determination that the interactions 144 are frequently, consistently, and / or routinely performed by a user 104 or a set of users 104 in a particular circumstance (e.g., when the need to perform a particular task 116 arises) and / or that performing the interactions 144 consistently generates a particular result as one or more objectives of the task 116 (e.g., that the user 104 or a group of users 104 appears to perform the task 116 in order to achieve a particular result). The determining of the task 116 may include a determining of information about the task 116, such as a title of the task 116, one or more objectives of the task 116, one or more circumstances in which the task 116 can be performed, or the like. Such information may be provided by the user 104 performing the task 116 (e.g., the user 104 may enter a title of the task 116 in a textbox before clicking a “Capture” or “Record” button) and / or may be inferred by a device 106 based on information about the task (e.g. , based on the type of web page 112, the type of user interface 114, and / or types of interactions 144 performed by the user 104 while performing the task 116).KNOWLEDGE CAPTURE - DETECTING ACTIONS
[0119] In some example embodiments, the step 502 of detecting a sequence 122 of interactions 144 performed by the user 104 while performing the task 116 within the user interface 114 may be performed in a variety of ways. As a first example, a web browser 110 of the device 106 may generate notifications of events associated with the user interface 114 (e.g., an event-based programming model in which events are presented to and may be handled by applications or scripts executing within the web browser 110). As one such example, a JavaScript engine that executes JavaScript embedded in a web page 112 may register various event-handler functions with the web browser 110 to handle events of one or more indicated event types, wherein the event-handler functions are to be invoked upon the occurrence of events of the indicated one or more event types. Code executing within the web browser 110 (e.g., as part of a web extension) may detect the sequence 122 of interactions 144 performed by the user 104 during the performance of a task 116 by registering event-handler functions for relevant event types (e.g., button-click event handlers, keyboard input event handlers, and drag event handlers), optionally with regard to particular elements 134 of the user interface 114 (e.g., click events for button-style user controls).When such an event occurs, the event handler may receive and analyze the raised event to determine one or more associated elements 134 of the user interface 114 and may determine that such events indicate interactions 144 that correspond to the actions 124 associated with the performance of a task 116 (e.g., during a period in which the user 104 has selected or activated a “Capture” or “Record” function of the web browser 110 in order to capture 102 a record of the steps 128 comprising a workflow 126). As a second such example, the device 106 may observe and evaluate user input received from the user 104 and map portions of the user input to interactions 144 performed by the user 104 to complete the task 116. For example, the device 106 may receive user input through one or more human input devices, such as a keyboard, mouse, touch-sensitive display, or the like, and may analyze the user input to detect the sequence 122 of interactions 144 performed by the user 104. The device 106 may also correlate each determined interaction 144 with at least one element 134 of the user interface 114, such as mouse-click and / or touch interactions 144 that occur at a location where an element 134 of the user interface 114 is shown and / or receives user input. Each interaction 144 may generate one or more events within the user interface 114 (e.g., an interaction 144 with a textbox may generate a set of events respectively corresponding to individual keystrokes). When the user completes an interaction 144 within the user interface 114, the device 106 may evaluate the events arising from the interaction 144 to determine the actions 124 of the steps 128 of the workflow 126. For example, the device 106 may consider the interactions 144 as candidates to create corresponding actions 124 in the steps 128 of the workflow 126.KNOWLEDGE CAPTURE - IDENTIFYING ELEMENTS
[0120] In some example embodiments, the step 504 of identifying, for each interaction 144 of the sequence 122, at least one element 134 of the user interface 114 that is associated with the interaction 144, may be performed in a variety of ways. In example embodiments, this identification involves three sub-steps: detecting an interaction 144 of the user 104 with the user interface 114; determining one or more elements 134 of the user interface 114 that are associated with the interaction 144; and recording information about the interaction 144 and associated elements 134.
[0121] The sub-step of determining an interaction 144 of the user 104 with the user interface may be performed in a variety of ways. As a first example, the device 106 may be configured to monitor user input of the user 104 provided by one or more human input devices (HIDs), such as a keyboard, mouse, and / or touch-sensitive display. The device 106 may determine that certain interactions 144 of the user 104 indicate and / or cause an interaction 144 with one or more elements of the user interface 114. For example, based on a detection of a mouse-button click of a mouse, the device 106 may determine that a pointer is currently located within the boundaries of a userinterface 114, and that the user 104 is therefore interacting with the user interface 114. Based on the detection of a touch (e.g., with a finger, stylus, or other pointing device) on a touch-sensitive display 108, the device 106 may determine that a location of the touch is within the boundaries of a user interface 114, and that the user 104 is therefore interacting with the user interface 114. As a second example, the device 106 may implement an event-based publication / subscription model, in which applications may subscribe to receive notifications that are of interest to the application. Applications (including web extensions of web browsers 110) may subscribe to various user interaction events that indicate interaction 144 by the user 104 with various components of the device 106. Upon the occurrence of an event of a particular event type, the device 106 may publish a notification to all applications that have subscribed to events of the event type. In such contexts, an application (including a web extension of a web browser 110) may subscribe to events that indicate interactions 144 by the user 104 with the device 106, and may evaluate any received notifications from such event types to determine whether the events indicate an interaction 144 with a user interface 114. As a third example, the device 106 may implement an event handler model, in which applications may provide event handler functions that are to be executed in response to particular events. Upon the occurrence of an event of a particular event type, the device 106 may determine one or more event handler functions that are registered with events of the particular event type, and may execute each such event handler function with information about the event that prompted such execution, publish a notification to all applications that have subscribed to events of the event type. The events may be generally associated with user input (e.g., keyboard input events, touch events, and mouse-click events) and / or may be associated with the occurrence of user input for certain elements 134 of a user interface 114 (e.g., HTML user controls, such as buttons and textboxes, may raise events for particular kinds of user input performed in relation to such user controls, such as mouse-click events when a pointer is over a button or keyboard input events when a textbox has an input focus). In such contexts, an application (including a web extension of a web browser 110) may provide an event handler function for any events that indicate interactions 144 by the user 104 with the device 106, wherein the event handler function determines that a interaction 144 by the user 104 with a user interface 114 is occurring and / or has occurred based on the details of the raised event. For instance, a web browser 110 may permit a web extension (e.g., a content process 210) to provide a JavaScript script that registers event handlers for various user input events, such as button-click events. When the web browser 110 receives user input that raises such a user input event, the JavaScript script may execute the event handler function provided by the web extension to notify the web extension of the occurrence of an interaction 144 by the user 104 with the user interface 114.
[0122] The sub-step of determining one or more elements 134 of the user interface 114 that are associated with the interaction 144 may be performed in a variety of ways. As a first example, an application may compare a location of the user input (e.g., a location of a pointer during a mouseclick event) with the locations of various elements 134 of the user interface 114 (e.g., locations of HTML user controls within a web page) to determine one or more elements 134 that are associated with the user input. When user input is received in relation to a user interface 114 (e.g., when a click event occurs when a pointer is within the boundaries of a user interface 114, or when a keyboard input event occurs when a textbox of a user interface 114 has input focus), the device 106 may identify one or more elements 134 of the user interface 114 that are associated with the user input. When an element 134 of the user interface 114 raises an event due to the receipt of user input, an event provided to a subscribing script or application may identify the element 134 that raised the event (e.g., by name, an identifier assigned to the element 134 by the device 106, an identifier assigned to the element 134 by an application or script, an identifier of the element 134 indicated in a declarative document such as an HTML document, a path of and / or to the element in a hierarchical document (e.g., an HTML or XML document), or the like.
[0123] As a second example, the device 106 may determine that a first element 134 associated with an input event is not responsive to the provided input. For example, a user 104 might click on a text label or image that is not responsive to click events, or may provide text input (e.g., via a keyboard) to a button that is not configured to receive text input. In such cases, the device 106 may endeavor to determine a second element 134 that is associated with the first element and that is responsive to the provided input. As a first example, the device 106 may examine a graphical layout of the user interface 114 to determine one or more other elements 134 that are near the first element 134 and that are responsive to the provided input. For example, a web page 112 may include an image that is not responsive to click user input, but that is overlapping, adjacent to, and / or located near a button that is responsive to click user input. In such cases, the device 106 may infer that the user 104 intended to click on the button but accidentally clicked on the image, and / or that the encoding of the web page 112 causes the button to be overlapped by an image that incorrectly intercepts the click user input. The examining of the graphical layout may be based on thresholds (e.g., considering other elements of the user interface 114 that are within 100 pixels of the first element 134). As a second example, the device 106 may examine a logical arrangement of the elements 134 to determine a second element that is logically related to the first element 134 and that is responsive to the provided input. For example, the device 106 may determine that the document object model 120 of a web page 112 includes an HTML DIV element 134 that is not responsive to click user input, but that includes a button that is responsive to click user input. The device 106 may determine that the click user input is intended for the button rather than the HTMLDIV element 134, and may select the button rather than the HTML DIV element 134 for the capture 102 of user input for the task 116. As another example, the device 106 may determine that the document object model 120 of a web page 112 associates a text label, which is not responsive to click user input, with a button that is responsive to click user input. The device 106 may determine that the click user input is intended for the button rather than the text label element 134, and may select the button rather than the text label element 134 for the capture 102 of user input for the task 116.
[0124] As a third example, the device 106 may translate user input into other user input that is to be processed as part of the task 116. As a first such example, a user 104 might click on an empty space of a user interface 114 (e.g., a location where no user controls or content exists). The device 106 may infer that the user 104 intended to click or activate another user control (e.g., a user control that is located near the empty space that the user 104 clicked). Accordingly, the device 106 may translate the user input into a click event at the location of the other user control. As a second such example, a device 106 may provide a set of accessibility features, such as a voice input device that receives voice commands from a user 104. Upon receiving a voice command (e.g., “click the ‘OK’ button”), the device 106 may translate the voice command into corresponding user input events (e.g., a mouse-click event on a button identified as showing an ‘OK’ caption). As a third such example, the device 106 may translate a user input event that an associated element 134 is not configured to handle into a different user input event that an associated element 134 is configured to handle. For instance, a button may be configured to receive click events, but not drag or touch user input events. Upon detecting a drag and / or touch user input event, the device 106 may determine that the button indicated by the event does not respond to drag or touch events, and may translate the user input event into a click event that the button is configured to handle. As a fourth such example, the device 106 may determine that a particular element 134 exhibits an appearance that resembles a user control that is responsive to the provided user input, but the particular element 134 may not be responsive to the provided user input (e.g., a button that is visually styled to resemble a checkbox, but that does not expressly include checkbox- style functionality). The device 106 may translate the user input into user input that corresponds to the element (e.g., translating user input to the user control into events that correspond to other common commands, such as “select” and “unselect” or “check” and “uncheck” events). As a fifth example, the device 106 may translate an event received from a user 104 into another event that is more likely to have been intended by the user 104. For instance, the user 104 may have intended to invoke a button by clicking on it, but may have accidentally doubleclicked on the button to generate a double-click event that is associated with a different type of interaction 144 and / or that the button is not configured to handle. The device 106 may determinethat a double-click event is likely to have been intended as a click event and may therefore substitute a click event for the double-click event. As a sixth example, the device 106 may consolidate two or more events into a particular event that an element 134 is configured to handle. For instance, the device 106 may receive a series of drag events, such as a first drag event indicating a drag motion from a first location to a second location, and a second drag event indicating a drag motion from the second location to a third location. The device 106 may consolidate the series of drag events into a single drag event from a start location to an end location.
[0125] The sub-step of recording, for each interaction 144, information about the interaction 144 and associated one or more elements 134 of the user interface 114 may be performed in a variety of ways. As a first example, the device 106 may record information about the user input and / or interaction 144, including information about one or more events associated with the interaction 144 (e.g., an event identifier of the event, a timestamp of the event, and / or a relative amount of time between the event and a preceding event), an event type (e.g., a click event), and / or user input that resulted in the event (e.g., a left-click of a mouse button). The device 106 may record information about how an event was handled (e.g., a list of event handler functions that handled the event). The device 106 may record, for each interaction 144, a response of the user interface 114 to the action of the user 104 (e.g., a result of a button-click operation, such as execution of a particular JavaScript function, or an invocation of a device behavior such as an alert or a presentation of a File-Open dialog).
[0126] As a second example, the device 106 may record, for the interaction 144, information about one or more elements 134 associated with the event. For example, the device 106 may record, for the interaction 144, an identification of each element 134 associated with the event, such as an identifier of the element (e.g., an ID or name of the element 134), a type of element 134 (e.g., a CSS class or HTML class of the element 134), a location of the element 134 in a visual layout of the user interface 114 (e.g., coordinates of the element 134 within a web page 112 and / or on a display 108), and / or a location of the element 134 within a declarative document (e.g., a path to the element within an HTML or XML document). The device 106 may record, for the interaction 144, metadata about each element 134 associated with the action (e.g., a list of HTML or XML attributes or enclosed content such as inner text of an HTML element).
[0127] As a third example, the device 106 may record, for each interaction 144, information about the user interface 114 at the time of the interaction 144. The device 106 may record, for the interaction 144, information about one or more events associated with the interaction 144 (e.g., an event name, event type, event timestamp, and / or event properties). The device 106 may record, for the interaction 144, a URL or URI of a source document in which the element 134 is specified,such as an HTML document, XML document, a JavaScript fragment injected into a web page 112 via an AJAX technique, and / or a script or web extension that programmatically created the element 134. The device 106 may record URLs and / or URIs of source documents relative to a base URL or base URI (e.g., a common or shared prefix of the URLs or URIs of all source documents). The device 106 may record, for the interaction 144, information about a source document in which the element 134 is specified (e.g., a snapshot of the document object model 120 that caused the element 134 to be created and / or to be associated with the event, wherein the snapshot may include URLs of each source document included in the document object model 120, optionally including a name, filename, timestamp, signature, and / or version indicator of each source document). The device 106 may record, for the workflow 126, information about the device 106 and the context in which the workflow 126 was performed, such as properties of the device 106 (e.g., a make, model, and / or computational resources of the device 106); a type and / or version of a web browser 110; properties of one or more input devices connected to the device 106 that were used during the capture 102 of the workflow 126; properties of one or more displays 108 on which the user interface 114 was presented to the user 104 during the capture 102 of the workflow 126 (e.g., a resolution of the display 108); and / or properties of one or more windows within which the user interface 114 was presented to the user 104 during the capture 102 of the workflow 126 (e.g., a location, width, height, and / or aspect ratio of the one or more windows).
[0128] In some example embodiments, a device 106 may include a UX capture system, wherein the capture system outputs visual documentation of each action that indicates the respective interactive element with which the user interacted with when the action was detected. For example, the capture system may be configured to output visual indicia (e.g., bounding box) over the interactive element that a user is potentially interacting with currently.
[0129] As a fourth example, in some example embodiments, a device 106 may record, for each interaction 144, one or more screenshots associated with the interaction 144. For example, the device 106 may capture 102 a screenshot of the web page 112 at the time of the event. The device may indicate, within a screenshot of the web page 112, an area that is associated with the event (e.g, a screenshot of the entire web page with a rectangle or circle drawn on the image to show the location of the click event). The device 106 may crop the screenshot to a representative portion of the web page 112 (e.g, a rectangle of + / - 100 pixels height and + / - 200 pixels width around a location of a click event, or around a region of a button associated with the click event). The device 106 may capture 102 a set of screenshots that show a video or animating image that depicts and / or explains the event (e.g., a cursor approaching a button and clicking on the button). The device 106 may record the one or more screenshots in a variety of formats, such as bitmaps, vector images, animating images such as GIFs, compressed images such as JPEGs, compressed or uncompressedvideos such as AVIs, or the like. The device 106 may store the screenshots in a record of the event along with the other information about the event.
[0130] In some example embodiments, a UX / UI of a device overlays visual indicia in relation to the interactive elements of the GUI in response to the interaction monitoring system identifying interactive elements with which the user is engaging and briefly interrupts display of the visual indicia in response to the action capture system executing a screenshot operation to capture an identified action performed using an interactive element of the GUI, wherein the action capture system overlays corresponding visual indicia over a screenshot of the GUI at a location in the screenshot depicting the interactive element used to perform the captured action. This may allow the visual indicia (e.g., "orange box") to be editable by the user in the screenshot. Otherwise, the visual indicia may be "burnt" into the screenshot, thereby making editing the visual indicia (e.g., moving or resizing the orange box) in relation to the rest of the screenshot much more challenging. The UX / UI may overlay visual indicia in relation to the interactive elements of the GUI in response to the interaction monitoring system identifying interactive elements with which the user is engaging and briefly interrupts display of the visual indicia in response to the action capture system executing a screenshot operation to capture an identified action performed using an interactive element of the GUI, wherein the action capture system overlays corresponding visual indicia over a screenshot of the GUI at a location in the screenshot depicting the interactive element used to perform the captured action. The "interrupt" portion may be short enough in duration that the user cannot see it but it is long enough to capture the screenshot without the indicia being depicted on the GUI.
[0131] In some example embodiments, the device 106 may be configured to capture 102 a screenshot of a user interface 114 during and / or after the occurrence of the event. For example, the web browser 110 may receive user input (e.g., a mouse click in a particular location), may initially process the user input (e.g., by determining that the user input occurs in the location of a button, and therefore showing a click animation of the button being pressed), and may raise an event for further processing by various applications (e.g., a button click event to be handled by one or more event handlers). In such cases, a screenshot of the user interface 114 may visually include a result of the event (e.g., the button-click animation). However, it may be desirable to capture 102 a screenshot of the user interface 114 just prior to the initiation of the event. For example, if the event causes visual changes to the user interface 114, such as a checkbox being checked, a button being hidden, or a textbox being cleared, the device 106 may instead be configured to capture 102 a screenshot of the user interface 114 just prior to the event and the visual changes to the user interface 114, such that the screenshot may show a user 104 a state of the user interface 114 just prior to the event being performed in order to guide the user to performthe action. In order to do so, a device 106 may be configured to perform an overlay interrupt that intercepts an event and that prevents processing of the event until a screenshot has been captured. For example, a content process 210 may register an event handler function with a click event, wherein the event handler function is designed to intercept and dispose of the event before any further processing of the event. The content process 210 may also be configured to initiate a screenshot of at least a portion of the user interface 114 before the event has completed processing. Once the screenshot has been captured, or after a period of time in which the screenshot is at least likely to have been captured (e.g., a 50-millisecond delay), the content process 210 may be configured to re-raise the event within the user interface 114 (e.g., raising a second button-click event with the same button). The event handler of the content process 210 may forgo processing of the re-raised event in order to permit the other event handlers to process the event.
[0132] In some example embodiments, while capturing a task 116 being performed by a user 104, a device 106 may insert visual indicators into the user interface 114 to indicate the elements 134 for which interactions 144 of the user interface 114 may be recorded in the workflow 126. For example, as the user 104 moves a cursor within a web page 112, the cursor may encounter various elements 134, such as HTML user controls (e.g., buttons, textboxes, and the like) that the user 104 may perform interactions 144 with during the task 116. At such times, the device 106 may highlight, color, or otherwise visually indicate the element 134 with which the user 104 may perform interactions 144 as part of the task 116. Such visual indicators may be helpful in informing the user 104 about the recording process (e.g., which element is currently being interacted with by the user). Further, the device 106 may be configured to capture 102 screenshots of the user interface 114 for respective actions 124 of the workflow 126. In such cases, the device 106 may be configured to show the visual indicator during the capture 102 of the screenshot (e.g, to indicate one or more elements 134 that are associated with the interaction 144). Alternatively, the device 106 may be configured to hide the visual indicator (e.g, to avoid obscuring the element 134 and / or other parts of the user interface 114). For example, if the device 106 is configured to register event handler functions with a web browser 110 for events that are associated with actions 124, the event handler function may both intercept the event for further processing and may briefly hide the visual indicator prior to capturing a screenshot of the user interface 114. Before or during re-raising the event, the event handler function may cause the visual indicator to be shown again. The event handler function may store one or more screenshots of the user interface 114 with and / or without the visual indicator of the one or more elements 134 associated with the interaction 144. For example, the event handler function may capture 102 a screenshot of the user interface 114 while the visual indicator is hidden, and may generate a second screenshot that adds the visual indicator to the screenshot. In this way, the device 106 may capture 102 a screenshot of the userinterface 114 that is not obscured by the visual indicator, and may also reserve the option of adding / overlaying the visual indicator to the screenshot to highlight the one or more elements 134 associated with the interaction 144.
[0133] FIG. 6 illustrates an example scenario featuring a capture 102 of a screenshot of an event arising within a user interface of a web browser of a device, according to aspects of the present disclosure. The example scenario of FIG. 6 may be performed by the example device 202 of FIG. 2.
[0134] In the example scenario of FIG. 6, at a first time 602, the web browser 110 of the device 106 executes a content process 210 (e.g., as a portion of a browser extension provided by a workflow server 220). During a capture 102 involving a web page 112, the content process 210 registers 604 to receive particular events arising within the web browser 110, such as click events arising within the user interface 114 of the web page 112. The content process 210 may provide and / or identify an event handler function 606 to be invoked for click events arising within the user interface 114 of the web page 112. The web browser 110 may store a record of the request to register 604 for click events (e.g., adding the event handler function 606 provided and / or indicated by the content process 210 to a list or stack of event handler functions for click events arising within the user interface 114). As the user explores the web page 112 and performs interactions 144 with the user interface 114, the content process 210 may perform various actions as part of the capture 102 for the task 116, such as showing a highlight around elements 134 with which the user 104 may perform interactions 144 to perform the task 116.
[0135] At a second time 608, the web browser 110 may receive user input that causes an event 612 to occur. The event 612 may be associated with a particular element 134 of the user interface 114, such as a button. Further, the content process 210 may have added a highlight 610 of the element 134 as a visual indicator of the elements 134 to be recorded in the actions 124 of the workflow 126 of the task 116. The web browser 110 may invoke the event handler function 606 of the content process 210 to handle the event 612. The event handler function 606 may receive the event 612 and may perform an overlay interrupt by disposing of the event 612, thereby postponing further processing of the event 612. The content process 210 may also hide the highlight 610 of the element 134. The content process 210 may then initiate a screenshot 614 of at least a portion of the user interface 114 that includes the element 134 associated with the event 612.
[0136] At a third time 616 (e.g., 50-100 milliseconds after the second time 608), the content process 210 may reinitiate the event 612 to enable further processing of the event 612. For example, the click event handler function 606 may re-raise the event as a replayed click event 618 and may forgo processing of the replayed click event 618 (e.g, refraining from performing anoverlay interrupt of the replayed click event 618). As a result, the web browser 110 may fulfill the processing of the event 612, optionally with a delay that is long enough to enable the capture 102 of the screenshot 614 of the user interface 114 prior to the complete processing of the event 612 while also short enough to be unnoticeable and / or inconsequential to the user 104. For example, the delay may be 50-100 milliseconds long. The content process 210 may also restore the highlight 610 of the element 134 to continue to visually assist the user 104 through the capture 102. In this manner, the content process 210 may capture 102 a screenshot 614 of the user interface 114 just prior to the processing of the click event 612 and also free of highlights 610 that may obscure the screenshot, in accordance with some aspects of the techniques presented herein.
[0137] In some example embodiments, as part of the information recorded about each action 124, the device 106 may generate one or more annotations that are associated with the action 124. For example, each action 124 may include an associated human-readable description of the action 124, wherein the description informs a user 104 about the manner of completing the action 124, a context of the action 124 within the workflow 126 of the task 116, one or more interactions 144 of the user 104 that are involved in performing the action 124, one or more forms of data that are involved in the action 124 (e.g., a type of information to enter into a textbox), and / or a response of the user interface 114 to the performance of the action 124 (e.g., how the user interface 114 responds and / or changes when the action 124 is performed). The annotation may describe, depict (e.g., via a screenshot), and / or indicate one or more elements 134 of the user interface 114 that are associated with the action 124, and / or one or more ways to perform interactions 144 with the one or more elements 134. The annotation may indicate different ways of performing an action 124 (e.g., clicking on a button with a cursor controlled by a mouse, touching the button on a touch- sensitive display, or pressing a key combination that activates the button). The annotation may include preconditions of performing the action 124 (e.g., other steps and / or actions to have been completed prior to the action 124). The annotation may include one or more alternatives to performing the action (e.g., instead of entering address information into a textbox, selecting a saved address from a drop-down list). The annotation may include one or more conditions under which the action 124 is to be performed (e.g., a reason for performing the action 124, a reason for skipping the action 124, and / or a reason for performing an alternative to the action 124). The annotation may include a description of an error condition that may occur during the action 124 (e.g., a type of warning that the user interface 114 may present if the action 124 is performed incorrectly and / or incompletely) and, optionally, one or more remedial actions 124 to be performed in response to the error condition. The annotation may include different stages and / or levels of presenting information about the action 124 (e.g., a brief description of the action 124 that is initially shown, and a more detailed description of the action 124 if the user 104 encountersdifficulty with the action 124, performs the action 124 incorrectly and / or incompletely, fails to perform the action 124 in a particular period of time, and / or requests help with the action 124).
[0138] In some example embodiments, a device may include a label generation component that automatically generates a respective text label for each respective action documented in a user- directed UX workflow, wherein each text label is generated based on the interactive element used to perform the object and any content provided by the user as input via the interactive element.
[0139] In some example embodiments, the device 106 may automatically generate one or more annotations of each action 124. The device 106 may automatically generate an annotation based on an event type associated with the action 124 (e.g., an action 124 including a click event involving a button may be automatically labeled as “Click the Button”). The device 106 may automatically generate an annotation based on one or more elements 134 associated with the action 124 (e.g., an action 124 involving a button labeled “Submit” may be automatically labeled as “Click the Submit button”). The device 106 may automatically generate an annotation based on predefined labels for particular types of elements 134 and / or particular events arising with respect to the elements 134 (e.g., a click event involving a checkbox element 134 may be automatically labeled “check the checkbox” if the click event causes the checkbox element 134 to be checked, or “clear the checkbox” if the click event causes the checkbox element 134 to be cleared). The device 106 may automatically generate an annotation based on one or more pieces of data or metadata associated with an event and / or element 134 associated with the action 124 (e.g., a name or ID attribute of the element 134; a text caption indicated for the element 134, such as a label shown on a button; and / or a text caption that is adjacent to the element 134, such as a text label shown to the left of a button). The device 106 may automatically generate an annotation based on one or more accessibility indicators of one or more elements 134 associated with the action 124 (e.g. , an aria label of a button that indicates that clicking the button performs a “Submit” action). The device 106 may automatically generate an annotation based on one or more other elements 134 that are near and / or associated with one or more elements 134 of the action 124 (e.g., an action 124 involving a click event on a non -interactive element, such as a non-interactive text label that serves as a label for a clickable button, may be automatically labeled “click the button”). The device 106 may automatically generate an annotation for an action 124 involving a first element 134 based on one or more related elements 134 (e.g., parent, grandparent, or the like elements of the first element; sibling elements of the first element; and / or one or more child, grandchild, or the like elements, according to a hierarchical arrangement of elements 134 such as a document object model 120). The device 106 may automatically generate an annotation for an action 124 involving a first element 134 based on one or more nearby elements 134 (e.g., elements that are within a particular graphical distance of the first element 134 in a visual layout of the userinterface 114, such as within 100 pixels of the first element 134 in a layout of a web page 112 as currently shown on a device 106). The device 106 may automatically generate an annotation for an action 124 based on analysis of code associated with the action 124 (e.g., based on a name, description, comment, result, or the like of an event handler function 606 that is executed to handle the action 124).
[0140] In some example embodiments, a device 106 may automatically alter one or more screenshots 614 and / or annotations of one or more actions 124. As a first example, an element 134 may include user-specific and / or sensitive information, such as a value of a password, a social security number, and / or a photo of the user 104 or another individual. While the device 106 is capturing a performance of a task 116, a user 104 may provide such information to complete the task 116, but it may be desirable to exclude the information from the workflow 126 presented to other users 104. Instead, the device 106 may remove, redact, obscure, or otherwise alter the information captured for the action 124. For example, if a textbox is indicated as receiving a password or a social security number of a user 104, the device 106 may remove the value of the textbox from information captured about the action 124, including one or more screenshots 614 that might otherwise reveal the value of the textbox. The device 106 may alter a descriptive annotation of the textbox (e.g., “enter your password” instead of “enter ‘hunter2’ in the password textbox”). The device 106 may overwrite, delete, blur, or otherwise obscure the value of the textbox in one or more captured screenshots 614 of the textbox. The device 106 may exclude the value of the textbox from metadata captured about the action 124, such as a state of the document object model that might otherwise include the value of the textbox. As a second example, the device 106 may determine that while a user 104 performs an action 124 in a certain way during capture 102 of the workflow 126, the presentation of the workflow 126 should indicate a different way of performing the action 124 (e.g., when presented to other users 104). For example, a user interface 114 may include a drop list of countries, and a user 104 who is performing the task 116 may be located in a first country. However, users to whom the workflow 126 is to be presented may be located in different countries, or in a specific second country. Instead of annotating the action 124 as “select the first country from the drop list” and a screenshot 614 showing the name of the first country in the drop list, the device 106 may instead present a generic version of the annotation (e.g., “select your country from the drop list” and a screenshot 614 showing a portion of the drop list of available countries), and / or may present a different version of the annotation (e.g., “select the second country from the drop list” and a screenshot 614 showing the name of the second country in the drop list). As another example, if the user 104 provides an image to the user interface 114 (e.g., uploading a personal image as part of a personal profile and / or identity verification process) during a capture 102 of the workflow 126, the device 106 may show a genericplaceholder image, a blank image, and / or a blurred image in a screenshot 614 of the action 124 instead of the personal image.
[0141] In some example embodiments, a device 106 may include a screenshot generation component that generates screenshots of a website GUI, overlays visual indicia over the screenshot to highlight an interactive element in the website GUI used to perform an action, receives user commands to edit one or more features of the overlaid indicia, such as size, location, and / or appearance of the visual indicia, and outputs an updated screenshot with an updated visual indicia in response to the user commands.
[0142] In some example embodiments, a device 106 may allow a user 104 to alter the information about an action 124 captured for the workflow 126. For example, during the capture 102 of a workflow 126, a web browser 110 may show a side panel 224 in a “Capture” mode that includes various controls and options for editing the captured workflow 126. The side panel 224 may show the user 104 the list of actions 124 so far captured. The side panel 224 may show the user 104 an annotation of each action 124, such as a name, label, type, description, screenshot, thumbnail, or the like. The side panel 224 may allow the user 104 to create, modify, and / or delete a name, label, type, description, screenshot, thumbnail, annotation, or the like for each captured action 124, including values and annotations that were automatically generated for the action 124 and / or values and annotations that were previously entered by the user 104. The side panel 224 may show, to the user 104, all of the information recorded for the action 124 (e.g., an event arising during the action and / or identifiers of one or more elements 134 of the user interface 114 that are involved in the action). The side panel 224 may allow the user 104 to delete an action 124 (e.g., if the action was performed in error or otherwise should not be included in the workflow 126). The side panel 224 may accept, from the user 104, one or more additional screenshots, animations, or the like to be associated with one or more elements 134 of the action 124, and may store the one or more additional screenshots, animations, or the like as part of the capture 102 of the action 124. The side panel 224 may allow the user 104 to edit one or more screenshots 614 of the action 124 (e.g., obscuring a portion of the screenshot 614; cropping, rotating, scaling, or otherwise modifying the screenshot 614; annotating the screenshot 614 with descriptive text or icons; and / or adding, relocating, or removing highlighting 610 in relation to one or more elements 134 of the user interface 114).KNOWLEDGE CAPTURE - DETERMINING WORKFLOW
[0143] In some example embodiments, the step 506 of determining a workflow 126 of the task 116 based on the actions 124 recorded by the device and / or application while the user performs the task within the user interface 114 may be performed in a variety of ways. As a first example, a device 106 may determine the workflow 126 of the task 116 as a sequence 122 of steps 128based on the sequence 122 of interactions 144 performed by the user 104. An order of the sequence 122 of the steps 128 may correspond to an order of the sequence 122 of the actions 124 that correspond to the interactions 144 performed by the user 104. The device 106 may determine the workflow 126 according to an order of steps 128 indicated by the user 104 (e.g., the user 104 may use controls in a side panel 224 to arrange or rearrange the steps 128 of the workflow 126 during and / or after capture 102). The device 106 may determine the workflow 126 as a sequence of steps 128 in a different order than the sequence 122 of the interactions 144 performed by the user 104. For instance, the device 106 may determine that one or more of the interactions 144 were performed out of a logical order, or that the workflow 126 may be more clearly presented by ordering the steps 128 in a different order than the interactions 144 performed by the user 104. As one such example, a user interface 114 may include an arrangement of textboxes as a vertical column, but the user 104 may complete the textboxes out of order, and the device 106 may choose to arrange the steps 128 and / or actions 124 of the workflow 126 according to their vertical order in the user interface 114 rather than in the order performed by the user 104. The device 106 may determine the workflow 126 to include steps 128 that may be performed in one or more different orders, or as an unordered set of steps 128 (e.g., a task 116 that involves completing form may be performed by completing different sections of the form in any order that a user 104 may choose). For example, if the user interface 114 includes a set of three textboxes arranged in a vertical order, and the user 104 performs interactions 144 with the textboxes in a different order than the vertical order of the arrangement (e.g., from bottom to top), the device 106 may generate a workflow 126 that includes steps 128 and actions 124 that correspond to interactions 144 with the textboxes according to the vertical order of the layout rather than the order of the interactions 144.
[0144] As a second example, the device 106 may omit one or more steps 128 and / or actions 124 from a workflow 126. That is, the device 106 may detect some interactions 144 of a user 104 during capture 102 of a workflow 126 that are not included in any step 128 of the workflow 126 determined by the device 106. As a first such example, the device 106 may determine that an action 124 performed by the user 104 was not associated with an element 134 that could respond to the action 124 (e.g., a click event 612 associated with non-clickable elements 134 or an empty area of the user interface 114, where no nearby and / or related element 134 is configured to receive and respond to the click event 612) and may omit the action 124 from the workflow 126. As a second such example, the device 106 may determine that an action 124 performed by the user 104 did not produce a consequential result (e.g., a click event 612 associated with a clickable button that has no associated button-click event handler or where the event handler for the click event produced no result; a list selection event where the user 104 did not select a different option in the list; or a textbox selection event where the user 104 did not enter or change any text) and mayomit the action 124 from the workflow 126. As a third such example, the device 106 may determine that a first action 124 performed by the user 104 was superseded by a second action 124 performed by the user 104 (e.g., a first text entry event 612 in a textbox during which the user 104 inserted a first value into the textbox, followed by a second text entry event 612 in a textbox during which the user 104 inserted a second value into the textbox that overwrite the first value) and may omit the first action 124 from the workflow 126. As a fourth such example, the device 106 may determine that a first action 124 performed by the user 104 was followed by a second action 124 that was redundant and cumulative with the first action 124 (e.g., clicking a button twice, where the second click produces the same result as the first click and / or is ignored) and may omit the second action 124 from the workflow 126. As a fifth such example, the device 106 may determine that an event 612 is intercepted by an element 134 that represented an entirety or near-entirety of the user interface 114 (e.g., an element 134 of a web page 112 that consumes 98% of the web page 112, such as an overlay) and may omit the action 124 from the workflow 126.
[0145] In some example embodiments, a device may include a multi-action capture component that detects multi-action steps performed by a user during a user-initiated capture session, wherein a multi-action step is a UX workflow step that comprises two or more distinct interactions 144 by the user with two or more interactive elements of a website GUI; and in response to detecting a multi-action step generates a single screenshot that depicts the two or more interactive elements and overlays one or more visual indicia over the single screenshot that highlight the two or more interactive elements. The screenshot that is used may be the last one in the list of actions / events. In some example embodiments, each interactive element is highlighted with a respective bounding box (e.g., Box A over "Phone" input element, Box B over "Home Phone" input element; and Box Cover Mobile). In some example embodiments, a device may include a multi-action capture component that detects multi-action steps performed by a user during a user-initiated capture session, wherein a multi-action step is a UX workflow step that comprises two or more distinct interactions 144 by the user 104 with two or more interactive elements of a website GUI, and generates multi-action step guidance data that is indicative of the two or more interactive elements implicated in the multi-action steps including target data that indicates one or more DOM attributes of the two or more interactive elements being combined into the multi-action step. The device may capture target data that is used to guide a user through multi-action steps.
[0146] As a third example, a device 106 may determine a workflow 126 in which two or more interactions 144 performed by the user 104 may be merged into a single step 128. For example, a task 116 may involve completing a form with several textboxes, such as “First Name,” “Middle Name,” and “Last Name,” and a user 104 may perform three distinct interactions 144 while entering text into each of the textboxes. Rather than creating a workflow 126 with three distinctsteps 128 for the respective elements 134, the device 106 may merge the interactions 144 into one step 128 including three actions 124 (e.g., a step 128 including separate actions 124 of entering the full name of the user into the “First Name,” “Middle Name,” and “Last Name” textboxes). In this example, the action may be highlighted with a single visual indicator that highlights all of the elements 134.
[0147] In some example embodiments, a device may include a multi -action capture module that detects multi-action steps performed by a user during a user-initiated capture session by maintaining a "cache" of recently detected actions performed by a user during the capture session, and applying a set of merge rules to the recently detected actions to determine whether two or more of the recently detected actions can be depicted as a multi-action step. In some example embodiments, a device may include a multi-action capture module that detects multi-action steps performed by a user during a user-initiated capture session by maintaining a "cache" of recently detected actions performed by a user during the capture session, and identifying two or more logically related interactive elements used to perform two or more of the recently detected actions based on a hierarchical analysis of the DOM. In this approach, the system may analyze the DOM to determine if two or more elements are logically related. As web code (e.g., HTML code) may be hierarchical in nature, elements that are logically grouped under a higher level element / object in the DOM (e.g., <username> and <password> are under and related to a modal for <user login> in the DOM). This relationship may be used to infer that the "children" are related. Such information may be used for determining a label for a multi-action step or for identifying a multiaction step (e.g., if actions were performed using two or more "related" elements).
[0148] In some example embodiments, a device 106 may determine whether to merge two or more actions 124 into one step 128 based on a set of merge rules, instructions, and / or heuristics. The set of merge rules, instructions, and / or heuristics may be applied to combinations of two or more actions 124 to determine whether to merge the two or more actions 124 into one step 128. The device 106 may apply the set of merge rules, instructions, and / or heuristics retrospectively (e.g, after receiving a set of actions 124, the device 106 may test various combinations of consecutive actions 124 to determine whether to merge two or more actions 124 into one step 128). The device 106 may apply the set of merge rules, instructions, and / or heuristics concurrently (e.g, after receiving a current action 124, the device 106 may test various combinations of the action 124 with one or more consecutive preceding actions 124 to determine whether to merge the current action and one or more two or more preceding actions 124 into one step 128). The device 106 may apply the set of merge rules, instructions, and / or heuristics prospectively. For example, after receiving a current action 124, the device 106 may consider whether the current action 124 can be merged with a list of preceding actions 124. If the current action 124 can be merged withthe preceding actions 124 of the list, the device 106 may refrain from creating a step 128 and may append the current action 124 to the list. If the current action 124 cannot be merged with the preceding actions 124 of the list, the device 106 may create a step 128 for all of the preceding actions 124 of the list, clear the list, and add the current action 124 to the list for inclusion in a subsequent step 128.
[0149] The set of merge rules, instructions, and / or heuristics may include a variety of considerations and / or tests. As a first such example, the device 106 may determine whether two or more consecutive actions 124 involve a same element 134 and / or elements of a same type, and if so, may merge the two or more consecutive actions 124 into one step 128. As a second such example, the device 106 may determine whether the elements 134 involved in two or more consecutive actions 124 remain visible in the user interface 114 and have not changed relative locations (e.g., after a scroll event), and if so, may merge the two or more consecutive actions 124 into one step 128. As a third such example, the device 106 may determine whether the elements 134 included in two or more consecutive actions 124 are located near each other in the user interface 114 (e.g., within a number of pixels of one another that is below a threshold number of pixels as measured by a metric such as Euclidean distance, or within a region of a size that is below a threshold region size), and if so, may merge the two or more consecutive actions 124 into one step 128. As a fourth such example, the device 106 may evaluate logical relationships between elements 134 (e.g., based on the hierarchy specified in a document object model 120) to determine whether the elements 134 included in two or more consecutive actions 124 are related and / or included in a particular group (e.g., radio buttons in a radio button array and / or elements included in a particular HTML DIV element 134), and if so, may merge the two or more consecutive actions 124 into one step 128. As a fifth such example, the device 106 may determine whether the elements 134 included in two or more consecutive actions 124 are in a same portion of a user interface 114 or in different portions of the user interface 114 (e.g., a first element 134 in a main portion of a web page 112 and a second element 134 in a different portion of the web page 112, such as a sub-window, popover, tab, or modal dialog), and if in a same portion of the user interface 114, may merge the two or more consecutive actions 124 into a merged step 706. As a sixth such example, the device 106 may determine whether the number of actions 124 included in a merged step 706 is above or below a threshold (e.g., a maximum of six actions 124 per merged step 706). If so, the merged step 706 may be limited to the threshold number of actions 124, and subsequent actions 124 may be added to a new following step 128 and / or a following merged step 706.
[0150] FIG. 7 illustrates an example scenario featuring a merging of actions within a step of a workflow, according to aspects of the present disclosure. The example scenario of FIG. 7 may be performed by the example device 202 of FIG. 2.
[0151] In the example scenario of FIG. 7, a web browser 110 of a device 106 presents a web page 112 including a user interface 114, wherein the user interface 114 includes a set of elements 134. In response to a user 104 requesting a capture 102 of a workflow 126 for a task 116, a content process 210 executed by the web browser 110 may track a set of interactions 144 and may generate steps 128 respectively including one or more actions 124. At a first time 702, the user 104 may perform a first interaction 144 with a first element 134 of the user interface 114 (e.g., a first textbox). The web browser 110 may inform the content process 210 of a first event 612 involving text entry into the first textbox. The content process 210 may generate a first step 128 including an action 124 corresponding to the event 612 (e.g., a step of entering text into the first textbox). The content process 210 may generate a description of the first step 128 to indicate the first action 124 (e.g., “enter text into the first textbox”).
[0152] At a second time 704, the user 104 may perform a second interaction 144 with a second element 134 of the user interface 114 (e.g., a second textbox). The web browser 110 may inform the content process 210 of a second event 612 involving text entry into the second textbox. The content process 210 may determine that the first action 124 and the second action 124 may be merged into a merged step 706 (e.g., because the textboxes are of a same type of element 134 and are close together in a layout of the web page 112, and because the user 104 performed consecutive actions 124 of interacting with the first textbox and the second textbox). The content process 210 may convert the first step 128 into a merged step 706 that includes the first action 124 of interacting with the first textbox and the second action 124 of interacting with the second textbox with the first action 124 of interacting with the first textbox. The content process 210 may update a description of the merged step 706 to indicate the merged actions 124 (e.g., “enter text into the first and second textboxes”).
[0153] At a third time 708, the user 104 may perform a third interaction 144 with a third element 134 of the user interface 114 (e.g., a list). The web browser 110 may inform the content process 210 of a third event 612 involving a selection of an item in the list. The content process 210 may determine that the third action 124 cannot be merged into the merged step 706 including the first action 124 and the second action 124 (e.g., because the list is of a different type of element than the textboxes and / or is not close enough to the first and second textboxes in a layout of the web page 112). Instead, the content process 210 may generate a second step 128 for the third action 124 of interacting with the list. The device 106 may add the second step 128 to the workflow, following the merged step 706 for the merged first action 124 and second action 124 for interacting with the first and second textboxes. In this manner, the device 106 may generate workflows with merged steps 706 in accordance with the techniques presented herein.
[0154] In some example embodiments, a device may include a multi-action step UI that allows a user to view and edit multi-action steps, including uncombining multi action steps and editing visual indicia overlayed on top of a screenshot depicting the multi-action step. The UI may allow a user to validate, reject, and / or edit multi-action steps that are detected by the capture system.
[0155] In some example embodiments, a device 106 may allow a user 104 to view the workflow 126 as a sequence of steps 128 that respectively correspond to one or more actions 124. For example, a side panel 224 presented alongside a web page 112 may show the workflow 126 and various controls for editing the workflow 126. The side panel 224 may play or replay the workflow 126 for the user 104 as a review of the actions 124 and steps 128 performed for the task 116. The side panel 224 may allow the user 104 to reorder two or more actions 124 in the workflow 126 (e.g., reversing an order of a first action 124 and a second action 124). The side panel 224 may allow the user 104 to create additional steps 128 in the workflow 126; duplicate one or more steps 128 in the workflow 126; merge two or more steps 128 of the workflow 126 into one step 128; split a multi-action step 128 including two or more actions 124 into two or more steps 128 respectively including a subset of the two or more actions 124; and / or delete one or more steps 128. The side panel 224 may allow the user to attach descriptions, screenshots, animated images, and / or video to various steps 128 of the workflow 126, and / or to edit one or more screenshots (e.g. , adding an annotation or cropping a screenshot to a selected portion of the user interface 114). The side panel 224 may allow the user 104 to save the workflow 126; apply, undo, and / or redo changes to the workflow 126; name or rename the workflow 126; add or edit a description of the workflow 126; and / or delete the workflow 126. Many such functions may be included in a side panel 224 to allow a user to review and edit a workflow 126.KNOWLEDGE CAPTURE - STORING RECORD OF WORKFLOW
[0156] In some example embodiments, the step 508 of storing a record of the workflow 126 may be performed in a variety of ways. For example, the record of the workflow 126 may include details of the task 116 recorded to capture 102 the workflow 126, such as a name, task type, a date and / or time on which the task 116 was performed, a URL of a web page 112 in which the user interface 114 of the task 116 was presented, an identifier of a local application in which the user interface 114 of the task 116 was presented, and a name and / or other identifier of the user 104 who performed the task 116. The record of the workflow 126 may include details of the sequence 122 of interactions 144 that were performed by the user 104, such as a timestamp of each interaction 144, a relative amount of time between the interaction 144 and a preceding interaction 144), an event type of one or more events included in and / or associated with the interaction 144 (e.g., one or more click or key input events), user input that resulted in the interaction 144 (e.g., a left-click of a mouse button), a title and / or description of each interaction 144, and / or one or morescreenshots 614 of each interaction 144. The record of the workflow 126 may include details of the steps 128 generated from the interactions 144, such as a mapping of the sequence 122 of interactions 144 to the sequence 122 of steps 128, one or more automatic actions performed by the device 106 to alter the sequence 122 of steps 128 (e.g., a merging of two or more actions 124 into a multi-action step 128), and / or an editing of the sequence of steps 128 by a user 104. The record of the workflow 126 may include details of the user interface 114 before, during, and / or after the capture 102 of the workflow 126 (e.g., initial, intermediate, and / or resulting values of user controls such as textboxes).
[0157] In some example embodiments, a device 106 may include a snapshot capture system that captures snapshots of a DOM corresponding to a website being accessed by a user during a capture-session being initiated by the user, wherein the snapshots are used to generate a user- directed UX workflow that guides subsequent through a series of steps performed by the user during the capture session. The device 106 may include a capture workflow that is triggered in response to receiving a user-initiated "capture instruction," wherein the capture workflow includes an interaction monitoring task that causes an interaction monitoring module to monitor interactions 144 by the user 104 with a website, identify potential actions being taken by the user, and determine which of the potential actions are documentable actions; a visual assistance task that causes an overlay module to display a visual indicia "over" the website to display an interactive element of the website that the user is potentially interacting with; and an action capture task that instructs the web device to capture a DOM snapshot in response to detecting a documentable action and transmit the DOM snapshot to a backend that generates a screenshot based on the DOM snapshot.
[0158] In some example embodiments, during a capture 102 of a workflow 126 for a task 116, a device 106 may store one or more representations or “snapshots” of at least a portion of a document object model 120 of the user interface 114 through which the task 116 is being performed. During the course of the performance of the task 116, the document object model 120 may change, e.g., due to the insertion, alteration, relocation, duplication, merging, and / or removal of various elements 134 and / or a layout of the user interface 114. For instance, some actions 124 performed through the user interface 114 may result in the execution of code (e.g., JavaScript) that causes one or more elements 134 to acquire, change, store, load, and / or remove a value (e.g., a value of a textbox) and / or to inject one or more additional elements into the document object model 120 (e.g., additional HTML elements and / or content received from a webserver 118, such as in an AJAX model). In some cases, server-initiated events may cause the document object model 120 to change asynchronously and / or independently of the actions 124 performed through the user interface 114 (e.g., the webserver 118 may transmit, to the device 106, instructions toinject elements and / or content into the document object model 120 after a delay in response to a user-initiated event and / or in a spontaneous manner). The device 106 may record, at one or more points in the performance of the task 116, a snapshot of the document object model 120. The device 106 may locally store the snapshot of the document object model 120 and / or may transmit the snapshot of the document object model 120 to a workflow server 220 (e.g., in a compressed or uncompressed form). The device 106 may include, with the snapshot of the document object model 120, one or more results of the capture 102, such as one or more screenshots 614 of at least a portion of the web page 112 associated with the snapshot of the document object model 120, a state of one or more elements 134 of the user interface 114 based on the document object model 120, and / or descriptions of one or more events 612 arising within the user interface 114 based on the document object model 120.
[0159] In some example embodiments, a device 106 may include a backend system that includes a snapshot processing system configured to execute, analyze, and / or manipulate snapshots of a DOM captured by a respective user device to generate editable screenshots corresponding to respective actions taken by a user of third party website during a user-initiated UX capture session. In some architectures, the device may offload some or all of the processing to the platform's backend compute resources. In some embodiments, the client (e.g., web extension) sends snapshots of the DOM to the backend and the backend executes a "headless" browser that renders the web-based GUI. Such architectures may allow the platform more flexibility to manipulate various aspects of the platform (e.g., higher resolution, editable screenshots) and / or may support the virtualization of corresponding native applications, which may enable the creation of workflows that depict corresponding steps / actions of a user-defined UX workflow also for the relating of native application GUIs.
[0160] In some example embodiments, a workflow server 220 may receive one or more snapshots of a document object model 120 captured during a performance of a task 116. The workflow server 220 may analyze the snapshots of the document object model 120 in various ways, including static and / or dynamic code analysis (e.g., a parser that traverses the structure of an HTML document, XML document, XAML document, CSS document, or the like, or combination thereof, to determine the elements 134 of the user interface 114 and their association with various actions 124 of the user 104). In some example embodiments, the workflow server 220 may re-render a user interface 114 (e.g. , a web page 112) based on a snapshot of the document object model. For example, the workflow server 220 may host a headless web browser 110 that is capable of rendering web pages 112 without a graphical representation to be shown on a display 108 of a device 106. The headless web browser 110 may include a particular software brand and / or version of the web browser; an instance of a web browser 110 on a particular device, such as amobile device with a display 108 having a particular resolution; and / or in a particular context, such as a presentation of the web page 112 while the web browser 110 was in a certain mode or state. The rendering may include executing code embedded in and / or referenced by the user interface 114, such as JavaScript embedded in and / or imported into an HTML document. The rendered web page 112 may include representations of the elements 134 of the user interface 114, and the workflow server 220 may replay the interactions 144 of the user 104 with the user interface 114 (e.g., replaying a sequence of events 612 detected during the capture 102) to step through the performance of the task 116 by the user 104. In this manner, the workflow server 220 may use snapshots of the document object model 120 to inform the determination and / or recording of the workflow 126 by the workflow server 220.
[0161] In some example embodiments, a platform may include a backend virtualization system that comprises a virtual web browser that recreates a user's UX journey through a website based on a series of snapshots of a DOM from a web browser executing the website, wherein the backend virtualization system generates editable user-directed UX workflows comprising screenshots output by the virtual web browser having overlaid visual indica to indicate respective steps undertaken during the UX journey. The platform may include a screenshot editor that allows a user to edit any visual indicia overlaid on the screenshots and / or content included in an interactive element highlighted by the visual indicia. The virtual web browser may output respective screenshots at a higher resolution than the web browser executing on the user device.
[0162] In some example embodiments, a workflow server 220 may analyze one or more snapshots of a document object model 120 at a set of time points during the performance of a task 116 by a user 104. For example, the workflow server 220 may analyze changes to the document object model 120 over the course of the task 116 as elements 134 are created, modified, and / or deleted from the user interface 114, either in response to actions 124 of the user 104, execution of code embedded in the web page 112 (e.g., JavaScript), and / or instructions received from the webserver 118 of the web page 112. The analysis of the document object model 120 may enable the workflow server 220 to understand and extract information about the performance of the task 116 by the user 104. As a first such example, analysis of the document object model 120 may inform the workflow server 220 of steps 128 in the workflow 126 where users 104 commonly make an error, such as a step 128 that is unclear, incomplete, and / or incompatible with a current version of the document object model 120. The workflow server 220 may modify the workflow 126 to clarify and / or correct the step 128 (e.g., by attaching a knowledge indicator 140 to one or more elements 134 associated with the step 128, wherein the knowledge indicator 140 indicates an availability of knowledge 142 that may aid users 104 in completing the step 128). As a second such example, analysis of the document object model 120 may inform the workflow server 220of steps 128 in the workflow 126 that result in unexpected changes to the document object model 120, such as unexpected results of executed code within the web page 112 in response to a step 128 and / or unexpected changes to one or more elements 134 of the document object model 120. These unexpected changes may indicate changes to the document object model 120 that alter a behavior of the user interface 114 (e.g., text input that a textbox previously accepted at a time of capture 102 may unexpectedly be rejected at a time of visual guidance 132). The unexpected results may inform the workflow server 220 of changes to the user interface 114, and the workflow server 220 may automatically update a workflow 126 associated with the user interface 114 (e.g., attaching a knowledge indicator 140 to a textbox element 134 to provide knowledge 142 that certain forms of text input are no longer accepted by the textbox element 134). As a third such example, the workflow server 220 may analyze one or more screenshots associated with a workflow 126 to provide annotations, clarifications, and / or updates to the one or more screenshots, such as resizing, cropping, and / or increasing or decreasing a resolution of the one or more screenshots.
[0163] In some example embodiments, the workflow server 220 may permit a user 104 who has created and / or managed a workflow 126 to edit the workflow 126 and information contained therein. For example, a side panel 224 included in a web browser 110 may feature controls that allow the user 104 to editthe workflow 126 (e.g., by adding, editing, rearranging, and / or removing steps 128 of the workflow 126) and / or to add, edit, or delete screenshots to one or more steps 128.
[0164] In some example embodiments, a platform may include a governance system that ensures that screenshots captured in connection with the generation of a user-directed UX workflow do not depict any sensitive information.
[0165] In some example embodiments, a workflow server 220 may review a workflow 126 and other information provided by users 104 and / or devices 106 based on a governance policy. For example, the policy may limit communication and images to prohibit sensitive, offensive, personal, or otherwise undesirable information in one or more workflows 126 from being presented to other users 104 and / or devices 106. The workflow server 220 may apply the governance policy by detecting content in a workflow 126 that is not consistent with the policy and may edit the workflow 126 to remove and / or replace the content. For example, if a screenshot included in a workflow 126 inadvertently includes personally identifying information (PII) about a user (e.g., name, address, social security number, or the like, or screenshot 614 including an image of a user 104), protected health information about a user (e.g., a portion of a medical record), or offensive or objectionable language, the workflow server 220 may redact the information from the workflow 126, screenshot 614, or the like, before presenting the workflow 126 to other users 104.
[0166] In some example embodiments, a device 106 may include a backend capture system that receives DOM snapshots captured by a respective web extension that captures DOM snapshots of a website being executed by a web browser during a user-initiated capture session and generates a series of screenshots that depict a respective series of actions taken by the user during the web extension, wherein the backend capture system overlays visual indicia to indicate a respective interactive element of the GUI used to perform a respective action in the series of actions. The backend capture system may render a virtual copy of a third-party website in a virtual browser based on a snapshot of a DOM of the website captured by a user device during a user-initiated capture session, and may generate an editable screenshot of the virtual copy of the website as output by the virtual browser, wherein the editable screenshot is used to generate visual documentation of a step in a user-directed UX workflow. The virtual browser may output screenshots of the virtual copy of the website at a higher resolution than a resolution at which the website was output by the user device of the user. In some example embodiments, a device 106 may include a backend capture system that renders a virtual copy of a third-party website in a virtual browser based on a snapshot of a DOM of the website as captured by a user device web browser during a capture session initiated by a user and generates guidance data that is used to visually guide subsequent users through a series of steps performed by the user during the capture session.
[0167] In some example embodiments, a workflow server 220 may use snapshots of a document object model 120 to render screenshots 614 of the user interface 114 in various states, such as before, during, and / or after the performance of certain actions 124 and / or steps 128 of a workflow 126. For example, the workflow server 220 may generate render a graphical representation of the web page 112 into a buffer (e.g., a bitmap surface) and may capture one or more screenshots 614 of the user interface 114 based on the snapshot of the document object model 120 at a point in time and / or in the sequence 122 of interactions 144. Even if the buffer is not shown on a display 108 of a device 106, the workflow server 220 may be capable of using the buffer to capture screenshots of the user interface 114 at various points during the capture 102 of the task 116. The graphical representation of the web page 112 may correspond to different browsers or browser types (e.g. , a desktop browser vs. a mobile browser); different scales and / or resolutions of the web page 112 (e.g., when presented on displays 108 with various resolutions, aspect ratios, refresh rates, display pitch characteristics, or the like); using various accessibility features, such as different-sized fonts and / or screen-reader capabilities; and / or using different input devices, such as a pointer-based input device, a touch-sensitive display, and a haptic and / or gesture-based input device. The capability of generating screenshots based on virtually rendered snapshots of the document object model 120 may enable the workflow server 220 to present versions of a workflow126 that are suitable to specific devices. For example, a workflow 126 of a task 116 may be based on a capture 102 of the task 116 on a workstation device 106 having a large display, but a user 104 may wish to view the workflow 126 of the same task 116 when performed on a mobile device 106 having a small display. The workflow server 220 may use the re-rendering of the snapshot of the document object model 120 to capture and present screenshots 614 that correspond to the steps 128 of the workflow 126 as they may appear on the mobile device 106 of the user 104. The workflow server 220 may also use the re-rendering of the snapshot to generate, regenerate, and / or modify screenshots 614 of various portions of the user interface 114 (e.g., to rescale the screenshots 614 for different dimensions; to insert, change, and / or remove annotations into the screenshots 614 as compared with screenshots 614 during the initial capture 102). As one such example, the workflow server 220 may use screenshots 614 based on the re-rendering of a snapshot of the document object model 120 to allow a user 104 to edit the screenshot (e.g., when presented at a different resolution and / or with a different set of visual indicators, such as highlights 610 of elements 134 of the user interface 114).
[0168] In some example embodiments, a device 106 may include a guidance documentation system that generates guidance data that indicates a series of actions taken by a user with respect to a graphical user interface of a website during a user-directed capture session, wherein the guidance data is structured to indicate a respective interactive element of the GUI that was used to perform a respective action during the capture session, and for each action, target data for the interactive element that was used to perform the action that indicates a respective set of attributes that are indicative of a portion of the DOM that was implicated when the respective action was captured.
[0169] In some example embodiments, the record of the workflow 126 generated by a device 106 and / or workflow server 220 may be represented in many ways to suit various applications. As a first such example, the record of the workflow 126 may be automatically structured as a guidance document, e.g., human-readable documentation that illustrates the steps 128 of the workflow 126 to perform the task 116. As a second such example, the record of the workflow 126 may be structured as a machine-readable and / or machine-executable document (e.g., a declarative document, script, executable application, or the like), which may instruct a device 106 to present the workflow 126 to guide one or more users 104 to perform the task 116 through the user interface 114. That is, a device 106 of the same user 104 or another user 104 may interpret and / or execute the document to present, to a user 104, a workflow 126 as a sequence 122 of steps 128 to perform the action. The workflow 126 may be adapted for a particular user 104 (e.g., an automated translation of human instructions in the workflow 126 into a native language of the user 104, and / or an adaptation of the workflow 126 based on an age, background, technical proficiency,and / or experience level of the user 104). The workflow 126 may be adapted to provide guidance of the task 116 on a particular device 106 (e.g., a workflow including screenshots that correspond to a browser type, a resolution of a display 108 of the device 106, and / or input devices that the user 104 may use to perform the task 116 on the device 106). The workflow 126 may include indicators, metadata, descriptions, and / or screenshots 614 of elements 134 of the user interface 114 that are associated with one or more steps 128 of the workflow 126, wherein such indicators, metadata, descriptions, and / or screenshots 614 of the elements 134 may enable the device 106 to identify corresponding elements 134 in a different version of the user interface 114 (e.g., due to changes in the document object model 120 of a web page 112). The workflow 126 may be adapted and / or personalized for a particular performance of the task 116 (e.g., the workflow 126 may include different information based on the context in which the task 116 is performed, such as an instruction to select a particular location in a list of locations based on a location of the user 104 who is performing the task 116). The workflow 126 may be adapted to include additional information based on a security context of the user 104, device 106, and / or task 116 (e.g., sharing a private workflow 126 generated during a performance of a task 116 by a first user 104 with a second user 104 only if the first user 104 and the second user 104 share a security context, such as membership in a particular private and / or governmental organization). In these and other ways, the device 106 and / or workflow server 220 may generate and alter representations of a workflow 126 for the performance of a task 116 in accordance with the techniques presented herein.
[0170] In some example embodiments, in addition to generating records of a workflow 126 to guide a user 104 to perform a task 116 (e.g., as automatically generated documentation for instructing other users 104 to perform the task 116), a device 106 may generate a record of a workflow 126 that may be consumed by one or more other devices 106. As a first such example, the record of the workflow 126 generated during the capture 102 may be received and processed by an instruction system that generates additional instruction and / or guidance for users 104 to perform the task 116. For instance, the instruction system may generate additional descriptions, animations, visualizations, or the like to provide further guidance of users 104 performing the workflow 126 to perform the task 116 (e.g., a large language model that is capable of providing interactive guidance for users 104 who encounter difficulty performing the steps 128 of the workflow 126). The instruction system may translate the record of the workflow 126 in order to personalize the workflow 126 based on an age, native language, technical proficiency, and / or experience level of a particular user 104. The instruction system may generate commentary, quizzes, and / or gamified presentations of the workflow 126 for users 104 who are performing the workflow 126. As a second such example, the record of the workflow 126 generated during the capture 102 may be received and processed by an automation system that may demonstrate, foranother user 104, the steps 128 of performing the workflow 126. That is, the automation system may automatically perform one or more steps 128 of the workflow 126 in order to instruct one or more other users 104 in how to perform the workflow 126 to complete the task 116. As a third such example, the record of the workflow 126 generated during the capture 102 may be received and processed by an automation system that may perform one or more steps 128 of the workflow 126 in order to aid a user 104 who is having difficulty performing the one or more steps 128 (e.g., as a backup or “autopilot” option) and / or to aid a user 104 who is performing other steps 128 of the workflow 126 (e.g., as a teammate or “copilot”). As a fourth such example, an organization may provide a standard operating procedure (“SOP”) for a task 116 that may be performed by an individual and / or an automated process. A device 106 may receive the SOP of a task 116 and translate it into a workflow 126 in order to provide visual guidance 132 for the individual and / or the automated process to perform the task 116 through the user interface 114. Many such uses of records of workflows 126 may be devised in accordance with the techniques presented herein.GUIDANCE
[0171] In some example embodiments, a device 106 may include a guidance UI that visually guides a current user through series of steps of a user-defined UX workflow that is performed using a GUI of a third-party website, wherein the user-defined UX workflow was captured by a previous user of the third party website during a capture session initiated by the previous user, and wherein the workflow guidance system is configured to analyze the DOM of the third-party website to determine which interactive GUI elements of the third party website to highlight to the current user based on a current state of the DOM and target data captured when generating the user-defined workflow.GUIDANCE - OVERVIEW
[0172] FIG. 8 illustrates a flowchart of an example method of presenting automatically generated user interface documentation according to aspects of the present disclosure. At least a portion of the example method 800 may be performed, for example, by the example device 202 of FIG. 2.
[0173] In general, the example method 800 of FIG. 8, which may be colloquially known as “Guidance,” enables a device 106 to display a step-by-step guided walkthrough or “wizard” for performing a user interface-based task, based on an automatically captured sequence of steps 128 that comprise a workflow 126 for the task 116. In some embodiments, the device 106 may display the steps 128 of the workflow 126 in a side panel shown adjacent to the user interface 114, and may include a title, description, screenshot, and / or visual highlight of an element associated with a current step 128. When the device detects that the user has completed the current step 128, thedevice 106 may advance to the next step 128 of the workflow in order to continue guiding the user 104 through the workflow.
[0174] The example method 800 includes a step 802 of receiving a workflow 126 for a task 116. The workflow 126 includes a sequence of steps 128 to be performed within the user interface 114. The user interface 114 may include a web page 112 of a website shown in a web browser 110, a window of an application executing on the device 106, or the like. Each step 128 may include one or more actions 124 and / or may be associated with one or more elements 134 of the user interface 114. The workflow 126 of the task 116 may have been generated during a capture 102 of a performance of the task 116 through the user interface 114 by a user 104.
[0175] The example method 800 includes a step 804 of presenting the sequence 122 of steps 128 of the workflow 126 during a performance of the task 116 by a user 104. For example, the device 106 may show, adjacent to the user interface 114, a side panel 224 that depicts a list of steps 128 of the workflow 126. Each depicted step 128 may include a title, a description, a screenshot and / or thumbnail of a screenshot, and / or a visual indication of one or more elements 134 associated with the step 128. The user 104 for whom the workflow is presented 126 as guidance in FIG. 8 may include the same user 104 who performed the task 116 during the capture 102 resulting in the workflow 126 and / or may be a different user 104.
[0176] The example method 800 includes a step 806 of presenting a current step 128 of the workflow 126, wherein the presenting guides the user 104 to perform at least one interaction 144 with at least one element 134 of the user interface 114, wherein the interaction 144 corresponds to at least one action 124 of the current step 128 of the workflow 126. For example, the workflow 126 may be based on a document object model 120, and the device 106 may display the visual highlight of at least one element 134 of the user interface 114. The device 106 may determine the one or more elements 134 associated with the one or more actions 124 of the current step 128 by identifying a portion of the document object model that matches or otherwise corresponds to a description of the actions 124 and / or elements 134 included in the current step 128 of the workflow 126.
[0177] In some example embodiments, the device 106 may display a highlight 610 of the one or more elements 134 of the user interface 114 as indicated by the identified portion of the document object model 120. For example, the device 106 may display a highlight 610 of the element 134 of the user interface 114 by identifying a first element 134 within the user interface that is associated with the current step 128, wherein the first element 134 is not responsive to an action 124 of the current step 128; determine a second element 134 within the user interface 114 that is associated with the first step 128, wherein the second element 134 is responsive to the action 124 of thecurrent step 128; and display the highlight 610 of the second element 134 of the user interface 114.
[0178] In some example embodiments, a current step 128 of the workflow 126 includes a description of one or more actions, and each of the one or more actions 124 may be associated with one or more elements 134 of the user interface 114. Each element 134 indicated in the workflow 126 may include at least one attribute indicated by the document object model. Each element 134 may be associated with a first interaction 144 to be performed by the user 104 within the user interface 114, wherein the first interaction 144 corresponds to a first step 128 of the task 116, and wherein the first action 124 is included in the current step 128 of the workflow 126. The device 106 may identify the portion of the document object model 120 that matches or best corresponds to the description of the elements 134 included in the current step 128 of the workflow 126. For example, the device 106 may determine a portion of the document object model 120 that matches or best corresponds to the at least one attribute of the at least one element 134 of the document object model 120. This matching may be helpful if the document object model 120 of the user interface 114, at the time of the capture 102 of the actions 124 of the workflow 126 associated with a particular element 134, differs from the document object model 120 of the user interface 114 at a later time when a user 104 is performing the steps 128 of the displayed workflow 126. In such cases, the device 106 may use the stored information to determine which element 134 of the document object model 120 corresponds to a particular action 124 and / or step 128 of the workflow 126.
[0179] In some example embodiments, at least one attribute of the at least one element 134 of the document object model 120 may include at least one of: a location of the at least one element 134 within a schema of the document object model 120, at least one attribute of the at least one element 134 defined by the document object model 120, or at least one attribute of at least one user control that is associated with the at least one element 134 of the document object model 120. The description of the current step 128 may include at least two attributes of at least one element 134 of the document object model 120, and the device 106 may identify the portion of the document object model 120 that matches or best corresponds to the description of the current step 128 included in the workflow 126 by performing a comparison of each attribute of the at least two attributes with at least one element 134 of the document object model 120. In some embodiments, this may include determining a score for each element 134 of the document object model 120 for each attribute of the at least two attributes, wherein the score is based on the comparison, and identifying the portion of the document object model 120 that matches or best corresponds to the description of the current step 128 included in the workflow 126. The identification may be based on determining which portion of the document object model 120, including an element 134, hasthe highest sum of scores for each attribute of the at least two attributes. In some embodiments, the comparisons may be evaluated as a set of “games” and / or “mini-games” for matching the information about an element 134 of the document object model 120 at the time of capturing the workflow 126 with one of the elements 134 of the document object model 120 at the time of guiding a user 104 to perform the workflow 126. The device 106 may perform the comparison of each attribute of the at least two attributes with at least one element 134 of the document object model 120 by comparing a label associated with at least one element 134 of the document object model 120 associated with the set of actions 124 performed by the user 104 while performing the task 116 within the user interface 114 with a label associated with at least one element 134 of the document object model 120 during the current step 128 of a current performance of the workflow 126 by a user 104; comparing at least one attribute included in at least one element 134 of the document object model 120 associated with the set of actions 124 performed by the user 104 while performing the task 116 within the user interface 114 with at least one attribute included in at least one element 134 of the document object model 120 during the current step 128 of a current performance of the workflow 126 by a user 104; comparing a geometry of at least one element 134 of the document object model 120 associated with the set of actions 124 performed by the user 104 while performing the task 116 within the user interface 114 with a geometry of at least one element 134 of the document object model 120 during the current step 128 of a current performance of the workflow 126 by a user 104; comparing a style selector associated with at least one element 134 of the document object model 120 associated with the set of actions 124 performed by the user 104 while performing the task 116 within the user interface 114 with a style selector associated with at least one element 134 of the document object model 120 during the current step 128 of a current performance of the workflow 126 by a user 104; and / or comparing a related element 134 that is associated with at least one element 134 of the document object model 120 associated with the set of actions 124 by the user 104 while performing the task 116 within the user interface 114 with a related element 134 that is associated with at least one element 134 of the document object model 120 during the current step 128 of a current performance of the workflow 126 by a user 104. In some embodiments, each attribute of the at least two attributes is associated with a weight, and the score for each element 134 of the document object model 120 for each attribute of the at least two attributes is further based on the weight associated with the attribute.
[0180] In some example embodiments, if a device 106 fails to identify a portion of the document object model 120 that matches the one or more elements 134 of the current step 128 of the workflow 126, the device 106 may present to the user 104 an alternative instruction associated with the current step 128 of the workflow 126. Alternatively or additionally, if a device 106 failsto identify a portion of the document object model 120 that matches the description of the current step 128 of the workflow 126, the device 106 may determine an alternative description for the current step 128, wherein the alternative description identifies at least one alternative element 134 within the user interface 114 that is associated with the current step 128 of the workflow 126. The device 106 may substitute the description of the current step 128 of the workflow 126 with the alternative description of the current step 128 (e.g., “healing” the workflow 126).
[0181] The example method 800 includes a step 808 of advancing to a next step 128 of the workflow 126 upon detecting a completion of the current step 128 by the user 104. The device 106 may determine the completion of the current step 128 by comparing the actions 124 associated with the step 128 with one or more actions 124 performed by the user 104, such as one or more events 612 involving one or more elements 134 of the user interface 114. When all actions 124 of the current step 128 are mapped to one or more events 612 arising during the performance of the task 116 by the user 104, the device 106 may determine that the current step 128 is complete and may advance the workflow 126 to the next step 128 of the workflow 126. In some embodiments, upon determining a completion of the current step 128 of the current performance of the workflow 126 by the user 104, the device 106 may remove a highlight 610 of at least one content element 134 of the user interface 114 that is associated with the current step 128 of the workflow 126 before advancing to the next step 128 of the current performance of the workflow 126. When all steps 128 are complete, the device 106 may determine that the workflow 126 is complete and that the user 104 has been successfully guided to complete the task 116. In some example embodiments, the device 106 may present a message indicating a completion of the task 116, hide the completed workflow 126, and / or close a side panel 224 in which the workflow 126 was presented.GUIDANCE - RECEIVING WORKFLOW
[0182] In some example embodiments, the step 802 of receiving a workflow 126 of a task 116 including a sequence of steps 128 involving a user interface 114 may be performed in various ways.
[0183] In some example embodiments, the device 106 may receive, from a user 104, a command for guidance to perform the task 116, and may request and receive the workflow 126 for the task 116 from a workflow server 220 in order to provide the requested guidance. For example, the user 104 may initiate a command to the device 106 to provide guidance with purchasing an airline ticket. The device 106 may retrieve, from a workflow server 220, a workflow 126 thatis associated with using a particular user interface 114 (e.g., a web page 112 of an airline) that enables the user 104 to purchase airline tickets. The device 106 may automatically open a web browser 110 to a web page 112 associated with the workflow 126 to begin the presentation of the workflow 126 tothe user 104. Alternatively or additionally, the device 106 may automatically acquire, install, configure, and / or start a native application that includes a user interface 114 associated with the workflow 126 in order to begin the presentation of the workflow 126 to the user 104. In some example embodiments, the device 106 may provide instructions to the user 104 that enable the device 106 to begin the presentation of the workflow 126 to the user 104 (e.g., requesting the user 104 to download and run an executable script or code that presents the workflow 126; requesting the user 104 to open a web browser 110 to a particular web page 112; and / or requesting the user 104 to download, install, and / or start a native application that includes the user interface 114 associated with the workflow 126).
[0184] In some example embodiments, the device 106 may store a workflow 126 before receiving a request from a user 104 for guidance (e.g., as part of a cache of workflows 126 involving a user interface 114 of an application that is installed on the device 106, and / or as part of a cache of workflows 126 of user interfaces 114 included in web pages 112 that the user 104 is likely to visit). In some example embodiments, the device 106 on which the capture 102 was performed may be the same device 106 on which the workflow 126 is presented in FIG. 8, and the workflow 126 may have been previously stored on the device 106 upon completion of the capture 102. In some example embodiments, the device 106 on which the capture 102 was performed may be a different device 106 than the device 106 that stores the workflow 126. In some example embodiments, upon receiving a request from a user 104 for guidance with a task 116, the device 106 may retrieve one of the previously stored workflows 126 and may initiate the presentation of the workflow 126 to the user 104.
[0185] In some example embodiments, the device 106 may encounter a user interface 114 that is associated with (e.g., embeds, imports, references, links to, or the like) one or more workflows 126. For example, a browser process 208 may monitor web pages 112 that a user 104 visits in a web browser 110 and may determine that one of the web pages 112 is associated with one or more workflows 126 (e.g., based on a reference to the workflow 126 included in a document object model 120 of the web page 112, such as an HTML reference element 134). Conversely, the browser process 208 may monitor workflows 126 during a browsing session of a user 104 and may identify and load a web page 112 associated with a selected workflow 126. In some example embodiments, a webserver 118 that provides a document object model 120 for a web page 112 may be the same as the workflow server 220 that provides a workflow 126 for a user interface 114 of the web page 112 (e.g., a collocation of the web page 112 and the workflow 126 on the same webserver 118). In some example embodiments, a webserver 118 that provides a document object model 120 for a web page 112 may be different than the workflow server 220 that provides a workflow 126 for a user interface 114 of the web page 112 (e.g., a workflow 126 provided by afirst party for a web page 112 associated with a second party). In some example embodiments, a workflow server 220 that provides a workflow 126 for a user interface 114 of a web page 112 may be associated with and / or personal to the user 104 (e.g., a workflow server 220 that provides private workflows 126 limited to members of a private and / or governmental organization that includes the user 104), or may be a public workflow server 220 that provides workflows 126 to any users 104 for any user interface 114. In some example embodiments, a workflow 126 may be included in the resources of a web page 112. For instance, when a device 106 requests the resources of the web page 112 or native application, a webserver 118 or application server may transmit a collection of resources to render the web page 112 or native application (e.g., one or more HTML files, CSS files, XML files, XAML files, JavaScript files, images, human-readable documents, or the like), and may also transmit a workflow 126 associated with one or more user interfaces 114 included in the web page 112 or native application.GUIDANCE - PRESENTING WORKFLOW STEPS
[0186] In some example embodiments, the step 804 of presenting the sequence of steps 128 during a performance of the workflow 126 may be performed in various ways. For example, a device 106 may present a workflow 126 of a task 116 as a sequence 122 of steps 128 in an area of the user interface 114, such as a window inserted into an empty area of the user interface 114. If the user interface 114 is included in a web page 112 of a web browser 110, the device 106 may expand a size of the web page 112 within the web browser 110 (e.g., creating additional space within the web page 112) for a new region (e.g., an HTML DIV) that includes a list of steps 128 comprising the workflow 126. The device 106 may reduce the size of the content of the web page 112 to create space for the new region to present the workflow 126. Alternatively or additionally, the web browser 110 may create a new browser window and / or a new tab of an existing browser window, and may show present the workflow 126 in the new browser window or tab. If the user interface 114 is included in a native application executing on a device 106, the device 106 may present the workflow 126 as a new window, such as a window in a separate application that presents the workflow 126. In either web browsers 110 or native applications, the device 106 may present the workflow 126 as a modal or modeless popover window, pop-under window, or other region inserted into a Z-order of the windows or other content panes of the device 106. In an augmented reality or virtual reality presentation, the device 106 may present the workflow 126 in a separate floating window that the user may size, locate, and / or otherwise manage in relation to the user interface 114 associated with the workflow 126.
[0187] In some example embodiments, a device 106 includes a guidance UI that includes a side panel that illustrates a series of steps of a user-directed UX workflow that the guidance UI is currently guiding a user through.
[0188] In the example architectural view of FIG. 2, the web browser 110 includes a side panel 224 that presents content alongside a web page 112. If the workflow 126 is related to a user interface 114 included in a web page 112, the side panel 224 may present the workflow 126 concurrently with and adjacent to the presentation of the user interface 114 in the web page 112. For example, a browser process 208 may monitor web pages 112 that are visited by a user 104. The browser process 208 may detect that a particular web page 112 includes (e.g., embeds, references, links to, or the like) a workflow 126 that is associated with a user interface 114 included in the web page 112. Alternatively or additionally, the browser process 208 may determine that he workflow server 220 stores a workflow 126 that is associated with the web page 112, and may retrieve the workflow 126 from the workflow server 220 for presentation in the side panel 224 alongside the web page 112. In some example embodiments, the side panel 224 may ordinarily be hidden, and the browser process 208 may cause the side panel 224 to be displayed in the web browser 110 in order to present the workflow 126 of the user interface 114 alongside the web page 112. In some example embodiments, the side panel 224 may ordinarily show other content (e.g., bookmarks, search tools, advertisements, or the like), and the browser process 208 may cause the side panel 224 to show the workflow 126 when the web page 112 is displayed by the web browser 110. The side panel 224 may present the workflow 126 as a sequence 122 of steps 128. For each step 128, the side panel 224 may show a name, title, description, image, thumbnail of an image, or the like. The side panel 224 may show all steps 128 of a workflow 126 at once (e.g., as the complete sequence 122 of steps 128 to be performed to complete the workflow 126 for the task 116); a subset of the steps 128 of the workflow 126 (e.g., a first page of the workflow 126 including a small number of steps 128, wherein the user 104 may browse through the pages of the workflow 126); and / or a single step 128 of the workflow 126 at a time (e.g., a current step 128 of the workflow 126, to be replaced by a next step 128 of the workflow 126 upon completion of the current step 128).
[0189] In some example embodiments, a device 106 includes a documentation system that generates user-directed user experience (UX) workflows that visually indicate a respective series of actions taken by a user with respect to a GUI of a website or web application during a user- initiated capture session, wherein a user-directed user experience (UX) workflow comprises a series of screenshots captured by a UX capture system during the user-initiated capture that highlight respective interactive elements used to perform respective actions of the UX workflow, and a respective text label for each action of the UX workflow. In some example embodiments, a device 106 includes a guidance UI that visually guides current users of a third-party application or website through user-directed UX workflows documented by a previous user of the third-party website by causing a web browser of a user device accessing the third-party website to overlayspecific visual guidance indicia at specific regions of a GUI of the third-party website based on a current state of the DOM and guidance data generated by a respective user device of the previous user during a user-initiated capture session.
[0190] In some example embodiments, each step 128 of a workflow 126 is associated with one or more elements 134 of the user interface 114. The device 106 may visually associate each step 128 with one or more elements 134 of the user interface 114. For example, if the user 104 is to click on or an element to perform a particular step 128 of the workflow 126, the device 106 may insert, into the user interface 114, a highlight 610 of one or more elements 134 of the user interface 114 that are associated with the particular step 128. Alternatively or additionally, the device 106 may present one or more screenshots 614 of the particular step 128 that depicts the one or more elements 134 of the user interface 114 that are associated with the particular step 128 (e.g., in the side panel). The one or more screenshots 614 may include an animating image or a video of the step 128 being performed using the one or more elements 134 of the user interface 114 (e.g., an animated depiction of a user moving a pointer to a list, clicking the list to view a drop-down set of options, moving the pointer to a particular option, and clicking on the option to complete the step 128). If a step 128 includes two or more elements 134, the device 106 may capture one screenshot 614 that depicts both elements 134, or may capture two or more screenshots 614 that respectively depict one element 134. If a step 128 includes two or more elements 134, the device 106 may include one visual indicator of both elements 134, such as a highlight 610 that encompasses two or more elements 134 (e.g., where the two or more elements 134 are positioned close together). Alternatively, the device 106 may present a first highlight 610 that encompasses the first element 134 of the user interface 114 and a second highlight 610 that encompasses the second element 134 of the user interface 114 (e.g., where the two elements 134 are positioned far apart in the user interface 114, or where a third element 134 that is not included in the step 128 is positioned between the two elements 134). The visual indicators of the elements 134 of the user interface 114 are not limited to highlights 610 and screenshots 614, but may also include, e.g., coloring or shading; different visual styles for the elements 134; a resizing and / or repositioning of the elements 134 within the user interface 114; an animation applied to the elements 134 (e.g., causing the elements 134 to grow, vibrate, or the like); and / or visual connectors between the step 128 and the associated elements 134 (e.g., lines drawn between a step 128 and the elements 134 involved in the step 128).
[0191] In some example embodiments, a workflow 126 may integrate one or more steps 128 of a workflow 126 with particular areas of a user interface 114. For example, the workflow 126 may indicate that for a particular step 128, a visual effect is to be applied to a particular portion of a web page 112, such as a region of an indicated shape, size, location, or the like. The visual effectmay be associated with one or more elements 134 of the user interface 114 (e.g., the specified location of the visual indicator of a step 128 may coincide with the location of an element 134 of the user interface 14 that is associated with the step 128). The visual effect for a step 128 may be associated with portions of a user interface 114 that do not include one or more elements 134 of the user interface 114 that are associated with the step 128 (e.g., a highlight of a text label of the web page 112 to highlight the type of information to be included in a textbox associated with the step 128). The workflow 126 may specify locations within the user interface 114 according to a visual layout of the web page 112 (e.g., absolute or relative coordinates within the web page) and / or according to one or more elements of a document object model 120 that are associated with the step 128 (e.g., a visual style specified in CSS for all elements of the document object model 120 having a specified class, such as all clickable elements 134 within a particular HTML DIV of a web page 112).
[0192] In some example embodiments, a device 106 includes a guidance UI that is configured to visually guide user through user-directed UX workflows that include multi-action steps.
[0193] In some example embodiments, a workflow 126 may include one or more merged steps 706 that are respectively associated with two or more actions 124, such as in the example scenario of FIG. 7. As previously discussed, two or more actions 124 may be merged if the actions 124 involve one element 134 of the user interface 114 (e.g., a click event on a textbox and a text entry event in the textbox); if the actions 124 involve locating an element 134 in the user interface 114 (e.g., scrolling to bring a button into view in a web page 112) and interacting with the element 134 (e.g., clicking the button once it is in view); if the actions 124 involve two or more elements 134 of a same type (e.g., consecutive actions 124 involving text entry into two or more textboxes); and / or if the actions 124 involve two or more elements 134 that are positioned close together and / or related (e.g., two or more elements 134 that included in one HTML DIV element and / or are positioned close to each other in a layout of the user interface 114). In such cases, the device 106 may present a workflow 126 including one or more merged steps 706. For example, a merged step 706 may indicate each of the two or more elements 134 that are included in the merged step 706 (e.g., with highlights 610 concurrently and / or consecutively applied to each element 134) and / or each of the actions 124 included in the merged step 706 (e.g., a first action 124 of text entry into a first element 134 and a second action 124 of text entry into a second element 134, or a combined action 124 of text entry into the first and second elements 134). In some example embodiments, the device 106 may present the merged step 706 as a set of sub-steps, each associated with one or more elements 134 of the user interface 114 and / or one or more actions 124. That is, the device 106 may present the workflow 126 with a hierarchical and / or multi-levelorganization, where one or more merged steps 706 include one or more sub-steps to be presented together.GUIDANCE - PRESENTING CURRENT STEP OF WORKFLOW
[0194] In some example embodiments, the step 806 of presenting a current step 128 of a workflow 126 may be performed in various ways.
[0195] In some example embodiments, a device 106 may include a guidance UI that displays, for a website, a sequence of steps to perform a workflow; identifies a current step of a user performing the workflow; identifies at least one element of the user interface that is associated with the current step; and presents visual guidance indicia associated with at least one element. The "visual guidance indicia" can be a highlight of the element, a pointer to the element, and / or a description or instruction related to the element shown in the side panel. The guidance UI may display, for a website, a sequence of steps of a user-directed UX workflow captured during a capture session initiated by a previous user; identify a current step of a user performing the workflow; identify a web page of the website that includes at least one interactive element that is associated with the current step; and visually guide the user to the at least one interactive element of the identified web page to complete the current step.
[0196] In some example embodiments, in addition to presenting the workflow 126 as a sequence 122 of steps 128, a device 106 may track a performance of the workflow 126 by the user 104 in order to guide the user through a current step 128 of the workflow 126. For example, the workflow 126 may associate each step 128 with one or more elements 134 of the user interface 114 and one or more actions 124 to be performed through the one or more elements 134 of the user interface 114. The device 106 may monitor the occurrence of user input and / or events that indicate that the user 104 has performed an interaction 144 with each of the one or more elements 134 of the user interface 114, wherein each interaction 144 corresponds to an action 124 of the workflow 126. For example, the device 106 may detect the occurrence of a button-click event that is associated with a button element 134 of a first step 128 of the workflow 126 and a text entry event that is associated with a textbox included in a second step 128 of the workflow 126. The device 106 may interpret the button-click event as a completion of the first step 128 of the workflow 126 and the text entry event as a completion of the second step 128 of the workflow 126. At a particular time, the device 106 may determine that a user 104 who is performing a task 116 is at a current step 128 of the workflow 126 (e.g., a first step 128 of the workflow 126 for which no steps 128 have yet been completed). The device 106 may specifically indicate, highlight 610, and / or otherwise provide guidance selectively for the elements 134 and / or the actions 124 included in the current step 128 (e.g., by hiding and / or removing highlights 610 or other guidance features for elements 134 associated with other steps 128 of the workflow 126). When the user 104 performs the actionsassociated with the one or more elements 134 and / or one or more actions 124 of the workflow 126, the device 106 may advance to a next step 128 of the workflow 126. The advancing may include hiding and / or removing the highlights 610 and other visual indicators associated with the elements 134 and / or actions 124 included in the previous current step 128 of the workflow 126and showing and / or adding highlights 610 and other visual indicators associated with the elements 134 and / or actions 124 included in the new current step 128 of the workflow 126. In this manner, the device 106 may incrementally and interactively guide the user 104 through the steps 128 of the workflow 126 on a just-in-time basis.
[0197] FIG. 9 illustrates an example scenario featuring a presentation of a current merged step of a workflow, according to aspects of the present disclosure. The example scenario of FIG. 9 may be performed by the example device 202 of FIG. 2.
[0198] In the example scenario of FIG. 9, a workflow 126 is presented in relation to a web page 112 presented in a web browser 110, wherein the web page 112 includes a user interface 114 through which a user 104 is to perform a task 116. The workflow 126 is presented in a side panel 224 of the web browser 110 as a sequence 122 of steps 128 arranged in the order in which the steps 128 are to be performed. The workflow begins with a first step 128, and at a first time 902, the device 106 selects the first step 128 as a current step 128 of the workflow 126. The side panel 224 highlights the first step 128 in the sequence 122 of steps 128 to indicate that it is the current step 128 of the workflow 126. Additionally, the device 106 inserts a highlight 610 of one or more elements 134 of the user interface 114 that are associated with the current step 128 of the workflow 126. More particularly, the first step 128 of the workflow 126 is also a merged step 706 that includes two actions 124 respectively associated with one of two elements 134 of the user interface 114. At the first time 902, the device 106 may insert the highlight 610 of a first element 134 that is associated with a first action 124 (e.g., a first sub-step) of the current merged step 706 of the workflow 126. The side panel 224 may also highlight a portion of the merged step 706 that indicates the current first action 124 and / or first sub-step to be performed by the user 104. Upon detecting a first event 612 that is associated with the first element 134 of the user interface 114 and the first action 124 of the current merged step 706, the device 106 may determine that the first action 124 of the current merged step 706 is complete, but that the current merged step 706 is not yet complete due to the inclusion of the second action 124 of the current merged step 706. Therefore, at a second time 904, the device 106 may insert the highlight 610 of a second element 134 that is associated with a second action 124 (e.g., a second sub-step) of the current merged step 706 of the workflow 126. The side panel 224 may also highlight a portion of the merged step 706 that indicates the current second action 124 and / or second sub-step to be performed by the user 104. Upon detecting a second event 612 that is associated with the second element 134 of the userinterface 114 and the second action 124 of the current merged step 706, the device 106 may determine that the second action 124 of the current merged step 706 is complete, and that the current merged step 706 is also complete. The device 106 may remove the highlight 610 of the second element 134 of the user interface 114, the highlight of the second action 124 or sub-step of the current merged step 706, and the highlight of the current merged step 706. The device may advance to a second step 128 of the workflow 126, the side panel 224 may highlight the second step 128 as the new current step 128, and the user interface 114 may include a highlight 610 of one or more elements 134 that are associated with the new current step 128 of the workflow 126. In this manner, the device 106 may provide visual guidance 132 by presenting a workflow 126 in a side panel 224 of a web browser 110, and by tracking and highlighting of a current step and the associated elements 134 of the user interface 114, including merged steps 706 and actions or substeps thereof.
[0199] In some example embodiments, a device 106 includes a description of an element of a user interface, wherein the element is associated with a step in a UX workflow captured from a previous DOM during a previous user-initiated capture session, and the guidance system uses the description to identify a corresponding element in a current version of the user interface and to present visual indicia associated with the corresponding element, wherein the visual indicia guide a current user to perform the step during a current instance of the UX workflow. The device 106 may include an object identification system that determines, for each step in the user-defined UX workflow captured by the previous user, which interactive element in the third party GUI to highlight to the current user based on the target data for the respective step and the current state of the GUI.
[0200] In some example embodiments, the step 806 of presenting a current step of the workflow may include a target determination step, wherein the device 106 determines one or more elements 134 of a user interface 114 correspond to a particular step 128. In some cases, the user interface 114 may not have significantly changed between the capture 102 of the workflow 126 during a performance of a task 116 by a user 104 and the presentation of the workflow 126 as visual guidance 132 for a user 104 to perform the task 116. As a result, each step 128 of the workflow 126 may indicate one or more elements 134 according to various features (e.g., an identifier, a name, a location in a document object model 120, or the like), and the features may uniquely and correctly identify the corresponding one or more elements 134 of the user interface 114 at the time of visual guidance 132. However, in other cases, the user interface 114 may have changed between the capture 102 of the workflow 126 during a performance of a task 116 by a user 104 and the presentation of the workflow 126 as visual guidance 132 for a user 104 to perform the task 116. For example, a developer or software process may have redesigned, reorganized, recoded, orotherwise altered the document object model 120, such that recorded features of the elements 134 of the workflow 126 may no longer uniquely and correctly identify the corresponding one or more elements 134 of the user interface 114 at the time of visual guidance 132. For instance, an action 124 may be associated with a textbox that, at the time of capture 102, is identified by a particular identifier, name, caption, appearance, location within the document object model 120, relationships with one or more other elements 134 of the document object model 120 (e.g., a child of a parent element), and / or geometric features within a rendering of the user interface 114 (e.g., a size and / or coordinates within a web page 112). However, at the time of visual guidance 132, the document object model 120 may include the textbox with different identifier, name, caption, appearance, path within the document object model 120, relationships with one or more other elements 134 of the document object model 120, and / or geometric features within a rendering of the user interface 114. Alternatively, the developer may have entirely removed the element 134 and / or replaced it with one or more elements 134 of a different type (e.g., changing a checkbox control to a radio button control). Due to such changes, the recorded features may ambiguously identify or correspond to two or more elements 134 of the user interface 114 at the time of visual guidance 132; may fail to identify or correspond to any elements 134 of the user interface 114 at the time of visual guidance 132; and / or may uniquely but incorrectly identify a different element 134 of the user interface 114. As a result, the workflow 126 cannot be followed in a simple and routine manner to identify the elements 134 of the user interface 114 for each step 128 of the workflow 126.
[0201] In some example embodiments, a device 106 may use a variety of additional techniques to present a current step of the workflow 126 to a user 104. As a first such example, if a current step 128 of the workflow 126 involves an element 134 of a web page 112 that is not visible through a viewport of the web page 112 (e.g., the visible portion of the web page 112 given a size of a display 108 and / or window of a web browser 110), the device 106 may automatically scroll the web page 112 in the direction of the element 134 until the element 134 is visible to the user 104. As a second such example, if a current step 128 of the workflow 126 involves an action 124 outside of the web page 112 (e.g., on a different web page 112 or provided outside of the web browser 110, such as a multi -factor authentication code that is transmitted by email or simple message service (SMS)), the device 106 may include, in the step 128 of the workflow 126, an instruction to perform the action 124 outside of the web page 112. The device 106 may also include, in the step 128 of the workflow 126, an option for the user 104 to indicate that the step 128 has been completed (e.g., “find the multi-factor authentication code sent by email and then click this button to continue”). Alternatively, upon presenting the instruction to the user 104, the device 106 may advance to the next step 128 and / or the next action 124 of the current step 128(e.g. , “after finding the multi-factor authentication code, input the multi-factor authentication code into the code textbox”). As a third such example, if the device 106 detects that the user 104 is diverting from the workflow 126 and / or is not following the workflow 126 (e.g., by performing the steps 128 and / or actions 124 out of order or incorrectly), the device 106 may offer to restart the workflow 126, to provide additional information about the workflow 126 (e.g., knowledge 142 associated with the current step 128 and / or action 124 of the workflow 126) to aid the user, or to terminate the workflow 126 if the user 104 finds the workflow 126 to be unhelpful, incorrect, and / or unnecessary.
[0202] In some example embodiments, a device 106 may include respective target data for each respective step of a series of steps in a UX workflow captured from a previous DOM during a previous user-initiated capture session, each respective target data having respective metadata from the previous DOM that is indicative of a specific interactive element in a GUI of a third- party website that was used by the previous user to perform a respective step in the UX workflow, wherein the guidance system uses respective target data to identify a respective interactive element in the third party GUI to highlight to a current user based on a current state of a current DOM generated by a web browser of the current user being guided through the UX workflow.
[0203] In some example embodiments, a device 106 may be configured to perform a target identification process, wherein a set of features of an element 134 included in a step 128 of a workflow 126 are compared with the elements 134 existing within a current version of the document object model 120 at the time of visual guidance 132. That is, the target identification process may seek to determine which element in a current version of the document object model 120 during visual guidance 132 most closely corresponds to an element 134 of a previous version of the document object model 120 during capture 102. The target identification process may involve comparing each of a variety of features, such as identifying data (e.g., an alphanumeric identifier, a name, a path, an element type, or the like), metadata (e.g., a caption associated with each element), relationships between the element 134 and other elements 134 of the document object model 120, a location and / or path of the element 134 in the document object model 120, and / or geometric features of the element in the user interface 114 (e.g., the coordinates and shape of the element 134 in a web page 112). The target identification process may prioritize matching or mismatching of some features of an element 134 over matching or mismatching other features of the element 134 (e.g., a matching element identifier and / or element name may be considered more reliable than matching locations and / or paths in the document object model 120, which in turn may be considered more reliable than matching geometric features of the elements 134 in a web page 112). Based on the comparison, for a current step 128 of the workflow 126, the device 106 may identify an element 134 of the current version of the document object model 120 at thetime of visual guidance 132 that most closely resembles an element 134 associated with the current step 128 in the previous version of the document object model 120 at the time of capture 102. The device 106 may then present the current step 128 with an indication of the determined element 134, track the actions 124 of the user for a completion of one or more actions 124 associated with the determined element 134 to detect a completion of the current step 128, and then advance to a next step 128 of the workflow 126.
[0204] In some example embodiments, a device 106 may include an element identification system that identifies an element in a current version of a web page that corresponds to a target element in a previous version of the web page, wherein the identifying is based on a weighted comparison of features of the element and corresponding features of the target element.
[0205] In some example embodiments, for a particular step 128 that includes an element 134, the previously described target evaluation process may include a score-based comparison of the features of the element 134 in a previous version of the document object model 120 at the time of capture 102 with the corresponding features of the element 134 in a current version of the document object model 120 at the time of visual guidance 132. For example, for each feature of the element 314, the device 106 may perform a comparison of the feature of the element in the previous version of the document object model 120 at the time of capture 102 with the corresponding feature of each element 134 in a current version of the document object model 120 at the time of visual guidance 132. In some example embodiments, the comparison may involve scoring. For example, at the time of capture 102, the document object model 120 may have identified an element 134 included in a step 128 by the name: “button submit form.” At the time of visual guidance 132, the document object model 120 may no longer include any element 134 with the name: “button submit form,” but may include a first element 134 with the name: “btn submit form” and a second element 134 with the name: “button erase form.” The device 106 may compare the names and may assign a score based on each comparison. For example, the device 106 may use Levenshtein distance scoring to determine a Levenshtein distance of 3 for the first element 134 and a Levenshtein distance of 6 for the second element 134. Because the Levenshtein distance is inversely proportional to similarity of strings, the scores for each representation of the element 134 may be determined according to a reciprocal of the Levenshtein distance, resulting in a score for the first element 134 of 0.33 and a lower score for the second element 134 of 0.19. In some example embodiments, the device 106 may perform scored comparisons of a variety of features, and may aggregate the scores of the comparisons (e.g., as a sum, mean, median, maximum, or the like). The device 106 may identify an element 134 of the current version of the document object model 120 at the time of visual guidance 132 that has a highest score among all of the elements 134 of the document object model 120 as the element 134with the highest correspondence and / or probability of matching the element 134 included in the step 128. In some example embodiments, the comparison may be based on a score and / or confidence threshold, e.g., an element 134 of the current version of the document object model 120 may be determined to correspond to an element 134 associated with a step 128 only if its score is above a threshold value, and / or only if its score can be determined within an acceptable degree of confidence.
[0206] In some example embodiments, the comparisons of elements 134 of the current version of the document object model 120 at the time of visual guidance 132 may not only be scored, but may be subjected to a weighted comparison. For example, in addition to calculating scores for respective features of each element 134, the device 106 may aggregate the scores according to respective weights of each feature. For example, for a first feature that is likely to be static and / or highly indicative of an identity of the element 134 (e.g., an identifier and / or name), the weight attributed to the first feature, and to scores based on the first feature, may be high. For a second feature that is likely to be dynamic and / or not highly indicative of an identity of the element 134 (e.g., the geometric features of the element 134 in a rendering of the user interface 114), the weight attributed to the second feature, and to scores based on the second feature, may be low. The device 106 may aggregate the scores resulting from the comparisons based on an aggregation (e.g., a sum, mean median, maximum, or the like) of the products of the score and the weight of respective features of the element 134. In this manner, the device 106 may perform a weighted and prioritized evaluation of the features of each element 134 to determine the element 134 in a current version of the document object model 120 at the time of visual guidance 132 that most closely corresponds to an element included in a step 128 of the workflow 126.
[0207] In some example embodiments, a device 106 may include an element identification system that performs an identification of an element in a current version of a web page that corresponds to a target element in a previous version of the web page, wherein the identification is based on a combination of aspect scores, and each aspect score is based on a weighted comparison of features of an aspect of the element and features of the aspect of the target element.
[0208] In some example embodiments, the features included in a weighted comparison may be organized into different aspects, wherein each aspect considers a particular type of features of the elements 134. Aspects may include, for example, features based on element attributes of the elements 134; features based on locations of the elements 134 within the document object model 120; features based on relationships between the element 134 and other elements 134 of the document object model 120; and / or features based on the geometry of the elements 134 within a rendering of the user interface 114 (e.g., a rendering of a web page 112 including the elements 134). Each aspect may include a weight, and each element 134 may be evaluated to determine aweighted score for each aspect. Alternatively or additionally, each feature of each aspect may include a weight, and each element 134 may be evaluated to determine a weighted score for each feature. The scores of the individual features and / or aspects may be aggregated for each element 134 (e.g., as a sum, mean, median, maximum, or the like) to determine which element 134 in a current version of the document object model 120 at a time of visual guidance 132 most closely corresponds to the features of the element 134 in the version of the document object model 120 at the time of capture 102 that are indicated in a step 128 of the workflow 126.
[0209] In some example embodiments, a device 106 may include an element identification system that identifies an element in a current version of a web page that corresponds to a target element in a previous version of the web page, wherein the identifying is based on a weighted comparison of element attributes of the element and corresponding element attributes of the target element (e.g., HTML element attributes like ID, name, class, etc. , HTML content attributes like label).
[0210] In some example embodiments, the comparison of features of respective elements 134 of a document object model 120 with the features of an element 134 corresponding to a step 128 may include a first aspect in which various attributes of each element 134 are considered, and optionally weighted. For example, the attributes of the elements 134 in an HTML document may include an identifier, name, caption, object class, visual properties (e.g., font or color), embedded CSS style, inner text and / or inner HTML content, accessibility features such as aria labels, the names and / or identities of associated event handlers (e.g., JavaScript events attached to the elements 134), initial and / or default values, or the like. The weights of the respective attributes may be determined according to their utility as identifying features (e.g., the identifier attribute may be considered the most identifying feature and may have a high weight, while the name attribute may be considered a less reliable identifying feature and may have a lower weight) and / or their rarity (e.g., a common element name such as “btn submit” may be considered to have a low weight, while a distinctive element name such as “btn launch local app on device” or “btn_0x6139c3a2” may be considered to have a high weight. The device 106 may adjust the weights of varying attributes based on their relative and marginal contribution to a successful identification of features of corresponding elements 134. Data and metadata features of the elements 134, as an aspect of the weighted comparison, may be assigned a particular weight that indicates the reliability of the data and metadata features as compared with other aspects and features of the elements 134.
[0211] In some example embodiments, a device 106 may include an element identification system that identifies an element in a current version of a document object model (DOM) that corresponds to a target element in a previous version of the DOM, wherein the identifying is basedon a comparison of a location of the element in the current version of the DOM and a location of the target element in the previous version of the DOM. In some example embodiments, "location" can be specified in various ways: CSS selector, XPath, regular expressions, etc.
[0212] In some example embodiments, the comparison of features of respective elements 134 of a document object model 120 with the features of an element 134 corresponding to a step 128 may include a second aspect in which an identifying location of the element 134 within the document object model 120 is considered. For example, the document object model 120 typically organizes elements 134 in a hierarchical manner, such as HTML DIV elements that include child HTML DIV elements as well as user controls such as HTML buttons and textboxes. Within the hierarchical arrangement, each element 134 is represented by, assigned, and / or reachable by a location in the document object model 120. Such locations may be indicated as a path through the document object model 120, beginning with a document root element and traversing a set of child elements to reach the element 134. The path of each element 134 within the document object model 120 may be specified in various ways, such as a path of a CSS style selector, an XPath expression, a twig query, a regular expression, or the like. The device 106 may compare the locations and / or paths of the elements 134 in a current version of the document object model 120 at the time of visual guidance 132 with the location and / or path of an element 134 in a step 128 of a workflow 126 according to a previous version of the document object model 120 at the time of capture 102 (e.g., based on a Levenshtein distance metric). Location features, as an aspect of the weighted comparison, may be assigned a particular weight that indicates the reliability of location as compared with other aspects and features of the elements 134.
[0213] In some example embodiments, a device 106 may include an element identification system that identifies an element in a current version of a document object model (DOM) that corresponds to a target element in a previous version of the DOM, wherein the identifying is based on one or more relationships between the element and other elements of the DOM in the current version of the DOM and one or more relationships between the target element and other elements in the previous version of the DOM. In some example embodiments, “relationship" can be a direct relationship in metadata (e.g. , one element targeting another element) or a hierarchical relationship (e.g., one element contains another, or is specified in a location relative to another element).
[0214] In some example embodiments, the comparison of features of respective elements 134 of a document object model 120 with the features of an element 134 corresponding to a step 128 may include a third aspect involving relationships between the element 134 and other elements 134 in the document object model 120. For example, in a particular version of the document object model 120, a first element 134 (e.g., a button element) may be arranged as a child element of a second element (e.g., an HTML DIV element); a sibling element of a third element (e.g., anotherbutton element); and a parent element of a fourth element (e.g., an image depicted on the button). While the identifiers, attributes, and location of the element 134 in the document object model 120 may change, the logical hierarchical organization of elements 134 in the document object model 120 may remain the same or similar, and so the element 134 may retain similar relationships with other elements 134 of the document object model 120 through several versions or changes. Thus, the device 106 may consider, as an identifying aspect of an element 134, the set of relationships with other elements 134 of the user interface 114. The device 106 may perform a comparison, may assign scores to the relationships, and may determine that an element 134 in a previous version of the document object model 120 at the time of capture 102 that is included in a step 128 of the workflow 126 correspond to an element 134 in a current version of the document object model 120 at the time of visual guidance 132 based on the similarities of relationships in the hierarchical arrangement of the document object model 120.
[0215] In some example embodiments, a device 106 may include an element identification system that identifies an element in a current version of a web page that corresponds to a target element in a previous version of the web page, wherein the identifying is based on a comparison of geometry of the element in the current version of the web page with a geometry of the target element in the previous version of the web page. In some example embodiments, “geometry” can be absolute position, relative position, size, Z-ordering, proximity to other elements (e.g., labels), etc.
[0216] In some example embodiments, the comparison of features of respective elements 134 of a document object model 120 with the features of an element 134 corresponding to a step 128 may include a fourth aspect involving geometric properties of the element 134 in a rendering of the user interface 114. Such geometric properties may include, e.g., a size, shape, and / or coordinates of an element 134 within a rendered web page 112. The geometric features may be determined according to various metrics (e.g., pixels, point sizes, relative units such as percentages, EM, REM, VW, and / or VH; and / or physical units such as millimeters or inches). The geometric features may include other visual features that relate to visual styles of the elements 134 (e.g., the presence, visual style, and / or thickness of boundaries, margins, containers, or the like). The geometric features may include layout features (e.g., flow layout properties, flexbox organizational features, and / or anchoring or distance with respect to other elements). The geometric features may include X and / or Y dimensions; a Z-order dimension (e.g., prioritized overlapping); a Z dimension in a three-dimensional interface such as an augmented reality and / or virtual reality environment; and / or an area, circumference, surface area, volume, or the like. The geometric features may be indicated in an absolute manner (e.g., 100 pixels), a standardized manner (e.g., a default size), and / or a relative manner (e.g., 10 pixels larger than a referenceelement, such as an embedded image). The device 106 may score the comparison of each geometric property of the element 134 corresponding to a step 128 as represented in the document object model 120 at the time of capture 102 with the corresponding geometric properties of the elements 134 of the document object model 120 at the time of visual guidance 132 to determine which element 134 most closely corresponds to the element 134 of the step 128.
[0217] FIG. 10 illustrates an example scenario featuring a weighted comparison of scores of features of various aspects of an element between versions of a document object model, according to aspects of the present disclosure. The example scenario of FIG. 10 may be performed by the example device 202 of FIG. 2 It is appreciated that while FIG. 10 is shown in connection with providing guiding a user through the performance of a task, the techniques may be applied in the context of knowledge distribution (as discussed elsewhere in the disclosure).
[0218] In the example scenario of FIG. 10, a workflow 126 includes a step 128 that in turn includes an element 134. The workflow 126 includes a record of a set of features of the element 134 in a user interface 114 implemented according to a version of the document object model 120 at the time of capture 102. For example, the workflow 126 may indicate the attributes of various HTML elements in a document object model of a web page, wherein the attributes enable the device 106 to determine the HTML elements that are associated with each action 124 of the workflow 126. In order to determine which element 134 of the document object model 120 at the time of visual guidance 132 corresponds to the record of the element 134 at the time of capture 102, the device 106 may perform a comparison 1002 as shown in the example scenario of FIG. 10. The record of the element 134 includes values for various features 1006 according to different aspects 1004 of the element 134. For example, as a first aspect 1004, the record of the element 134 includes features 1006 based on attributes of the element 134, such as a value of an “id” attribute and a value of a “css class” attribute. As a second aspect 1004, the record of the element 134 includes a feature 1006 based on the location of the element 134 in the document object model 120, such as a hierarchical path to the element 134. As a third aspect 1004, the record of the element 134 includes features 1006 based on relationships of the element 134 with other elements of the user interface 114, such as the name and / or identifier of a parent element 134 and a child element 134. As a fourth aspect 1004, the record of the element 134 includes features 1006 based on the geometry of the element 134 in a rendering of the user interface 114, such as a web page 112 presented on a display 108 of a particular device 106.
[0219] At a time of visual guidance 132, the device 106 may compare the aspects 1004 and features 1006 included in the record of the element 134 with corresponding aspects 1004 and features 1006 of each element 134 in the version of the document object model 120 at the time of visual guidance 132. The example scenario of FIG. 10 shows one such comparison, in whichrespective aspects 1004 and features 1006 of one such element 134 are compared with the corresponding aspects 1004 and features 1006 in the record of the element 134 for the step 128. For each comparison 1008, the device 106 may determine a score 1010 indicating a similarity of the feature 1006 (e.g., based on a standard scale of 1.0, indicating completely identical features 1006, to 0.0, indicating completely dissimilar features 1006). Also, for each comparison 1008, the device may store a weight 1012 that indicates a degree of contribution and / or confidence of the feature 1006 to the overall comparison for identification. The device 106 may determine a product of each score 1010 and each corresponding weight 1012. Additionally or alternatively, the device 106 may determine a total score 1014 based on an aggregation of the products (e.g., as a sum of the products of the respective scores 1010 and corresponding weights 1012). The device 106 may compare the total score 1014 for the element 134 with the total scores 1014 of other elements 134 of the user interface 114 to determine the element 134 that most closely corresponds to the record of the element 134 included in the step 128. The device 106 may complete the target identification by using the determined element 134 as the element 134 for a current step 128 of the workflow 126.
[0220] In some example embodiments, a score-based comparison 1008 of elements 134 according to aspects 1004 and / or features 1006 may be performed in various ways. As a first example, the comparison 1008 may include a particular comparison (e.g., determining whether a name and / or identifier of an element 134 in the current version of the document object model 120 at the time of visual guidance 132 exactly matches the name and / or identifier of the element 134 at the time of capture 102). If so, the device 106 may add a fixed amount to the total score 1014 of the element 134. Conversely, the comparison 1008 may determine whether a feature 1006 of the element 134 at the time of capture 102 is consistent with the feature of the element 134 in the current version of the document object model 120 at the time of visual guidance 132 (e.g., whether the element 134 is responsive to the type of event that is associated with the action 124 and / or step 128). If not, the device 106 may add a penalty to the total score 1014 of the element 134. As a second example, the comparison 1008 may use different weights 1012 for different degrees of stringency of a comparison of a feature 1006. For example, the comparison may include a tight comparison of a name of an element 134 at a time of capture 102 and of an element 134 in the document object model 120 at the time of visual guidance 132 (e.g., a case-sensitive comparison with a small tolerance of Levenshtein distance between the names), and may increase the total score 1014 of the element 134 by a large amount if the tight comparison succeeds. The comparison 1008 may also include a loose comparison of the name of an element 134 at a time of capture 102 and of an element 134 in the document object model 120 at the time of visual guidance 132 (e.g., a case-insensitive comparison with a larger tolerance of Levenshtein distance between the names),and may increase the total score 1014 of the element 134 by a smaller amount if the loose comparison succeeds. As a third example, the comparison 1008 may involve various types of representations of the elements 134 and features 1006 thereof. For example, the comparison 1008 may examine a value of the feature 1006 of a particular element 134 (e.g., a string representation of a name) and / or a hash value of the feature 1006 (e.g., a hash value of an image included in an element 134) as an identifier of the element 134.
[0221] In some example embodiments, a device 106 may include a description of an interactive element of a graphical user interface (GUI), wherein the interactive element is associated with a respective step in a user-directed UX workflow captured from a previous DOM during a previous user-initiated capture session, and the guidance system may update the description to identify a corresponding element in a current version of the GUI based on actions of a current user with the current version of the user interface while performing the step during a current instance of the UX workflow. The updating may include updating the description with more accurate information about the corresponding element.
[0222] In some example embodiments, a device 106 or a workflow server 220 may determine that the document object model 120 for a particular user interface 114 is changing over time. For example, the document object model 120 may be changed to rename elements 134, rearrange a hierarchical structure in a manner that alters paths of certain elements 134, reorganize the elements 134 to include different sets of relationships thereamong, and / or may restyle the user interface 114 in a manner that changes the geometric properties of the elements 134. In addition to performing various target identification processes to determine correspondence between elements 134 in different versions of the document object model 120, the device 106 or workflow server 220 may update and / or annotate the record of the elements 134 in the workflow 126 to indicate the correspondence and / or changes in different versions of the document object model 120. As a first such example, the workflow 126 may include, for an element 134 included in a step 128, a first record of the element 134 in a version of the document object model 120 at a first point in time (e.g., at the time of capture) and a second record of the element 134 in a version of the document object model 120 at a second point in time (e.g., at a time of first guidance). As a second such example, the workflow 126 may include, for an element 134 included in a step 128, a base record of the element 134 in a version of the document object model 120 at the time of capture 102 and one or more differential records corresponding to different times and different versions of the document object model 120, wherein each differential record indicates which features 1006 of the element 134 changed with respect to the base record and / or another version of the document object model 120. The updating and / or annotation of the record of the element 134 may provide a “self- healing” feature, e.g., where changes to the element 134 over the course of evolution of thedocument object model 120 are incrementally accumulated and incorporated in the comparisons to inform target identification for visual guidance 132 based on the aggregated information over the lifespan of the user interface 114.GUIDANCE - ADVANCING TO NEXT STEP OF WORKFLOW
[0223] In some example embodiments, the step 808 of advancing to a next step 128 of a workflow 126 upon detecting a completion of a current step 128 of the workflow 126 may be performed in various ways. For example, the device 106 may monitor the actions 124 included in a current step 128, e.g., by mapping events arising within the user interface 114 to the performance of actions 124. When the device 106 determines that all of the actions 124 included in a current step 128 have been performed (e.g., an event has been raised that corresponds to a completion of each action 124 with a corresponding element 134 of the user interface 114), the device 106 may advance to a next step 128 of the workflow 126.
[0224] In some example embodiments, a device 106 may include an interaction monitoring system that monitors interactions 144 of the current user to determine when the user has completed a current step in the user-defined UX workflow corresponding to a currently highlighted interactive element and instructs the guidance UI to highlight a next interactive element corresponding to a subsequent step in the user-defined UX workflow.
[0225] In some example embodiments, a device 106 may respond to a completion of a current step 128 of a workflow 126, and may advance to a next step 128 of the workflow 126, by removing details and / or a highlight 610 of the previous current step 128 in a side panel 224 and may add details and / or a highlight 610 of the new current step 128 in the side panel 224. The device 106 may remove a highlight 610 of one or more elements 134 of the user interface 114 associated with the previous current step 128 and may add a highlight 610 to each of one or more elements 134 of the user interface 114 associated with the new current step 128. In some example embodiments, the device 106 may remove the previous current step 128 from the side panel 224 (e.g., turning a page of the workflow 126 from the previous current step 128 to the new current step 128) and / or may insert, into the side panel 224, one or more additional steps 128 that follow the new current step 128. In some example embodiments, if the completed current step 128 is the last step in the workflow 126, the device 106 may indicate to the user and / or record a completion of the workflow 126 for the task 116. The device 106 may remove the workflow 126 from the side panel 224 and / or may close the side panel 224 in the web browser 110.KNOWLEDGE DISTRIBUTION
[0226] In some example embodiments, the distribution of knowledge may refer to a function that allows users to initiate and engage in collaboration over the user interface of a user interface 114 of a third-party application / website. This may include the collaboration by sharing ofknowledge relating to a user-defined UX workflow “on top of’ the user interface to which the workflow pertains or sharing information relating to the subject matter of a third party application / website (e.g., “read here for more info on XYZ”). In some example embodiments, the device may be configured to allow users to create, view, and / or add to collaboration elements that are pinned in relation to respective elements of the UI (interactive or non-interactive elements alike).KNOWLEDGE DISTRIBUTION - OVERVIEW
[0227] FIG. 11 illustrates another flowchart of another method of presenting automatically generated user interface documentation according to aspects of the present disclosure. At least a portion of the example method 1100 may be performed, for example, by the example device 202 of FIG. 2.
[0228] In general, the method 1100 of FIG. 11 enables the device to integrate, in a user interface 114, knowledge that relates to the user interface 114. In some example embodiments, the device 106 may display an interactive knowledge indicator 140 (e.g., near and in-relation to an element of the user interface 114 selected by a user) in order to signal the availability of knowledge 142. Activation of the interactive knowledge indicator 140 may cause the device 106 to supplement the user interface 114 with additional information about the element 134, such as a submitted user comment about the element 134 or at least a portion of a conversation within a conversational interface that relates to the element 134. In some example embodiments, the device 106 may show user comments about an element 134 of the user interface 114 within a side panel 224 adjacent to the user interface 114, and / or at least a portion of a conversation within a conversational interface that relates to the element 134. The device 106 may receive user comments from a user 104 relating to an element 134 of the user interface 114 and may store the user comment for later presentation as a knowledge indicator 140, and / or may add the user comment to a conversation about the element 134 within a conversational interface.
[0229] The example method 1100 includes a step 1102 of detecting, by the processor, a presentation of at least a portion of the user interface 114 that is associated with at least one action 124 of the sequence 122 of steps 128. For example, the device 106 may determine that a particular web page 112 within a website includes a user control.
[0230] The example method 1100 includes a step 1104 of presenting, by the processor, a knowledge indicator 140 in the user interface 114. In some embodiments, each action 124 of the sequence 122 of steps 128 includes at least one user comment that is associated with at least one element 134 of the user interface. The device 106 may supplement the presentation of the at least a portion of the user interface 114 by displaying the knowledge indicator 140 adjacent to the at least one element 134 included in the user interface 114.
[0231] The example method 1100 includes a step 1106, responsive to detecting an interaction 144 of the user 104 with the knowledge indicator 140, of presenting, by the processor, knowledge that is associated with the knowledge indicator 140. For example, in some example embodiments, the user interface 114 further comprises a website including at least one web page 112. The device may present knowledge that is associated with the at least a portion of the user interface 114 by displaying the knowledge in response to an interaction 144 of the user 104 with the knowledge indicator 140. For example, at least one action 124 of the sequence 122 of steps 128 may include at least one user comment that is associated with at least one element 134 of the user interface 114. The device 106 may present the at least one user comment that is associated with the at least one element of the user interface 114 in response to the interaction 144 of the user 104 with the knowledge indicator 140.
[0232] In some example embodiments, the knowledge that is associated with the knowledge indicator 140 may include a workflow 126 associated with a task 116. For example, when a user 104 visits a web page 112, the device 106 of the user 104 may determine that the web page is associated with a workflow 126. The device 106 may include a knowledge indicator 140 that indicates an availability of a workflow 126 relating to a task 116 that may be performed using the user interface 114 of the web page 112. When the user 104 interacts with the knowledge indicator 140, the device 106 may initiate visual guidance 132 based on the workflow 126 to guide the user 104 through performing the task 116.
[0233] Alternatively or additionally, in some example embodiments, the knowledge that is associated with the knowledge indicator 140 may include information about the user interface 114, such as comments, questions, tips, or links to other resources that relate to the user interface 114. For example, a first user 104 who visits a web page 112 may provide information, such as creating a comment about a portion of the web page 112. The information of the first user 104 may be stored as knowledge 142 that is associated with the web page 112. When a second user 104 visits the web page 112, a device 106 of the second user 104 may determine and / or receive an indication of an availability of the knowledge 142 relating to the web page 112. The indication of the availability of knowledge 142 may be based on information embedded in the web page 112 provided by the webserver 118 and / or may be based on information received from another server (e.g., the device 106 may send the URL or URI of the web page 112 to a knowledge server, and the knowledge server may transmit an indication that knowledge is available for the web page 112). The device 106 of the second user 104 may insert, into the web page 112, a knowledge indicator 140 that indicates the availability of knowledge about the web page 112. When the second user 104 interacts with the knowledge indicator 140, the device 106 of the second user 104may retrieve and / or receive the knowledge 142 and insert the knowledge 142 into the web page 112.
[0234] In some example embodiments, the knowledge 142 may include an interactive element, such as a chat interface of a chat service. For example, a first user 104 may create a first chat message about the web page 112. The chat interface may be shown on the second device 106 of the second user 104, and may include the first chat message from the first user 104 about the web page 112. The device 106 of the second user 104 may insert the chat interface into the web page 112 when the second user 104 interacts with the knowledge indicator 140. When the second user 104 provides a second chat message in the chat interface, the device 106 of the second user 104 may transmit the second chat message to the chat service, and the second chat message from the second user 104 may be shown in a chat interface included in a presentation of the web page 112 on the first device 106 of the first user 104.
[0235] In some example embodiments, knowledge 142 that is available through interaction with a knowledge indicator 140 in a user interface 114 may include various types of information that relates to the user interface 114. As a first such example, the information may explain a context, process, role, mechanism, and / or feature of a user interface 114 (e.g., explaining features of a user interface 114, or explaining how to use the user interface 114). As a second such example, the information may explain a manner of using a user interface 114 (e.g., kinds of data that may be used with and / or generated by a user interface 114, or a circumstance in which a user interface 114 may be used). As a third such example, the information may include a problem occurring with a user interface 114 (e.g., the existence of errors that may arise while using a user interface 114). As a fourth such example, the information may include a solution to a problem occurring within a user interface 114 (e.g., a manner of operating a user interface 114 to avoid and / or resolve an error). As a fifth such example, the information may include a question or request for further information related to a user interface 114 (e.g., a request to explain a feature, use, and / or purpose of a user interface 114, or a question about using a user interface 114). As a sixth such example, the information may include an opinion, anecdote, experience, review, and / or observations about a user interface 114 (e.g., a personal anecdote or testimony about using a user interface 114). In a seventh such example, the inform...
Claims
CLAIMS1. A method, comprising: recording, by a processor, a workflow during a capture of a task by a user through a user interface, wherein the workflow includes a sequence of steps of the task performed through the user interface; providing, by the processor, a visual guidance to a user to perform, in the user interface, each step of the sequence of steps of the workflow; and presenting, by the processor, at least one knowledge indicator in the user interface, wherein each knowledge indicator indicates an availability of knowledge related to the user interface.
2. A system, comprising: a web browser process that executes within a context of a web browser of a device during a presentation of a user interface of a web page, wherein the web browser process is configured to, record, by a processor of the device, a workflow during a capture of a task by a user through the user interface of the web page, wherein the workflow includes a sequence of steps of the task performed through the user interface of the web page; provide, by the processor of the device, a visual guidance to a user to perform, in the user interface of the web page, each step of the sequence of steps of the workflow; and present, by the processor of the device, at least one knowledge indicator in the user interface of the web page, wherein each knowledge indicator indicates an availability of knowledge related to the user interface of the web page.
3. A method, comprising: detecting, by a processor, a sequence of interactions by a user while performing a task within a user interface; identifying, by a processor, at least one element of the user interface that is associated with each interaction of the sequence of interactions; determining, by the processor, a workflow of the task based on the sequence of interactions performed by the user within the user interface; wherein the workflow includes a sequence of at least one step, and each step includes at least one action; and storing, by the processor, a record of the workflow of the task, wherein the record of the workflow is usable to inform users to perform the task within the user interface.
4. The method of claim 3, wherein, the user interface further comprises a website including at least one web page,each interaction by the user is associated with at least one content item of the at least one web page, and each content item includes at least one of, at least one content element included in at least one web page of the website, or at least one user control included in at least one web page of the website.
5. The method of claim 3, wherein identifying the at least one element of the user interface that is associated with each interaction of the sequence of interactions includes, associating, by the processor, at least one event arising within the user interface with at least one event handler, detecting, by the processor, an invocation of an event handler in response to at least one event caused by at least one interaction of the user with the user interface, and determining, by the processor, at least one event detail associated with the invocation of the at least one event, wherein each interaction based on the at least one event detail associated with the invocation of the at least one event.
6. The method of claim 3, wherein identifying the at least one element of the user interface that is associated with each interaction includes, identifying, by the processor, at least one user control within the user interface that is associated with an interaction, and recording, by the processor, at least one attribute of the at least one user control to identify the at least one user control within the user interface that is associated with the interaction.
7. The method of claim 3, wherein identifying the at least one element of the user interface that is associated with each interaction includes, identifying, by the processor, a first content item within the user interface that is associated with an interaction, wherein the first content item is not responsive to the interaction, determining, by the processor, a second content item within the user interface that is associated with the first content item, wherein the second content item is responsive to the interaction, and recording, by the processor, at least one attribute of the second content item to identify at least one content item within the user interface that is associated with the interaction.
8. The method of claim 3, wherein identifying the at least one element of the user interface that is associated with an interaction includes, responsive to detecting an interaction by the user while performing the task, identifying, by the processor, at least one user control within the user interface that is associated with the interaction,capturing, by the processor, a screenshot of the at least one user control, and recording, by the processor, the screenshot of the at least one user control within the user interface that is associated with the interaction.
9. The method of claim 8, wherein capturing the screenshot of the at least one user control includes, intercepting, by the processor, at least one event arising within the user interface with at least one event handler, wherein the at least one event is associated with the interaction by the user, capturing, by the processor, the screenshot of the at least one user control while intercepting the at least one event associated with the interaction, and after capturing the screenshot, re-raising, by the processor, the at least one event arising within the user interface.
10. The method of claim 8, wherein identifying the at least one element of the user interface that is associated with each interaction includes, displaying, by the processor, a visual highlight associated with the at least one user control, and capturing the screenshot of the at least one user control includes removing, by the processor, the visual highlight associated with the at least one user control while capturing the screenshot of the at least one user control.
11. The method of claim 3, wherein identifying the at least one element of the user interface that is associated with each interaction includes, presenting, by the processor, a side panel adjacent to the user interface, responsive to detecting an interaction by the user while performing the task, presenting, by the processor, a description of the interaction in the side panel, receiving, by the processor, at least one information item provided by the user as input to the description of the interaction in the side panel, wherein the at least one information item further describes the interaction, and recording, by the processor, the at least one information item provided by the user for the interaction.
12. The method of claim 3, wherein, the user interface is based on a document object model, and identifying the at least one element of the user interface that is associated with each interaction includes,identifying, by the processor, a portion of the document object model that describes at least one user control within the user interface that is associated with each interaction, and recording, by the processor, at least one attribute of the portion of the document object model, wherein the at least one attribute identifies at least one element within the document object model that defines the at least one user control.
13. The method of claim 12, wherein the at least one attribute that identifies at least one element of the document object model includes at least one of, a location of the at least one element within a schema of the document object model, at least one attribute of the at least one element defined by the document object model, or at least one attribute of at least one user control that is associated with the at least one element of the document object model.
14. The method of claim 3, wherein generating the workflow for the task includes, determining, by the processor, at least two consecutive interactions of the sequence of interactions by the user while performing the task within the user interface, wherein at least one interaction of the at least two consecutive interactions is redundant with at least one other interaction of the at least two consecutive interactions, and removing, by the processor, the at least one interaction from the sequence of interactions.
15. The method of claim 3, wherein generating the workflow for the task includes, determining, by the processor, at least two consecutive interactions of the sequence of interactions by the user while performing the task within the user interface, wherein at least one interaction of the at least two consecutive interactions renders moot at least one other interaction of the at least two consecutive interactions, and removing, by the processor, the at least two consecutive interactions from the sequence of interactions.
16. The method of claim 3, wherein generating the workflow for the task includes, determining, by the processor, at least two consecutive interactions of the sequence of interactions by the user while performing the task within the user interface, wherein the at least two consecutive interactions are associated with a user control within the user interface, and substituting, by the processor, an aggregated interaction in the sequence of interactions for the at least two consecutive interactions associated with the user control within the user control.
17. The method of claim 3, wherein generating the workflow for the task includes,determining, by the processor, at least one first interaction of the sequence of interactions by the user while performing the task within the user interface, wherein the at least one first interaction is associated with a first user control within the user interface, determining, by the processor, at least one second interaction of the sequence of interactions by the user while performing the task within the user interface, wherein the at least one second interaction is associated with a second user control within the user interface, and substituting, by the processor, an aggregated interaction in the sequence of interactions for the at least one first interaction and the at least one second interaction, wherein the aggregated interaction is associated with both the first user control and the second user control.
18. A method, comprising: receiving, by a processor, a workflow for a task, wherein the workflow includes a sequence of steps to be performed by a user within a user interface; presenting, by the processor, the sequence of steps of the workflow to the user during a performance of the task by the user; presenting, by the processor, a current step of the workflow, wherein presenting the current step guides the user to perform at least one interaction with at least one element of the user interface, and the at least one interaction corresponds to at least one action of the current step; and advancing, by the processor, to a next step of the workflow upon detecting a completion by the user of the current step of the workflow.
19. The method of claim 18, wherein, the user interface further comprises a website including at least one web page, respective steps in the sequence of steps are associated with at least one content item, each content item includes at least one of, at least one content element included in at least one web page of the website, or at least one user control included in at least one web page of the website, and presenting the sequence of steps includes displaying, by the processor, the sequence of steps of the workflow in a side panel adjacent to the user interface.
20. The method of claim 18, wherein, the user interface is based on a document object model, and presenting the current step of the workflow includes, identifying, by the processor, a portion of the document object model that matches a description of a current step included in the workflow, and displaying, by the processor, a visual highlight of at least a portion of the user interface indicated by the portion of the document object model.
21. The method of claim 20, wherein, the description of the current step includes at least one attribute of at least one element of the document object model, the at least one element is associated with a first interaction by a user while performing the task within the user interface, and the first interaction corresponds to the current step of the workflow, and identifying the portion of the document object model that matches the description of the current step included in the workflow includes determining the portion of the document object model that matches the at least one attribute of the at least one element of the document object model.
22. The method of claim 21, wherein the at least one attribute of the at least one element of the document object model includes at least one of, a location of the at least one element within a schema of the document object model, at least one attribute of the at least one element defined by the document object model, or at least one attribute of at least one user control that is associated with the at least one element of the document object model.
23. The method of claim 21, wherein, the description of the current step includes at least two attributes of at least one element of the document object model, and identifying the portion of the document object model that matches the description of the current step included in the workflow includes, performing, by the processor, a comparison of each attribute of the at least two attributes with at least one element of the document object model, determining, a score for each element of the document object model for each attribute of the at least two attributes, wherein the score is based on the comparison, and identifying, by the processor, the portion of the document object model that matches the description of the current step included in the workflow based on the portion of the document object model including an element having a highest sum of scores for each attribute of the at least two attributes.
24. The method of claim 23, wherein the comparison of each attribute of the at least two attributes with at least one element of the document object model includes at least one of, comparing, by the processor, a label associated with at least one element of the document object model during a capture of the task to generate the workflow with a label associated with at least one element of the document object model,comparing, by the processor, at least one attribute included in at least one element of the document object model during a capture of the task to generate the workflow with at least one attribute included in at least one element of the document object model during the current step of the performance of the task by the user, comparing, by the processor, a geometry of at least one element of the document object model during a capture of the task to generate the workflow within the user interface with a geometry of at least one element of the document object model during the current step of the performance of the task by the user, comparing, by the processor, a style selector associated with at least one element of the document object model during a capture of the task to generate the workflow with a style selector associated with at least one element of the document object model during the current step of the performance of the task by the user, or comparing, by the processor, a related element that is associated with at least one element of the document object model during a capture of the task to generate the workflow with a related element that is associated with at least one element of the document object model during the current step of the performance of the task by the user.
25. The method of claim 23, wherein, each attribute of the at least two attributes is associated with a weight, and the score for each element of the document object model for each attribute of the at least two attributes is further based on the weight associated with each attribute.
26. The method of claim 21, further comprising: responsive to failing to identify a portion of the document object model that matches the description of the current step included in the workflow, presenting, by the processor, an alternative instruction associated with the current step.
27. The method of claim 21, further comprising: responsive to failing to identify a portion of the document object model that matches the description of the current step included in the workflow, determining, by the processor, an alternative description for the current step, wherein the alternative description identifies at least one alternative content item within the user interface that is associated with the current step, and substituting, by the processor, the description for the current step within the workflow with the alternative description for the current step.
28. The method of claim 18, wherein presenting the current step of the workflow includes,identifying, by the processor, a first content item within the user interface that is associated with the current step, wherein the first content item is not responsive to the current step, determining, by the processor, a second content item within the user interface that is associated with the first content item, wherein the second content item is responsive to the current step, and displaying a visual highlight of the second content item within the user interface.
29. The method of claim 18, further comprising: determining, by the processor, a completion of the current step of the performance of the workflow by a user; responsive to the completion of the current step, removing, by the processor, a visual highlight of at least one content item within the user interface that is associated with the current step; and advancing, by the processor, the current step of the performance of the workflow to a next step of the performance of the workflow.
30. The method of claim 18, wherein, the user interface further comprises a website including at least one web page, each content item includes at least one of, at least one content element included in at least one web page of the website, or at least one user control included in at least one web page of the website, and presenting the sequence of steps of the workflow includes displaying, by the processor, the sequence of steps of the workflow in a side panel adjacent to the user interface.
31. A device configured to automatically generate documentation for a user interface, the device comprising: a processor; and a memory storing instructions that, when executed by the processor, cause the device to, determine a task to be performed within the user interface, determine a sequence of interactions by a user while performing the task within the user interface, record a set of details for each interaction of the sequence of interactions, wherein the set of details identifies at least one content item within the user interface that is associated with each interaction, generate a workflow for the task, wherein the workflow includes a sequence of descriptions generated by the processor for respective actions of the workflow, and each description for each action is based on the set of details associated with each interaction, andresponsive to a request to describe the workflow for the task, present the sequence of descriptions generated by the processor.
32. The device of claim 31, wherein, the user interface further comprises a website including at least one web page, the sequence of interactions by the user further comprises a sequence of interactions respectively associated with at least one content item, each content item includes at least one of, at least one content element included in at least one web page of the website, or at least one user control included in at least one web page of the website, and execution of instructions by the processor further causes the device to present the sequence of descriptions in a side panel adjacent to the user interface.
33. A system for automatically generating documentation for a user interface, the system comprising: a content process configured to, determine a task to be performed within the user interface, determine a sequence of interactions by a user while performing the task within the user interface, record a set of details for each action of the sequence of interactions, wherein the set of details for each action identify at least one content item within the user interface that is associated with each interaction, and generate a workflow for the task, wherein the workflow includes a sequence of descriptions generated by the content process for each action of the workflow, and each description for each action is based on the set of details associated with each interaction; and a display process configured to, responsive to a request to describe the workflow for the task, present the sequence of descriptions generated by the content process.
34. The system of claim 33, wherein, the user interface further comprises a website including at least one web page, the sequence of interactions by the user further comprises a sequence of interactions respectively associated with at least one content item, each content item includes at least one of, at least one content element included in at least one web page of the website, or at least one user control included in at least one web page of the website, and the display process is further configured to present the sequence of descriptions generated by the content process in a side panel adjacent to the user interface.
35. The system of claim 33, further comprising:a background process configured to, transmit the workflow to a web application, and responsive to the request to describe the workflow for the task, retrieve the workflow from the web application.
36. The system of claim 33, further comprising: a web application configured to, receive the workflow from the content process, store the workflow, and display the workflow in association with the user interface.
37. The system of claim 36, wherein the display process is further configured to present the sequence of descriptions generated by the content process responsive to a selection within the web application of the workflow displayed in association with the user interface.
38. A method, comprising: detecting, by a processor, a presentation of at least a portion of a user interface; presenting, by the processor, a knowledge indicator in the user interface; and responsive to detecting an interaction with the knowledge indicator, presenting, by the processor, knowledge associated with the knowledge indicator.
39. The method of claim 38, wherein, the knowledge indicator includes at least one user comment that is associated with at least one content item included in the user interface, wherein the at least one content item is associated with least one interaction of a sequence of interactions with the user interface, and presenting the knowledge indicator includes displaying, by the processor, the knowledge indicator adjacent to the at least one content item included in the user interface.
40. The method of claim 38, wherein, the knowledge indicator includes at least one user comment that is associated with at least one content item included in the user interface, wherein the at least one content item is associated with at least one interaction of a sequence of interactions with the user interface, and presenting the knowledge includes displaying, by the processor, the at least one user comment associated with the at least one content item.
41. The method of claim 38, further comprising: responsive to receiving, from a user, at least one user comment associated with at least one content item included in the user interface, wherein the at least one content item is associated with at least one interaction of a sequence of interactions with the user interface; storing, by the processor, the at least one user comment in a description for each interaction of the sequence of interactions; andassociating, by the processor, the at least one user comment with the at least one content item included in the user interface.
42. The method of claim 38, wherein, at least one content item included in the user interface is associated with at least one conversation in a conversational interface associated with the user interface, and presenting the knowledge includes displaying, by the processor, the at least one conversation in the conversational interface associated with the user interface.
43. The method of claim 38, further comprising: responsive to receiving, from a user, at least one user comment associated with at least one content item included in the user interface, wherein the at least one content item is associated with at least one interaction of a sequence of interactions with the user interface, adding, by the processor, the at least one user comment to a conversation in a conversational interface associated with the user interface.
Citation Information
Patent Citations
Application training simulation system and methods
US20040046792A1
analytics
US20150134694A1
System and method for real-time collaboration
US20200159950A1
System for providing intelligent part of speech processing of complex natural language
US20200226324A1
Automatic data transfer between a source and a target using semantic artificial intelligence for robotic process automation
US20230107316A1