Automatic Flow Implementation from Text Input
A machine learning-based system generates computerized workflows from natural language inputs, addressing the complexity of graphical interfaces by predicting and converting workflow steps into API calls, enhancing efficiency and adherence to best practices.
Patent Information
- Application Number
- JP2023083598
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-05-24
- Filing Date
- 2023-05-22
- Publication Date
- 2025-07-24
- Estimated Expiration
- 2043-05-22
AI Technical Summary
Developers face challenges in creating computer workflows due to the complexity of graphical user interfaces, which require significant learning and may not promote best practices, especially for novice users.
A system that uses machine learning to generate computerized workflows from natural language descriptions, leveraging a text-to-text model to predict workflow steps and convert them into API calls, reducing the need for manual intervention and promoting best practices.
Enables efficient and adaptable workflow generation, reducing the user learning curve and ensuring adherence to best practices, even for novice users, while minimizing computational and training costs.
Smart Images

Figure 0007712978000001 
Figure 0007712978000002 
Figure 0007712978000003
Abstract
Description
Background Art
[0001] Machine-assisted development of computer instructions enables developers to generate executable sequences of computer actions without requiring significant knowledge of computer languages. Computer instructions can be in the form of an automated process, a computer program, or a set of instructions that tell a computer how to operate. To develop computer instructions, for example, a developer can interact with the graphical user interface of a development tool. However, sometimes, a developer can be faced with the difficulty of learning how to use the development tool. The developer may be overwhelmed by the many options within the development tool and may not use best practices. Therefore, there is a need for a technology that supports developers in this regard.
Brief Description of the Drawings
[0002] In the following detailed description and the accompanying drawings, various embodiments of the present invention are disclosed.
[0003]
Figure 1
[0004]
Figure 2
[0005]
Figure 3
[0006]
Figure 4
[0007]
Figure 5A
Figure 5B
Figure 5C
[0008]
Figure 6
[0009]
Figure 7
DETAILED DESCRIPTION OF THE INVENTION
[0010] The present invention can be implemented in various forms, including a process, an apparatus, a system, a composition of matter, a computer program product embodied on a computer-readable storage medium, and / or a processor (a processor configured to execute instructions stored in and / or provided by a memory connected to the processor). In this specification, these embodiments or any other form that the present invention can take may be referred to as a technique. Generally, the order of the steps of the disclosed processes may be changed within the scope of the present invention. Unless otherwise specified, components such as processors or memories described as being configured to perform tasks may be implemented as general components temporarily configured to perform tasks at a certain time or as specific components manufactured to perform tasks. As used in this specification, the term "processor" refers to one or more devices, circuits, and / or processing cores configured to process data such as computer program instructions.
[0011] Hereinafter, while referring to the drawings showing the principles of the present invention, a detailed description of one or more embodiments of the present invention will be given. The present invention is described in relation to such embodiments, but is not limited to any embodiment. The scope of the present invention is limited only by the claims, and the present invention includes many alternatives, modifications, and equivalents. In the following description, many specific details are set forth in order to provide a complete understanding of the present invention. These details are for illustrative purposes only, and the present invention can be practiced according to the claims even without some or all of these specific details. For the sake of simplicity, technical matters well-known in the technical field related to the present invention are not described in detail so that the present invention is not made more difficult to understand than necessary.
[0012] An automatic flow implementation from text input is disclosed. A user-provided text description of at least a part of a desired workflow is received. Context information associated with the desired workflow is determined. Machine learning input based at least in part on the text description and the context information is provided to a machine learning model to determine an implementation prediction for the desired workflow. One or more processors are used to automatically implement the implementation prediction as at least a part of a computerized workflow implementation of the desired workflow.
[0013] Many low-code environments rely on a graphical user interface, which hides the executable code associated with the workflow. As used herein, a workflow may also be referred to as a "computerized workflow", "computerized flow", "automated flow", "action flow", "flow", etc., and is an automated process (e.g., executed by a programmed computer system) composed of a sequence of actions. The sequence of actions may also be referred to as a sequence of steps, a sequence of action steps, etc. Often, a workflow also includes triggers for the sequence of actions. Examples of flows are shown in FIGS. 2, 3, and 5C. The techniques disclosed herein solve the technical problem of enabling efficient computer workflow generation in scenarios where the use of a graphical user interface is difficult. The techniques disclosed herein enable a user to instantiate workflow steps using free-form natural language, and the steps can be displayed as visual steps within a graphical user interface. In other words, an automated flow can be generated from a natural language description. The use of a graphical user interface can be difficult, especially for novice users, because there may be hundreds of available steps. The techniques disclosed herein have many advantages, such as reducing the user learning curve, promoting best practices for novice users, and helping experienced users discover new features.
[0014] In various embodiments, the techniques disclosed herein lead to a highly adaptable system that requires only a few labeled samples by leveraging large pre-trained language models, and as a result, is robust to variations in user input compared to rule-based techniques. The techniques disclosed herein are widely applicable to different types of automated flow builder applications. In various embodiments, as further detailed herein, a trained machine learning model receives a natural language description of a flow and then predicts all of the actions for the flow in the appropriate order. Since the machine learning model has learned to perform this task from training examples, manual processing and feature engineering are not required. Further, since user input can be taken as is, no preprocessing is required. In various embodiments, the output of the machine learning model is converted into application programming interface (API) calls that are sent to the flow builder application. These techniques are described in further detail below.
[0015] FIG. 1 is a block diagram showing one embodiment of a system for automatically implementing a computerized flow based on text input. In the example of the figure, the text-to-flow unit 100 includes an input aggregator 102, a flow-to-text converter 104, a context-to-text converter 106, an embedding selector 108, a text-to-text model 110, and a text-to-API (text-to-API) converter 112. In the example of the figure, the flow builder application 114 is separate from the text-to-flow unit 100, but in other embodiments, the flow builder application 114 can also be incorporated as a component of the text-to-flow unit 100. In some embodiments, the flow-to-text unit 100 (including its components) is composed of computer program instructions executed by a general-purpose processor (e.g., a central processing unit (CPU)) of a programmed computer system. FIG. 7 shows an example of a programmed computer system. Also, the logic of the flow-to-text unit 100 can also be executed on other hardware (e.g., executed using an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA)). In the example of the figure, the text-to-flow unit 100 provides an output to the flow builder application 114. In various embodiments, the flow builder application 114 includes software that can be interfaced via an API. In the example of the figure, the input data to the text-to-flow unit 100 is a flow description 120, a builder current state 122, a context 124, and a flow embedding 126.
[0016] In various embodiments, the flow description 120 is a required text input that describes an overall flow, a sub - flow, or a single step (e.g., a single action within a flow). The flow description 120 includes step descriptions that can be known descriptions or any other semantically equivalent descriptions (e.g., "create a record" and "add a record to a table" both produce the same flow step). In the example of the figure, the flow description 120 is received by the input aggregator 102. In various embodiments, the input aggregator 102 creates input text for the text - to - text model 110. In various embodiments, the input aggregator 102 does not modify the flow description 120. In various embodiments, since the flow description 120 is required, the input aggregator 102 checks to ensure that the flow description 120 is a non - empty string, but other inputs to the text - to - flow unit 100 are not required. In the example of the figure, the input aggregator 102 further receives text inputs from the flow - to - text converter 104 and the context - to - text converter 106. In various embodiments, the input aggregator 102 determines a flow description based on the flow description 120 and the output of the flow - to - text converter 104 and combines this with context information that is the text output of the context - to - text converter 106. In some embodiments, there is a specific order in which the information is combined because starting with elements that are more strongly influenced by the output of the text - to - text model 110 can lead to better results.
[0017] The technology disclosed in this specification is applicable to other information media such as audio or video. In other words, the flow description may be provided in another format (e.g., audio, video, etc.). For embodiments in which an audio or video flow description is received, the text-to-flow unit 100 may comprise a media-to-text converter that receives the audio and / or video and converts the audio and / or video to text. For example, to convert audio to text, any of a variety of speech recognition technologies well known to those skilled in the art may be used to generate a text format of the audio input (e.g., in the same format as the flow description 120). The text-to-flow unit 100 may then utilize the text format in the same manner as described for the flow description 120. Similarly, video-to-text technologies well known to those skilled in the art may be used to generate a text format from the video input.
[0018] In the example of the figure, the builder current state 122 is an optional input used when predicting a partial flow. Predicting the partial flow adds steps to the incomplete flow by specifying either a single step or multiple steps. The case of using a single step can enable the technology disclosed herein to function within a system such as a chatbot that provides an interaction for the user to create a flow step by step. In various embodiments, the builder current state 122 includes the following two items. 1) Existing steps: steps already created by the user in the builder using either a user interface or a previous call to the text-to-flow unit 100, and 2) Current position: the position where the user requested the generation of the flow (in other words, the index within the list of existing steps). In the example of the figure, the builder current state 122 is received by the flow-to-text converter 104. The flow-to-text converter 104 converts the existing flow and the current position into text format. In some embodiments, the existing steps in the builder current state 122 are in either Extensible Markup Language (XML) or JavaScript Object Notation (JSON) format. Thus, in some embodiments, the flow-to-text converter 104 converts XML or JSON data into text. The builder current state 122 can be in any known data format (e.g., XML, JSON, etc.). Examples of XML and JSON are merely illustrative and not limiting.
[0019] Assume that the existing steps are "Send Email" (where "Send Email" is the step name) and "Create Incident Record" (where "Create Record" is the step name and "Incident" is a step parameter representing the table name). The current position is the 3rd position, indicating the insertion point. In various embodiments, the flow-to-text converter 104 first serializes the existing steps by converting each step from a name to a description using a one-to-one mapping and extracting any existing step parameters. The output generated may be in the following format: "Existing Steps: Step 1 [Parameter 1], Step 2,..., Step N [Parameter N] Current Position: X". In this format, "Existing Steps" and "Current Position" are prefixes that distinguish the existing steps from the current position. This is necessary because the output of the flow-to-text converter 104 is in text format. In various embodiments, when predicting a partial flow, the text-to-text model 110 adjusts the output using the existing steps, which is useful for steps that are affected by previously created steps. The text-to-text model 110 can verify the likelihood of a particular step occurring, assuming the previous steps learned during the training phase of the text-to-text model 110. In a scenario where the user has not created any steps before calling the text-to-flow unit 100, the builder current state 122 has no meaningful information and the output of the flow-to-text converter 104 is an empty string.
[0020] In the example of the figure, context 124 is another optional input. Context 124 may be used to adjust the output of the text-to-text model 110. This adjustment aims to improve the prediction of the text-to-text model 110 based on external factors other than the flow description provided by the user. Regarding context 124, the text-to-text model 110 can adjust its output using specific available context items. For example, depending on the flow creator or business unit, the processing of some cases, such as error handling, logging, or approval management, may vary. For example, when the flow fails, a user or group of users may send an email to the administrator, while other users record the error message and end the flow. Context 124 affects the text-to-text model 110 through the patterns of the training data used to train the text-to-text model 110. In some embodiments, context 124 includes the following items: 1) Application metadata: application properties (application name, business unit, creator, etc.); 2) Flow metadata: flow properties (title, creation date, creator, etc.); 3) Preferences (e.g., enabling one-to-one prediction (as detailed later), setting a list of steps to use (as detailed later regarding out-of-domain step descriptions), embedding an ID (as detailed later regarding flow embedding)). Context 124 may be in XML, JSON, or other well-known data formats. The above may be regarded as adjustment parameters of the text-to-text model 110.
[0021] In various implementations, the context-to-text converter 106 receives a context 124 (e.g., in XML or JSON format) and encodes all elements of the context except the flow embedding ID into text format. In various embodiments, all elements of the context are represented as a list of key / value pairs, which means that the text output of the context-to-text converter 106 can be formatted as follows: "Preferences: key1[value1],..., keyn[valuen] App metadata: key1[value1],..., keyn[valuen] Flow metadata: key1[value1],..., keyn[valuen]". "Preferences", "App metadata", and "Flow metadata" are prefixes that distinguish each part of the serialized text. The context-to-text converter 106 does not need to make all context items available. For example, if the flow metadata is missing, it means that the flow-to-text converter 106 serializes only the other available items. In a scenario where the entire context is not available, the context-to-text converter 106 outputs an empty string.
[0022] In the example of the figure, the flow embedding 126 is an additional optional input. The flow embedding 126 includes a list of previously learned flow embeddings. In various embodiments, the flow embedding 126 is based on an existing flow and includes embeddings that can be used to adjust the text-to-text model 110. Such adjustment can adjust the output of the text-to-text model 110 to resemble a previously created flow. In some embodiments, the flow embedding 126 is a list of fixed-size tensors that are individually learned during the training of the text-to-text model 110. Each flow embedding can be associated with a single dataset and may be stored on disk or in memory. The embeddings can be regarded as a way to describe the difference in training between one dataset and another, or as a way to factorize model weights. For example, two different datasets may be used during the training phase to train a single machine learning model that has a single set of model weights and two different embeddings for each of the datasets. Then, during the deployment of the machine learning model in inference mode, the embeddings can be swapped to match each training dataset without interrupting the machine learning model. Similar results can be achieved by training two different models without using embeddings, but using embeddings reduces the computational and other costs because only a single model needs to be created, deployed, and maintained. Nevertheless, since the embedding function is optional, it is also possible to train two different models and use the techniques disclosed herein. This can be useful in scenarios where it is necessary to separate datasets (e.g., for reasons of confidentiality). In the example of the figure, the embedding selector 108 selects and loads an embedding (e.g., a tensor) from the flow embedding 126. This can be achieved by indicating the selection using the flow embedding ID of the context 124. In scenarios where no flow embedding ID is provided, the embedding selector 108 outputs a NULL tensor.In some embodiments, the selected embedding tensor is loaded into a processor (e.g., a central processing unit (CPU), a graphics processing unit (GPU), etc.) that implements the text-to-text model 110 when the embedding is selected.
[0023] In various embodiments, the text-to-text model 110 predicts an overall flow or a partial flow based at least in part on the flow description 120. The text-to-text model 110 may also utilize inputs other than the user-provided flow description depending on the builder current state 122 and the context 124. Specifically, the prediction of the text-to-text model 110 may be adjusted by inputs such as context parameters, existing steps, and / or flow embeddings. In some embodiments, the flow is predicted by the text-to-text model 110 in a text format such as the following: "Step description 1, Step description 2 [parameter 1], ···, Step description N [parameter 1, ···, parameter M]". Depending on the user input, the output may include a single parameter, multiple parameters, or zero parameters. The architecture of the text-to-text model 110 may be based on various machine learning architectures configured to perform end-to-end learning of semantic mapping from input to output, such as transformers and recurrent neural networks (RNNs) (large language models (LLMs)). The text-to-text model 110 is trained on text examples and is configured to receive text inputs and generate text outputs. In various embodiments, the text-to-text model 110 is trained using transfer learning. Transfer learning means first pre-training the model on a data-rich task and then fine-tuning the model on a downstream task. In some embodiments, the text-to-text model 110 is an LLM with an encoder-decoder architecture. In various embodiments, the text-to-text model 110 is pre-trained with a multi-task mixture of unsupervised and supervised tasks where each task is converted into a text-to-text format. An example of an LLM model with an encoder-decoder architecture is the T5 model.
[0024] The text-to-text model 110 may be configured to predict more than what the user requested within the flow description. For example, even if the user only requests the "request approval" step, the text-to-text model 110 can predict the pattern of an "if" step following the "request approval" step because this pattern frequently occurs in the training data. The purpose of this feature is to help the user follow best practices and shorten the length of the flow description. This feature may be beneficial to novice users, but experienced users may find it distracting. Thus, in various embodiments, this feature may be disabled by enabling one-to-one prediction (the text-to-text model 110 is configured to accurately predict what the user is describing) as a preference within the context 124 (as described above).
[0025] Also, the context 124 may be used to control out-of-domain step descriptions by the text-to-text model 110. To ensure that the text-to-text model 110 does not output any out-of-domain steps at all, the user can provide a list of possible steps via the context 124. If there are no steps in the provided list, the text-to-text model 110 does not predict that step in this mode. For example, the model output for "Send an email, purchase milk, create an incident record" will exclude the "Purchase milk" step if the "Purchase milk" step is not in the list of possible steps, resulting in "Send an email, create a table record [incident]". On the other hand, if the user has not provided a list of possible steps, the model output for this example will be "Send an email, purchase milk, create a table record [incident]". However, since the text-to-API converter 112 may not have an API call mapping for this step, it may remove and record the "Purchase milk" step. In this position, the system administrator analyzes the user request using the recorded information to determine what the text-to-flow unit 100 should be configured to handle. For example, if the text-to-flow unit 100 only handles "Send an email" for communicating with individuals, but the user is trying to send information via the Short Message Service (SMS), the system administrator may use this information to configure the "Send SMS" step. Also, the user can provide a new step (a step not seen in the training data) to the list of possible steps. Thus, instead of generating a new step, the text-to-text model 110 can utilize the newly added step that matches the flow description. For example, assume the flow description includes "Communicate via SMS". If the user adds the "Send SMS" step to the list of possible steps, the text-to-text model 110 will output "Send SMS".Alternatively, in the absence of a user-provided step, the text-to-text model 110 can predict "communicate via SMS" on its own (generate this step). This feature eliminates the need to retrain the text-to-text model 110 and reduces the computational cost for training the text-to-text model 110. In various embodiments, the user can also manually edit the results inaccurately output by the text-to-text model 110 from the user interface of the flow builder.
[0026] In some embodiments, in addition to predicting the flow steps, the text-to-text model 110 also extracts the slot values of the steps. For example, in the input "Create an incident record", the slot value is "incident" and represents the table name where the record is created. FIG. 5C shows another example of a slot value representing a table. The text-to-text model 110 does not require the user to provide the exact name of the value having an identifier name (e.g., table name, application name, etc.). For example, the text-to-text model 110 may predict the same table name for "Create an incident record" and "Create a case record" if the "case" table does not exist (extrapolating from the incident record to the case record). Thus, due to the architecture and training of the text-to-text model 110, the variations in the ways the user can describe the slot values are automatically processed. For example, from the perspective of the model, "Create an incident record" is similar to "Create a record in the incident table" due to natural language similarity. Variations in the slot value format of date and time values can be processed because they are seen during training. For example, in various embodiments, the text-to-text model 110 is trained to extract the same time value from "8pm", "20h", and "8p.m.". Also, it is possible to obtain correct results without seeing all possible combinations during training (e.g., since similar typos are seen during training, the typos can be corrected without encountering a pattern with a typo). By leveraging pre-training on natural language, the amount of training data required to achieve reliable performance can be reduced.
[0027] With the text-to-text architecture utilized, the text-to-text model 110 can adapt to a new set of steps without making any changes to the model. This reduces the required modeling and experimental effort. Further, this feature is extremely important from the perspective of use cases as the user has a variety of effective steps that can be rapidly deployed. Thus, using the techniques disclosed herein, the cost is reduced as creating a new model for each user can be avoided. As described above, the text-to-text model 110 can handle variations in the way the user describes the flow, which is more powerful than existing match-based systems. For example, the text-to-text model 110 can determine whether "lookup record" and "search for an entry in the table" can refer to the same thing depending on the use case. In contrast, it can be very difficult to handle this with either a rule-based system or a classical natural language processing (NLP) model. Another advantage is that the text-to-text model 110 can understand the positionalities of the steps and the composition of the flow. In the flow example disclosed herein, the flow starts with a trigger (e.g., "when an email is received"). However, the user does not have to start the flow description with a trigger description. For example, the text-to-text model 110 generates the same flow for "create an incident record when an email is received" and "when an email is received, create an incident record". This advantage becomes very important as the number of steps increases (e.g., for a 10-step flow length). The text-to-text model 110 is pre-trained without a teacher on a large-scale language dataset before being fine-tuned for text-to-automation flow tasks, so it does not need to look at all possible combinations or ways to describe the same thing semantically, which reduces the data requirements necessary for fine-tuning the text-to-text model 110.
[0028] From a model deployment perspective, the use of a text-to-text architecture reduces the effort required to deploy a new model because there is no need to repeat performance testing or hardware compatibility testing. In contrast, other machine learning models that perform classification using a single classification layer with a fixed number of output categories need to be modified when the number of classification classes changes, which can push the model's response time beyond the upper limit required by the application. Further, when the number of classes increases exponentially, traditional models can no longer be adapted to the available hardware. In traditional machine learning implementations, it may also be necessary to add new classification components for each task in a multi-task setting. In contrast, in the techniques disclosed herein, only a single configuration needs to be modified.
[0029] In the example of the figure, the text-to-API converter 112 converts the text output of the text-to-text model 110 into the API format. In some embodiments, the text-to-API converter 112 uses a defined one-to-one mapping from step descriptions to API calls. In various embodiments, if the output of the text-to-text model 110 includes out-of-domain step descriptions, the text-to-API converter 112 removes the out-of-domain steps and records events for the monitoring system. Thus, the text-to-API converter 112 only executes API calls for valid steps.
[0030] In the example of the figure, the text-to-flow unit 100 does not create the final flow. Instead, it outputs an API call to the flow builder application 114 to complete the conversion of the predicted flow of the text format to the actual flow. The flow builder application 114 is instructed to create a flow via the API call. An example of a flow builder application is the ServiceNow® Flow Designer Platform. However, this example is merely illustrative and not limiting. The technology disclosed herein is not limited to this builder and can be used with any other automated flow builder application. Processing of a new builder application involves creating step descriptions used as the output of the text-to-text model 110 and updating the text-to-API converter 112 with a new API call (e.g., creating a one-to-one mapping update). Further, either the text-to-flow unit 100 or one or more of the flow builder applications may be integrated into a single combined unit without changing the technology disclosed herein.
[0031] In the example of the figure, a portion of the communication path between components is shown. Other communication paths may exist and the example of FIG. 1 is simplified for clarity of this example. Components are shown one by one for simplicity of the drawing, but any of the components shown in FIG. 1 may additionally exist. The number of components and connections shown in FIG. 1 are merely illustrative. Components not shown in FIG. 1 may exist.
[0032] FIG. 2 is a diagram showing an example of a computerized flow. In this example, flow 200 is used to automate an information technology management process. In the example of the figure, flow 200 is composed of a trigger and 13 operation steps that respond to the trigger. As shown in the figure, the steps can be actions (single steps), sub-flows (groups of steps), or flow logic (e.g., "if" statements, etc.). As another example, the flow can automate the following process: "When an email is received, create an incident record." Here, the email reception is the trigger, and the creation of the incident record is the step that is executed when the trigger condition is met. Also, flow 200 is an example of a computerized flow created using a graphical user interface environment. To create a flow in such an environment, a user can rely on a set of buttons to add or modify steps. For example, in the example of the figure, button 202 can be used to add an action, a sub-flow, or flow logic.
[0033] As described above in this specification, there are constraints when generating a flow using a graphical user interface. For example, the user must be familiar with the platform tables and fields used by the application or process and must know all available steps. These requirements make it more difficult to learn to use a graphical user interface, especially for new users and especially when the design environment has many available steps and configurations. Another constraint is that experienced users may not notice new features and may be able to construct a flow using old functions and methods. In some cases, not using the latest features can affect the performance, stability, or security of the generated flow. User training is a way to overcome this constraint, but it requires a significant amount of human effort. Another constraint is that less experienced users may not follow best practices when constructing a flow. Some steps require further processing (such as checking for errors or edge cases that affect the quality and stability of the flow execution). The techniques disclosed in this specification for text-input-based automatic flow generation address the above constraints.
[0034] FIG. 3 is a diagram showing an example of an automatically implemented computerized flow that promotes best practices. Generally speaking, creating a computerized flow is similar to writing computer code, except that the user can create the computerized flow by interacting with a user interface instead of writing computer code in a text editor. Similar to writing computer code, it is preferable for the user to follow best practices while constructing the flow to ensure quality and robustness. Examples of best practices include handling errors or edge cases, as well as following how a user group handles some specific use cases. For example, the step of reporting a runtime error could be that a certain group sends an email to an administrator or sends an SMS to another group. However, inexperienced users may not follow these practices. Using the techniques disclosed herein, a machine learning model can predict best practices based on patterns in the data, eliminating the need for any feature engineering or manual intervention. Furthermore, since this function is data-driven, it can be adapted to a single user or user group without changing the training procedures and usage methods at all. For example, assume the following redundant flow description. "When an incident record is created, if the sender includes the chief, update the incident record; otherwise, if the sender's country is Brazil, update the incident record; otherwise, if the sender is a VP, update the incident record. Then, classify the incident case. If the confidence level exceeds 80%, update the incident record." Flow 300 is the flow predicted using the techniques disclosed herein for the above flow description. Flow 300 does not include a set of if-else statements (underlined above). Instead, a decision table step is predicted and included in Flow 300, which is a better way to handle such use cases compared to a set of if-else statements that are more difficult to maintain.This is an example that promotes best practices for improving the flow from a technical perspective and further teaches inexperienced users how to properly construct a computerized flow. This reduces the time required for training inexperienced users, especially when a large portion of the target users are non-technical contributors.
[0035] FIG. 4 is a block diagram showing an embodiment of a system for synthetically generating training data. In some embodiments, the synthetically generated training data is used to train the text-to-text model 110 of FIG. 1. The text-to-text model 110 of FIG. 1 can be trained with little or no labeled training data as a result of using the synthetically generated training data. Also, it is possible to perform conventional machine learning model training using labeled training data without using the synthetically generated training data. Since the text-to-text model 110 is a deep learning model, it is important to have a large amount of training data. However, finding an acceptable amount of labeled data is a difficult task. There may be little previous work where data is available, and labeling data is costly because it requires manual labor. To solve these problems, in some embodiments, the text-to-text model 110 is trained in the following two stages. A first stage in which the model is trained on a large set of synthetically generated data (the goal of this stage is to learn how to predict the flow from a natural language description regardless of the actual steps, order, flow length, etc.), and a second stage in which the model is trained on a smaller set of data that reflects the patterns of actual users. These training techniques improve the performance of the model and reduce the need for labeled data. In some scenarios, these training techniques are necessary because machine learning-based techniques perform worse than rule-based systems without a significant amount of data. In other words, data augmentation techniques may be needed due to a lack of training data.
[0036] In the example of the figure, the paraphrasing model 404 receives the description 402. In various embodiments, the description 402 is a known text input. These may be known steps of a flow previously generated by the flow builder application 114 of FIG. 1. In various embodiments, the paraphrasing model 404 is a text-to-text machine learning model trained to generate paraphrased text of the description 402. In some embodiments, the paraphrasing model 404 is an LLM (e.g., the T5 model) specially trained for the paraphrasing task. The parameters of the paraphrasing model 404 (e.g., seed, output length, fluency, etc.) can be controlled and varied to generate one set of variations 406 for each input of the description 402. After a sufficiently reasonable description is generated, an automated flow can be generated by the synthetic data generator 408. In some embodiments, the synthetic data generator 408 is composed of one or more processors configured to construct a flow based on input parameters and steps of the flow. In the example of the figure, the synthetic data generator 408 receives the variations 406, the step distribution 410, and the values 412 to generate the synthetic flow 414. In various embodiments, the synthetic data generator 408 first randomly determines the flow length according to a known length distribution that matches the existing usage characteristics. The flow length indicates the number of steps for the synthetic data generator 408 to construct a synthetic flow from the variations 406. Then, in various embodiments, for each step, the synthetic data generator 408 randomly selects a step according to the step distribution 410, randomly selects one of the reasonable descriptions of the step with equal probability, and randomly or with equal probability generates slot values from the values 412 for dates, times, email, etc. In various embodiments, the step distribution 410 is a known distribution indicating the likelihood of a particular step occurring in an actual situation. The step distribution 410 can be obtained from actual data to match a particular user's pattern with an existing flow, or can be randomly generated.In various embodiments, the value 412 is a set of known values for a table, name, system name, etc. The synthetic data generator 408 can be called multiple times to generate multiple synthetic flows.
[0037] Figures 5A - 5C are diagrams showing examples of user interfaces associated with automatically implementing computerized flows based on text input. In the user interface window 500 of Figure 5A, the user can enter various information related to an information technology (IT) case (such as a text description of the IT case) into the text box 502. The window 510 of Figure 5B is an example of an interface element utilized by a developer after the user submits information to the user interface window 500. The text within the text box 512 is used for flow generation. In the example of the figure, this text is as follows: "When a case submission is created, for all affected areas, create a compliance case task record and send a notification." In the example of the figure, the user can click the submit button 514 to submit the flow description within the text box 512 for automatic generation of a flow based on the submitted flow description. In some embodiments, the flow description is the flow description 120 of Figure 1 that is submitted to the text - to - flow unit 100. In some embodiments, the text - to - text model 110 of Figure 1 is used to convert the flow description into an API call to generate the corresponding flow. In the example of the figure, the submission of the flow description is performed within the same user interface used to graphically generate the flow. In other words, in this example, instead of submitting a text description to the text box 512, the same flow can be generated by adding triggers and actions via the graphical user interface using the trigger - add button 516 and the action - add button 518.
[0038] The flow 520 in FIG. 5C is a flow generated based on the flow description submitted via the text box 512. In the example of the figure, the format of the flow 520 is the same as the formats of the flow 200 in FIG. 2 and the flow 300 in FIG. 3. In the example of the figure, the part of the flow description "when the case submission was created" corresponds to the trigger part 522 of the flow 520, and the other parts of the flow description correspond to the corresponding action steps. In the example shown in the figure, the user can graphically and manually add additional action steps to the generated flow by clicking the action addition button 524. Also, in the example of the figure, the user can examine the details of the trigger by clicking the trigger part 522. Thereby, a user interface element 526 including information about the trigger (such as the slot value generated for the trigger) is displayed. In this example, the slot value is the table 528.
[0039] FIG. 6 is a flowchart showing an embodiment of a process for automatically implementing a computerized flow based on text input. In some embodiments, the process of FIG. 6 is executed by the text-to-flow unit 100 of FIG. 1.
[0040] In step 602, a user-provided text description of at least a part of the desired workflow is received. In some embodiments, the user-provided text description is the description of the flow in step 120 of FIG. 1. In various embodiments, the user-provided text description is a natural language input from the user.
[0041] In operation 604, context information associated with a desired workflow is determined. In some embodiments, the context information is determined from received input other than the received user-provided text description. Examples of such input include the builder current state 122, context 124, and flow embedding 126 of FIG. 1. The context information is in the form of various types of information that can be used to adjust and fine-tune the generated final computerized flow, including existing steps of the flow (in the scenario where the generated flow is a partial flow that is added to existing steps), metadata associated with the flow and flow generation preferences, and flow embedding selection that weights and / or biases a basic machine learning model used to predict the computerized flow, but is not limited thereto.
[0042] In operation 606, machine learning input based at least in part on the text description and the context information is provided to a machine learning model to determine an implementation prediction for the desired workflow. In some embodiments, the machine learning model is the text-to-text model 110 of FIG. 1. In various embodiments, the machine learning model outputs the implementation prediction in text format.
[0043] In operation 608, one or more processors are used to automatically implement the implementation prediction as at least a partial computerized workflow implementation of the desired workflow. In some embodiments, the implementation prediction is converted from text format to an API call to a flow builder application. In some embodiments, the flow builder application is the flow builder application 114 of FIG. 1.
[0044] FIG. 7 is a functional diagram showing a programmed computer system. In some embodiments, the processing of FIG. 6 is performed by computer system 700. Computer system 700 is an example of a processor.
[0045] In the example of the figure, computer system 700 comprises various subsystems as described below. Computer system 700 includes at least one microprocessor subsystem (also referred to as a processor or central processing unit (CPU)) 702. Computer system 700 may be a physical system or a virtual system (such as a virtual machine). For example, processor 702 may be implemented by a single-chip processor or a multiprocessor. In some embodiments, processor 702 is a general-purpose digital processor that controls the operation of computer system 700. Using instructions retrieved from memory 710, processor 702 controls the reception and manipulation of input data, as well as the output and display of data on an output device (e.g., display 718).
[0046] Processor 702 is bidirectionally connected to memory 710, and memory 710 may include a first primary storage (typically random access memory (RAM)) and a second primary storage area (typically read-only memory (ROM)). As is well known to those skilled in the art, primary storage is available as a general storage area and as a scratchpad memory, and is also available for storing input data and processed data. Primary storage can further store programming instructions and data in the form of data objects and text objects, in addition to other data and instructions for processing executed on processor 702. Also, as is well known to those skilled in the art, primary storage typically comprises the basic operating instructions, program code, data, and objects used by processor 702 to execute functions (e.g., programmed instructions). For example, memory 710 may include any suitable computer-readable storage medium described below, depending on whether data access needs to be bidirectional or unidirectional. For example, processor 702 can directly and very quickly store and retrieve frequently needed data in a cache memory (not shown).
[0047] A persistent memory 712 (e.g., a removable mass storage device) provides additional data storage capacity to the computer system 700 and is connected to the processor 702 in a bidirectional (read / write) or unidirectional (read-only) manner. For example, the persistent memory 712 may also include computer-readable media such as magnetic tape, flash memory, PC cards, portable mass storage devices, holographic storage devices, and other storage devices. The fixed mass storage 720 can also provide additional data storage capacity, for example. The most common example of the fixed mass storage 720 is a hard disk drive. The persistent memory 712 and the fixed mass storage 720 generally store additional programming instructions, data, etc. that are not typically used much by the processor 702. It is understood that the information held in the persistent memory 712 and the fixed mass storage 720 can be incorporated in a standard manner into a part of the memory 710 (e.g., RAM) as virtual memory if necessary.
[0048] In addition to enabling the processor 702 to access the storage subsystem, a bus 714 may be used to enable access to other subsystems and devices. As shown in the figure, these may include a display monitor 718, a network interface 716, a keyboard 704, and a pointing device 706, and, optionally, an auxiliary input / output device interface, a sound card, speakers, and other subsystems. For example, the pointing device 706 may be a mouse, a stylus, a trackball, or a tablet, and is useful for interacting with the graphical user interface.
[0049] As shown in the figure, network interface 716 enables processor 702 to be connected to another computer, computer network, or telecommunications network using a network connection. For example, through network interface 716, processor 702 can receive information (such as data objects or program instructions) from another network or output information to another network during the execution of a method / processing step. Information is often represented as a series of instructions executed on a processor, can be received from another network, and can be output to another network. Using an interface card (or similar device) and appropriate software implemented (e.g., executed / implemented) by processor 702, computer system 700 can be connected to an external network and data can be transferred according to standard protocols. The processing may be executed on processor 702 or may be executed on a network (such as the Internet, intranet, or local area network) together with a remote processor that shares part of the processing. An additional mass storage device (not shown) may be connected to processor 702 through network interface 716.
[0050] An auxiliary I / O device interface (not shown) may be used with computer system 700. The auxiliary I / O device interface can include a general-purpose interface and a customized interface that enable processor 702 to send data and, more typically, receive data from other devices (such as microphones, touch sensor-based displays, transducer card readers, tape readers, voice or handwriting recognition devices, biometric readers, cameras, portable mass storage devices, and other computers).
[0051] Furthermore, various embodiments disclosed herein further relate to a computer storage product including a computer-readable medium having program code for performing various computer-implemented operations. A computer-readable medium is any data storage device that can store data that can later be read by a computer system. Examples of computer-readable media include, but are not limited to, all of the above media. Magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD-ROM disks, magneto-optical media such as optical disks, and specially configured hardware devices such as application specific integrated circuits (ASICs), programmable logic devices (PLDs), and ROM / RAM devices. Examples of program code include, for example, machine code generated by a compiler, or files containing high-level code (e.g., scripts) that can be executed using an interpreter.
[0052] The computer system shown in FIG. 7 is merely an example of a computer system suitable for use with various embodiments disclosed herein. Other computer systems suitable for such use may include more or fewer subsystems. Further, bus 714 is an example of any interconnect scheme that functions to connect the subsystems. Other computer architectures with different configurations of subsystems may be utilized.
[0053] The above embodiments have been described in some detail for ease of understanding, but the present invention is not limited to the provided details. There are many alternative ways to implement the present invention. The disclosed embodiments are illustrative and not intended to be limiting. [Application Example 1] A method comprising: receiving a user-provided text description of at least a part of a desired workflow; determining context information associated with the desired workflow; providing machine learning input based at least in part on the text description and the context information to a machine learning model to determine an implementation prediction for the desired workflow; automatically implementing the implementation prediction as at least a part of a computerized workflow implementation of the desired workflow using one or more processors. A method comprising the above steps. [Application Example 2] The method according to Application Example 1, wherein the user-provided text description includes natural language input. [Application Example 3] The method according to Application Example 1, wherein the desired workflow is configured to be executed on a computer and comprises a trigger condition and one or more action steps configured to be executed in response to a determination that the trigger condition has occurred. [Application Example 4] The method according to Application Example 1, wherein the context information includes processing data associated with existing steps within the desired workflow. [Application Example 5] The method according to Application Example 1, wherein the context information includes processing data associated with adjustment parameters of the machine learning model. [Application Example 6] The method according to Application Example 1, wherein determining the context information includes converting data in an Extensible Markup Language (XML) or JavaScript Object Notation (JSON) format to a text format. [Application Example 7] The method according to Application Example 1, wherein determining the context information includes selecting a tensor data object from a list of tensor data objects, and each tensor data object in the list of tensor data objects is associated with different embedding weights for the machine learning model. [Application Example 8] The method according to Application Example 1, wherein the machine learning model is a text-to-text pre-trained language model. [Application Example 9] The method according to Application Example 8, wherein the text-to-text pre-trained language model has an encoder-decoder architecture. [Application Example 10] The method according to Application Example 1, wherein the machine learning model is pre-trained on a large language dataset and then fine-tuned for a text-to-workflow prediction task. [Application Example 11] The method according to Application Example 1, wherein the machine learning model is trained based at least in part on synthetically generated training data. [Application Example 12] The method according to Application Example 11, wherein the synthetically generated training data includes a plurality of variations of non-synthetically generated workflow descriptions. [Application Example 13] The method according to Application Example 1, wherein automatically performing the implementation prediction as the computerized workflow implementation using the one or more processors includes converting the implementation prediction from a text format into one or more application programming interface messages. [Application Example 14] The method according to Application Example 13, further comprising sending the one or more application programming interface messages to an application configured to generate the computerized workflow implementation. [Application Example 15] The method according to Application Example 1, wherein the computerized workflow implementation is part of the desired workflow. [Application Example 16] The method according to Application Example 1, further comprising displaying the desired workflow including the computerized workflow implementation on a graphical user interface. [Application Example 17] The method according to Application Example 16, further comprising receiving a user request to add steps to the desired workflow via the graphical user interface. [Application Example 18] The method according to Application Example 1, wherein the user-provided text description is associated with an information technology case. [Application Example 19] A system, one or more processors, receiving a user-provided text description of at least a part of a desired workflow, Determine context information associated with the desired workflow, To determine an implementation prediction for the desired workflow, provide machine learning input based at least in part on the text description and the context information to a machine learning model, One or more processors configured to automatically perform the implementation prediction as at least part of a computerized workflow implementation of the desired workflow; A memory connected to at least one of the one or more processors and configured to provide instructions to at least one of the one or more processors; A system comprising the same. [Application Example 20] A computer program product embodied in a non-transitory computer-readable medium, Computer instructions for receiving at least a user-provided text description of at least a part of a desired workflow, Computer instructions for determining context information associated with the desired workflow, Computer instructions for providing machine learning input based at least in part on the text description and the context information to a machine learning model to determine an implementation prediction for the desired workflow, Computer instructions for automatically performing the implementation prediction as at least part of a computerized workflow implementation of the desired workflow using one or more processors; A computer program product comprising the same.
Claims
1. A method comprising: a processing system receiving a user-provided text description of at least a portion of a desired workflow; the processing system determining context information associated with the desired workflow based on selecting one of a plurality of embeddings generated for a pre-trained machine learning model; the processing system providing a machine learning input based at least in part on the text description and the context information to the pre-trained machine learning model to determine a predicted implementation of the desired workflow; the processing system generating a computerized workflow implementation of at least a portion of the desired workflow based on the predicted implementation of the desired workflow. A method.
2. The method of claim 1, wherein the user-provided text description includes natural language input.
3. The method of claim 1, wherein the computerized workflow implementation is configured to execute on a computer and comprises a trigger condition and one or more action steps configured to execute in response to a determination that the trigger condition has occurred.
4. The method of claim 1, wherein the context information includes processing data associated with existing steps within the desired workflow.
5. The method of claim 1, wherein the context information includes processing data associated with adjustment parameters of the machine learning model.
6. The method of claim 1, wherein determining the context information includes converting data in an extensible markup language (XML) or JavaScript object notation (JSON) format to a text format.
7. The method of claim 1, wherein determining the context information includes selecting a tensor data object from a list of tensor data objects, each tensor data object in the list of tensor data objects being associated with different embedding weights for the machine learning model.
8. The method according to claim 1, wherein the machine learning model is a text-to-text pre-trained language model.
9. The method according to claim 8, wherein the text-to-text pre-trained language model comprises an encoder-decoder architecture.
10. The method according to claim 1, wherein the pre-trained machine learning model is pre-trained on a large-scale language dataset and then fine-tuned for a text-to-workflow prediction task.
11. The method according to claim 1, wherein the pre-trained machine learning model is trained at least partially based on synthetically generated training data.
12. The method according to claim 1, wherein the processing system generating the computerized workflow implementation comprises converting the predicted implementation of the desired workflow from a text format into one or more application programming interface messages.
13. The method according to claim 12, further comprising the processing system sending the one or more application programming interface messages to an application configured to generate the computerized workflow implementation.
14. The method according to claim 1, wherein the computerized workflow implementation is part of the desired workflow.
15. The method according to claim 1, further comprising the processing system displaying the computerized workflow implementation on a graphical user interface.
16. The method according to claim 15, further comprising the processing system receiving, via the graphical user interface, a user request to add steps to the desired workflow.
17. The method according to claim 1, wherein the user-provided text description is associated with an information technology case. **Claim 18**: The method according to claim 1, wherein the plurality of embeddings comprise dynamically exchangeable options for selectively adjusting the pre-trained machine learning model using different training data sets. **Claim 19** A system comprising: one or more processors configured to: receive a user-provided text description of at least a portion of a desired workflow; determine context information associated with the desired workflow based on selecting one of a plurality of embeddings generated for a pre-trained machine learning model; provide machine learning input based at least in part on the text description and the context information to the pre-trained machine learning model to determine a predicted implementation of the desired workflow; generate at least a partial computerized workflow implementation of the desired workflow based on the predicted implementation of the desired workflow; and a memory connected to at least one of the one or more processors and configured to provide instructions to at least one of the one or more processors. A system. **Claim 20**: A non-transitory computer-readable medium comprising instructions that, when executed by a processing circuit, cause the processing circuit to perform operations comprising: receiving a user-provided text description of at least a portion of a desired workflow; determining context information associated with the desired workflow based on selecting one of a plurality of embeddings generated for a pre-trained machine learning model; providing machine learning input based at least in part on the text description and the context information to the pre-trained machine learning model to determine a predicted implementation of the desired workflow; and generating at least a partial computerized workflow implementation of the desired workflow based on the predicted implementation of the desired workflow.
Citation Information
Patent Citations
Generation program, generation device, and generation method
JP2019159602A
Task automation by support robots for robotic process automation (RPA)
WO2022081380A1