Methods and systems for contextual chatbot operation
The chatbot engine addresses the limitations of conventional chatbots by using context prompts and task history to provide appropriate contextual information to the LLM, improving response accuracy and user experience while reducing resource consumption.
Patent Information
- Application Number
- PCT/CA2024/050260
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-18
- Filing Date
- 2024-03-01
- Publication Date
- 2025-05-22
AI Technical Summary
Conventional chatbots struggle to understand user navigational context, leading to inaccurate responses and inefficient user interactions, as they fail to maintain historical context and provide options for previewing or confirming tasks.
A chatbot engine that automatically extracts contextual information from the current page and provides it to a large language model (LLM) via context prompts, checks the suitability of the context for requested tasks, maintains task history, and offers users the option to preview or confirm operations before execution.
This solution enhances the accuracy and relevance of chatbot responses, reduces user frustration, and conserves computing resources by ensuring that the LLM receives appropriate contextual information and that users have control over task execution.
Smart Images

Figure CA2024050260_22052025_PF_FP_ABST
Abstract
Description
METHODS AND SYSTEMS FOR CONTEXTUAL CHATBOT OPERATIONCROSS-REFERENCE TO RELATED APPLICATION
[0001] The present disclosure claims priority from U.S. provisional patent application no. 63 / 622,334, filed January 18, 2024, entitled “METHODS AND SYSTEMS FOR CONTEXTUAL CHATBOT OPERATION”; and U.S. provisional patent application no. 63 / 598,803, filed November 14, 2023, entitled “METHODS AND SYSTEMS FOR CONTEXTUAL CHATBOT OPERATION”, all of which are hereby incorporated by reference in their entireties.FIELD
[0002] The present disclosure relates to machine learning and large language models (LLMs), and, more particularly, to contextual operation of a LLM-based chatbot.BACKGROUND
[0003] A large language model (LLM) is a deep learning algorithm that can process natural language to summarize, translate, predict and generate text and other content. A LLM may be trained to leam billions of parameters in order to model how words relate to each other in a textual sequence. Inputs to a LLM may be referred to as prompts. A prompt is a natural language input that includes instructions to cause the LLM to generate a desired output.
[0004] A chatbot is a type of artificial intelligence that typically provides assistance to a user via a conversational interaction. Some chatbots make use of LLMs to carry out user interactions. Chatbots may also be referred to as virtual assistants, conversational agents or smart assistants.SUMMARY
[0005] Conventionally, chatbots may be provided, for example on a website or application, to assist a user such as by providing information in a conversational manner. However, conventional chatbots often are unable to understand the user’s navigational context (e.g., unable to understand which page a user is currently at or which tab or window of the application theuser currently has active) and unable to correctly navigate within the application or website based on the user’s input, such as when the user’s input is a request to perform a task at a particular page or tab. This limitation often results in inaccurate responses or inappropriate actions, because the chatbot provides output that does not account for the user’s context. If the user is not aware that the output is inaccurate or inappropriate, this may result in the user inputting commands or performing actions that are invalid and / or that must be later undone, thus wasting computing resources and leading to user frustration.
[0006] Moreover, if the user has navigated from one page or tab to another, conventional chatbots typically do not have the ability to maintain the historical context of that navigation, leading to an inefficient and disjointed user experience (e.g., the user must repeat the same request that was previously inputted in an earlier page). The result is a poor user experience and wasted computing resources because user input must be repeated and processed again.
[0007] As well, conventional chatbots that help a user to perform a task typically perform the requested task without providing the user with the option to preview or confirm the task. This inability to preview or confirm a task can lead to user dissatisfaction and potential errors, particularly where the task involves a coding change or other operational change that may not be easily understood or undone by the user manually. Again, this may lead to erroneous operation and / or wasted computing resources (e.g., processing power and memory resources are consumed to perform the erroneous task and then again consumed to undo the task).
[0008] In various examples, the present disclosure provides a technical solution for implementing a chatbot engine that addresses at least some of the above drawbacks. Examples of the disclosed chatbot engine are configured to automatically extract contextual information (which may be coded into a page) from a current page. The disclosed chatbot engine provides the LLM with the contextual information (e.g., via a context prompt) and then prompts the LLM to check whether the context is suitable for a requested task. This provides a technical advantage in that the LLM is provided with suitable and appropriate contextual information to enable the LLM to generate appropriate output and avoid having to perform multiple iterations of prompting (which would consume computing resources with each iteration) in order to get the desired result from the LLM.
[0009] Examples of the disclosed chatbot engine may maintain a task history and may automatically include the task history in a task prompt to the LLM. This provides a technicaladvantage in that, even if the user has navigated from one page to another, the user does not need to repeat previous input (e.g., the user does not need to repeat the same request that was previously inputted in an earlier page), thus reducing latency in the user experience and saving computing resources required to process the user input again.
[0010] Examples of the disclosed chatbot engine may provide the user with the option to preview or confirm an operation prior to executing the operation command. This provides the technical advantage that the user is provided with the opportunity to spot potential errors prior to implementing a change, thus avoiding possible erroneous operation and / or wasted computing resources (e.g., avoiding use of processing power and memory resources to perform the erroneous task and then again consumed to undo the task).
[0011] In an example aspect, the present disclosure describes a computing system including a processing unit configured to execute computer-readable instructions to cause the system to: while a user interface (UI) is at a page, provide a context prompt to a large language model (LLM), the context prompt providing contextual information including information about the page; provide a check context prompt to the LLM instructing the LLM to determine a suitable context for performing a task; receive output from the LLM based on the check context prompt, the output including a confirmation that the page offers a suitable context for performing the task; provide a task prompt to the LLM instructing the LLM to generate an operation command for performing the task using a functionality of the page; and receive output from the LLM based on the task prompt, the output including the operation command for performing the task using the functionality of the page.
[0012] In an example of the preceding example aspect of the system, the page may be a target page, and the processing unit may be further configured to execute computer-readable instructions to cause the computer system to, prior to the UI being at the target page: while the UI is at a prior page different from the target page, provide another context prompt to the LLM, the another context prompt providing contextual information including information about the prior page; provide another check context prompt to the LLM instructing the LLM to determine a suitable context for performing the task; and receive output from the LLM based on the another check context prompt, the output including a navigation command to navigate to the target page.
[0013] In an example of the preceding example aspect of the system, the processing unit may be further configured to execute computer-readable instructions to cause the computer system to:provide a navigation option in the UI for executing the navigation command to navigate the UI to the target page; and navigate the UI to the target page responsive to selection of the navigation option.
[0014] In an example of the preceding example aspect of the system, the context prompt providing contextual information including information about the target page may be automatically provided responsive to navigating the UI to the target page.
[0015] In an example of a preceding example aspect of the system, the processing unit may be further configured to execute computer-readable instructions to cause the computer system to: while the UI is at the prior page and prior to providing the other check context prompt, receive a task request via user input to the UI, wherein the other check context prompt is provided responsive to receiving the task request; and after the UI has navigated to the target page and responsive to receiving the confirmation of the target page, automatically provide the task prompt to the LLM instructing the LLM to generate an operation command for performing the task using the functionality of the target page.
[0016] In an example of any of the preceding example aspects of the system, the processing unit may be further configured to execute computer-readable instructions to cause the computer system to: responsive to navigation of the UI to the page, automatically extract the contextual information from the page and automatically provide the context prompt to the LLM.
[0017] In an example of the preceding example aspect of the system, the processing unit may be further configured to execute computer-readable instructions to cause the computer system to: generate the context prompt using the contextual information extracted from the page.
[0018] In an example of a preceding example aspect of the system, the contextual information may be automatically extracted by automatically executing code embedded in the page.
[0019] In an example of any of the preceding example aspects of the system, the processing unit may be further configured to execute computer-readable instructions to cause the computer system to: search a task history database for a historical LLM session portion relevant to the task; and include the historical LLM session portion in the task prompt.
[0020] In an example of any of the preceding example aspects of the system, the processing unit may be further configured to execute computer-readable instructions to cause the computer system to: provide an operation option in the UI for executing the operation command; andexecute the operation command responsive to selection of the operation option.
[0021] In an example of the preceding example aspect of the system, the processing unit may be further configured to execute computer-readable instructions to cause the computer system to: prior to executing the operation command, provide a preview of a result of executing the operation command.
[0022] In an example of any of the preceding example aspects of the system, the context prompt may provide contextual information including information about the functionality of the page.
[0023] In another example aspect, the present disclosure describes a method including: while a user interface (UI) is at a page, providing a context prompt to a large language model (LLM), the context prompt providing contextual information including information about the page; providing a check context prompt to the LLM instructing the LLM to determine a suitable context for performing a task; receiving output from the LLM based on the check context prompt, the output including a confirmation that the page offers a suitable context for performing the task; providing a task prompt to the LLM instructing the LLM to generate an operation command for performing the task using a functionality of the page; and receiving output from the LLM based on the task prompt, the output including the operation command for performing the task using the functionality of the page.
[0024] In an example of the preceding example aspect of the method, the page may be a target page, and the method may include, prior to the UI being at the target page: while the UI is at a prior page different from the target page, providing another context prompt to the LLM, the another context prompt providing contextual information including information about the prior page; providing another check context prompt to the LLM instructing the LLM to determine a suitable context for performing the task; and receiving output from the LLM based on the another check context prompt, the output including a navigation command to navigate to the target page.
[0025] In an example of the preceding example aspect of the method, the method may include: providing a navigation option in the UI for executing the navigation command to navigate the UI to the target page; and navigating the UI to the target page responsive to selection of the navigation option.
[0026] In an example of the preceding example aspect of the method, the context promptproviding contextual information including information about the target page may be automatically provided responsive to navigating the UI to the target page.
[0027] In an example of a preceding example aspect of the method, the method may include: while the UI is at the prior page and prior to providing the other check context prompt, receiving a task request via user input to the UI, wherein the other check context prompt is provided responsive to receiving the task request; and after the UI has navigated to the target page and responsive to receiving the confirmation of the target page, automatically providing the task prompt to the LLM instructing the LLM to generate an operation command for performing the task using the functionality of the target page.
[0028] In an example of any of the preceding example aspects of the method, the method may include: responsive to navigation of the UI to the page, automatically extracting the contextual information from the page and automatically providing the context prompt to the LLM.
[0029] In an example of the preceding example aspect of the method, the method may include: generating the current context prompt using the contextual information extracted from the page.
[0030] In an example of a preceding example aspect of the method, the contextual information may be automatically extracted by automatically executing code embedded in the page.
[0031] In an example of any of the preceding example aspects of the method, the method may include: searching a task history database for a historical LLM session portion relevant to the task; and including the historical LLM session portion in the task prompt.
[0032] In an example of any of the preceding example aspects of the method, the method may include: providing an operation option in the UI for executing the operation command; and executing the operation command responsive to selection of the operation option.
[0033] In an example of the preceding example aspect of the method, the method may include: prior to executing the operation command, providing a preview of a result of executing the operation command.
[0034] In an example of any of the preceding example aspects of the method, the context prompt may provide contextual information including information about the functionality of the page.
[0035] In another example aspect, the present disclosure provides a non-transitory computer- readable medium storing instructions that, when executed by a processing unit of a computingsystem, cause the computing system to: while a user interface (UI) is at a page, provide a context prompt to a large language model (LLM), the context prompt providing contextual information including information about the page; provide a check context prompt to the LLM instructing the LLM to determine a suitable context for performing a task; receive output from the LLM based on the check context prompt, the output including a confirmation that the page offers a suitable context for performing the task; provide a task prompt to the LLM instructing the LLM to generate an operation command for performing the task using a functionality of the page; and receive output from the LLM based on the task prompt, the output including the operation command for performing the task using the functionality of the page.
[0036] In some examples, the computer-readable medium may store instructions that, when executed by the processor of the computing system, cause the computing system to perform any of the example aspect of the methods described above.
[0037] In another example aspect, the present disclosure provides a computer program including processor-executable instructions that, when executed by a processor of a computing system, cause the computing system to perform any of the example aspect of the methods described above.BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Reference will now be made, by way of example, to the accompanying drawings which show example embodiments of the present application, and in which:
[0039] FIG. 1 A is a block diagram of a simplified convolutional neural network, which may be used in examples of the present disclosure;
[0040] FIG. IB is a block diagram of a simplified transformer neural network, which may be used in examples of the present disclosure;
[0041] FIG. 2 is a block diagram of an example computing system, which may be used to implement examples of the present disclosure;
[0042] FIGS. 3A-3C are signalling diagrams illustrating example communications performed by modules of an example chatbot engine, in accordance with examples of the present disclosure;
[0043] FIG. 4 is a flowchart illustrating an example method for contextual operation of an example chatbot engine, in accordance with examples of the present disclosure;
[0044] FIGS. 5A-5C illustrate a simplified example user interface showing contextual operation of an example chatbot engine, in accordance with examples of the present disclosure;
[0045] FIG. 6 is a block diagram of an example e-commerce platform, which may be an example implementation of the examples disclosed herein; and
[0046] FIG. 7 is an example homepage of an administrator, which may be accessed via the e- commerce platform of FIG. 6.
[0047] Similar reference numerals may have been used in different figures to denote similar components.DETAILED DESCRIPTION
[0048] In various examples, the present disclosure describes methods and systems for implementing a chatbot engine, which may include a chatbot user interface (UI) (which may provide outputs to a user and receiver inputs from the user) and a chatbot backend (which may interface with system tools, remote servers, etc.). The chatbot engine generates prompts to a large language model (LLM) and receives output from the LLM, such as navigation output that includes a command to navigate among pages of a website or portal, or among tabs, panels or windows of an application. For simplicity, the present disclosure may describe examples in the context of navigating pages (e.g., web pages, or sub-pages of an administrative page or portal), however this is not intended to be limiting. For example, the navigation may be between windows, panels or tabs of an application, among other possibilities.
[0049] Examples of the disclosed chatbot engine are configured to automatically extract contextual information (which may be coded into a page) from a current page. The contextual information may be provided to the LLM via a context prompt. When a task request is inputted by the user, the disclosed chatbot engine prompts the LLM to check whether the current context is suitable for the requested task. This helps to ensure that, prior to prompting the LLM to generate output for the requested task, the LLM is getting suitable and appropriate contextual information to perform the requested task. If the LLM is prompted to generate output for the requested task without ensuring that the current context is suitable for the task, the LLM maygenerate incorrect output (e.g., generating a command for an incorrect or invalid operation, generating a command for an operation in the wrong context, requesting information that is not available in the current context, etc.) and / or may generate excessively complex / long output (e.g., generating a long command that involves navigation steps). Both incorrect output and excessive output negatively impact the performance of the chatbot engine, provide a poor user experience and is a drain on computing resources (e.g., consuming communication bandwidth, processing power, number of tokens, etc.). Additionally, if the user is not aware that the output is inaccurate or inappropriate, this may result in the user executing commands or performing actions that are invalid and / or that must be later undone, thus wasting computing resources and leading to user frustration. These drawbacks may be avoided and computing efficiency may be improved by examples disclosed herein.
[0050] Examples of the disclosed chatbot engine may maintain a task history and may automatically include the task history in a task prompt to the LLM. For example, after navigating to a new page that provides a suitable context for a requested task, the chatbot engine may automatically repeat a previous task prompt to the LLM. This provides a technical advantage in that, even if the user has navigated from one page to another, the user does not need to repeat previous input (e.g., the user does not need to repeat the same task request that was previously inputted in an earlier page), thus reducing latency in the user experience and saving computing resources required to process the user input again.
[0051] Examples of the disclosed chatbot engine may provide the user with the option to preview or confirm an operation prior to executing the operation command. This provides the technical advantage that the user is provided with the opportunity to spot potential errors prior to implementing a change, thus avoiding possible erroneous operation and / or wasted computing resources (e.g., avoiding use of processing power and memory resources to perform the erroneous task and then again consumed to undo the task).
[0052] As will be discussed further below, examples of the disclosed chatbot engine may send prompts to and receive output from a LLM, which is a type of deep neural network.
[0053] To assist in understanding the present disclosure, some concepts relevant to neural networks and machine learning (ML) are first discussed.
[0054] Generally, a neural network comprises a number of computation units (sometimes referred to as “neurons”). Each neuron receives an input value and applies a function to the inputto generate an output value. The function typically includes a parameter (also referred to as a “weight”) whose value is learned through the process of training. A plurality of neurons may be organized into a neural network layer (or simply “layer”) and there may be multiple such layers in a neural network. The output of one layer may be provided as input to a subsequent layer. Thus, input to a neural network may be processed through a succession of layers until an output of the neural network is generated by a final layer. This is a simplistic discussion of neural networks and there may be more complex neural network designs that include feedback connections, skip connections, and / or other such possible connections between neurons and / or layers, which need not be discussed in detail here.
[0055] A deep neural network (DNN) is a type of neural network having multiple layers and / or a large number of neurons. The term DNN may encompass any neural network having multiple layers, including convolutional neural networks (CNNs), recurrent neural networks (RNNs), and multilayer perceptrons (MLPs), among others.
[0056] DNNs are often used as ML-based models for modeling complex behaviors (e.g., human language, image recognition, object classification, etc.) in order to improve accuracy of outputs (e.g., more accurate predictions) such as, for example, as compared with models with fewer layers. In the present disclosure, the term “ML-based model” or more simply “ML model” may be understood to refer to a DNN. Training a ML model refers to a process of learning the values of the parameters (or weights) of the neurons in the layers such that the ML model is able to model the target behavior to a desired degree of accuracy. Training typically requires the use of a training dataset, which is a set of data that is relevant to the target behavior of the ML model. For example, to train a ML model that is intended to model human language (also referred to as a language model), the training dataset may be a collection of text documents, referred to as a text corpus (or simply referred to as a corpus). The corpus may represent a language domain (e.g., a single language), a subject domain (e.g., scientific papers), and / or may encompass another domain or domains, be they larger or smaller than a single language or subject domain. For example, a relatively large, multilingual and non-subject-specific corpus may be created by extracting text from online webpages and / or publicly available social media posts. In another example, to train a ML model that is intended to classify images, the training dataset may be a collection of images. Training data may be annotated with ground truth labels (e.g. each data entry in the training dataset may be paired with a label), or may be unlabeled.
[0057] Training a ML model generally involves inputting into an ML model (e.g. an untrained ML model) training data to be processed by the ML model, processing the training data using the ML model, collecting the output generated by the ML model (e.g. based on the inputted training data), and comparing the output to a desired set of target values. If the training data is labeled, the desired target values may be, e.g., the ground truth labels of the training data. If the training data is unlabeled, the desired target value may be a reconstructed (or otherwise processed) version of the corresponding ML model input (e.g., in the case of an autoencoder), or may be a measure of some target observable effect on the environment (e.g., in the case of a reinforcement learning agent). The parameters of the ML model are updated based on a difference between the generated output value and the desired target value. For example, if the value outputted by the ML model is excessively high, the parameters may be adjusted so as to lower the output value in future training iterations. An objective function is a way to quantitatively represent how close the output value is to the target value. An objective function represents a quantity (or one or more quantities) to be optimized (e.g., minimize a loss or maximize a reward) in order to bring the output value as close to the target value as possible. The goal of training the ML model typically is to minimize a loss function or maximize a reward function.
[0058] The training data may be a subset of a larger data set. For example, a data set may be split into three mutually exclusive subsets: a training set, a validation (or cross-validation) set, and a testing set. The three subsets of data may be used sequentially during ML model training. For example, the training set may be first used to train one or more ML models, each ML model, e.g., having a particular architecture, having a particular training procedure, being describable by a set of model hyperparameters, and / or otherwise being varied from the other of the one or more ML models. The validation (or cross-validation) set may then be used as input data into the trained ML models to, e.g., measure the performance of the trained ML models and / or compare performance between them. Where hyperparameters are used, a new set of hyperparameters may be determined based on the measured performance of one or more of the trained ML models, and the first step of training (i.e., with the training set) may begin again on a different ML model described by the new set of determined hyperparameters. In this way, these steps may be repeated to produce a more performant trained ML model. Once such a trained ML model is obtained (e.g., after the hyperparameters have been adjusted to achieve a desired level of performance), a third step of collecting the output generated by the trained ML model applied to the third subset (the testing set) may begin. The output generated from the testing set may becompared with the corresponding desired target values to give a final assessment of the trained ML model’s accuracy. Other segmentations of the larger data set and / or schemes for using the segments for training one or more ML models are possible.
[0059] Backpropagation is an algorithm for training a ML model. Backpropagation is used to adjust (also referred to as update) the value of the parameters in the ML model, with the goal of optimizing the objective function. For example, a defined loss function is calculated by forward propagation of an input to obtain an output of the ML model and comparison of the output value with the target value. Backpropagation calculates a gradient of the loss function with respect to the parameters of the ML model, and a gradient algorithm (e.g., gradient descent) is used to update (i.e., “learn”) the parameters to reduce the loss function. Backpropagation is performed iteratively, so that the loss function is converged or minimized. Other techniques for learning the parameters of the ML model may be used. The process of updating (or learning) the parameters over many iterations is referred to as training. Training may be carried out iteratively until a convergence condition is met (e.g., a predefined maximum number of iterations has been performed, or the value outputted by the ML model is sufficiently converged with the desired target value), after which the ML model is considered to be sufficiently trained. The values of the learned parameters may then be fixed and the ML model may be deployed to generate output in real-world applications (also referred to as “inference”).
[0060] In some examples, a trained ML model may be fine-tuned, meaning that the values of the learned parameters may be adjusted slightly in order for the ML model to better model a specific task. Fine-tuning of a ML model typically involves further training the ML model on a number of data samples (which may be smaller in number / cardinality than those used to train the model initially) that closely target the specific task. For example, a ML model for generating natural language that has been trained generically on publically-available text corpuses may be, e.g., fine-tuned by further training using the complete works of Shakespeare as training data samples (e.g., where the intended use of the ML model is generating a scene of a play or other textual content in the style of Shakespeare).
[0061] FIG. 1 A is a simplified diagram of an example CNN 10, which is an example of a DNN that is commonly used for image processing tasks such as image classification, image analysis, object segmentation, etc. An input to the CNN 10 may be a 2D RGB image 12.
[0062] The CNN 10 includes a plurality of layers that process the image 12 in order togenerate an output, such as a predicted classification or predicted label for the image 12. For simplicity, only a few layers of the CNN 10 are illustrated including at least one convolutional layer 14. The convolutional layer 14 performs convolution processing, which may involve computing a dot product between the input to the convolutional layer 14 and a convolution kernel. A convolutional kernel is typically a 2D matrix of learned parameters that is applied to the input in order to extract image features. Different convolutional kernels may be applied to extract different image information, such as shape information, color information, etc.
[0063] The output of the convolution layer 14 is a set of feature maps 16 (sometimes referred to as activation maps). Each feature map 16 generally has smaller width and height than the image 12. The set of feature maps 16 encode image features that may be processed by subsequent layers of the CNN 10, depending on the design and intended task for the CNN 10. In this example, a fully connected layer 18 processes the set of feature maps 16 in order to perform a classification of the image, based on the features encoded in the set of feature maps 16. The fully connected layer 18 contains learned parameters that, when applied to the set of feature maps 16, outputs a set of probabilities representing the likelihood that the image 12 belongs to each of a defined set of possible classes. The class having the highest probability may then be outputted as the predicted classification for the image 12.
[0064] In general, a CNN may have different numbers and different types of layers, such as multiple convolution layers, max-pooling layers and / or a fully connected layer, among others. The parameters of the CNN may be learned through training, using data having ground truth labels specific to the desired task (e.g., class labels if the CNN is being trained for a classification task, pixel masks if the CNN is being trained for a segmentation task, text annotations if the CNN is being trained for a captioning task, etc.), as discussed above.
[0065] Some concepts in ML-based language models are now discussed. It may be noted that, while the term “language model” has been commonly used to refer to a ML-based language model, there could exist non-ML language models. In the present disclosure, the term “language model” may be used as shorthand for ML-based language model (i.e., a language model that is implemented using a neural network or other ML architecture), unless stated otherwise. For example, unless stated otherwise, “language model” encompasses LLMs.
[0066] A language model may use a neural network (typically a DNN) to perform natural language processing (NLP) tasks such as language translation, image captioning, grammaticalerror correction, and language generation, among others. A language model may be trained to model how words relate to each other in a textual sequence, based on probabilities. A language model may contain hundreds of thousands of learned parameters or in the case of a large language model (LLM) may contain millions or billions of learned parameters or more.
[0067] In recent years, there has been interest in a type of neural network architecture, referred to as a transformer, for use as language models. For example, the Bidirectional Encoder Representations from Transformers (BERT) model, the Transformer-XL model and the Generative Pre-trained Transformer (GPT) models are types of transformers. A transformer is a type of neural network architecture that uses self-attention mechanisms in order to generate predicted output based on input data that has some sequential meaning (i.e., the order of the input data is meaningful, which is the case for most text input). Although transformer-based language models are described herein, it should be understood that the present disclosure may be applicable to any ML-based language model, including language models based on other neural network architectures such as recurrent neural network (RNN)-based language models.
[0068] FIG. IB is a simplified diagram of an example transformer 50, and a simplified discussion of its operation is now provided. The transformer 50 includes an encoder 52 (which may comprise one or more encoder layers / blocks connected in series) and a decoder 54 (which may comprise one or more decoder layers / blocks connected in series). Generally, the encoder 52 and the decoder 54 each include a plurality of neural network layers, at least one of which may be a self-attention layer. The parameters of the neural network layers may be referred to as the parameters of the language model.
[0069] The transformer 50 may be trained on a text corpus that is labelled (e.g., annotated to indicate verbs, nouns, etc.) or unlabelled. LLMs may be trained on a large unlabelled corpus. Some LLMs may be trained on a large multi-language, multi-domain corpus, to enable the model to be versatile at a variety of language-based tasks such as generative tasks (e.g., generating human-like natural language responses to natural language input).
[0070] An example of how the transformer 50 may process textual input data is now described. Input to a language model (whether transformer-based or otherwise) typically is in the form of natural language as may be parsed into tokens. It should be appreciated that the term “token” in the context of language models and NLP has a different meaning from the use of the same term in other contexts such as data security. Tokenization, in the context of languagemodels and NLP, refers to the process of parsing textual input (e.g., a character, a word, a phrase, a sentence, a paragraph, etc.) into a sequence of shorter segments that are converted to numerical representations referred to as tokens (or “compute tokens”). Typically, a token may be an integer that corresponds to the index of a text segment (e.g., a word) in a vocabulary dataset. Often, the vocabulary dataset is arranged by frequency of use. Commonly occurring text, such as punctuation, may have a lower vocabulary index in the dataset and thus be represented by a token having a smaller integer value than less commonly occurring text. Tokens frequently correspond to words, with or without whitespace appended. In some examples, a token may correspond to a portion of a word. For example, the word “lower” may be represented by a token for [low] and a second token for [er]. In another example, the text sequence “Come here, look!” may be parsed into the segments [Come], [here], [,], [look] and [!], each of which may be represented by a respective numerical token. In addition to tokens that are parsed from the textual sequence (e.g., tokens that correspond to words and punctuation), there may also be special tokens to encode non-textual information. For example, a [CLASS] token may be a special token that corresponds to a classification of the textual sequence (e.g., may classify the textual sequence as a poem, a list, a paragraph, etc.), a [EOT] token may be another special token that indicates the end of the textual sequence, other tokens may provide formatting information, etc.
[0071] In FIG. IB, a short sequence of tokens 56 corresponding to the text sequence “Come here, look!” is illustrated as input to the transformer 50. Tokenization of the text sequence into the tokens 56 may be performed by some pre-processing tokenization module such as, for example, a byte pair encoding tokenizer (the “pre” referring to the tokenization occurring prior to the processing of the tokenized input by the LLM), which is not shown in FIG. IB for simplicity. In general, the token sequence that is inputted to the transformer 50 may be of any length up to a maximum length defined based on the dimensions of the transformer 50 (e.g., such a limit may be 2048 tokens in some LLMs). Each token 56 in the token sequence is converted into an embedding vector 60 (also referred to simply as an embedding). An embedding 60 is a learned numerical representation (such as, for example, a vector) of a token that captures some semantic meaning of the text segment represented by the token 56. The embedding 60 represents the text segment corresponding to the token 56 in a way such that embeddings corresponding to semantically -related text are closer to each other in a vector space than embeddings corresponding to semantically -unrelated text. For example, assuming that the words “look”,“see”, and “cake” each correspond to, respectively, a “took” token, a “see” token, and a “cake” token when tokenized, the embedding 60 corresponding to the “took” token will be closer to another embedding corresponding to the “see” token in the vector space, as compared to the distance between the embedding 60 corresponding to the “took” token and another embedding corresponding to the “cake” token. The vector space may be defined by the dimensions and values of the embedding vectors. Various techniques may be used to convert a token 56 to an embedding 60. For example, another trained ML model may be used to convert the token 56 into an embedding 60. In particular, another trained ML model may be used to convert the token 56 into an embedding 60 in a way that encodes additional information into the embedding 60 (e.g., a trained ML model may encode positional information about the position of the token 56 in the text sequence into the embedding 60). In some examples, the numerical value of the token 56 may be used to look up the corresponding embedding in an embedding matrix 58 (which may be learned during training of the transformer 50).
[0072] The generated embeddings 60 are input into the encoder 52. The encoder 52 serves to encode the embeddings 60 into feature vectors 62 that represent the latent features of the embeddings 60. The encoder 52 may encode positional information (i.e., information about the sequence of the input) in the feature vectors 62. The feature vectors 62 may have very high dimensionality (e.g., on the order of thousands or tens of thousands), with each element in a feature vector 62 corresponding to a respective feature. The numerical weight of each element in a feature vector 62 represents the importance of the corresponding feature. The space of all possible feature vectors 62 that can be generated by the encoder 52 may be referred to as the latent space or feature space.
[0073] Conceptually, the decoder 54 is designed to map the features represented by the feature vectors 62 into meaningful output, which may depend on the task that was assigned to the transformer 50. For example, if the transformer 50 is used for a translation task, the decoder 54 may map the feature vectors 62 into text output in a target language different from the language of the original tokens 56. Generally, in a generative language model, the decoder 54 serves to decode the feature vectors 62 into a sequence of tokens. The decoder 54 may generate output tokens 64 one by one. Each output token 64 may be fed back as input to the decoder 54 in order to generate the next output token 64. By feeding back the generated output and applying selfattention, the decoder 54 is able to generate a sequence of output tokens 64 that has sequential meaning (e.g., the resulting output text sequence is understandable as a sentence and obeysgrammatical rules). The decoder 54 may generate output tokens 64 until a special [EOT] token (indicating the end of the text) is generated. The resulting sequence of output tokens 64 may then be converted to a text sequence in post-processing. For example, each output token 64 may be an integer number that corresponds to a vocabulary index. By looking up the text segment using the vocabulary index, the text segment corresponding to each output token 64 can be retrieved, the text segments can be concatenated together and the final output text sequence (in this example, “Viens ici, regarde!”) can be obtained.
[0074] Although a general transformer architecture for a language model and its theory of operation have been described above, this is not intended to be limiting. Existing language models include language models that are based only on the encoder of the transformer or only on the decoder of the transformer. An encoder-only language model encodes the input text sequence into feature vectors that can then be further processed by a task-specific layer (e.g., a classification layer). BERT is an example of a language model that may be considered to be an encoder-only language model. A decoder-only language model accepts embeddings as input and may use auto-regression to generate an output text sequence. Transformer-XL and GPT-type models may be language models that are considered to be decoder-only language models.
[0075] Because GPT-type language models tend to have a large number of parameters, these language models may be considered LLMs. An example GPT-type LLM is GPT-3. GPT-3 is a type of GPT language model that has been trained (in an unsupervised manner) on a large corpus derived from documents available to the public online. GPT-3 has a very large number of learned parameters (on the order of hundreds of billions), is able to accept a large number of tokens as input (e.g., up to 2048 input tokens), and is able to generate a large number of tokens as output (e.g., up to 2048 tokens). GPT-3 has been trained as a generative model, meaning that it can process input text sequences to predictively generate a meaningful output text sequence. ChatGPT is built on top of a GPT-type LLM, and has been fine-tuned with training datasets based on text-based chats (e.g., chatbot conversations). ChatGPT is designed for processing natural language, receiving chat-like inputs and generating chat-like outputs.
[0076] A computing system may access a remote language model (e.g., a cloud-based language model), such as ChatGPT or GPT-3, via a software interface (e.g., an application programming interface (API)). Additionally or alternatively, such a remote language model may be accessed via a network such as, for example, the Internet. In some implementations such as,for example, potentially in the case of a cloud-based language model, a remote language model may be hosted by a computer system as may include a plurality of cooperating (e.g., cooperating via a network) computer systems such as may be in, for example, a distributed arrangement. Notably, a remote language model may employ a plurality of processors (e.g., hardware processors such as, for example, processors of cooperating computer systems). Indeed, processing of inputs by an LLM may be computationally expensive / may involve a large number of operations (e.g., many instructions may be executed / large data structures may be accessed from memory) and providing output in a required timeframe (e.g., real-time or near real-time) may require the use of a plurality of processors / cooperating computing devices as discussed above.
[0077] Inputs to an LLM may be referred to as a prompt, which is a natural language input that includes instructions to the LLM to generate a desired output. A computing system may generate a prompt that is provided as input to the LLM via its API. As described above, the prompt may optionally be processed or pre-processed into a token sequence prior to being provided as input to the LLM via its API. A prompt can include one or more examples of the desired output, which provides the LLM with additional information to enable the LLM to better generate output according to the desired output. Additionally or alternatively, the examples included in a prompt may provide inputs (e.g., example inputs) corresponding to / as may be expected to result in the desired outputs provided. A one-shot prompt refers to a prompt that includes one example, and a few-shot prompt refers to a prompt that includes multiple examples. A prompt that includes no examples may be referred to as a zero-shot prompt.
[0078] FIG. 2 illustrates an example computing system 200, which may be used to implement examples of the present disclosure. For example, the computing system 200 may be used to generate a prompt to an LLM to cause the LLM to generate output that includes text in a tokenefficient language as disclosed herein. Additionally or alternatively, one or more instances of the example computing system 200 may be employed to execute the LLM. For example, a plurality of instances of the example computing system 200 may cooperate to provide output using an LLM in manners as discussed above.
[0079] The example computing system 200 includes at least one processing unit and at least one physical memory 204. The processing unit may be a hardware processor 202 (simply referred to as processor 202). The processor 202 may be, for example, a central processing unit,a microprocessor, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a dedicated logic circuitry, a dedicated artificial intelligence processor unit, a graphics processing unit (GPU), a tensor processing unit (TPU), a neural processing unit (NPU), a hardware accelerator, or combinations thereof. The memory 204 may include a volatile or non-volatile memory (e.g., a flash memory, a random access memory (RAM), and / or a read-only memory (ROM)). The memory 204 may store instructions for execution by the processor 202, to cause the computing system 200 to carry out examples of the methods, functionalities, systems and modules disclosed herein.
[0080] The computing system 200 may also include at least one network interface 206 for wired and / or wireless communications with an external system and / or network (e.g., an intranet, the Internet, a P2P network, a WAN and / or a LAN). The network interface 206 may enable the computing system 200 to carry out communications (e.g., wireless communications) with systems external to the computing system 200, such as a LLM residing on a remote system.
[0081] The computing system 200 may optionally include at least one input / output (I / O) interface 208, which may interface with optional input device(s) 210 and / or optional output device(s) 212. Input device(s) 210 may include, for example, buttons, a microphone, a touchscreen, a keyboard, etc. Output device(s) 212 may include, for example, a display, a speaker, etc. In this example, optional input device(s) 210 and optional output device(s) 212 are shown external to the computing system 200. In other examples, one or more of the input device(s) 210 and / or output device(s) 212 may be an internal component of the computing system 200.
[0082] A computing system, such as the computing system 200 of FIG. 2, may access a remote system (e.g., a cloud-based system) to communicate with a remote language model or LLM hosted on the remote system such as, for example, using an application programming interface (API) call. The API call may include an API key to enable the computing system to be identified by the remote system. The API call may also include an identification of the language model or LLM to be accessed and / or parameters for adjusting outputs generated by the language model or LLM, such as, for example, one or more of a temperature parameter (which may control the amount of randomness or “creativity” of the generated output) (and / or, more generally some form of random seed as serves to introduce variability or variety into the output of the LLM), a minimum length of the output (e.g., a minimum of 10 tokens) and / or a maximumlength of the output (e.g., a maximum of 1000 tokens), a frequency penalty parameter (e.g., a parameter which may lower the likelihood of subsequently outputting a word based on the number of times that word has already been output), a “best of’ parameter (e.g., a parameter to control the number of times the model will use to generate output after being instructed to, e.g., produce several outputs based on slightly varied inputs). The prompt generated by the computing system is provided to the language model or LLM and the output (e.g., token sequence) generated by the language model or LLM is communicated back to the computing system. In other examples, the prompt may be provided directly to the language model or LLM without requiring an API call. For example, the prompt could be sent to a remote LLM via a network such as, for example, as or in message (e.g., in a payload of a message).
[0083] In the example of FIG. 2, the computing system 200 may store in the memory 204 computer-executable instructions, which may be executed by a processing unit such as the processor 202, to implement one or more embodiments disclosed herein. For example, the memory 204 may store instructions for implementing a chatbot engine 250, which may include a chatbot UI 252 and a chatbot backend 254, as discussed further below.
[0084] In some examples, the computing system 200 may be a server of an online platform that provides the chatbot engine 250 as a web-based or cloud-based service that may be accessible by a user device (e.g., via communications over a wireless network). Other such variations may be possible without departing from the subject matter of the present application.
[0085] As will be discussed further below, the present disclosure describes an example chatbot engine that provides contextual information (in particular information about a navigational context) to a LLM prior to prompting the LLM for generate output (e.g., an operation command) for performing a requested task.
[0086] FIGS. 3A-3C are signalling diagrams that illustrate example communications performed by the example chatbot engine 250. FIGS. 3A-3C illustrate computing components including the chatbot UI 252, chatbot backend 254, a LLM 260 and a set of system tools 270, each of which may be implemented using any suitable hardware and / or software.
[0087] The chatbot UI 252 and chatbot backend 254 may be part of the chatbot engine 250 implemented by the computing system 200. The LLM 260 may be hosted by a remote system external to the computing system 200, and the chatbot backend 254 may send prompts to the LLM 260 via API calls, for example. The system tools 270 may provide functions accessible bythe chatbot engine 250, such as navigation. In some examples, the chatbot engine 250 and system tools 270 may be hosted on a common software platform (e.g., a common software as a service (SaaS) platform) and the chatbot backend 254 may send function calls to the system tools 270, for example.
[0088] The LLM 260 may be pre-trained or provided with navigational information in an initial prompt to enable the LLM 260 to generate appropriate navigation commands. For example, the LLM 260 may, prior to the signalling illustrated in FIGS. 3A-3C, be provided with an initial prompt that includes an outline of the navigation menu and / or valid navigation paths. An example initial prompt providing navigational information to the LLM 260 may be as follows:The structure of the typical left navigation menu may be helpful in getting a sense of the available URLs that you can navigate to (each entry is the path that you should use):- (empty string: home)- orders: draft_orders, checkouts- products: products / inventory- catalogs (product catalogs)- reportsThere are also many settings available. Each of the following are available under the path "settings" (e.g. settings / general). Some have sub-paths (like settings / domains / buy).* general (Store Details)* plan* billing (and sub-paths history). Edit credit cards for plan billing.* account (Users and permissions)Here are non-admin resources you can navigate to if a user wants.Changelog: changelog.website.comHelp center: help.website.com
[0089] In some examples, the LLM 260 may be provided with such navigational information as part of the context prompt (discussed below) or in the first instance of the context prompt in asession.
[0090] FIG. 3A is discussed first. In this example, the user has navigated to a current page in a website (e.g., using a browser) that provides the chatbot UI 252. As previously mentioned, navigation among pages of a website is only provided as an example and is not intended to be limiting. The present disclosure encompasses navigation among pages or sub-pages of a website or portal, navigation among panels, windows or tabs of an application, among other possibilities. For example, the chatbot UI 252 may be provided in an application or portal, among other possibilities. When the user navigates to the current page, the chatbot backend 254 performs operation 302 to extract contextual information from the current page. Contextual information may be coded into each page, and may provide contextual information about the page (e.g., functionality available on the page, name of the page, etc.) as well as other contextual information (e.g., date and time, user’s authorization level, user’s preferred language, etc.). Contextual information may be dynamic that varies over time (e.g., current date and time) as well as static information that is not expected to vary (e.g., name of the page), and may include information other than information specific to the current page (e.g., user information, system information). In some examples, the contextual information may be obtained indirectly from the page (e.g., using information on the page to reference another source for the contextual information). For example, contextual information that is static may be obtained by referring to a lookup table (or other reference document, reference database or reference system) based on the content of the current page. In general, the contextual information may include any information that is relevant to the functionality of the page. The following is an example of code that may be embedded in a page and that may be extracted by the chatbot backend 254 to obtain contextual information:# Current Date / Time{{ now }}# Page Context{%- for item in page %}{{ item[0] }}: {{ item[l] }}{%- endfor %}{% if themeName and themeVersion %}# Theme{ {themeName} } { { theme V ersion} }{% endif %}
[0091] The chatbot backend 254 sends a communication 304 (e.g., API call) that provides a context prompt to the LLM 260, where the context prompt provides the extracted contextual information. It may be noted that the contextual information that is embedded in the page may be designed to provide sufficient and relevant information for the page (e.g., sufficient and relevant information to enable the functionality available on the page). This means that the context prompt provided to the LLM 260 provides sufficient and relevant information to enable the LLM 260 to understand the current context, and without providing excessive or irrelevant information. The context prompt thus provides information that is tailored to the current context (e.g., current page of the website), rather than containing generic information about the website and functions of all pages on the website. In this way, the context prompt provides suitable contextual information to the LLM 260 so that the LLM 260 is not misdirected and / or does not waste tokens generating unnecessary output. As previously mentioned, in some examples the context prompt (or the first instance of a context prompt in a session with the LLM 260) may include navigational information, if the LLM 260 has not already been provided with the navigational information in an initial prompt.
[0092] The chatbot UI 252 receives user input indicating a task request (e.g., a natural language task request such as “I want to change my preferred language”). The task request is passed to the chatbot backend 254 in a communication 306. The chatbot backend 254 generates a check context prompt instructing the LLM 260 to determine whether the current page is a suitable context for performing the requested task, and sends the check context prompt in a communication 308 (e.g., another API call) to the LLM 260. An example check context prompt for the task request “I want to change my preferred language” is as follows:Much of your context and documentation comes from the page that the user is currently looking at. So, if the user asks you to do something related to a different page, you should navigate there first to get all the required information before completing the request.Do not ask the user if they want to navigate, output a navigate command instead.If you do not need to navigate, you can confirm no navigation is needed.User asks: I want to change my preferred language
[0093] FIG. 3B illustrates an example in which the current page does not provide a suitable context for the requested task (e.g., the current page does not provide the functionality required to perform the requested task, or if multiple functionalities are required the current page does not provide any of the functionalities required to perform the requested task). In this example, in response to the context check prompt, the LLM 260 sends a communication 322 of an output that includes a navigation command to navigate to a target page that does provide suitable context for the requested task. The navigation command included in the output from the LLM 260 may be in the form of code (e.g., JSON code, XML code, etc.) that can be parsed and processed by the chatbot backend 254. In some examples, the LLM 260 may be configured (e.g., pre-trained) to generate output using a token-efficient code (e.g., a specialized, proprietary token-efficient code) that can be parsed and processed by the chatbot backend 254 into a conventional coding language. The chatbot backend 254 sends a communication 324 to the chatbot UI 252 including the navigation command, which causes the chatbot UI 252 to present a navigation option (e.g., a soft button presented in the chatbot UI 252) to the user.
[0094] If the user selects the navigation option, then the chatbot UI 252 sends a communication 326 to the chatbot backend 254 indicating the user input (e.g., user selection of the navigation option) confirming the navigation command. The chatbot backend 254 then sends a communication 328 (e.g., function call) to the system tools 270 to execute the navigation command (e.g., the chatbot backend 254 may make a function call using the navigation command generated by the LLM 260). The system tools 270 performs operation 330 to navigate the browser and the chatbot UI 252 to the target page.
[0095] After navigating to the target page, again the chatbot backend 254 automatically performs operation 332 to extract contextual information from the current page (which is now the target page) and sends a communication 334 to the LLM 260 to provide the contextual information in a context prompt, as previously described. Without requiring the user to input the task request again, the chatbot backend 354 also provides, in a communication 336 to the LLM 260, the check context prompt instructing the LLM 260 to check whether the current page is a suitable context for performing the requested task. For example, the chatbot backend 354 may maintain a history of outstanding tasks, which may cause the chatbot backend 354 to automatically send the check context prompt to check if the current page is a suitable context foran outstanding task, without requiring the user to input a previously inputted task request. In some examples, instead of or in addition to the chatbot backend 354 maintaining a history of outstanding tasks, the previous output from the LLM 260 (i.e., the previous output including the navigation command) may also include a reissue command that causes the chatbot backend 354 to, after the navigation command is executed, automatically reissue the previous check context prompt (that includes the previously requested task) or the navigation command generated by the LLM 260 may be a single navigate-and-reissue command.
[0096] If the current page is not a suitable context for the requested task (e.g., the user has manually navigated to a different page other than the target page), then the signalling illustrated in FIG. 3B may be repeated. This may be performed multiple times (e.g., until the chatbot UI 252 has finally been navigated to the target page with the suitable context for the requested task, or until some maximum number is reached).
[0097] FIG. 3C illustrates an example in which the chatbot UI 252 has been navigated to the target page providing a suitable context for the requested task. In this example, in response to the context check prompt, the LLM 260 sends a communication 342 of an output that includes a confirmation that the page offers a suitable context for performing the requested task. The chatbot backend 254 sends a communication 344 that provides a task prompt to the LLM 260 instructing the LLM 260 to generate operation command(s) for performing the requested task. In some examples, the chatbot backend 254 may be configured to automatically generate the task prompt including the previously requested task in response to receiving the communication 342 including the confirmation from the LLM 260. In some examples, the confirmation from the LLM 260 may include a command to the chatbot backend 254 to send the task prompt.
[0098] The task prompt may, in some examples, include a task history (e.g., the history of the most recent inputs / outputs of the chatbot UI 252, or the history of the most recent communications between the chatbot backend 254 and the LLM 260). For example, the chatbot engine 250 may maintain the most recent task history by storing all communications in a first-in first-out (FIFO) local memory. The task history may then be retrieved from the local memory and inserted into the task prompt. In some examples the task history may be retrieved from a task history database (e.g., using retrieval augmented generation (RAG) techniques). For example, a historical portion of the text in a chatbot session may be encoded as an embedding vector (e.g., using any suitable embedding encoder, such as the encoder 52 described previously) and thehistorical text may be stored as a task history portion with the corresponding embedding in the task history database. Later, another embedding may be generated (e.g., using any suitable embedding encoder, such as the encoder 52 described previously) from a current text in a chatbot session and used as a query embedding to search the task history database. The stored embedding that is most similar to the query embedding is identified and the task history portion corresponding to the identified stored embedding is retrieved. The retrieved task history portion may then be included in the task prompt to the LLM 260. The task history database may be stored locally at the chatbot engine 250 or be otherwise accessible by the chatbot engine 250 (e.g., may be a component of the system tools 270 that may be accessed by a function call). In other examples, the task history database may be maintained remotely. For simplicity, the task history database as well as the communications and operations of the chatbot backend 254 and / or system tools 270 to query the task history database are not illustrated in FIG. 3C.
[0099] In response to the task prompt, the LLM 260 sends, in a communication 346 to the chatbot backend 254, an output including an operation command for performing the requested task (e.g., an operation command that uses a functionality of the current page). In some examples, the LLM 260 may be configured (e.g., pre-trained) to generate the operation command using a particular grammar and / or a particular language that may be more computer-efficient (e.g., using fewer tokens for the output), for example using a specialized token-efficient coding language rather than conventional coding language. In some examples, the output from the LLM 260 may be a request for more information (e.g., if the original task request was to “change my theme”, the output may be a request for more information such as “how do you want to change your theme?”). In some examples, the output from the LLM 260 may be a request for a task history (e.g., if the task request is “undo yesterday’s change”, the output may be a request or command to the chatbot backend 254 to retrieve the relevant task history from the task history database, for example using a query embedding as described above).
[0100] The chatbot backend 254 sends a communication 348 to the chatbot UI 252 including the operation command, which causes the chatbot UI 252 to present an operation option and / or a preview option (e.g., as a soft button presented in the chatbot UI 252) to the user. If the user selects the preview option, a preview of the result of executing the operation command is provided (e.g., displayed within the chatbot UI 252). The preview may be generated using functions of the system tools 270. For example, a function call may be made to the system tools 270 to execute the operation command but not to commit to the change. If the user selects theoperation option, then the chatbot UI 252 sends a communication 350 to the chatbot backend 254 indicating the user input (e.g., user selection of the operation option) confirming the operation command. The chatbot backend 254 then sends a communication 360 (e.g., function call) to the system tools 270 to execute the operation command (e.g., the chatbot backend 254 may make a function call using the operation command generated by the LLM 260). The system tools 270 then execute the operation command at 362.
[0101] The signalling described above and shown in FIGS. 3A-3C are only exemplary and are not intended to be limiting. As well, there may be greater or fewer computing components involved in enabling contextual operation of the chatbot engine 250, as disclosed herein.
[0102] FIG. 4 is a flowchart of an example method 400 for an example embodiment of the present disclosure, which may be performed by a computing system, in accordance with examples of the present disclosure. For example, a processing unit of a computing system (e.g., the processor 202 of the computing system 200 of FIG. 2) may execute instructions (e.g., instructions of the chatbot engine 250) to cause the computing system to carry out the example method 400. The method 400 may, for example, be implemented by an online platform or a server. The method 400 may enable contextual chatbot operation, using a LLM. The LLM may be a generative pre-trained transformer LLM, such as LLaMA, Falcon 40B, GPT-3, GPT-4 or ChatGPT, among others. It should be understood that the present disclosure is not intended to be limited to any particular LLM. The operations of the chatbot UI 252 and chatbot backend 254 in the example of FIGS. 3A-3C may illustrate an example implementation of the method 400.
[0103] The LLM may be configured (e.g., pre-trained or primed by an initial prompt) to be able to generate valid navigational commands. For example, the LLM may be pre-trained or initially prompted with information about the navigational structure or options available and / or how to structure a navigational command.
[0104] At an operation 402, contextual information is extracted from the current page. For example, contextual information may be coded into the current page and automatically extracted (e.g., by a chatbot backend of the chatbot engine) when the UI (e.g., a chatbot UI of the chatbot engine) is navigated to the current page. As previously mentioned, the contextual information may include dynamic information (e.g., current date and time) and / or static information (e.g., name of current page), and may include information specific to the current page (e.g., functionality available on current page) as well as other information (e.g., authorization level ofuser). The contextual information may be automatically extracted from the current page by executing code embedded in the page (e.g., the code may be executed as part of operations for rendering the page).
[0105] Optionally, at an operation 404, a context prompt is generated using the contextual information extracted from the page. In some examples, the context prompt may be included in the contextual information extracted from the page and may not need to be separately generated.
[0106] At an operation 406, the context prompt is provided to the LLM. The context prompt includes the contextual information, including information about the page (such as the functionality available on the page). For example, the context prompt may be provided to a remote LLM via an API call to a remote system hosting the LLM. The context prompt may be converted to a set of tokens (e.g., by a tokenizer, which may be specific to the LLM) and the set of tokens may be provided to the LLM in sequential order. In some examples, tokenization of the context prompt may occur prior to the API call, such that the set of tokens is provided via the API call.
[0107] The operations 402-406 may occur automatically when the UI is navigated to the page, without requiring explicit input from the user. The operations 402-406 may occur without the user’s knowledge.
[0108] Optionally, at an operation 408, a task request may be received via user input to the UI. If there is an outstanding task (e.g., the UI was navigated to the current page in order to provide a suitable context for a previously requested task), the method 400 may proceed to an operation 410 without receiving a task request via user input.
[0109] At the operation 410, a check context prompt is provided to the LLM (e.g., via an API call as described above). The check context prompt instructs the LLM to determine if the current context is a suitable context for performing a task (such as the task requested at the operation 408 or a previously requested outstanding task). As described previously, the check context prompt may include the requested task as well as instructions to the LLM to confirm the suitability of the current context or generate a navigation command to navigate to a suitable context if the current context is unsuitable.
[0110] If the current page provides a suitable context for the requested task, then the method400 proceeds to an operation 412.
[0111] At the operation 412, output is received from the LLM confirming that the current page is a suitable context for performing the requested task (e.g., the current page provides a functionality for performing the requested task). In some examples, the output from the LLM may be a request for additional information. The additional information may require further user input, such as to clarify the task request, and the UI may output to the user a request for the additional information. In some examples, the output from the LLM may be a request for task history. The requested task history may be extracted from within the current LLM session or may be retrieved from a task history database at an optional operation 414.
[0112] Optionally, at the operation 414, a historical LLM session portion may be retrieved from a task history database. For example, a query embedding may be generated from the most recent communications with the LLM and the query embedding may be used to search a task history database for a similar stored embedding. After identifying the most similar stored embedding, the historical LLM session portion associated with that stored embedding may be retrieved. It should be noted that the operation 414 may be performed with or without a request for additional information from the LLM.
[0113] At an operation 416, a task prompt is generated and provided to the LLM (e.g., via an API call). The task prompt may be automatically provided in response to receiving confirmation from the LLM that the current page is a suitable context for the requested task. The task prompt includes instructions to the LLM to generate an operation command for performing the requested task (e.g., using a functionality of the page). If the operation 414 was performed to retrieve a historical LLM session portion, the retrieved historical LLM session portion may be included in the task prompt.
[0114] At an operation 418, output from the LLM including an operation command for performing the task (e.g., using the functionality of the page) is received. As previously mentioned, the operation command may be represented in a token-efficient coding language, which may be parsed and converted into a conventional coding language before being executed.
[0115] Optionally, at an operation 420, a user selectable operation option may be provided in the UI (e.g., in the form of a one-click soft button) for executing the operation command. In response to selection of the operation option, the operation command is executed. In some examples, the operation option may also provide a preview of the result of executing the operation command.
[0116] Returning to the operation 410, if the current page does not provide a suitable context for the requested task, then the method 400 proceeds to an operation 422.
[0117] At the operation 422, output is received from the LLM including a navigation command to navigate to a target page (e.g., another page that provides a suitable context). As previously mentioned, the navigation command may be represented in a token-efficient coding language, which may be parsed and converted into a conventional coding language.
[0118] Optionally, at an operation 424, a user selectable operation option may be provided in the UI (e.g., in the form of a one-click soft button) for executing the navigation command. In response to selection of the navigation option, the navigation command is executed. The UI is navigated to the target page and the method 400 returns to the operation 402.
[0119] It should be noted that, even if the UI is not navigated to the target page (e.g., instead of selecting the navigation option in the UI, the user manually navigates to a different page), the method 400 returns to the operation 402. This means that, regardless of how the user navigates, contextual information about the page that the UI is currently at is extracted and provided to the LLM. In this way, the LLM is always provided with current and relevant information about the current context of the user, to enable the LLM to generate appropriate output and avoid invalid or spurious output, thus avoiding unnecessary consumption of computing resources (e.g., memory resources, processing power, tokens, etc.).
[0120] FIGS. 5A-5C illustrate an example of a simplified chatbot UI, which may be implemented by an example of the chatbot engine as disclosed herein (e.g., using the example method 400). It should be understood that this example is not intended to be limiting.
[0121] In this simple example, a user is viewing and navigating through an administrative portal 70 that has multiple pages or tabs, as indicated in the navigation bar 72.
[0122] A chatbot UI 500 (e.g., provided by the disclosed chatbot engine) is presented to the user. The chatbot UI 500 includes a chat history portion 502 displaying the most recent inputs and outputs in the chat history and an input portion 504 in which the user may enter text input, such as a task request. In some examples, the user may provide input by other means, such as voice input and / or touch input.
[0123] In FIG. 5 A, the user has navigated to an Accounts page 74. When the user navigated to the Accounts page 74, the chatbot engine automatically extracted contextual information fromthe Accounts page 74 and provided this contextual information to a LLM via a context prompt.
[0124] The user has provided a natural language task request 512 in the chatbot UI 500. In response to receiving the task request 512 via user input, the chatbot engine generates and provides a check context prompt to the LLM to check whether the current context (e.g., the current Accounts page 74) is suitable for performing the requested task. In this example, the current context is not suitable and the response from the LLM is a navigation command to navigate to the payments page. The chatbot UI 500 presents a response 514 indicating that the user should navigate to the payments page. Notably, instead of automatically navigating the user to the payments page (which may be confusing to the user and removes the user’s agency) or requiring the user to use the navigation bar 72 to find and select the payments page (which places a greater burden on the user, may be inconvenient and assumes the user has knowledge to use the navigation bar 72), the chatbot UI 500 presents a selectable navigation option 516 that the user can conveniently click to perform the navigation operation.
[0125] In this example, the user selects the navigation option 516, causing the navigation command to be executed and the chatbot UI 500 is navigated to the Payments page 76, as shown in FIG. 5B. The selection of the navigation option 516 may be indicated by visually updating the navigation option 516 to include a checkmark (or other visual indicator).
[0126] When the chatbot UI 500 was navigated to the Payments page 76, the chatbot engine again automatically extracted contextual information from the Payments page 76 and provided this contextual information to a LLM via another context prompt. Since the user’s requested task is still outstanding, the chatbot engine automatically sends the check context prompt again to the LLM. The output from the LLM confirms that the current context is suitable for the requested task. Accordingly, the chatbot engine automatically sends a task prompt to the LLM instructing the LLM to generate an operation command for performing the requested task, and the operation command is received as output from the LLM. The chatbot UI 500 presents a response 518 including a selectable operation option 520, which the user may click in order to execute the operation command generated by the LLM. Again, instead of automatically performing the operation, presentation of the selectable operation option 520 gives the user greater agency and allows the user an opportunity to check if the operation is actually what is desired (which may help to reduce errors, such as when the LLM misinterprets the task request). In some examples, the chatbot UI 500 may, instead of or in addition to the selectable operation option 520, present apreview option that the user may select to cause a preview of the operation to be shown (e.g., a preview of what the Payments page 76 would look like after performing the operation).
[0127] In this example, the user selects the operation option 520, causing the operation command to be executed, which updates the Payments page 76, as shown in FIG. 5C. The selection of the operation option 520 may be indicated by visually updating the operation option 520 to include a checkmark (or other visual indicator). The chatbot UI 500 may provide further output 522 (e.g., asking if the user needs further help).
[0128] Examples of the present disclosure may enable contextual chatbot operation, where the chatbot may help a user with various tasks. A chatbot engine as disclosed herein may be used in various implementations, such as on a website, a portal, a software application, etc. In an example, the disclosed chatbot engine may be implemented on an e-commerce platform, for example to help a user (e.g., a merchant, store owner or store employee) with tasks on an administrative webpage or portal of an online store (e.g., as shown in FIG. 7).An example e-commerce platform
[0129] Although integration with a commerce platform is not required, in some embodiments, the methods disclosed herein may be performed on or in association with a commerce platform such as an e-commerce platform. Therefore, an example of a commerce platform will be described.
[0130] FIG. 6 illustrates an example e-commerce platform 100, according to one embodiment. The e-commerce platform 100 may be used to provide merchant products and services to customers. While the disclosure contemplates using the apparatus, system, and process to purchase products and services, for simplicity the description herein will refer to products. All references to products throughout this disclosure should also be understood to be references to products and / or services, including, for example, physical products, digital content (e.g., music, videos, games), software, tickets, subscriptions, services to be provided, and the like.
[0131] While the disclosure throughout contemplates that a ‘merchant’ and a ‘customer’ may be more than individuals, for simplicity the description herein may generally refer to merchants and customers as such. All references to merchants and customers throughout this disclosure should also be understood to be references to groups of individuals, companies, corporations,computing entities, and the like, and may represent for-profit or not-for-profit exchange of products. Further, while the disclosure throughout refers to ‘merchants’ and ‘customers’, and describes their roles as such, the e-commerce platform 100 should be understood to more generally support users in an e-commerce environment, and all references to merchants and customers throughout this disclosure should also be understood to be references to users, such as where a user is a merchant-user (e.g., a seller, retailer, wholesaler, or provider of products), a customer-user (e.g., a buyer, purchase agent, consumer, or user of products), a prospective user (e.g., a user browsing and not yet committed to a purchase, a user evaluating the e-commerce platform 100 for potential use in marketing and selling products, and the like), a service provider user (e.g., a shipping provider 112, a financial provider, and the like), a company or corporate user (e.g., a company representative for purchase, sales, or use of products; an enterprise user; a customer relations or customer management agent, and the like), an information technology user, a computing entity user (e.g., a computing bot for purchase, sales, or use of products), and the like. Furthermore, it may be recognized that while a given user may act in a given role (e.g., as a merchant) and their associated device may be referred to accordingly (e.g., as a merchant device) in one context, that same individual may act in a different role in another context (e.g., as a customer) and that same or another associated device may be referred to accordingly (e.g., as a customer device). For example, an individual may be a merchant for one type of product (e.g., shoes), and a customer / consumer of other types of products (e.g., groceries). In another example, an individual may be both a consumer and a merchant of the same type of product. In a particular example, a merchant that trades in a particular category of goods may act as a customer for that same category of goods when they order from a wholesaler (the wholesaler acting as merchant).
[0132] The e-commerce platform 100 provides merchants with online services / facilities to manage their business. The facilities described herein are shown implemented as part of the platform 100 but could also be configured separately from the platform 100, in whole or in part, as stand-alone services. Furthermore, such facilities may, in some embodiments, may, additionally or alternatively, be provided by one or more providers / entities.
[0133] In the example of FIG. 6, the facilities are deployed through a machine, service or engine that executes computer software, modules, program codes, and / or instructions on one or more processors which, as noted above, may be part of or external to the platform 100.Merchants may utilize the e-commerce platform 100 for enabling or managing commerce with customers, such as by implementing an e-commerce experience with customers through anonline store 138, applications 142A-B, channels 110A-B, and / or through point of sale (POS) devices 152 in physical locations (e.g., a physical storefront or other location such as through a kiosk, terminal, reader, printer, 3D printer, and the like). A merchant may utilize the e-commerce platform 100 as a sole commerce presence with customers, or in conjunction with other merchant commerce facilities, such as through a physical store (e.g., ‘brick-and-mortar’ retail stores), a merchant off-platform website 104 (e.g., a commerce Internet website or other internet or web property or asset supported by or on behalf of the merchant separately from the e-commerce platform 100), an application 142B, and the like. However, even these ‘other’ merchant commerce facilities may be incorporated into or communicate with the e-commerce platform 100, such as where POS devices 152 in a physical store of a merchant are linked into the e- commerce platform 100, where a merchant off-platform website 104 is tied into the e-commerce platform 100, such as, for example, through ‘buy buttons’ that link content from the merchant off platform website 104 to the online store 138, or the like.
[0134] The online store 138 may represent a multi-tenant facility comprising a plurality of virtual storefronts. In embodiments, merchants may configure and / or manage one or more storefronts in the online store 138, such as, for example, through a merchant device 102 (e.g., computer, laptop computer, mobile computing device, and the like), and offer products to customers through a number of different channels 110A-B (e.g., an online store 138; an application 142A-B; a physical storefront through a POS device 152; an electronic marketplace, such, for example, through an electronic buy button integrated into a website or social media channel such as on a social network, social media page, social media messaging system; and / or the like). A merchant may sell across channels 110A-B and then manage their sales through the e-commerce platform 100, where channels 110A may be provided as a facility or service internal or external to the e-commerce platform 100. A merchant may, additionally or alternatively, sell in their physical retail store, at pop ups, through wholesale, over the phone, and the like, and then manage their sales through the e-commerce platform 100. A merchant may employ all or any combination of these operational modalities. Notably, it may be that by employing a variety of and / or a particular combination of modalities, a merchant may improve the probability and / or volume of sales. Throughout this disclosure the terms online store 138 and storefront may be used synonymously to refer to a merchant’s online e-commerce service offering through the e- commerce platform 100, where an online store 138 may refer either to a collection of storefronts supported by the e-commerce platform 100 (e.g., for one or a plurality of merchants) or to anindividual merchant’s storefront (e.g., a merchant’s online store).
[0135] In some embodiments, a customer may interact with the platform 100 through a customer device 150 (e.g., computer, laptop computer, mobile computing device, or the like), a POS device 152 (e.g., retail device, kiosk, automated (self-service) checkout system, or the like), and / or any other commerce interface device known in the art. The e-commerce platform 100 may enable merchants to reach customers through the online store 138, through applications 142A-B, through POS devices 152 in physical locations (e.g., a merchant’s storefront or elsewhere), to communicate with customers via electronic communication facility 129, and / or the like so as to provide a system for reaching customers and facilitating merchant services for the real or virtual pathways available for reaching and interacting with customers.
[0136] In some embodiments, and as described further herein, the e-commerce platform 100 may be implemented through a processing facility. Such a processing facility may include a processor and a memory. The processor may be a hardware processor. The memory may be and / or may include a non-transitory computer-readable medium. The memory may be and / or may include random access memory (RAM) and / or persisted storage (e.g., magnetic storage). The processing facility may store a set of instructions (e.g., in the memory) that, when executed, cause the e-commerce platform 100 to perform the e-commerce and support functions as described herein. The processing facility may be or may be a part of one or more of a server, client, network infrastructure, mobile computing platform, cloud computing platform, stationary computing platform, and / or some other computing platform, and may provide electronic connectivity and communications between and amongst the components of the e-commerce platform 100, merchant devices 102, payment gateways 106, applications 142A-B , channels 110A-B, shipping providers 112, customer devices 150, point of sale devices 152, etc. In some implementations, the processing facility may be or may include one or more such computing devices acting in concert. For example, it may be that a plurality of co-operating computing devices serves as / to provide the processing facility. The e-commerce platform 100 may be implemented as or using one or more of a cloud computing service, software as a service (SaaS), infrastructure as a service (laaS), platform as a service (PaaS), desktop as a service (DaaS), managed software as a service (MSaaS), mobile backend as a service (MBaaS), information technology management as a service (ITMaaS), and / or the like. For example, it may be that the underlying software implementing the facilities described herein (e.g., the online store 138) is provided as a service, and is centrally hosted (e.g., and then accessed by users via a web browseror other application, and / or through customer devices 150, POS devices 152, and / or the like). In some embodiments, elements of the e-commerce platform 100 may be implemented to operate and / or integrate with various other platforms and operating systems.
[0137] In some embodiments, the facilities of the e-commerce platform 100 (e.g., the online store 138) may serve content to a customer device 150 (using data 134) such as, for example, through a network connected to the e-commerce platform 100. For example, the online store 138 may serve or send content in response to requests for data 134 from the customer device 150, where a browser (or other application) connects to the online store 138 through a network using a network communication protocol (e.g., an internet protocol). The content may be written in machine readable language and may include Hypertext Markup Language (HTML), template language, JavaScript, and the like, and / or any combination thereof.
[0138] In some embodiments, online store 138 may be or may include service instances that serve content to customer devices and allow customers to browse and purchase the various products available (e.g., add them to a cart, purchase through a buy -button, and the like). Merchants may also customize the look and feel of their website through a theme system, such as, for example, a theme system where merchants can select and change the look and feel of their online store 138 by changing their theme while having the same underlying product and business data shown within the online store’s product information. It may be that themes can be further customized through a theme editor, a design interface that enables users to customize their website's design with flexibility. Additionally or alternatively, it may be that themes can, additionally or alternatively, be customized using theme-specific settings such as, for example, settings as may change aspects of a given theme, such as, for example, specific colors, fonts, and pre-built layout schemes. In some implementations, the online store may implement a content management system for website content. Merchants may employ such a content management system in authoring blog posts or static pages and publish them to their online store 138, such as through blogs, articles, landing pages, and the like, as well as configure navigation menus. Merchants may upload images (e.g., for products), video, content, data, and the like to the e- commerce platform 100, such as for storage by the system (e.g., as data 134). In some embodiments, the e-commerce platform 100 may provide functions for manipulating such images and content such as, for example, functions for resizing images, associating an image with a product, adding and associating text with an image, adding an image for a new product variant, protecting images, and the like.
[0139] As described herein, the e-commerce platform 100 may provide merchants with sales and marketing services for products through a number of different channels 110A-B, including, for example, the online store 138, applications 142A-B, as well as through physical POS devices 152 as described herein. The e-commerce platform 100 may, additionally or alternatively, include business support services 116, an administrator 114, a warehouse management system, and the like associated with running an on-line business, such as, for example, one or more of providing a domain registration service 118 associated with their online store, payment services 120 for facilitating transactions with a customer, shipping services 122 for providing customer shipping options for purchased products, fulfillment services for managing inventory, risk and insurance services 124 associated with product protection and liability, merchant billing, and the like. Services 116 may be provided via the e-commerce platform 100 or in association with external facilities, such as through a payment gateway 106 for payment processing, shipping providers 112 for expediting the shipment of products, and the like.
[0140] In some embodiments, the e-commerce platform 100 may be configured with shipping services 122 (e.g., through an e-commerce platform shipping facility or through a third-party shipping carrier), to provide various shipping-related information to merchants and / or their customers such as, for example, shipping label or rate information, real-time delivery updates, tracking, and / or the like.
[0141] FIG. 7 depicts a non-limiting embodiment for a home page of an administrator 114. The administrator 114 may be referred to as an administrative console and / or an administrator console. The administrator 114 may show information about daily tasks, a store’s recent activity, and the next steps a merchant can take to build their business. In some embodiments, a merchant may log in to the administrator 114 via a merchant device 102 (e.g., a desktop computer or mobile device), and manage aspects of their online store 138, such as, for example, viewing the online store’s 138 recent visit or order activity, updating the online store's 138 catalogue, managing orders, and / or the like. In some embodiments, the merchant may be able to access the different sections of the administrator 114 by using a sidebar, such as the one shown on FIG. 7. Sections of the administrator 114 may include various interfaces for accessing and managing core aspects of a merchant’s business, including orders, products, customers, available reports and discounts. The administrator 114 may, additionally or alternatively, include interfaces for managing sales channels for a store including the online store 138, mobile application(s) made available to customers for accessing the store (Mobile App), POS devices, and / or a buy button.The administrator 114 may, additionally or alternatively, include interfaces for managing applications (apps) installed on the merchant’s account; and settings applied to a merchant’s online store 138 and account. A merchant may use a search bar to find products, pages, or other information in their store.
[0142] More detailed information about commerce and visitors to a merchant’s online store 138 may be viewed through reports or metrics. Reports may include, for example, acquisition reports, behavior reports, customer reports, finance reports, marketing reports, sales reports, product reports, and custom reports. The merchant may be able to view sales data for different channels 110A-B from different periods of time (e.g., days, weeks, months, and the like), such as by using drop-down menus. An overview dashboard may also be provided for a merchant who wants a more detailed view of the store's sales and engagement data. An activity feed in the home metrics section may be provided to illustrate an overview of the activity on the merchant’s account. For example, by clicking on a ‘view all recent activity’ dashboard button, the merchant may be able to see a longer feed of recent activity on their account. A home page may show notifications about the merchant’s online store 138, such as based on account status, growth, recent customer activity, order updates, and the like. Notifications may be provided to assist a merchant with navigating through workflows configured for the online store 138, such as, for example, a payment workflow, an order fulfillment workflow, an order archiving workflow, a return workflow, and the like.
[0143] The e-commerce platform 100 may provide for a communications facility 129 and associated merchant interface for providing electronic communications and marketing, such as utilizing an electronic messaging facility for collecting and analyzing communication interactions between merchants, customers, merchant devices 102, customer devices 150, POS devices 152, and the like, to aggregate and analyze the communications, such as for increasing sale conversions, and the like. For instance, a customer may have a question related to a product, which may produce a dialog between the customer and the merchant (or an automated processorbased agent / chatbot representing the merchant), where the communications facility 129 is configured to provide automated responses to customer requests and / or provide recommendations to the merchant on how to respond such as, for example, to improve the probability of a sale.
[0144] The e-commerce platform 100 may provide a financial facility 120 for secure financialtransactions with customers, such as through a secure card server environment. The e-commerce platform 100 may store credit card information, such as in payment card industry data (PCI) environments (e.g., a card server), to reconcile financials, bill merchants, perform automated clearing house (ACH) transfers between the e-commerce platform 100 and a merchant’s bank account, and the like. The financial facility 120 may also provide merchants and buyers with financial support, such as through the lending of capital (e.g., lending funds, cash advances, and the like) and provision of insurance. In some embodiments, online store 138 may support a number of independently administered storefronts and process a large volume of transactional data on a daily basis for a variety of products and services, for example, in an analytics facility 132. Transactional data may include any customer information indicative of a customer, a customer account or transactions carried out by a customer such as. for example, contact information, billing information, shipping information, retums / refund information, discount / offer information, payment information, or online store events or information such as page views, product search information (search keywords, click-through events), product reviews, abandoned carts, and / or other transactional information associated with business through the e-commerce platform 100. In some embodiments, the e-commerce platform 100 may store this data in a data facility 134. Referring again to FIG. 6, in some embodiments the e-commerce platform 100 may include a commerce management engine 136 such as may be configured to perform various workflows for task automation or content management related to products, inventory, customers, orders, suppliers, reports, financials, risk and fraud, and the like. In some embodiments, additional functionality may, additionally or alternatively, be provided through applications 142A-B to enable greater flexibility and customization required for accommodating an evergrowing variety of online stores, POS devices, products, and / or services. Applications 142A may be components of the e-commerce platform 100 whereas applications 142B may be provided or hosted as a third-party service external to e-commerce platform 100. The commerce management engine 136 may accommodate store-specific workflows and in some embodiments, may incorporate the administrator 114 and / or the online store 138.
[0145] Implementing functions as applications 142A-B may enable the commerce management engine 136 to remain responsive and reduce or avoid service degradation or more serious infrastructure failures, and the like.
[0146] Although isolating online store data can be important to maintaining data privacy between online stores 138 and merchants, there may be reasons for collecting and using cross-store data, such as, for example, with an order risk assessment system or a platform payment facility, both of which require information from multiple online stores 138 to perform well. In some embodiments, it may be preferable to move these components out of the commerce management engine 136 and into their own infrastructure within the e-commerce platform 100.
[0147] Platform payment facility 120 is an example of a component that utilizes data from the commerce management engine 136 but is implemented as a separate component or service. The platform payment facility 120 may allow customers interacting with online stores 138 to have their payment information stored safely by the commerce management engine 136 such that they only have to enter it once. When a customer visits a different online store 138, even if they have never been there before, the platform payment facility 120 may recall their information to enable a more rapid and / or potentially less-error prone (e.g., through avoidance of possible mis-keying of their information if they needed to instead re-enter it) checkout. This may provide a crossplatform network effect, where the e-commerce platform 100 becomes more useful to its merchants and buyers as more merchants and buyers join, such as because there are more customers who checkout more often because of the ease of use with respect to customer purchases. To maximize the effect of this network, payment information for a given customer may be retrievable and made available globally across multiple online stores 138.
[0148] For functions that are not included within the commerce management engine 136, applications 142A-B provide a way to add features to the e-commerce platform 100 or individual online stores 138. For example, applications 142A-B may be able to access and modify data on a merchant’s online store 138, perform tasks through the administrator 114, implement new flows for a merchant through a user interface (e.g., that is surfaced through extensions / API), and the like. Merchants may be enabled to discover and install applications 142A-B through application search, recommendations, and support 128. In some embodiments, the commerce management engine 136, applications 142A-B, and the administrator 114 may be developed to work together. For instance, application extension points may be built inside the commerce management engine 136, accessed by applications 142A and 142B through the interfaces 140B and 140A to deliver additional functionality, and surfaced to the merchant in the user interface of the administrator 114.
[0149] In some embodiments, applications 142A-B may deliver functionality to a merchant through the interface 140A-B, such as where an application 142A-B is able to surface transactiondata to a merchant (e.g., App: “Engine, surface my app data in the Mobile App or administrator 114”), and / or where the commerce management engine 136 is able to ask the application to perform work on demand (Engine: “App, give me a local tax calculation for this checkout”).
[0150] Applications 142A-B may be connected to the commerce management engine 136 through an interface 140A-B (e.g., through REST (REpresentational State Transfer) and / or GraphQL APIs) to expose the functionality and / or data available through and within the commerce management engine 136 to the functionality of applications. For instance, the e- commerce platform 100 may provide API interfaces 140A-B to applications 142A-B which may connect to products and services external to the platform 100. The flexibility offered through use of applications and APIs (e.g., as offered for application development) enable the e-commerce platform 100 to better accommodate new and unique needs of merchants or to address specific use cases without requiring constant change to the commerce management engine 136. For instance, shipping services 122 may be integrated with the commerce management engine 136 through a shipping or carrier service API, thus enabling the e-commerce platform 100 to provide shipping service functionality without directly impacting code running in the commerce management engine 136.
[0151] Depending on the implementation, applications 142A-B may utilize APIs to pull data on demand (e.g., customer creation events, product change events, or order cancelation events, etc.) or have the data pushed when updates occur. A subscription model may be used to provide applications 142A-B with events as they occur or to provide updates with respect to a changed state of the commerce management engine 136. In some embodiments, when a change related to an update event subscription occurs, the commerce management engine 136 may post a request, such as to a predefined callback URL. The body of this request may contain a new state of the object and a description of the action or event. Update event subscriptions may be created manually, in the administrator facility 114, or automatically (e.g., via the API 140A-B). In some embodiments, update events may be queued and processed asynchronously from a state change that triggered them, which may produce an update event notification that is not distributed in real-time or near-real time.
[0152] In some embodiments, the e-commerce platform 100 may provide one or more of application search, recommendation and support 128. Application search, recommendation and support 128 may include developer products and tools to aid in the development of applications,an application dashboard (e.g., to provide developers with a development interface, to administrators for management of applications, to merchants for customization of applications, and the like), facilities for installing and providing permissions with respect to providing access to an application 142A-B (e.g., for public access, such as where criteria must be met before being installed, or for private use by a merchant), application searching to make it easy for a merchant to search for applications 142A-B that satisfy a need for their online store 138, application recommendations to provide merchants with suggestions on how they can improve the user experience through their online store 138, and the like. In some embodiments, applications 142A-B may be assigned an application identifier (ID), such as for linking to an application (e.g., through an API), searching for an application, making application recommendations, and the like.
[0153] Applications 142A-B may be grouped roughly into three categories: customer-facing applications, merchant-facing applications, integration applications, and the like. Customerfacing applications 142A-B may include an online store 138 or channels 110A-B that are places where merchants can list products and have them purchased (e.g., the online store, applications for flash sales) (e.g., merchant products or from opportunistic sales opportunities from third- party sources), a mobile store application, a social media channel, an application for providing wholesale purchasing, and the like). Merchant-facing applications 142A-B may include applications that allow the merchant to administer their online store 138 (e.g., through applications related to the web or website or to mobile devices), run their business (e.g., through applications related to POS devices), to grow their business (e.g., through applications related to shipping (e.g., drop shipping), use of automated agents, use of process flow development and improvements), and the like. Integration applications may include applications that provide useful integrations that participate in the running of a business, such as shipping providers 112 and payment gateways 106.
[0154] As such, the e-commerce platform 100 can be configured to provide an online shopping experience through a flexible system architecture that enables merchants to connect with customers in a flexible and transparent manner. A typical customer experience may be better understood through an embodiment example purchase workflow, where the customer browses the merchant’s products on a channel 110A-B, adds what they intend to buy to their cart, proceeds to checkout, and pays for the content of their cart resulting in the creation of an order for the merchant. The merchant may then review and fulfill (or cancel) the order. Theproduct is then delivered to the customer. If the customer is not satisfied, they might return the products to the merchant.
[0155] In an example embodiment, a customer may browse a merchant’s products through a number of different channels 110A-B such as, for example, the merchant’s online store 138, a physical storefront through a POS device 152; an electronic marketplace, through an electronic buy button integrated into a website or a social media channel). In some cases, channels 110A-B may be modeled as applications 142A-B. A merchandising component in the commerce management engine 136 may be configured for creating, and managing product listings (using product data objects or models for example) to allow merchants to describe what they want to sell and where they sell it. The association between a product listing and a channel may be modeled as a product publication and accessed by channel applications, such as via a product listing API. A product may have many attributes and / or characteristics, like size and color, and many variants that expand the available options into specific combinations of all the attributes, like a variant that is size extra-small and green, or a variant that is size large and blue. Products may have at least one variant (e.g., a "default variant") created for a product without any options. To facilitate browsing and management, products may be grouped into collections, provided product identifiers (e.g., stock keeping unit (SKU)) and the like. Collections of products may be built by either manually categorizing products into one (e.g., a custom collection), by building rulesets for automatic classification (e.g., a smart collection), and the like. Product listings may include 2D images, 3D images or models, which may be viewed through a virtual or augmented reality interface, and the like.
[0156] In some embodiments, a shopping cart object is used to store or keep track of the products that the customer intends to buy. The shopping cart object may be channel specific and can be composed of multiple cart line items, where each cart line item tracks the quantity for a particular product variant. Since adding a product to a cart does not imply any commitment from the customer or the merchant, and the expected lifespan of a cart may be in the order of minutes (not days), cart objects / data representing a cart may be persisted to an ephemeral data store.
[0157] The customer then proceeds to checkout. A checkout object or page generated by the commerce management engine 136 may be configured to receive customer information to complete the order such as the customer’s contact information, billing information and / or shipping details. If the customer inputs their contact information but does not proceed topayment, the e-commerce platform 100 may (e.g., via an abandoned checkout component) transmit a message to the customer device 150 to encourage the customer to complete the checkout. For those reasons, checkout objects can have much longer lifespans than cart objects (hours or even days) and may therefore be persisted. Customers then pay for the content of their cart resulting in the creation of an order for the merchant. In some embodiments, the commerce management engine 136 may be configured to communicate with various payment gateways and services 106 (e.g., online payment systems, mobile payment systems, digital wallets, credit card gateways) via a payment processing component. The actual interactions with the payment gateways 106 may be provided through a card server environment. At the end of the checkout process, an order is created. An order is a contract of sale between the merchant and the customer where the merchant agrees to provide the goods and services listed on the order (e.g., order line items, shipping line items, and the like) and the customer agrees to provide payment (including taxes). Once an order is created, an order confirmation notification may be sent to the customer and an order placed notification sent to the merchant via a notification component. Inventory may be reserved when a payment processing job starts to avoid over-selling (e.g., merchants may control this behavior using an inventory policy or configuration for each variant). Inventory reservation may have a short time span (minutes) and may need to be fast and scalable to support flash sales or “drops”, which are events during which a discount, promotion or limited inventory of a product may be offered for sale for buyers in a particular location and / or for a particular (usually short) time. The reservation is released if the payment fails. When the payment succeeds, and an order is created, the reservation is converted into a permanent (longterm) inventory commitment allocated to a specific location. An inventory component of the commerce management engine 136 may record where variants are stocked, and may track quantities for variants that have inventory tracking enabled. It may decouple product variants (a customer-facing concept representing the template of a product listing) from inventory items (a merchant-facing concept that represents an item whose quantity and location is managed). An inventory level component may keep track of quantities that are available for sale, committed to an order or incoming from an inventory transfer component (e.g., from a vendor).
[0158] The merchant may then review and fulfill (or cancel) the order. A review component of the commerce management engine 136 may implement a business process merchant’s use to ensure orders are suitable for fulfillment before actually fulfilling them. Orders may be fraudulent, require verification (e.g., ID checking), have a payment method which requires themerchant to wait to make sure they will receive their funds, and the like. Risks and recommendations may be persisted in an order risk model. Order risks may be generated from a fraud detection tool, submitted by a third-party through an order risk API, and the like. Before proceeding to fulfillment, the merchant may need to capture the payment information (e.g., credit card information) or wait to receive it (e.g., via a bank transfer, check, and the like) before it marks the order as paid. The merchant may now prepare the products for delivery. In some embodiments, this business process may be implemented by a fulfillment component of the commerce management engine 136. The fulfillment component may group the line items of the order into a logical fulfillment unit of work based on an inventory location and fulfillment service. The merchant may review, adjust the unit of work, and trigger the relevant fulfillment services, such as through a manual fulfillment service (e.g., at merchant managed locations) used when the merchant picks and packs the products in a box, purchase a shipping label and input its tracking number, or just mark the item as fulfilled. Alternatively, an API fulfillment service may trigger a third-party application or service to create a fulfillment record for a third-party fulfillment service. Other possibilities exist for fulfilling an order. If the customer is not satisfied, they may be able to return the product(s) to the merchant. The business process merchants may go through to "un-sell" an item may be implemented by a return component. Returns may consist of a variety of different actions, such as a restock, where the product that was sold actually comes back into the business and is sellable again; a refund, where the money that was collected from the customer is partially or fully returned; an accounting adjustment noting how much money was refunded (e.g., including if there was any restocking fees or goods that weren't returned and remain in the customer’s hands); and the like. A return may represent a change to the contract of sale (e.g., the order), and where the e-commerce platform 100 may make the merchant aware of compliance issues with respect to legal obligations (e.g., with respect to taxes). In some embodiments, the e-commerce platform 100 may enable merchants to keep track of changes to the contract of sales over time, such as implemented through a sales model component (e.g., an append-only date-based ledger that records sale-related events that happened to an item).
[0159] In some examples, the applications 142A-B may include an application that enables a user interface (UI) to be displayed on the customer device 150. In particular, the e-commerce platform 100 may provide functionality to enable content associated with an online store 138 to be displayed on the customer device 150 via a UI.
[0160] In various examples, the present disclosure provides a technical solution that enables more efficient operation of a LLM-based chatbot by enabling the LLM to understand the current context of the chatbot. The use of context prompts designed for each page is an efficient mechanism for providing contextual information to the LLM. Notably, the context prompt provides contextual information each time the user navigates to a new page, even if there is no explicit user input to the chatbot UI. The context prompts are maintained in the history of the LLM session, such that the LLM is provided with not only the current context but the navigation history (e.g., prior context prompts). This may enable the LLM to generate more accurate output (e.g., user might request “I want to make this page look like the last page” and the LLM would have navigation history to understand what is meant by “last page”).
[0161] The context prompt provides the LLM with sufficient and relevant information for the functionality of the page, so that the LLM is not attempting to collect incorrect information or generating output with insufficient information. The context prompt that contains contextual information extracted from the current page is a more efficient way to provide information to the LLM than a generic prompt that provides the LLM with all possible information (including information irrelevant to the functioning of the current page; e.g., information about a theme color may be irrelevant when the user is currently at an inventory page). This also improves the user experience by relieving the user of the burden of knowing, finding and providing the necessary information to the LLM.
[0162] In examples described herein, the chatbot engine may allow the user to be involved in the navigation steps (e.g., user confirms each navigation step and is navigated to each page), which may help to increase user confidence in the chatbot operation and may also help to educate the user. This provides an improved user experience over conventional chatbots that simply tell the user what to do without providing the means (e.g., requiring the user to seek out how to perform the required navigation), or other conventional chatbots that perform the operation in the backend without user involvement. For example, a conventional LLM-based chatbot may (e.g., using plugins to system tools) perform the navigation and operation commands automatically without requiring the user to confirm each step. This may be undesirable because there is a risk of the LLM generating navigation and operation commands that are incorrect for performing the requested task (e.g., due to the hallucination phenomenon or due to the LLM having insufficient information).
[0163] Although the present disclosure has described a LLM in various examples, it should be understood that the LLM may be any suitable language model (e.g., including LLMs such as LLaMA, Falcon 40B, GPT-3, GPT-4 or ChatGPT, as well as other language models such as BART, among others).
[0164] Although the present disclosure describes methods and processes with operations (e.g., steps) in a certain order, one or more operations of the methods and processes may be omitted or altered as appropriate. One or more operations may take place in an order other than that in which they are described, as appropriate.
[0165] Although the present disclosure is described, at least in part, in terms of methods, a person of ordinary skill in the art will understand that the present disclosure is also directed to the various components for performing at least some of the aspects and features of the described methods, be it by way of hardware components, software or any combination of the two. Accordingly, the technical solution of the present disclosure may be embodied in the form of a software product. A suitable software product may be stored in a pre-recorded storage device or other similar non-volatile or non-transitory computer readable medium, including DVDs, CD- ROMs, USB flash disk, a removable hard disk, or other storage media, for example. The software product includes instructions tangibly stored thereon that enable a processing device (e.g., a personal computer, a server, or a network device) to execute examples of the methods disclosed herein.
[0166] The present disclosure may be embodied in other specific forms without departing from the subject matter of the claims. The described example embodiments are to be considered in all respects as being only illustrative and not restrictive. Selected features from one or more of the above-described embodiments may be combined to create alternative embodiments not explicitly described, features suitable for such combinations being understood within the scope of this disclosure.
[0167] All values and sub-ranges within disclosed ranges are also disclosed. Also, although the systems, devices and processes disclosed and shown herein may comprise a specific number of elements / components, the systems, devices and assemblies could be modified to include additional or fewer of such elements / components. For example, although any of the elements / components disclosed may be referenced as being singular, the embodiments disclosed herein could be modified to include a plurality of such elements / components. The subject matterdescribed herein intends to cover and embrace all suitable changes in technology.
Claims
CLAIMS1. A method comprising: while a user interface (UI) is at a page, providing a context prompt to a large language model (LLM), the context prompt providing contextual information including information about the page; providing a check context prompt to the LLM instructing the LLM to determine a suitable context for performing a task; receiving output from the LLM based on the check context prompt, the output including a confirmation that the page offers a suitable context for performing the task; providing a task prompt to the LLM instructing the LLM to generate an operation command for performing the task using a functionality of the page; and receiving output from the LLM based on the task prompt, the output including the operation command for performing the task using the functionality of the page.
2. The method of claim 1, wherein the page is a target page, the method further comprising, prior to the UI being at the target page: while the UI is at a prior page different from the target page, providing another context prompt to the LLM, the another context prompt providing contextual information including information about the prior page; providing another check context prompt to the LLM instructing the LLM to determine a suitable context for performing the task; and receiving output from the LLM based on the another check context prompt, the output including a navigation command to navigate to the target page.
3. The method of claim 2, further comprising: providing a navigation option in the UI for executing the navigation command to navigate the UI to the target page; and navigating the UI to the target page responsive to selection of the navigation option.
4. The method of claim 3, wherein the context prompt providing contextual information including information about the target page is automatically provided responsive to navigating the UI to the target page.
5. The method of claim 3 or 4, further comprising: while the UI is at the prior page and prior to providing the other check context prompt, receiving a task request via user input to the UI, wherein the other check context prompt is provided responsive to receiving the task request; and after the UI has navigated to the target page and responsive to receiving the confirmation of the target page, automatically providing the task prompt to the LLM instructing the LLM to generate an operation command for performing the task using the functionality of the target page.
6. The method of any one of claims 1 to 5, further comprising: responsive to navigation of the UI to the page, automatically extracting the contextual information from the page and automatically providing the context prompt to the LLM.
7. The method of claim 6, further comprising: generating the current context prompt using the contextual information extracted from the page.
8. The method of claim 6 or 7, wherein the contextual information is automatically extracted by automatically executing code embedded in the page.
9. The method of any one of claims 1 to 8, further comprising: searching a task history database for a historical LLM session portion relevant to the task; andincluding the historical LLM session portion in the task prompt.
10. The method of any one of claims 1 to 9, further comprising: providing an operation option in the UI for executing the operation command; and executing the operation command responsive to selection of the operation option.
11. The method of claim 10, further comprising: prior to executing the operation command, providing a preview of a result of executing the operation command.
12. The method of any one of claims 1 to 11, wherein the context prompt provides contextual information including information about the functionality of the page.
13. A computer system comprising: a processor configured to execute computer-readable instructions to cause the computer system to perform the method of any one of claims 1 to 12.
14. A non-transitory computer-readable medium storing instructions that, when executed by a processor of a computing system, cause the computing system to perform the method of any one of claims 1 to 12.
15. A computer program comprising processor-executable instructions that, when executed by a processor of a computing system, cause the computing system to perform the method of any one of claims 1 to 12.
Citation Information
Patent Citations
Session data processing method and device, equipment and storage medium
CN116644145A
Voice assistant-enabled client application with user view context
US20220308828A1