Constructing cue information by dynamic selection from contextual information for submission to language model
By using dynamic prompting management technology, the content units of the language model are selectively reduced, which solves the problems of high language model resource consumption and response latency, improves processing efficiency and response quality, and adapts to different applications and environments.
Patent Information
- Application Number
- CN202480015735.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-19
- Filing Date
- 2024-04-30
- Publication Date
- 2025-10-31
AI Technical Summary
Language models consume high resources and have high response latency when processing prompts, making it difficult to correctly interpret context and affecting the fluency and performance of dialogue.
By using dynamic prompt management technology, the number of content units submitted to the language model is reduced. The most relevant information items are selected using candidate context information to construct prompt information, and the prompt size is dynamically adjusted to reduce resource consumption and response latency.
It improves the processing efficiency of language models, reduces resource and time consumption, maintains response quality, adapts to different applications and execution environments, and has strong flexibility and scalability.
Smart Images

Figure CN120883202A_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This disclosure claims the rights to U.S. Provisional Application 63 / 468,195 ('195 Application'), filed May 22, 2023. The entire contents of '195 Application are incorporated herein by reference. Background Technology
[0003] Executing a language model typically requires significant processing and memory resources. Therefore, applications using language models may be forced to limit the number of user requests they accept at any given time to avoid exceeding the resource capacity of their execution platform. To further address this issue, applications often limit the size of prompts that can be fed into the language model. A prompt is the input information fed to the language model in each dialogue turn, based on which the language model generates a response. The application progressively increases the size of the prompts with each turn of the dialogue. This is because each prompt is constructed by appending the user's current query to the last generated model response, which in turn is appended to previous prompts. Thus, the current prompt represents the current query along with recent dialogue history (up to the maximum number of tokens allowed by the application).
[0004] In some cases, even unrelated to the resource-related challenges described above, applications using language models may exhibit substandard performance. For example, in some situations, the language model struggles to correctly interpret the context conveyed in a prompt. Furthermore, applications may require relatively long timeframes (e.g., several seconds) to generate responses using the language model. This response latency can hinder the natural flow of conversation and / or may degrade application performance in other context-specific ways. Summary of the Invention
[0005] This paper describes a technique for interacting with machine-trained language models using dynamic prompt management. During dialogue, this technique reduces the number of content units submitted to the language model. Content units refer to linguistic information units (such as words, phrases, word fragments, etc.) and / or any other type of information unit (such as image information). This reduction in the number of content units allows the execution platform running the language model to process each input query efficiently. Efficiency is reflected in the reduction of resources and time required to process each input query. Simultaneously, this technique does not degrade the quality of the language model's response because the information submitted to the language model is selected based on its evaluation relevance to each input query.
[0006] A language model is a machine-trained model capable of processing language-based input information, and optionally any other type of input information (including video, image, and audio information). Therefore, a language model can be compared to a multimodal machine-trained model.
[0007] According to one illustrative aspect, the technique includes: receiving an input query; accessing a state data repository providing candidate context information; segmenting the candidate context information into multiple parts; determining the semantic relevance of the input query to each of the multiple parts by performing vector-based analysis to select target context information from the candidate context information; creating a cue message that includes the input query and the target context information; submitting the cue message to a machine-trained language model and receiving a response from the machine-trained language model based on the cue message; and generating output information based on the response. The selection operation reduces the size of the cue message by selecting a subset of the candidate context information that is less than all of the candidate context information, which reduces the amount of resources consumed by the language model when processing the cue message and reduces the latency of the language model in providing the response.
[0008] According to another illustrative aspect, candidate contextual information includes the dialogue history preceding the input query. More specifically, the dialogue history includes previous input queries submitted to the language model and previous responses generated by the language model for those previous input queries.
[0009] According to another illustrative aspect, candidate contextual information also includes knowledge information obtained from at least one knowledge source other than the dialogue history.
[0010] According to another illustrative aspect, the technique also includes: assessing the level of complexity associated with the task of processing the input query; and determining the size of the prompt message based on the level of complexity.
[0011] This summary is provided to introduce some concepts in a simplified form; these concepts are further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description
[0012] Figure 1 A computational system including a prompt-management component is shown, which is used to dynamically generate efficient prompts for submission to the language model.
[0013] Figure 2 A graphical representation of how the tooltip management component constructs tooltips by selecting from candidate context information.
[0014] Figure 3A graphical representation of how the prompt management component dynamically changes the size of prompt message instances during a conversation.
[0015] Figure 4 The following diagram illustrates the implementation. Figure 1 The rule-based logic of the computing system.
[0016] Figure 5 The following diagram illustrates the implementation. Figure 1 The logic of machine training in the computing system.
[0017] Figure 6 This illustrates one implementation of a complexity analysis component, which is... Figure 1 Tip - A part of the management component.
[0018] Figure 7 This illustrates one implementation of a dialogue history-selection component, which is... Figure 1 Tip - Another part of the management components.
[0019] Figure 8 This illustrates one implementation of a knowledge-supplementing component, which is... Figure 1 Tip - Another part of the management components.
[0020] Figure 9 This illustrates one implementation of a compression component, which is Figure 1 Tip - Another part of the management components.
[0021] Figure 10 This illustrates one implementation of a content unit replacement component, which is... Figure 9 It is a part of the compression component.
[0022] Figure 11 This illustrates one implementation of a deduplication component, which is... Figure 1 This is another part of the prompt management component.
[0023] Figure 12 This illustrates one implementation of a redundancy information identification component, which is... Figure 11 It is a part of the deduplication component.
[0024] Figure 13 This illustrates one implementation of a data structure reformatting component, which is... Figure 11 Another part of the deduplication component.
[0025] Figure 14 It shows Figure 1 One way to implement a language model.
[0026] Figure 15 It shows the representation Figure 1 This is an overview of the operation mode of a computing system.
[0027] Figure 16 It shows the representation Figure 1 An overview of another operating method of the computing system.
[0028] Figure 17 This shows how it is used in some implementation methods. Figure 1 The computing device of a computing system.
[0029] Figure 18 A computational system of an illustrative type is shown, which in some implementations is used to achieve any aspect of the features shown in the foregoing figures.
[0030] The same numbers are used throughout the disclosure and accompanying drawings to refer to similar components and features. Sequence 100 refers to the initial numbering in... Figure 1 The features found, sequence number 200 refers to the initial sequence number in Figure 2 The features found, sequence number 300 refers to the initial sequence number in Figure 3 The features found in [the context] can be used to further illustrate this. Detailed Implementation
[0031] A. Overview of Computing Systems
[0032] This chapter provides Figure 1 An overview of the computing system 102 is shown. The computing system 102 includes a dialogue system 104 that provides responses to user queries through one or more dialogue turns. The dialogue system 104 works in conjunction with a language model 106 to perform this service. This section provides an overview of the computing system 102. Sections B through G provide additional illustrative details about the various components of the computing system 102.
[0033] In terminology, as used herein, a "machine-trained model" refers to the computer-implemented logic used to perform a task using machine-trained weights generated during training operations. "Weights" refers to any type of parameter value produced by iterative training operations. In some contexts, terms such as "component," "module," "engine," and "tool" refer to the computer-based technical parts that perform the corresponding functions. The following description... Figure 17 and Figure 18 Examples of illustrative computing devices for performing these functions are provided.
[0034] In some implementations, the dialogue system 104 and the language model 106 are provided and maintained by the same entity. In other implementations, the dialogue system 104 and the language model 106 are provided by different corresponding entities.
[0035] The terms "content unit" and "token" refer to units of linguistic information (including words, word parts, phrases, etc.) and / or any other type of information unit (such as image information). Specifically, the term "token" refers to a unit of information processed by the language model 106 itself. For example, in some implementations, the language model 106 includes a tokenizer that breaks down received linguistic segments into a sequence of units called tokens, which are then processed using a transformer-based neural network (e.g., as described in section G). The term "content unit" is used to quantify the information processed by the dialogue system 104. For example, the term "content unit" refers to a unit of information sent to the language model 106 for processing. It is important to note that the above definitions are independent of where tokenization is formally performed in the computing system 102. For example, the dialogue system 104 may perform tokenization instead of the language model 106. In any such case, the term "content unit" is used when discussing information before the language model 106 processes it. To simplify the definition of the problem, in an illustrative case, there is a one-to-one correspondence between the sequence of content units and the corresponding sequence of tokens processed by the language model 106, and each unit / token corresponds to a single word.
[0036] Application system 108 uses dialogue system 104 in the process of providing overall services. For example, one application system performs a booking function with the help of dialogue system 104. Another application performs a question-and-answer function with the help of dialogue system 104. Yet another application performs an online shopping function based on the guidance provided by dialogue system 104, and so on. Figure 1 Typically, application system 108 includes application logic 110 for performing its native functions. For example, a reservation system includes programs for checking the availability of items (including vehicles, airline flights, hotel rooms, etc.), programs for processing user payments, etc.
[0037] In some implementations, the computational system 102 relies on an “off-the-shelf” language model 106 with given fixed weights 112 generated by others using pre-training operations. The publicly available transformer-based model used for performing pattern completion is the BLOOM model, which is available from HUGGING FACE in New York City, New York, with one version being version 1.3 released on July 6, 2022.
[0038] In some implementations, a pre-training system (not shown) trains a language model 106 for one or more general language modeling tasks independent of the specific function performed by the dialogue system. (It should be noted that developers typically receive the language model 106 after others have performed the pre-training.) For example, in a first language modeling task, the pre-training system randomly masks tokens in a sequence of input tokens fed to the language model 106. The pre-training system evaluates the extent to which the language model 106 can successfully predict the identities of the masked tokens and updates the weights 112 of the language model 106 accordingly. In a second language modeling task, the pre-training system feeds two concatenated sentences into the language model 106, including a first sentence and a second sentence. The pre-training system then measures the extent to which the language model 106 can successfully predict whether the second sentence appropriately follows the first sentence (referencing ground-truth information indicating whether the second sentence appropriately follows the first sentence) and updates the weights of the language model accordingly. The paper “BERT: Pre-training of DeepBidirectional Transformers for Language Understanding” by Devlin et al. (arXiv, Cornell University, arXiv:1810.04805v2[cs.CL], May 24, 2019, 16 pages) provides background techniques for the general task of pre-training language models.
[0039] Once trained, the language model 106 operates as a pattern completion engine. That is, the language model 106 autoregressively predicts the tags most likely to follow the initial tag set. The language model 106 performs this function based on its ability to capture statistical patterns exhibited by the training examples processed in the pre-training operation. Background technical information on the general topic of autoregression in language models can be found in the paper “Language Models are Few-Shot Learners” by Brown et al. (arXiv, Cornell University, arXiv:2005.14165v4[cs.CL], July 22, 2020, p. 75).
[0040] More specifically, language model 106 performs autoregression in the following manner. Assume the initial sequence of text tags (…T) N-3 T N-2 T N-1 T N ) is input into language model 106, where T NThis is the last submitted text tag. For example, the initial sequence of text tags corresponds to an instance of the prompt message. Language model 106 maps the model input information to output information that identifies the next text tag (T) that may follow the text tag sequence. N+1 The agent will generate the tag (T). N+1 ) is appended to the end of the previous labeled sequence, and then the updated model input information (...T) is added. N-3 T N-2 T N-1 T N T N+1 The data is fed into the language model. The agent continues the autoregressive process until the language model 106 generates a stop marker. The agent interprets the stop marker as an instruction to stop generating markers in the manner described above.
[0041] In some implementations, language model 106 includes attention-based logic. Attention-based logic is a function that evaluates the relevance of each part of the input information fed to it to the interpretation of each other part of the input information (and to the same parts). More specifically, in some implementations, language model 106 is implemented as a series of transformer blocks. The following section combines... Figure 14 Further details about this model are presented in Chapter G. Other implementations of the language model 106 use models trained on other types of machines, including fully connected feedforward neural networks (FFN), convolutional neural networks (CNN), recurrent neural networks (RNN), and any combination thereof.
[0042] In some implementations, language model 106 processes only language-based content units provided to it. Language-based content units correspond to any linguistic information unit, including words, parts of words, etc. In other implementations, language model 106 is a multimodal machine-trained model capable of processing content units of any type or combination of types. For example, in some implementations, language model 106 processes input information that includes any combination of language-based content units, video-based content units, image-based content units, audio-based content units, etc. Here, a content unit corresponds to any part of a larger whole, such as an n×m pixel portion of an image for the case of an image content unit. However, for ease of explanation, the following explanation presents an example of language model 106 processing language-based content units.
[0043] Training system 114 trains models trained on one or more other machines used by dialogue system 104. Additional details about these other machine-trained models will be provided in later sections. However, it is important to note at this point that the weights 112 of language model 106 are fixed. This means that when training system 114 trains models trained on other machines, it does not need to fine-tune the weights 112 of language model 106 itself (although in other implementations it may perform some retraining on language model 106).
[0044] Referring now to the dialogue system 104 itself, the user interface component 116 provides an interface through which a user or other entity interacts with the dialogue system 104. For example, in some cases, the user interface component 116 receives an input query 118 from a user. The input query 118 comprises one or more words that convey a question or other information to which the language model 106 is asked to respond. The user interface component 116 receives input queries in any form, such as text-based, speech-based, etc. If received in speech-based form, the user interface component 116 uses a speech recognition system (not shown) to convert the input query into a text-based form.
[0045] The dialogue system 104 generates output information 120 in response to an input query 118. The output information 120 may, in part, express or otherwise depend on the final response provided by the language model 106. The user interface component 116 delivers the output information 120 to the user in any form, such as text-based, speech-based, etc.
[0046] The prompt-management component 122 manages the generation of prompt messages 124 for each round of the dialogue session. As explained above, prompt messages 124 correspond to the input information fed to the language model 106 by the dialogue system 104 and consist of a sequence of content units (e.g., words). The language model 106 processes the prompt messages 124 to generate a response 126, which is fed back to the prompt-management component 122. More specifically, the prompt messages 124 express the user's input query 118 and target context information expressing the context of the user's input query 118. As will be explained in more detail below, the target context information includes content units selected from the dialogue history and / or content units selected from knowledge information obtained from at least one knowledge source. The pool of information from which the target context information is selected is generally referred to below as candidate context information. (Although not shown, at the beginning of the dialogue session, the prompt-management component 122 adds an initial set of content units to the prompt messages 124. This initial set of content units is generally referred to as seed prompts and is generally used to inform the language model 106 of the task it is expected to perform.)
[0047] The dynamic prompt generation component 128 operates as a control agent for the prompt management component 122. For example, the dynamic prompt generation component 128 coordinates interactions with different analysis components 130. The dynamic prompt generation component 128 also assembles information provided by the individual analysis components 130 into prompt information 124.
[0048] In summary, when creating the prompt message 124, the prompt-management component 122 selects some, but not all, context information from the candidate context information. Additionally or alternatively, the prompt-management component 122 selects information from the input query 118 when constructing the prompt message 124.
[0049] By performing target selection, over many dialogue rounds, the dynamic prompt-generation component 128 reduces the number of content units in the prompt information instances submitted to the language model 106 compared to including the entire candidate context information and input query 118. Simultaneously, the dynamic prompt-generation component 128 operates intelligently by selecting the context information most suitable for answering the input query; therefore, the dynamic prompt-generation component 128 does not degrade the quality of the response generated by the language model 106. However, it should be noted that the prompt-generation component 128 does not always need to eliminate content units; for example, in other cases, the dynamic prompt-generation component 128 concludes that it is appropriate for the prompt information 124 to include a relatively large number of content units because the question asked is complex and requires a relatively lengthy prompt information instance to describe it. In other cases, the dynamic prompt-generation component 128 determines that it is appropriate to include all candidate context information at the beginning of the dialogue.
[0050] The prompt-management component 122 offers numerous technical advantages. For example, it reduces the number of content units in the prompt message 124 during the dialogue. Compared to shorter prompt message instances, the execution platform requires more time and resources to process lengthy prompt message instances. Therefore, the prompt-management component 122 has the overall effect of reducing resource consumption in the execution platform implementing the language model 106 and improving the latency-related performance of the entire computing system 102. Here, "resources" include processing resources, memory resources, communication resources, bandwidth, power, etc. More specifically, consider the case where the first prompt message has a first number of content units and the second prompt message has a second number of content units, where the second number of content units is greater than the first number. Compared to the first prompt message, processing and storing the second prompt message during processing requires more memory and processing resources (such as CPU resources). More network and bus-related resources are also required to transfer the second prompt message from one destination to another compared to the first prompt message. In contrast, previous dialogue systems progressively increased the size of the prompt messages as the dialogue progressed. In such a system, the execution platform that implements the language model therefore requires an increasing amount of resources to process each query in the dialogue.
[0051] As explained above, the reduction in content units implemented by the dialogue system 104 does not excessively degrade the quality of the responses generated by the language model 106. This is because the prompt-management component 122 intelligently selects the information item most appropriately suited to answer each input query.
[0052] In some cases, the prompt-management component 122 also improves the quality of the responses generated by the language model 106. This is because the prompt-management component 122 eliminates or reduces the occurrence of irrelevant information items that are not related to answering the question. This reduces the chance that irrelevant contextual information in the prompt message 124 will mislead the language model 106, causing it to generate incorrect responses.
[0053] Furthermore, the prompt-management component 122 does not follow the practice of automatically clearing older context information when the maximum number of content units is reached. Instead, the prompt-management component 122 intelligently selects from all parts of the candidate context information for each input query, without necessarily disregarding older context items. However, in other implementations, one or more components of the prompt-management component 122 may consider the duration of existence of a particular context item as a factor among others when determining whether a particular context item should be included in the composed prompt information 124.
[0054] As another advantage, the dialogue system 104 is easily adaptable to different applications without requiring significant (or any) modifications to its infrastructure. This makes the dialogue system 104 flexible and scalable, and reduces the workload and cost associated with its maintenance. For example, the dialogue system 104 can be easily modified at least by: (1) adjusting the amount of content units used to compose the prompt message 124; and / or (2) adjusting the criteria used to compose the prompt message 124. Developers can take advantage of this flexibility by maintaining one type of contextual information for one application and another type of contextual information for another application. In some cases, adapting to a new application environment only requires modifying one or more parameter values that control the operation of the prompt-management component 122. For example, developers can adjust a parameter value that determines the amount of content units to be used when constructing the prompt message 124. Alternatively or additionally, developers can adjust another parameter value that specifies the weight assigned to a particular type of contextual information when constructing the prompt message 124. For example, dialogue systems used in conjunction with shopping-related applications can advantageously weight terms related to product names and product attributes in candidate context information, while navigation applications can advantageously weight terms related to map-related entities.
[0055] Similarly, the dialogue system 104 flexibly adapts to different execution environments. For example, the dialogue system 104 can adapt the way it constructs the prompt message 124 based on the complexity level set by the user or application provider and / or the processing power of the execution platform that runs the application system 108 and / or the language model 106.
[0056] The dialogue system 104 is effective in handling many different scenarios, examples of which are provided at the end of Section A. For example, in some examples, the dialogue system 104 dynamically produces a prompt 124 that reflects the user's changing search focus, without necessarily including content from previous conversational turns that are irrelevant to the current focus of interest. Alternatively or additionally, the dialogue system 104 compresses the source information used to construct the prompt 124, for example, by selecting prominent terms from the source information and / or removing redundant information from the source information.
[0057] Now, one implementation of the dialogue system 104 will be explained in more detail. The analysis component 130 includes one or more of the following: a complexity-analysis component 132, a dialogue history-selection component 134, a knowledge-supplementation component 136, a compression component 138, and a deduplication component 140. The complexity-analysis component 132 determines the complexity level to be assigned to each input query submitted in the dialogue session. Based on this, the complexity-analysis component 132 determines the appropriate number of content units to include in the prompt information 124 when processing each input query. Section B provides further information on the operation of the complexity-analysis component 132.
[0058] The Dialogue History - Selection component 134 selects the portion of the dialogue history most relevant to the task of answering the input query 118, and the Dynamic Prompt - Generation component 128 uses these portions to compose the prompt message 124. In many cases, the Dialogue History - Selection component 134 will select only a portion of the entire dialogue history, rather than the entire dialogue history. Section C provides further information on the operation of the Dialogue History - Selection component 134.
[0059] The knowledge-supplementing component 136 acquires knowledge information from one or more external knowledge sources 142. In this context, "external" refers to some knowledge source that expresses information generated outside of the dialogue system 104. One exemplary knowledge source corresponds to a dictionary or encyclopedia-type resource, such as the Wikipedia website. Another exemplary knowledge source corresponds to a customer review repository. Another exemplary knowledge source corresponds to a repository that provides information about users, such as a website or data repository that provides user profiles. Yet another knowledge source represents any information obtainable through a search via a general search engine.
[0060] Regardless, the knowledge information consists of multiple knowledge items. The knowledge-supplement component 136 selects those knowledge items most relevant to the task of answering the user's input query 118. The dynamic prompt-generation component 128 uses this part to compose a prompt message 124 for the current input query 118. Section D provides further information about the operation of the knowledge-supplement component 136.
[0061] Compression component 138 selects the concepts that best represent the source information from the source information. "Source information" refers to any of the input query 118 and / or candidate context information (including dialogue history and knowledge information). For example, compression component 138 selects one or more keywords from the source information. Alternatively or additionally, compression component 138 selects one or more named entities from the source information. Alternatively or additionally, compression component 138 uses topic analysis to identify one or more topics related to the source information. Compression component 138 has the effect of compressing the source information by describing it using selected terms. (It should be noted that other components of the prompt-management component 122 also perform compression functions, but in different corresponding ways.) Dynamic prompt-generation component 128 uses the selected terms to compose a prompt message 124 for the current input query 118. Alternatively or additionally, compression component 138 includes a content unit replacement component ( Figure 1 (Not shown in the text), which replaces certain terms in the source information with abbreviations of those terms as a one-to-one mapped compressed representation. Section E provides further information about the operation of compression component 138.
[0062] Deduplication component 140 identifies redundant information and removes it from the source information to further compress the source information. Deduplication component 140 performs this task by identifying a group of information items embedded within each other at a specified distance in the vector space, selecting representative information items from this group, and / or using data structures to represent some of the source information in a way that reduces the amount of redundant information items contained therein. Dynamic hint generation component 128 uses portions of the compressed source information to compose hint information 124 for the current input query 118. Section F provides further information about the operation of deduplication component 140.
[0063] State data repository 144 stores state information 146. State information 146 describes various aspects of the current state of the dialogue between the user and language model 106 mediated by dialogue system 104. For example, state information 146 includes any of the following: a) the current input query 118; b) knowledge information identified by knowledge-supplementary component 136 in the current and previous dialogue turns; c) the current response (or responses) generated by language model 106 in response to the current input query 118; d) dialogue history information about any previous turns of the dialogue prior to the current dialogue turn in which the user submitted input query 118; and e) at least the last submitted prompt information. In summary, state information 146 describes input query 118 and overall candidate context information that may be associated with input query 118. In other implementations, the candidate context information expressed in state information 146 may also include other contextual factors, such as location information (identifying the user's location in the dialogue session), user behavior information, etc. Different implementations of dialogue system 104 use different strategies to control the retention duration of each of the above information.
[0064] Figure 2 This summarizes the operating principle of the prompt-management component 122. Figure 2 Specifically, the prompt information 124 includes content units (e.g., words or word segments) expressed in the current query 118 and target context information 202. In some implementations, the prompt-management component 122 selectively composes the target context information 202 from candidate context information 204 provided in the state data repository 144. The candidate context information 204 includes at least complete dialogue information (including all previous input queries and model responses generated by the language model 106) and all information obtained by the knowledge-supplementation component 136 for the current dialogue turn and any previous dialogue turns. It should be noted that the candidate context information 204 includes a first number of content units, and the target context information 202 includes a second number of content units. For many dialogue turns, the second number of content units is less than the first number of content units. Although Figure 2 As not shown in the diagram, note that management component 122 can also be selected from a portion of input query 118, instead of including the complete input query 118 as provided.
[0065] As used herein, “source information” 206 refers to either the input query 118 and / or the candidate context information 204. Another way to express the functionality of the prompt management component 122 is as follows: the prompt management component 112 constructs the prompt information 124 by selectively compressing the source information 206.
[0066] Figure 3Another operating principle of the prompt-management component 122 is summarized below. The horizontal axis represents the number of content units in a user query or response. The vertical axis represents time. As explained above, the prompt-management component 122 dynamically composes each instance of the prompt message by selecting only those parts of the dialogue history and external knowledge information relevant to the current input query and relating to a specific topic. (Other implementations may also consider other factors, such as the assumed goal of the dialogue, when selecting information items.) Additionally or alternatively, the prompt-management component 122 selects a portion of the input query 118, rather than the entire input query 118. Given this behavior, unlike traditional language model solutions, the prompt-management component 122 does not need to increase the size of the prompt message in a constant manner as the conversation progresses. This improves the efficiency of the dialogue system 104 for the reasons explained above.
[0067] exist Figure 3 In a specific scenario, suppose the user delves deeper into a particular question during the first three dialogue rounds, but then, in the fourth dialogue round, initiates a new inquiry related to a new topic. In response to this query behavior, the prompt-management component 122 progressively increases the size of the prompt message instance during the first three dialogue rounds, but then creates a relatively small prompt message instance for the fourth dialogue round. This fourth-round behavior is performed because the prompt-management component 122 determines that the information conveyed in the first three dialogue rounds is irrelevant to the topic raised in the fourth dialogue round, and therefore does not need to be expressed in the prompt message (at least not fully). Although not shown, it is assumed that the user returns to the topic of the first three dialogue rounds later in the conversation. In response, when constructing new prompt messages for this dialogue round, the prompt-management component 122 selectively extracts content units generated in the first three dialogue rounds of the conversation.
[0068] refer to Figure 4 and Figure 5 The computing system 102 relies on any type of function or any combination of different types of functions to achieve the functions described above. For example, Figure 4 An example is shown in which algorithm component 402 uses one or more rules provided in data repository 404 to map input information to output information. These rules can be expressed as discrete if-then rules and / or any other (multiple) types of rules. Alternatively or additionally, the rules can be expressed as algorithms, such as programs that execute subroutines. Figure 5An example is shown of a machine-trained model 502 mapping input information to output information. The machine-trained model 502 includes weights generated by the training system 114 during initial training operations. For example, the training system 114 iteratively processes the set of training examples in the data repository 504, for example, using stochastic gradient descent in conjunction with backpropagation. The training system 114 uses a loss function to compute the error for each training iteration.
[0069] As explained above, the dialogue system 104 is effective in handling many different scenarios. The following are representative scenarios in which the dialogue system 104 improves the efficiency of the execution platform running the language model 106, and in the process improves the overall performance of the computing system 102.
[0070] Scenario A. Application system 108 hosts an e-commerce website. A user submits a user query inquiring about a phone number, and language model 106 delivers a response based on this user query. In the next conversation round, the user submits an input query on the topic of battery health. In the next conversation round, the user submits an input query on details about the phone's camera. Thus, as the user explores different topics of interest, the user's focus shifts continuously over the first three rounds. If the user's interests continue to change, the content of previous user queries and language model responses may not be entirely relevant to the current input query. To address this issue, dialogue system 104 selects the most relevant parts of the candidate context information that have the greatest impact on the user's current focus of interest.
[0071] Scenario B. Application system 108 includes functionality that enables users to interact with online resources containing relevant information items in linked graphs. In a specific conversation turn, the user submits an input query that includes information extracted from the graph. Alternatively, the knowledge-supplementing component 136 extracts information from the graph. It is assumed that the content extracted from the graph comprises multiple key-value pairs, where the same key appears multiple times for at least one key. For example, a portion of calendar-related content includes redundant labels for attributes such as name, ID, email, location, etc. The deduplication component 140 addresses this situation by eliminating or reducing the amount of redundant labels in the extracted information. Without this provision, the information extracted from the graph would waste a significant portion of the tagging budget allocated to the current conversation turn, potentially even exceeding the allocated budget.
[0072] Scenario C. User queries, language model responses, and / or external knowledge information include unique entities with names, some of which may be verbose. Compression component 138 addresses this issue by using placeholder abbreviations to replace these names.
[0073] Scenario D. Application System 108 hosts a video conferencing application. Assume a user has participated in a one-hour meeting, speaking approximately 8000 words, expressed in about 800 sentences. Assume the user submits a user query referencing the meeting minutes. For example, the user submits an input query asking, "Who is the project leader responsible for the Atlanta Delta project, as discussed in the meeting minutes?" Dialogue System 104 addresses this by identifying at least one portion of the record relevant to the input query, for example, by identifying sentences in the record most closely related to the concepts of "project leader," "Delta project," "Atlanta," etc., expressed in the input query. Dialogue System 104 can perform the same function for any referenced document.
[0074] B. Descriptive Complexity Analysis Components
[0075] Figure 6 One implementation of the complexity analysis component 132 is illustrated. The complexity analysis component 132 performs the task of determining the complexity level associated with the user's input query 118. The complexity level of the input query 118 typically reflects the complexity of the task of the language model 106 in interpreting the input query 118. The complexity level of the input query 118 also affects the complexity of the prompting information 124 developed by the prompting management component 122 to express the input query 118. Generally, the number of content units required to express the query increases with the complexity of the input query. Overall, the complexity analysis component 132 allows the dialogue system 104 to flexibly adapt to different application and execution environments.
[0076] Complexity-analysis component 132 relies on one or more components and associated technologies to perform its tasks, including explicit input-receiving component 602, query complexity-evaluation component 604, and resource availability-evaluation component 606. Content unit quantity-evaluation component 608 coordinates interactions with the components identified above. Further, content unit quantity-evaluation component 608 determines the maximum number of content units that should be included in the prompt message 124 for the current dialogue turn based on the evaluated complexity level. In other cases, content unit quantity-evaluation component 608 sets a scaling factor that operates to reduce (or expand) the number of content units in prompt message 124 without setting a maximum number of content units. Alternatively or additionally, each of the other components (134, 136, 138, and 140) of prompt-management component 122 uses the evaluated complexity level to determine the degree to which it should compress the source information (corresponding to input query 118 and / or candidate context information); this implementation does not require defining an explicit maximum number of content units. Typically, the content unit quantity-evaluation component 608 is considered to provide output results that control the scale of the prompt information 124 to be formulated.
[0077] The explicit input-receiving component 602 receives an explicit instruction specifying a complexity level from a user or other entity, such as a developer or application provider. For example, the instruction specifies any level—low, medium, or high (e.g., it could be relabeled as "economic," "standard," and "advanced"). In other examples, the instruction specifies a complexity level within a consecutive range. In one implementation, the user selects the complexity level for the entire dialogue (or multiple dialogues) via a configuration interface provided by the computing system 102. Subsequently, the prompt-management component 122 constrains the size of each instance of the prompt information it generates based on the complexity level. In some applications, different complexity levels are associated with different costs.
[0078] Alternatively or additionally, the user specifies the complexity level for each turn of the dialogue. The explicit input-receiver component 602 allows the user to type instructions for each query in different ways. In one example, the explicit input-receiver component 602 receives input signals in response to user interaction with user interface controls provided by the dialogue system 104 or when the user verbally types an input command. In another case, the explicit input-receiver component 602 receives user instructions specified in the input query 118 itself. In this last mentioned case, in addition to controlling the size of the prompt message 124, the complexity level specified in the input query 118 also instructs the language model 106 to generate a response with a specific level of detail.
[0079] The query complexity evaluation component 604 uses rule-based logic and / or machine-trained models and / or other functionalities to map the input query 118 to a complexity level. The rule-based logic uses one or more rules, based on one or more factors, to determine the complexity level. These factors include any one or a combination of the following: a) the length of the input query 118; b) the number of clauses in the input query 118; c) the number of distinct named entities (e.g., products, places, or people) specified in the input query 118; d) the complexity of the logical relationships expressed in the input query 118, etc. Other factors depend on the overall complexity of the conversational session in which the input query 118 occurs. For example, other factors reflect the level of complexity of the topic the user appears to be asking in one or more conversational turns. For example, consider a first user who wants to know when a plane departs from a particular airport to a particular destination, versus a second user who, given several explicit preferences, asks for the cheapest flight to that destination. The second user's topic is more complex than the first user's. This is partly because the first user's topic has fewer variables than the second user's topic, and the first user's topic requires less context to answer compared to the second user's topic.
[0080] With the query complexity evaluation component 604 implemented as a machine-trained model, the machine-trained model maps the input query 118 to an embedding, and then maps the embedding to a classification result. The machine-trained model can rely on any type of neural network to perform this task, including feedforward neural networks, convolutional neural networks, transformer-based neural networks, etc. The training system 114 uses a training example corpus to train this machine-trained model, where each training example specifies an illustrative input query coupled to the ground truth complexity level. The training system 114 attempts to minimize the difference between its predictions and the ground truth labels.
[0081] Resource availability assessment component 606 receives an input signal indicating the current processing capacity of the execution platform running language model 106. Processing capacity depends on one or more factors, including any combination of: a) the number of queued input requests; b) the amount of available processing resources; c) the amount of available memory resources, etc. Additionally or alternatively, resource availability assessment component 606 receives an input signal indicating the current processing capacity of application system 108 using dialogue system 104. Resource availability assessment component 606 uses rule-based systems and / or machine-trained models and / or other functionalities to map these factors to complexity levels. Here, the complexity level does not assess the conceptual complexity of the input query 118 itself, but rather the amount of resources that can be devoted to processing the input query 118.
[0082] The content unit quantity evaluation component 608 consults environment-specific rules to determine the final complexity level based on the various complexity levels specified by the components (602, 604, 606) identified above. In one case, the content unit quantity evaluation component 608 selects the lowest complexity level specified by the components (602, 604, 606) identified above. In another case, the content unit quantity evaluation component 608 calculates the final complexity level as a weighted combination of the complexity levels specified by the components (602, 604, 606) identified above, or as a transformation of the machine training based on the complexity levels specified by the components identified above. In some implementations, the content unit quantity evaluation component 608 then uses an environment-specific lookup table, a machine-trained model, etc., to map the final complexity level to the number of content units to be used to compose the prompt message 124. The prompt management component 122 uses the number of content units specified by the complexity analysis component 132 to control the amount of candidate content information selected to be included in the prompt message. In other words, the prompt-management component 122 uses the number of content units to determine the strength of compression of the source information used to construct the prompt message 124.
[0083] C. Explanatory Dialogue History Selection Component
[0084] Figure 7One implementation of the dialogue history selection component 134 is shown. The dialogue history selection component 134 performs the task of selecting the portion of the dialogue history 702 most relevant to the input query 118. For some dialogue rounds, this has the effect of reducing the size of the prompt message 124, which in turn contributes to the efficiency-related effects stated above. The dialogue history selection component 134 includes a segmentation component 704 that segments the dialogue history 702 into portions with specific ranges based on one or more factors. The range determines the size of each portion. In one implementation, the segmentation component 704 treats each user query in the dialogue as a distinct portion and each response in the dialogue as a distinct portion. In another implementation, the segmentation component 704 treats each paragraph or sentence or each distinct clause or each key-value pair in each input query and each response as a distinct portion. For example, a portion may correspond to a part of an input query or a response. In another case, the segmentation component 704 dynamically selects the range of portions for each dialogue round based on one or more factors, including the complexity of the query (determined by the query complexity evaluation component 604), explicit user instructions, etc. A section consists of one or more content units. For example, a sentence-level section consists of a sequence of words in a sentence, and each word can be considered a content unit.
[0085] The segmentation component 704 performs its "chunking" operation in response to environment-specific triggering events. For example, in some implementations, the segmentation component 704 performs its operation when any new information item (such as a new input query, response, or information item) is introduced. The resulting partitions persist in subsequent dialogue rounds. In other cases, the segmentation component 704 re-performs the segmentation operation on all or some of the candidate context information at each dialogue round. Depending on the nature of the question posed in the current dialogue round, the new segmentation may be appropriate.
[0086] Mapping component 706 maps the input query to the query embedding (e.g., distributed vector V). Q In this context, each dialogue part is mapped to a dialogue part embedding (e.g., the distributed vector V of the first dialogue part). DP1 The mapping component 706 together provides the embedding set 708. A distributed vector is a vector whose information is distributed over its d dimensions, as opposed to a one-hot vector that assigns a specific concept to each dimension. The proximity of two vectors in a vector space specifies how well the two vectors describe similar concepts. The mapping component 706 can use any neural network to perform the mapping described above, such as the language model 106 itself (described in more detail below in section G). In other cases, the mapping component 706 uses feedforward neural networks, convolutional neural networks, etc., to perform the mapping.
[0087] The relevance-evaluation component 710 determines the proximity of the query embedding to each dialogue part embedding. The relevance-evaluation component 710 can use any metric to perform this evaluation, such as cosine similarity, inner product, Euclidean distance, etc. The relevance-evaluation component 710 then selects zero, one, or more dialogue parts that satisfy a specified context-specific relevance test. For example, in one implementation, the relevance-evaluation component 710 selects N dialogue parts that are closest to the current input query 118, where each of these parts satisfies a context-specific threshold in its proximity to the input query 118. The dynamic prompt-generation component 128 then selectively includes these dialogue parts in the prompt information 124 it generates. In summary, the analysis performed by the mapping component 708 and the relevance-evaluation component 710 is considered vector-based because it relies on comparisons of vectors in a vector space.
[0088] D. Supplementary Explanatory Knowledge Components
[0089] Figure 8 One implementation of the knowledge-supplementing component 136 is shown. The knowledge-supplementing component 136 uses an acquisition engine 802 to acquire knowledge information 804 from a knowledge source 142. Then, the knowledge-supplementing component 136 selects portions of the knowledge information 804 for use in the prompt information 124. For some dialogue rounds, this has the effect of reducing the amount of knowledge information conveyed by the prompt information 124, which in turn contributes to the efficiency-related effects described above.
[0090] Regarding the first phase, the retrieval engine 802 is configured to initiate retrieval operations based on different environment-specific trigger events. In one example, the retrieval engine 802 performs a retrieval operation for each query submitted by the user. For instance, when a new input query 118 is submitted, the compression component 138 (described below) identifies keywords, named entities, and / or topics in the input query 116 and / or candidate context information. In response, the retrieval engine 108 performs a search to find supplementary information related to the identified concepts and then adds it to the candidate context information.
[0091] In another example, the retrieval engine 802 uses rule-based logic and / or machine-trained models and / or other functionalities to determine whether to perform a retrieval operation on the current input query 118. For example, in some cases, the retrieval engine 802 performs a retrieval operation when the input query 118 contains named entities and / or the input query 118 specifies a particular topic.
[0092] Alternatively or additionally, the acquisition engine 802 performs an acquisition operation when it is initially determined that no other dialogue portion or existing knowledge item is sufficiently relevant to the input query 118 (as assessed using environment-specific thresholds).
[0093] The acquisition engine 802 uses the vector-based analysis specified in Section C above to evaluate the relevance of the knowledge information instance to the input query 118 (e.g., by mapping the knowledge information instance and the input query 118 to two distributed vectors and then evaluating the distance between the two vectors in the vector space). The acquisition engine 802 also uses environment-specific rules to determine the knowledge sources(s) from which to extract knowledge information 804. For example, in cases where the input query 118 is soliciting advice about a product, the acquisition engine 802 consults a customer review database for knowledge information. In some implementations, the acquisition engine 802 uses an application programming interface (API) to interact with knowledge source 142.
[0094] Knowledge information 804 consists of multiple knowledge items. More specifically, knowledge information 804 may include information acquired in response to the current input query 118 and / or one or more previous input queries and / or in response to other triggering events. The knowledge-supplementing component 136 uses the segmentation component 806 to determine the scope of each knowledge item based on any of the factors described above in Section C. For example, the segmentation component 806 treats individual sentences or individual paragraphs as separate knowledge items. The mapping component 808 maps the input query 118 and each knowledge item to corresponding embeddings to provide an embedding set 810. The mapping component 808 is implemented in any of the ways specified above in Section C. The relevance-evaluation component 812 selects zero, one, or more knowledge items based on any of the considerations specified above in Section C. For example, the relevance-evaluation component 812 selects N knowledge items that are closest to the current input query 118 or satisfy any other context-specific relevance test. The proximity of a knowledge item to the input query 118 can be evaluated in any of the ways described above, for example, by representing the knowledge item and the input query 118 as two distributed vectors and using cosine similarity or any other distance metric to evaluate the distance between the two vectors. After selecting a knowledge item, the dynamic suggestion-generation component 128 selectively includes the selected knowledge item in the suggestion information 124 it generates.
[0095] E. Explanatory Compression Components
[0096] Figure 9One implementation of compression component 138 is illustrated. Compression component 138 has the effect of reducing the number of content units in a more inclusive set of content units, which in turn contributes to the efficiency-related benefits stated above. In some cases, compression component 138 compresses the content in candidate context information 902, including dialogue history and / or knowledge information acquired by knowledge-supplementing component 136. Alternatively or additionally, compression component 138 compresses the content of user input query 118. For brevity, the content operated on by compression component 138 is referred to herein as “source information” 904. That is, source information 904 refers to input query 118, candidate context information 902, etc., or any combination thereof. Compression component 138 maps source information 904 to compressed source information.
[0097] Compression component 138 uses different components and associated technologies to perform different types of compression. Typically, each technology provides a scaled-down representation of the source information that retains at least some of the semantic content of the source information in its original form. The scaled-down representation of the source information is included in prompt information 124 in place of the source information in its original form.
[0098] Compression component 138 comprises a keyword extraction component 906, a NER extraction component 908 (where "NER" is an abbreviation for Named Entity Recognition), a topic modeling component 910, and a content unit replacement component 912. Compression management component 914 uses rule-based logic and / or machine-trained logic and / or other functionalities to determine when to invoke the individual compression components (906, 908, 910, and 912). In one scenario, once compression component 138 is invoked, compression management component 914 invokes all individual compression components (906, 908, 910, and 912), allowing these components to operate in parallel.
[0099] More specifically, in some cases, the dialogue history-selection component 134 and the knowledge-supplement component 136 perform a first-level, relatively coarse compression. Then, the compression component 138 performs a more detailed level of compression. In other cases, the compression component 138 is executed first, and the concepts extracted by it are used to trigger the operation of the knowledge-supplement component 136. In still other cases, the prompt-management component 122 applies the compression component 138 instead of the operations of the dialogue history-selection component 134 and / or the knowledge-supplement component 136. Alternatively or additionally, when the specified prompt size constraint level is below a prescribed environment-specific threshold, the prompt-management component 122 invokes the compression component 138, which requires special measures to use content units wisely. Other strategies for invoking the knowledge-supplement component 136 are also possible.
[0100] The keyword extraction component 906 uses any rule-based logic (e.g., any algorithm) or machine-trained model to detect prominent keywords or named entities associated with the source information 904. For example, the keyword extraction component 906 can use term frequency-inverse document frequency (TF-IDF) or TextRank algorithms to identify prominent words in the source information 904. Alternatively or additionally, the keyword extraction component uses any type of machine-trained model (such as a neural network of classifier type) to identify keywords in the source information 904.
[0101] Similarly, the NER extraction component 908 uses any rule-based logic (e.g., any algorithm) and / or a machine-trained model to identify named entities associated with the source information 904. For example, in one implementation, the NER extraction component 909 uses a Conditional Random Field (CFR) classifier to identify entity references within the text content unit stream. In another implementation, the NER extraction component 908 uses any type of neural network to identify named entities. For example, a transformer-based encoder maps a sequence of text content units to a corresponding sequence of hidden state embeddings. A post-processing classifier neural network then maps the hidden state embeddings to probabilistic information. The probabilistic information specifies whether each content unit in the content unit sequence is part of an entity reference. In some implementations, the post-processing classifier neural network comprises a machine-trained linear neural network followed by a Softmax operation (e.g., a normalized exponential function).
[0102] The topic-modeling component 910 can also use various rule-based logic and / or machine-trained models to extract topics associated with the source information 904, including Latent Dirichlet Allocation (LDA), Nonnegative Matrix Factorization (NMF), etc. Background technical information on general topics regarding neural network techniques for performing topic extraction and summarization can be found in the following literature: “Topic Modelling Meets Deep Neural Networks: A Survey” by Zhao et al. (arXiv, Cornell University, arXiv:2103.00498v1 [cs.LG], February 28, 2021, 8 pages); and “A Survey on Neural Network-Based Summarization Methods” by Dong and Yue (arXiv, Cornell University, arXiv:1804.04589v1 [cs.CL], March 19, 2018, 16 pages).
[0103] In some implementations, the compression component 138 further weights the relevance of the selected terms (keywords, named entities, topics, etc.) based on one or more weighting factors, and uses these weighting factors when determining which terms to include in the prompt information 124. For example, the compression component 138 determines the degree to which the selected terms are relevant to the user's interest information, such as as specified in the user profile. In some implementations, the compression component 138 makes this determination by performing a lexical and / or semantic comparison between the selected terms and the user's interest information. In some cases, the compression component 138 selects the top K terms. By advantageously weighting the selected terms, the compression component 138 will prioritize these terms compared to other terms without similar weighting, and increases the likelihood that the selected terms will be included in the top K information items.
[0104] Figure 10 The operation of the content unit replacement component 912 is summarized. Content unit replacement component 912 applies one or more transformation rules to map certain strings in source information 1002 to abbreviated strings in reformatted source information 1004. The transformation rules specify the types of text strings to be abbreviated and the manner of abbreviation. Some transformation rules are implemented as mapping lookup tables. For example, suppose the original source information 1002 includes the string "Bill_Gates@microsoft.com". Content unit replacement component 912 consistently replaces all occurrences of "Bill_Gates" (or similar expressions) with "BG". Similarly, content unit replacement component 912 replaces the above email address of Bill Gates with the abbreviated string "BG-email". In another example, suppose the original source information 1002 includes the GUID "F9168C5E-CEB2-4faa-B6BF-329BF39FA1E4". Content unit replacement component 912 abbreviates this code to "F916". The language model 106 was trained to find patterns in text, so it is likely to correctly interpret the meaning of abbreviation strings based on the meaning conveyed by the abbreviation and its surrounding context.
[0105] The Content Unit Replacement Component 912 performs a supplementary restoration operation when it receives a response from the Language Model 106 that includes one or more abbreviations from its previously defined abbreviations. For example, suppose the Language Model 106 delivers a response 126 containing a string that corresponds to an abbreviation expressed in the prompt message 124. The Content Unit Replacement Component 912 resolves this by mapping the abbreviation back to its original form, for example, by mapping “BG-email” back to “Bill_Gates@microsoft.com”.
[0106] F. Descriptive Deduplication Component
[0107] Figure 11 One implementation of the deduplication component 140 is illustrated. The deduplication component 140 performs another aspect of compression by identifying and removing (or reducing) redundant information in the input query 118 and / or candidate context information 1102 (which in turn corresponds to dialogue history and / or external knowledge information). This information is again referred to herein as “source information” 1104. A portion of the source information 1104 is referred to herein as an information item. The deduplication component 140 maps the source information 1104 to compressed source information. Overall, the deduplication component reduces the amount of information conveyed by the prompt information 124, which in turn contributes to the efficiency-related effects stated above.
[0108] Redundant Information Identification Component 1106 identifies a group of information items in source information 1104 that are considered to convey the same or closely related concepts. Then, Redundant Information Identification Component 1106 selects at least one representative member of this group to include in prompt information 124. In other words, this member represents the group as a whole and is used to represent the entire group. For example, Redundant Information Identification Component 1106 selects the group member most similar to the input query 118, for example, by performing a vector-based comparison.
[0109] In some implementations, the redundant information identification component 1106 identifies a set of eligible information items by using a mapping component (as explained in Section C), for example, by mapping the information items to corresponding embeddings in a vector space using a neural network. The redundant information identification component 1106 then determines whether there exists at least one group in which at least two embeddings are within a radius of a predetermined size. The redundant information identification component 1106 then selects one or more representative embeddings (and corresponding information items) from each such group.
[0110] In some implementations, the redundancy information-identification component 1106 selects the radius of candidate groups based on the sparsity of embeddings in the vector space. In some implementations, the redundancy information-identification component 1106 typically uses a smaller radius for densely filled vector spaces compared to sparsely filled vector spaces. In some implementations, the redundancy information-identification component 1106 calculates the density of a cluster by generating the average diameter of the candidate embedding cluster. Here, the radius of the cluster is also defined by its average diameter.
[0111] In some implementations, the redundant information-identification component 1106 uses Mahalanobis distance, Kullback-Leibler (KL) divergence, etc., to identify one or more qualified clusters and evaluate the characteristics of these clusters. Mahalanobis distance assesses the difference between a point and a distribution, while KL divergence assesses the difference between two probability distributions. For example, in some implementations, the redundant information-identification component 1106 uses Mahalanobis distance or KL divergence to calculate the distance between two embeddings associated with two information items. If the calculated metric meets a specified environment-specific threshold, the redundant information-identification component 1106 concludes that the two information items belong to the same cluster. Alternatively or additionally, the information-identification component 1106 uses any of a variety of nonparametric methods to identify one or more qualified clusters and evaluate the characteristics of these clusters. One such nonparametric method uses k-nearest neighbor analysis.
[0112] Different implementations use the redundancy information-identification component 1106 in different ways. In some implementations, the redundancy information-identification component 1106 first finds the information item closest to the user's current input query 118 in the source information, for example, by performing the vector-based analysis described in Section C. Then, the redundancy information-identification component 1106 determines whether the closest information item is a member of a cluster with redundant (or closely related) information items. If so, the redundancy information-identification component 1106 selects representative members of that group, such as the information item closest to the input query 118. To find a second information item related to the input query, the redundancy information-identification component 1106 repeats the above analysis, except that information items in the first mentioned cluster are now excluded from the feasible information items. That is, the redundancy information-identification component 1106 finds the information item closest to the user's current input query 118, excluding information items in the first mentioned cluster. Then, the redundancy information-identification component 1106 determines whether the newly identified closest information item is a member of a cluster with redundant information items. If so, the redundancy information identification component 1106 selects representative members of the group. The redundancy information identification component 1106 can repeat this operation M times to select M information items. Through the behavior described above, the prompting-management component 122 ensures that the M information items convey different facts relevant to the input query 118, rather than restating a single fact. The redundancy information identification component 1106 can achieve the same effect by first dividing the space of the source information into different clusters, and then selecting the M representative information items most relevant to the input query 118 from the M different clusters.
[0113] Alternatively or additionally, the redundancy information identification component 1106: (1) examines the entire source information without referring to the input query; (2) identifies redundant clusters; and (3) replaces the redundant clusters with representative information items. The redundancy information identification component 1106 may perform this function periodically or in response to any type of triggering event.
[0114] Alternatively or additionally, whenever a new information item is submitted to the state data repository 144, the redundancy information-identification component 1106 is triggered to perform its function. The redundancy information-identification component 1106 ensures that the new information item is different from or not closely related to pre-existing information items. If they are the same or closely related, the redundancy information-identification component 1106 again selects a single representative information item for the concept under consideration, which may correspond to the new information item or to a pre-existing information item. Other strategies using the redundancy information-identification component 1106 are also possible.
[0115] The data structure-reformatting component 1108 modifies the format of at least a portion of the source information 1104 to reduce redundant information contained therein. For example, consider an example where the original source information 1104 describes a set of objects by specifying the category(s) to which each object belongs. Further assume that two or more objects share the same category(s). The data structure-generation component 1108 reformatts this source information such that it specifies the shared category(s) only once for two or more objects, rather than repeating the information for each of those objects.
[0116] Compression-management component 1110 determines when to invoke redundant information-identification component 1106 and data structure-generation component 1108. In some implementations, compression-management component 1110 invokes these two components (1106, 1108) for each dialogue round. In other implementations, compression-management component 1110 invokes these two components (1106, 1108) when operating under a restrictive tag budget, and / or when compression-management component 1110 detects that redundant information is included in the source information (redundant information-identification component 1106 and data structure-reformatting component 1108 can be advantageously applied to this source information).
[0117] Figure 12An example of the operation of the redundancy information-identification component 1106 is illustrated. The mapping component 1202 maps multiple information items from the source information to corresponding embeddings. Assume that five of these embeddings reside in a cluster 1204 defined by radius 1206. The selection component 1208 selects a representative member from this cluster 1204, which subsequently represents the entire cluster 1204. In one scenario, when a user submits input query 118, the redundancy information-identification component 1106 initiates its operation. The mapping component 1202 maps input query 118 to embedding 1210. In some implementations, the selection component 1208 selects the embedding in cluster 1204 that is closest to embedding 1210.
[0118] Figure 13 The operation of the data structure-reformatting component 1108 is illustrated. In the first example, instances of the original source information 1302 identify three information items (P1, P2, and P3) of type "A". The original source information 1302 specifically copies the label "A" for all three information items. The data structure-reformatting component 1108 produces reformatted source information 1304, in which the redundant label "A" appears only once, accompanied by any context-specific symbols that convey that the label applies to all three information items (P1, P2, and P3).
[0119] In the second example, the original source information 1302 identifies four information items (P1, P2, P3, and P4) of type "T1". At the next level, the original source information 1302 identifies two information items (P1 and P2) associated with category "A" and two information items (P3 and P4) associated with category "B". Each original information item is labeled with a tag applicable to its corresponding information item. The data structure-reformatting component 1108 then generates reformatted source information 1304, where redundant tags appear only once. In this example, the content unit count has been reduced from 12 to 7 (delimiter characters are not counted). In this case, the data structure-reformatting component 1108 uses a hierarchical tree structure to reduce redundant content.
[0120] In many applications, the data structure-reformatting component 1108 helps reduce redundant information in the content referenced by the input query 118. For example, suppose the input query 118 references calendar content with redundant calendar-related tags. The data structure-reformatting component 1108 is effective in reducing the occurrence of such redundant information.
[0121] G. Descriptive Language Model
[0122] Figure 14 This illustrates one implementation of language model 1402, which can be used as Figure 1The language model 106. The language model 1402 is partly composed of a pipeline of converter components, including a first converter component 1404. Figure 14 Details are provided regarding one method of implementing the first transformer component 1404. Although not specifically illustrated, other transformer components of the language model 1402 have the same architecture as the first transformer component 1404 and perform the same functions (but are controlled by different sets of weights).
[0123] The language model 1402 begins by receiving model input information, for example, corresponding to the prompt information 124. The model input information is expressed as a series of language tokens 1406. As previously explained, a "token" or "text token" refers to a unit of text with any granularity, such as a single word, a word segment generated by byte-pair encoding (BPE), a character n-gram, a word segment identified by the WordPiece algorithm or the SentencePiece algorithm, etc. For ease of explanation, it is assumed that each token corresponds to a complete word.
[0124] Next, the embedding component 1408 maps the label sequence 1406 to the corresponding embedding vectors. For example, the embedding component 1408 generates one-hot vectors describing the labels and then maps these one-hot vectors to the embedding vectors using a linear transformation trained by the machine. Then, the embedding component 1408 adds positional information to the corresponding embedding vectors to produce position-padded embedding vectors 1410. The positional information added to each embedding vector describes the position of the embedding vector within the embedding vector sequence.
[0125] The first transformer component 1404 operates on the position-padded embedding vector 1410. In some implementations, the first transformer component 1404 includes, in sequence, an attention component 1412, a first residual connection and normalization component 1414, a feedforward neural network (FFN) component 1416, and a second residual connection and normalization component 1418.
[0126] Attention component 1412 uses the following equation to perform attention analysis:
[0127]
[0128] Attention component 1412 multiplies the position-padded embedding vector 1410 (or, in some applications, only the last position-padded embedding vector associated with the last received tag) by the query weighting matrix W. Q This generates query information Q. Similarly, the attention component 1412 generates the position-padded embedding vector by correspondingly multiplying it by the key-weighted matrix W. K Sum-weighted matrix W VThis generates key information K and value information V. To execute equation (1), the attention component 1412 takes the dot product of Q and the transpose of K, and then divides the dot product by the scaling factor. This produces a scaling result. The symbol d represents the dimensions of Q and K. The attention component 1412 performs a Softmax (normalized exponential function) operation on the scaling result, and then multiplies the result of the Softmax operation by V to produce the attention output information. More generally, the attention component 1412 determines how much importance should be given to certain parts of the input information when interpreting other parts of the input information. In some cases, the attention component 1412 can be said to perform masked attention as long as it masks output label information that has not yet been determined at any given time. Background technical information on the general concept of attention is provided in the paper "Attention Is All You Need" (9 pages) presented by Vaswani et al. at the 31st Conference on Neural Information Processing Systems (NIPS2017) in 2017.
[0129] It is important to note that Figure 14 Attention component 1412 is shown to consist of multiple attention heads, including a representative attention head 1420. Each attention head performs a computation specified by equation (1), but for a specific representation subspace that differs from the subspaces of the other attention heads. To accomplish this, the attention heads perform the computations described above using different sets of responses to the query, keywords, and value weight matrices. Although not shown, attention component 1412 concatenates the outputs of the individual attention heads of the attention component and then multiplies the concatenated result by another weight matrix W. O .
[0130] The residual connection and normalization component 1414 includes residual connections that combine (e.g., summation) the input information fed to the attention component 1412 with the output information generated by the attention component 1412. The residual connection and normalization component 1414 then normalizes the output information generated by the residual connections, for example, by normalizing the values in the output information based on the mean and standard deviation of the values. Another residual connection and normalization component 1418 performs the same function as the first mentioned residual connection and normalization component 1414. The FFN component 1416 uses a feedforward neural network with any number of layers to transform the input information into output information.
[0131] The first transformer component 1404 produces an output embedding 1422. A series of other transformer components (1424, ..., 1426) perform the same function as the first transformer component 1404, each operating on the output embedding produced by its predecessor. Each transformer component uses its own level-specific machine-trained set of weights. The final transformer component 1426 in the language model 1402 produces the final output embedding 1428.
[0132] Post-processing component 1430 performs post-processing operations on the final output embedding 1428 to produce final output information 1432. For example, in one case, post-processing component 1430 performs a machine-trained linear transformation on the final output embedding 1428 and uses a Softmax component (not shown) to process the result of the transformation. Post-processing component 1430 may optionally use a beam search method to decode the output of the Softmax component.
[0133] In some implementations, the language model 1402 operates in an autoregressive manner. To operate in this manner, the post-processing component 1430 uses a softmax operation to predict the next tag (or, in some cases, predict the most likely set of next tags). The language model 1402 then appends the next tag to the end of the input tag sequence 1406 to provide an updated tag sequence. In the next round, the language model 1402 processes the updated tag sequence to generate the next output tag. The language model 1402 repeats the above process until it generates the specified stop tag.
[0134] It is important to note that Figure 14 The language model 106 shown corresponds to a decoder-only implementation of a machine-trained language model. In other examples, language model 106 encompasses any combination of encoding, decoding, and / or any other functionality. For instance, in other cases, language model 106 uses a decoder model that receives encoded information from a separate encoder model. In some implementations, both the encoder and decoder models include corresponding chains of transformer components and / or other types of attention-based logic.
[0135] H. Explanatory Process
[0136] Figure 15 and Figure 16 Together they show the representation Figure 1The dialog system 104 is described in one embodiment by two processes (1502, 1602). Each process (1502, 1602) is expressed as a series of operations executed in a specific order. However, the order of these operations is merely representative, and these operations can vary in other implementations. Furthermore, any two or more operations described below can be executed in parallel. In one implementation, the boxes shown in the processes (1502, 1602) related to the processing-related functions are combined... Figure 17 and Figure 18 The described computing device is implemented.
[0137] More specifically, Figure 15 The process for interacting with a machine-trained language model (e.g., language model 106) is illustrated. In box 1504, the dialogue system 104 receives an input query (e.g., input query 118). In box 1506, the dialogue system 104 accesses a state data repository (e.g., state data repository 144) that provides candidate contextual information (e.g., candidate contextual information 104). The candidate contextual information includes the dialogue history prior to the input query. The dialogue history, in turn, includes previous input queries submitted to the language model and previous responses generated by the language model for the previous input queries. In box 1508, the dialogue system 104 segments the candidate contextual information into multiple parts, each part comprising one or more content units. In box 1510, the dialogue system 104 selects target contextual information (e.g., target contextual information 202) from the candidate contextual information by performing vector-based analysis to determine the semantic relevance of the input query to each of the multiple parts. In box 1512, the dialogue system 104 creates a prompt message (e.g., prompt message 124) that includes the input query and the target contextual information. In box 1514, the dialogue system 104 submits a prompt to a machine-trained language model and receives a response (e.g., response 126) from the machine-trained language model based on the prompt. A selection operation (in box 1510) reduces the size of the prompt by selecting a subset of candidate context information that is less than all of the candidate context information. This reduces the amount of resources consumed by the language model when processing the prompt and reduces the latency of the language model in providing the response. In box 1516, the dialogue system 104 generates output information (e.g., output information 120) based on the response. Loop 1518 indicates that the operations described above are repeated in each round of the dialogue.
[0138] Figure 16Another process 1602 for interacting with a machine-trained language model (e.g., language model 106) is illustrated. In box 1604, the dialogue system 104 receives an input query (e.g., input query 118). In box 1606, the dialogue system 104 creates a prompt message (e.g., prompt message 124) expressing the input query and target contextual information (e.g., target contextual information 202), which is selected from candidate contextual information (e.g., candidate contextual information 202). Further, the source information, including the input query and / or candidate contextual information, is compressed by reducing the number of content units in the source information to form part of the prompt message. More specifically, the compression applies one or more techniques to provide a scaled-down representation of the source information that retains at least some of the semantic content of the source information in its original form. In box 1608, the dialogue system submits the prompt message to the machine-trained language model and receives a response (e.g., response 126) from the machine-trained language model based on the prompt message. In box 1610, the dialogue system 104 generates output information (e.g., output information 120) based on the response. The amount of time the execution platform implementing the machine-trained language model takes to deliver the response depends on the number of content units in the prompt, and the amount of resources consumed also depends on the number of content units in the prompt. The compression operation reduces the number of content units in the prompt, which reduces the amount of resources consumed by the language model when processing the prompt and reduces the latency of the language model in providing the response. Loop 1612 indicates that the operations of receiving, compressing, creating, submitting, and generating are repeated in each round of the dialogue.
[0139] I. Explanatory Calculation Functionality
[0140] Figure 17 A computing device 1702 is shown, which, in some implementations, is used to implement Figure 1 The computing system 102. The computing device 1702 includes a set of local devices 1704 coupled to a set of servers 1706 via a computer network 1708. Each local device corresponds to any type of computing device, including any of the following: desktop computing devices, laptop computing devices, any type of handheld computing device (e.g., smartphones or tablet computers), mixed reality devices, smart appliances, wearable computing devices (e.g., smartwatches), Internet of Things (IoT) devices, gaming systems, immersive "caves," media devices, in-vehicle computing systems, any type of robotic computing systems, computing systems in manufacturing systems, etc. In some implementations, the computer network 1708 is implemented as a local area network (LAN), a wide area network (e.g., the Internet), one or more peer-to-peer links, or any combination thereof.
[0141] Figure 17 The dashed box in the figure indicates that the functionality of computing system 102 can be distributed across local device 1704 and / or server 1706 in any way. For example, in some cases, each local device or a group of associated local devices implements the entire computing system 102. In other implementations, server 1706 implements the entire computing system 102. Here, individual users interact with server 1706 via browser applications or other local functions provided by local devices. In other implementations, the functionality of computing system 102 is distributed between each local device and server 1706. For example, in one case, server 1706 provides the execution platform for implementing language model 106, and each local device implements... Figure 1 The remaining functions are shown.
[0142] Figure 18 A computational system 1802 is shown in some implementations of any aspect of the mechanism described in the figures above. For example, in some implementations, Figure 18 The type of computing system 1802 shown is used to implement Figure 17 Any local computing device or any server shown. Furthermore, Figure 18 The computing system 1802 of the type shown is used to implement any of the dialogue system 104, language model 106, application system 108, etc. In all cases, the computing system 1802 represents a physical and tangible processing mechanism.
[0143] The computing system 1802 includes a processing system 1804, which includes one or more processors. The processors include one or more central processing units (CPUs), and / or one or more graphics processing units (GPUs), and / or one or more application-specific integrated circuits (ASICs), and / or one or more neural processing units (NPUs), and / or one or more tensor processing units (TPUs), etc. More generally, any processor corresponds to a general-purpose processing unit or a dedicated processor unit.
[0144] The computing system 1802 also includes a computer-readable storage medium 1806 corresponding to one or more computer-readable medium hardware units. The computer-readable storage medium 1806 retains information 1808 of any kind, such as machine-readable instructions, settings, model weights, and / or other data. In some implementations, the computer-readable storage medium 1806 includes one or more solid-state devices, one or more magnetic hard disks, one or more optical disks, magnetic tapes, etc. Any instance of the computer-readable storage medium 1806 uses any technology to store and retrieve information. Further, any instance of the computer-readable storage medium 1806 represents a fixed or movable unit of the computing system 1802. Further, any instance of the computer-readable storage medium 1806 provides volatile and / or non-volatile retention of information.
[0145] More generally, any storage resource or any combination of storage resources described herein will be considered a computer-readable medium. In many cases, a computer-readable medium represents some form of physical and tangible entity. The term computer-readable medium also covers propagated signals, for example, signals transmitted or received via physical conduits and / or air or other wireless media. However, the specific terms “computer-readable storage medium” or “storage device” explicitly exclude the propagated signal itself in transit, while including all other forms of computer-readable medium; for this purpose, a computer-readable storage medium or storage device is “non-transient.”
[0146] The computing system 1802 utilizes any instance of the computer-readable storage medium 1806 in different ways. For example, in some implementations, any instance of the computer-readable storage medium 1806 represents a hardware memory unit (such as random access memory (RAM)) for storing information during program execution by the computing system 1802, and / or a hardware storage unit (such as a hard disk) for more permanently retaining / archiving information. In the latter case, the computing system 1802 also includes one or more drive mechanisms 1810 (such as a hard disk drive mechanism) for storing and retrieving information from instances of the computer-readable storage medium 1806.
[0147] In some implementations, when processing system 1804 executes computer-readable instructions stored in any instance of computer-readable storage medium 1806, computing system 1802 performs any of the functions described above. For example, in some implementations, computing system 1802 executes computer-readable instructions to perform reference... Figure 15 and Figure 16 Each box of the described process. Figure 18 This typically indicates that the hardware logic circuit device 1812 includes any combination of the processing system 1804 and the computer-readable storage medium 1806.
[0148] Additionally or alternatively, the processing system 1804 includes one or more other configurable logic units that use a set of logic gates to perform operations. For example, in some implementations, the processing system 1804 includes a fixed configuration of hardware logic gates, created and set at manufacturing time and not subsequently changed. Additionally or alternatively, the processing system 1804 includes a set of programmable hardware logic gates configured to perform different application-specific tasks. This latter type of device includes programmable array logic devices (PALs), general-purpose array logic devices (GALs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), etc. In these implementations, the processing system 1804 effectively includes a storage device for storing computer-readable instructions, provided that the configurable logic units are configured to execute instructions and thus implement or store those instructions.
[0149] In some cases (e.g., where computing system 1802 represents a user computing device), computing system 1802 also includes an input / output interface 1814 for receiving various inputs (via input device 1816) and providing various outputs (via output device 1818). Illustrative input devices include keyboard devices, mouse input devices, touchscreen input devices, digitizers, one or more still image cameras, one or more video cameras, one or more depth camera systems, one or more microphones, speech recognition mechanisms, any location determination device (e.g., a GPS device), any motion detection mechanism (e.g., an accelerometer and / or gyroscope), etc. In some implementations, a particular output mechanism includes a display device 1820 and an associated graphical user interface (GUI) presentation 1822. Display device 1820 corresponds to a liquid crystal display device, a light-emitting diode display (LED) device, a cathode ray tube device, a projection mechanism, etc. Other output devices include printers, one or more speakers, haptic output mechanisms, archiving mechanisms (for storing output information), etc. In some implementations, the computing system 1802 also includes one or more network interfaces 1824 for exchanging data with other devices via one or more communication channels 1826. One or more communication buses 1828 communicatively couple the units described above together.
[0150] The (multiple) communication channels 1826 are implemented in any manner, such as through a local area computer network, a wide area computer network (e.g., the Internet), a point-to-point connection, or any combination thereof. The (multiple) communication channels 1826 include any combination of hardwired links, wireless links, routers, gateway functions, name servers, etc., controlled by any protocol or combination of protocols.
[0151] Figure 18The computing system 1802 is shown to consist of a discrete set of individual units. In some cases, the set of units corresponds to discrete hardware units provided in a computing device rack with any form factor. Figure 18 The bottom portion illustrates the illustrative form factor. In other cases, the computing system 1802 includes a hardware logic unit that integrates… Figure 18 The functions of two or more units shown. For example, in some implementations, computing system 1802 includes a system-on-a-chip (SoC or SOC), corresponding to a combination of Figure 18 An integrated circuit that functions as two or more of the units shown.
[0152] The following summary provides a collection of illustrative examples of the techniques described in this article.
[0153] (A1) According to one aspect, a method (e.g., process 1502) for interacting with a machine-trained language model (e.g., language model 106) is described. The method includes: receiving (e.g., in box 1504) an input query (e.g., input query 118); and accessing (e.g., in box 1506) a state data repository (e.g., state data repository 144) that provides candidate contextual information (e.g., candidate contextual information 204). The candidate contextual information includes a dialogue history prior to the input query, which includes previous input queries submitted to the language model and previous responses generated by the language model for the previous input queries. The method further includes: segmenting candidate context information (e.g., in box 1508) into multiple parts, each part comprising one or more content units; determining the semantic relevance of the input query to each of the multiple parts by performing vector-based analysis, thereby selecting target context information (e.g., target context information 202) from the candidate context information (e.g., in box 1510); creating (e.g., in box 1512) a prompt message (e.g., prompt message 124) comprising the input query and the target context information; submitting the prompt message (e.g., in box 1514) to a machine-trained language model; and receiving a response from the machine-trained language model based on the prompt message (e.g., response 126). This selection reduces the size of the prompt message by selecting a subset of the candidate context information that is less than all parts of the candidate context information, which reduces the amount of resources consumed by the language model when processing the prompt message and reduces the latency of the language model in providing the response. The method also includes generating output information (e.g., output information 120) based on the response. The operations of receiving, accessing, segmenting, selecting, creating, submitting, and generating are repeated in each round of the dialogue, as follows: Figure 15 As indicated by loop 1518 in the middle.
[0154] According to an illustrative characteristic, this method reduces the number of content units sent to the language model. Reducing the number of content units decreases the amount of work the language model is required to perform. As a further result, reducing the number of content units reduces the resource consumption of the language model and improves the latency of the language model's response delivery. This is because the language model consumes resources and time to process each content unit. In some cases, this method also improves the quality of the language model's response.
[0155] (A2) According to some implementations of the method in A1, each content unit is a word or a part of a word.
[0156] (A3) According to some implementations of the methods of A1 or A2, the machine-trained language model is a transformer-based model that includes attention logic for evaluating the relevance assigned to a part of the input information fed to the attention logic when interpreting each part of the input information.
[0157] (A4) Based on some implementation of any of the methods in A1 to A3, in the process of exploring dialogues on multiple different topics, the target context information is re-evaluated on a per-query basis, and when processing specific input queries related to a specific topic, the selection operation selects only the part of the candidate context information that is related to the specific topic.
[0158] (A5) According to some implementation of any of the methods in A1 to A4, the input query references a document, and the selection operation identifies at least one part of the document that is relevant to the input query.
[0159] (A6) According to some implementation of any of the methods in A1 through A5, the method further includes: evaluating the level of complexity associated with the task of processing the input query; and determining the size of the prompt message based on the level of complexity. The method uses the size of the prompt message to control the amount of candidate content information selected and incorporated into the prompt message.
[0160] (A7) Based on some implementations of the method in A6, the evaluation operation is based on an explicit instruction received by the method, which specifies the complexity level of the dialogue as a whole or a specific dialogue turn in the dialogue.
[0161] (A8) Based on some implementations of the methods in A6 or A7, evaluate the operation based on the resource capabilities of the execution platform that implements the language model and / or the number of requests queued at the execution platform.
[0162] (A9) According to some implementation of any of the methods in A6 to A8, the evaluation operation is performed by rule-based logic or a machine-trained model based on determining the complexity of the input query, which is evaluated based on the length of the input query, or the number of clauses in the input query, or the complexity of the logical relations expressed in the input query, or the number of named entities in the input query, or the complexity of the dialogue, or any combination thereof.
[0163] (A10) According to some implementation of any of the methods in A1 to A9, the dialogue history is segmented into dialogue parts, and the selection operation includes: mapping an input query to a query embedding; mapping a dialogue part to a corresponding dialogue part embedding using a neural network; evaluating the distance between the query embedding in the vector space and the dialogue part embedding associated with a particular dialogue part; identifying a particular dialogue part as relevant to the current query when the distance satisfies a prescribed relevance test; and including the particular dialogue part in the prompt message when the particular dialogue part is determined to be relevant.
[0164] (A11) According to some implementation of any of the methods in A1 to A10, the subject of the candidate context information also includes knowledge information obtained from at least one knowledge source other than the dialogue history based on the input query or previous input queries in the dialogue.
[0165] (A12) According to some implementations of the method in A11, knowledge information is segmented into knowledge items, and the selection operation further includes: mapping the input query to the query embedding; using a neural network to map the knowledge item to the corresponding knowledge item embedding; evaluating the distance between the query embedding in the vector space and the knowledge item embedding associated with the specific knowledge item; identifying the specific knowledge item as relevant to the current query when the distance satisfies a specified relevance test; and including the specific knowledge item in the prompt information when a specific dialogue part is determined to be relevant.
[0166] (A13) According to some implementation of any of the methods in A1 to A12, the resulting parts are generated by splitting and correspond to individual input queries and responses or parts thereof in the dialogue.
[0167] (A14) Depending on some implementation of any of the methods in A1 to A13, the segmentation operation is performed in each dialogue round based on the complexity level of the specific input query associated with each dialogue round.
[0168] In another aspect, some implementations of the techniques described herein include a computing system (e.g., computing system 1802) that includes a processing system (e.g., processing system 1804) having a processor. The computing system also includes a storage device (e.g., computer-readable storage medium 1806) for storing computer-readable instructions (e.g., information 1808). The processing system executes the computer-readable instructions to perform any of the methods described herein (e.g., any individual method of methods A1 through A14).
[0169] In another aspect, some implementations of the techniques described herein include a computer-readable storage medium (e.g., computer-readable storage medium 1806) for storing computer-readable instructions (e.g., information 1808). A processing system (e.g., processing system 1804) executes the computer-readable instructions to perform any operation of the operations described herein (e.g., operations in any individual method of the methods A1 to A14).
[0170] More generally, any individual element and step described herein can be combined into any logically consistent permutation or subset. Furthermore, any such combination can be embodied as a method, apparatus, system, computer-readable storage medium, data structure, article of manufacture, graphical user interface presentation, etc. The technique can also be expressed in the claims as a series of means-plus-format elements; however, this format should not be considered an invocation unless the phrase "means for..." is explicitly used in the claims.
[0171] Regarding the terminology used in this specification, the phrase "configured to" encompasses various physical and tangible mechanisms for performing the identified operations. These mechanisms can be configured to use... Figure 18 The hardware logic circuit device 1812 is used to perform operations. The term "logic" also encompasses the various physical and tangible mechanisms used to perform tasks. For example, Figure 15 and Figure 16 Each processing-related operation illustrated in the flowchart corresponds to a logical component used to perform that operation.
[0172] This specification may identify one or more features as optional. Statements of this type should not be construed as exhaustive indications of features considered optional; generally, unless otherwise stated, any feature will be considered exemplary, even if not explicitly identified herein. Furthermore, any reference to a single entity is not intended to exclude the use of multiple such entities; similarly, the description of multiple entities in the specification is not intended to exclude the use of a single entity. Therefore, a statement that an apparatus or method has feature X does not exclude the possibility that it may have additional features. Furthermore, unless otherwise stated, any feature described as an alternative way of performing the identified function or implementing the identified mechanism may also be combined in any combination.
[0173] With regard to specific terminology, the terms "plurality" or "plural," or the plural form of any term (unless explicitly stated otherwise), refer to two or more items and do not necessarily imply "all" items of a particular kind. The term "at least one of..." refers to one or more items; unless otherwise stated, a reference to a single item without explicitly stating "at least one of..." is not intended to exclude the inclusion of multiple items. Furthermore, the descriptors "first," "second," "third," etc., are used to distinguish different items and do not imply an ordering between items unless otherwise stated. The phrase "A and / or B" refers to A, or B, or A and B. The phrase "any combination thereof" refers to any combination of two or more elements in a list of elements. Furthermore, the terms "comprising," "including," and "having" are open-ended terms used to identify at least one part of a larger whole, but not necessarily all parts of the whole. A "set" is a group that includes one or more members. The phrase "A corresponds to B" in some contexts means "A is B". Finally, the terms "exemplary" or "illustrative" refer to one implementation among many potential implementations.
[0174] Finally, the functionality described herein employs various mechanisms to ensure that any user data is processed in a manner that complies with applicable laws, social norms, and the expectations and preferences of individual users. For example, the functionality is configurable to allow users to explicitly opt in (and then explicitly opt out) to the functionality's terms. The functionality can also be configured to provide appropriate security mechanisms to ensure the privacy of user data (such as data sanitization, encryption, and / or password protection).
[0175] Furthermore, this specification may state various concepts in the context of an illustrative challenge or problem. This interpretation is not intended to suggest that others have understood and / or articulated the challenge or problem in the manner specified herein. Furthermore, this interpretation is not intended to suggest that the subject matter enumerated in the claims is limited to solving the identified challenge or problem; that is, the subject matter in the claims may be applied in the context of challenges or problems other than those described herein.
[0176] Although the subject matter has been described in language specific to structural features and / or methodological actions, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are disclosed as exemplary forms for implementing the claims.
Claims
1. A computer-implemented method for interacting with a machine-trained language model, comprising: Receive input queries; Access a state data repository that provides candidate context information, which includes dialogue history prior to the input query, including previous input queries submitted to the language model and previous responses generated by the language model for the previous input queries; The candidate context information is divided into multiple parts, each part including one or more content units; The semantic relevance of the input query to each of the plurality of parts is determined by performing vector-based analysis, and the target context information is selected from the candidate context information; Create a prompt message that includes the input query and the target context information, wherein the selection reduces the size of the prompt message by selecting a subset of the candidate context information that is less than all of the portions of the candidate context information; The prompt information is submitted to the machine-trained language model, and a response is received from the machine-trained language model based on the prompt information; as well as Output information is generated based on the response. The processes of receiving, accessing, segmenting, selecting, creating, submitting, and generating are repeated for each round of the dialogue.
2. The method of claim 1, wherein the content unit is a word or a portion thereof.
3. The method of claim 1, wherein during the exploration of the dialogues on multiple different topics, the target context information is re-evaluated on a per-query basis, and wherein when processing a specific input query related to a specific topic, the selection picks only the portion of the candidate context information related to the specific topic.
4. The method of claim 1, wherein the input query references a document, and wherein the selection identifies at least one portion of the document that is related to the input query.
5. The method according to claim 1, wherein the method further comprises: Assess the level of complexity associated with the task of processing the input query; as well as The size of the prompt message is determined based on the complexity level. The method uses the size of the prompt information to control the amount of candidate content information that is selected and incorporated into the prompt information.
6. The method of claim 5, wherein the evaluation is based on an explicit instruction received by the method, the instruction specifying a complexity level for the entire dialogue or for a specific dialogue turn within the dialogue.
7. The method of claim 5, wherein the evaluation is based on the determination of the resource capabilities of the execution platform implementing the language model and / or the number of requests queued at the execution platform.
8. The method of claim 5, wherein the evaluation is performed by a rule-based logic or a machine-trained model based on a determination of the complexity of the input query, the complexity being evaluated based on: the length of the input query, or the number of clauses in the input query, or the complexity of the logical relations expressed in the input query, or the number of named entities in the input query, or the complexity of the dialogue, or any combination thereof.
9. The method of claim 1, wherein the dialogue history is segmented into dialogue portions, and wherein the selection includes: Map the input query to a query embedding; The dialogue portion is mapped to the corresponding dialogue portion embedding using a neural network; Evaluate the distance between the query embedding and the dialogue part embedding associated with a specific dialogue part in the vector space; When determining that the distance meets the specified relevance test, the specific dialogue portion is identified as relevant to the current query; as well as When a particular dialogue segment is determined to be relevant, that particular dialogue segment is included in the prompt message.
10. The method of claim 1, wherein the candidate context information further includes knowledge information obtained from at least one knowledge source other than the dialogue history based on the input query or a previous input query in the dialogue.
11. The method of claim 10, wherein the knowledge information is segmented into knowledge items, and wherein the selection further comprises: Map the input query to a query embedding; The knowledge items are mapped to corresponding knowledge item embeddings using a neural network. Evaluate the distance between the query embedding and the knowledge item embedding in the vector space, wherein the knowledge item embedding is associated with a specific knowledge item; When determining that the distance meets the specified relevance test, the specific knowledge item is identified as relevant to the current query; as well as When it is determined that a particular dialogue section is relevant, the specific knowledge item is included in the prompt message.
12. The method of claim 1, wherein the portion generated by the segmentation corresponds to individual input queries and responses and their portions in the dialogue.
13. The method of claim 1, wherein the segmentation is performed for each dialogue round in the dialogue based on the complexity level of a specific input query associated with each dialogue round.
14. A processing system having a processor and a storage device storing machine-readable instructions, the processing system executing the machine-readable instructions to perform the method according to any one of claims 1 to 13.
15. A computer-readable storage medium for storing computer-readable instructions that, when executed by a processing system, perform the method according to any one of claims 1 to 13.