Large language model utterance enhancement
By using pre-trained large language models in intention detection and slot extraction to generate semantic-related phrase variants, the problem of process tedious, time-consuming and non-scaling in the prior art is solved, and more efficient and flexible intention detection and slot extraction is achieved.
Patent Information
- Application Number
- CN202380071754.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-12
- Filing Date
- 2023-09-22
- Publication Date
- 2025-05-16
AI Technical Summary
Existing intention detection mechanisms rely on complex regular expressions or a large number of feature engineering, resulting in the process being tedious, time-consuming and unscalable, making it difficult to effectively interpret and generate semantic-related phrase variants.
By providing seed data to a pre-trained large language model (LLM), semantic-related phrase variants are generated, using the "pre-training, prompting, and prediction" paradigm to improve intention detection and slot extraction, increasing flexibility in interpretation and generation of discourses.
It has achieved the reduction of human input needs, improved the efficiency and flexibility of intention detection and slot extraction, generated semantic-related phrase variants, and improved the operation efficiency of the computing system.
Smart Images

Figure CN120019432A_ABST
Abstract
Description
Background Art
[0001] Today’s intent detection mechanisms rely on rule-based regular expressions or supervised machine learning (ML) techniques with extensive feature engineering like named entity recognition (NER). Such mechanisms require brainstorming complex regular expressions or picking out large, labeled datasets containing an exhaustive set of possible utterances that map to every “intent” of the system (i.e., what a user can say to trigger a command). Along with this list of utterances comes an even larger list of examples of “slot” values.
[0002] Supervised slot extraction requires manual labeling of slots in an inside-outside-begin (IOB) format. As a result, this process of intent detection for slot commands is very tedious, time consuming, and non-scalable. Therefore, what is needed is an improved technique that differs from the traditional "pre-train then fine-tune" paradigm and adopts a new paradigm. In addition, techniques for generating variants of phrases are needed to increase the flexibility of interpreting utterances using new paradigms. There is also a need for improved techniques that facilitate speech-based transcription for specific domains. It is hoped that these various techniques will provide improved results to users and increase the operating efficiency of computing systems.
[0003] The subject matter claimed herein is not limited to embodiments that solve any disadvantages or that operate only in environments such as those described above. Rather, this background is merely provided to illustrate one exemplary technology area in which some embodiments described herein may be practiced. Summary of the invention
[0004] Embodiments disclosed herein relate to systems, devices, and methods for causing a large language model (LLM) to generate semantically related phrase variants for phrases included in seed data that is provided to the LLM to generate the semantically related phrase variants.
[0005] Some embodiments access an LLM that is typically pre-trained on an arbitrary corpus of language training data. An embodiment feeds as input seed data comprising a plurality of phrases that are semantically related to each other and that describe a particular command. When any one of the phrases is received as an utterance input, the utterance input triggers execution of the command. These embodiments cause the LLM to generate a plurality of phrase variants based on the phrase, wherein each phrase variant is semantically related to the other phrases. When any one of the phrase variants is received as a new utterance input, the new utterance input also triggers execution of the command. An embodiment stores the phrases and phrase variants as a phrase list in a data store. The phrases / variants in the list are identified as being semantically related to each other and as triggers for executing the command.
[0006] This Summary is provided to introduce some concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.
[0007] Additional features and advantages will be set forth in the following description, and in part may be apparent from the description, or may be learned by practicing the teachings herein. The features and advantages of the present invention may be realized and obtained by the means and combinations particularly pointed out in the appended claims. The features of the present invention will become more fully apparent from the following description and the appended claims, or may be learned by the practice of the present invention as described below. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] In order to describe the manner in which the above-recited and other advantages and features can be obtained, a more particular description of the subject matter briefly described above will be rendered by reference to specific embodiments that are illustrated in the accompanying drawings. Understanding that these drawings depict only typical embodiments and are therefore not to be considered limiting of scope, the embodiments will be described and explained with additional specificity and detail through the use of the accompanying drawings, in which:
[0009] Figure 1 An example architecture for performing contextualized intent detection and slot extraction based on a pre-training, hint, and prediction paradigm is shown.
[0010] Figure 2 Examples of hints that may be provided to a large language model (LLM) are shown.
[0011] Figure 3 An extension of the architecture is shown, where the architecture now provides utterance input for analysis.
[0012] Figure 4 Another example of a prompt is shown, where the prompt includes text generated by a speech-to-text engine.
[0013] Figure 5 Another example of a prompt is shown, where the prompt is constructed to include additional contextual information to be provided to the LLM.
[0014] Figure 6 A flowchart of an example method for performing contextualized intent and slot extraction is shown.
[0015] Figure 7 An example architecture constructed to generate variations of a phrase is shown, including a service and an LLM, and a feedback loop between the service and the ML model.
[0016] Figure 8 An example of seed data that may be used to prompt the LLM is shown.
[0017] Fig. 9 A flow chart illustrating an example method for causing an LLM to generate semantically related phrase variants for phrases included in seed data.
[0018] Fig.10 An example architecture for facilitating speech-based dictation of programming code is shown.
[0019] Fig.11 A list of phrase pairs is shown.
[0020] Fig.12 A flow diagram of an example method for facilitating speech-based dictation of programming code within the context of an integrated development environment (IDE) is shown.
[0021] Fig.13 An example computer system is shown that may be configured to perform any of the disclosed operations. DETAILED DESCRIPTION
[0022] Some embodiments relate to systems, devices and methods for causing an LLM to generate semantically related phrase variants for phrases included in seed data, which are provided to the LLM to generate semantically related phrase variants. For example, some embodiments access an LLM that is typically pre-trained on an arbitrary corpus of language training data. These embodiments feed seed data including one or more phrases that are semantically related to each other and describe a specific command as input. When any of the phrases is received as an utterance input, the utterance input triggers the execution of the command. The LLM then generates multiple phrase variants based on the phrase, each of which is semantically related to the other phrases. When any of the phrase variants is received as a new utterance input, the new utterance input also triggers the execution of the command. An embodiment stores phrases and phrase variants as a phrase list in a data store (e.g., a permanent data store so that previously generated progress is not lost). The phrases in the list (including any phrase variants) are identified as semantically related to each other and are identified as triggers for executing commands.
[0023] As used herein, the term "utterance" refers to the speech of a user saying something to trigger the execution of an intent. In other words, an "utterance" is a set of possible spoken phrases that map to an intent that provides a command or instruction about an activity to be performed. As used herein, the term "intent" refers to an identified command embedded in or included in an utterance. In other words, an "intent" corresponds to an action that satisfies a user's verbal request. An intent may optionally have an independent variable called a "slot". As used herein, the term "slot" refers to a parameter or value associated with an intent.
[0024] It should be noted that while much of this disclosure provides examples in the context of an integrated development environment (IDE), it should be understood how the disclosed principles can be practiced in other environments and contexts, not limited to IDEs. Examples will be helpful.
[0025] Suppose a user utters the following phrase: "Go to line 5" in the context of an IDE. The spoken phrase "Go to line 5" is an example of an utterance. The "intent" or command associated with the utterance is a "go to" command or action that the computer can perform. The "slot" or parameter associated with the utterance is "line 5", meaning that the computer is to navigate to line 5 of the code.
[0026] Embodiments are able to parse an utterance into its component parts, which include "intents" and "slots." From a machine learning perspective, this parsing process can be viewed as a two-part process. To illustrate, for an incoming utterance, the machine learning engine is first presented with a classification problem, such as "what intent does this utterance belong to." In other words, the machine learning engine maps intents to the incoming utterance. Once this classification is achieved, the second problem faced by the machine learning engine is an extraction problem. To illustrate, if there are one or more slots associated with the identified intent, the machine learning engine is able to extract these slots from the utterance. Therefore, the ability to analyze an utterance is optionally viewed (from the context of machine learning) as a two-part problem involving classification and extraction. As described below, these embodiments improve upon these methods.
[0027] Examples of technology benefits, improvements and real-world applications
[0028] The following section summarizes some example improvements and practical applications provided by the disclosed embodiments. However, it should be understood that these are only examples and the embodiments are not limited to these improvements.
[0029] As previously mentioned, traditional techniques for identifying intents and slots are too labor-intensive and add a lot of manual work. Therefore, these traditional techniques are prone to errors due to manual mistakes. In addition, traditional techniques are not easy to scale. Although these techniques can work for the specific input label data provided, when unknown utterances are provided, such as utterances that deviate from the input label data, those techniques perform very poorly.
[0030] The disclosed embodiments improve upon these traditional techniques in a number of ways. A significant benefit of the disclosed embodiments is that they effectively remove the manual requirement aspect. These embodiments can now operate with significantly reduced human input. For example, the techniques described herein remove the need for human users to provide large amounts of labeled data or large amounts of input modeling data. Despite the reduced amount of input, the disclosed model is still able to learn from the data provided and provide improved results compared to traditional techniques. The embodiments also provide a generalization system that can learn and adapt over time. Utilizing this generalization system, the embodiments can perform quite well even when unknown utterances are provided.
[0031] As another benefit, the disclosed embodiments differ from the traditional "pre-train and fine-tune" approach. Instead, the embodiments have shifted to a "pre-train, prompt, and predict" paradigm, as will be described in greater detail herein. By leveraging this new paradigm, the embodiments are able to significantly improve how utterances are analyzed, how intents are determined from those utterances, and how slots are extracted from those utterances. Furthermore, the embodiments are significantly more flexible in their ability to identify utterances, intents, and slots than traditional techniques.
[0032] These embodiments can also advantageously generate variations of phrases that can be formulated. By doing so, an expanded set of related phrases can be stored in an accessible manner. Such phrases can be manipulated as a set of input-output relationships. These relationships can optionally be used as training data for other ML models. In another scenario, when a user formulates a phrase, the stored phrases can be consulted to determine whether the formulated phrase is associated with a particular intent. Significant improvements in speed and processing can be achieved by practicing these principles.
[0033] The embodiments also provide significant improvements in how utterances are transcribed, particularly in the context of an integrated development environment (IDE). Certain words in a programming language have inherent, executable meanings in the context of an IDE. Traditional speech-to-text models fail to assign appropriate meanings to those terms when dictation occurs. The embodiments provide various advantages and benefits in how utterances are analyzed so that appropriate, contextual meanings are imposed on the terms included in the utterances. Accordingly, these and many other benefits will now be described in more detail in the remainder of this disclosure.
[0034] Introduction to pre-training, hinting, and prediction examples
[0035] Using a large language model
[0036] A "Large Language Model" (LLM) is a type of machine learning (ML) algorithm that can recognize human language input and then predict and create variations of that language input. LLMs are typically tens of gigabytes in size (although they can be smaller) and can sometimes be trained using petabytes of input data (although less training data can be used). LLMs can also use a large number of parameters. Parameters are values that can change as the model learns and grows. In other words, parameters are the parts of the model that are learned over time from historical training data. Parameters typically define the basis or trick of the model with respect to a specific problem, such as a language analysis problem. Various examples of LLMs can include, but are not limited to, GPT-3LLM, BERT LLM, OPT-175B LLM, and the upcoming GPT-4LLM. Of course, there are other types of LLMs.
[0037] After the LLM is trained using the initial training data, the LLM can be used in "zero-shot scenarios" as well as "few-shot scenarios". In these cases, very little domain-customized training data (different from the initial training data provided to the LLM) is provided to the LLM. Despite this small amount of domain-customized input data, the LLM is still able to generate output based on very few different input prompts. The phrase "few shot" means that minimal data is provided as training data, while "zero-shot" means that the LLM can learn, grow, and recognize new patterns or things that the model was not exposed to before the training phase. The performance of the LLM can scale as new parameters are added to the LLM and new data is provided to the LLM.
[0038] Utilizing the “pre-training, prompting, and prediction” paradigm, the LLMs of the present disclosure are available in a pre-trained state. For example, LLMs can be used to facilitate the operation of the present disclosure. As previously described, for “pre-training,” these LLMs are trained using a large amount of training data. It should be noted that the pre-training of these LLMs is very generalized, in that the training process has no specific goals; rather, it is performed in a generalized manner. On the other hand, the pre-training phase of conventional techniques is very targeted and particularly focused on intent and slot extraction (e.g., if a machine learning observes an utterance, it is trained to identify a specific intent and a corresponding slot). Therefore, the LLMs used herein are generalized pre-trained LLMs, in which a large number of different types of language inputs are provided as training data, and in which the LLMs are trained using not only utterances, intents, and slots. In other words, the disclosed LLMs are pre-trained using one or more arbitrary libraries of language data.
[0039] During the "prompt" phase of the paradigm, an embodiment can perform calls to these pre-trained LLMs. After the prompts are fed to the pre-trained LLMs, those LLMs will then generate predictions about the intent and slots of the utterance. With respect to this prompt phase, an embodiment can use a "few-shot" learning approach. With this approach, the user provides a selected or limited number of expected input and output samples for a specific use case to the system / service (e.g., an API that feeds input to the LLMs). A sampling trigger is also provided to produce the desired output.
[0040] A two-branch approach can be adopted for intent detection of slot commands. One branch includes the ability to identify intent using a masked LLM such as BERTLLM or a library such as NLP.js. The other branch includes the ability to extract slots from utterances by querying an LLM (e.g., GPT-3).
[0041] Using the above method, an embodiment provides a selected set of prompts to the LLM. In some embodiments, the size of the prompts can be limited or restricted. That is, the size of the prompts is generally designed to be less than a maximum size threshold. In some cases, the size of the prompt depends on the complexity of the determined intent. More complex intents can utilize larger prompts, while less complex intents can utilize smaller prompts.
[0042] The LLM learns from these prompts to determine intents and slots. When a new or previously unseen utterance is provided as input, the LLM is still able to map intents to the utterance and is also able to extract slots. The LLM can also associate the intent with the intent provided in the prompt, if they are associated. In addition, the LLM is able to identify the context associated with the utterance and customize its output based on the context. As an example, assume that an utterance is received as input in the context of an integrated development environment (IDE). Here, the LLM can recognize that the utterance is received in the context of the IDE, and the LLM can customize its output based on the identified context. As a specific example, the context may include the grammatical language of the IDE, the file extensions used by the IDE, and the like.
[0043] Regarding the prediction phase, the LLM is able to receive a previously unknown utterance and then predict the intent of the utterance based on the determined context associated with the utterance. Similarly, the LLM is able to predict which slots are included in the utterance. These predictions are performed based on a limited number of prompts used to help generalize the LLM's understanding. In addition, even if the utterance does not match a previous recording of an utterance known to the LLM, the LLM is still able to extract the intent and slots of the unknown utterance.
[0044] Example Architecture
[0045] Having just described the new paradigm in a general manner, we now turn our attention to Figure 1 , which shows an example architecture 100 that can be used to implement the above-mentioned pre-training, prompting, and prediction paradigms. Architecture 100 is shown as including service 105. Service 105 can be any type of service. For example, service 105 can be a cloud service running in a cloud environment. In some cases, service 105 can be a local service operating locally on a computer. In some cases, service 105 can even be a hybrid or distributed service that is partially implemented in the cloud and partially implemented locally.
[0046] Service 105 is shown as including or at least associated with LLM 110. Service 105 may include an API for communicating with LLM 110.
[0047] LLM 110 may operate in a cloud or data center. In some cases, LLM 110 may be dedicated for use by service 105. In some cases, LLM 110 may be a shared resource. In some cases, LLM 110 may operate locally on a computer.
[0048] LLM 110 is a pre-trained LLM. That is, LLM 110 is pre-trained in the general manner described above.
[0049] According to the disclosed principles, the service 105 can receive a prompt 115, which can optionally include multiple prompts or a batch of prompts. Recall that the size of the prompt 115 is set to not exceed a maximum size threshold. As will be shown in more detail later, the prompt includes any number of prompt phrases for providing additional contextual knowledge to the LLM 110. These prompt phrases include a variety of different text or words to describe a common intent. The prompt phrase also indicates what part of the phrase constitutes a slot.
[0050] The service 105 then provides the hint 115 to the LLM 110, which analyzes the hint 115 to identify semantic relationships 120 between different text bodies and optionally generates additional (multiple) intent sets 125 and (multiple) slot sets 130. These (multiple) intents 125 and (multiple) slots 130 are designed to have the same semantics as those included in the hint 115. That is, in some cases, the hint 115 may not include meanings similar to the test data. Fine-tuning of the LLM can be used to teach this. Hints can be used to guide the LLM to perform intent detection and slot extraction, even if the hint and extraction parts are different.
[0051] That is, the LLM 110 may generate additional phrases that include these intent(s) 125 and slot(s) 130. These new phrases may use different vocabularies, but the semantic meaning of these phrases corresponds to the semantic meaning of the phrases included in the prompt 115. In other words, the intents (e.g., those generated by the LLM 110 and those included in the prompt 115) are aligned and matched to each other. Optionally, these intent(s) 125 and slot(s) 130 may be stored in a repository 135 for subsequent reference or use. An example will help provide some better context. Figure 2 Such an example is provided.
[0052] Figure 2 An example prompt 200 is shown. The prompt 200 is a text-based document that can be picked up by a human user. The prompt 200 includes multiple text phrases that are semantically related to each other and all correspond to the same intent.
[0053] To illustrate, prompt 200 is shown as including the following phrases: "Replace all occurrences of %searchTerm% with %replaceTerm%"; "Find and replace %searchTerm% to %replaceTerm%"; "Replace %searchTerm% in the project with %replaceTerm%"; and "Substitute %searchTerm% throughout the project to %replaceTerm%". These various phrases correspond to utterances that a user may optionally speak in the context of the IDE. Of course, different phrases may be spoken in different contextual scenarios.
[0054] In this example scenario, prompt 200 includes four specified variations. However, depending on the complexity of the intent, prompt 200 may include more or less than four different summary phrases. Therefore, the complexity of the prompt may optionally depend on the complexity of the intent that the LLM is summarizing. In this example scenario, four phrases are sufficient to enable the LLM to be summarized with respect to generating predictions. Using traditional machine learning techniques, those machine learning algorithms would require thousands of examples to produce a workable output result.
[0055] Note that all of these phrases are semantically related to each other in that they are all associated with the same "intent" or "command". In this example scenario, these phrases all represent various different techniques for executing the "find and replace" command / intent. The words surrounded by the "%" distinguisher tokens represent slots. That is, %search term% and %replace term% are both considered slots or parameters of the intent.
[0056] These phrases are fed as input to Figure 1 The service 105 may include an API for communicating with the LLM 110. The service 105 passes these phrases as input to the LLM 110. The LLM 110 reviews these phrases and identifies the correlation between the semantic meaning and these different variations of the "find and replace" intent. The LLM 110 can then learn from the prompt 200. That is, given the semantic meaning identified from within the prompt 200, the LLM 110 is able to generate additional text / phrases that may also conform to the semantic meaning of the phrase provided in the prompt 200.
[0057] Prompt 200 is also shown to include an "input utterance" text field, which operates as an example of LLM 110. Here, the input utterance includes the following text: "Replace all occurrences of hello with world.".
[0058] Prompt 200 identifies one slot as "hello" (eg, a search term). Prompt 200 identifies a second slot as "world" (eg, a replacement term). This prompt 200 effectively informs LLM 110 what the slots are in the "input utterance" provided above.
[0059] Figure 2 The portion of the prompt 200 is labeled "Show" 205. In other words, based on this prompt 200, the user is showing the LLM 110 what he wants the LLM 110 to do. It is then the task of the LLM 110 to further complete or fill in this prompt 200 with additional summaries / phrases that are generated based on the new utterances formulated by the user. That is, the last two lines in the portion of the prompt 200 labeled "Extract" 210 correspond to data generated by the LLM (another text entered or prompted by the user), and this data can optionally be inserted into the prompt 200 to further fill it in.
[0060] By way of further illustration, in this specific example, the user utters the following phrase: “change robust utterances to weak commands.” The service 105 receives the utterance and converts it from speech to text. The utterance text is then provided to the LLM 110 .
[0061] The LLM 110 analyzes the utterance text and attempts to identify intents and slots. In this case, the LLM 110 determines that the utterance text is as follows: "Substitute all occurrences of %searchTerm% with %replaceTerm%". The "intent" is "find and replace". The LLM also identifies slots. In this case, the %searchTerm% slot has the value "robustutterances" and the %replaceTerm% slot has the value "weak commands".
[0062] like Figure 2 As shown, the LLM may optionally append this information to the prompt 200, as shown in the extract 210 portion of the prompt 200. In this way, the prompt 200 may optionally be used as a running log for recording various alternative techniques used to stimulate or trigger the "find and replace" command or intent. The prompt 200 may optionally be stored in the repository 135.
[0063] As more utterances are generated, the LLM 110 is able to determine the semantics of those utterances, extract intents, and extract slots. The LLM 110 can then associate that utterance with other utterances that share the same semantics. Thus, even if the actual language / vocabulary used in one utterance is different from the language / vocabulary in other utterances, the LLM is still able to identify the relationship between the utterances because the intents of those utterances are determined to correspond to each other.
[0064] By way of further illustration, in this example, the user formulates the phrase "change robust utterances to weak commands." None of the previous prompt phrases included this exact language. Despite the fact that none of the previous prompt phrases had this exact language, the LLM is still able to determine that the underlying intent of the phrase "change robust utterances to weak commands" corresponds to a find and replace command. The LLM then forms a relationship between this new utterance and the prompt phrase included in prompt 200. Furthermore, the LLM supplements, augments, or adds to prompt 200 by including the new phrase, its determined intent (e.g., "Substitute all occurrences of %searchTerm% with %replaceTerm%"), and its determined slot. In this way, the LLM is able to generalize the find and replace command such that different utterances, or different methods of triggering the same command, will be associated with each other in prompt 200.
[0065] It should be understood how other prompts can be provided for other purposes, particularly for specific contexts. By way of example only, another prompt may be generated for a "close" action, such as a close window action. Another prompt may be generated for a turn action, etc. The benefit provided by LLM is the ability to form relationships between different phrases or vocabularies, even though these phrases or vocabularies are different. That is, even though the combination of words may be different, LLM can still associate different combinations of words together based on their underlying semantics. In this sense, LLM can discover variants in the vocabulary that people express, and LLM can form relationships between these variants. Using these newly formed relationships, LLM can fill in prompts / documents to record a variety of different relationships. Therefore, embodiments of the present disclosure relate to scenarios in which a user can "show" LLM what to do, rather than a "do this" type of model.
[0066] Figure 3 Shows the Figure 1 The example architecture 300 is a supplement to the architecture 100 of FIG. The architecture 300 includes a service 305 and an LLM 310, which respectively represent Figure 1 305 and LLM 110 in . Service 305 can receive utterance 315, convert the utterance into text, and then pass the text to LLM 310. Based on the above principles, LLM 310 can then determine intent 320 and slot 325 for utterance 315. Optionally, LLM 310 can attach intent 320 and slot 325 to Figure 1 The prompt 115 can be stored in the repository 135. Architecture 3 also shows the use of a speech-to-text (STT) engine 330 or STT module. Further details about the STT engine 330 will be provided later. However, in brief, the STT engine 330 receives an audio input and transcribes the audio input to generate a text output. The text output is what is provided to the LLM 310.
[0067] Speech to Text
[0068] When speech is used as the modality, there is a large amount of variability in utterances. For example, using speech-based commands (i.e., utterances) brings up a lot of examples that may not make sense, which will be described in more detail shortly. Such problems do not usually arise when using text (unless there are spelling errors) because text has a well-defined structure.
[0069] In particular, it is desirable not to extract slots that are meaningless to the current context. For example, suppose a user wants to open a file named "main.py". The current state-of-the-art speech-to-text (STT) module may recognize the utterance as "main dot pie". It is understandable that this transcription problem does not arise in a text input scenario.
[0070] While traditional STT modules are very good at transcribing text, they are deficient in interpreting what is being said from the context in which the words are said. For example, if the phrase "openmain.py" is said in the context of an IDE, what should happen is that a file named "main.py" should be opened. However, traditional STT modules cannot correctly interpret statements based on context, and will determine an incorrect slot. The cost of using an incorrect slot can be quite high. Therefore, it is desirable to provide a correct identification of slots from an utterance, particularly taking into account the context in which those utterances are said. Embodiments of the present disclosure help facilitate this operation.
[0071] Figure 4 An example prompt 400 is shown that includes a number of prompt utterance phrases typed by a user. These phrases include the following phrases: "Search for file %searchTerm%" and "Look up file %searchTerm%".
[0072] In addition to these phrases, prompt 400 also includes an STT-generated phrase, such as the following: "Search for file heylo dot pai". For clarity, this phrase is the phrase that will be generated by the STT module (e.g., from Figure 3 The STT module 330 generates language based on the actual utterances expressed by the user. The actual phrase expressed by the user is as follows: "Search for filehello.py". The STT module misinterprets the spoken phrase and generates the following text: "Search for file heylo dotpai". An embodiment can provide additional context to the LLM via prompt 400, where the context is the language that the STT engine may generate and how the language should be correctly interpreted, as described below.
[0073] In the display 405 portion of prompt 400, the user is providing instructions to the LLM that even when the slot "heylodot pai" is received as input, the input should be recognized as a variation of the actual slot "hello.py". For clarity, the user has entered what the slot should actually read using the following line: "searchTerm:hello.py". This particular prompt 400 is generated to incorporate situations where the STT module may not have correctly transcribed the user's utterance. This new line item in prompt 400 is considered additional context that can be provided to the LLM.
[0074] Now, in this example, a previously unseen utterance is provided to the service and LLM. The unseen utterance is as follows (also an output generated by STT): "Look for intex dot jay less". The actual language expressed by the user is as follows: "Look for index.js".
[0075] In this example scenario, the LLM has correctly determined the intent, which is "Find file %searchTerm%". The LLM has also correctly identified which text in the utterance corresponds to that slot; in this case, the %searchTerm% slot corresponds to the text "intex dot jay less".
[0076] Unfortunately, however, LLM incorrectly interprets, maps, predicts, or generalizes the slot language "intext dot jayless" and generates the following incorrect slot: "index.html". LLM correctly predicts "intex" as "index", but LLM incorrectly predicts "dot jay less" as ".html". The correct prediction should be "dot jay less" as ".js". This prediction problem arises because LLM does not utilize knowledge of the files that are currently available or usable in the context of the IDE. This incorrect prediction is shown by the extraction 410 portion of prompt 400. This extraction 410 should be deleted or not included in prompt 400
[0077] When generating predictions related to file usage in an IDE, it is desirable to be able to consult those files. What LLM did previously was to generate its predictions based on an arbitrary set of files that were not constrained or prioritized based on domains (e.g., IDE domains). Therefore, LLM was unrestricted in searching or generating variant predictions that could potentially map to the slot "dot jay less".
[0078] Embodiments of the present disclosure can advantageously take advantage of information available in the context of the utterance. For example, if the utterance is expressed in the context of an IDE, embodiments can take advantage of information available from the IDE to summarize predictions. For example, embodiments can take advantage of the fact that a particular set of files exists in the IDE's working directory, and the IDE is (currently) able to open only those files. Thus, these embodiments can guide the LLM's predictions based on the identified context, which can be added to prompts for delivery to the LLM. For example, embodiments can cause the LLM to identify a particular set of files in the working directory and make predictions based on that information.
[0079] Figure 5A modification to the previous prompt is shown in the form of prompt 500. Prompt 500 now displays a line item labeled filename 505. Filename 505 includes the following prompt statement: "Available file names:hello.py,hello.js,test.py,main.java". Filename 505 is a list of file names available in the IDE, and prompt 500 instructs the LLM to consult the list of files in the IDE when generating its predictions. Thus, supplemental prompt 500 provides additional knowledge or "context" for the LLM to use. Note that the LLM's output now shows that the identified search term is correctly predicted as "index.js". Note that the embodiments do not limit the search scope of the LLM; rather, the embodiments provide enhanced contextual information to enable the LLM to use a more complete set of knowledge to provide better predictions.
[0080] That is, embodiments may supplement the contextual understanding of the LLM by providing additional context within the prompt 500. In other words, additional context may be provided to the LLM by adding information to the prompt 500. In fact, the context included in the prompt 500 may be used to supplement any context that has been identified by the LLM (e.g., perhaps an IDE).
[0081] In this example, the prompt informs the LLM that the LLM should generally not extract file name information from its general knowledge base, information, or context. Instead, the prompt indicates that the LLM should give a weighted preference to the additional context provided within the scope of prompt 500 relative to the LLM's general knowledge base. In this case, the LLM should prioritize "Available file names." Thus, the LLM will weigh the contents of prompt 500 disproportionately against its general knowledge base or context.
[0082] It should be noted that the LLM can identify slots even if no slot value is specified in prompt 500. For example, assuming that prompt 500 omits the "hello.js" filename from the "Available file names" portion, the LLM will still be able to generate an "index.js" output. Although the ".js" extension is not included in prompt 500, the LLM is still aware of a variety of different file extensions. The LLM can look up these filename extensions and then identify the next similar filename extension that will map to the "dot jay less" text. About Figure 4 In prompt 400, the LLM previously used the ".html" filename extension because, based on its current knowledge base, the ".html" filename extension was the most common filename extension during the rollout. However, with the enhanced supplemental information in prompt 500, the LLM can prioritize other filename extensions over the filename extension that ranks highest based on its general knowledge.
[0083] In this way, embodiments relate to techniques for providing contextualized intent and slot extraction, wherein additional context is provided within a prompt sent to the LLM. These embodiments provide a mechanism for performing prompt writing, wherein prompts can be constructed or designed in a manner to provide supplemental, enhanced, or increased context awareness to the LLM. Additionally, the LLM does not need to be trained on a particular format for the prompt. In fact, these embodiments provide a high degree of flexibility in enabling one prompt to be swapped out for another while still enabling the LLM to recognize the new prompt.
[0084] The principles of the present disclosure can be used in a variety of different scenarios. As just one example, consider a scenario where a user wants to generate code for an application. Here, the user is currently working in an IDE. The principles of the present disclosure can simulate or operate as a virtual programming assistant that allows users to write code in a collaborative manner. Traditional coding techniques, even voice-activated coding techniques, are extremely rigid and require strict adherence to the specific syntax of the programming language.
[0085] On the other hand, embodiments of the present disclosure provide an enhanced level of flexibility with respect to input (e.g., in this case, actual code). For example, a human developer can express a variety of different utterances. The service is able to receive these utterances, convert them to text, and then feed the text as input to the LLM. Based on prompts previously provided to the LLM, the LLM is able to analyze the text-based utterances, extract intent from the utterances, and also extract slots from the utterances. Under the guidance of the human user, the LLM can then generate lines of code based on the utterances. Thus, the service and the LLM can operate as a virtual programming assistant for the user. In addition, the technology can also be used to control a code editor.
[0086] It is understandable how such an assistant can provide substantial benefits to users and to various technical fields, such as perhaps programming technology. For example, a human user now does not need to fully understand the strict syntax and programming rules that a development language may require. Instead, as long as the user understands the basic mechanics of programming logic, the user can provide his / her "intentions" to the service, and the service and LLM can help generate the actual code. Therefore, the user's understanding of programming syntax can essentially be a language-agnostic understanding.
[0087] These embodiments also allow for a very natural and intuitive interaction between the user and the service. Thus, these embodiments enable the user to maintain control of the programming experience while also providing collaborative tools to aid in the programming activity. The services of the present disclosure provide options for the user to choose from, thereby enabling the user to maintain control of the process.
[0088] In addition, embodiments of the present disclosure can help users who may have physical impairments, such as those who may have slurred speech or typing mobility issues. In scenarios where a user slurs his / her speech, LLM can be generalized via prompts to understand the user's speech patterns and generate or predict output based on a limited set of prompts.
[0089] As another example, suppose a user uses the service to help generate multiple lines of code. The user can then formulate an utterance, which may be, for example, the following: "Explain what is happening in lines 3 to 5". An embodiment can receive the utterance, convert it to text, and then pass it to the LLM. The LLM can extract the intent from the utterance. In this example scenario, the intent is an "explain" command, detailing what is happening programmatically in a certain portion of the code. The slot values can be the numbers 3 and 5. In other words, the slots are the numbers 3 and 5, and the intent is "explain between the lines" (e.g., between lines 3 and 5). The service (including the LLM; for brevity, references to the "service" should be considered to also include the LLM) can then generate an explanation of what is happening programmatically in lines 3 to 5.
[0090] Example method(s)
[0091] The following discussion now relates to a number of methods and method actions that may be performed. Although method actions may be discussed in a particular order or shown in a flowchart as occurring in a particular order, no particular order is required unless specifically stated or because an action depends on another action being completed before the action is performed.
[0092] Now turn your attention to Figure 6 , which shows a flow chart of an example method 600 for enabling a large language model (LLM) to perform contextualized intent and slot extraction. The method 600 may be used Figure 1 Architecture 100 and Figure 3 ; in addition, method 600 can be performed by service 105 / 305.
[0093] Method 600 includes accessing an LLM (e.g., Figure 1 The LLM 110 is typically pre-trained on an arbitrary corpus of language training data (action 605). Any type of general, non-specific training data may be used to train the LLM.
[0094] Action 610 includes providing a prompt (e.g., prompt 115) to the LLM that includes a limited number of prompt phrases. In some cases, the number of prompt phrases included in the prompt is based on a determined complexity level of the intent described in the prompt. In some cases, the number of prompt phrases is between 1 and about 20.
[0095] Prompt phrases share a semantic relationship with each other. That is, they correspond to the intent of the prompt description. As an example, it could be the case that different prompt phrases correspond to the following command / intent: "go to line x". The prompt phrases use different vocabularies to describe the intent of the prompt description. There are different ways to verbally recount this action using different vocabularies. Some example ways include "navigate to"; "make line x active"; "emphasize line x"; and so on.
[0096] Action 615 includes accessing a transcription of the utterance.The utterance may have been received at or processed by an STT engine, and the STT engine generates a transcription of the utterance.
[0097] Action 620 includes providing the transcript to the LLM. For example, Figure 3 Utterance 315 is shown as being provided by service 305 to LLM 310 .
[0098] Action 625 includes causing the LLM to extract the extracted intent and the extracted slot from the transcription. Figure 3 The intent 320 and slot 325 are representational.
[0099] Action 630 includes determining that the extracted intent is related to the intent of the prompt description included in the prompt. Referring to the previous "go to" example, it may be the case that the transcription of the utterance includes the following text: "place the cursor at line x". The LLM is able to analyze the text and determine the "intent" of the text. In this case, the LLM may predict that the intent appears to be a "go to" command.
[0100] Action 635 includes supplementing the prompt by adding the extracted intent and the extracted slot to the prompt, resulting in the extracted intent being identified as sharing a semantic relationship with other prompt phrases included in the prompt. For example, a prompt including "go to" language can now be supplemented with "place the cursor at line x" language. Doing so expands the knowledge base or context of the LLM and will further enable the LLM to analyze other phrases that depict similar intent.
[0101] Thus, embodiments of the present disclosure relate to various techniques for performing contextualized intent and slot extraction. Additional context can be provided to the LLM within the prompt itself, as previously described with the filename extension example. These embodiments can permanently build the LLM's knowledge base, allowing the LLM to generalize even more predictions and variants, and subsequently identify those variants.
[0102] LLM-based Discourse Enhancement
[0103] As described above, one of the benefits provided by embodiments of the present disclosure is the ability to generate variants of a body of text. One of the problems with traditional machine learning techniques is that these techniques require a large amount of user interaction in order to check and verify that the variants are correct. Embodiments of the present disclosure improve those techniques by significantly minimizing the level of human involvement. Embodiments of the present disclosure also solve problems associated with the lack of discourse input-output relationships that can be fed to other ML models that are different from LLM models. Historically, these other ML models require a large amount of input-output relationship data in order to be fully trained to generate new variants. These input-output relationships previously had to be handcrafted by human users. Embodiments of the present disclosure are able to use LLMs to generate an initial set of discourse input-output relationships, which can then be fed as input to different ML models. Optionally, an embodiment can feed seed data to the same LLM itself, and the LLM can recursively generate more variants until a stopping criterion is reached (e.g., similar suggestions after a certain point).
[0104] Embodiments of the present disclosure are capable of providing "seed data" to the LLM. "Seed data" represents a baseline descriptor or single-line descriptor for a particular command. The LLM receives this seed data and then generates any number of different variations of how a user might issue the command. As a simple example, assume that the seed data includes the following text: "cut line 8 and paste it in line 3". In accordance with the principles of the present disclosure, the LLM is constructed to generate multiple different variations of how a user might potentially issue the command. As some non-limiting examples, some variations include (but are certainly not limited to) the following: "move line 8 to line 3"; "copy line 8 and paste it in line 3 then delete line 8"; "remove line 8 and paste it in line 3". There are many different ways of expressing this same semantics or command. The LLM is constructed to generate these various potential utterances and then record them in a repository. These variations are all linked to each other and share common relationships. In this way, the embodiments are capable of performing utterance enhancement, or the generation of utterance variations. Figure 7 and 8 It is representative.
[0105] Figure 7 An example architecture 700 is shown, which is similar to the previously mentioned architecture in that it includes a service 705 and an LLM 710, both of which are configured to operate in the manner previously described. In this example scenario, the service 705 is now provided with seed data 715, which may include one or more phrases associated with a particular command. The service 705 passes the seed data 715 to the LLM 710 along with the purpose of the LLM 710 generating different variations 720 of the phrases included in the seed data 715. In some cases, the LLM 710 is tasked with generating a selected number of phrases, such as 5, 10, 15, 20, or more than 20 per iteration. In some cases, the LLM 710 is tasked with generating as many phrases as it can generate in a specified time period (e.g., run for 5 seconds and output how many variations it can have in those 5 seconds). The service 705 may optionally cause the LLM 710 to run multiple different times under the guidance of the user. For example, the LLM 710 can generate multiple variants 720. The service 705 can then submit these phrases for review by the user to ensure that the user agrees with the variants 720 generated by the LLM 710. If the user wishes to generate more variants, the service 705 can trigger the LLM 710 to generate more variants. These variants 720 can then be stored in a repository 725, forming links or relationships between the variants 720. As will be discussed in more detail later, a machine learning (ML) model 730 different from the LLM 710 can also be included in the architecture 700.
[0106] Figure 8 An example prompt 800 is shown, which may be optionally constructed or formatted, or may operate in the same manner as previously described prompts. The prompt 800 is shown to include seed data 805 that specifies the following command: "Cut line". The LLM has been provided with this seed data 805 and has generated the following variants 810: "Cut the line"; "Cut this line"; "Cut selected line"; "Cut the chosen line"; and "Cut the highlighted line". These variants may optionally be appended or added to the prompt 800 as a log record of the variants and the seed data for the command. That is, the prompt 800 may operate as a scalable or extensible record of the seed data and the variants.
[0107] Thus, the LLM generates a list of available phrases or utterances that a user can utter in order to invoke the "cut line" command. Similar operations may be performed for any other command. In this sense, embodiments are able to generate a rich repository of numerous different methods or speech modes that may be used to trigger the execution of a particular command. Thus, later, when a user is speaking and wishes to invoke the "cut line" command, any one of the phrases listed above (as well as any other phrases previously generated or optionally interpreted on the fly) may be used to trigger the execution of that command.
[0108] In some cases, after the LLM is used to generate these different phrases, an embodiment may optionally avoid further reliance on the LLM when receiving and analyzing speech input from the user. For example, if the LLM has established a list of variants of a command, then when the user actually speaks, the embodiment may avoid further use of the LLM because it is likely that one of the phrases the user is about to speak has already been generated by the LLM and can already be used to trigger the execution of the command. An embodiment may pre-generate and retain tens, hundreds, or even thousands of variants. These pre-generated variants may then be consulted when the user issues a command. In one example scenario, the goal of speech enhancement may be to generate enough variants of speech for intent detection; the goal need not be to produce an exhaustive set.
[0109] Thus, embodiments are able to generate different ways of expressing the same intent or command. Once a threshold number of these variations have been generated, embodiments may choose not to consult the LLM when an utterance is received because embodiments may have an understanding of what command the user is trying to invoke. That is, these embodiments account for these different variations of this particular utterance during the above-mentioned generation phase. Thus, embodiments are able to avoid relying on external models for verification. Instead, embodiments effectively create a huge hash map of commands. That is, a list of phrases may optionally represent a hash map of commands, where the hash map reflects different utterance inputs that can be used to trigger command execution. If an utterance is detected to be included in the list of generated phrases, these embodiments can advantageously reduce processing time by avoiding having to further consult the LLM.
[0110] Optionally, the list of utterance variants can also be used to train a smaller machine learning model. The smaller ML model (e.g., from Figure 7 The ML model 730 of the machine-generated seed can be used to generate even more variants. Thus, the LLM can be used to generate an initial set of base variants. In other words, the variant of the machine-generated seed can serve as the next seed for the LLM (or possibly to another LLM) to generate even more variants in a recursive manner.
[0111] Optionally, subsequent ML models different from the LLM model can then use the initial set of base variants to optionally generate even more variants. In other words, the LLM can be used to generate an initial set of input-output relationships in the form of an initial list of variables. These input-output relationships can then be fed as input to different ML models to further generate variants. Previously, human users were required to generate these input-output relationships. However, now LLM can be used to generate input-output relationships for ML models to operate on them. Now, humans can act as verifiers of outputs instead of generators of inputs.
[0112] When the task of the LLM is to generate additional variants after some variants have been generated, some embodiments can use the variants generated by the LLM as new seed data for the LLM. For example, assume that the LLM has generated the following variants: "Cut theline"; "Cut this line"; and "Cut selected line". Obviously, these phrases are phrases that the LLM has already generated. If the service requests the LLM to generate additional variants, some embodiments will feed in the previous first variant generated by the LLM as seed data to stimulate or trigger the generation of new variants. For example, the phrase "Cut the line" can be provided as seed data. In addition, the other two phrases "Cut this line" and "Cut selected line" can also be fed to the LLM as seed data. Therefore, in some cases, the LLM's own output can be fed as input to the LLM to trigger the generation of additional variants.
[0113] In some cases, embodiments perform fuzzy checking and may perform filtering to remove repeated phrases. Fuzzy checking or fuzzy searching involves a search technique for finding strings that have a specific pattern or whose pattern is sufficiently similar to a specified pattern. Fuzzy checking may be performed to check for spelling errors, grammatical errors, or syntax errors.
[0114] The data can also be stored persistently. If the user finishes but later wishes to generate more variants, the data file containing these variants can still be accessed and used as seed data for subsequent iterations using LLM.
[0115] Thus, a seed file may be made to include a command and a descriptor for the command, wherein the descriptor is one or more phrases that can be used to trigger the execution of the command. The seed file is fed to the LLM as input. Optionally, prompts may also be provided to the LLM to instruct the LLM how to process the seed file. For example, a prompt may be customized to instruct the LLM to generate variant phrases that, when expressed by a user, may also be used to trigger the execution of a command. Some example languages that may be included in the prompt may include the following: (i) "generate 5 other natural language ways of expressing utterances of 10 words or less related to IDE actions"; (ii) a command; and (iii) one or more example phrases of the command. Of course, prompts may be made in an alternative manner, such as by modifying the number of desired alternative phrases and / or by modifying word counts.
[0116] These variant phrases or phrase arrangements are semantically related to the descriptor phrases included in the seed file. Any number of variants may be generated. In fact, embodiments generate a large number of input-output relationships that may optionally be used as input to other ML models.
[0117] Optionally, some embodiments also include the use of stop criteria. For example, if the LLM repeatedly produces the same output, an embodiment can detect this condition and stop the LLM from continuing to process. A human user can also determine when to stop the LLM process.
[0118] Based on the above description, one can also observe how the disclosed principles can be optionally used in the field of image processing. That is, the above-mentioned fields are generally focused on text analysis. That is, the disclosed principles can also be used in the field of image analysis. For example, the principles can be used for facial recognition, object recognition, or image segmentation. Images can be transformed, for example, by changing saturation, hue, and other characteristics. A single image can be provided as a seed image, but variants of the image can be generated. For example, assume that a visible light image is provided as seed data. An embodiment can optionally use the seed data to generate images reflecting different camera modalities, such as images that may be generated by a low-light camera or images generated by a thermal imager. By feeding a visible light image, an embodiment can generate corresponding images that appear as if they were generated by different camera types or modalities. Similarly, other characteristics can also be modified. Perspective, viewpoint, or even coordinate relationships (e.g., vertical flip, horizontal flip, rotation, etc.) can also be modified by LLM, or more precisely, by a model similar to LLM but suitable for image analysis. For example, an embodiment can use the DALL-E2 model, which takes text cues and generates images from them. Feeding textual cues as seed data will generate similar textual cues, which in turn can be fed to DALL-E to generate variants of the data.
[0119] Example method for performing LLM-based utterance enhancement
[0120] The following discussion now relates to various methods and method actions that may be performed. Although method actions may be discussed in a particular order or shown in a flowchart as occurring in a particular order, no particular order is required unless specifically stated or because an action depends on another action being completed before the action is performed.
[0121] Now turn your attention to Fig. 9 , which shows a flow chart of an example method 900 for causing a large language model (LLM) to generate semantically related phrase variants for phrases included in seed data that is provided to the LLM to generate semantically related phrase variants. The method 900 may be used Figure 7 In addition, the method 900 may be performed by a service 705, which may include the LLM 710 or may be associated with the LLM 710.
[0122] Method 900 includes an action of accessing an LLM that is typically pre-trained on an arbitrary corpus of language training data (action 905). Action 910 includes feeding as input seed data including one or more phrases that are semantically related and describe a specific command. It is noteworthy that when a phrase is received as an utterance input, the utterance input triggers the execution of the specific command. In some cases, a single phrase is fed as input, while in other cases, multiple phrases are fed as input.
[0123] Action 915 includes causing the LLM to generate a plurality of phrase variants based on the phrase, wherein each phrase variant is semantically related to the phrase in the seed data. When any of the phrase variants is received as a new utterance input, the new utterance input also triggers execution of the specific phrase.
[0124] Action 920 then includes storing the phrases in the seed data and the multiple phrase variants in the data store as a phrase list. The phrase list is identified as being semantically related to each other and as a trigger for executing a particular command. Thus, the list operates as an utterance input-output relationship mapping that can be optionally provided to another ML model. In some cases, the list can also be consulted during subsequent events of receiving the utterance. These embodiments can determine whether any received utterance is included in the list. If so, the embodiment can trigger the execution of the relevant command / intent.
[0125] Optionally, the phrase list can be provided to the LLM as new seed data. The LLM can then recursively generate additional phrase variations until a stopping criterion is reached (e.g., a threshold number of phrases may be generated, or a duration may expire). The phrase variations represent different ways a user may issue a command. In some implementations, the phrase list includes more than 10, 20, 30, 40, 50, 100, or more than 1000 phrase variations.
[0126] Thus, an embodiment is able to generate training data through a semi-automatic mechanism using a large language model (LLM), while putting checks in place to ensure excellent data quality. To this end, a user can upload a seed dataset file containing a list of commands (e.g., possibly IDE commands) and a single-line description of the commands. The service / application can then iterate over the dataset, reading each command and its description. The service uses the description to generate a set of utterances that are most suitable for descriptions using the LLM. During this process, an embodiment is able to remove or filter out any repeated utterances produced by the LLM. In addition, utterances made by the LLM in the past can optionally be fuzzy checked so that duplications can be removed during the current iteration. The service can display the generated utterance variants and any fuzzy copies. The user can then choose to edit any language, or even delete language that does not make sense or is detected as being too similar to other languages.
[0127] The user can then continue to use the newly generated utterances as a seed set for more utterances that can be generated by the LLM. The LLM can continue to generate utterances for a particular command indefinitely until all possibilities have been exhausted or until a stopping criterion has been reached. Once an utterance has been generated for a particular command, the user can move to the next command, and the process continues until all commands have been covered.
[0128] At any point, if the user exits the application, the user's progress can be saved so that the user can restart generating utterances at a later point in time without losing progress. The user can also download the data in a format of his / her choice with the language-command mappings in place.
[0129] The disclosed techniques rely on minimal seed data provided to the LLM in the form of prompts to generate new utterances. Various checks and balances, such as in the form of duplicate elimination and fuzziness checks, can be implemented to help improve the quality of the data generated. The system also generates utterances exhaustively and has the ability to generate unique utterances that may not simply be paraphrases of the seed set. Thus, many benefits can be achieved by practicing the disclosed principles.
[0130] Robust speech-based language dictation in IDE
[0131] Traditional speech-to-text (STT) models trained on a specific language (e.g., English) often fail when used to attempt to construct programming source code using dictation. Typical STT models cannot easily understand the custom coding language (i.e., domain or context) inherent to software programming languages. As a result, when using STT models to attempt to generate source code through dictation techniques, the STT models produce considerable errors.
[0132] As an example, consider the following scenario. Assume that a user expresses the following phrase using an STT model in the IDE domain: "COUThello world". In this example scenario, the word "COUT" is a specific coded word that has a programming meaning or code vocabulary meaning associated with it. The STT model will likely generate the following text based on the spoken utterance: "see out hello world". The command associated with the word "COUT" will not be executed, and the use of the transcription will result in a programming error. From this example, one can easily observe how traditional STT models fail to capture the essence or meaning of a phrase that has inherent programming meaning and is expressed in the context or domain of an IDE.
[0133] Embodiments of the present disclosure are configured to address the problems faced by traditional STT models, particularly when used to transcribe specific vocabulary with potential meanings, such as programming code. To this end, embodiments access a repository of existing programming code written in the same programming language. The repository can include any number of different programs.
[0134] Then, the embodiment uses a text-to-speech (TTS) model to generate an audio recording of the written program code. In other words, the programming code included in the repository is originally in text form. The embodiment provides the text as input to the TTS model. The TTS model consumes the text and generates an audio output from the text.
[0135] Subsequently, the embodiment accesses the audio output produced by the TTS model, which is then fed to a speech-to-text (STT) model. The STT model then generates a transcription of the audio file. The embodiment then forms a link or relationship pairing between each line of actual code included in the repository and the resulting TTS to STT generated text. For clarity, each line of code is mapped to a corresponding set of text generated based on a combination of the TTS and STT models. The code pairs are therefore linked together. The first item in the pairing is actual, real programming code. The second item in the pairing represents a transcription of how the first item in the pairing sounds when read aloud. In fact, these embodiments simulate how a user dictates code without actually having anyone involved in the process of dictating the code.
[0136] Optionally, the STT model can be configured in various different ways to support different speech modes. For example, the STT model can include different language models or different dialect models to generate text. Different accents can also be used. Moreover, if the user has various different types of speech disorders or unclear speech, the embodiment can still operate using that language.
[0137] The resulting phrase pairing can then be fed into the LLM in the form of a prompt. The LLM can then use the prompt to perform any previously disclosed operations described herein. It is noteworthy that the LLM can form a specific connection or link between the programming language vocabulary included in the utterance and how the vocabulary may be expressed. Therefore, when a user expresses code with command-based meanings, those utterances can be appropriately interpreted in the context of the IDE, and when programming vocabulary is identified in the user's utterances, the programming vocabulary can be accurately generated. Therefore, an embodiment is able to map utterances to the most recent working code snippet within the context of the IDE. The phrase pairing can be provided to the LLM in a prompt. In some cases, the phrase pairing can be used to fine-tune the LLM model (e.g., it may be a GPT-3 model) to improve its accuracy and performance.
[0138] Embodiments of the present disclosure may be advantageously used to perform transcription correction based on the domain currently being operated. However, the principles may be practiced in any domain or context and are not limited to IDEs. For example, the medical field has a language that is generally considered to be highly convoluted, complex, and difficult to pronounce, particularly for pharmaceutical products. The principles of the present disclosure may operate in this domain to resolve or correct a variety of transcription errors.
[0139] In other words, these embodiments can be used to automatically correct transcripts, or more precisely, can be used to contextualize transcripts based on identified domains. In yet other words, the embodiments can impose meaning on specific vocabulary based on the identified context in which the vocabulary is used. In the context or domain of an IDE, an embodiment can perform programming language detection on a repository to determine what specific programming languages are stored in the repository. Thus, associations or relationships can be formed for specific types of programming languages.
[0140] Example architecture for speech-based language dictation in an IDE
[0141] Now turn your attention to Fig.10, which shows an exemplary architecture 1000, which may be related or may be an extension of the architectures mentioned so far. The architecture 1000 is shown as including a repository 1005, which includes source code 1005A or programming code for at least one programming language. Optionally, a variety of different programs in a variety of different programming languages can be stored in the repository 1005. Some embodiments perform language identification on the code to determine in which language the code is written.
[0142] The service 1010 accesses the code 1005A and feeds the code 1005A into a text-to-speech (TTS) model 1015. The TTS model 1015 generates one or more audio files including an audio version of the text-based code 1005A. The service 1010 then feeds these audio files into a speech-to-text (STT) model 1020.
[0143] The STT model 1020 then generates at least one file that includes a transcription of the audio recording generated by the TTS model 1015. The service 1010 then generates a file that includes various phrase pairings 1025. In particular, the lines of code from code 1005A are paired with the lines of code generated based on the output of the STT model 1020. Thus, the phrase pairings 1025 include the actual line of code and a transcription of how the line of code would sound if it were read aloud. Fig.11 It is representative.
[0144] Fig.11 A document including phrase pairs 1100 is shown, which represent phrases from Fig.10 Phrase pairing 1025. Note that this document includes 5 different pairs.
[0145] Each pairing includes an "actual" line of code and an "stt_output" transcription of the code. To illustrate, consider the first pairing. The "actual" line of code is as follows: "#Function to check whether the given\n". Previously, this line of code passed through the TTS model, and then passed through the STT model. The output of the STT model is the following phrase: "hashtagfunction to check whether the given." The service pairs this "stt_output" with the "actual" line of code.
[0146] Return to Fig.10, the service 1010 then passes the phrase pairing 1025 to the LLM 1030. The task of the LLM 1030 is to identify words that have a specific meaning in the context of the IDE and link or associate the transcription of that word to the actual meaning. Thus, if a user expresses a context-specific word, the corresponding meaning (within the IDE) can be entered into that word. Examples will be helpful.
[0147] refer to Fig.11 , one of the "actual" phrase pairing values includes the following statement: "def isArmstrong(s):\n". Another value contained in this statement: "Deaf is Armstrong,x.". In the context of the IDE, the word "def" has a specific executable or programmable meaning. When a user utters the word "def", it is desired to impose a programming language-specific executable meaning on the word. LLM is able to make such associations to ensure that the user's programming-specific vocabulary utterances are given their correct meaning in the context of the IDE.
[0148] Example method for performing speech-based language dictation
[0149] The following discussion now relates to various methods and method actions that can be performed. Although method actions may be discussed in a particular order or shown in a flowchart as occurring in a particular order, no particular order is required unless specifically stated or because an action depends on another action being completed before the action is performed.
[0150] Now turn your attention to Fig.12 , Fig.12 A flow chart of an example method 1200 for facilitating speech-based dictation of programming code in the context of an integrated development environment (IDE) so that vocabulary specific to the programming code is identifiable is shown. The method 1200 may be performed by Fig.10 1010 as shown in FIG.
[0151] Action 1205 includes feeding the programming code to a text-to-speech (TTS) model. The TTS model generates at least one audio file associated with the programming code.
[0152] Action 1210 includes feeding at least one audio file to a speech-to-text (STT) model. The STT model generates at least one transcription file associated with the at least one audio file.
[0153] Act 1215 includes mapping each corresponding line of code included in the programming code to a corresponding line of code included in the at least one transcription file, thereby generating a list of phrase pairs. The phrase pairs represent a relationship between the actual code and how the actual code sounds when read aloud.
[0154] Act 1220 includes ingesting the phrase pairing list with a large language model (LLM). The LLM identifies correlations between programming words that have a particular meaning in the context of the IDE and how the programming words sound when read aloud.
[0155] The method 1200 may optionally include an action of transcribing an utterance including programming vocabulary. The programming vocabulary identified within the utterance is converted from the STT language into a language having a specific meaning within the context of the IDE. The conversation may be based on the relevance identified by the LLM. Therefore, the disclosed embodiments are advantageously able to attribute specific contextual meanings to words identified within the utterance, where these words are identified by the embodiments as having a specific meaning.
[0156] Example Computer / Computer System
[0157] Now turn your attention to Fig.13 , which illustrates an example computer system 1300 that may include and / or be used to perform any of the operations described herein. The computer system 1300 may take a variety of different forms. For example, the computer system 1300 may be implemented as a tablet, a desktop computer, a laptop computer, a mobile device, or a stand-alone device, such as those described throughout this disclosure. The computer system 1300 may also be a distributed system including one or more connected computing components / devices that communicate with the computer system 1300.
[0158] In its most basic configuration, computer system 1300 includes a variety of different components. Fig.13 Computer system 1300 is shown including one or more processors 1305 (also referred to as “hardware processing units”) and memory 1310 .
[0159] With respect to processor(s) 1305, it should be appreciated that the functionality described herein may be performed, at least in part, by one or more hardware logic components (e.g., processor(s) 1305). For example, and without limitation, illustrative types of hardware logic components / processors that may be used include field programmable gate arrays (“FPGAs”), application specific or application specific integrated circuits (“ASICs”), application specific standard products (“ASSPs”), systems on chips (“SOCs”), complex programmable logic devices (“CPLDs”), central processing units (“CPUs”), graphics processing units (“GPUs”), or any other type of programmable hardware.
[0160] As used herein, the terms "executable module", "executable component", "component", "module" or "engine" may be a hardware processing unit or a software object, routine or method that can be executed on the computer system 1300. The different component modules, engines and services described herein may be implemented as objects or processors that execute on the computer system 1300 (e.g., as separate threads).
[0161] Memory 1310 may be physical system memory, which may be volatile, non-volatile, or some combination of the two. The term "memory" may also be used herein to refer to non-volatile mass storage devices, such as physical storage media. If computer system 1300 is distributed, processors, memory, and / or storage capabilities may also be distributed.
[0162] Memory 1310 is shown as including executable instructions 1315. Executable instructions 1315 represent instructions executable by processor(s) 1305 of computer system 1300 to perform operations of the present disclosure (eg, those described in the various methods).
[0163] Embodiments of the present disclosure may include or utilize a special or general-purpose computer including computer hardware, such as one or more processors (e.g., (multiple) processors 1305) and system memory (e.g., memory 1310), as discussed in more detail below. Embodiments also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. Such computer-readable media can be any available media that can be accessed by a general or special-purpose computer system. Computer-readable media that store computer-executable instructions in the form of data are "physical computer storage media" or "hardware storage devices." In addition, computer-readable storage media including physical computer storage media and hardware storage devices exclude signals, carriers, and propagation signals. On the other hand, computer-readable media that carry computer-executable instructions are "transmission media" and include signals, carriers, and propagation signals. Therefore, by way of example and not limitation, the current embodiment may include at least two distinct types of computer-readable media: computer storage media and transmission media.
[0164] Computer storage media (also called "hardware storage devices") are computer-readable hardware storage devices such as RAM, ROM, EEPROM, CD-ROM, RAM-based solid-state drives ("SSD"), flash memory, phase-change memory ("PCM"), or other types of memory, or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other media that can be used to store the desired program code means in the form of computer-executable instructions, data or data structures and that can be accessed by a general purpose or special purpose computer.
[0165] The computer system 1300 may also be connected (via a wired or wireless connection) to external sensors (e.g., one or more remote cameras) or devices via the network 1320. For example, the computer system 1300 may communicate with any number of devices or cloud services to obtain or process data. In some cases, the network 1320 may itself be a cloud network. In addition, the computer system 1300 may also be connected to a remote / separate computer system via one or more wired or wireless networks that is configured to perform any of the processing described with respect to the computer system 1300.
[0166] A "network" similar to network 1320 is defined as one or more data links and / or data switches capable of transmitting electronic data between computer systems, modules and / or other electronic devices. When information is transmitted or provided to a computer via a network (hardwired, wireless, or a combination of hardwired and wireless), the computer appropriately regards the connection as a transmission medium. Computer system 1300 will include one or more communication channels for communicating with network 1320. Transmission media include networks that can be used to carry data or desired program code devices in the form of computer executable instructions or in the form of data structures. In addition, these computer executable instructions can be accessed by general or special-purpose computers. Combinations of the above should also be included in the scope of computer-readable media.
[0167] Program code means in the form of computer executable instructions or data structures may be automatically transferred from transmission media to computer storage media (or vice versa) upon arrival at various computer system components. For example, computer executable instructions or data structures received over a network or data link may be cached in RAM within a network interface module (e.g., a network interface card or "NIC") and then ultimately transferred to computer system RAM and / or to less volatile computer storage media at the computer system. Thus, it should be understood that computer storage media may be included in computer system components that also (or even primarily) utilize transmission media.
[0168] Computer executable (or computer interpretable) instructions include, for example, instructions that cause a general-purpose computer, a special-purpose computer, or a special-purpose processing device to perform a specific function or group of functions. Computer executable instructions can be, for example, binary code, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in a language specific to structural features and / or method actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the above-described features or actions. On the contrary, the described features and actions are disclosed as example forms of implementing the claims.
[0169] Those skilled in the art will appreciate that the embodiments can be practiced in a network computing environment with many types of computer system configurations, including personal computers, desktop computers, laptop computers, message processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, pagers, routers, switches, etc. The embodiments can also be practiced in a distributed system environment, where local and remote computer systems linked by a network (by hardwired data links, wireless data links, or by a combination of hardwired and wireless data links) each perform tasks (e.g., cloud computing, cloud services, etc.). In a distributed system environment, program modules may be located in local and remote memory storage devices.
[0170] The present invention may be implemented in other specific forms without departing from the features of the present invention. The described embodiments are considered to be illustrative and non-restrictive in all respects. Therefore, the scope of the present invention is indicated by the appended claims rather than by the preceding description. All variations within the meaning and scope of the equivalents of the claims are included within their scope.
Claims
1. A method for causing a large language model (LLM) to generate semantically related phrase variants for phrases included in seed data, the seed data being provided to the LLM to generate the semantically related phrase variants, the method comprising: Access LLMs pre-trained on arbitrary corpora of language training data; feeding as input seed data comprising a plurality of phrases, the plurality of phrases being semantically related and describing a particular command, wherein when any one of the plurality of phrases is received as utterance input, the utterance input triggers execution of the particular command; causing the LLM to generate a plurality of phrase variants based on the plurality of phrases, wherein each phrase variant included in the plurality of phrase variants is semantically related to the plurality of phrases, and wherein when any one of the phrase variants in the plurality of phrase variants is received as a new utterance input, the new utterance input also triggers the execution of the specific phrase; as well as The plurality of phrases and the plurality of phrase variations are stored as a phrase list in a data store, wherein the plurality of phrases and the plurality of phrase variations in the phrase list are identified as being semantically related to each other and as triggers for executing the particular command.
2. The method of claim 1, wherein the phrase list is provided to the LLM as new seed data.
3. The method of claim 2, wherein the LLM recursively generates additional phrase variants until a stopping criterion is reached. The method of claim 1 , wherein the seed data represents a baseline descriptor for the particular command.
5. The method of claim 1, wherein the plurality of phrase variations represent different ways that the particular command can potentially be phrased by a user.
6. The method of claim 1, wherein causing the LLM to generate the plurality of phrase variations comprises causing the LLM to generate a selected number of phrase variations.
7. The method of claim 1, wherein causing the LLM to generate the plurality of phrase variations comprises causing the LLM to continue generating phrase variations until a specified time period expires.
8. The method of claim 1, wherein the plurality of phrase variations are submitted for user review.
9. The method of claim 1, wherein the phrase list is stored in a prompt that operates as a scalable record of seed data and variations.
10. The method of claim 1, wherein the method further comprises, after storing the list of phrases, suppressing further use of the LLM upon receipt of a subsequent utterance.
11. The method of claim 1, wherein the phrase list includes more than 10 phrases or phrase variations.
12. The method of claim 11, wherein a filtering operation is performed on the phrase list to remove duplicate phrases.
13. The method of claim 1, wherein a number of phrases included in the plurality of phrases is less than a preselected threshold number.
14. The method of claim 1, wherein the phrase list represents a hash map for the particular command, wherein the hash map reflects different utterance inputs that may be used to trigger execution of the particular command.
15. The method according to claim 1, wherein the method further comprises: The list of phrases is used to train a machine learning model tasked with generating additional phrase variations.
16. A computer system for causing a large language model (LLM) to generate semantically related phrase variants for phrases included in seed data, the seed data being provided to the LLM to generate the semantically related phrase variants, the computer system comprising: at least one processor; as well as at least one hardware storage device storing instructions executable by the at least one processor, the instructions causing the computer system to: Access LLMs pre-trained on arbitrary corpora of language training data; feeding as input seed data comprising a plurality of phrases, the plurality of phrases being semantically related and describing a particular command, wherein when any one of the plurality of phrases is received as utterance input, the utterance input triggers execution of the particular command; causing the LLM to generate a plurality of phrase variants based on the plurality of phrases, wherein each phrase variant included in the plurality of phrase variants is semantically related to the plurality of phrases, and wherein when any one of the phrase variants in the plurality of phrase variants is received as a new utterance input, the new utterance input also triggers the execution of the specific phrase; as well as The plurality of phrases and the plurality of phrase variations are stored as a phrase list in a data store, wherein the plurality of phrases and the plurality of phrase variations in the phrase list are identified as being semantically related to each other and as triggers for executing the particular command.
17. The computer system of claim 16, wherein the phrase list operates as a set of input-output relationships for a machine learning model to operate to generate additional phrase variations.
18. The computer system of claim 16, wherein a fuzzy check is performed on the phrase list to remove duplicate phrases.
19. The computer system of claim 16, wherein a filtering operation is performed on the phrase list to remove duplicate phrases.
20. A computer system for causing a large language model (LLM) to generate semantically related phrase variants for phrases included in seed data, the seed data being provided to the LLM to generate the semantically related phrase variants, the computer system comprising: at least one processor; as well as at least one hardware storage device storing instructions executable by the at least one processor, the instructions causing the computer system to: Access an LLM pre-trained on a corpus of language training data; feeding as input seed data comprising a phrase describing a command, wherein when the phrase is received as spoken input, the spoken input triggers execution of the command; causing the LLM to generate a plurality of phrase variants based on the phrase, wherein each phrase variant included in the plurality of phrase variants is semantically related to the phrase, and wherein when any one of the phrase variants in the plurality of phrase variants is received as a new utterance input, the new utterance input also triggers the execution of the phrase; as well as The phrase and the plurality of phrase variations are stored as a phrase list in a data store, wherein the phrase and the plurality of phrase variations in the phrase list are identified as being semantically related to each other and as triggers for executing the command.