Context-based attribution of meanings to utterance terms

By transcribing the audio files of the programming code and forming a phrase pairing list, and combining the large language model to identify the meaning of programming vocabulary, the cumbersomeness of intention detection and slot extraction and the accuracy of programming code transcription in the prior art are solved, and more efficient and flexible intention detection and programming code interpretation are achieved.

CN120051825APending Publication Date: 2025-05-27MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380071727.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-01-12
Filing Date
2023-09-22
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Existing intent detection mechanisms rely on regular expressions based on rules or heavy feature engineering, resulting in the process of manually labeling slots being tedious, time-consuming and unscalable, and the traditional speech-to-text model cannot capture the inherent meaning of programming languages ​​when transcribing programming code.

Method used

Generate an audio file by inputting programming code into text to speech model, and then inputting the audio file into speech to text model to generate transcript files, forming a phrase pairing list, and using a large language model to identify programming vocabulary with specific meanings in an integrated development environment.

Benefits of technology

It achieves reduced human input, improved efficiency and flexibility in intention detection and slot extraction, able to correctly interpret specific vocabulary in programming code, and improve the accuracy of speech transcription.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120051825A_ABST
    Figure CN120051825A_ABST
Patent Text Reader

Abstract

Techniques are disclosed for facilitating voice-based dictation of programming code within the context of an IDE. The programming code is fed to a text-to-speech (TTS) model. The TTS model generates an audio file associated with the code. The audio file is then fed to a voice-to-text (STT) model. The STT model generates a transcription file associated with the audio file. Each respective code line included in the programming code is mapped to a corresponding code line included in the transcription file, thereby causing generation of a phrase pair list. The phrase pairs represent a relationship between the actual code and how the actual code listens to when read out by loud sound. The LLM then ingests a list of phrase pairs. The LLM identifies a correlation between a programming vocabulary having a particular meaning within the context of the IDE and how the programming vocabulary listens to when read out by loud sound.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Current intent detection mechanisms rely on rule-based regular expressions or on supervised machine learning (ML) techniques with heavy feature engineering such as named entity recognition (NER). Such mechanisms require brainstorming complex regular expressions or organizing large labeled datasets containing an exhaustive set of possible utterances mapped to each "intent" of the system (i.e., the ways in which a user can speak to trigger a command). Along with this list of utterances, a larger list of "slot" value examples is obtained.

[0002] Supervised slot extraction requires manually labeling slots in the inside-outside-beginning (IOB) format. As a result, this process for intent detection for slot commands is very tedious, time-consuming, and non-scalable. Thus, what is needed are improved techniques that move away from the traditional "pre-train and then fine-tune" paradigm and adopt a new paradigm. Additionally, what is needed are techniques for generating variations of phrases using the new paradigm to increase the flexibility in interpreting utterances. Also needed are improved techniques for facilitating speech-based transcription for certain domains. It is expected that these various techniques provide improved results to users and increase the operational efficiency of computing systems.

[0003] The subject matter claimed herein is not limited to embodiments that solve any disadvantages or operate only in environments such as those described above. Rather, this background art is only provided to illustrate one exemplary technical field in which some embodiments described herein may be practiced. Summary of the Invention

[0004] Embodiments disclosed herein relate to systems, devices, and methods for facilitating speech-based dictation of programming code within the context of an integrated development environment (IDE) such that programming code-specific vocabulary is recognizable.

[0005] Embodiments feed programming code into a text-to-speech (TTS) model. The TTS model generates at least one audio file associated with the programming code. Embodiments feed the audio file into a speech-to-text (STT) model. The STT model generates at least one transcription file associated with the audio file. Embodiments map each respective code line included in the programming code to a corresponding code line included in the transcription file, resulting in the generation of a list of phrase pairings, where a phrase pairing represents the relationship between the actual code and how that actual code sounds when read aloud. Embodiments cause a large language model (LLM) to ingest the list of phrase pairings. The LLM identifies the correlation between programming vocabulary that has a specific meaning within the context of the IDE and how that programming vocabulary sounds when read aloud. Doing so provides various advantages, such as through improved dictation capabilities.

[0006] The present invention content is provided to introduce a selection of concepts in a simplified form, which is further described in the following detailed description. The present invention content is not intended to identify the key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.

[0007] Additional features and advantages will be set forth in the following description, and in part will be obvious from the description, or may be learned by practice of the teachings herein. The features and advantages of the present invention may be realized and obtained by means of the instrumentalities and combinations particularly pointed out in the appended claims. The features of the present invention will become more fully apparent from the following description and the appended claims, or may be learned by the practice of the invention as set forth hereinafter. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] To describe the manner in which the above and other advantages and features can be obtained, a more particular description of the subject matter briefly described above will be presented by reference to specific embodiments shown in the drawings. It should be understood that these drawings depict only typical embodiments and are not therefore to be considered limiting of its scope, and the embodiments will be described and explained with additional specificity and detail by use of the drawings, in which:

[0009] Figure 1 An example architecture for performing contextual intent detection and slot extraction based on pre-training, prompting, and prediction paradigms is shown.

[0010] Figure 2 An example of a prompt that can be provided to a large language model (LLM) is shown.

[0011] Figure 3 An extension of the architecture is shown, where now a discourse input for analysis is provided to the architecture.

[0012] Figure 4 Another example of a prompt is shown, where the prompt includes text generated by a speech-to-text engine.

[0013] Figure 5 Another example of a prompt is shown, where the prompt is constructed to include additional context information to be provided to the LLM.

[0014] Figure 6 A flowchart of an example method for performing contextual intent and slot extraction is shown.

[0015] Figure 7 An example architecture constructed to generate phrase variations is shown, including a feedback loop between a service and the LLM and between the service and an ML model.

[0016] Figure 8Shows examples of seed data that can be used to prompt an LLM.

[0017] Figure 9 Shows a flowchart of an example method for causing an LLM to generate semantically related phrase variations for phrases to be included in seed data.

[0018] Figure 10 Shows an example architecture for facilitating voice-based dictation of programming code.

[0019] Figure 11 Shows a list of phrase pairings.

[0020] Figure 12 Shows a flowchart of an example method for facilitating voice-based dictation of programming code within the context of an integrated development environment (IDE).

[0021] Figure 13 Shows an example computer system that can be configured to perform any of the disclosed operations. Detailed Description

[0022] Some embodiments relate to systems, devices, and methods for facilitating voice-based dictation of programming code within the context of an integrated development environment (IDE) such that programming code-specific vocabulary is recognizable. For example, some embodiments feed programming code into a text-to-speech (TTS) model. The TTS model generates an audio file associated with the programming code. The audio file is then fed into a speech-to-text (STT) model. The STT model generates a transcription file associated with the audio file. Embodiments map each respective code line included in the programming code to a corresponding code line included in the transcription file, resulting in the generation of a list or inventory of phrase pairings. The phrase pairings represent the relationship between the actual code and how that actual code sounds when read aloud. A large language model (LLM) then ingests the list of phrase pairings. The LLM identifies the correlation between programming vocabulary that has a particular meaning within the context of the IDE and how that programming vocabulary sounds when read aloud.

[0023] As used herein, the term "utterance" refers to a user saying something to trigger a voice for an intended execution. In other words, an "utterance" is a set of possible spoken phrases that are mapped to an intention that actively provides a command or instruction to execute. As used herein, the term "intention" refers to an identified command that is embedded or included in an utterance. In other words, an "intention" corresponds to an action that satisfies a user's spoken request. An intention may optionally have arguments referred to as "slots". As used herein, the term "slot" refers to a parameter or value associated with that intention.

[0024] Note that while most of the present disclosure provides examples within the context of an integrated development environment (IDE), it should be understood how the disclosed subject matter may be practiced in other environments and contexts without limitation. The examples may be helpful.

[0025] Suppose a user utters the following phrase: "Go to line 5" within the IDE context. The spoken phrase "Go to line 5" is an example of an utterance. The "intention" or command associated with that utterance is the "go to" command or action that the computer can perform. The "slot" or parameter associated with that utterance is "line 5", meaning the computer will navigate to line 5 of the code.

[0026] Embodiments are capable of parsing an utterance into its components, which include "intention" and "slot". From a machine learning perspective, this parsing process can be viewed as two sub-processes. For illustration, for an incoming utterance, a classification problem is first presented to the machine learning engine, such as "What intention does this utterance belong to". In other words, the machine learning engine maps an intention to the incoming utterance. Once this classification is achieved, the second problem faced by the machine learning engine is an extraction problem. For illustration, if there are one or more slots associated with the identified intention, the machine learning engine is able to extract those slots from the utterance. Thus, the ability to analyze an utterance can optionally (from the machine learning context) be viewed as a two-part problem involving classification and extraction. As described below, embodiments improve these processes.

[0027] Examples of Technical Benefits, Improvements, and Practical Applications

[0028] The following sections outline example improvements and practical applications provided by the disclosed embodiments. However, it will be understood that these are merely examples and the embodiments are not limited to these improvements.

[0029] As previously mentioned, traditional techniques for identifying intentions and slots are overly labor-intensive and add a significant amount of manual work. Thus, those traditional techniques are error-prone due to manual slips. Additionally, traditional techniques do not scale well. While those techniques may work for the specific input labeled data provided, when an unknown utterance is provided, such as an utterance that deviates from the input labeled data, those techniques perform extremely poorly.

[0030] The disclosed embodiments improve upon these traditional techniques in a variety of ways. A significant benefit of the disclosed embodiments is that they effectively remove the manual requirements aspect. Now, the embodiments can operate with significantly reduced human input. For example, the techniques described herein remove the requirement for a human user to provide large amounts of labeled data or large amounts of input modeling data. Despite the reduced amount of input, the disclosed models are still able to learn from the provided data and provide improved results compared to traditional techniques. The embodiments also provide a general system that can learn and adapt over time. With this general system, the embodiments can perform quite well even when provided with unknown utterances.

[0031] As another benefit, the disclosed embodiments move away from traditional "pre-training and fine-tuning" methods. Instead, the embodiments have shifted to a "pre-training, prompting, and prediction" paradigm, as will be described in more detail in this document. By leveraging this new paradigm, the embodiments are able to significantly improve how utterances are analyzed, how intents are determined from those utterances, and how slots are extracted from those utterances. Additionally, compared to traditional techniques, the embodiments are significantly more flexible in their ability to identify utterances, intents, and slots.

[0032] The embodiments are also able to advantageously generate variations of phrases that can be spoken. By doing so, an extended set of related phrases can be stored in an accessible manner. Such phrases can be operated on as a set of input-output relationships. These relationships can optionally be operated on as training data for other ML models. In another scenario, when a user speaks a phrase, the stored phrases can be consulted to determine if the spoken phrase is associated with a particular intent. Significant improvements in speed and processing can be achieved by practicing these principles.

[0033] The embodiments also significantly improve how utterances are transcribed, particularly in the context of an integrated development environment (IDE). Certain words in a programming language have an inherent, executable meaning within the context of an IDE. When dictation occurs, traditional speech-to-text models are unable to ascribe the appropriate meaning to those terms. The embodiments provide various advantages and benefits in how utterances are analyzed such that the appropriate contextual meaning is imposed on the terms included in the utterance. Accordingly, these and many other benefits will now be described in more detail in the remaining sections of this disclosure.

[0034] Introduction to Pre - training, Prompts, and Prediction Paradigms

[0035] Using Large Language Models

[0036] A "Large Language Model" is a machine learning (ML) algorithm that can recognize human language input and then predict and create variations of that language input. LLM sizes are typically in the tens of gigabytes (although they can be smaller) and can sometimes be trained using petabytes of input data (although less training data can be used). LLMs can also use a large number of parameters. A parameter is a value that the model can change as it learns and grows. In other words, a parameter is part of the model that learns from historical training data over time. Parameters generally define the basis or techniques of the model regarding a specific problem (such as a language analysis problem). Various examples of LLMs can include but are not limited to the GPT-3 LLM, BERT LLM, OPT-175B LLM, and the upcoming GPT-4 LLM. Of course, there are other types of LLMs.

[0037] After an LLM has been trained using initial training data, the LLM can be used in "zero-shot scenarios" as well as in "few-shot scenarios". In these scenarios, very little domain-customized training data (which is different from the initial training data provided to the LLM) is provided to the LLM. Despite this small amount of domain-customized input data, the LLM is still able to generate outputs based on very few different input prompts. The term "few-shot" means that the least amount of data is provided as training data, while "zero-shot" means that during the training phase, the LLM can learn, grow, and recognize new patterns or things that the model has not previously been exposed to. As new parameters are added to the LLM and as new data is provided to the LLM, the performance of the LLM can scale.

[0038] With the "pretrain, prompt, and predict" paradigm, the disclosed LLMs are available in a pre-trained state. For example, the LLMs can be used to facilitate the disclosed operations. As previously mentioned, to "pretrain", these LLMs are trained using a large amount of training data. It should be noted that the pre-training for these LLMs is very general because there is no specific target for the training process; instead, it is performed in a general way. On the other hand, the pre-training phase for traditional techniques is very targeted and specifically focuses on intent and slot extraction (e.g., if a machine learning observes an utterance, it is trained to identify a specific intent and the corresponding slot). Thus, the LLMs used herein are generally pre-trained LLMs, where a large amount of different types of language input are provided as training data, and where the LLMs are not trained using just utterances, intents, and slots. In other words, the disclosed LLMs are pre-trained using one or more arbitrary corpora of language data.

[0039] During the "prompting" phase of the example, the embodiments are capable of performing calls to these pre-trained LLMs. The prompts have been fed into the pre-trained LLMs, and then those LLMs will generate predictions regarding the intent and slots for the utterance. Regarding this prompting phase, the embodiments may use a "few-shot" learning method. With this method, the user provides a selected or limited number of expected input and output samples for a particular use case to the system / service (e.g., an API that feeds the input into the LLM). Sample triggers are also provided to generate the desired output.

[0040] A two-pronged approach can be adopted for intent detection of slot commands. One approach includes the ability to identify intents using a masked LLM such as BERT LLM or a library such as NLP.js. Another approach includes the ability to extract slots from the utterance by querying an LLM (e.g., GPT-3).

[0041] With the above method, the embodiments provide a selected set of prompts to the LLM. In some embodiments, the size of the prompt can be constrained or limited. That is, typically the size of the prompt is designed to be less than a maximum size threshold. In some cases, the size of the prompt depends on the determined complexity for the intent. More complex intents can utilize larger prompts, while less complex intents can utilize smaller prompts.

[0042] The LLM learns from these prompts to determine the intent and slots. When a new or previously unseen utterance is provided as input, the LLM is still able to map the intent to the utterance and is also able to extract the slots. The LLM can also associate the intent with the intents provided in the prompts if they are relevant. Additionally, the LLM is able to identify the context associated with the utterance and customize its output based on that context. As an example, assume that an utterance is received as input in the context of an integrated development environment (IDE). Here, the LLM can identify that the utterance is received within the context of the IDE, and the LLM can customize its output based on the identified context. As a specific example, the context can include a syntax-specific language for the IDE, file extensions used by the IDE, etc.

[0043] Regarding the prediction phase, the LLM is able to receive a previously unknown utterance and then predict the intent for that utterance based on the determined context associated with the utterance. Similarly, the LLM is able to predict which slots are included in the utterance. These predictions are performed based on a limited number of prompts that are used to help generalize the understanding of the LLM. Additionally, even if the utterance does not match the previous records of utterances known to the LLM, the LLM is still able to extract the intent and slots for the unknown utterance.

[0044] Example Architectures

[0045] The new paradigm has just been described in a general manner. Now, attention will be turned to Figure 1 , which shows an example architecture 100 that can be used to implement the above-described pre-training, prompting, and prediction paradigms. Architecture 100 is shown as including a service 105. Service 105 can be any type of service. For example, service 105 can be a cloud service operating in a cloud environment. In some cases, service 105 can be a local service operating on a computer. In some cases, service 105 can even be a hybrid or distributed service that is partially implemented in the cloud and partially implemented locally.

[0046] Service 105 is shown as including or at least being associated with an LLM 110. Service 105 can include an API for communicating with LLM 110.

[0047] LLM 110 can operate in the cloud or in a data center. In some cases, LLM 110 can be dedicated to being used by service 105. In some cases, LLM 110 can be a shared resource. In some cases, LLM 110 can operate locally on a computer.

[0048] LLM 110 is a pre-trained LLM. That is, LLM 110 is pre-trained in the general manner described previously.

[0049] According to the disclosed principles, service 105 is capable of receiving a prompt 115, which can optionally include multiple prompts or a batch of prompts. Recall that the size of prompt 115 is set to not exceed a maximum size threshold. The prompt includes any number of prompt phrases that are used to provide additional context knowledge to LLM110, as will be shown in more detail later. These prompt phrases include various different texts or vocabularies to describe a common intention. The prompt phrases also indicate what part of the phrase constitutes a slot.

[0050] Then, service 105 provides prompt 115 to LLM 110, and LLM 110 analyzes prompt 115 to identify semantic relationships 120 between different text bodies and optionally generates an additional set of intents 125 and a set of slots 130. These (multiple) intents 125 and (multiple) slots 130 are designed to have the same semantic meaning as the semantic meaning included in prompt 115. That is, in some cases, prompt 115 may not include a meaning similar to the test data. Fine-tuning of the LLM can be used to teach this. The prompt can be used to guide the LLM to perform intent detection and slot extraction, even if the prompt and the extraction parts are different.

[0051] That is, the LLM 110 can generate additional phrases that include these (multiple) intents 125 and (multiple) slots 130. These new phrases can use different vocabulary, but the semantic meanings used for those phrases correspond to the semantic meanings of the phrases included in the prompt 115. In other words, the intents (e.g., those generated by the LLM 110 and those included in the prompt 115) are all aligned and match each other. Optionally, these (multiple) intents 125 and (multiple) slots 130 can be stored in the repository 135 for subsequent reference or use. Examples will help provide some better context. Figure 2 Such an example is provided.

[0052] Figure 2 Example prompt 200 is shown. The prompt 200 is a text-based document that can be authored by a human user. The prompt 200 includes multiple text phrases that are semantically related to each other and all correspond to the same intent.

[0053] For illustration, the prompt 200 is shown as including the following phrases: "Replace all occurrences of %searchTerm% with %replaceTerm%"; "Find and replace %searchTerm% to %replaceTerm%"; "replace %searchTerm% in the project with %replaceTerm%"; and "Substitute all occurrences of %searchTerm% with %replaceTerm%". These various different phrases correspond to utterances that a user can optionally speak within the context of an IDE. Of course, different phrases can be spoken in different context scenarios.

[0054] In this example scenario, the prompt 200 includes four specified variations. However, depending on the complexity of the intent, the prompt 200 can include more or fewer than four different generalization phrases. Thus, the complexity of the prompt can optionally depend on the complexity of the intent on which the LLM is generalized. In this example scenario, four phrases are sufficient for the LLM to be generalized with respect to generating predictions. With traditional machine learning techniques, those machine learning algorithms would require thousands of examples to produce a working output result.

[0055] Note that all of these phrases are semantically related to each other because they are all associated with the same "intention" or "command". In this example scenario, these phrases all represent various different techniques for performing the "find and replace" command / intention. The terms surrounded by the "%" differentiator tokens represent slots. That is, both %searchTerm% and %replaceTerm% are considered slots or parameters of the intention.

[0056] These phrases are fed as input into Figure 1 service 105. Service 105 may include an API that communicates with LLM 110. Service 105 passes these phrases as input to LLM 110. LLM 110 examines these phrases and identifies the correlation between the semantic meanings and these different variations of the "find and replace" intention. Then LLM 110 can learn from prompt 200. That is, given the semantic meanings identified within prompt 200, LLM 110 is capable of generating additional text / phrases that can also conform to the semantic meanings of the phrases provided in prompt 200.

[0057] Prompt 200 is also shown as including an "input utterance" text field that operates as an example for LLM 110. Here, the input utterance includes the following text: "Replace all occurrences of hello with world".

[0058] Prompt 200 identifies one slot as "hello" (e.g., the search term). Prompt 200 identifies a second slot as "world" (e.g., the replacement term). This prompt 200 effectively notifies LLM 110 of what slots are in the "input utterance" provided above.

[0059] Figure 2 A segment of prompt 200 is labeled "display" 205. In other words, from this prompt 200, the user is showing LLM110 what is expected of LLM 110. Then the task of LLM 110 is to further complete or fill in this prompt 200 with additional generalizations / phrases that are generated based on new utterances spoken by the user. That is, the last two lines in the segment of prompt 200 labeled "extract" 210 correspond to data generated by the LLM (the other text is input or prompted by the user), and this data can optionally be inserted into prompt 200 to further fill it in.

[0060] By way of further illustration, in this specific example, the user has uttered the following phrase: "change robust utterances to weak commands". Service 105 receives the utterance and converts it from speech to text. The utterance text is then provided to LLM 110.

[0061] LLM 110 analyzes the utterance text and attempts to identify the intent and slots. In this case, LLM 110 determines the utterance text to be as follows: "Substitute all occurrences of %searchTerm% with %replaceTerm%". The "intent" is "find and replace". The LLM also identifies the slots. In this case, the %searchTerm% slot has the value "robust utterances", and the %replaceTerm% slot has the value "weak commands".

[0062] As Figure 2 shown, the LLM can optionally append this information to prompt 200, as shown in the extraction 210 segment of prompt 200. In this way, prompt 200 can optionally serve as a run log for recording various alternative techniques used to elicit or trigger the "find and replace" command or intent. Prompt 200 can optionally be stored in repository 135.

[0063] When more utterances are generated, LLM 110 is able to determine the semantic meaning of those utterances, extract the intent, and extract the slots. Then LLM 110 can associate that utterance with other utterances that share the same semantic meaning. Thus, even though the actual language / vocabulary used in one utterance is different from the language / vocabulary in other utterances, the LLM is still able to identify the relationship between the utterances because the intents of those utterances are determined to correspond to each other.

[0064] By further illustration, in this example, the user says the phrase "change robust utterances to weak commands". None of the previous prompt phrases include this exact language. Although none of the previous prompt phrases have this exact language, the LLM is still able to determine that the potential intent of the phrase "change robust utterances to weak commands" corresponds to a find and replace command. The LLM then forms a relationship between this new utterance and the prompt phrases included in Prompt 200. Additionally, the LLM supplements, augments, or adds to Prompt 200 by including this new phrase, its determined intent (e.g., "Substitute all occurrences of %searchTerm% with %replaceTerm%"), and its determined slots. In this way, the LLM is able to generalize the find and replace command such that different utterances or different ways of triggering the same command will be related to each other in Prompt 200.

[0065] It should be understood how to provide other prompts for other intents, particularly for specific contexts. By way of example only, another prompt can be generated for a "close" action such as a close window action. Another prompt can be generated for a go-to action and so on. The benefit provided by the LLM is the ability to form relationships between different phrases or vocabulary, even though those phrases or vocabulary are different. That is, even though the combination of words may be different, the LLM can still relate different combinations of words based on their underlying semantic meanings. In this sense, the LLM can discover variations in the vocabulary people say, and the LLM can form relationships between those variations. With these newly formed relationships, the LLM can populate the prompt / document to record the various different relationships. Thus, the disclosed embodiments relate to scenarios where the user can "show" the LLM what to do rather than a "do this" type of model.

[0066] Figure 3 Shows an example architecture 300 that supplements the Figure 1 architecture 100. Architecture 300 includes a service 305 and an LLM 310, which respectively represent the service 105 and the LLM 110 from Figure 1 . The service 305 is able to receive an utterance 315, convert the utterance to text, and then pass the text to the LLM 310. According to the above principles, the LLM 310 is then able to determine the intent 320 and slots 325 for the utterance 315. Optionally, the LLM 310 can append the intent 320 and slots 325 to Figure 1prompt 115 and can store the prompt 115 in the repository 135. Architecture 3 also shows the use of a speech-to-text (STT) engine 330 or STT module. Further details about the STT engine 330 will be provided shortly. However, in brief, the STT engine 330 receives an audio input and transcribes the audio input to generate a text output. The text output is the content that is provided to the LLM 310.

[0067] Speech - to - Text

[0068] When speech is used as a modality, there is a significant amount of variability in the utterances. For example, using voice-based commands (i.e., utterances) opens the door for many examples that may not make sense, as will be described in more detail shortly. When text is used (unless there are spelling mistakes), this problem generally does not surface because text has a definite structure.

[0069] In particular, it is desirable not to extract slots that have no meaning in the current context. As an example, assume the user wants to open a file named "main.py". The current state of the speech-to-text (STT) module may recognize the utterance as "main dot". It can be understood how this transcription problem does not surface in a text input scenario.

[0070] Although traditional STT modules are very good at transcribing text, they are defective in interpreting what is being said in view of the context from which the words are spoken. For example, if the phrase "open main.py" is spoken in the context of an IDE, what should happen is that the file named "main.py" should be opened. However, traditional STT modules are not able to correctly interpret the statement in view of the context, and incorrect slots will be determined. The cost of using incorrect slots can be quite high. Therefore, it is desirable to provide correct identification of slots from utterances, especially in view of the context in which those utterances are spoken. The disclosed embodiments help facilitate such an operation.

[0071] Figure 4 An example prompt 400 is shown that includes multiple prompt utterance phrases typed by a user. These phrases include the following phrases: "Search for file%searchTerm%" and "Look up file%searchTerm%".

[0072] In addition to those phrases, the prompt 400 also includes STT-generated phrases, as follows: "Search for fileheylo dot pai". For clarity, this phrase is what would be generated by the STT module (e.g., from Figure 3The language generated by the STT engine 330) and based on the actual utterance spoken by the user. The actual phrase spoken by the user is as follows: "Search for file hello.py". The STT module misinterprets the spoken phrase and generates the following text: "Searchfor file heylo dot pai". The embodiment can provide additional context to the LLM via the prompt 400, where the context is the language that the STT engine may generate and how that language should be correctly interpreted, as discussed below.

[0073] In the display 405 portion of the prompt 400, the user is providing an instruction to the LLM that even when the slot "heylodot pai" is received as input, the input should be recognized as a variant of the actual slot "hello.py". For clarity, the user has entered the following line for what the slot should actually be read as: "Search term: hello.py". This particular prompt 400 is generated to account for the situation where the STT module may not have correctly transcribed the utterance spoken by the user. This new line item in the prompt 400 is considered additional context that can be provided to the LLM.

[0074] Now, in this example, a previously unseen utterance is provided to the service and the LLM. The unseen utterance is as follows (which is also the output generated by STT): "Look for intex dot jay less". The actual language spoken by the user is as follows: "Lookfor index.js".

[0075] In this example scenario, the LLM has correctly determined the intent, which is "Find file %searchTerm%". The LLM has also correctly identified which text in the utterance corresponds to the slot; in this case, the %searchTerm% slot corresponds to the text "intex dot jay less".

[0076] However, unfortunately, the LLM incorrectly interprets, maps, predicts, or generalizes the slot language "intext dot jayless" and generates the following incorrect slot: "index.html". The LLM correctly predicts "intex" as "index", but the LLM incorrectly predicts "dot jay less" as ".html". The correct prediction should be "dot jay less" as ".js". This prediction problem occurs because the LLM does not utilize the knowledge of the files that are currently available or can be used within the context of the IDE. This incorrect prediction is shown in the extraction 410 portion of the prompt 400. The extraction 410 should be removed or should not be included in the prompt 400.

[0077] It is desired to be able to consult those files when generating predictions typically associated with the use of files in an IDE. What the LLM has done previously is that it generates its predictions based on an arbitrary set of files, which is not constrained or prioritized based on a domain (e.g., the IDE domain). As a result, the LLM searches or generates variation predictions that might potentially map to the slot "dot jay less" without limitation.

[0078] The disclosed embodiments can advantageously utilize the information available within the context of the uttered speech. For example, if the speech is uttered within the context of an IDE, the embodiments can utilize the information available from the IDE to generalize the predictions. For example, the embodiments can utilize the fact that a specific set of files exists in the working directory of the IDE, and the IDE is (currently) able to open only those files. Thus, the embodiments can guide the LLM's predictions based on the identified context, and the identified context can be added to the prompt for delivery to the LLM. For example, the embodiments can cause the LLM to identify the specific set of files in the working directory and make predictions based on that information.

[0079] Figure 5 A modification to the previous prompt is shown in the form of prompt 500. Prompt 500 now shows line items for the marked file names 505. The file name 505 includes the following prompt statement: "Available file names:hello.py,hello.js,test.py,main.java". The file name 505 is a list of file names available in the IDE, and prompt 500 instructs the LLM to consult the list of files in the IDE when generating its predictions. Thus, the supplementary prompt 500 provides additional knowledge or "context" for the LLM to use. Note that the output of the LLM now shows that the identified search term is correctly predicted as "index.js". Note that the embodiments do not limit the search scope for the LLM; rather, the embodiments provide enhanced context information to enable the LLM to use a more complete knowledge set to provide better predictions.

[0080] That is, the embodiments can supplement the LLM's context understanding by providing additional context within prompt 500. In other words, additional context can be provided to the LLM by adding information to prompt 500. In fact, the context included in prompt 500 can be used to supplement any context that the LLM has already identified (e.g., potentially the IDE).

[0081] With this example, the prompting notifies the LLM that it generally should not pull file name information from its general knowledge base, information, or context. Instead, the prompt indicates that the additional context provided within the range of prompt 500 should be given a weighted preference over the LLM's general knowledge base. In this case, "Available file names" should be prioritized by the LLM. Thus, the LLM will disproportionately weigh the content of prompt 500 over its general knowledge base or context.

[0082] It should be noted that even when no slot values are specified in prompt 500, the LLM can identify slots. For example, assume that prompt 500 omits the "hello.js" file name from the "Available file names" segment. Then the LLM will still be able to generate an "index.js" output. Although the ".js" extension is not included in prompt 500, the LLM still knows various different file extensions. The LLM can look up these file name extensions and then identify the next similar file name extension that will map to the "dot jay less" text. Regarding Figure 4 prompt 400, the LLM previously used the ".html" file name extension because, based on its current knowledge base, the ".html" file name extension was the most common file name extension during the first presentation. However, with the enhanced supplementary information in prompt 500, the LLM can prioritize other file name extensions over the file name extension that ranks highest based on its general knowledge.

[0083] In this way, the embodiments relate to techniques for providing context intent and slot extraction, where additional context is provided within the prompt sent to the LLM. The embodiments provide a mechanism for performing prompt writing, where the prompt can be engineered or designed in a way to provide supplementary, enhanced, or increased context awareness to the LLM. Additionally, the LLM does not need to be trained on a specific format of the prompt. In fact, the embodiments provide increased flexibility in enabling one prompt to be swapped out for another while still enabling the LLM to recognize the new prompt.

[0084] The disclosed principles can be used in a variety of different scenarios. Just as an example, consider the scenario where a user wants to generate code for an application. Here, the user is currently working in an IDE. The disclosed principles can simulate or operate as a virtual programming assistant that allows the user to write code in a collaborative manner. Traditional coding techniques, even voice-activated coding techniques, are extremely rigid and require strict adherence to the specific syntax of the programming language.

[0085] On the other hand, the disclosed embodiments provide an enhanced level of flexibility with respect to typed input (e.g., in this case, actual code). For example, a human developer can say various different utterances. The service can receive these utterances, convert them to text, and then feed the text as input to the LLM. Based on the prompts previously provided to the LLM, the LLM can analyze the text-based utterance, extract the intent from the utterance, and also extract the slots from the utterance. Under the guidance of the human user, the LLM can then generate lines of code based on the spoken utterance. Thus, the service and the LLM can operate as a virtual programming assistant for the user. Additionally, this technology can also be used to control a code editor.

[0086] It can be understood how such an assistant can provide substantial benefits to users and various technical fields, such as the possible programming technical field. For example, a human user no longer needs to fully understand the strict syntax and programming rules that a development language may require. Instead, as long as the user understands the basic mechanisms of programming logic, the user can provide his / her "intent" to the service, and the service and the LLM can help generate the actual code. Thus, the user's knowledge of programming syntax can essentially be a language-agnostic understanding.

[0087] The embodiments also enable a very natural and intuitive interaction between the user and the service. Thus, the embodiments enable the user to maintain control over the programming experience while also providing collaborative tools to assist in that programming activity. The disclosed service provides options for the user to choose from, enabling the user to maintain control over the process.

[0088] Furthermore, the disclosed embodiments can assist users who may have physical disabilities, such as those who may have slurred speech or input movement problems. In scenarios where the user has slurred his / her speech, the LLM can be generalized via prompts to understand the user's speech patterns and generate or predict output based on a limited set of prompts.

[0089] As another example, assume the user uses the service to help generate multiple lines of code. Then the user can say an utterance such as the following: "Explain what is happening in lines 3 to 5". The embodiment can receive this utterance, convert it to text, and then pass it to the LLM. The LLM can extract the intent from the utterance. In this example scenario, the intent is the "explain" command to elaborate on what is happening programmatically in a part of the code. The slot values can be the number 3 and the number 5. In other words, the slots are the number 3 and the number 5, and the intent is "explain between lines" (e.g., between line 3 and line 5). Then the service (which includes the LLM; for the sake of brevity, references to "service" should be considered to also include the LLM) can generate an explanation of what is happening programmatically at lines 3 to 5.

[0090] (Multiple) Example Methods

[0091] The following discussion now pertains to multiple methods and method acts that may be performed. Although method acts may be discussed in a certain order or shown in a flowchart as occurring in a particular order, a particular order is not required unless specifically stated, or required because an act depends on another act being completed before performing the act.

[0092] Attention is now turned to Figure 6 , which shows a flowchart of an example method 600 for performing context intent and slot extraction using a large language model (LLM). Method 600 may be implemented using Figure 1 architecture 100 of Figure 3 and

[0093] architecture 300 of Figure 1 ; in addition, method 600 may be performed by service 105 / 305.

[0094] Method 600 includes an act (act 605) of accessing an LLM (e.g.,

[0095] LLM 110 of Figure 1

[0094] ) that is pre-trained on any corpus of language training data. Any type of general, non-specific training data may be used to train the LLM.

[0095] Act 610 includes providing a prompt (e.g., prompt 115) to the LLM that includes a limited number of prompt phrases. In some cases, the number of prompt phrases included in the prompt is based on the determined complexity level of the intent being described in the prompt. In some cases, the number of prompt phrases is between 1 and about 20.

[0096] The prompt phrases share a semantic relationship with each other. That is, they correspond to the intent described by the prompt. As an example, it may be the case that different prompt phrases correspond to the following command / intent: "go to line x". The prompt phrases use different vocabulary to describe the intent described by the prompt. There are different ways to verbally recalculate the act using different vocabulary. Some example ways include "navigate to"; "make line x active"; "emphasize line x"; etc.

[0097] Act 615 includes accessing a transcription of the utterance. The utterance may have been received at or processed by an STT engine, and the STT engine generates a transcription of the utterance. Figure 3 Act 620 includes providing the transcription to the LLM. For example, Figure 3 utterance 315 of Figure 1 Figure 3 is shown as being provided by service 305 to LLM 310.

[0098] Action 625 includes causing the LLM to extract the extracted intent and the extracted slots from the transcription. The intent 320 and slots 325 from Figure 3 are representative.

[0099] Action 630 includes determining that the extracted intent is related to the intent described in the prompt and included in the prompt. Referring to the previous "go to" example, it could be the case that the transcription of the utterance includes the following text: "place the cursor at line x". The LLM is able to analyze this text and determine the "intent" of the text. In this case, the LLM may predict that the intent appears to be a "go to" command.

[0100] Action 635 includes supplementing the prompt by adding the extracted intent and the extracted slots to the prompt, resulting in the extracted intent being identified as sharing a semantic relationship with other prompt phrases included in the prompt. For example, a prompt that includes "go to" language can now be supplemented with "place the cursor at line x" language. Doing so expands the knowledge base or context of the LLM and will further enable the LLM to analyze other phrases that depict similar intents.

[0101] Accordingly, the disclosed embodiments relate to various techniques for performing context intent and slot extraction. Additional context can be provided to the LLM within the prompt itself, such as previously described with the file name extension example. The embodiments are capable of permanently building the knowledge base of the LLM such that the LLM can generalize even more predictions and variations and subsequently identify those variations.

[0102] LLM - based Utterance Enhancement

[0103] As indicated above, one of the benefits provided by the disclosed embodiments is the ability to generate variations of a body of text. One of the problems with traditional machine learning techniques is that these techniques require a large amount of user interaction to check and verify that the variations are correct. The disclosed embodiments improve those techniques by significantly minimizing the level of human involvement. The disclosed embodiments also address the problem associated with the lack of discourse input-output relationships that can be fed into other ML models different from the LLM model. Historically, these other ML models have required a large amount of input-output relationship data to be fully trained to generate new variations. Those input-output relationships had to be handcrafted by human users previously. The disclosed embodiments are capable of using the LLM to generate an initial set of discourse input-output relationships, which can then be fed as input into different ML models. Optionally, the embodiments can feed seed data into the same LLM itself, and the LLM can recursively generate more variations until a stop criterion is reached (e.g., similar suggestions after a point).

[0104] The disclosed embodiments are capable of providing "seed data" to the LLM. "Seed data" represents a baseline descriptor or one-liner descriptor for a particular command. The LLM receives the seed data and then generates any number of different variations of how that command can be spoken by the user. As a simple example, assume the seed data includes the following text: "cut line 8 and paste it in line 3". According to the disclosed principles, the LLM is configured to generate multiple different variations of how that command can potentially be spoken by the user. As some non-limiting examples, some variations include (but are of course not limited to) the following: "move line 8 to line 3"; "copy line 8 and paste it in line 3 then delete line 8"; "emove line 8 and paste it in line 3". There are many different ways to express the same semantic meaning or command. The LLM is configured to generate these various different possible utterances and then record them in a repository. These variations are all linked to each other and share a common relationship. In this way, the embodiments are capable of performing discourse enhancement or the generation of discourse variations. Figure 7 and Figure 8 are representative.

[0105] Figure 7Shows an example architecture 700, which is similar to the previously mentioned architectures in that the architecture 700 includes a service 705 and an LLM 710, both of which are configured to operate in the manner previously described. In this example scenario, seed data 715 is now provided to the service 705, which may include one or more phrases associated with a specific command. The service 705 passes the seed data 715 along with the purpose that the LLM 710 will generate different variations 720 of the phrases included in the seed data 715 to the LLM 710. In some cases, the task of the LLM 710 is to generate a selected number of phrases, such as 5, 10, 15, 20, or more than 20 per iteration. In some instances, the task of the LLM 710 is to generate as many phrases as it can generate within a specified time period (e.g., run for 5 seconds and output how many variations it can have in those 5 seconds). The service 705 may optionally cause the LLM 710 to run for multiple different times, possibly under the guidance of the user. For example, the LLM 710 may generate multiple variations 720. Then the service 705 may submit those phrases for the user to view to ensure that the user agrees with the variations 720 generated by the LLM 710. If the user desires that more variations be generated, the service 705 may trigger the LLM 710 to generate more variations. Then those variations 720 may be stored in a repository 725, forming associations or relationships between the variations 720. As will be discussed in more detail later, a machine learning (ML) model 730 different from the LLM 710 may also be included in the architecture 700.

[0106] Figure 8 Shows an example prompt 800, which may optionally be constructed or formatted, or may operate in the same manner as the previously described prompts. The prompt 800 is shown to include seed data 805, which specifies the following command: "Cutline". The LLM is provided with the seed data 805 and has generated the following variations 810: "Cut the line"; "Cutthis line"; "Cut selected line"; "Cut the chosen line"; and "Cut the highlightedline". These variations may optionally be appended or added to the prompt 800 as a log of the variations and the seed data for the command. That is, the prompt 800 may operate as a scalable or extensible record of the seed data and variations.

[0107] Accordingly, the LLM generates a list of available phrases or utterances that a user can speak to invoke the "Cut line" command. Similar operations can be performed for any other command. In this sense, the embodiments can generate a rich repository of many different methods or vocalizations that can be used to trigger the execution of a particular command. Thus, later, when the user is speaking and wishes to invoke the "Cut line" command, any one of the phrases listed above (and any other phrases previously generated or optionally interpreted on the fly) can be used to trigger the execution of that command.

[0108] In some cases, after the LLM is used to generate these different phrases, the embodiments can optionally suppress further reliance on the LLM when receiving and analyzing the user's spoken input. For example, if the LLM constructs a list of variations for a command, then when the user actually speaks, the embodiments can suppress further use of the LLM because it is likely that one of the phrases the user is going to speak has already been generated by the LLM and can already be used to trigger the execution of the command. Dozens, hundreds, or possibly even thousands of variations can be pre-generated and retained by the embodiments. These pre-generated variations can then be consulted when the user speaks the command. In one example scenario, the goal of utterance augmentation can be to generate sufficient variations of the utterances for intent detection; the goal does not have to be to generate an exhaustive set.

[0109] Accordingly, the embodiments can generate different ways of saying the same intent or command. Once a threshold number of those variations are generated, the embodiments can choose not to consult the LLM when receiving the utterance because the embodiments may have an understanding of what command the user is trying to invoke. That is, the embodiments have vetted these different variations of the particular utterance during the above-mentioned generation phase. As a result, the embodiments can suppress reliance on an external model for verification. Instead, the embodiments effectively create a large hash map for the command. If the utterance is detected to be included in the list of generated phrases, the embodiments can advantageously reduce the processing time by avoiding having to consult the LLM further.

[0110] Optionally, the list of utterance variations can also be used to train a smaller-scale machine learning model. The task of this smaller ML model (e.g., the ML model 730 from Figure 7 can be to generate even more variations. Thus, the LLM can be used to generate an initial basic set of variations. In other words, the machine-generated variations of the seeds can serve as the next seeds for the LLM (or possibly another LLM) to recursively generate even more variations.

[0111] Optionally, subsequent ML models different from the LLM model can optionally generate even more variations using the varying initial base set. In other words, the LLM can be used to generate an initial set of input-output relationships in the form of a varying initial inventory. Those input-output relationships can then be fed as input to different ML models for further generation of variations. Previously, human users were required to generate these input-output relationships. However, now the LLM can be used to generate input-output relationships for the ML models to operate on. Now, the human can act as a validator of the output rather than a generator of the input.

[0112] When the task of the LLM is to generate additional variations after some variations have already been generated, some embodiments can use the variations generated by the LLM as new seed data for the LLM. For example, assume that the LLM has generated the following variations: "Cutthe line"; "Cut this line"; and "Cut selected line". For clarity, these phrases are phrases that the LLM has generated. If the service requests the LLM to generate additional variations, then some embodiments will feed the previously generated first variations by the LLM as seed data to encourage or trigger the generation of new variations. For example, the phrase "Cutthe line" can be provided as seed data. Additionally, the other two phrases "Cut this line" and "Cutselected line" can also be fed as seed data to the LLM. Thus, in some cases, the LLM's own output can be fed as input to the LLM to trigger the generation of additional variations.

[0113] In some cases, embodiments perform fuzzy checks and can perform filtering to remove duplicate phrases. Fuzzy checking or fuzzy search involves search techniques for finding strings that have a particular pattern or whose pattern is sufficiently similar to a specified pattern. Fuzzy checks can be performed to check for spelling errors, grammar errors, or syntactic errors.

[0114] The data can also be persistently stored. If the user ends but then subsequently desires to generate more variations, the data file containing the variations can still be accessed and used as seed data for subsequent iterations with the LLM.

[0115] Thus, the seed file can be crafted to include a command and a descriptor for the command, where the descriptor is one or more phrases that can be used to trigger the execution of the command. The seed file is fed as input to the LLM. Optionally, a prompt can also be provided to the LLM to indicate how to process the seed file. For example, the prompt can be customized to instruct the LLM to generate variation phrases that can also be used to trigger the execution of the command when spoken by the user. Some example language that can be included in the prompt can include the following: (i) "Generate 5 other natural language ways to say the following utterances relating to IDE actions in under 10 words"; (ii) the command; and (iii) one or more example phrases for the command. Of course, the prompt can be crafted in alternative ways, such as by modifying the number of desired alternative phrases and / or by modifying the word count.

[0116] These variation phrases or phrase substitutions are semantically related to the descriptor phrase(s) included in the seed file. Any number of variations can be generated. In fact, embodiments generate a large number of input-output relationships that can optionally be used as input for other ML models.

[0117] Optionally, some embodiments also incorporate the use of a stopping criterion. For example, if the LLM repeatedly generates the same output, then the embodiment is able to detect this condition and stop the LLM from continuing to process. The human user can also determine when to stop the LLM processing.

[0118] Based on the above description, it can also be observed how the disclosed principles can optionally be used in the field of image processing. That is, the above fields typically focus on text analysis. That is to say, the disclosed principles can also be used in the field of image analysis. For example, the principles can be used for face recognition, object recognition, or image segmentation. Images can be transformed, such as by changing saturation, hue, and other characteristics. A single image can be provided as a seed image, but variations of the image can be generated. For example, assume that a visible light image is provided as seed data. Embodiments can optionally use this seed data to generate images that reflect different camera modalities, such as images that might be generated by a low-light camera or by a thermal camera. By inputting a visible light image, embodiments can generate corresponding images that appear as if they were generated by different camera types or modalities. Similarly, other characteristics can also be modified. Perspective, viewpoint, or even coordinate relationships (e.g., vertical flip, horizontal flip, rotation, etc.) can also be modified by the LLM, or rather, by a model similar to the LLM but applicable to image analysis. For example, embodiments can use the DALL-E 2 model, which takes a text prompt and generates an image from it. Feeding the text prompt as seed data will produce a similar text prompt, which in turn can be fed to DALL-E to generate variations of the data.

[0119] Example Methods for Performing LLM - based Utterance Enhancement

[0120] The following discussion now relates to a number of methods and method acts that can be performed. Although method acts may be discussed in a certain order or shown in a flowchart as occurring in a particular order, no particular order is required unless specifically stated, or because an act depends on another act that is completed before the act is performed.

[0121] Attention is now turned to Figure 9 , which shows a flowchart of an example method 900 for causing a large language model (LLM) to generate semantically related phrase variations for phrases included in seed data, where the seed data is provided to the LLM to generate semantically related phrase variations. Method 900 can be implemented using the architecture 700 of Figure 7 . In addition, method 900 can be performed by a service 705, which can include the LLM 710 or can be associated with the LLM 710.

[0122] Method 900 includes the act of accessing an LLM that is typically pre-trained on any corpus of language training data (act 905). Act 910 includes feeding the seed data as input, where the seed data includes one or more phrases that are semantically related and describe a particular command. Notably, when a phrase is received as an utterance input, the utterance input triggers the execution of a particular command.

[0123] Action 915 includes causing the LLM to generate multiple phrase variations based on a phrase, where each phrase variation is semantically related to the phrase in the seed data. When any one of the phrase variations is received as a new utterance input, the new utterance input also triggers the execution of a specific phrase.

[0124] Then action 920 includes storing the phrase in the seed data and the multiple phrase variations as a phrase list in a data store. The phrase list is identified as being semantically related to each other and as a trigger for executing a specific command. Thus, the list operates as a mapping of utterance input-output relationships that can optionally be provided to another ML model. In some cases, the list can also be consulted during subsequent events of receiving an utterance. An embodiment can determine whether any of the utterances in the received utterance are included in the list. If so, the embodiment can trigger the execution of the relevant command / intention.

[0125] Thus, embodiments are able to use a large language model (LLM) to generate training data via a semi-automatic mechanism while putting checks in place to ensure excellent data quality. To this end, a user can upload a seed dataset file containing a list of commands (e.g., perhaps IDE commands) and a one-line description of the commands. Then, the service / application can iterate over the dataset, reading each command and its description. The service uses the description to generate a set of utterances that are most suitably described using the LLM. During this process, embodiments are able to eliminate or filter out any duplicate utterances generated by the LLM. Additionally, a fuzzy check can optionally be performed on utterances made by the LLM in the past such that during the current iteration, duplicates can be removed. The service can display the generated utterance variations and any fuzzy duplicates. Then, the user has the option to select and edit any of the utterances or even delete utterances that are meaningless or detected as being too similar to other utterances.

[0126] Then, the user can continue to use the newly generated utterances as a seed set for more utterances that can be generated by the LLM. The LLM can continue to generate utterances for a specific command indefinitely until all possibilities have been exhausted or until a stop criterion has been reached. Once utterances have been generated for a specific command, the user can move on to the next command, and the process continues until all commands have been covered.

[0127] At any point in time, if the user exits the application, the user's progress can be saved such that the user can resume generating utterances at a later point in time without loss of progress. The user can also download the data in the format of his / her choice, where the utterance-command mapping is in place.

[0128] The disclosed technology relies on minimal seed data provided to the LLM in the form of prompts to generate new utterances. Various checks and balances (such as in the form of deduplication and fuzz checking) can be implemented to help improve the quality of the generated data. The system also exhaustively generates utterances and has the ability to generate unique utterances that may be more than just paraphrases of the seed set. Thus, many benefits can be achieved by practicing the disclosed principles.

[0129] Robust Speech - based Language Dictation in an IDE

[0130] Using a traditional speech-to-text (STT) model trained on a specific language (e.g., English) typically fails when used to attempt to construct programming source code using dictation. A typical STT model cannot simply understand the custom coding language (i.e., domain or context) inherent in software programming languages. As a result, when using an STT model to attempt to generate source code via dictation techniques, the STT model produces significant errors.

[0131] As an example, consider the following scenario. Suppose a user says the following phrase using an STT model in the IDE domain: "COUNThello world". In this example scenario, the term "COUT" is a specific coding term with a programming meaning or code vocabulary meaning associated with it. The STT model will likely generate the following text based on the spoken utterance: "see out helloworld". The command associated with the term "COUT" will not be executed, and the use of the transcription will result in a programming error. From this example, it can be easily observed how a traditional STT model fails to capture the essence or meaning of a phrase spoken with an inherent programming meaning and within the context or domain of an IDE.

[0132] The disclosed embodiments are configured to address the problems faced by traditional STT models, particularly when used to transcribe specific vocabulary with potential meaning, such as programming code. To this end, the embodiments access a repository of existing programming code written in the same programming language. The repository can include any number of different programs.

[0133] Then, the embodiments use a text-to-speech (TTS) model to generate an audio recording of the written program code. In other words, the programming code included in the repository is initially in text form. The embodiments feed this text as input to the TTS model. The TTS model consumes the text and generates an audio output from the text.

[0134] Subsequently, the embodiment accesses the audio output generated by the TTS model and then feeds the audio output to a speech-to-text (STT) model. The STT model then generates a transcription of the audio file. The embodiment then forms a link or relationship pair between each line of actual code included in the actual code stored in the repository and the resulting TTS-to-STT generated text. For clarity, each line of code in the code is mapped to a corresponding set of text generated based on the combination of the TTS and STT models. The code pairs are thus linked together. The first item in the pair is the actual, real programming code. The second item in the pair represents the transcription of how the first item in the pair would sound when read out loud. In effect, the embodiment simulates how a user would dictate code without anyone actually participating in dictating the code.

[0135] Optionally, the STT model can be configured in a variety of different ways to support different speech styles. For example, the STT model can include different pronunciation models or different dialect models to generate text. Different accents can also be used. Moreover, if the user has various different types of speech impairments or slurred speech, the embodiment can still operate using the language.

[0136] The resulting phrase pairs can then be fed into the LLM in the form of a prompt. The LLM can then use the prompt to perform any of the previously disclosed operations described herein. Notably, the LLM can form a specific connection or link between the programming language vocabulary included in the utterance and how that vocabulary might be spoken. Thus, when the user speaks code with a command-based meaning, those utterances can be appropriately interpreted within the context of the IDE, and when programming vocabulary is recognized in the user's utterance, that programming vocabulary can be accurately generated. Thus, the embodiment is able to map the utterance to the nearest snippet of working code within the context of the IDE. The phrase pairings can be provided to the LLM in the prompt. In some cases, the phrase pairings can be used to fine-tune the LLM model (e.g., perhaps the GPT-3 model) to improve its accuracy and performance.

[0137] The disclosed embodiments can be advantageously used to perform transcription correction based on the domain of the current operation. However, the principle can be practiced in any domain or context and is not limited to the IDE. For example, the medical field has a language that is generally considered highly convoluted, complex, and difficult to pronounce, especially for drug products. The disclosed principle can operate in that domain to address or correct various transcription errors.

[0138] In other words, the embodiments can be used to automatically correct a transcription, or rather, can be used to place the transcription in context based on the identified domain. In yet other words, the embodiments can impose meaning on a particular word based on the identified context in which the word is used. In the context or domain of an IDE, the embodiments can perform programming language detection on a repository to determine what particular programming language is stored in the repository. Thus, an association or relationship can be formed for a particular type of programming language.

[0139] Example Architecture for Speech - based Language Dictation in an IDE

[0140] Attention is now turned to Figure 10 , which shows an example architecture 1000 that can be related or may be an extension of the architectures mentioned so far. Architecture 1000 is shown as including a repository 1005 that includes source code 1005A or programming code for at least one programming language. Optionally, multiple different programs in multiple different programming languages can be stored in repository 1005. Some embodiments perform language identification on the code to determine in what language the code is written.

[0141] Service 1010 accesses code 1005A and feeds the code 1005A into a text-to-speech (TTS) model 1015. Optionally, the code can include multiple different programs written in the same programming language.

[0142] TTS model 1015 generates one or more audio files that include an audio version of text-based code 1005A. Optionally, different audio files can be generated for each of the multiple different programs. In another embodiment, a single audio file is generated for all of the multiple different programs. Service 1010 then feeds those audio files into a speech-to-text (STT) model 1020.

[0143] Then STT model 1020 generates at least one file that includes a transcription of the audio recording generated by TTS model 1015. Optionally, if there are multiple audio files, different transcription files are generated for each of the different audio files. Service 1010 then generates a file that includes various phrase pairings 1025. In particular, the lines of code from code 1005A are paired with the lines of code generated based on the output of STT model 1020. Thus, phrase pairings 1025 include the actual lines of code and the transcription of how the code line would sound if it were read out loud. Figure 11 is representative.

[0144] Figure 11 shows a file that includes phrase pairings 1100, which represent from Figure 10The phrase pairings 1025. Note that the file includes 5 different pairings.

[0145] Each pairing includes the "actual" line of code and the "stt_output" transcription for the code. For illustration, consider the first pairing. The "actual" line of code is as follows: "#Function to check whether the given\n". Previously, this line of code passed through the TTS model and then through the STT model. The resulting output of the STT model is the following phrase: "hashtag function to check whether the given". The service pairs this "stt_output" with the "actual" line of code. Thus, a phrase pairing can include a first phrase and a second phrase. The first phrase represents the actual code, and the second phrase represents how the actual code sounds when read aloud. Notably, as Figure 11 shown, the second phrase is different from the first phrase.

[0146] Returning to Figure 10 , the service 1010 then passes the phrase pairings 1025 to the LLM 1030. The task of the LLM 1030 is to identify the vocabulary that has a specific meaning within the context of the IDE and link or associate the transcription of that vocabulary with the actual meaning. Thus, if the user utters a context-specific vocabulary, the corresponding meaning (within the IDE) can be assigned to that vocabulary. That is, the embodiment is capable of receiving an utterance from the user, where the utterance includes a vocabulary that has a specific meaning within the context of the IDE. Then the embodiment can assign a meaning to the vocabulary. For example, the meaning can be one of an IDE command, a variable, or even a comment. Examples will be helpful.

[0147] Referring to Figure 11 , one of the "actual" phrase pairing values in the "actual" phrase pairing values includes the following statement: "def isArmstrong(s):\n". Other values include this statement: "Deaf is Armstrong,x". Within the context of the IDE, the term "def" has a specific executable or programming meaning. When the user utters the term "def", it is expected that the programming language-specific executable meaning will be imposed on that term. The LLM is capable of making this association to ensure that the utterance of the user's program-specific vocabulary is attributed to its correct meaning within the context of the IDE. Then the identified relevance made by the LLM is included in the prompt, similar to the prompt described herein.

[0148] Example Methods for Performing Speech - based Language Dictation

[0149] The following discussion now refers to a number of methods and method acts that may be performed. Although the method acts may be discussed in a certain order or shown in a flowchart as occurring in a particular order, no particular order is required unless specifically stated, or because an act depends on another act being completed before the act is performed.

[0150] Attention is now turned to Figure 12 , Figure 12 which shows a flowchart of an example method 1200 for facilitating speech-based dictation of programming code within the context of an integrated development environment (IDE) such that programming code-specific vocabulary is recognizable. Method 1200 may be implemented by the service 1010 shown in Figure 10 .

[0151] Act 1205 includes feeding programming code to a text-to-speech (TTS) model. The TTS model generates at least one audio file associated with the programming code.

[0152] Act 1210 includes feeding the at least one audio file to a speech-to-text (STT) model. The STT model generates at least one transcription file associated with the at least one audio file.

[0153] Act 1215 includes mapping each respective code line included in the programming code to a corresponding code line included in the at least one transcription file, resulting in the generation of a phrase pairing list. The phrase pairings represent the relationship between the actual code and how that actual code sounds when read aloud.

[0154] Act 1220 includes causing a large language model (LLM) to ingest the phrase pairing list. The LLM identifies the correlation between programming vocabulary that has a specific meaning within the context of the IDE and how that programming vocabulary sounds when read aloud.

[0155] Method 1200 may optionally include an act of transcribing utterances that include programming vocabulary. The programming vocabulary identified within the utterance is converted from the STT language to a language that has a specific meaning within the context of the IDE. The conversion may be based on the correlation identified by the LLM. Thus, the disclosed embodiments are advantageously able to ascribe a particular contextual meaning to terms identified within an utterance, where those terms are recognized by the embodiments as having a specific meaning. In some cases, an embodiment is able to receive an utterance that includes a set of programming vocabulary and determine the meaning of that set of programming vocabulary based on the context of the IDE. The embodiment may then assign the meaning to the programming vocabulary within the received utterance.

[0156] Example Computer / Computer System

[0157] Attention is now turned to Figure 13 ,Figure 13 Shown is an example computer system 1300 that can include and / or be used to perform any of the operations described herein. The computer system 1300 can take a variety of different forms. For example, the computer system 1300 can be implemented as a tablet computer, a desktop computer, a laptop computer, a mobile device, or a stand-alone device, such as those described throughout this disclosure. The computer system 1300 can also be a distributed system that includes one or more connected computing components / devices in communication with the computer system 1300.

[0158] In its most basic configuration, the computer system 1300 includes various different components. Figure 13 Illustrated, the computer system 1300 includes one or more processors 1305 (also referred to as "hardware processing units") and a storage device 1310.

[0159] Regarding the (multiple) processors 1305, it will be appreciated that the functions described herein can be performed, at least in part, by one or more hardware logic components (e.g., the (multiple) processors 1305). For example, but not limited to, illustrative types of hardware logic components / processors that can be used include field programmable gate arrays ("FPGAs"), application specific or special purpose integrated circuits ("ASICs"), application specific standard products ("ASSPs"), systems on a chip ("SOCs"), complex programmable logic devices ("CPLDs"), central processing units ("CPUs"), graphics processing units ("GPUs"), or any other type of programmable hardware.

[0160] As used herein, the terms "executable module", "executable component", "component", "module", or "engine" can refer to a hardware processing unit or a software object, routine, or method that can be executed on the computer system 1300. The different components, modules, engines, and services described herein can be implemented as objects or processors (e.g., as separate threads) executing on the computer system 1300.

[0161] The storage device 1310 can be a physical system memory, which can be volatile, non-volatile, or some combination of both. The term "memory" can also be used herein to refer to non-volatile mass storage devices, such as physical storage media. If the computer system 1300 is distributed, then the processing, memory, and / or storage capabilities can also be distributed.

[0162] The storage device 1310 is shown as including executable instructions 1315. The executable instructions 1315 represent instructions executable by the (multiple) processors 1305 of the computer system 1300 to perform the disclosed operations (such as those described in the various methods).

[0163] The disclosed embodiments may include or utilize a special-purpose or general-purpose computer including computer hardware, such as one or more processors (such as (multiple) processors 1305) and system memory (such as storage device 1310), as discussed in more detail below. The embodiments also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. Such computer-readable media can be any available media accessible by a general-purpose or special-purpose computer system. A computer-readable medium that stores computer-executable instructions in data form is a "physical computer storage medium" or "hardware storage device". In addition, computer-readable storage media including physical computer storage media and hardware storage devices exclude signals, carriers, and propagated signals. On the other hand, a computer-readable medium that carries computer-executable instructions is a "transmission medium" and includes signals, carriers, and propagated signals. Thus, by way of example and not limitation, the current embodiments may include at least two distinctly different types of computer-readable media: computer storage media and transmission media.

[0164] Computer storage media (also referred to as "hardware storage devices") are computer-readable hardware storage devices such as RAM, ROM, EEPROM, CD-ROM, RAM-based solid state drives ("SSDs"), flash memory, phase change memory ("PCM"), or other types of memory, or other optical disc storage, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store the desired program code means in the form of computer-executable instructions, data, or data structures and that can be accessed by a general-purpose or special-purpose computer.

[0165] Computer system 1300 may also be connected (via a wired or wireless connection) to external sensors (e.g., one or more remote cameras) or devices via network 1320. For example, computer system 1300 may communicate with any number of devices or cloud services to obtain or process data. In some cases, network 1320 itself may be a cloud network. In addition, computer system 1300 may also be connected to remote / separate (multiple) computer systems via one or more wired or wireless networks, the remote / separate (multiple) computer systems being configured to perform any of the processing described with respect to computer system 1300.

[0166] A "network" similar to network 1320 is defined as one or more data links and / or data switches capable of transferring electronic data between computer systems, modules, and / or other electronic devices. When information is transmitted or provided to a computer via a network (wired, wireless, or a combination of wired and wireless), the computer appropriately views the connection as a transmission medium. Computer system 1300 will include one or more communication channels used to communicate with network 1320. The transmission medium includes a network that can be used to carry data or desired program code means in the form of computer-executable instructions or in the form of a data structure. In addition, these computer-executable instructions can be accessed by a general-purpose or special-purpose computer. The above combination should also be included within the scope of computer-readable media.

[0167] Upon reaching various computer system components, program code means in the form of computer-executable instructions or data structures can be automatically transferred from the transmission medium to computer storage media (and vice versa). For example, computer-executable instructions or data structures received via a network or data link can be cached in RAM within a network interface module (e.g., a network interface card or "NIC") and then ultimately transferred to computer system RAM and / or more non-volatile computer storage media at the computer system. Thus, it should be understood that computer storage media can be included in computer system components that also (or even primarily) utilize the transmission medium.

[0168] Computer-executable (or computer-interpretable) instructions include, for example, instructions that cause a general-purpose computer, a special-purpose computer, or a special-purpose processing device to perform certain functions or groups of functions. Computer-executable instructions can be, for example, binary code, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in language specific to structural features and / or method acts, it will be understood that the subject matter defined in the appended claims need not be limited to the described features or the above acts. Instead, the described features and acts are disclosed as example forms for implementing the claims.

[0169] Those skilled in the art will appreciate that the embodiments can be practiced in network computing environments with many types of computer system configurations, including personal computers, desktop computers, laptop computers, messaging processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, pagers, routers, switches, and the like. The embodiments can also be practiced in distributed system environments where local and remote computer systems that are network-linked (either by hardwired data links, wireless data links, or by a combination of hardwired and wireless data links) each perform tasks (such as cloud computing, cloud services, etc.). In a distributed system environment, program modules can be located in both local and remote memory storage devices.

[0170] The present invention may be embodied in other specific forms without departing from its characteristics. The described embodiments are to be considered in all respects only as illustrative and not restrictive. Thus, the scope of the present invention is indicated by the appended claims rather than by the foregoing description. All changes that come within the meaning and range of equivalency of the claims are to be embraced within their scope.

Claims

1. A method for facilitating speech-based dictation of programming code within the context of an integrated development environment (IDE) such that vocabulary specific to the programming code is recognizable, the method include: feeding the programming code to a text-to-speech (TTS) model, wherein the TTS model generates at least one audio file associated with the programming code; feeding the at least one audio file to a speech-to-text (STT) model, wherein the STT model generates at least one transcription file associated with the at least one audio file; mapping each respective line of code included in the programming code to a corresponding line of code included in the at least one transcription file, thereby causing generation of a list of phrase pairs, wherein the phrase pairs represent a relationship between an actual code and how the actual code sounds when read aloud; as well as A large language model (LLM) is ingested into the phrase pairing list, wherein the LLM identifies correlations between programming words that have a particular meaning within the context of the IDE and how the programming words sound when read aloud. 2 . The method of claim 1 , wherein the programming code is stored in a repository, and wherein the repository includes a plurality of different programs in a plurality of different programming languages. 3 . The method of claim 2 , wherein the method further comprises performing language identification on the programming code to determine in what language the programming code is written.

4. The method of claim 1, wherein the phrase pair comprises a first phrase and a second phrase, the first phrase representing the actual code and the second phrase representing how the actual code sounds when read aloud, and wherein the second phrase is different from the first phrase.

5. The method according to claim 1, wherein the method further include: receiving an utterance, wherein the utterance includes words having a specific meaning within the context of the IDE; as well as The specific meaning is assigned to the word.

6. The method according to claim 5, wherein the specific meaning is one of the following: an IDE command, a variable, or a comment. The method according to claim 5 , wherein the specific meaning is a programming language executable meaning.

8. A computer system that facilitates speech-based dictation of programming code within the context of an integrated development environment (IDE) such that vocabulary specific to the programming code is recognizable, the computer system include: at least one processor; as well as at least one hardware storage device storing instructions executable by the at least one processor to cause the computer system to: feeding the programming code to a text-to-speech (TTS) model, wherein the TTS model generates at least one audio file associated with the programming code; feeding the at least one audio file to a speech-to-text (STT) model, wherein the STT model generates at least one transcription file associated with the at least one audio file; mapping each respective line of code included in the programming code to a corresponding line of code included in the at least one transcription file, thereby causing generation of a list of phrase pairs, wherein the phrase pairs represent a relationship between an actual code and how the actual code sounds when read aloud; as well as A large language model (LLM) is ingested into the phrase pairing list, wherein the LLM identifies correlations between programming words that have a particular meaning within the context of the IDE and how the programming words sound when read aloud.

9. The computer system of claim 8, wherein the programming code comprises a plurality of different programs written in the same programming language.

10. The computer system of claim 9, wherein a different audio file is generated for each of the plurality of different programs.

11. The computer system of claim 10, wherein a different transcription file is generated for each of the different audio files.

12. The computer system of claim 9, wherein a single audio file is generated for all of the plurality of different programs.

13. The computer system of claim 8, wherein the identified dependencies made by the LLM are included in a hint.

14. The computer system of claim 8, wherein execution of the instructions further causes the computer system to: receiving an utterance including a set of programming words, the meaning of the set of programming words being determined based on the context of the IDE; and The meaning is assigned to the programming words in the received utterance.

15. The computer system of claim 8, wherein the programming code is stored in a repository, and wherein the repository includes a plurality of different programs in a plurality of different programming languages.

16. The computer system of claim 15, wherein the computer system is further caused to perform language identification on the programming code to determine in what language the programming code is written.

17. At least one hardware storage device, the at least one hardware storage device comprising instructions executable by at least one processor of a computer system to cause the computer system to: feeding the programming code to a text-to-speech (TTS) model, wherein the TTS model generates at least one audio file associated with the programming code; feeding the at least one audio file to a speech-to-text (STT) model, wherein the STT model generates at least one transcription file associated with the at least one audio file; mapping each respective line of code included in the programming code to a corresponding line of code included in the at least one transcription file, thereby causing generation of a list of phrase pairs, wherein the phrase pairs represent a relationship between an actual code and how the actual code sounds when read aloud; as well as The phrase pair list is ingested by a large language model (LLM), wherein the LLM identifies correlations between programming words that have a particular meaning within the context of an integrated development environment (IDE) and how the programming words sound when spoken aloud.

18. The at least one hardware storage device of claim 17, wherein the phrase pair comprises a first phrase and a second phrase, the first phrase representing the actual code and the second phrase representing how the actual code sounds when read aloud, and wherein the second phrase is different from the first phrase.

19. The at least one hardware storage device of claim 17, wherein the programming code comprises a plurality of different programs written in the same programming language.

20. The at least one hardware storage device of claim 17, wherein the identified dependencies made by the LLM are included in a hint.