Revising big language model cues
By evaluating and revising the prompt words in LLM, the problem of the response quality of large language models being affected by prompt words was solved, resulting in higher quality and more consistent response generation and simplifying user operations.
Patent Information
- Application Number
- CN202480022170.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-05-23
- Filing Date
- 2024-04-06
- Publication Date
- 2025-11-14
AI Technical Summary
The usability of responses from Large Language Models (LLMs) is affected by the quality of cue words. Novice and expert users find it difficult to construct appropriate cue words, resulting in LLMs generating inappropriate or useless responses, which hinders their widespread adoption.
A computing system is provided that evaluates and revises prompts in an LLM using a processor, including receiving user input, generating responses, evaluation reports, and revision instructions, and iteratively optimizing the prompts to generate a high-quality final response.
It improves the quality and consistency of LLM responses, simplifies the process for users to construct effective prompts, and enhances the responsiveness of LLM.
Smart Images

Figure CN120958459A_ABST
Abstract
Description
Cross-references to related applications
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 499,045, filed April 28, 2023, the entire contents of which are incorporated herein by reference for all purposes. Background Technology
[0002] Recently, large language models (LLMs) capable of generating natural language responses in response to prompts from user input have been developed. Many state-of-the-art LLMs are based on a Transformer architecture, which utilizes word segmentation and word embeddings to represent words in the input sequence and employs a self-attention mechanism, allowing each lexical unit to potentially attend to every other lexical unit in the input sequence during neural network training. Examples of such LLMs include Generative Pre-trained Transformers (GPTs), such as GPT-3, GPT-4, and GPT-J, as well as BLOOM, LLAMA, and others. Typically, these LLMs are sequence transformation transformer models trained on a next-word prediction task. These types of LLMs are generative language models that generate output sequences for a given input sequence by repeatedly performing next-word predictions. Such models are trained on natural language corpora containing billions of words and have a parameter scale exceeding one billion parameters. These parameters are the weights in the trained neural network of the transformer. Some of these models are based on ground-valued examples and fine-tuned using human reinforcement learning or one-shot or few-shot learning. Due to their large parameter scale and, in some cases, fine-tuning, these LLMs have achieved excellent results in generative tasks, such as generating a series of chat-like responses to user prompts that substantially respond to the instructions in the prompts, in a particular writing style or format specified by the prompts, for a particular audience, and / or from the perspective of a particular author.
[0003] One drawback of these models is that the usability of the response is heavily influenced by the quality of the cue words. Both novice and expert users face the technical challenge of constructing appropriate cue words so that the LLM can respond with the level of detail, precision, perspective, and reasoning expected by the user. Sometimes, users become disappointed with the LLM when it deviates from its goals by responding to overly generalized cue words and outputting inappropriate or useless responses. Therefore, the adoption of generative LLMs has not reached the level it should have achieved in overcoming this technical challenge. Summary of the Invention
[0004] This paper provides a computational system for revising input prompts for a Large Language Model (LLM). In one example, the computational system includes at least one processor configured to render a prompt interface for a trained LLM and receive prompts from a user via the prompt interface, the prompts including instructions for the LLM to generate output. In this example, at least one processor is configured to provide the LLM with a first input including the prompts and, in response to the first input, generate a first response to the prompts via the LLM. At least one processor is configured to perform the evaluation and revision of the prompts at least in part by: evaluating the first response via the LLM according to evaluation criteria to generate an evaluation report for the first response; providing the LLM with a second input including the first prompts, the first response, the evaluation report, and prompt revision instructions to revise the prompts based on the evaluation report; and, in response to the second input, generating the revised prompts via the LLM. At least one processor is configured to provide the LLM with a final input including the revised prompts; in response to the final input, generate a final response to the revised prompts via the LLM; and output the final response to the user.
[0005] This summary is provided to introduce some concepts in a simplified form, which will be further described in the following detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Furthermore, the claimed subject matter is not limited to the implementation of solutions to any or all the shortcomings mentioned in any part of this disclosure. Attached Figure Description
[0006] Figure 1A This is a schematic diagram illustrating a computational system implemented according to a first example, which is used to revise input prompts for a Large Language Model (LLM) using a semantic function flow, which has evaluation and revision performed by an LLM program to evaluate the LLM's response and revise the prompts for the LLM.
[0007] Figure 1B This is a schematic diagram illustrating a computing system implemented according to a second example, which is used to revise LLM input prompts, including server computing devices and client computing devices.
[0008] Figure 1C This is a schematic diagram illustrating a computing system implemented according to the third example, which is used to revise LLM input prompts, wherein two LLMs are used to revise and respond to input prompts.
[0009] Figure 2 This is a schematic diagram illustrating an extended version of the semantic function flow, which has the following characteristics: Figures 1A-1CEvaluation and revision of the LLM program implementation of the computing system.
[0010] Figure 3 It shows Figure 2 A detailed view of the first part of the assessment and revisions shown.
[0011] Figure 4 yes Figure 3 The second part, which continues the discussion, provides a detailed view of the evaluation and revisions.
[0012] Figure 5 yes Figure 4 The continuation, and showed Figure 3 and 4 The third part of the assessment and revision.
[0013] Figure 6 It shows Figures 1A-1C The example graphical user interface of the computing system shows background response evaluation and prompt word revision.
[0014] Figure 7 It shows Figures 1A-1C Another example of a graphical user interface for a computing system shows response evaluation and prompt word revision in response to user input.
[0015] Figure 8 A flowchart is shown, based on an example implementation, of a method for revising LLM input prompts.
[0016] Figure 9 It shows Figures 1A-1C A schematic diagram of an example computing environment in which the computing system can be implemented. Detailed Implementation
[0017] To solve the above problems, Figure 1AA schematic diagram of a computing system 10 according to a first example implementation is shown. This computing system 10 is used to revise Large Language Model (LLM) input prompts using evaluation and revision logic provided by a semantic function flow 12. The computing system 10 includes a computing device 14 having at least one processor 16, a memory 18, and a storage device 20. In the first example implementation, the computing system 10 takes the form of a single computing device 14, with an LLM program 22 stored in the storage device 20. This program can be executed by at least one processor 16 to perform various functions (including evaluation and revision of LLM input prompts) according to the semantic function flow 12. At least one processor 16 can be configured to render a prompt interface 24 for a trained LLM 26. The LLM 26 may include, for example, a generative pre-trained transformer 26A. The generative pre-trained transformer may be a sequence-to-sequence transformer 26A including an encoder and a decoder, trained based on a next-word prediction task to predict the next word in a sequence. LLM 26 may also include an embedding module 88 and an image model 90 configured to generate input vectors for the LLM, which include input word sequences and relevant embeddings generated from the input text and images into the model, as described below. In some cases, the prompt word interface 24 may be part of a graphical user interface (GUI) 28 for accepting user input and presenting information to the user. In other cases, the prompt word interface 24 may be presented in a non-visual form, such as an audio interface for receiving and / or outputting audio, as might be used in a digital assistant. In yet another example, the prompt word interface 24 may be implemented as a prompt word interface application programming interface (API). In this configuration, input to the prompt word interface can be implemented via API calls from the software program to the prompt word interface API, and output can be returned in the API response from the prompt word interface API to the calling software program. It should be understood that the distributed processing strategy can be implemented to execute the software described herein, and therefore at least one processor 16 may include multiple processing devices, such as cores of a central processing unit, coprocessors, graphics processing units, field-programmable gate array (FPGA) accelerators, tensor processing units, etc., and these multiple processing devices may reside within one or more computing devices and may be connected, for example, via interconnect (when within the same device) or via packet-switched network links (when in multiple computing devices). Thus, at least one processor 16 can be configured to execute a prompt word interface API (e.g., prompt word interface 24) for a trained LLM 26.
[0018] Typically, at least one processor 16 can be configured to receive a prompt 30 from the user via a prompt word interface 24 (or, in some implementations, a prompt word interface API), the prompt word 30 including instructions for generating output for the LLM 26, which will be referenced below. Figures 2 to 5 A more detailed description follows. It should be understood that the prompt words can also be generated and received by a software program, rather than directly from a human user. Briefly, LLM 26 can be configured to receive prompt word 30 and generate a first response 32. Optionally, the first response 32 can be output to the user, and if the user is not satisfied with the first response 32, they can optionally request a revision 33. Alternatively, LLM program 22 can be configured to self-evaluate and revise without further input from the user, for example, by performing a predetermined number of iterations. If evaluation and revision are to be performed, LLM 26 evaluates the first response 32 based on evaluation criteria, thereby generating an evaluation report 34. The evaluation report 34, the first response 32, and the prompt word 30 are fed back to LLM 26 with instructions to revise the prompt word 30, thereby improving the first response 32. For example, if the user is not satisfied with the next response generated based on the revised prompt word, or the predetermined number of iterations has not been performed, or the evaluation report 34 of the current response does not meet a predefined evaluation threshold, then evaluation and revision will be performed at least one more iteration. However, if the current response is deemed acceptable by the user or predefined criteria, the revised response 36 is output to the user.
[0019] Go to Figure 1B The diagram illustrates a computing system 110 implemented according to a second example, wherein the computing system 110 includes a server computing device 38 and a client computing device 40. Here, the server computing device 38 and the client computing device 40 may include corresponding processors 16, memory 18, and storage devices 20. Figure 1A Descriptions of identical components will not be repeated. Client computing device 40 can be configured to execute client program 42 via its processor 16, thereby presenting prompt interface 24. Client computing device 40 can be responsible for communication between the user operating client computing device 40 and server computing device 38 executing LLM program 22 and containing LLM 26 via application programming interface (API) 44 of LLM program 22. Client computing device 40 can take the form of a personal computer, laptop computer, tablet computer, smartphone, smart speaker, etc. (See above reference) Figure 1A The same evaluation and revision process described can be performed, except that, in this case, prompt 30, request 33, first response 32 and revised response 36 can be transmitted between server computing device 38 and client computing device via a network such as the Internet.
[0020] Go to Figure 1C A schematic diagram of a computing system 210 implemented according to the third example is shown. For simplicity, it is... Figure 1A and Figure 1BDescriptions of similar components will not be repeated. Here, two LLM revisions are used to provide improved responses instead of... Figure 1A and 1B In contrast, only one LLM is used. It should be understood that... Figure 1A single device implementation or Figure 1B Both client-server implementations can use one or two LLMs. Figure 1C A single computing device 14 is shown only as an example.
[0021] Here, similar to Figure 1A At least one processor 16 can be configured to render a cue word interface 24 for the first trained LLM 26 and to receive cue words 30 from a user via the cue word interface 24, the cue words 30 including instructions for the first LLM 26 to generate output. That is, the first LLM 26 can be a model responsible for generating output (e.g., responses 32 and 36) for the user in response to a given received cue word 30. To efficiently utilize resources, the first LLM 46 can be a conventional model or a computationally less intensive model that requires fewer resources to run and can be more limited in its capabilities than the second trained LLM 48. More specifically, the second LLM 26 can have a larger parameter scale than the first LLM 46, meaning there are more weights between the nodes of the model, and the second LLM 48 may have a higher average computational cost during inference than the first LLM 46. The same evaluation and revision described above can be performed by the second LLM 48 on the first response 32 generated by the first LLM 46 to output a revised and improved prompt, which is then input to the first LLM 46 to generate the next response 36. In this way, the relatively large capacity of the second LLM 48 can be reserved for improving the prompt 30, and the first LLM 46 is able to fully process the prompt 30 to generate an acceptable response 36, without wasting resources by having the second LLM 48 perform the entire process, and without accepting substandard output from the first LLM 46.
[0022] Figure 2 The diagram shows... Figure 1A-1CThe evaluation and revision shown are performed through 1 to N iterations, thereby generating corresponding first-generation to N-generation responses. As shown, the first iteration 50 includes a first response generation phase, a response evaluation phase, and a cue word revision phase, and the remaining iterations (the second iteration 52 shown here) repeat these phases as revised response generation phases, response evaluation phases, and cue word revision phases, until the Nth iteration 54, in which the final response 56 (Nth generation) is output in response to the final input 58. In the first response generation phase, at least one processor 16 is configured to provide a first input 60 including cue word 30 to the LLM 26, and in response to the first input 60, generate a first response 32 to cue word 30 via the LLM 26 (which can be output as a first-generation response 62). At least one processor 16 is configured to perform the evaluation of response 32 and the revision of prompt 30 at least in part during the response evaluation phase and the prompt revision phase by: (a) evaluating the first response 32 according to evaluation criteria 64 to generate an evaluation report 34 of the first response 32 via LLM 26; and (b) providing the LLM 26 with a second input 66 including a prompt revision instruction 68 to generate a revised prompt 69 based on the evaluation report 34. For example, in addition to the prompt revision instruction 68, the second input 66 may also include the initial prompt 30, the first response 32, and the evaluation report 34. In the above process, in response to the second input 66, at least one processor 16 executing the LLM 26 is configured to generate the revised prompt 69 via the LLM 26.
[0023] One or more intermediate iterations of these stages can be performed, as shown in the second iteration 52. Similar to the first iteration 50, at least one processor 16 is configured to provide a response revision instruction 70 to the LLM 26 to generate a revised response 36 (which can be output as the second-generation response 72) based on the revised prompt 69; evaluate the generated revised response 32 according to evaluation criteria 64 to generate an evaluation report 34 via the LLM 26 for the previously revised response 36; and provide a prompt revision instruction 68 to the LLM 26 to generate a revised prompt 69. It should be understood that the prompt revision instruction 68, response revision instruction 70, evaluation report 34, revised prompt 69, and revised response 36 will typically differ between iterations. In the final (Nth) iteration 54, at least one processor 16 is configured to provide the LLM 26 with a final input 58 including a recently generated revised prompt 69 (e.g., from the second iteration 52), and in response to the final input 58, generate a final response 56 to the revised prompt 69 via the LLM 26, and output the final response 56 to the user (in some implementations, via the prompt interface API). If the user decides to conduct further evaluation and revision after reviewing the final response 56, the user can restart. Figure 2 The process is shown.
[0024] Typically, the evaluation and revision of prompts are performed iteratively over multiple iterations, as will be discussed below. Figure 6 and Figure 7 As shown, the number of iterations can be user-defined. Alternatively, the number of iterations can be a predefined number, such as 1, 2, 3, 4, or 5. The number of iterations can also be set programmatically, or iterations can continue until an evaluation threshold is met, for example, for a response, a certain evaluation criterion exceeds a specific value. For example, a response can be iteratively optimized for a politeness evaluation criterion until it meets, for example, a politeness threshold. Of course, to avoid wasting computational resources, a maximum number of iterations can be set, which can vary based on user level (paid vs. non-paid customers, developers vs. end-users, etc.). This will be discussed below. Figure 6 As shown, in one implementation, at least one processor 16 is configured to output a final response 56 generated after multiple iterations (e.g., via a display or voice) to the user, without outputting any intermediate responses. In another implementation, as will be discussed below... Figure 7 As shown, intermediate responses can be presented to the user, such as... Figure 2 The dashed lines indicate the first-generation response 62 and the second-generation response 72.
[0025] Figures 3 to 5 Detailed illustration Figure 2The overall assessment and revisions are shown in three corresponding views. First, go to... Figure 3 The diagram illustrates a cue word generation module 74 with an evaluation and revision engine 76 of the LLM 26. The cue word generation module 74 is configured to present a cue word interface 24 displayed in a GUI 28 and receive input data from a user that constitutes cue words 30. Cue words 30 include text instructions 78 from the user, providing control input 80 for the LLM 26. Cue words 30 also include contextual input 82, which may take the form of, for example, an image 84 and / or text 86. Additionally or alternatively, cue words may include audio and / or video. It should be understood that the LLM 26 can be multimodal, capable of receiving at least two input modes. For example, the LLM 26 can be configured to receive a primary input mode (such as the text mode described above) and one or more secondary input modes (such as an image mode, an audio mode, or a video mode). To achieve this, the LLM 26 can be trained on a corpus of text and image data (and / or appropriate audio and / or video data) using a cross-modal encoder, as described below. Figure 7 An example of multimodal input for an LLM is shown, which illustrates prompts 30 including article text, article images, and text instructions.
[0026] Next, the cue word 30 is passed to the embedding module 88, which computes an embedding for each input pattern. The embedding module 88 is depicted as part of the LLM 26, but in an alternative implementation, it may be partially or completely incorporated into the cue word generation module, such that the embedding representation is output from the cue word generation module to the LLM 26. An image model 90 is used to convert the context image 84 into a context image embedding 92. A tokenizer 94 is provided to convert the context text 86 into a context text embedding 96. The tokenizer 94 also generates a text instruction embedding 98 based on the text instruction 78. The context image embedding 92, the context text embedding 96, and the text instruction embedding 98 are concatenated to form a concatenated cue word input vector 100, which is fed to the LLM 26 as a first input 60. In response to the first input 60, the LLM 26 generates a first response 32. The first response 32 is passed back to the prompt word generation module 74, where it can be displayed or otherwise presented to the user, or simply stored in memory for background processing. During the response evaluation phase, the first response 32 is passed as context 102 to the next prompt word 104. In one implementation shown by solid lines, the next prompt word 104 may also include the previous context 82 and previous instruction 78 from the first response generation phase, which the tokenizer converts into a previous context / instruction text embedding 107. Alternatively, to avoid recalculating the embeddings of these data items, as shown by dashed lines, the concatenated prompt word input vector 100 can be directly merged into the concatenated prompt word input vector 106 used for the response evaluation phase.
[0027] Furthermore, the evaluation and revision engine 76 of the prompt generation module 74 is configured to generate a text instruction 108 including an evaluation instruction 112 to evaluate the response 32. It is understood that the text instruction 108 can be input by the user via the prompt interface 24. The response 32 and the evaluation instruction 112 are processed by the tokenizer 94 to produce corresponding response text embeddings 114 and evaluation instruction text embeddings 116, which are then concatenated with the previous prompt input vector 110 to form a concatenated prompt input vector 106 for the response evaluation phase. The concatenated prompt input vector 106 for the response evaluation phase is fed to the LLM 26 to generate a response 118 including an evaluation report 34, which may include content such as those discussed above.
[0028] Now go to Figure 4As shown by the solid line at (A1), the previous context 102 and the previous instruction 108 can be passed as text input (or appropriate multimodal input) to the context 122 of the prompt 124, which is then converted by the tokenizer 94 into a previous context / instruction text embedding 105 and incorporated into the concatenated prompt input vector 120. Alternatively, as shown by the dashed line at (A2), the concatenated prompt input vector 106 from the response evaluation phase can be passed as a previous prompt input vector 106 to be incorporated into the concatenated prompt input vector 120 for the prompt revision phase. Furthermore, as shown at (B), the response 118 with the evaluation report 34 is passed to be incorporated into the context 122 of the prompt 124 for the prompt revision phase. The evaluation report 34 is tokenized by the tokenizer 94 to produce an evaluation report text embedding 126. Furthermore, the evaluation and revision engine 76 is configured to generate a prompt revision instruction 68 (in the form of a text instruction 128) and process it through the tokenizer 94 of the embedding module 88 to produce a prompt revision instruction embedding 130. It is understood that the text instruction 128 can be input by the user via the prompt interface 24. The evaluation report text embedding 126 and the prompt revision instruction embedding 130 are concatenated with the previous prompt input vector 106 to form a concatenated prompt input vector 120 for the prompt revision stage. The concatenated prompt input vector 120 for the prompt revision stage is fed to the LLM 26 to generate a response 132 including the revised prompt 69.
[0029] Now go to Figure 5As shown by the solid line at (C1), the previous context 102 and the previous instruction 108 can be passed as text input (or appropriate multimodal input) to the revised prompt 69 generated by the prompt generation module 74, and then segmented by the tokenizer 94 of the LLM 26 to be included in the concatenated prompt input vector 134 as the previous context instruction text embedding 105. Alternatively, to save computational resources, as shown by the dashed line at (C2), the concatenated prompt input vector 120 from the prompt revision stage can be passed as the previous prompt input vector 120 to be incorporated into the concatenated prompt input vector 134 for the revised response generation stage. Furthermore, as shown at (D), the revised prompt 69 with the revised text instruction 136 is passed to be used as the prompt for the revised response generation stage. The evaluation and revision engine 76 is configured to provide a response revision instruction 70 (which may be displayed in the GUI 28 or instantiated in the background without being displayed) to the revised prompt 69 in the prompt interface 24 in text form, to instruct the generation of a revised response using the revised text instruction 136 during the revised response generation phase. It should be understood that the delivery of the revised prompt 69, including the revised text instruction 136, can be programmable, or the user can manually submit the revised prompt 69 in an attempt to improve the first response 32.
[0030] As shown by the dashed line, the user can provide the original context 82 again, such as Figure 3 As shown, processing is performed via image model 90 and word segmenter 94. Alternatively, the previous cue input vector 120 can be directly incorporated into the concatenated cue input vector 134 to provide the previous context 122 and the previous instruction 128. The revised text instruction 136 is processed by word segmenter 94 of embedding module 88 to produce a revised text instruction embedding 138. The revised text instruction embedding 138, along with the previous cue input vector 120 (optionally along with context image embedding 92 and context text embedding 96), is incorporated into the concatenated input vector 134 for the revised response generation stage. The concatenated input vector 134 for the revised response generation stage is fed into LLM 26 to generate the revised response 36. As described above, Figures 3 to 5 The evaluation and revision process shown may be iterated once or multiple times, and the revised response of the final iteration (Nth iteration 54) is referred to herein as the final response 56 of the final iteration.
[0031] Now go to Figure 6 , showed Figures 1A-1CA first example of a GUI 28 for a computing system 10, 110, or 210. In this example, a prompt word evolution settings interface 140 is provided. It should be understood that at least one processor 16 can also be configured to display a prompt word revision element and, in response to user input selecting a prompt word revision element, output a revised prompt word 69 to the user. In the prompt word evolution settings interface 140, a selector 142 is presented, through which the user can provide user input indicating whether a prompt word should be optimized (used as a prompt word revision element), and, if necessary, can use input field 144 to indicate the number of user-specified iterations for evaluation and revision. Furthermore, a selector 146 is presented, through which the user can specify whether evaluation and revision should occur in the background, preventing intermediate revised prompt words and responses from being displayed, or whether intermediate prompt words and responses should be displayed. It should be understood that this setting can be configured by the user or, on the server side, programmatically. In the example shown, the user selects 4 iterations to optimize the prompt word and does not display intermediate results. As shown in the figure, in the prompt word interface 24 of GUI 28, the user inputs prompt word 30, which includes article 148 with article body 150 as context 82, and instruction 152 "summarize the above article for fifth-grade elementary school students". Therefore, LLM program 22 performs 4 iterations of evaluation and revision in the background and outputs final response 56 including final response text 154.
[0032] exist Figure 7 The example provided is a second instance of GUI 28. In this example, the prompt evolution settings interface 140 is shown as including a selector 142, through which the user has indicated that the prompt should be optimized for one iteration. A second selector 146 is presented, through which the user has indicated that intermediate prompts and responses should be displayed, and a third selector 156 is presented, through which the user has indicated that user-specified evaluation criteria should be used. Figure 7 In the prompt word interface 24 of the GUI 28, the prompt word 30 entered by the user is multimodal, including an article 148 with article text 150 and article image 158 as context 82. This input will be used as the base input 160 (see...). Figure 3 The user also entered the same text instruction 152, which will be used by the LLM 26 as control input 80 (see...). Figure 3 It should be understood that many LLMs are configured to utilize the base input 160 and control input 80 during training, for example, by using different attention mechanisms and / or different loss terms for each input, to fine-tune the model to generate responses that respond to textual instructions (control input 80) and also take into account information in the context (base input 160) in style or manner.
[0033] Figure 6 and Figure 7 The dashed lines in the processing flow indicate user gating, where user input is requested before prompt generation proceeds to the next stage. In response to the input prompt 30, LLM 26 is configured to display a response 32 with response text 162 and a gating control that asks the user, "Do you want to evaluate the response and revise your prompt?" or similar wording, based on the prompt evolution settings. A "Yes" selector 164 and a "No" selector 166 are displayed, allowing the user to stop or continue revisions by inputting a command. When the "Yes" selector 164 is selected, the evaluation and revision engine 76 of the prompt generation module 74 displays an evaluation criteria text input pane 168, where the user can enter evaluation criteria 64. That is, in one implementation, evaluation criteria 64 can be received from the user. In another implementation, evaluation criteria 64 can be predetermined. For example, a set of evaluation criteria including conciseness, audience suitability, sufficient detail, providing citations, readability, and style of request can be used. In the example shown, suggested evaluation criteria 170 are displayed to the user. By clicking on one of the suggested evaluation criteria 170, the user can select the suggested evaluation criterion 170 to be used. The user can press the Continue button 172, causing the prompt word generation module 74 to pass the evaluation criterion 64 from the response evaluation instruction 112 to the LLM 26, as described above in the response evaluation phase. Each of the "Yes" and "Continue" selectors discussed in this example GUI 28 can be used as a prompt word revision element discussed above. Next, an evaluation report 34, including the evaluation report text 174, is displayed to the user, along with a gating control asking the user whether to generate a revised response. For example, the evaluation report 34 may include a numerical score calculated for each evaluation criterion 64 on a scale of 1-10, and a natural language (text) description of the reason for the score for each evaluation criterion 64. In some cases, at least one processor 16 may also be configured to request and receive information from the user to further specify the prompt word 30. This information can be requested based on the evaluation report 34, and / or can be used to generate the evaluation criterion 64 to be used in future response evaluation phases. For example, if evaluation report 34 includes a low rating for evaluation criterion 64 regarding the "intended audience acceptability" of response 32, the system can request further information from the user about the intended audience of response 32. Additionally, if the user specifies that the intended audience is a university mathematics professor or something similar, the evaluation criteria can be revised to include "acceptability to a university mathematics professor audience," etc. As shown in the example of evaluation criterion text input pane 168, the user can freely edit the input or choose from preset answers. Upon receiving a "yes" selection from the "yes" selector 176, the prompt generation module is configured to pass the evaluation report 34 and the prompt revision instruction 68 to the LLM 26, as described above in the prompt revision phase. Thus, as shown in the figure, the LLM 26 outputs the revised prompt 69. In this example, the user can freely edit the revised prompt 69 as needed, and upon satisfaction, the user can process the "continue" button 178, causing the response revision instruction 70 to indicate that the revised prompt 178 is fed back into the LLM 26 again, as described above in the revised response generation phase. Thus, in the final iteration of an iteration specified by the user, the final response 56 generated by the LLM 26 in response to the revised prompt 69 is displayed, and the final response includes the final response text 154.
[0034] To illustrate how evaluation and revision simultaneously lead to improved prompts and improved final responses, the following will be elaborated using the online article 148 about pandas mentioned above as an example. The first response 32 to the initial prompt 30 "Summarize the above article for fifth-grade primary school students" can be "The article is about a bear called the giant panda. These bears live in central China and mainly eat bamboo. People are worried about pandas because there are not many pandas left in the wild. But the good news is that the number of pandas in the wild seems to be increasing! People in China and around the world are working hard to prevent pandas from going extinct". Then the LLM 26 is instructed to perform a self-evaluation of the previous prompt 30 (labeled PREVIOUS_PROMPT) and the response 32 (labeled PREVIOUS_RESPONSE) using the following evaluation instruction 112: Complete the following survey regarding PREVIOUS_PROMPT: <Audience description = "Based on PREVIOUS_PROMPT, who is the intended audience for this material?">Answer here< / Audience> Rate the PREVIOUS_RESPONSE on a scale of 1-10 for each of the following criteria: <Appropriateness description = "Appropriateness for the intended audience" rating = "<1-10>" / > <Readability description = "Readability and understandability of PREVIOUS_RESPONSE" rating = "<1-10>" / > <Simplicity description = "How well the response is summarized. Is it concise?" rating = "<1-10>" / > <Inclusivity description = "How inclusive is the language and terminology in PREVIOUS_RESPONSE?" rating = "<1-10>" / > <Fun Description: How appealing is PREVIUS_RESPONSE to the target audience? Rating: "<1-10>" / >
[0035] As provided in this example, assessment instruction 112 may include a mixture of plain language and markers. In this example, the user specifies five user assessment criteria 64. In response, LLM 26 may output the following first assessment report 34, which may include one or both of the following: a score and a written description of how the first response 32 meets assessment criteria 64. Audience: The target audience appears to be fifth-grade elementary school students, as specified in the cue words. Appropriateness: 8- The response breaks the article down into terms that elementary school students can understand, but it could be simplified further. Readability: 8 - The response is generally easy to understand, but some words (such as "extinction") may be difficult for some fifth graders. Conciseness: 9-Response summarizes the key points of the article well without getting bogged down in details. Inclusivity: 10- The language used in the response is acceptable to all readers. Fun factor: 7- The response summarizes the article well, but it could use more vivid language to capture the attention of fifth-grade students.
[0036] Then, PREVIOUS_PROMPT and evaluation report 34 are fed back into LLM 26 with further instructions: "Based on these scores, create improved prompts to produce better results." The improved prompt 69 output by LLM 26 could be: "Summarize the above article for fifth-grade students, using simple vocabulary and vivid language to make it appealing to young readers." This concludes the first iteration of optimization 50.
[0037] For the second iteration 52, the improved prompt 69 and original article 148 were input into LLM 26. This time, the second-generation response 72 could be: "The above article is about giant pandas, which are bears that live in China. Giant pandas are known for their black and white fur and their love of eating bamboo. Unfortunately, due to habitat destruction by humans, giant pandas are now a vulnerable species. But the good news is that people are working hard to help pandas, and their numbers are starting to grow again. Giant pandas are so important to China that they are often used as a national symbol." The same self-evaluation instruction 112 as in the first iteration could be submitted, and the result might be as follows: Appropriateness: 8 - The response is appropriate for the intended audience, using simple vocabulary and providing a basic summary of the article. However, some more scientific terms (such as "vulnerable species") may require further explanation for fifth-grade students. Readability: 9 - Very readable, with short sentences and simple vocabulary. Conciseness: 8-Response summarizes the article well in a short paragraph, although it could be more concise if some less important details were omitted. Inclusivity: 10 - The response did not use any exclusive language or terminology that could affect understanding. Fun factor: 8 - The response is engaging and includes some interesting facts about giant pandas, but it would be more vivid if it included more sensory details or a stronger sense of urgency about the pandas' plight.
[0038] After being instructed to further revise the revised prompt 69, LLM 26 can output "Summarize the above article for fifth graders, using simple vocabulary and vivid language to make it appealing to young readers. Add sensory details to make the summary more interesting and explain any scientific terms in a way that fifth graders can understand. Keep the summary short and concise, focusing only on the most important information." This concludes the second iteration.
[0039] Using the latest prompt word 69, LLM 26 can output the third-generation response: "Giant pandas are large, black-and-white bears that live only in China. They mainly eat bamboo, but sometimes they eat other things. Because people are constantly encroaching on their habitat, there are not many giant pandas left in the wild. Scientists are trying to help them by keeping some giant pandas in zoos. They are also trying to build more habitats for wild pandas. Giant pandas are a symbol of China and were the mascot of the 2008 Olympic Games." The evaluation report 34 for the third iteration can be as follows: Appropriateness: 8- The vocabulary used is primarily intended for fifth-grade students, but scientific terms can be explained in a simpler way. Readability: 8- Response is easy to read and understand, but it could be more engaging with more vivid language and sensory details. Simplicity: The 9-response is well summarized and concise, focusing on the most important information. Inclusivity: 10 - The language and terminology used in the response are inclusive. Fun factor: 7 - The response is fun, but it could be made more engaging with sensory details and vivid language.
[0040] After being instructed to further revise the revised prompt 69, LLM 26 can output: "Summarize the above article for fifth graders, using simple vocabulary and vivid language to make it appealing to young readers. Use sensory details to make the summary relatable and explain any scientific terms in a way that fifth graders can understand. Keep the summary short and concise, focusing only on the most important information. Ensure that the language and terminology used are inclusive and friendly to all students." This concludes the third iteration.
[0041] Using the latest and final revised prompt 69, LLM 26 can output the final response 56: “The giant panda is a bear that lives in China. It has black and white fur and a round body. It mainly eats bamboo, but sometimes it also eats other plants and even meat. Due to the encroachment of their habitat by people for farming and construction, there are not many giant pandas left in the wild. However, people are trying to help protect pandas, and the number of wild pandas is increasing. The giant panda is also a symbol of China and is the mascot of the Olympic Games.” If needed, the final response 56 can be evaluated to generate scores of 8, 9, 10, 10, and 7 for the same category as before. The responses of each iteration can be compared by total score or average score, or another suitable comparison method can be used. In this example, the response scores are 42, 43, 42, and 44, respectively, showing that the final prompt 69 and response 56 have been improved under the provided evaluation criteria 64. Before accepting the final result, the revised cue word 69 can be used in large projects to generate higher-quality results more efficiently by iterating through multiple iterations using resources. For example, if a website hosting online articles about giant pandas also hosts a large repository of other articles and wants to generate summaries for children for each article, it is better to ensure that the globally used cue words are thoroughly tested and generate acceptable responses before having LLM 26 generate all summaries at once, rather than relying on the expertise of the user who drafted the initial cue word 30 to do well on the first attempt.
[0042] Figure 8 A flowchart of a method 600 for revising LLM input prompts is shown. Method 600 can be performed by... Figure 1A The computing system 10, 110, or 210 shown in -C is implemented.
[0043] At 602, method 600 may include presenting a cue word interface for the trained LLM. This interface may be, for example, an audio interface allowing the user to provide audio input, or a graphical user interface (GUI) allowing the user to input text or graphical input. At 604, method 600 may include receiving cue words from a user via the cue word interface, the cue words including instructions for the LLM to generate output. The cue words may be initial cue words from the user to produce expected output, such as text output, audio output, or graphical output. That is, the LLM may be multimodal. At 606, method 600 may include providing the LLM with a first input including the cue words.
[0044] At 608, method 600 may include generating a first response to the prompt word via the LLM in response to the first input. The first response may be user-acceptable. However, in some cases, the user may not have written the prompt word in a way that would allow them to obtain the expected output from the LLM. The user may be unfamiliar with the operation of the LLM, make incorrect assumptions, or omit helpful information. Therefore, to improve the response and / or the prompt word, in some implementations, at 610, method 600 may include receiving evaluation criteria from the user. Alternatively, at 612, method 600 may include requesting further information from the user to specify the prompt word. That is, if the user can indicate the response they want, the user may be more willing to submit the evaluation criteria directly, but the computational system may generate appropriate evaluation criteria on behalf of the user after requesting and receiving contextual information (such as who the intended audience for the output is). Even if the user is unfamiliar with operating the LLM, higher quality revisions may be obtained by progressively asking the user for further information. Accordingly, the evaluation criteria may be generated by the LLM at least based on the intended audience for the output, which may be provided by the user or inferred by the LLM. As described below, by using this information to generate relevant evaluation criteria, including appropriateness for the target audience, and then using these criteria to evaluate previous responses, LLM can be better able to determine whether previous responses were appropriate for the intended audience.
[0045] At 614, method 600 may include performing the evaluation and revision of the prompt word at least in part by: at 616, evaluating a first response via an LLM according to evaluation criteria to generate a first evaluation report for the first response; at 618, providing the LLM with a second input including the first prompt word, the first response, the first evaluation report, and prompt word revision instructions to revise the prompt word in light of the first evaluation report; and at 620, generating the revised prompt word via an LLM in response to the second input. In some implementations, the evaluation report may include one or both of the following: a score and a written description of how the first response meets the evaluation criteria. The score may allow for mathematical analysis and summarization of the acceptability of the first response, while the written description may allow for providing the LLM with a clear path to revise the prompt word in light of the evaluation. It should be understood that the evaluation may be a self-evaluation performed by a single LLM, or one LLM may be responsible for generating a response based on the prompt word, while another LLM is responsible for evaluating the response and revising the prompt word. In this case, the LLM responsible for the evaluation may be a large LLM with more parameters, which tends to require more resources to run and may have higher demands and / or costs. Using a more expensive LLM to revise the prompts that will be run on the response generation LLM (which may be an older, legacy model) allows for the generation of responses with fewer resources while achieving a higher standard than when the older LLM model was generated alone.
[0046] In some implementations, prompt evaluation and revision are performed iteratively over multiple iterations, whereby responses to previous prompts are evaluated and revised to produce improved responses. In this way, the prompts themselves are improved, resulting in better responses for both casual users and experts. Furthermore, the number of iterations can be customized by the user. This allows users to freely decide whether to invest more or less resources to improve the prompts based on their needs and available resources.
[0047] At 622, method 600 may include providing the LLM with a final input including the revised cue words. That is, if the cue words are iteratively revised, the final input may be the last input after all iterations have been run. At 624, method 600 may include generating a final response to the revised cue words via the LLM in response to the final input. At 626, method 600 may include outputting the final response to the user. In this way, the user can receive a final response that meets the evaluation criteria, where the first response may have failed to meet the requirements or received a lower score, and is therefore more likely to be considered acceptable by the user. In some implementations, the final response generated after multiple iterations may be output to the user without outputting any intermediate responses. Accordingly, the system may be able to present the user with the best impression of being highly capable and immediately and accurately generating the content the user wants.
[0048] The systems and methods described above offer potential technical advantages: reducing computational resources while increasing usability and effectiveness for users during the generation of LLM responses. For example, these systems and methods can reduce the number of times users repeatedly provide prompts to the LLM during trial and error to extract useful information by optimizing user prompts more quickly and efficiently. This is applicable to developers working on software utilizing LLM. These developers can configure the system by providing a test dataset, which can serve as contextual input for evaluating responses from the LLM when using the software. In this way, developers can provide evaluation criteria for assessing prompt responses, helping the system generate responses to user prompts more effectively. Another group of users for whom the systems and methods offer technical advantages is end-users. The systems and methods described above can be configured to programmatically and dynamically revise prompts input by users to evaluate LLM responses based on evaluation criteria that meet user needs, and to evolve those prompts based on the evaluation criteria to improve responses and better meet user expectations. This helps save computational resources because it reduces the trial-and-error cycle users spend searching for prompts that yield useful responses from the LLM. In some implementations, it also enables resource-efficient, computationally inexpensive LLMs to achieve or exceed the response levels of larger, more expensive models when responding to user prompts, thereby saving computational resources.
[0049] In some embodiments, the methods and processes described herein can be attached to a computing system of one or more computing devices. In particular, such methods and processes can be implemented as computer applications or services, application programming interfaces (APIs), libraries, and / or other computer program products.
[0050] Figure 9 A non-limiting embodiment of a computing system 700 is schematically shown, which can perform one or more methods or processes described above. The computing system 700 is shown in a simplified form. The computing system 700 can embody the above and... Figures 1A to 1C The computing systems 10, 110, and 210 are shown. The computing system 700 may take the form of one or more personal computers, server computers, tablet computers, home entertainment computers, network computing devices, gaming devices, mobile computing devices, mobile communication devices (e.g., smartphones), and / or other computing devices, as well as wearable computing devices such as smartwatches and head-mounted augmented reality devices.
[0051] The computing system 700 includes a logic processor 702, volatile memory 704, and non-volatile storage device 706. The computing system 700 may optionally include a display subsystem 708, an input subsystem 710, a communication subsystem 712, and / or... Figure 9Other components not shown.
[0052] The logic processor 702 includes one or more physical devices configured to execute instructions. For example, the logic processor may be configured to execute instructions that are part of one or more applications, programs, routines, libraries, objects, components, data structures, or other logical constructs. Such instructions may be implemented to perform tasks, implement data types, transform the state of one or more components, achieve technical effects, or otherwise achieve desired results.
[0053] A logic processor may include one or more physical processors (hardware) configured to execute software instructions. Additionally or alternatively, a logic processor may include one or more hardware logic circuits or firmware devices configured to execute hardware-implemented logic or firmware instructions. The processor of logic processor 702 may be single-core or multi-core, and the instructions executed thereon may be configured for sequential, parallel, and / or distributed processing. The various components of the logic processor may optionally be distributed across two or more separate devices, which may be located remotely and / or configured for coordinated processing. Aspects of the logic processor may be virtualized and executed by remotely accessible networked computing devices configured in a cloud computing configuration. It should be understood that, in this case, these virtualized aspects run on different physical logic processors on various different machines.
[0054] The non-volatile storage device 706 includes one or more physical devices configured to hold instructions executable by a logic processor to implement the methods and processes described herein. When such methods and processes are implemented, the state of the non-volatile storage device 706 can be transformed, for example, to retain different data.
[0055] Non-volatile storage device 706 may include removable and / or built-in physical devices. Non-volatile storage device 706 may include optical memory (e.g., CD, DVD, HD-DVD, etc.), semiconductor memory (e.g., ROM, EPROM, EEPROM, flash memory, etc.), and / or magnetic memory (e.g., hard disk drive, floppy disk drive, magnetic tape drive, MRAM, etc.) or other high-capacity storage device technologies. Non-volatile storage device 706 may include non-volatile, dynamic, static, read / write, read-only, sequential access, location-addressable, file-addressable, and / or content-addressable devices. It should be understood that non-volatile storage device 706 is configured to retain instructions even when non-volatile storage device 706 is powered off.
[0056] Volatile memory 704 may include a physical device containing random access memory. Volatile memory 704 is typically used by logic processor 702 to temporarily store information during the processing of software instructions. It should be understood that when volatile memory 704 is powered off, volatile memory 704 typically does not continue storing instructions.
[0057] The logic processor 702, volatile memory 704, and non-volatile storage device 706 can be integrated together into one or more hardware logic components. For example, such hardware logic components may include field-programmable gate arrays (FPGAs), programmable and application-specific integrated circuits (PASICs / ASICs), programmable and application-specific standard products (PSS / ASSPs), system-on-a-chip (SOCs), and complex programmable logic devices (CPLDs).
[0058] The terms "module," "program," and "engine" can be used to describe aspects of a computing system 700 typically implemented in software by a processor to perform specific functions using portions of volatile memory, involving transformation processing specifically configured for the processor to perform those functions. Thus, a module, program, or engine can be instantiated using portions of volatile memory 704 via a logical processor 702 executing instructions held by non-volatile memory 706. It should be understood that different modules, programs, and / or engines can be instantiated from the same applications, services, code blocks, objects, libraries, routines, APIs, functions, etc. Similarly, the same module, program, and / or engine can be instantiated from different applications, services, code blocks, objects, routines, APIs, functions, etc. The terms "module," "program," and "engine" can encompass individuals or groups of executable files, data files, libraries, drives, scripts, database records, etc.
[0059] When included, the display subsystem 708 can be used to present a visual representation of the data held by the non-volatile storage device 706. This visual representation can take the form of a graphical user interface (GUI). Since the methods and processes described herein change the data held by the non-volatile storage device, and thus transform the state of the non-volatile storage device, the state of the display subsystem 708 can also be transformed to visually represent the changes in the underlying data. The display subsystem 708 may include one or more display devices utilizing virtually any type of technology. Such display devices may be combined with the logic processor 702, the volatile memory 704, and / or the non-volatile storage device 706 within a shared housing, or such display devices may be peripheral display devices.
[0060] When included, the input subsystem 710 may include or be connected to one or more user input devices such as a keyboard, mouse, touchscreen, or game controller. In some embodiments, the input subsystem may include or be connected to selected Natural User Input (NUI) components. Such components may be integrated or peripheral, and the translation and / or processing of input actions may be handled in-vehicle or non-vehicle. Example NUI components may include microphones for voice and / or speech recognition; infrared, color, stereo, and / or depth cameras for machine vision and / or gesture recognition; head trackers, eye trackers, accelerometers, and / or gyroscopes for motion detection and / or intent recognition; and / or any other suitable sensors.
[0061] When included, the communication subsystem 712 can be configured to communicatively couple the various computing devices described herein to each other and to other devices. The communication subsystem 712 may include wired and / or wireless communication devices compatible with one or more different communication protocols. As a non-limiting example, the communication subsystem can be configured to communicate via a wireless telephone network or a wired or wireless local area network or wide area network (such as HDMI over a Wi-Fi connection). In some embodiments, the communication subsystem may allow the computing system 700 to send messages to and / or receive messages from other devices via a network such as the Internet.
[0062] The following paragraphs provide additional support for the claims of this application. One aspect provides a computational system for revising input prompts for a Large Language Model (LLM). The computational system includes at least one processor configured to present a prompt interface for a trained LLM, receive prompts from a user via the prompt interface, the prompts including instructions for generating output for the LLM, provide a first input including the prompts to the LLM, and generate a first response to the prompts via the LLM in response to the first input. The at least one processor is configured to perform the evaluation and revision of the prompts at least in part by: evaluating the first response via the LLM according to evaluation criteria to generate an evaluation report for the first response; providing the LLM with a second input including the first prompts, the first response, the evaluation report, and prompt revision instructions to revise the prompts in view of the first evaluation report; and generating the revised prompts via the LLM in response to the second input. The at least one processor is configured to provide the LLM with a final input including the revised prompts, generate a final response to the revised prompts via the LLM in response to the final input, and output the final response to the user. In this respect, additionally or alternatively, the evaluation and revision of the cue words are performed iteratively for multiple iterations. In this respect, additionally or alternatively, the number of iterations can be user-customizable. In this respect, additionally or alternatively, the at least one processor can also be configured to output the final response generated after multiple iterations to the user, without outputting any intermediate responses to the user. In this respect, additionally or alternatively, the LLM can be multimodal. In this respect, additionally or alternatively, the evaluation criteria can be received from the user. In this respect, additionally or alternatively, the at least one processor can be further configured to request further information from the user to specify the cue words. In this respect, additionally or alternatively, the evaluation criteria can be generated by the LLM at least based on the intended audience of the output, the intended audience being provided by the user or inferred by the LLM. In this respect, additionally or alternatively, the evaluation report can include one or both of the following: a score and a written description of how the first response meets the evaluation criteria. In this respect, additionally or alternatively, the at least one processor may also be configured to display a prompt revision element and, in response to user input that selects the prompt revision element, output a revised prompt to the user.
[0063] On the other hand, a method for revising input prompts for a Large Language Model (LLM) is provided. The method includes: presenting a prompt interface for a trained LLM; receiving prompts from a user via the prompt interface, the prompts including instructions for the LLM to generate output; providing the LLM with a first input including the prompts; generating a first response to the prompts via the LLM in response to the first input; and performing the evaluation and revision of the prompts at least in part by: evaluating the first response via the LLM according to evaluation criteria to generate an evaluation report for the first response; providing the LLM with a second input including the first prompts, the first response, the evaluation report, and prompt revision instructions to revise the prompts based on the evaluation report; and generating the revised prompts via the LLM in response to the second input. The method further includes: providing the LLM with a final input including the revised prompts; generating a final response to the revised prompts via the LLM in response to the final input; and outputting the final response to the user. In this aspect, additionally or alternatively, the evaluation and revision of the prompts are performed iteratively for multiple iterations. In this respect, additionally or alternatively, the number of iterations may be user-defined. In this respect, additionally or alternatively, the final response generated after multiple iterations may be output to the user without any intermediate responses. In this respect, additionally or alternatively, the LLM is multimodal. In this respect, additionally or alternatively, the method may further include receiving evaluation criteria from the user. In this respect, additionally or alternatively, the method may further include requesting the user for further information specifying cue words. In this respect, additionally or alternatively, the evaluation criteria are generated by the LLM based at least on the intended audience of the output, which is provided by the user or inferred by the LLM. In this respect, additionally or alternatively, the evaluation report may include one or both of the following: a score and a written description of how the first response meets the evaluation criteria.
[0064] On the other hand, a computational system for revising input prompts for a Large Language Model (LLM) is provided. The computational system includes at least one processor configured to: render a prompt interface for a first trained LLM; receive prompts from a user via the prompt interface, the prompts including instructions for generating output for the first LLM; provide a first input including the prompts to the first LLM; generate a first response to the prompts via the first LLM in response to the first input; perform prompt evaluation and revision at least in part by: evaluating the first response via a second LLM according to evaluation criteria, generating an evaluation report for the first response, the second LLM having a larger parameter scale than the first LLM; providing a second input to the LLM including the first prompts, the first response, the evaluation report, and prompt revision instructions to revise the prompts based on the evaluation report; and generating revised prompts via the second LLM in response to the second input, providing a final input including the revised prompts to the first LLM, generating a final response to the revised prompts via the first LLM in response to the final input, and outputting the final response to the user.
[0065] On the other hand, a computational system for revising input prompts for a Large Language Model (LLM) is provided. The computational system includes at least one processor configured to: execute a prompt interface application programming interface (API) for a trained LLM; receive prompts via the prompt interface API, the prompts including instructions for the LLM to generate output; provide the LLM with a first input including the prompts; generate a first response to the prompts via the LLM in response to the first input; perform prompt evaluation and revision at least in part by: evaluating the first response via the LLM according to evaluation criteria and generating an evaluation report for the first response; providing the LLM with a second input including the first prompts, the first response, the evaluation report, and prompt revision instructions to revise the prompts based on the first evaluation report; generating revised prompts via the LLM in response to the second input; providing the LLM with a final input including the revised prompts; generating a final response to the revised prompts via the LLM in response to the final input; and outputting the final response via the prompt interface API.
[0066] It should be understood that the configurations and / or methods described herein are exemplary in nature, and these specific embodiments or examples should not be considered limiting, as many variations are possible. The specific routines or methods described herein may represent one or more of any number of processing strategies. Therefore, the various actions shown and / or described may be performed in the shown and / or described order, in another order, in parallel, or omitted. Similarly, the order of the above processes may also be changed.
[0067] The subject matter of this disclosure includes all novel and non-obvious combinations and sub-combinations of the various processes, systems and configurations of this disclosure, as well as any and all equivalents thereof.
Claims
1. A computational system (10, 110, 210) for revising input prompts in a large language model (LLM), the computational system (10, 110, 210) comprising: At least one processor (16) is configured to: The prompt word (30) interface (24) for the trained LLM (26) is presented; The prompt (30) is received from the user via the prompt (30) interface (24), the prompt (30) including instructions for the LLM (26) to generate output; Provide the LLM (26) with a first input (60) including the prompt word (30); In response to the first input (60), a first response (32) to the prompt word (30) is generated via the LLM (26); The evaluation and revision of the prompt word (30) shall be performed at least in part by the following means: The first response (32) is evaluated via the LLM (26) according to the evaluation criteria (64), and an evaluation report (34) for the first response (32) is generated; The LLM (26) is provided with a second input (66) including the first prompt word (30), the first response (32), the evaluation report (34), and the prompt word revision instruction (68) to revise the prompt word (30) in view of the first evaluation report (34); as well as In response to the second input (66), a revised prompt word (69) is generated via the LLM (26); Provide the LLM (26) with the final input (58) including the revised prompt (69); In response to the final input (58), a final response (56) to the revised prompt (69) is generated via the LLM (26); as well as The final response (56) is output to the user.
2. The computing system of claim 1, wherein the evaluation and revision of the prompt words are performed iteratively for multiple iterations.
3. The computing system of claim 2, wherein the number of iterations is a user-customizable number.
4. The computing system of claim 2, wherein the at least one processor is further configured to output the final response generated after the multiple iterations to the user, without outputting any intermediate responses to the user.
5. The computing system according to claim 1, wherein the LLM is multimodal.
6. The computing system of claim 1, wherein the evaluation criteria are received from the user.
7. The computing system of claim 1, wherein the at least one processor is further configured to request the user to specify further information regarding the prompt word.
8. The computing system of claim 1, wherein the evaluation criteria are generated by the LLM based at least on the expected audience of the output, the expected audience being provided by the user or inferred by the LLM.
9. The computing system of claim 1, wherein the evaluation report includes one or both of the following: a score and a written description of how the first response meets the evaluation criteria.
10. The computing system of claim 1, wherein the at least one processor is further configured to: Make the prompt word revision element displayed; and In response to user input that selects the prompt word revision element, the revised prompt word is output to the user.
11. A method (600) for revising input prompt words in a large language model (LLM), the method (600) comprising: (602) The prompt word interface for the trained LLM is presented; The prompt word is received from the user (604) via the prompt word interface, the prompt word including instructions for the LLM generation output; Provide the LLM with (606) a first input including the prompt word; In response to the first input, a first response to the prompt word is generated (608) via the LLM; The evaluation and revision of the prompt words described in (614) shall be performed at least in part by the following means: The first response is evaluated (616) via the LLM according to the evaluation criteria, and an evaluation report for the first response is generated. The LLM is provided with (618) a second input including the first prompt word, the first response, the evaluation report, and a prompt word revision instruction, to revise the prompt word in view of the evaluation report; as well as In response to the second input, a revised prompt word is generated (620) via the LLM; Provide the LLM with (622) final input including the revised prompt words; In response to the final input, a final response to the revised prompt word is generated (624) via the LLM; as well as Output the final response (626) to the user.
12. The method of claim 11, wherein the evaluation and revision of the prompt words are performed iteratively for multiple iterations.
13. The method of claim 12, wherein the number of iterations is a user-customizable number.
14. The method of claim 12, wherein the final response generated after the multiple iterations is output to the user, without outputting any intermediate responses to the user.
15. The method of claim 11, wherein the LLM is multimodal.
16. The method of claim 11, further comprising receiving the evaluation criteria from the user.
17. The method of claim 11, further comprising requesting the user to specify further information regarding the prompt word.
18. The method of claim 11, wherein the evaluation criteria are generated by the LLM based at least on the expected audience of the output, the expected audience being provided by the user or inferred by the LLM.
19. The method of claim 11, wherein the evaluation report comprises one or both of the following: a score and a written description of how the first response meets the evaluation criteria.
20. A computational system (10, 110, 210) for revising input prompts in a large language model (LLM), the computational system (10, 110, 210) comprising: At least one processor (16) is configured to: Execute the prompt word interface application programming interface (API) (44) for the trained LLM (26); The prompt word (30) is received via the prompt word interface API (44), the prompt word (30) including instructions for the LLM (26) to generate output; Provide the LLM (26) with a first input (60) including the prompt word (30); In response to the first input (60), a first response (32) to the prompt word (30) is generated via the LLM (26); The evaluation and revision of the prompt words are performed at least in part by the following means (30): The first response (32) is evaluated via the LLM (26) according to the evaluation criteria (64), and an evaluation report (34) for the first response (32) is generated; Provide the LLM (26) with a second input (66) including the first prompt word (30), the first response (32), the evaluation report (34) and the prompt word revision instruction (68) to revise the prompt word (30) based on the first evaluation report (34); as well as In response to the second input (66), a revised prompt word (69) is generated via the LLM (26); Provide the LLM (26) with the final input (58) including the revised prompt (69); In response to the final input (58), a final response (56) to the revised prompt (69) is generated via the LLM (26); as well as The final response (56) is output via the prompt word interface API (44).