Provide suitability hints to evaluate and improve the output of generative models without fine-tuning.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2026-08-14
Smart Images

Figure CN122580667A_ABST
Abstract
Description
Background Technology
[0001] Machine learning models, especially generative models, may require fine-tuning before they are suitable for use with specific tasks in a business environment. However, such fine-tuning typically requires computational resources and high-quality data that provides good examples of the results to be trained on. Attached Figure Description
[0002] Details of one or more aspects of the subject matter described in this disclosure are set forth in the accompanying drawings and the description below. However, the drawings illustrate only some typical aspects of this disclosure and should not be considered as limiting the scope of this disclosure. Other features, aspects, and advantages will become apparent from the description, drawings, and claims.
[0003] Figure 1 Example book applications based on some aspects of this technology are illustrated.
[0004] Figure 2 An example system layout for generating applications based on some aspects of this technology is illustrated.
[0005] Figure 3 An example process is illustrated, based on some aspects of this technique, for interacting with a generative model to obtain a response with improved quality without fine-tuning a large language model for use cases.
[0006] Figure 4 A table illustrating the types or categories of questions that facilitate conversational reading, outlining some aspects of this technology.
[0007] Figure 5 Examples of problems based on this technique are illustrated, along with tables explaining why appropriate determinations might be made regarding these examples and the reasoning behind those conclusions.
[0008] Figure 6A Examples of book applications based on some aspects of this technology at time t=0 are shown.
[0009] Figure 6B Examples of book applications based on some aspects of this technique are illustrated at time t=1, which occurs after the appropriate response has been provided by the generative model.
[0010] Figure 7 Examples of systems used to implement some aspects of this technology are shown. Detailed Implementation
[0011] Various embodiments of this disclosure are discussed in detail below. While specific implementations are discussed, it should be understood that this is for illustrative purposes only. Those skilled in the art will recognize that other components and configurations can be used without departing from the spirit and scope of this disclosure.
[0012] Additional features and advantages of this disclosure will be set forth in the following description and will be apparent in part from that description, or may be learned by practicing the principles disclosed herein. The features and advantages of this disclosure may be realized and obtained by means of the instruments and combinations particularly pointed out in the appended claims. These and other features of this disclosure will become more fully apparent from the following description and the appended claims, or may be learned by practicing the principles set forth herein.
[0013] Machine learning models, especially generative models, may require fine-tuning before they are suitable for use in a specific task within a business environment. However, this fine-tuning typically requires computational resources and high-quality data that provides good examples of the results to be trained on. Often, fine-tuning also requires experts to create or curate suitable data for fine-tuning and to evaluate the model's performance to provide reinforcements or corrections (e.g., reinforcement learning). Unfortunately, such resources are unavailable to many, which limits what can be achieved by individuals and small groups.
[0014] This technique provides a mechanism to obtain results with quality similar to that achievable by fine-tuning a generative model from a base model without fine-tuning. Specifically, this technique can provide suitability cues to evaluate and improve the output of a generative model without fine-tuning. Suitability cues are hints provided to the design of the generative model, prompting it to evaluate candidate responses already generated by the model. Typically, suitability cues may include indications of one or more properties of a quality candidate response. When the generative model provides a response to a suitability cue indicating that a candidate response is a quality response, the candidate response can be considered good enough to be returned to the user. Alternatively, when the generative model provides a response to a suitability cue indicating that a candidate response is not a quality response, revision cues can be provided to allow the generative model to attempt to generate a higher quality response again.
[0015] Throughout this specification, the technology will be described in the context of an example question-generating application that generates questions to facilitate conversational reading based on books read by novice readers. In conversational reading, adults guide contextual dialogue with children by asking specific types of story-related questions that address vocabulary development, phonemic awareness, recall, expressive fluency, and building connections between the story and the child's life. While the question-generating application uses carefully designed prompts that instruct the generative model to produce questions that achieve goals such as those listed, considerable variability can exist in the output from the generative model, as not all questions that appear to achieve these goals are considered high-quality responses that would generate meaningful dialogue.
[0016] As noted above, conventional practice in this field would be to fine-tune generative models to distinguish between high-quality and low-quality responses. However, this may be impractical because identifying high-quality responses would require expert educators to invest significant time ranking the responses output by the generative model, and then substantial time and computational resources would be needed to fine-tune the model to achieve an acceptable level of performance. For this use case, these requirements are considered impractical.
[0017] Conversely, question generation applications can create follow-up prompts (called suitability prompts) that require the generative model to evaluate candidate responses already generated by the generative model to provide answers regarding whether the candidate responses meet one or more quality criteria. These quality criteria can be based on quality evaluation metric used to evaluate candidate responses.
[0018] Using this approach, this technique avoids fine-tuning the generative model by alternatively requiring the generative model to evaluate the quality of its output. Throughout this specification, examples are provided that offer multiple prompts for generating candidate responses and then evaluating them; however, the initial prompt can also request the generative model to generate a problem that both achieves the stated objective and satisfies the quality criteria. While this is possible and explicitly envisioned to fall within the scope of this technique, the approach of interacting with the generative model through multiple prompts is used in most examples for two reasons. The first reason is that this interpretive approach makes the purpose of each prompt clearer. The second reason is that it has been shown that some generative models struggle to direct their attention to all the key aspects of lengthy prompts. Therefore, while the primary concept of this technique is that when a generative model is given information about the quality criteria that can be used to evaluate the quality of a response, it can be prompted to perform a specific task with a quality typically only visible through fine-tuning, this technique also envisions a method of interacting with the generative model in a way that forces it to focus on the problem generation criteria and the quality criteria.
[0019] This technique offers significant advantages, as it eliminates the need for fine-tuning the base model to work with even a specific task. This drastically reduces the human and computational burden of developing high-quality training data and decreases computation time from the training process. Interestingly, in some cases, some of these benefits may actually erode over time. More specifically, while there are clear savings in computational resources by avoiding training the machine learning model, there can be an increasing burden on the use of generative models, resulting from providing additional or longer, more complex cues to them. The net savings in computational resources from avoiding fine-tuning can vary depending on the size of the generative model (how many parameters it has) and its deployment (whether it runs on expensive cloud resources or less expensive personal computing devices).
[0020] Figure 1 Example book applications based on some aspects of this technology are illustrated. For example... Figure 1 As illustrated, the book application 102 displays story content 104, which in this case includes text and images. Through interaction with this technology, the book application 102 can receive high-quality questions 106 from a generative model to facilitate conversational reading and present those questions within the book application 102.
[0021] The question generation application can be provided with templates to create generation hints to be provided to the generative model to generate questions. In some implementations, the generation hints may correspond to a type of question that is easy to read conversationally (for more information on question types, see...). Figure 4 When a generative model provides candidate responses, the question generation application can provide one or more suitability cues created from a quality response evaluation metric. For example, suitability cues can guide the generative model to evaluate candidate responses for aspects such as wording, realism, and complexity. In some cases, suitability cues vary based on the type of question requested in the generated cues.
[0022] Figure 2 An example system arrangement for generating applications based on some aspects of the present technology is illustrated. Although this example system depicts specific system components and arrangements of such components, this description is for the purpose of facilitating the discussion of the present technology and should not be considered limiting unless specified in the appended claims. For example, some components illustrated as separate may be combined with other components, and some components may be divided into separate components.
[0023] Figure 2An example system layout for a question generation application 202 based on some aspects of this technology is illustrated. While the question generation application 202 is named in the context of a specific use case, this technology relates to any application that can provide suitability hints to a generative model for the purpose of evaluating the appropriateness of the output of the generative model for a particular purpose. Although this technology is described throughout this specification in the context of a question generation application 202 for providing questions that facilitate conversational reading, this is merely an example use case. The output of this technology is not limited to outputting questions.
[0024] like Figure 2 As illustrated, the book application 102 is invoked at time t=0, and then later at time t=1. At time t=0, the book application 102 can be configured to request questions suitable for conversational reading of the story content 104. The book application 102 can automatically request and then present questions 106 to facilitate conversational reading, or the user can interact with the UI button 204 to cause the book application 102 to request and then present questions 106 to facilitate conversational reading.
[0025] The question generation application 202 includes a prompt generation service 208 configured to combine a selected prompt template 206 with story content 104 extracted from books displayed in the book application 102. As described herein, the question generation application 202 can send prompts created by the prompt generation service 208 to a generative model 214 to receive candidate responses, such as candidate questions that can facilitate conversational reading within the context of the story content 104. In some implementations, the prompt template can be varied according to the question type. In some implementations, the question type is randomly selected. In some implementations, the question type is selected using logic based on factors such as parental preference, child age, and the context of the story.
[0026] Although Figure 2 Generative model 214 is illustrated separately from problem-generating application 202, but this is merely for illustrative purposes, and the generative model can be part of problem-generating application 202. Similarly, problem-generating application 202 can be part of book application 102, even though they are shown as separate entities.
[0027] The prompt generation service 208 is also configured to receive candidate responses from the generative model 214 and pass the candidate responses to the response suitability service 210. The response suitability service 210 is configured to prompt the generative model 214 to evaluate the suitability of candidate responses based on suitability criteria provided by the quality response evaluation metric 212. Depending on the response provided by the generative model 214, the response suitability service 210 may deem the candidate response suitable, or it may provide feedback to the prompt generation service 208 to further interact with the generative model 214 to receive a revised candidate response.
[0028] A response deemed appropriate can be returned to the book application 102 at time t=1 to be displayed as questions 106 and story content 104 to facilitate conversational reading.
[0029] In some implementations, generative model 214 can be any service capable of generating appropriate output in response to received prompts. For example, generative model 204 can be a machine learning model or other techniques with natural language processing capabilities and / or image generation capabilities. This technique is independent of the specific generative model 214 utilized, and different generative models 214 can be utilized simply by calling the application programming interfaces of different generative models 214.
[0030] In some examples, the generative model 214 can be a large language model, such as OPEN AI's CHATGPT, Google's BARD, or ANTHROPIC's CLAUDE.
[0031] Large Language Models (LLMs), exemplified by GPT (Generative Pre-trained Transformer) and BERT (Bidirectional Encoder Representation from Transformer), are built on deep neural network architectures and trained on massive datasets of internet text. The training process involves immersing the model in massive amounts of text, allowing it to internalize the complexities of language, including nuances of syntax, semantics, and contextual understanding. LLMs probabilistically generate content because they sample from probability distributions to determine the most likely next word or word sequence based on their training.
[0032] Figure 3 Example processes according to some aspects of the present invention are illustrated for interacting with a generative model to obtain responses with improved quality without fine-tuning the generative model for use cases. Although the example processes depict a specific sequence of operations, this sequence may be changed without departing from the scope of this disclosure. For example, some of the operations depicted may be performed in parallel or in a different order that does not substantially affect the functionality of the process. In other examples, different components of the example device or system implementing the process may perform functions substantially simultaneously or in a specific order.
[0033] In Figure 2 Explained in the context Figure 3 .
[0034] In the context of the example use case being described in this technology, the book application 102 can be configured to request questions suitable for facilitating conversational reading of the story content 104. The book application 102 can automatically request and then present questions 106 to facilitate conversational reading, or a user can interact with the UI button 204 to cause the book application 102 to request and then present questions 106 to facilitate conversational reading.
[0035] Based on some examples, regardless of whether the user initiates a request for an issue, the method includes generating a prompt at box 302 from a combination of a first template and a portion of the content. For example, Figure 2 The illustrated prompt generation service 208 can generate prompts from a combination of a first template from prompt template 206 and a portion of story content 104.
[0036] In some implementations, the book application 102 runs on a personal computing device such as a laptop, smartphone, tablet, e-reader, or spatial computing device. The problem-generating application 202 may operate on the same device as the book application 102 or may run in a computing cloud. In some implementations, the problem-generating application 202 may be part of the book application 102. In some implementations, the book application 102 and the problem-generating application 202 are logically separate applications (e.g., the problem-generating application 202 may operate as a service-type application), and the book application 102 may interact with the problem-generating application 202 via an application programming interface.
[0037] The first template is one of several templates designed to make the output of the generative model facilitate discussion of the content—a type of question. For example, Figure 4 A table illustrating the types or categories of questions that facilitate conversational reading, outlining some aspects of this technology. Figure 4 It also includes descriptions of these types of questions, the objectives of these types of questions, and examples of these types of questions. Based on, for example, in Figure 4 The information displayed in the table indicates that the designed prompt template 206 can be created.
[0038] Specifically, Figure 4This document outlines the questions, categorized into five groups mapped to the acronym CROWD. CROWD identifies five types of prompts that can initiate a conversation—Completion, Recall, Open-ended, Wh-, and Distancing. Each prompt type has a specific educational objective. The final column includes examples for each prompt type in the story "The Three Little Pigs."
[0039] The prompt template 206 includes a reference to the story content 104, allowing the prompt template to be combined with portions of the story content 104 to generate a prompt. In some examples, the story content 104 may include information from a page range, or may include information from a specific paragraph. In some examples, the story content 104 may include text from the story, images from the story, or a combination of text and images.
[0040] Example prompts could be: “ As an early childhood reading instructor, I generate prompts to encourage conversation and participation in 'dialogue-based reading' of the text. A {hint type} hint, which has a {hint target}. Remember the language you used to create [PROMPT] and the text you extracted. The chosen topic must be suitable for children aged 4 to 6. Ensure the presentation is relevant and concise. Read the following text and... Use it to better understand the roles and events of the main text block. Do not use any text within the text for hints. The following template happens to format the response in JSON format: { "prompt": PROMPT } Using this context, generate a hint of type {hint type} for this main text: {current page text} " Therefore, as an example, we can have: {Prompt Type} = Open {Prompt Objective} = Encourage children to express their own thoughts and opinions about the story. This prompt should allow for creativity and imagination. Avoid questions that can be answered with simple yes or no.
[0041] {Previous Page Text} = Text from the previous page {Current Page Text} = Text from the current page of the story In some examples, the prompt can be used with " This is for [CURRENT_PAGE] and [PREVIOUS_PAGES] Some examples of inappropriate prompts: "And a list of issues that failed the quality response evaluation rubric and the reasons for correction."
[0042] Therefore, generating a prompt by the prompt generation service 208 may include selecting a prompt from the prompt template 206 and combining the selected prompt template with the relevant story content 104.
[0043] Based on some examples, the method includes providing the generated hints to the generative model at box 304. For example, Figure 2The illustrated suggestion generation service 208 can provide generated suggestions to a generative model. In some implementations, the generative model is a large language model.
[0044] Generative model 214 can operate on the same device (personal computing device) as the book application 102, or it can operate in a cloud environment. Generative model 214 is a base model or other model that has not been fine-tuned for the purpose of generating questions 106 to facilitate conversational reading. Hint generation service 208 and response suitability service 210 can interact with generative model 214 via one or more application programming interfaces.
[0045] Based on some examples, the method includes receiving candidate responses from the generative model at box 306. For example, Figure 2 The illustrated question generation application 202 can receive candidate responses from a generative model. These candidate responses are created by the generative model in response to generated prompts. As described above, these candidate responses may not be of sufficient quality because the generative model 214 has not been fine-tuned. To check whether the candidate responses are of sufficient quality, this technique can leverage the capabilities of the generative model 214 to evaluate the quality of the candidate responses against specific criteria.
[0046] According to some examples, the method includes providing the generative model with an initial suitability cue at box 308 to evaluate the output based on quality criteria. For example, Figure 2 The illustrated response suitability service 210 can provide a first suitability suggestion to the generative model to evaluate the output based on quality criteria. The suitability suggestion is created by selecting an appropriate suitability suggestion template from the quality response evaluation metric 212 and combining it with a candidate response. The first suitability suggestion can be one of several suitability suggestions provided to the generative model to evaluate the characteristics of the candidate response. The several suitability suggestions correspond to factors in the quality response evaluation metric.
[0047] The Quality Response Evaluation metric outlines the factors used to assess whether story-based questions will facilitate discussion of the content. Quality Response Evaluation metric 212 includes appropriateness prompt templates for each type of generation prompt (e.g., the CROWD prompt type for questions used to facilitate conversational reading examples). Multiple appropriateness prompts may exist for each generation prompt type, and these prompts can be configured to be presented to the generative model one at a time. For example, for the generation prompt type, there may be three appropriateness prompts—one for the wording of the candidate response, one for the authenticity of the candidate response, and one for the complexity of the candidate response. These appropriateness prompts may be delivered one at a time, in parallel, or all as part of a single appropriateness prompt.
[0048] The Quality Response Evaluation metric 212 can be created by one or more experts. In this way, expert input can be obtained and applied more efficiently, allowing experts to label thousands of samples for fine-tuning the generative model. The quality criteria in the Quality Response Evaluation metric incorporate expert knowledge. This approach is not only more efficient, but the expert knowledge is explicitly applied to all outputs of the generative model, whereas with fine-tuning, it can be difficult to verify that expert knowledge is being used for any particular candidate response.
[0049] According to some examples, the method includes receiving a suitability response from the generative model at box 310, where the suitability response describes whether a candidate response meets a quality criterion. For example, Figure 2 The illustrated response suitability service 210 can receive suitability responses from a generative model, where the suitability responses describe whether candidate responses meet quality criteria.
[0050] Based on some examples, the method includes evaluating the appropriateness of the response at decision box 312 to determine whether a candidate response is appropriate. For example, Figure 2 The illustrated response suitability service 210 can evaluate the suitability of responses to determine whether a candidate response is appropriate.
[0051] For example, suitability prompts are designed to generate suitability responses that provide a clear understanding of whether a candidate response meets a specific suitability criterion, and if not, provide a clear understanding of why. In some implementations, this can involve multiple prompts and responses, where a first response may indicate that a response is inappropriate, and subsequent prompts and responses may explain why a candidate response is considered inappropriate.
[0052] Once a candidate response is determined to be appropriate within the context of the appropriateness criteria explicitly expressed in the appropriateness cues, the method may proceed to determine at decision box 314 whether any other appropriateness cues should be presented. If other appropriateness cues exist to be presented, the method includes providing additional appropriateness cues to the generative model at box 316 to evaluate the output based on additional quality criteria.
[0053] When it is determined at decision box 314 that no other suitability prompts need to be presented, the method includes presenting a candidate response in the user interface at box 318. For example, question generation application 202 may provide suitable candidate responses to book application 102 to be presented as question 106 for facilitating conversational reading.
[0054] However, if a candidate response to any criterion explicitly expressed in a suitability suggestion according to a suitability suggestion is deemed inappropriate at decision box 312, the method includes generating a revision suggestion from the combination of the generation suggestion, the candidate response, and the suitability response at box 320. A candidate response is determined to be inappropriate when the suitability response includes a reason why a candidate response to a quality criterion included in the suitability suggestion is inappropriate. For example, Figure 2 The illustrated response suitability service 210 can generate a revision suggestion from a combination of a generation suggestion, a candidate response, and a second response. The revision suggestion request is followed by a revised response, which explains why the candidate response is unsuitable according to the quality criteria that the candidate response is deemed unsatisfactory.
[0055] Figure 5 Examples of problems based on this technique are illustrated, along with tables explaining why appropriate determinations might be made regarding these examples and the reasoning behind those conclusions. Figure 5 The questions in the table are based on the well-known story of "Goldilocks and the Three Bears." In this table, some example questions have been indicated as appropriate or inappropriate. Questions deemed inappropriate are due to reasons such as involving details unimportant to the plot, prompting speculation about things unrelated to the story's theme, or being easily evaded with a simple answer.
[0056] As asserted in this paper, this technique can provide similar quality responses with a much greater frequency than the un-fine-tuned base model, and in some respects, it can approach similar quality as if the generative model had been fine-tuned for a specific task. This assertion can be verified by evaluating the quality of questions generated by this technique and by evaluating the quality of questions generated by a generative model using the same prompts but without appropriateness response assessment. Questions generated by both methods were presented to four elementary school educators, who were asked to rate the likelihood of each question facilitating contextualized dialogue between parents and children on a scale from 1 (very unlikely) to 5 (very likely). Inferring that a question's score depends on both the bias of the generating system and the rater was performed using ordinal logistic regression, where the ratings assigned by the educators were the dependent variable, and the system and rater were the independent variables. Controlling for raters, it was found that questions generated by this technique (including appropriateness response assessment) were 1.64 times more likely to be rated than questions generated by a generative model lacking appropriateness response assessment. Of the 165 questions generated by this technique (including appropriateness response assessment), educators assigned scores of 3 or higher to 135 questions, representing an overall appropriateness rate of 79%. In contrast, generative models lacking appropriateness response assessment had an appropriateness rate of 69%.
[0057] While the above approach was proposed in the context of generating questions to facilitate conversational reading, this technique is not limited to such use cases. This technique can be adapted to any use case where sufficient quality response criteria for the use case can be reflected in the suitability tips. The higher the quality of the tips and the criteria they contain, the better the output of this technique should be.
[0058] An example of an additional use case might be using this technology to generate high-quality descriptions of apps or other content in an online store. In such an example, the app could provide generation prompts to allow the generative model to output a description of the content. The Response Suitability Service could then interact with the generative model to determine if the description of the content is of sufficiently high quality. Factors such as the use of marketing-appropriate language, the avoidance of a critical view of the product, and clarity—among others—might be considered. A suitable description can be published, while other descriptions can be revised to overcome their shortcomings.
[0059] Figure 6A An example of a book application at time t=0 is illustrated according to some aspects of this technology. The book application 102 includes story content 104, which comprises images and text associated with children's books. Additionally, the book application 102 may present a UI button 204 that can receive selections from a user requesting the generation of questions that can facilitate conversational reading.
[0060] Figure 6B An example of a book application at time t=1, based on some aspects of this technique, is illustrated, occurring after a suitable response has been provided by the generative model. The book application 102 displays the story content 104 again, but the UI button 204 has been replaced with a question 106 that promotes conversational reading, which has been generated by the generative model 104 and deemed appropriate by the question-generating application 202.
[0061] Figure 7 An example of a computing system 700 is shown. This computing system can be any computing device that constitutes, for example, a book application 102, a problem-generating application 202, or any component thereof, wherein the components of the system communicate with each other using a connection 702. The connection 702 can be a physical connection via a bus or a direct connection to a processor 704, such as in a chipset architecture. The connection 702 can also be a virtual connection, a networking connection, or a logical connection.
[0062] In some embodiments, computing system 700 is a distributed system, wherein the functions described herein may be distributed across a data center, multiple data centers, a peer-to-peer network, etc. In some embodiments, one or more system components described represent a plurality of such components that each perform some or all of the functions described for which the component is used. In one embodiment, these components may be physical devices or virtual devices.
[0063] The example computing system 700 includes at least one processing unit (CPU or processor) 704 and a connection 702 that couples various system components, including system memory 708, such as read-only memory (ROM) 710 and random access memory (RAM) 712, to the processor 704. The computing system 700 may include a cache of high-speed memory 706 that is directly connected to, close to, or integrated into the processor 704.
[0064] Processor 704 may include any general-purpose processor and hardware or software services (such as services 716, 718, and 720 stored in storage device 714), which are configured to control processor 704 as well as dedicated processors where software instructions are incorporated into the actual processor design. Processor 704 can essentially be a completely independent computing system, containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors can be symmetric or asymmetric.
[0065] To enable user interaction, the computing system 700 includes an input device 726, which can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice input, etc. The computing system 700 may also include an output device 722, which can be one or more of a plurality of output mechanisms known to those skilled in the art. In some cases, a multimodal system allows the user to provide multiple types of input / output to communicate with the computing system 700. The computing system 700 may include a communication interface 724, which typically governs and manages user input and system output. Operation is not limited to any particular hardware arrangement; therefore, the basic features described herein can be readily replaced with improved hardware or firmware arrangements after such arrangements have been developed.
[0066] Storage device 714 may be a non-volatile memory device and may be a hard disk or other type of computer-readable medium that can store data accessible by a computer, such as magnetic tape cassettes, flash memory cards, solid-state memory devices, digital universal optical discs, cartridges, random access memory (RAM), read-only memory (ROM), and / or some combination of these devices.
[0067] Storage device 714 may include software services, servers, etc., which enable the system to perform functions when the code defining such software is executed by processor 704. In some embodiments, hardware services that perform specific functions may include software components stored in a computer-readable medium and combined with necessary hardware components (such as processor 704, connection 702, output device 722, etc.) to perform functions.
[0068] For clarity, in some cases, the technology may be presented as comprising individual functional blocks, including functional blocks comprising devices, device components, steps or routines in methods embodied in software, or combinations of hardware and software.
[0069] Any of the steps, operations, functions, or processes described herein may be performed or implemented by a combination of hardware and software services, or individually or in combination with other devices. In some embodiments, a service may be software residing in the memory of a client device and / or one or more servers of a content management system, and performing one or more functions while the processor executes the software associated with the service. In some embodiments, a service is a program or set of programs that performs a specific function. In some embodiments, a service may be considered a server. The memory may be a non-transitory computer-readable medium.
[0070] In some implementations, computer-readable storage devices, media, and memories may include cables or wireless signals containing bit streams. However, by reference, non-transitory computer-readable storage media explicitly excludes media such as energy, carrier signals, electromagnetic waves, and the signal itself.
[0071] The methods described in the examples above can be implemented using computer-executable instructions stored in or otherwise obtainable from a computer-readable medium. Such instructions may include, for example, instructions and data that cause a general-purpose computer, a special-purpose computer, or a special-purpose processing device, or otherwise configure such device to perform a function or a set of functions. The portion of the computer resources used may be accessible via a network. The executable computer instructions may be, for example, binary, intermediate format instructions such as assembly language, firmware, or source code. Examples of computer-readable media that can be used to store instructions, information used, and / or information generated during the methods according to the described examples include disks or optical discs, solid-state storage devices, flash memory, USB devices with non-volatile memory, networked storage devices, etc.
[0072] Devices implementing the methods disclosed herein may include hardware, firmware, and / or software, and may take any of a variety of form factors. Typical examples of such form factors include servers, laptops, smartphones, small form factor personal computers, personal digital assistants, etc. The functionality described herein may also be embodied in peripheral devices or expansion cards. As another example, this functionality may also be implemented in different processes executed between different chips on a circuit board or within a single device.
[0073] Instructions, media for transmitting such instructions, computing resources for executing such instructions, and other structures for supporting such computing resources are components for providing the functionality described in these disclosures.
[0074] aspect: This technology includes a computer-readable storage medium for storing instructions, and a system for performing any of the methods embodied in the instructions set forth in the various aspects of this technology presented below: Aspect 1: A method for interacting with a large language model to obtain responses with improved quality without fine-tuning the large language model for use cases, the method comprising: generating a generated prompt from a combination of a first template and a portion of content by a prompt generation service, wherein the content is a story, wherein the first template is one of a plurality of templates designed to cause a generative model to output a question of a type that will facilitate discussion of the content, wherein the generative model is a large language model; providing the generated prompt to the generative model by the prompt generation service; and receiving candidate responses from the generative model by a response suitability service, wherein the candidate responses are responses of the generative model to the generated prompt. The process involves: providing a first suitability suggestion to the generative model to evaluate the output based on quality criteria, wherein the first suitability suggestion is one of a plurality of suitability suggestions provided to the generative model to evaluate the characteristics of the candidate response, wherein the plurality of suitability suggestions correspond to factors in a quality response evaluation metric, wherein the quality generation matrix describes factors used to evaluate whether a question based on the story would facilitate the discussion of the content; receiving a suitability response from the generative model by the response suitability service, wherein the suitability response describes whether the candidate response meets the quality criteria; and evaluating the suitability response by the response suitability service to determine whether the candidate response is appropriate.
[0075] Aspect 2: According to the method of Aspect 1, the method includes: when the evaluation of the suitability response results in the candidate response being determined to be unsuitable when evaluated according to the first suitability prompt, wherein the candidate response is determined to be unsuitable when the suitability prompt includes a reason for the candidate response being unsuitable according to the quality criteria included in the first suitability prompt, the prompt generation service generates a revision prompt from a combination of the generated prompt, the candidate response, and the suitability response, wherein the revision prompt request is revised in response, the revised response explaining the reason for the candidate response being unsuitable according to the quality criteria included in the first suitability prompt.
[0076] Aspect 3: The method according to any one of Aspects 1 to 2, the method comprising: when the evaluation of the suitability response results in the candidate response being determined to be suitable when evaluated in accordance with the first suitability cue, providing a second suitability cue to the generative model to evaluate the output based on a second quality criterion.
[0077] Aspect 4: The method according to any one of Aspects 1 to 3, the method according to claim 1, the method comprising: presenting the candidate response in a user interface when the evaluation of the suitability response results in the candidate response being determined to be suitable when evaluated in accordance with the first suitability prompt.
[0078] Aspect 5: The method according to any one of Aspects 1 to 4, wherein the revision suggestion is generated from a combination of the generation suggestion, the candidate response, and the suitability response.
[0079] Aspect 6: The method according to any one of Aspects 1 to 5, wherein the plurality of suitability tips correspond to quality criteria in a quality response evaluation metric.
[0080] Aspect 7: The method according to any one of Aspects 1 to 6, wherein the quality criteria in the quality response evaluation metric include expert knowledge.
Claims
1. A method, the method comprising: Candidate responses are received from the generative model by the response suitability service; Provide the generative model with a first suitability hint to evaluate the output based on quality criteria; The response suitability service receives a suitability response from the generative model, wherein the suitability response describes whether the candidate response meets the quality criteria; as well as The appropriateness of the response is evaluated by the appropriateness service to determine whether the candidate response is appropriate.
2. The method according to claim 1, further comprising: Generate prompts from a combination of the first template and a portion of the content; The generative model is provided with the generative hints, whereby the candidate responses are created by the generative model in response to the generative hints.
3. The method according to claim 1, wherein the method comprises: When the assessment of the suitability response leads to the determination that the candidate response is unsuitable, wherein the candidate response is determined to be unsuitable when the suitability suggestion includes a reason for the unsuitability of the candidate response according to the quality criteria included in the first suitability suggestion. Generate a revision suggestion, wherein the revision suggestion request is revised to a response, and the revised response explains the reasons why the candidate response is unsuitable according to the quality criteria included in the first suitability suggestion.
4. The method of claim 3, wherein the revision suggestion is generated from a combination of the generation suggestion, the candidate response, and the suitability response.
5. The method according to claim 1, wherein the method comprises: When the assessment of the suitability response leads to the candidate response being determined to be suitable when assessed according to the first suitability cues, A second suitability cues are provided to the generative model to evaluate the output based on a second quality criterion.
6. The method of claim 1, wherein evaluating the suitability response leads to determining that the candidate response is suitable, the method comprising: The candidate response is presented in the user interface.
7. The method of claim 1, wherein the first suitability suggestion is one of a plurality of suitability suggestions provided to the generative model to evaluate the characteristics of the candidate response.
8. The method of claim 7, wherein the plurality of suitability tips correspond to quality criteria in a quality response evaluation metric.
9. The method of claim 8, wherein the quality criteria in the quality response evaluation metric include expert knowledge.
10. The method of claim 2, wherein the first template is one of a plurality of templates, wherein the template is designed to make the generative model output a type of question that will facilitate discussion of the content.
11. A computing system, the computing system comprising: At least one processor; and The memory stores instructions that, when executed by the at least one processor, configure the system to: Candidate responses are received from the generative model by the response suitability service; Provide the generative model with a first suitability hint to evaluate the output based on quality criteria; The response suitability service receives a suitability response from the generative model, wherein the suitability response describes whether the candidate response meets the quality criteria; as well as The appropriateness of the response is evaluated by the appropriateness service to determine whether the candidate response is appropriate.
12. The computing system of claim 11, wherein the instructions further configure the system to: Generate prompts from a combination of the first template and a portion of the content; The generative prompts are provided to the generative model, whereby the candidate responses are created by the generative model in response to the generative prompts.
13. The computing system according to claim 11, wherein the instructions include: When the assessment of the suitability response leads to the determination that the candidate response is unsuitable, wherein the candidate response is determined to be unsuitable when the suitability suggestion includes a reason for the unsuitability of the candidate response according to the quality criteria included in the first suitability suggestion. Generate a revision suggestion, wherein the revision suggestion request is revised to a response, and the revised response explains the reasons why the candidate response is unsuitable according to the quality criteria included in the first suitability suggestion.
14. The computing system of claim 13, wherein the revision suggestion is generated from a combination of the generation suggestion, the candidate response, and the suitability response.
15. The computing system according to claim 11, wherein the instructions include: When the assessment of the suitability response leads to the candidate response being determined to be suitable when assessed according to the first suitability cues, A second suitability cues are provided to the generative model to evaluate the output based on a second quality criterion.
16. A non-transitory computer-readable storage medium, the non-transitory computer-readable storage medium comprising instructions that, when executed by at least one processor, cause the at least one processor to: Candidate responses are received from the generative model by the response suitability service; Provide the generative model with a first suitability hint to evaluate the output based on quality criteria; The response suitability service receives a suitability response from the generative model, wherein the suitability response describes whether the candidate response meets the quality criteria; as well as The appropriateness of the response is evaluated by the appropriateness service to determine whether the candidate response is appropriate.
17. The computer-readable storage medium of claim 16, wherein the instructions further configure the at least one processor to: Generate prompts from a combination of the first template and a portion of the content; The generative prompts are provided to the generative model, whereby the candidate responses are created by the generative model in response to the generative prompts.
18. The computer-readable storage medium of claim 16, wherein the instructions comprise: When the assessment of the suitability of the response leads to the determination that the candidate response is unsuitable... Generate a revision suggestion, wherein the revision suggestion request is revised to a response, and the revised response explains why the candidate response is unsuitable according to the quality criteria included in the first suitability suggestion.
19. The computer-readable storage medium of claim 16, wherein the instructions comprise: When the assessment of the suitability response leads to the candidate response being determined to be suitable when assessed according to the first suitability cues, A second suitability cues are provided to the generative model to evaluate the output based on a second quality criterion.
20. The computer-readable storage medium of claim 17, wherein the first template is one of a plurality of templates, wherein the plurality of templates are designed such that the output of the generative model will facilitate a discussion of the content as a type of question.