Prompt generating device and method

By using a prompt generator and reinforcement learning training, the optimal prompt is determined automatically, solving the time-consuming prompt design problem in existing technologies and improving the efficiency of personalized responses for large language models.

CN121658591APending Publication Date: 2026-03-13HON HAI PRECISION INDUSTRY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

In existing technologies, the process of designing prompts and evaluating personalized responses from large language models is time-consuming and inefficient, and cannot effectively induce personalized responses.

Method used

The prompt generator converts situational questions into multiple candidate prompts, and then uses reinforcement learning training based on a large language model and a feedback model to automatically determine the best prompt to guide the large language model to generate a personalized response.

Benefits of technology

It reduces the time cost of designing prompts and improves the efficiency of generating personalized responses from large language models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121658591A_ABST
    Figure CN121658591A_ABST
Patent Text Reader

Abstract

The invention discloses a prompt generating device and method. The cue generating device receives a contextual question. The prompt generation device transmits a target personality category in a plurality of personality categories and the situation question to a prompt generator to generate a plurality of candidate prompts corresponding to the target personality category and a plurality of reward signals corresponding to the plurality of candidate prompts, wherein the cue generator is trained based on a large language model and feedback models corresponding to the plurality of personality categories. The cue generation device determines an optimal cue from among the plurality of candidate cues based on the plurality of reward signals corresponding to the plurality of candidate cues. According to the prompt generation technology provided by the invention, the time cost of prompt design is reduced, and a large language model is guided to effectively generate personalized replies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a prompt generation device and control method. Specifically, this invention relates to a prompt generation device and control method capable of automatically generating optimal prompts to guide a large language model in producing personalized responses. Background Technology

[0002] In recent years, the demand for dialogue with users based on large language models has been increasing in various industrial fields (such as autonomous driving). Among them, technologies that design prompts to induce large language models to generate personalized responses have been proposed one after another.

[0003] In existing technologies, prompts input to large language models are generated manually, and the personalized responses induced by these prompts are also evaluated manually. However, in this scenario, the process of manually designing prompts and manually evaluating responses is quite time-consuming and cannot effectively induce personalized responses from large language models.

[0004] In view of this, the industry urgently needs to develop a prompt generation technology that can automatically design prompts to reduce the time cost of prompt design and guide large language models to effectively generate personalized responses. Summary of the Invention

[0005] One object of the present invention is to provide a prompt generation device, comprising a transceiver interface and a processor. The transceiver interface is used to receive contextual questions, and the processor is electrically connected to the transceiver interface. The processor transmits a target personality category from a plurality of personality categories and the contextual question to the prompt generator to generate a plurality of candidate prompts corresponding to the target personality category. The prompt generator is trained based on a large language model and a feedback model corresponding to the plurality of personality categories. The processor determines the optimal prompt from the plurality of candidate prompts based on a plurality of reward signals corresponding to the plurality of candidate prompts.

[0006] In one embodiment of the present invention, the cue generator is trained based on the following operations: transmitting a plurality of catalytic cues to the large language model to generate a plurality of evoked responses corresponding to the plurality of catalytic cues, wherein the plurality of catalytic cues are generated based on the cue generator; transmitting the plurality of evoked responses to the feedback model to generate a plurality of personalized feedback information corresponding to the plurality of evoked responses, wherein each of the plurality of personalized feedback information includes a plurality of personalized feedback scores corresponding to the plurality of personalized categories; and training the cue generator based on the plurality of catalytic cues and the plurality of personalized feedback information.

[0007] In one embodiment of the present invention, the feedback model is trained based on the following operation: the feedback model is trained based on multiple training data corresponding to the multiple personality categories.

[0008] In one embodiment of the invention, the processor further performs the following operations: transmitting the plurality of training data corresponding to the plurality of personality categories to an augmented large language model to generate a plurality of augmented training data corresponding to the plurality of personality categories; and training the feedback model based on the plurality of augmented training data corresponding to the plurality of personality categories.

[0009] In one embodiment of the invention, the processor is further configured to perform the following operations: transmit the optimal suggestion to the large language model to generate a response message corresponding to the target personality category.

[0010] In one embodiment of the invention, the processor further performs the following operations: controlling the human-machine interface to display the response message; and receiving a confirmation signal from the human-machine interface and updating the prompt generator based on the confirmation signal.

[0011] In one embodiment of the present invention, the large language model is trained based on the following operation: the large language model is trained based on multiple enhancement instructions and multiple test labels corresponding to the multiple enhancement instructions.

[0012] In one embodiment of the present invention, the plurality of enhancement instructions are generated based on the following operation: transmitting a plurality of test instructions and an evolutionary enhancement method to the large language model to generate the plurality of enhancement instructions, wherein the evolutionary enhancement method is generated based on the initial enhancement method.

[0013] In one embodiment of the present invention, the evolutionary enhancement method is generated based on the following operations: transmitting a plurality of training instructions and the initial enhancement method to the large language model to generate a plurality of enhancement training instructions, wherein the initial enhancement method is used to indicate the enhancement type corresponding to the plurality of training instructions; transmitting the plurality of enhancement training instructions to the large language model to generate a plurality of enhancement training responses; and transmitting enhancement prompts and a plurality of enhancement feedback information to the large language model to generate the evolutionary enhancement method based on the initial enhancement method.

[0014] In one embodiment of the present invention, the plurality of augmented feedback information is generated based on the following operation: analyzing the plurality of augmented training instructions based on the large language model, and comparing the plurality of augmented training responses with the plurality of training labels corresponding to the plurality of training instructions to generate the plurality of augmented feedback information.

[0015] Another object of the present invention is to provide a prompt generation method for use in an electronic device. The prompt generation method includes the following steps: transmitting a target personality category from a plurality of personality categories and a contextual question to a prompt generator to generate a plurality of candidate prompts corresponding to the target personality category, wherein the prompt generator is trained based on a large language model and a feedback model corresponding to the plurality of personality categories; and determining the optimal prompt from the plurality of candidate prompts based on a plurality of reward signals corresponding to the plurality of candidate prompts.

[0016] In one embodiment of the present invention, the cue generator is trained based on the following steps: transmitting multiple catalytic cues to the large language model to generate multiple evoked responses corresponding to the multiple catalytic cues, wherein the multiple catalytic cues are generated based on the cue generator; transmitting the multiple evoked responses to the feedback model to generate multiple personalized feedback information corresponding to the multiple evoked responses, wherein each of the multiple personalized feedback information includes multiple personalized feedback scores corresponding to the multiple personalized categories; and training the cue generator based on the multiple catalytic cues and the multiple personalized feedback information.

[0017] In one embodiment of the present invention, the feedback model is trained based on the following steps: training the feedback model based on multiple training data corresponding to the multiple personality categories.

[0018] In one embodiment of the present invention, the method further includes the following steps: transmitting the plurality of training data corresponding to the plurality of personality categories to an augmented large language model to generate a plurality of augmented training data corresponding to the plurality of personality categories; and training the feedback model based on the plurality of augmented training data corresponding to the plurality of personality categories.

[0019] In one embodiment of the present invention, the method further includes the following step: transmitting the optimal suggestion to the large language model to generate a response message corresponding to the target personality category.

[0020] In one embodiment of the present invention, the method further includes the following steps: controlling the human-machine interface to display the response message; and receiving a confirmation signal from the human-machine interface and updating the prompt generator based on the confirmation signal.

[0021] In one embodiment of the present invention, the large language model is trained based on the following steps: training the large language model based on multiple enhancement instructions and multiple test labels corresponding to the multiple enhancement instructions.

[0022] In one embodiment of the present invention, the plurality of enhancement instructions are generated based on the following steps: transmitting a plurality of test instructions and an evolutionary enhancement method to the large language model to generate the plurality of enhancement instructions, wherein the evolutionary enhancement method is generated based on the initial enhancement method.

[0023] In one embodiment of the present invention, the evolutionary enhancement method is generated based on the following steps: transmitting a plurality of training instructions and the initial enhancement method to the large language model to generate a plurality of enhancement training instructions, wherein the initial enhancement method is used to indicate the enhancement type corresponding to the plurality of training instructions; transmitting the plurality of enhancement training instructions to the large language model to generate a plurality of enhancement training responses; and transmitting enhancement prompts and a plurality of enhancement feedback information to the large language model to generate the evolutionary enhancement method based on the initial enhancement method.

[0024] In one embodiment of the present invention, the plurality of augmented feedback information is generated based on the following steps: based on the large language model, analyzing the plurality of augmented training instructions, and comparing the plurality of augmented training responses with the plurality of training labels corresponding to the plurality of training instructions to generate the plurality of augmented feedback information.

[0025] The prompt generation technology (including at least an apparatus and method) provided by this invention converts a situational question into multiple candidate prompts through a prompt generator, and determines the optimal prompt from among the multiple candidate prompts. Then, based on a large language model and a feedback model, this invention trains the prompt generator through reinforcement learning. Furthermore, this invention transmits the optimal prompt to the large language model to generate a response message corresponding to the target personality category. Since this invention can automatically generate prompts through a prompt generator trained based on reinforcement learning, the prompt generation technology provided by this invention reduces the time cost of designing prompts and guides the large language model to effectively generate personalized responses.

[0026] The following detailed description of the technology and embodiments of the present invention, in conjunction with the accompanying drawings, enables those skilled in the art to understand the technical features of the claimed invention. Attached Figure Description

[0027] Figure 1 A schematic diagram illustrating the architecture of the prompt generation device according to the first embodiment;

[0028] Figure 2 A schematic diagram illustrating the memory of the first embodiment;

[0029] Figure 3 A schematic diagram illustrating the prompt generation process of the first embodiment;

[0030] Figure 4A schematic diagram illustrating the training process of a prompt generator in some implementations;

[0031] Figure 5 A schematic diagram illustrating the training process of a feedback model in certain implementations;

[0032] Figure 6 A schematic diagram illustrating the generation process of augmented training instructions, augmented training responses, and evolutionary augmentation methods in certain implementations;

[0033] Figure 7 A schematic diagram illustrating the enhanced instruction generation process for certain implementation methods; and

[0034] Figure 8 A flowchart illustrating the prompt generation method of the second embodiment. Detailed Implementation

[0035] The following description of embodiments will explain the prompt generation apparatus provided by the present invention. However, these embodiments are not intended to limit the implementation of the invention to any of the environments, applications, or methods described in the embodiments. Therefore, the description of the embodiments is for illustrative purposes only and is not intended to limit the scope of the invention. It should be understood that in the following embodiments and drawings, elements not directly related to the present invention have been omitted and are not shown, and the dimensions of each element and the proportions between elements are merely illustrative and not intended to limit the scope of the invention.

[0036] A prompt generation device according to a first embodiment of the present invention is schematically depicted in Figure 1 .like Figure 1 As shown, the prompt generating device 1 includes a processor 11, a memory 12 and a transceiver interface 13, and the processor 11 is electrically connected to the memory 12 and the transceiver interface 13.

[0037] The prompt generation device 1 of the present invention can be used in any environment where a user can interact with a large language model. For example, the prompt generation device 1 can be installed inside a vehicle, and the user (e.g., driver or passenger) can interact with the prompt generation device 1.

[0038] In this embodiment, such as Figure 2As shown, memory 12 stores a prompt generator (PG), a large language model (LLM), and a feedback model (RM). The prompt generator (PG) is a language model that can convert input prompts into advanced prompts (e.g., converting input prompts into prompts with different tones or more complex grammar). The large language model (LLM) is a large language model that can understand input prompts and generate responses corresponding to them. The feedback model (RM) is a regression model that can provide feedback on input information (e.g., a score indicating the degree to which the input text corresponds to a specific personality trait).

[0039] In this embodiment, the transceiver interface 13 is used to receive contextual questions. For example, the prompt generating device 1 also includes a sensor (e.g., a microphone). The contextual questions received by the transceiver interface 13 are the contextual questions in text form that are generated when the user speaks and the processor 11 translates the voice signal sensed by the sensor into text form.

[0040] In some embodiments, the prompt generating device 1 further includes a human-machine interface (HMI), where the contextual question received by the transceiver interface 13 is a contextual question input by the user through the HMI. For example, the user can input the contextual question in text form by operating a touchscreen interface. It should be noted that the HMI is an interface that can be controlled by the user to generate input signals and report system status to the user.

[0041] It should be noted that processor 11 may be various processing units, central processing units (CPUs), microprocessors, or other computing devices known to those skilled in the art to which this application pertains. Memory 12 may be memory, universal serial bus (USB), hard disk, optical disk, removable storage, or any other storage medium or circuitry known to those skilled in the art to which this application pertains and having the same function. Transceiver interface 13 is an interface capable of receiving and transmitting data, or other interfaces capable of receiving and transmitting data known to those skilled in the art to which this application pertains. Transceiver interface 13 may receive data from sources such as external devices, external web pages, external applications, etc.

[0042] In this invention, the main approach is to transmit the target personality category and contextual question to the prompt generator (PG) to generate an optimal prompt, and then guide the large language model (LLM) to generate a response message corresponding to the target personality category based on the optimal prompt. The following paragraphs will describe in detail the implementation details related to this invention.

[0043] In the first embodiment, for ease of explanation, please refer to Figure 3The prompt generation process is illustrated in diagram 300. Specifically, the processor 11 transmits the target personality category TPC from multiple personality category PCs and the contextual question SQ to the prompt generator PG to generate multiple candidate prompts CP corresponding to the target personality category TPC and multiple reward signals RS corresponding to the multiple candidate prompts CP.

[0044] First, the memory 12 stores multiple personality category PCs, which can be determined based on the personalized text dataset used in this invention. For example, the personalized text dataset can be the Big Five Personality Dataset, and the multiple personality category PCs are the five personality traits described in the Big Five Personality Dataset: Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism.

[0045] It should be noted that the personalized text dataset used in this invention is not limited to the Big Five personality dataset mentioned above, but can be any dataset containing text corresponding to different personality traits.

[0046] Next, the target personality category TPC is one of the plurality of personality category PCs, and the target personality category TPC can be determined by the user. For example, the processor 11 controls the human-machine interface to display a plurality of option buttons corresponding to the plurality of personality category PCs, and the user operates the human-machine interface to select one of the plurality of option buttons to determine the target personality category TPC.

[0047] In this embodiment, the prompt generator PG is a language model that can convert input prompts into more advanced prompts. The processor 11 transmits the target personality category (TPC) and the contextual question (SQ) to the prompt generator PG to generate multiple candidate prompts (CP) corresponding to the target personality category (TPC). In some embodiments, the target personality category (TPC) and the contextual question (SQ) can be input to the prompt generator PG in text form.

[0048] It should be noted that since the multiple candidate prompts (CPs) are generated by the prompt generator (PG) based on the target personality category (TPC) and the situational question (SQ), the content described by the multiple candidate prompts (CPs) and the situational question (SQ) is similar, but the way they are described is different.

[0049] For example, the target personality category (TPC) is "Openness," the situational question (SQ) is "Hello, how's the weather today?", and the prompt generator (PG) generates multiple candidate prompts (CPs), including a first candidate prompt and a second candidate prompt. The first candidate prompt is "Hi! How do you think the weather is today?", and the second candidate prompt is "Hello, I really want to know today's weather!". From the situational question (SQ) and the multiple candidate prompts (CPs), it can be seen that the content of both the situational question (SQ) and the multiple candidate prompts (CPs) is "asking about today's weather," but the narration style of the multiple candidate prompts (e.g., tone, interjections, punctuation, etc.) is different.

[0050] Furthermore, the cue generator PG can also generate multiple reward signals RS corresponding to the plurality of candidate cue CPs, wherein each of the multiple reward signals RS is presented numerically, and the multiple reward signals RS are used to indicate the degree of conformity of each of the multiple candidate cue CPs with the target personality category TPC. In some embodiments, the multiple reward signals RS are used to indicate the degree of conformity of each of the multiple candidate cue CPs with the multiple personality category PCs.

[0051] It should be noted that the initial model of the prompt generator PG is a pre-trained language model. The prompt generator PG needs to be fine-tuned and trained to improve its ability to generate prompts. Specifically, the prompt generator PG is trained based on the large language model LLM and the feedback model RM corresponding to the multiple personality categories PC.

[0052] In a first embodiment, specifically, the processor 11 determines the optimal prompt BP from the plurality of candidate prompt CPs based on the plurality of reward signals RS corresponding to the plurality of candidate prompt CPs.

[0053] It should be noted that the plurality of candidate cues (CPs) are multiple different discrete cues explored by the cue generator (PG) in a multi-dimensional vocabulary space, and the cue generator (PG) tests the plurality of candidate cues (CPs) to select the best-performing one as the best cue (BP). For example, the processor 11 selects the one with the highest corresponding reward signal (RS) from the plurality of candidate cues (CPs) as the best cue (BP).

[0054] It should be noted that the optimal feedback (BP) is used to instruct the large language model (LLM) to generate a response message corresponding to the target personality category (TPC). In other words, the optimal feedback (BP) is used to enhance the personality characteristics of the generated response message by the large language model (LLM), making the content of the response message more consistent with the target personality category (TPC).

[0055] In some implementations, the cue generator PG is trained using reinforcement learning. For clarity, please refer to [link to relevant documentation]. Figure 4 The training process diagram 400 for the cue generator is shown below. First, the processor 11 transmits multiple catalytic cues (CTPs) to the large language model (LLM) to generate multiple evoked responses (ERs) corresponding to the multiple catalytic cues (CTPs). In other words, the multiple evoked responses (ERs) are personalized responses generated by the large language model (LLM) after receiving the multiple catalytic cues (CTPs).

[0056] It should be noted that the multiple catalytic prompts (CTPs) are generated based on the prompt generator (PG). The prompt generator (PG) is a language model; therefore, this invention can generate the multiple catalytic prompts (CTPs) by giving the prompt generator (PG) simple instructions. For example, giving the prompt generator (PG) the instruction "Please generate multiple prompts related to weather conditions" will allow the prompt generator (PG) to generate multiple catalytic prompts (CTPs) corresponding to the instruction in text form.

[0057] Next, the processor 11 transmits the plurality of evoked responses (ERs) to the feedback model RM to generate a plurality of personalized feedback information PRIs corresponding to the plurality of evoked responses (ERs), wherein each of the plurality of personalized feedback information PRIs includes a plurality of personalized feedback scores corresponding to the plurality of personalized categories (PCs).

[0058] For example, taking the five personality traits in the Big Five personality dataset as an example of the multiple personality categories PC, each of the multiple personality feedback information PRI includes feedback scores corresponding to the five personality categories of openness, cognitive, extraversion, harmony, and neuroticism.

[0059] It should be noted that each of the multiple feedback scores can be a value within a range (e.g., a maximum value of 1 and a minimum value of 0), where the value is used to indicate the degree of conformity of the personality category PC. For example, among the multiple personality feedback information PRI of the evoked response ER, the personality feedback score corresponding to openness is 0.9, indicating that the evoked response ER has a high degree of conformity to openness; the personality feedback score corresponding to harmony is 0.1, indicating that the evoked response ER has a low degree of conformity to harmony.

[0060] Finally, the processor 11 trains the cue generator PG using reinforcement learning based on the multiple catalytic cues (CTPs) and the multiple personality feedback information (PRIs). In other words, the feedback model (RM) is used to evaluate the multiple evoked responses (ERs) generated by the large language model (LLM) according to different personality categories (PCs), thereby training the cue generator PG to enhance its understanding of "which personality category PC evoked response ER can be generated by the large language model (LLM) from the multiple catalytic cues (CTPs) generated by the cue generator PG itself". Furthermore, the training process of the cue generator PG can be repeated to further enhance its capabilities.

[0061] In some embodiments, for ease of explanation, please refer to Figure 5 The feedback model training process is illustrated in diagram 500. Specifically, the processor 11 trains the feedback model RM based on multiple training data TDs corresponding to the multiple personality category PCs. It should be noted that the multiple training data TDs contain multiple training texts, and each training text corresponds to one of the multiple personality category PCs. The multiple training data TDs can be generated in various ways. For example, the multiple training data TDs can be provided by open-source personalized text datasets (e.g., the Big Five Personality Dataset, PANDORA). Alternatively, the multiple training data TDs can also be generated by manually inputting text and manually labeling it with the corresponding personality category PC.

[0062] In some embodiments, the present invention further increases the amount and diversity of the training data TDs by performing data augmentation on the plurality of training data TDs. Specifically, the processor 11 transmits the plurality of training data TDs corresponding to the plurality of personality category PCs to an augmentation large language model ALLM to generate a plurality of augmented training data ATDs corresponding to the plurality of personality category PCs. Then, the processor 11 trains the feedback model RM based on the plurality of augmented training data ATDs corresponding to the plurality of personality category PCs.

[0063] It should be noted that the large language model ALLM used for augmentation can be any large language model that can transform input text into advanced text (e.g., transform input text into text with more complex grammar).

[0064] In some implementations, augmentation instructions can also be manually input into the large language model ALLM for augmentation, wherein the augmentation instructions are used to indicate the augmented form of the augmented training data ATD generated by the large language model ALLM for augmentation. For example, the augmentation instructions may be "use a more complex grammar to augment the input text" or "use a different narrative order to augment the input text," where "more complex grammar" and "different narrative order" are both augmented forms indicated by the augmentation instructions.

[0065] In some implementations, specifically, the processor 11 transmits the Best Hint (BP) to the Large Language Model (LLM) to generate the response message corresponding to the Target Personality Category (TPC).

[0066] It should be noted that if the processor 11 transmits the contextual question SQ to the large language model LLM, the large language model LLM will generate a basic response message. Since the basic response message is generated based on the contextual question SQ, and the response message is generated based on the best prompt BP, and the best prompt BP is generated by the prompt generator PG based on the contextual question SQ, the content described by the basic response message and the response message has a certain degree of similarity.

[0067] Next, since the optimal suggestion (BP) is used to enhance the personality characteristics of the large language model (LLM) in generating response messages, the response message is more consistent with the target personality category (TPC) than the basic response message.

[0068] For example, the target personality category TPC is "Openness," and the situational question SQ is "You are attending a concert or live performance. How much do you enjoy the crowd atmosphere and the energy of the event?" The basic response message induced by the situational question SQ is "I enjoy the atmosphere of the event, but I don't enjoy the crowd." The prompt generation device 1 generates the optimal prompt BP based on the situational question SQ, and the response message generated based on the optimal prompt BP is "I enjoy the crowd atmosphere and the energy of the event." Obviously, the response message induced by the optimal prompt BP is more consistent with the target personality category TPC of "Openness" than the basic response message induced by the situational question SQ.

[0069] In some embodiments, the prompt generating device 1 can display the response message through the human-machine interface during interaction with the user, and ask questions to the user through the human-machine interface. The user can operate the human-machine interface to generate a confirmation signal, and further train and update the prompt generator PG based on the feedback.

[0070] For example, processor 11 controls the touchscreen interface to display the response message and also controls the touchscreen interface to display a question message (e.g., "Which personality category do you think this response fits?"). Next, processor 11 controls the touchscreen interface to display multiple feedback buttons corresponding to the multiple personality category PCs (e.g., a button for "openness," a button for "cognition," etc.). Assuming the user believes the response message fits "openness," the user can operate the touchscreen interface to select the "openness" feedback button to generate the confirmation signal. This confirmation signal instructs the user to select the feedback button corresponding to which personality category PC, as feedback to the prompt generator PG.

[0071] Accordingly, the processor 11 can train and update the cue generator PG based on the feedback to enhance the cue generator PG so that the optimal cue BP generated by the cue generator PG can better induce the large language model LLM to generate response messages that conform to the multiple personality categories PC.

[0072] Specifically, the processor 11 controls the human-machine interface to display the response message. Then, the processor 11 receives a confirmation signal from the human-machine interface and updates the prompt generator PG based on the confirmation signal.

[0073] In addition, in some embodiments, the user can also operate the human-machine interface to input feedback suggestions in the form of messages to generate the confirmation signal, wherein the confirmation signal includes the feedback suggestions input by the user.

[0074] In some embodiments, to enable the large language model (LLM) to understand more complex input instructions, the prompt generation device 1 further enhances multiple training instructions based on an initial enhancement method, evaluates the initial enhancement method to generate an evolutionary enhancement method, generates multiple enhanced instructions through the evolutionary enhancement method, and trains the large language model (LLM) based on the multiple enhanced instructions. Accordingly, the large language model (LLM) can understand more complex input instructions.

[0075] It should be noted that the prompt generating device 1 can communicate with a cloud database, and the cloud database stores a command training dataset. The command training dataset includes multiple training commands, multiple training labels corresponding to the multiple training commands, multiple test commands, and multiple test labels corresponding to the multiple test commands. The multiple training commands or the multiple test commands can be any command. The multiple training labels are responses to the multiple training commands, and the multiple test labels are responses to the multiple training commands. All commands and responses are in text format. For example, the training command is "How's the weather today?", and the corresponding training label is "The weather is very good today."

[0076] It should be noted that the instruction training dataset can be an open-source dataset from the network (e.g., AlpacaDataset, OASST1, etc.) or it can be generated manually.

[0077] First, in order to enhance the multiple training instructions, the generation device 1 is prompted to use an initial enhancement method to generate multiple enhanced training instructions. For clarity, please refer to... Figure 6 The flowchart 601 illustrates the process of generating augmented training instructions. Specifically, the processor 11 transmits the multiple training instructions TRI and the initial augmentation method IEM to the large language model LLM to generate multiple augmented training instructions ETRI.

[0078] It should be noted that the initial augmentation method (IEM) can be manually defined text, and the initial augmentation method (IEM) is used to indicate the augmentation type corresponding to the plurality of training instructions (TRIs). For example, the initial augmentation method (IEM) is "Please augment the input instructions with more complex grammar," where "more complex grammar" is the augmentation type corresponding to the plurality of training instructions (TRIs). Accordingly, the large language model (LLM) will generate the plurality of augmentation training instructions (ETRIs) that conform to "more complex grammar."

[0079] Next, please refer to Figure 6 The flowchart 603 illustrates the augmented training response generation process. Specifically, the processor 11 transmits the multiple augmented training instructions ETRI to the large language model LLM to generate multiple augmented training responses ETRR corresponding to the multiple augmented training instructions ETRI.

[0080] In some implementations, in order to measure the quality of the plurality of augmented training instructions (ETRI) and the plurality of augmented training responses (ETRR), specifically, the processor 11 analyzes the plurality of augmented training instructions (ETRI) based on the large language model (LLM) and compares the plurality of augmented training responses (ETRR) and the plurality of training labels corresponding to the plurality of training instructions (TRI) to generate the plurality of augmented feedback information (EFI).

[0081] It should be noted that in this embodiment, feedback role prompts can be manually given to the large language model LLM. The feedback role prompts are used to instruct the large language model LLM to measure the quality of the multiple reinforcement training instructions (ETRI) and the multiple reinforcement training responses (ETRR) in a specific role, task, style, scoring method, etc.

[0082] For example, the feedback role prompt could be: "Your role is a professional article reviewer, responsible for evaluating the quality of the rewritten article. Your scores must be objective and consistent each time. Your task is to analyze all the instructions I provide and compare the responses generated by these instructions with the corresponding tags. Your style should be professional and easy to understand, with a clear and organized structure. Your scoring method is to analyze each instruction one by one according to the following evaluation criteria, assigning a score from 0 to 10, and explaining the reasons in detail. These evaluation criteria include clarity, retention of original meaning, depth, diversity of viewpoints, and vocabulary diversity."

[0083] Next, please refer to Figure 6 The evolutionary enhancement method generation process diagram 605 shows that, specifically, the processor 11 transmits enhancement prompts (EP) and multiple enhancement feedback information (EFI) to the large language model (LLM) to generate an evolutionary enhancement method (EEM) based on the initial enhancement method (IEM).

[0084] It should be noted that the enhancement prompt EP is used to guide the large language model LLM to perform the task of enhancing the initial enhancement method IEM. For example, the enhancement prompt EP is "Please further optimize the initial enhancement method according to the multiple enhancement feedback information" to instruct the large language model LLM to optimize the initial enhancement method IEM to generate the evolutionary enhancement method EEM.

[0085] In some embodiments, for ease of explanation, please refer to Figure 7 The flowchart 700 illustrates the enhancement instruction generation process. Specifically, the processor 11 transmits multiple test instructions TEI and the evolutionary enhancement method EEM to the large language model LLM to generate the multiple enhancement instructions EI, wherein the evolutionary enhancement method EEM is generated based on the initial enhancement method IEM.

[0086] It should be noted that the Evolutionary Enhancement Method (EEM) is used to indicate the enhancement type corresponding to the plurality of test instructions (TEI). Since the EEM is generated based on the optimization of the initial enhancement method (IEM) (e.g., changing the narrative structure of the text, increasing the difficulty of the grammar, etc. based on the initial enhancement method IEM), the EEM has a stronger "enhancement instruction capability" compared to the initial enhancement method IEM.

[0087] Next, the processor 11 trains the large language model LLM based on the multiple enhancement instructions EI. Accordingly, the large language model LLM can understand more complex input instructions.

[0088] As described above, the prompt generation device 1 provided by this invention converts a situational question into multiple candidate prompts through a prompt generator, and determines the optimal prompt from among these candidate prompts. Next, based on a large language model and a feedback model, this invention trains the prompt generator using reinforcement learning. Furthermore, this invention transmits the optimal prompt to the large language model to generate a response message corresponding to the target personality category. Since this invention can automatically generate prompts through a prompt generator trained using reinforcement learning, the prompt generation device 1 provided by this invention reduces the time cost of designing prompts and guides the large language model to effectively generate personalized responses.

[0089] The second embodiment of the present invention is a prompt generation method, the flowchart of which is depicted in... Figure 8 The prompt generation method 800 is applicable to electronic devices, such as the prompt generation device 1 described in the first embodiment. The prompt generation method 800 generates a prompt through steps S801 to S803.

[0090] First, in step S801, the electronic device transmits the target personality category and the contextual question from multiple personality categories to the prompt generator to generate multiple candidate prompts corresponding to the target personality category and multiple reward signals corresponding to the multiple candidate prompts, wherein the prompt generator is trained based on a large language model and a feedback model corresponding to the multiple personality categories.

[0091] Next, in step S803, the electronic device determines the best prompt from the plurality of candidate prompts based on the plurality of reward signals corresponding to the plurality of candidate prompts.

[0092] In some embodiments, the cue generator is trained based on the following steps: transmitting multiple catalytic cues to the large language model to generate multiple evoked responses corresponding to the multiple catalytic cues, wherein the multiple catalytic cues are generated based on the cue generator; transmitting the multiple evoked responses to the feedback model to generate multiple personalized feedback information corresponding to the multiple evoked responses, wherein each of the multiple personalized feedback information includes multiple personalized feedback scores corresponding to the multiple personalized categories; and training the cue generator based on the multiple catalytic cues and the multiple personalized feedback information.

[0093] In some implementations, the feedback model is trained based on the following steps: training the feedback model based on multiple training data corresponding to the multiple personality categories.

[0094] In some embodiments, the prompt generation method 800 further includes the following steps: transmitting the plurality of training data corresponding to the plurality of personality categories to an augmented large language model to generate a plurality of augmented training data corresponding to the plurality of personality categories; and training the feedback model based on the plurality of augmented training data corresponding to the plurality of personality categories.

[0095] In some embodiments, the prompt generation method 800 further includes the step of: transmitting the optimal prompt to the large language model to generate a response message corresponding to the target personality category.

[0096] In some embodiments, the prompt generation method 800 further includes the following steps: controlling a human-machine interface to display the response message; and receiving an acknowledgment signal from the human-machine interface and updating the prompt generator based on the acknowledgment signal.

[0097] In some implementations, the large language model is trained based on the following steps: training the large language model based on multiple augmentation instructions and multiple test labels corresponding to the multiple augmentation instructions.

[0098] In some implementations, the plurality of enhancement instructions are generated based on the following steps: transmitting a plurality of test instructions and an evolutionary enhancement method to the large language model to generate the plurality of enhancement instructions, wherein the evolutionary enhancement method is generated based on an initial enhancement method.

[0099] In some embodiments, the evolutionary enhancement method is generated based on the following steps: transmitting a plurality of training instructions and the initial enhancement method to the large language model to generate a plurality of enhancement training instructions, wherein the initial enhancement method is used to indicate the enhancement type corresponding to the plurality of training instructions; transmitting the plurality of enhancement training instructions to the large language model to generate a plurality of enhancement training responses; and transmitting enhancement prompts and a plurality of enhancement feedback information to the large language model to generate the evolutionary enhancement method based on the initial enhancement method.

[0100] In some implementations, the plurality of augmentation feedback information is generated based on the following steps: analyzing the plurality of augmentation training instructions based on the large language model, and comparing the plurality of augmentation training responses with the plurality of training labels corresponding to the plurality of training instructions to generate the plurality of augmentation feedback information.

[0101] In addition to the steps described above, the second embodiment can also perform all the operations and steps of the prompt generation device 1 described in the first embodiment, have the same function, and achieve the same technical effect. Those skilled in the art to which this invention pertains can directly understand how the second embodiment performs these operations and steps based on the first embodiment described above, has the same function, and achieves the same technical effect, so it will not be described in detail here.

[0102] In summary, the prompt generation technology (including at least an apparatus and method) provided by this invention converts a situational question into multiple candidate prompts through a prompt generator, and determines the optimal prompt from among these candidate prompts. Next, this invention trains the prompt generator using reinforcement learning, based on a large language model and a feedback model. Furthermore, this invention transmits the optimal prompt to the large language model to generate a response message corresponding to the target personality category. Since this invention can automatically generate prompts through a prompt generator trained using reinforcement learning, the prompt generation technology provided by this invention reduces the time cost of designing prompts and guides the large language model to effectively generate personalized responses.

[0103] The above embodiments are merely illustrative of some implementations of the present invention and to explain the technical features of the present invention, and are not intended to limit the scope and range of protection of the present invention. Any changes or equivalent arrangements that can be easily made by those skilled in the art to which this invention pertains are within the scope of the present invention, and the scope of protection of the present invention is determined by the claims.

[0104] [Symbol Explanation]

[0105] 1: Prompt Generating Device

[0106] 11: Processor

[0107] 12: Memory

[0108] 13: Send and receive interface

[0109] PG: Prompt Generator

[0110] LLM: Large Language Model

[0111] RM: Feedback Model

[0112] 300: Prompt Generation Flowchart

[0113] PC: Personality Category

[0114] TPC: Target Personality Category

[0115] SQ: Situational Questioning

[0116] CP: Candidate Hint

[0117] RS: Reward Signal

[0118] BP: Best Tip

[0119] 400: Schematic diagram of the prompt generator training process

[0120] CTP: Catalysis Indication

[0121] ER: Evoked Response

[0122] PRI: Personalized Feedback Information

[0123] 500: Schematic diagram of the feedback model training process

[0124] TD: Training data

[0125] ALLM: Large Language Model for Augmentation

[0126] ATD: Augmenting Training Data

[0127] 601: Schematic diagram of the reinforcement training instruction generation process

[0128] TRI: Training Instructions

[0129] IEM: Initial Enhancement Method

[0130] ETRI: Enhanced Training Instructions

[0131] 603: Schematic diagram of the reinforcement training response generation process

[0132] ETRR: Enhanced Training Response

[0133] 605: Schematic diagram of the evolutionary enhancement method generation process

[0134] EP: Enhanced Tips

[0135] EFI: Enhanced Feedback Information

[0136] EEM: Evolutionary Enhancement Methods

[0137] 700: Schematic diagram of enhanced instruction generation process

[0138] TEI: Test Command

[0139] EI: Enhanced Instructions

[0140] 800: Prompt generation method

[0141] S801~S803: Steps.

Claims

1. A prompt generating device, characterized in that, Include: The send / receive interface is used to receive contextual questions; Memory, used to store the cue generator, large language model, and feedback model; and The processor, electrically connected to the transceiver interface and the memory, performs the following operations: The target personality category from multiple personality categories and the contextual question are transmitted to the prompt generator to generate multiple candidate prompts corresponding to the target personality category and multiple reward signals corresponding to the multiple candidate prompts, wherein the prompt generator is trained based on the large language model and the feedback model corresponding to the multiple personality categories; as well as Based on the multiple reward signals corresponding to the multiple candidate prompts, the best prompt is determined from the multiple candidate prompts.

2. The prompt generating device according to claim 1, characterized in that, The prompt generator mentioned above is trained based on the following operation: Multiple catalytic cues are transmitted to the large language model to generate multiple evoked responses corresponding to the multiple catalytic cues, wherein the multiple catalytic cues are generated based on the cue generator; The plurality of evoked responses are transmitted to the feedback model to generate a plurality of personalized feedback information corresponding to the plurality of evoked responses, wherein each of the plurality of personalized feedback information includes a plurality of personalized feedback scores corresponding to the plurality of personalized categories; as well as The prompt generator is trained based on the multiple catalytic prompts and the multiple personalized feedback information.

3. The prompt generating device according to claim 2, characterized in that, The feedback model mentioned above is trained based on the following operation: The feedback model is trained based on multiple training data corresponding to the multiple personality categories.

4. The prompt generating device according to claim 3, characterized in that, The processor also performs the following operations: The training data corresponding to the multiple personality categories are transmitted to a large language model for augmentation to generate multiple augmented training data corresponding to the multiple personality categories; as well as The feedback model is trained based on the multiple augmented training data corresponding to the multiple personality categories.

5. The prompt generating device according to claim 1, characterized in that, The processor is also used to perform the following operations: The optimal suggestion is transmitted to the large language model to generate a response message corresponding to the target personality category.

6. The prompt generating device according to claim 5, characterized in that, The processor also performs the following operations: The human-machine interface is controlled to display the response message; and The system receives a confirmation signal from the human-machine interface and updates the prompt generator based on the confirmation signal.

7. The prompt generating device according to claim 1, characterized in that, The large language model mentioned above is trained based on the following operation: The large language model is trained based on multiple enhancement instructions and multiple test labels corresponding to the multiple enhancement instructions.

8. The prompt generating device according to claim 7, characterized in that, The aforementioned multiple enhancement instructions are generated based on the following operation: Multiple test instructions and evolutionary enhancement methods are transmitted to the large language model to generate the multiple enhancement instructions, wherein the evolutionary enhancement methods are generated based on the initial enhancement method.

9. The prompt generating device according to claim 8, characterized in that, The aforementioned evolutionary enhancement method is based on the following operation: Multiple training instructions and the initial enhancement method are transmitted to the large language model to generate multiple enhancement training instructions, wherein the initial enhancement method is used to indicate the enhancement type corresponding to the multiple training instructions; The multiple augmentation training instructions are transmitted to the large language model to generate multiple augmentation training responses; as well as Enhancement prompts and multiple enhancement feedback messages are transmitted to the large language model to generate the evolutionary enhancement method based on the initial enhancement method.

10. The prompt generating device according to claim 9, characterized in that, The aforementioned multiple enhanced feedback messages are generated based on the following operation: Based on the large language model, the multiple augmentation training instructions are analyzed, and the multiple augmentation training responses and the multiple training labels corresponding to the multiple training instructions are compared to generate the multiple augmentation feedback information.

11. A method for generating a prompt, characterized in that, For use in an electronic device, wherein the electronic device receives contextual questions, the electronic device stores a prompt generator, a large language model, and a feedback model, wherein the prompt generation method includes the following steps: The target personality category from multiple personality categories and the contextual question are transmitted to the prompt generator to generate multiple candidate prompts corresponding to the target personality category and multiple reward signals corresponding to the multiple candidate prompts, wherein the prompt generator is trained based on the large language model and the feedback model corresponding to the multiple personality categories; as well as Based on the multiple reward signals corresponding to the multiple candidate prompts, the best prompt is determined from the multiple candidate prompts.

12. The prompt generation method according to claim 11, characterized in that, The prompt generator is trained based on the following steps: Multiple catalytic cues are transmitted to the large language model to generate multiple evoked responses corresponding to the multiple catalytic cues, wherein the multiple catalytic cues are generated based on the cue generator; The plurality of evoked responses are transmitted to the feedback model to generate a plurality of personalized feedback information corresponding to the plurality of evoked responses, wherein each of the plurality of personalized feedback information includes a plurality of personalized feedback scores corresponding to the plurality of personalized categories; as well as The prompt generator is trained based on the multiple catalytic prompts and the multiple personalized feedback information.

13. The prompt generation method according to claim 12, characterized in that, The feedback model mentioned above is trained based on the following steps: The feedback model is trained based on multiple training data corresponding to the multiple personality categories.

14. The prompt generation method according to claim 13, characterized in that, It also includes the following steps: The training data corresponding to the multiple personality categories are transmitted to a large language model for augmentation to generate multiple augmented training data corresponding to the multiple personality categories; as well as The feedback model is trained based on the multiple augmented training data corresponding to the multiple personality categories.

15. The prompt generation method according to claim 11, characterized in that, It also includes the following steps: The optimal suggestion is transmitted to the large language model to generate a response message corresponding to the target personality category.

16. The prompt generation method according to claim 15, characterized in that, It also includes the following steps: The human-machine interface is controlled to display the response message; and The system receives a confirmation signal from the human-machine interface and updates the prompt generator based on the confirmation signal.

17. The prompt generation method according to claim 11, characterized in that, The large language model mentioned above is trained based on the following steps: The large language model is trained based on multiple enhancement instructions and multiple test labels corresponding to the multiple enhancement instructions.

18. The prompt generation method according to claim 17, characterized in that, The aforementioned multiple enhancement instructions are generated based on the following steps: Multiple test instructions and evolutionary enhancement methods are transmitted to the large language model to generate the multiple enhancement instructions, wherein the evolutionary enhancement methods are generated based on the initial enhancement method.

19. The prompt generation method according to claim 18, characterized in that, The evolutionary enhancement method described above is based on the following steps: Multiple training instructions and the initial enhancement method are transmitted to the large language model to generate multiple enhancement training instructions, wherein the initial enhancement method is used to indicate the enhancement type corresponding to the multiple training instructions; The multiple augmentation training instructions are transmitted to the large language model to generate multiple augmentation training responses; as well as Enhancement prompts and multiple enhancement feedback messages are transmitted to the large language model to generate the evolutionary enhancement method based on the initial enhancement method.

20. The prompt generation method according to claim 19, characterized in that, The aforementioned multiple enhanced feedback messages are generated based on the following steps: Based on the large language model, the multiple augmentation training instructions are analyzed, and the multiple augmentation training responses and the multiple training labels corresponding to the multiple training instructions are compared to generate the multiple augmentation feedback information.