Prompt word template updating method and device of large language model and storage medium

By combining the evaluation content of the test sample set with the actual prediction results, supplementary prompt information is generated from test samples of different accuracy categories, which solves the problem of inaccurate prompt word templates in large language models and improves the output performance of the model.

CN121457598APending Publication Date: 2026-02-03ALIBABA CLOUD COMPUTING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411061094.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-02
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

The prompt word templates of existing large language models are inaccurate in practical applications, resulting in outputs that do not meet user expectations. Furthermore, the manual update method is time-consuming and inflexible.

Method used

The first prompt word is generated by using an initial prompt word template combined with the evaluation content in the test sample set. The accuracy category is determined based on the accuracy of the actual prediction results. Test samples of different accuracy categories are selected, supplementary prompt information is generated, and the prompt word template is updated.

Benefits of technology

It enables accurate, automatic, and efficient updating of prompt word templates for large language models, thereby improving the model's output performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121457598A_ABST
    Figure CN121457598A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a cue word template updating method and device for a large language model and a storage medium, and the method comprises the steps: combining an initial cue word template with the evaluation content in each test sample in a test sample set to form a first cue word, and inputting the first cue word into a first large language model, enabling the first large language model to output actual prediction results corresponding to the test examples respectively; according to the accuracy of the actual prediction result, determining an accuracy category corresponding to each test sample in the test sample set; screening out a plurality of test samples corresponding to different accuracy categories from the test sample set; determining output reason information of the actual prediction results corresponding to the plurality of test examples; generating supplementary prompt information of the initial prompt word template according to the output reason information; and updating the initial cue word template according to the supplementary cue information to obtain an updated cue word template. According to the scheme, accurate supplementary prompt information is obtained by analyzing a plurality of test samples covering different accuracy categories.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, device, and storage medium for updating prompt word templates for a large language model. Background Technology

[0002] Large Language Models (LLMs) are large-parameter language models learned using deep learning frameworks and large-scale corpora. They can be used to handle various natural language tasks such as text classification, question answering, and dialogue. In practice, LLMs understand user intent through pre-set prompts and output predictions. However, in real-world applications, the pre-set prompt templates can be inaccurate, causing the LLM's output to deviate from user expectations. Therefore, it is usually necessary to optimize and update the pre-set prompt templates to improve the LLM's performance in applications. Summary of the Invention

[0003] This invention provides a method, device, and storage medium for updating prompt word templates of a large language model, so as to accurately update the prompt word templates of the large language model.

[0004] In a first aspect, embodiments of the present invention provide a method for updating prompt word templates in a large language model, the method comprising:

[0005] Using an initial prompt word template, combined with the evaluation content of each test sample in the test sample set, a first prompt word is input into the first large language model, so that the first large language model outputs the actual prediction results corresponding to each test sample.

[0006] Based on the accuracy of the actual prediction results corresponding to each test sample in the test sample set, determine the accuracy category corresponding to each test sample in the test sample set.

[0007] Multiple test samples are selected from the test sample set, and the multiple test samples correspond to different accuracy categories;

[0008] Determine the output reason information for the actual prediction results corresponding to the multiple test samples;

[0009] Based on the output reason information, supplementary prompt information corresponding to the initial prompt word template is generated;

[0010] The initial prompt word template is updated based on the supplementary prompt information to obtain the updated prompt word template.

[0011] Secondly, embodiments of the present invention provide a prompt word template updating device for a large language model, the device comprising:

[0012] The prediction module is used to combine the initial prompt word template with the evaluation content in each test sample in the test sample set to form a first prompt word, which is then input into the first large language model so that the first large language model outputs the actual prediction results corresponding to each test sample.

[0013] The processing module is configured to determine the accuracy category of each test sample in the test sample set based on the accuracy of the actual prediction results corresponding to each test sample in the test sample set; select multiple test samples from the test sample set, the multiple test samples corresponding to different accuracy categories; determine the output reason information of the actual prediction results corresponding to each of the multiple test samples; and generate supplementary prompt information corresponding to the initial prompt word template based on the output reason information.

[0014] The update module is used to update the initial prompt word template according to the supplementary prompt information to obtain the updated prompt word template.

[0015] Thirdly, embodiments of the present invention provide an electronic device, including: a memory, a processor, and a communication interface; wherein, the memory stores executable code, and when the executable code is executed by the processor, the processor can at least implement the prompt word template update method of the large language model as described in the first aspect.

[0016] Fourthly, embodiments of the present invention provide a non-transitory machine-readable storage medium storing executable code, wherein when the executable code is executed by a processor of an electronic device, the processor is able to at least implement the prompt word template update method for a large language model as described in the first aspect.

[0017] Fifthly, embodiments of the present invention provide a computer program product, comprising: a computer program that, when executed by a processor of an electronic device, enables the processor to at least implement the prompt word template update method for a large language model as described in the first aspect.

[0018] In the solution provided by this embodiment of the invention, the first language model is pre-configured with an initial prompt word template. This initial prompt word template is used to generate prompt words so that the first language model can understand the user's intent based on the prompt words and complete the prediction task. In practical applications, optimizing and updating the initial prompt word template is beneficial for generating more accurate prompt words, resulting in better output performance of the first language model in applications. In this scheme, during the update of the initial prompt word template of the first language model, firstly, the initial prompt word template is combined with the evaluation content of each test sample in the test sample set to form the first prompt word, which is then input into the first language model so that the first language model outputs the actual prediction results corresponding to each test sample. Then, based on the accuracy of the actual prediction results corresponding to each test sample in the test sample set, the accuracy category corresponding to each test sample in the test sample set is determined, and multiple test samples corresponding to different accuracy categories are selected from the test sample set. Next, the output reason information of the actual prediction results corresponding to the multiple test samples is determined, and supplementary prompt information corresponding to the initial prompt word template is generated based on the output reason information. Finally, the initial prompt word template is updated based on the supplementary prompt information to obtain the updated prompt word template, so that the first language model can subsequently generate new prompt words based on the updated prompt word template.

[0019] In this embodiment of the invention, the test samples in the test sample set are first divided into different accuracy categories according to the actual prediction results of the first language model. Then, multiple test samples corresponding to different accuracy categories are selected from the test sample set, so that the output reason information corresponding to multiple test samples covers the reasons for the actual prediction results of the first language model for various accuracy categories. As a result, the supplementary prompt information generated based on the output reason information can more comprehensively and accurately cover various features, and achieve accurate updating of the initial prompt word template. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 A schematic diagram illustrating a scenario for updating prompt word templates in a large language model, as provided in an embodiment of the present invention.

[0022] Figure 2 A flowchart illustrating a method for updating prompt word templates in a large language model, as provided in an embodiment of the present invention;

[0023] Figure 3 A flowchart of a test sample screening method provided in an embodiment of the present invention;

[0024] Figure 4 A flowchart illustrating a method for generating supplementary prompt information provided in an embodiment of the present invention;

[0025] Figure 5 A flowchart illustrating another method for updating prompt word templates in a large language model, as provided in an embodiment of the present invention;

[0026] Figure 6 A schematic diagram of a prompt word template provided in an embodiment of the present invention;

[0027] Figure 7 A flowchart illustrating another method for updating prompt word templates in a large language model, as provided in an embodiment of the present invention;

[0028] Figure 8 This is a schematic diagram of the structure of a prompt word template updating device for a large language model provided in an embodiment of the present invention;

[0029] Figure 9 To and Figure 8 The illustrated embodiment provides a schematic diagram of the electronic device corresponding to the prompt word template updating device for the large language model. Detailed Implementation

[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0031] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in the embodiments of the present invention are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0032] Furthermore, the timing of the steps in the following method embodiments is merely an example and not a strict limitation.

[0033] To facilitate understanding, the relevant concepts involved in the embodiments of the present invention will be explained first.

[0034] Large Language Models (LLMs) are large-scale parametric language models learned using deep learning frameworks and large-scale corpora. They can be used to handle various natural language tasks such as classification, extraction, summarization, generation, question answering, and dialogue. Large language models are often also referred to as foundation models or large models.

[0035] The classification task emphasizes the identification and understanding of features and the accurate selection within a given set of categories. When performing classification tasks, large language models utilize existing knowledge bases and semantic understanding capabilities to determine the most suitable category for the text.

[0036] Extraction tasks emphasize identifying and extracting key information elements from text, such as entities, relationships, and events. Generally, extraction tasks have structured outputs, such as returning specific data fragments from text.

[0037] Summarizing tasks involve extracting the main content of a long text into a shorter summary, retaining only the core information and key points. This allows users to quickly grasp the main idea of ​​the long text without having to read the entire document. In summarizing tasks, large language models typically generate concise summaries by understanding the theme, key points, and structure of the long text.

[0038] The generation task emphasizes that the large language model generates new text content based on the input text content, and the generated text content is grammatically accurate, logically coherent with the input text, and topically relevant.

[0039] A prompt template is a structured text used to provide a framework for creating specific prompts, enabling efficient prompt creation. Typically, a prompt template includes: a task description of the target task, prior knowledge of the target task, the type or format of the output content, placeholders or variables, etc. The placeholders or variables can be replaced with actual text or data in specific applications. For example, a prompt template might be "Please explain the basic concept of [placeholder]". When it's necessary to explain a certain noun i, the placeholder can be replaced with the noun i to generate the corresponding prompt.

[0040] A prompt is a piece of text or instruction given to a large language model. It is generated based on a prompt template and contains the actual text or data used in the specific application to guide the large language model in generating the corresponding response.

[0041] It should be noted that the embodiments of the present invention involve multiple large language models for performing different tasks, and each of these large language models has a prompt word template that matches the task it performs. For ease of description, the various language models are distinguished by terms such as "first" and "second".

[0042] During use, large language models understand user intent through prompts and output prediction results. In practical applications, prompts are generated based on prompt templates. If the prompt templates are inaccurate, the generated prompts will also be inaccurate, leading to problems where the output of the large language model does not meet user expectations. For example, in classification tasks, the classification result output by the large language model may not match the actual category of the classified object.

[0043] Generally, professionals will rewrite the prompt word template of the large language model based on the erroneous output results, thus optimizing and updating the prompt word template. By updating the prompt word template, the prompt words can be updated, thereby enabling the large language model to have better output performance in applications.

[0044] However, due to the diverse application areas of large language models and the complexity of writing prompt word templates, manually writing and updating prompt word templates based on the erroneous outputs generated by the large language model in each application area is both time-consuming and inflexible. Furthermore, relying solely on erroneous outputs to update prompt word templates can lead to a situation where the output performance of the large language model worsens after the update.

[0045] To address at least one of the aforementioned technical problems, embodiments of the present invention provide a method for updating prompt word templates for a large language model. Figure 1 A scenario illustration of a prompt word template update method for a large language model provided for the implementation of this invention, as shown in the figure. Figure 1 As shown, an initial prompt word template is first used, combined with the evaluation content from each test sample in the test sample set, to form the first prompt word, which is then input into the first large language model. This causes the first large language model to output the actual prediction results corresponding to each test sample. Next, based on the accuracy of the actual prediction results corresponding to each test sample, the accuracy category corresponding to each test sample is determined, and multiple test samples corresponding to different accuracy categories are selected from the test sample set. Then, based on the output reason information of the actual prediction results corresponding to multiple test samples, supplementary prompt information corresponding to the initial prompt word template is generated. Finally, the initial prompt word template is updated using the supplementary prompt information. This method determines the supplementary prompt information used to update the prompt word template by comprehensively considering the reasons for various types of actual prediction results output by the large language model, and then uses the supplementary prompt information to accurately, automatically, and efficiently update the prompt word template of the large language model. The following is a detailed explanation.

[0046] The method for updating prompt word templates for a large language model provided in this invention can be executed by an electronic device, such as a PC, laptop, or smartphone, or a server. The server can be a physical server containing an independent host, a virtual server, a cloud server, or a server cluster.

[0047] Figure 2 A flowchart illustrating a method for updating prompt word templates in a large language model, as provided in an embodiment of the present invention, is shown below. Figure 2 As shown, it may include the following steps:

[0048] 201. Using the initial prompt word template, combined with the evaluation content in each test sample in the test sample set, the first prompt word is input into the first large language model so that the first large language model outputs the actual prediction results corresponding to each test sample.

[0049] 202. Based on the accuracy of the actual prediction results corresponding to each test sample in the test sample set, determine the accuracy category corresponding to each test sample in the test sample set.

[0050] 203. Select multiple test samples from the test sample set, with each test sample corresponding to a different accuracy category.

[0051] 204. Determine the output reason information for the actual prediction results corresponding to multiple test cases.

[0052] 205. Based on the output reason information, generate supplementary prompt information corresponding to the initial prompt word template.

[0053] 206. Update the initial prompt word template based on the supplementary prompt information to obtain the updated prompt word template.

[0054] In this embodiment, the first large language model is the large language model that needs to have its prompt word template updated. The initial prompt word template is the prompt word template to be updated corresponding to the first large language model. It can be understood that updating the prompt word template of the first large language model can be updating the initialized prompt word template (i.e., the manually written original prompt word template), or it can be updating the updated prompt word template again. Therefore, the initial prompt word template can be understood as the prompt word template used by the first large language model when performing the task, which can be either the initialized prompt word template or the updated prompt word template.

[0055] The test sample set, also known as the test dataset, contains several test samples, each with evaluation content. The evaluation content from each test sample in the test sample set is combined with the initial prompt word template to generate the first prompt word input to the first language model. Then, the first language model can output the actual prediction results corresponding to each test sample in the test sample set based on the first prompt word, so that supplementary prompt information for updating the initial prompt word template can be determined based on the actual prediction results.

[0056] The process of outputting actual prediction results based on the test sample set, which is also the process of using the first language model, involves replacing the placeholders in the initial prompt word template used to fill the actual test samples with the evaluation content from each test sample in the test sample set. This creates a first prompt word corresponding to each test sample. Then, the first prompt words corresponding to each test sample are input into the first language model, enabling it to understand the prediction task based on the prompt words and output the actual prediction results corresponding to each test sample. The actual prediction results are the inference results of the first language model.

[0057] It is understandable that the actual prediction results corresponding to each test sample output by the first language model may be correct or incorrect. For example, the actual prediction result corresponding to test sample 1 may be correct, while the actual prediction result corresponding to test sample 2 may be incorrect. In this embodiment, in order to improve the accuracy of the initial prompt word template update, we no longer focus solely on analyzing the reasons for the incorrect prediction results output by the first language model, but comprehensively consider the reasons for each actual prediction result output by the first language model.

[0058] To facilitate subsequent cause analysis and processing, after obtaining the actual prediction results corresponding to each test sample, the test samples in the test sample set are first classified based on the actual prediction results corresponding to each test sample, that is, the accuracy category corresponding to each test sample in the test sample set is determined.

[0059] Optionally, the accuracy category corresponding to a test sample can include positive examples, negative examples, and intermediate examples. The classification of positive examples, intermediate examples, and negative examples is related to the correctness (i.e., accuracy) of the actual prediction result. Following the order of positive examples, intermediate examples, and negative examples, the correctness of the actual prediction result corresponding to the test sample gradually decreases. For example, when the actual prediction result of a test sample is completely correct, the accuracy category of that test sample can be determined as a positive example; when the actual prediction result of a test sample is partially correct, the accuracy category of that test sample can be determined as an intermediate example; and when the actual prediction result of a test sample is completely wrong, the accuracy category of that test sample can be determined as a negative example.

[0060] In practical applications, the accuracy of actual prediction results can be determined manually or automatically. The choice of method depends on whether the test samples in the test case set are pre-configured with reference prediction results for judging the accuracy of the actual prediction results; these reference prediction results are preset correct results. In practice, whether the test samples in the test case set contain only the evaluation content or both the evaluation content and the corresponding reference prediction results can be customized.

[0061] In an optional embodiment, if each test sample in the test sample set does not contain a reference prediction result corresponding to the evaluation content, the correctness of the actual prediction result can be determined by human assistance. Optionally, the accuracy category corresponding to each test sample can be determined based on the accuracy category labeling result fed back by the user based on the actual prediction result corresponding to each test sample. The accuracy category labeling result fed back by the user can be whether the actual prediction result meets expectations, that is, the degree of conformity between the actual prediction result and the correct result. If the user determines that the actual prediction result corresponding to the test sample completely meets expectations, the accuracy category of the test sample is determined to be a positive example; if the user determines that the actual prediction result corresponding to the test sample generally meets expectations, for example, only some information in the actual prediction result is correct, the accuracy category of the test sample is determined to be an intermediate example; if the user determines that the actual prediction result corresponding to the test sample does not meet expectations, the accuracy category of the test sample is determined to be a negative example.

[0062] In another optional embodiment, if each test sample in the test sample set contains a reference prediction result corresponding to the evaluation content, the correctness of the actual prediction result can be determined by human assistance, or automatically. The automated determination process can be implemented by: determining the accuracy category corresponding to each test sample based on the matching degree between the actual prediction result corresponding to each test sample and the reference prediction result corresponding to each test sample.

[0063] Optionally, the aforementioned automated judgment process can be implemented using machine learning. Specifically, a large language model (referred to as the second large language model) can be pre-trained for classifying test samples based on the actual prediction results and reference prediction results of the test samples, and corresponding prompt word templates can be configured for the second large language model. The prompt word templates corresponding to the second large language model can include a task description, output results, and content to be analyzed.

[0064] For example, the prompt word template for the second largest language model could be:

[0065] ##Task Description: Please compare and analyze the [Correct Result] and the [Output Result] to determine whether the two results are the same. The criterion for judgment is the proportion of the content that is the same in the [Correct Result] and the [Output Result] to the total content of the [Correct Result].

[0066] ## Output results: Output scores from 1 to 10, where 10 is completely identical and 1 is completely different.

[0067] ## Content to be analyzed: [Correct Result] Placeholder 1, [Output Result] Placeholder 2.

[0068] Placeholder 1 is used to be replaced by the reference prediction result, and placeholder 2 is used to be replaced by the actual prediction result.

[0069] When using the second language model, for any test sample in the test sample set, a second prompt word is first generated based on the prompt word template, the actual prediction result of any test sample, and the reference prediction result. This second prompt word contains the actual prediction result and the reference prediction result corresponding to any test sample. Then, the second prompt word is input into the second language model so that the second language model can determine the matching degree between the actual prediction result and the reference prediction result corresponding to any test sample. Finally, based on the matching degree between the actual prediction result and the reference prediction result corresponding to any test sample, the accuracy category corresponding to any test sample is determined.

[0070] Optionally, matching thresholds for positive examples, negative examples, and intermediate examples can be predefined. For example, it can be defined that the matching degree of test samples belonging to positive examples is greater than a first threshold, the matching degree of test samples belonging to intermediate examples is between the first threshold and a second threshold, the matching degree of test samples belonging to negative examples is less than the second threshold, and the first threshold is greater than the second threshold.

[0071] In practical applications, if the matching degree is represented by a score, the matching degree score corresponding to the first threshold and the matching degree score corresponding to the second threshold can be preset. Therefore, after determining the matching degree score between the actual prediction result and the reference prediction result corresponding to the test sample, the accuracy category of the test sample can be determined based on the relationship between this matching degree score and the first and second thresholds. For example, using scores from 1 to 10 to represent the matching degree between the actual prediction result and the reference prediction result corresponding to the test sample, with the first threshold preset to 9 points and the second threshold to 5 points, if the matching degree score of a test sample is 9.1 points, the accuracy category of the test sample is determined to be a positive example; if the matching degree score of a test sample is 7 points, the accuracy category of the test sample is determined to be an intermediate example; and if the matching degree score of a test sample is 4 points, the accuracy category of the test sample is determined to be a negative example.

[0072] After classifying the test samples in the test sample set, we can further analyze the reasons for the actual prediction results of each type of test sample output by the first language model. However, in practical applications, on the one hand, the number of test samples in the test sample set is often large. If all test samples are analyzed, the amount of data to be processed is large and the analysis efficiency is low. On the other hand, for the test sample set, the number of test samples contained in each accuracy category often differs. For example, the number of test samples contained in positive examples and the number of test samples contained in negative examples correspond to different orders of magnitude. In this case, even if we analyze the reasons for the actual prediction results of each type of test sample output by the first language model, the large differences in the number of test samples and the uneven data distribution may lead to errors in the analysis results, thus affecting the accuracy of the initial prompt word template update.

[0073] Therefore, in this embodiment, after classifying the test samples in the test sample set, sample filtering is performed on the test samples in the test sample set. Specifically, multiple test samples corresponding to different accuracy categories can be filtered from the test sample set, and the number of test samples corresponding to each accuracy category in the multiple test samples is the same or similar. Through sample filtering, the amount of data to be processed can be reduced, processing efficiency can be improved, and the number of test samples for each accuracy category being analyzed can be the same or similar.

[0074] Next, the output reason information for the actual prediction results corresponding to multiple test samples of different accuracy categories is determined. The output reason information for each test sample may include: where the model's reasoning process was correct or incorrect, the reason for the correctness or error, and reasoning suggestions summarized based on the current test sample.

[0075] As an optional method to obtain output cause information, machine learning can be used to determine the output cause information of the actual prediction results corresponding to multiple test samples of different accuracy categories using a third language model. The third language model is a pre-trained large language model configured with corresponding prompt word templates. These prompt word templates can include the task description of the cause extraction task, the output results, and the content to be analyzed.

[0076] Optionally, the prompt word templates for the third language model can be configured separately based on the accuracy category of the test samples and / or the task category of the prediction task performed by the first language model. The task category includes, but is not limited to, classification tasks, extraction tasks, etc.

[0077] For example, when the prediction task performed by the first language model is a classification task, and the accuracy category of the test samples is positive, the prompt word template corresponding to the third language model can be:

[0078] ##Task Description: 1) Analyze the correct reasoning process and / or correct result in the positive example, and the reasons for the correctness, and output them as "key elements"; 2) Summarize the points that need attention to answer similar questions correctly, and summarize "improvement suggestions" based on the example.

[0079] ## Output Results: "Key Elements": Analyze the reasons for the correct result, and summarize which important content or elements led to the correct result, in no more than 50 words; "Improvement Suggestions": Propose specific suggestions and points of focus, which can be specific strategies, methods, or analytical perspectives, in no more than 50 words.

[0080] ## Content to be analyzed: Test sample placeholders, actual prediction result placeholders corresponding to the test samples, and reference prediction result placeholders.

[0081] Among them, the test sample placeholder is used to replace the evaluation content of the test sample that is a positive example in the classification task, and the actual prediction result placeholder and the reference prediction result placeholder are used to replace the actual prediction result and the reference prediction result corresponding to the test sample that is a positive example in the classification task.

[0082] For example, when the prediction task performed by the first language model is a classification task, and the accuracy category of the test samples is negative, the prompt word template corresponding to the third language model can be:

[0083] ##Task Description: 1) Analyze the incorrect reasoning process and / or incorrect results in the negative examples, and the reasons for the errors, and use them as "key elements"; 2) Summarize the areas that need attention to answer the negative examples correctly, and summarize "improvement suggestions" based on the examples.

[0084] ## Output Results: "Key Elements": Analyze the cause of the error, and summarize which important content or elements were omitted / confused / not given sufficient attention, leading to the error result (no more than 50 words); "Improvement Suggestions": Propose specific improvement measures, suggestions, or points of focus to supplement the overlooked content. The points of focus can be specific strategies, methodologies, or analytical perspectives (no more than 50 words).

[0085] ## Content to be analyzed: Test sample placeholders, actual prediction result placeholders corresponding to the test samples, and reference prediction result placeholders.

[0086] Among them, the test sample placeholder is used to replace the evaluation content of the test sample that is a negative example in the classification task, and the actual prediction result placeholder and the reference prediction result placeholder are used to replace the actual prediction result and the reference prediction result corresponding to the test sample that is a negative example in the classification task.

[0087] For example, when the prediction task performed by the first language model is an extraction task, and the accuracy category of the test samples is positive, the prompt word template corresponding to the third language model can be:

[0088] ##Task Description: 1) Tags: Ignore empty strings if the tag value is empty. Find all tags whose tag value is not empty and output "tag"; 2) Notes Extraction: For the tags in step 1), independently identify the reasons that led to the correct extraction, list the specific successful practices one by one, and output them as "Notes"; 3) Root Cause Analysis: Explore the principles or core factors behind the success factors, extract how key information or methods lead to the correct conclusion, and output them as "Key Elements"; 4) Improvement Suggestions: Summarize the aspects that should be focused on to correctly answer similar questions or analyze similar content. This may include specific methods or analytical perspectives, and output them as "Improvement Suggestions".

[0089] ## Output results: Output the tags involved in the extraction task, as well as the corresponding "focus", "key elements" and "improvement suggestions" for each tag.

[0090] ## Content to be analyzed: Test sample placeholders, actual prediction result placeholders corresponding to the test samples, and reference prediction result placeholders.

[0091] Among them, the test sample placeholder is used to replace the evaluation content of the test samples that are positive examples in the extracted task, and the actual prediction result placeholder and the reference prediction result placeholder are used to replace the actual prediction result and the reference prediction result corresponding to the test samples that are positive examples in the extracted task.

[0092] The above provides an example illustrating the prompt word templates for the third major language model. In practical applications, prompt word templates can also be pre-built for intermediate examples in classification tasks, negative examples in extraction tasks, and intermediate examples in extraction tasks, etc.

[0093] In the specific application of the third language model, for a target test sample among multiple test samples of different accuracy categories, a prompt word template corresponding to the accuracy category of the target test sample can be selected first. The evaluation content of the target test sample is then combined with this prompt word template to generate a third prompt word corresponding to the target test sample. This third prompt word includes the evaluation content of the target test sample, the actual prediction result and reference prediction result corresponding to the target test sample, and the causal extraction task description information corresponding to the accuracy category of the target test sample. Then, the third prompt word is input into the third language model, causing the model to output the output causal information for the actual prediction result corresponding to the target test sample. The target test sample is any one of multiple test samples of different accuracy categories.

[0094] After obtaining the output reason information for the actual prediction results corresponding to multiple test cases, supplementary prompt information corresponding to the initial prompt word template is generated based on the output reason information. Optionally, the output reason information corresponding to multiple test cases can be summarized to generate supplementary prompt information.

[0095] Finally, the initial prompt template is updated based on the supplementary information to obtain the updated prompt template. In practical applications, the updated prompt template can be generated by adding the supplementary information as a note to the initial prompt template.

[0096] During the use of the first major model, new first prompt words can be generated based on the updated prompt word template. Compared to the first prompt words generated based on the initial prompt word template, the new first prompt words may contain more or more accurate information, which helps the first major language model to understand the prediction task more accurately, thus resulting in better model output performance.

[0097] In summary, in the process of updating the initial prompt word template of the first large language model, this embodiment uses the initial prompt word template and combines it with the evaluation content of each test sample in the test sample set to form the first prompt word, which is then input into the first large language model so that the first large language model outputs the actual prediction results corresponding to each test sample. Then, based on the accuracy of the actual prediction results corresponding to each test sample in the test sample set, the accuracy category corresponding to each test sample in the test sample set is determined. Then, multiple test samples corresponding to different accuracy categories are selected from the test sample set. This ensures that the multiple test samples used to determine the output cause information cover various accuracy categories, and also ensures that the number of test samples in different accuracy categories is the same or similar, avoiding errors caused by uneven numbers of test samples in different accuracy categories. Because multiple test samples corresponding to different accuracy categories are selected from the test sample set, the accuracy categories are comprehensively covered and the number of test samples corresponding to different accuracy categories is the same or similar. Finally, based on the output reason information of the actual prediction results corresponding to multiple test samples, the supplementary prompt information corresponding to the initial prompt word template has high accuracy. Thus, the initial prompt word template can be accurately updated based on the supplementary prompt information.

[0098] In the large language model prompt word update method provided in the foregoing embodiments, multiple test samples corresponding to different accuracy categories are selected from the test sample set based on the accuracy category corresponding to each test sample. This is explained from the dimension of the number of samples selected for each test sample; that is, the sample selection ensures that the number of test samples for each accuracy category is the same or similar among the determined multiple test samples. In practical applications, a corresponding number of test samples for each accuracy category can be sampled through methods such as random sampling.

[0099] However, it is understandable that even within the same accuracy category, different test samples will reflect different sample characteristics. If random sampling or other methods are used to filter test samples for each accuracy category, the filtered test samples may not comprehensively reflect the sample characteristics of all test samples in that accuracy category. Therefore, even if the number of test samples in different accuracy categories is the same or similar, the selected test samples may not fully reflect the sample characteristics of the corresponding accuracy category, leading to inaccurate supplementary prompt information and preventing accurate updates to the initial prompt word template.

[0100] To this end, the present invention provides a test sample screening method. By performing clustering processing on each test sample in the test sample set, it ensures that among the multiple screened test samples, the test samples corresponding to each accuracy category can accurately and comprehensively reflect the sample characteristics of the test samples included in that accuracy category.

[0101] Figure 3 A flowchart of a test sample screening method provided in an embodiment of the present invention is shown below. Figure 3 As shown, it may include the following steps:

[0102] 301. Perform clustering on each test sample in the test sample set to obtain multiple clusters.

[0103] 302. Sample multiple clusters to obtain multiple test cases.

[0104] Among them, clustering algorithms such as k-means clustering can be used to cluster the test samples in the test sample set based on the semantics of the test samples.

[0105] In practical applications, you can optionally cluster all test samples in the test sample set directly, or you can cluster test samples of different accuracy categories separately.

[0106] Since the accuracy category corresponding to each test sample in the test sample set has been predetermined in this embodiment, and the final screening result needs to ensure that the number of test samples selected in each accuracy category is the same or similar, clustering the test samples for different accuracy categories can more efficiently select multiple suitable test samples.

[0107] In detail, when clustering test samples for different accuracy categories is performed separately, the test sample selection process may include: First, determining multiple test sample groups based on the accuracy category corresponding to each test sample in the test sample set, wherein a test sample group consists of test samples corresponding to the same accuracy category; then, performing clustering on the multiple test sample groups to obtain multiple clusters corresponding to each test sample group; finally, sampling the multiple clusters corresponding to each of the multiple test sample groups to obtain multiple test samples.

[0108] To make it easier to understand, let's take an example. Suppose that the test sample set contains test sample 1, test sample 2, ..., test sample m, where the accuracy categories corresponding to test sample 1 to test sample n are positive examples, the accuracy categories corresponding to test sample n+1 to test sample g are intermediate examples, and the accuracy categories corresponding to test sample g+1 to test sample m are negative examples, where m, n, and g are positive integers, and m is greater than g and g is greater than n.

[0109] Based on this assumption, test samples 1 to n are designated as test sample group 1, test samples n+1 to g as test sample group 2, and test samples g+1 to m as test sample group 3. Then, for test sample group 1, based on the preset expected number of clusters N and the number of test samples M contained in each cluster, a clustering algorithm is used to cluster test samples 1 to n into N classes, resulting in N clusters, where each cluster contains less than or equal to M test samples. Next, samples are taken from each of the N clusters, for example, G samples from each cluster, to obtain the test sample selection results corresponding to test sample group 1, where M, N, and G are positive integers. Similarly, the test sample selection results corresponding to test sample groups 2 and 3 can be obtained respectively. Finally, the test sample selection results corresponding to test sample groups 1, 2, and 3 are merged to obtain multiple test samples used to determine the output cause information.

[0110] In this embodiment, multiple test sample groups are determined based on the accuracy category corresponding to each test sample in the test sample set. By performing clustering processing on multiple test sample groups separately and sampling the multiple clusters obtained from the clustering of each test sample group, it is possible to ensure that the number of test samples of different accuracy categories in the multiple sampled test samples is the same or similar, and that the sampling results corresponding to each accuracy category come from different clusters in the test samples of the same accuracy category, which can comprehensively reflect the sample characteristics of the corresponding accuracy category.

[0111] As mentioned in the preceding embodiments, after obtaining the output reason information of the actual prediction results corresponding to multiple test cases, supplementary prompt information corresponding to the initial prompt word template can be generated based on the output reason information. For example, the output reason information corresponding to multiple test cases can be directly summarized, and the summarized result can be used as supplementary prompt information.

[0112] Directly summarizing and outputting the cause information ensures that the supplementary prompt information is detailed and comprehensive. However, it results in duplicate content in the prompt word template updated based on the supplementary prompt information, and the prompt words generated based on the updated prompt word template are too long.

[0113] Therefore, this embodiment provides a method for generating supplementary prompt information. The supplementary prompt information is obtained by aggregating the output cause information corresponding to multiple test samples, making the obtained supplementary prompt information more concise.

[0114] Figure 4 A flowchart of a method for generating supplementary prompt information provided in an embodiment of the present invention is shown below. Figure 4 As shown, it may include the following steps:

[0115] 401. Aggregate the output reason information of the actual prediction results corresponding to multiple test cases.

[0116] 402. Based on the aggregation results, generate supplementary prompt information corresponding to the initial prompt word template.

[0117] In the specific implementation process, optionally, the output cause information of the actual prediction results corresponding to multiple test samples can be directly aggregated to obtain the aggregated result; alternatively, based on the accuracy categories corresponding to multiple test samples, the output cause information of the actual prediction results corresponding to multiple test samples can be aggregated separately according to the accuracy category to obtain the aggregated result corresponding to each accuracy category, and then the aggregated results corresponding to each accuracy category can be merged to obtain the final aggregated result. Among them, the scheme of directly performing aggregation processing has higher aggregation efficiency and can obtain the aggregated result faster because it does not require merging the aggregated results corresponding to each accuracy category.

[0118] As an optional method for generating supplementary prompts, machine learning can be used to generate supplementary prompts corresponding to the initial prompt word template using a fourth language model. This fourth language model is a pre-trained large language model capable of performing summarization and generation tasks, and is configured with corresponding prompt word templates. These prompt word templates can include the task description, output requirements, and content to be analyzed for the cause aggregation task.

[0119] For example, for a classification task, the prompt word template corresponding to the fourth language model could be:

[0120] ##Task Description: As an information aggregation and analysis expert, your task is to perform cluster analysis on a large number of collected optimization suggestions. Based on the provided "key elements" and corresponding "improvement suggestions," you need to categorize and organize this information, ensuring that each category is unique and the content is comprehensive and concise. When integrating information, be sure to maintain the integrity, simplicity, and refinement of each category, and then directly present the results of the clustering summary.

[0121] ## Output Requirements: Directly output the summary results after classification and induction, reflecting the differences and characteristics between different categories. Do not output other information.

[0122] ## Content to be analyzed: "Key elements" placeholder 3 and "Improvement suggestions" placeholder 4.

[0123] Placeholders 3 and 4 are used to replace the corresponding content in the output reason information.

[0124] For example, for extraction tasks, the prompt word template corresponding to the fourth language model can be:

[0125] ##Task Description: As an information aggregation and analysis expert, your task is to perform cluster analysis on a large number of collected optimization suggestions. Based on the provided "tags," "focuses," "key elements," and "improvement suggestions," you need to categorize and organize this information, ensuring that each category is unique and the content is comprehensive and concise. When integrating information, be sure to maintain the integrity, simplicity, and refinement of each category, and then directly present the results of the clustering summary.

[0126] ## Output Requirements: Directly output the summary results after classification and induction, reflecting the differences and characteristics between different categories. Do not output other information.

[0127] ## Content to be analyzed: Placeholder 1 for "Tags", Placeholder 2 for "Focus Points", Placeholder 3 for "Key Elements" and Placeholder 4 for "Improvement Suggestions".

[0128] Placeholders 1 to 4 are used to replace the corresponding content in the output reason information.

[0129] When using the fourth language model, you can first select the appropriate prompt word template according to the task type. Then, based on the prompt word template of the fourth language model and the output reason information of the actual prediction results corresponding to multiple test samples, generate the fourth prompt word corresponding to multiple test samples. The fourth prompt word includes the output reason information of the actual prediction results corresponding to multiple test samples and the task description information of the reason aggregation. After that, input the fourth prompt word into the fourth language model so that the fourth language model outputs the aggregated result of the output reason information of the actual prediction results corresponding to multiple test samples.

[0130] In this solution, by aggregating the output reason information, on the one hand, it helps to simplify the output reason information corresponding to multiple test cases, making the obtained supplementary prompt information more concise; on the other hand, based on the summarizing effect of aggregation processing, it also helps to improve the generalization of the obtained supplementary prompt information.

[0131] Figure 5 A flowchart of another method for updating prompt word templates for a large language model provided in an embodiment of the present invention is shown below. Figure 5 As shown, this may include the following steps:

[0132] 501. Identify the target information in the supplementary prompts that is different from the initial prompt template.

[0133] 502. Add the target information to the initial prompt word template to obtain the updated prompt word template.

[0134] After generating the supplementary prompt information corresponding to the initial prompt word template, you can directly replace the initial prompt word template with the supplementary prompt information to update the initial prompt word template; or, you can directly merge the supplementary prompt word information with the information in the initial prompt word template to update the initial prompt word template.

[0135] One update method, which directly replaces the initial prompt template with supplementary prompt information, discards information from the initial prompt template during the update. In application scenarios with multiple rounds of iterative optimization of the prompt template, this can lead to the loss of historical feedback and optimization information, resulting in lower accuracy of the updated prompt template, or higher accuracy only in specific application areas. Another update method, which directly merges the supplementary prompt information with the information in the initial prompt template, results in information redundancy in the updated prompt template.

[0136] Therefore, in this embodiment, when updating the initial prompt word template, firstly, the target information that is different from the initial prompt word template in the supplementary prompt information is determined; then, the target information is added to the initial prompt word template, for example, the target information is added as a note to the initial prompt word template to obtain the updated prompt word template.

[0137] For ease of understanding, combined with Figure 6 To illustrate, Figure 6 This is a schematic diagram of a prompt word template provided in an embodiment of the present invention, such as... Figure 6 As shown in the left image, the initial prompt word template includes:

[0138] ##Task Description: Based on the given trending event name and blog post content, determine whether the blog post content is related to the trending event. Regarding the definition of "related": both descriptions should be consistent in their main idea.

[0139] ## Output: Output: Relevant or irrelevant. Do not output anything else.

[0140] ## Content to be analyzed: < <content>>

[0141] Among them, < <content>The ">" symbol is a placeholder used to replace the first prompt word when it is generated by each test sample in the test sample set.

[0142] Assuming that the determined supplementary prompt information includes the target information "xxx" in addition to the information contained in the initial prompt template, then the initial prompt template is updated based on the target information "xxx", resulting in the updated prompt template, as follows: Figure 6 The right figure in the image includes:

[0143] ##Task Description: Based on the given trending event name and blog post content, determine whether the blog post content is related to the trending event. Regarding the definition of "related": both descriptions should be consistent in their main idea.

[0144] ## Output: Output: Relevant or irrelevant. Do not output anything else.

[0145] ## Content to be analyzed: < <content>>

[0146] ##Notes: xxx.

[0147] In this solution, by first identifying the target information in the supplementary prompt information that distinguishes it from the initial prompt template, and then adding the target information to the update method of the initial prompt template, we can ensure that the updated prompt template contains all the information of the initial prompt template, thus preserving the integrity and generalizability of historical feedback optimization information. On the other hand, we can also ensure that the content updated to the new prompt template does not overlap with the content in the initial prompt template, meaning that the updated prompt template is clear, concise, and free of redundant information.

[0148] Figure 7 A flowchart of another method for updating prompt word templates for a large language model provided in an embodiment of the present invention is shown below. Figure 7 As shown, this may include the following steps:

[0149] 701. Based on the test sample set, determine the test performance metrics for the initial prompt word template and the updated prompt word template.

[0150] 702. Determine the target prompt word template for use in the first language model based on the test performance indicators. The target prompt word template is the prompt word template with better test performance indicators between the initial prompt word template and the updated prompt word template.

[0151] Understandably, in practical applications, not every updated prompt word template will improve the performance of the primary language model; there are instances where the primary language model performs worse based on the updated prompt word template. Therefore, during the update process, it is essential to select the target prompt word template that is more beneficial to improving the output performance of the primary language model from both the initial and updated prompt word templates.

[0152] In the specific implementation process, firstly, a test performance metric is determined to evaluate the accuracy of the initial and updated prompt word templates. For example, the test performance metric could be the accuracy of the prediction results output by the first language model when using the corresponding prompt word template. Next, a first prompt word is generated based on the test sample set and the initial prompt word template, and input into the first language model. Based on the prediction results output by the first language model, the test performance metric 1 corresponding to the first language model when using the initial prompt word template is determined. Then, a first prompt word is generated based on the test sample set and the updated prompt word template, and input into the first language model. Based on the prediction results output by the first language model, the test performance metric 2 corresponding to the first language model when using the updated prompt word template is determined. If test performance metric 1 is better than test performance metric 2, the target prompt word template used by the first language model is determined to remain the initial prompt word template; if test performance metric 2 is better than test performance metric 1, the target prompt word template used by the first language model is determined to be the updated prompt word template. The test sample set can be the test sample set used in the process of updating the initial prompt word template, or it can be other test sample sets.

[0153] In this embodiment, by comparing the test performance indicators of the prompt word template before and after the update, and selecting the prompt word template with better test performance indicators as the target prompt word template used by the first language model, it is possible to obtain a prompt word template that meets the expected effect and has high accuracy more quickly.

[0154] The above provides a detailed description of the prompt word update method for the large language model provided in the embodiments of the present invention. The following examples illustrate the prompt word update method for the first large language model in specific application scenarios, using classification task scenarios and extraction task scenarios as examples.

[0155] 1. Update of prompt words for the first major language model in classification task scenarios.

[0156] Assuming a classification task scenario, the initial prompt word template for the first major language model is as follows:

[0157] ##Task Description: Based on the given trending event name and blog post content, determine whether the blog post content is related to the trending event. Regarding the definition of "related": both descriptions should be consistent in their main idea.

[0158] ## Output: Output: Relevant or irrelevant. Do not output anything else.

[0159] ## Content to be analyzed: < <content>>

[0160] The test case set includes:

[0161] Test Sample 1 - Evaluation Content: Hot Event Name: Three Company Executives Hold Press Conference to Apologize; Blog Content: "The culture of xx company is that bowing ≠ admitting mistakes and apologizing, bowing = compensation and resolving the matter", "Three Company Executives Hold Press Conference to Apologize".

[0162] Test Sample 2 – Evaluation Content: Hot Event Name: TV Series Z; Blog Post Content: "Three consecutive ancient costume female-oriented dramas on a certain video streaming platform have exceeded 10,000 views", "This incredible wealth has finally come to a certain video streaming platform", "Even with fierce competition, TV series X, TV series Y, and TV series Z still have exceeded 10,000 views. Thumbs up for the certain video streaming platform!"

[0163] It should be noted that in practical applications, a test case set may contain more than two test cases. In this embodiment, test case 1 and test case 2 are used as examples for illustration only, and the number of test cases in the test case set is not limited.

[0164] When updating the initial prompt word template of the first major language model, firstly, the placeholders < in the initial prompt word template are removed. <content>Replace the evaluation content of Test Sample 1 and Test Sample 2 respectively to generate the first prompt words corresponding to Test Sample 1 and Test Sample 2 respectively. Assume that the first language model outputs the actual prediction result of Test Sample 1 as "relevant" based on the first prompt word corresponding to Test Sample 1; the first language model outputs the actual prediction result of Test Sample 2 as "relevant" based on the first prompt word corresponding to Test Sample 2.

[0165] If the actual prediction result corresponding to test sample 1 is correct and the actual prediction result corresponding to test sample 2 is incorrect, then based on the actual prediction results corresponding to test sample 1 and test sample 2 respectively, the accuracy category corresponding to test sample 1 is determined to be positive and the accuracy category corresponding to test sample 2 is determined to be negative. Then, test sample 1 and test sample 2 are selected from the test sample set as multiple test samples for further analysis and processing.

[0166] Next, for test sample 1, the prompt word template used in the third language model to predict the output cause information for positive examples in the classification task is obtained. The test sample placeholder in the prompt word template is replaced with the evaluation content of test sample 1 and the actual prediction result and reference prediction result corresponding to test sample 1 to generate the third prompt word. The third prompt word is then input into the third language model so that the third language model outputs the output cause information 1 that is "relevant" to the actual prediction result corresponding to test sample 1.

[0167] Key elements: The trending events and the themes mentioned in the blog post are consistent, both involving apologies from corporate executives. The blog post connects the trending events through cultural interpretation.

[0168] Suggestions for improvement: Pay attention to the consistency of keywords between blog posts and trending events. When analyzing, the context should be considered to ensure the relevance and consistency of the topics.

[0169] For test sample 2, obtain the prompt word template from the third language model used to predict the output cause information for negative examples in classification tasks, and replace the test sample placeholders in the prompt word template with the evaluation content of test sample 2 and the actual prediction result and reference prediction result corresponding to test sample 2 to generate the third prompt word; input the third prompt word into the third language model so that the third language model outputs the output cause information 2 that is "relevant" to the actual prediction result corresponding to test sample 2.

[0170] Key elements: The definition of "the two descriptions have the same theme" was ignored, and the main theme of the blog post was not carefully compared.

[0171] Suggestions for improvement: Carefully read and understand the names of trending events and the content of blog posts to ensure that their main themes are consistent before making a judgment, and avoid concluding that a topic is relevant simply because of the appearance of keywords.

[0172] After determining output reason information 1 and output reason information 2, they are aggregated. Specifically, the prompt word template of the fourth language model is obtained, and the "key element" placeholder 3 and "improvement suggestion" placeholder 4 in the prompt word template are replaced with the "key element" and "improvement suggestion" from output reason information 1 and output reason information 2, respectively, to generate the fourth keyword. This fourth prompt word is then input into the fourth language model, so that the fourth language model outputs the aggregated result of the above output reason information 1 and output reason information 2, so as to generate supplementary prompt information corresponding to the initial prompt word template based on the aggregated result.

[0173] The aggregated results may contain at least one cause category summarized by the fourth language model, and each cause category contains corresponding specific descriptive information. For example, the aggregated results of output cause information 1 and output cause information 2 include:

[0174] Reason Category: Theme Consistency and Keyword Matching

[0175] Alignment with core theme: Ensure that the name of the trending event is consistent with the main theme of the blog post, and avoid judging relevance solely based on the appearance of keywords.

[0176] Keyword in-depth understanding: Pay attention to the context of keywords, understand their potential meanings, and ensure accurate relevance assessment even when the expression is not a direct description.

[0177] Finally, the clustering results are used as supplementary prompts to the initial prompt word template of the first language model. These supplementary prompts are then used as notes to update the initial prompt word template, resulting in the updated prompt word template.

[0178] ##Task Description: Based on the given trending event name and blog post content, determine whether the blog post content is related to the trending event. Regarding the definition of "related": both descriptions should be consistent in their main idea.

[0179] ## Output: Output: Relevant or irrelevant. Do not output anything else.

[0180] ## Content to be analyzed: < <content>>

[0181] ##Important Notes: Main Theme Consistency and Keyword Matching. Core Theme Alignment: Ensure the name of the trending event aligns with the main theme of the blog post; avoid judging relevance solely based on keyword appearance. In-depth Keyword Understanding: Pay attention to the context of keywords, understand their potential meanings, and ensure accurate relevance assessment even if the expression is not a direct description.

[0182] After obtaining the updated prompt word template from the initial prompt word template, a new first prompt word is generated based on the updated prompt word template, thereby updating the prompt word of the first language model. The new first prompt word contains information related to the precautions.

[0183] 2. Method for updating prompt words of the first major language model in the task scenario.

[0184] Assuming an extraction task scenario, the initial prompt word template for the first major language model is as follows:

[0185] ##Task Description: Please extract the names of car sales-related entities from the following dialogue log. Entities include: customer age, place of residence, car usage scenario, competitors' interests, and key selling points. If there are no car sales-related entity names in the dialogue log, output an empty string.

[0186] ## Output results: Please output the results in the format of

Entity Name: "Extracted Content"

[0187] ## Content to be analyzed: < <content>>

[0188] The test case set includes:

[0189] Test Sample 1 - Evaluation Content: Salesperson: "This actually brings you a different experience. Our four-wheel drive performance can also accommodate your self-driving tours. In fact, all six of these cars can be used for self-driving tours." Customer: "Yes, yes, what I mainly want to know is that your car also has a zero-gravity seat."

[0190] Test Sample 2 - Evaluation Content: Customer: "This shouldn't be announced yet. Basically, for example, the AA car, you can think about its supply chain, which is provided by us in xx. For example, its seats are quite good." Salesperson: "Hmm."

[0191] When updating the initial prompt word template of the first major language model, firstly, the placeholders < in the initial prompt word template are removed. <content>Replace the evaluation content of test sample 1 and test sample 2 respectively to generate the first prompt words corresponding to test sample 1 and test sample 2 respectively, and input the first prompt words into the first language model.

[0192] Assuming the first language model is based on the first prompt word corresponding to test sample 1, the actual prediction result for test sample 1 is:

[0193] Customer age: ""

[0194] Place of residence: ""

[0195] Use case: "Road trip"

[0196] Monitor competitors: ""

[0197] Key selling point: "Zero gravity seat".

[0198] The first language model, based on the first prompt word corresponding to test sample 2, outputs the actual prediction result for test sample 2 as follows:

[0199] Customer age: ""

[0200] Place of residence: ""

[0201] Usage scenario: "

[0202] Monitor competitors: "aa"

[0203] Focus on selling points: "".

[0204] If the actual prediction result corresponding to test sample 1 is correct, and the actual prediction result corresponding to test sample 2 is incorrect, then based on the actual prediction results corresponding to test sample 1 and test sample 2 respectively, the accuracy category corresponding to test sample 1 is determined to be positive and the accuracy category corresponding to test sample 2 is determined to be negative. Test sample 1 and test sample 2 are selected from the test sample set as multiple test samples for further analysis and processing.

[0205] Next, for test sample 1, the prompt word template used in the third language model to predict the output cause information for positive examples in the extraction task is obtained. The test sample placeholder in the prompt word template is replaced with the evaluation content of test sample 1 and the actual prediction result and reference prediction result corresponding to test sample 1 to generate the third prompt word. The third prompt word is then input into the third language model so that the third language model outputs the actual prediction result corresponding to test sample 1, which outputs cause information 1:

[0206] Tags: Car usage scenarios

[0207] Key points: During the correct reasoning process, attention was paid to the car usage scenarios mentioned by the salesperson in the conversation, and "self-driving tour" was correctly identified as the customer's car usage scenario.

[0208] Key element: The salesperson mentioned "we can also accommodate your road trips" during the conversation, implying that the customer might need the car for road trips, which is an important clue to determine the usage scenario.

[0209] Improvement suggestion: When analyzing the conversation, pay attention to the statements made by the salesperson that are related to the car usage scenario. Even if these statements do not directly answer the question about the car usage scenario, they may still be useful information for inferring the customer's car usage scenario.

[0210] Tags: Focus on selling points

[0211] Focus: The correct reasoning process focused on the zero-gravity seat mentioned by the customer as a potential selling point.

[0212] Key element: When a customer mentions in the conversation, "What I mainly want to know is that your car has a zero-gravity seat," it indicates that the customer may be interested in the zero-gravity seat.

[0213] Improvement suggestion: When extracting information, when customers actively mention certain product features, we should capture the customer's implicit points of interest in the conversation.

[0214] For test sample 2, obtain the prompt word template from the third language model used to predict the output cause information of negative examples in the extraction task, and replace the test sample placeholders in the prompt word template with the evaluation content of test sample 2 and the actual prediction result and reference prediction result corresponding to test sample 2 to generate the third prompt word; input the third prompt word into the third language model so that the third language model outputs the output cause information 2 of the actual prediction result corresponding to test sample 2:

[0215] Tags: permanent residence

[0216] Key point: In the flawed reasoning process, the location information mentioned by the customer in the conversation was ignored, and "Changchun" was not correctly identified as the customer's place of residence.

[0217] Key element: The customer mentions "we are from xx place" in the conversation, which implies that the customer may live in xx place. This is an important clue to determine the customer's permanent residence.

[0218] Improvement suggestion: When analyzing the conversation, pay special attention to statements related to geographical location, even if these statements do not directly answer questions about the customer's place of residence, they may still be useful information for inferring the customer's place of residence.

[0219] Tags: Focus on selling points

[0220] Key point: The flawed reasoning process overlooked the customer's mention of AA Auto's supply chain and seats as potential selling points.

[0221] Key elements: The customer mentioned in the conversation that part of AA Auto's supply chain comes from xx, and specifically mentioned the seats, indicating that the customer may be interested in these aspects.

[0222] Improvement suggestions: When extracting information, we should not only focus on the needs or preferences expressed directly, but also capture the customer's implicit interests in the conversation, especially when the customer actively mentions certain product features, which often reflect the selling points that the customer is interested in.

[0223] After determining output reason information 1 and output reason information 2, they are aggregated. Specifically, the prompt word template of the fourth language model is obtained, and the placeholders "tag" 1, "focus" 2, "key element" 3, and "improvement suggestion" 4 in the prompt word template are replaced with "tag", "focus", "key element", and "improvement suggestion" from output reason information 1 and output reason information 2, respectively, to generate the fourth keyword. This fourth prompt word is then input into the fourth language model, so that the fourth language model outputs the aggregated result of the above output reason information 1 and output reason information 2, so as to generate supplementary prompt information corresponding to the initial prompt word template based on the aggregated result.

[0224] The aggregated results may contain at least one cause category summarized by the fourth language model, and each cause category contains corresponding specific descriptive information. For example, the aggregated results of output cause information 1 and output cause information 2 include:

[0225] Reason Category 1: Identification of Geographic Location and Permanent Residence

[0226] Key takeaway: Accurately capture geographical location information mentioned in the conversation, such as city name and region name, to determine the customer's place of residence.

[0227] Improvement suggestions: Enhance keyword sensitivity to ensure no location-related details are missed; combine contextual understanding to ensure that the extracted location information truly represents the customer's usual residence.

[0228] Reason Category Two: Dialogue Content Analysis and Logical Reasoning

[0229] Key takeaway: Conduct in-depth analysis of the dialogue content and uncover implicit information from the customer's words and actions.

[0230] Improvement suggestions: Enhance the ability to analyze dialogue content, pay attention to details, and capture clues about car usage scenarios and key selling points from the context of the dialogue; learn to infer possible car usage scenarios or key selling points from the customer's attitude, tone, and level of interest in specific topics; even if not directly mentioned in the dialogue, try to reasonably infer from the context to avoid missing information.

[0231] Finally, the clustering results are used as supplementary prompts to the initial prompt word template of the first language model. These supplementary prompts are then used as notes to update the initial prompt word template, resulting in the updated prompt word template.

[0232] ##Task Description: Please extract the names of car sales-related entities from the following dialogue log. Entities include: customer age, place of residence, car usage scenario, competitors' interests, and key selling points. If there are no car sales-related entity names in the dialogue log, output an empty string.

[0233] ## Output results: Please output the results in the format of

Entity Name: "Extracted Content"

[0234] ## Content to be analyzed: < <content>>

[0235] ##Important Notes: 1) Identification of Geographic Location and Residence. Key Point: Accurately capture geographic location information mentioned in the conversation, such as city names and region names, to determine the customer's residence. Improvement Suggestions: Strengthen keyword sensitivity to ensure no details related to geographic location are missed; combine contextual understanding to ensure that the extracted location information truly represents the customer's residence. 2) Dialogue Content Analysis and Logical Reasoning. Key Point: In-depth analysis of dialogue content to uncover implicit information from the customer's words and actions. Improvement Suggestions: Enhance dialogue content analysis capabilities, pay attention to details, and capture clues about usage scenarios and key selling points from the context of the conversation; learn to infer possible usage scenarios or key selling points from the customer's attitude, tone, and level of interest in specific topics; even if not directly mentioned in the conversation, try to reasonably infer from the context to avoid missing information.

[0236] After obtaining the updated prompt word template from the initial prompt word template, a new first prompt word is generated based on the updated prompt word template, thereby updating the prompt word of the first language model. The new first prompt word contains information related to the precautions.

[0237] The following will describe in detail one or more embodiments of the prompt word update apparatus for a large language model according to the present invention. Those skilled in the art will understand that these apparatuses can all be configured using commercially available hardware components through the steps taught in this solution.

[0238] Figure 8 A schematic diagram of a prompt word template updating device for a large language model provided in this embodiment of the invention is shown below. Figure 8 As shown, the device includes: a prediction module 11, a processing module 12, and an update module 13.

[0239] The prediction module 11 is used to combine the initial prompt word template with the evaluation content in each test sample in the test sample set to form a first prompt word, which is then input into the first large language model so that the first large language model outputs the actual prediction results corresponding to each test sample.

[0240] Processing module 12 is configured to determine the accuracy category of each test sample in the test sample set based on the accuracy of the actual prediction results corresponding to each test sample in the test sample set; select multiple test samples from the test sample set, the multiple test samples corresponding to different accuracy categories; determine the output reason information of the actual prediction results corresponding to the multiple test samples; and generate supplementary prompt information corresponding to the initial prompt word template based on the output reason information.

[0241] The update module 13 is used to update the initial prompt word template according to the supplementary prompt information to obtain the updated prompt word template.

[0242] Optionally, the processing module 12 is specifically configured to: determine the accuracy category corresponding to each test sample based on the accuracy category labeling result fed back by the user based on the actual prediction result corresponding to each test sample; or, determine the accuracy category corresponding to each test sample based on the matching degree between the actual prediction result corresponding to each test sample and the reference prediction result corresponding to each test sample; the accuracy category includes positive examples, negative examples, and intermediate examples, the matching degree of test samples belonging to positive examples is greater than a first threshold, the matching degree of test samples belonging to intermediate examples is between the first threshold and a second threshold, the matching degree of test samples belonging to negative examples is less than the second threshold, and the first threshold is greater than the second threshold.

[0243] Optionally, the processing module 12 is further configured to: generate a second prompt word containing the actual prediction result and the reference prediction result corresponding to any test sample in the test sample set; input the second prompt word into a second large language model so that the second large language model determines the matching degree between the actual prediction result and the reference prediction result corresponding to any test sample; and determine the accuracy category corresponding to any test sample based on the matching degree between the actual prediction result and the reference prediction result corresponding to any test sample.

[0244] Optionally, the processing module 12 is further configured to perform clustering processing on each test sample in the test sample set to obtain multiple clusters; and to sample the multiple clusters to obtain the multiple test samples.

[0245] Optionally, the processing module 12 is further configured to: determine multiple test sample groups according to the accuracy categories corresponding to each test sample in the test sample set, wherein a test sample group consists of test samples corresponding to the same accuracy category; perform clustering processing on the multiple test sample groups respectively to obtain multiple clusters corresponding to each of the multiple test sample groups; and sample the multiple clusters corresponding to each of the multiple test sample groups respectively to obtain the multiple test samples.

[0246] Optionally, the processing module 12 is further configured to, for a target test sample among the plurality of test samples, generate a third prompt word corresponding to the target test sample based on the accuracy category of the target test sample. The third prompt word includes the evaluation content of the target test sample, the actual prediction result and reference prediction result corresponding to the target test sample, and the reason extraction task description information corresponding to the accuracy category of the target test sample. The target test sample is any one of the plurality of test samples. The third prompt word is then input into a third language model so that the third language model outputs the output reason information of the actual prediction result corresponding to the target test sample.

[0247] Optionally, the processing module 12 is further configured to aggregate the output reason information of the actual prediction results corresponding to the multiple test samples respectively; and generate supplementary prompt information corresponding to the initial prompt word template based on the aggregation result.

[0248] Optionally, the processing module 12 is further configured to generate a fourth prompt word corresponding to the plurality of test samples, wherein the fourth prompt word includes output reason information of the actual prediction results corresponding to the plurality of test samples and reason aggregation task description information; and input the fourth prompt word into the fourth language model so that the fourth language model outputs the aggregated result of the output reason information of the actual prediction results corresponding to the plurality of test samples.

[0249] Optionally, the update module 13 is specifically used to determine the target information in the supplementary prompt information that is different from the initial prompt word template; and add the target information to the initial prompt word template to obtain the updated prompt word template.

[0250] Optionally, the update module 13 is further configured to: determine the test performance indicators of the initial prompt word template and the updated prompt word template based on the test sample set; and determine a target prompt word template for use by the first large language model based on the test performance indicators, wherein the target prompt word template is the prompt word template with better test performance indicators between the initial prompt word template and the updated prompt word template.

[0251] Figure 8 The device shown can perform the steps described in the foregoing embodiments. For detailed execution process and technical effects, please refer to the description in the foregoing embodiments, which will not be repeated here.

[0252] In one possible design, the above Figure 8 The structure of the prompt word template update device for the large language model shown can be implemented as an electronic device, such as... Figure 9 As shown, the electronic device may include: a memory 21, a processor 22, and a communication interface 23. The memory 21 stores executable code, which, when executed by the processor 22, enables the processor 22 to at least implement the prompt word template update method for the large language model provided in the foregoing embodiments.

[0253] In addition, embodiments of the present invention provide a non-transitory machine-readable storage medium storing executable code, which, when executed by a processor of an electronic device, enables the processor to at least implement the prompt word template update method for a large language model as provided in the foregoing embodiments.

[0254] This invention provides a computer program product, including: a computer program that, when executed by a processor of an electronic device, causes the processor to execute the prompt word template update method for a large language model as provided in the foregoing embodiments.

[0255] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separate. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0256] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of a necessary general-purpose hardware platform, or by a combination of hardware and software. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a computer product. The present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0257] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.< / content> < / content> < / content> < / content> < / content> < / content> < / content> < / content> < / content>

Claims

1. A method for updating a prompt word template of a large language model, characterized in that, The method comprises the following steps: using an initial prompt word template, combining the evaluation content in each test example in a test example set to combine into a first prompt word input into a first large language model, so that the first large language model outputs the actual prediction results corresponding to each test example respectively; determining the accuracy categories corresponding to each test example in the test example set according to the accuracy of the actual prediction results corresponding to each test example in the test example set; selecting a plurality of test examples from the test example set, the plurality of test examples corresponding to different accuracy categories; determining the output reason information of the actual prediction results corresponding to the plurality of test examples respectively; generating supplementary prompt information corresponding to the initial prompt word template according to the output reason information; updating the initial prompt word template according to the supplementary prompt information to obtain an updated prompt word template.

2. The method of claim 1, wherein, The method comprises the following steps: determining the accuracy categories corresponding to each test example in the test example set according to the accuracy of the actual prediction results corresponding to each test example in the test example set; or determining the accuracy categories corresponding to each test example in the test example set according to the matching degree between the actual prediction results corresponding to each test example and the reference prediction results corresponding to each test example; The accuracy categories include positive examples, negative examples, and intermediate examples, the matching degree of the test examples belonging to the positive examples is greater than a first threshold value, the matching degree of the test examples belonging to the intermediate examples is between the first threshold value and a second threshold value, and the matching degree of the test examples belonging to the negative examples is less than the second threshold value, the first threshold value is greater than the second threshold value.

3. The method of claim 2, wherein, The method comprises the following steps: generating a second prompt word containing the actual prediction result and the reference prediction result corresponding to any test example in the test example set for the any test example; inputting the second prompt word into a second large language model to make the second large language model determine the matching degree between the actual prediction result and the reference prediction result corresponding to the any test example; determining the accuracy category corresponding to the any test example according to the matching degree between the actual prediction result and the reference prediction result corresponding to the any test example.

4. The method of claim 1, wherein, The method comprises the following steps: performing clustering processing on each test example in the test example set to obtain a plurality of clustering clusters; sampling the plurality of clustering clusters respectively to obtain the plurality of test examples.

5. The method of claim 4, wherein, The method comprises the following steps: determining a plurality of test example groups according to the accuracy categories corresponding to each test example in the test example set, wherein a test example group is composed of test examples corresponding to the same accuracy category; respectively, to obtain a plurality of clustering clusters corresponding to each of the plurality of test example groups; The sampling of the plurality of clustering clusters respectively to obtain the plurality of test examples comprises: Sampling the plurality of clustering clusters corresponding to each of the plurality of test example groups respectively to obtain the plurality of test examples.

6. The method of claim 1, wherein, The determining of the output reason information of the actual prediction result corresponding to the plurality of test examples respectively comprises: For a target test example in the plurality of test examples, a third prompt word corresponding to the target test example is generated according to the accuracy category of the target test example, the third prompt word including the evaluation content of the target test example, the actual prediction result and the reference prediction result corresponding to the target test example, and the reason extraction task description information corresponding to the accuracy category of the target test example, the target test example being any one of the plurality of test examples; The third prompt word is input into a third large language model to make the third large language model output the output reason information of the actual prediction result corresponding to the target test example.

7. The method of claim 1, wherein, The generating of the supplementary prompt information corresponding to the initial prompt word template according to the output reason information comprises: Aggregating the output reason information of the actual prediction result corresponding to the plurality of test examples respectively; Generating the supplementary prompt information corresponding to the initial prompt word template according to the aggregation result.

8. The method of claim 7, wherein, The aggregating of the output reason information of the actual prediction result corresponding to the plurality of test examples respectively comprises: Generating a fourth prompt word corresponding to the plurality of test examples, the fourth prompt word including the output reason information of the actual prediction result corresponding to the plurality of test examples respectively and the reason aggregation task description information; The fourth prompt word is input into a fourth large language model to make the fourth large language model output the aggregation result of the output reason information of the actual prediction result corresponding to the plurality of test examples.

9. The method of claim 1, wherein, The updating of the initial prompt word template according to the supplementary prompt information to obtain an updated prompt word template comprises: Determining target information in the supplementary prompt information that is different from the initial prompt word template; Adding the target information to the initial prompt word template to obtain an updated prompt word template.

10. The method of claim 1, wherein, The method further comprises: Determining a test performance index of the initial prompt word template and the updated prompt word template according to a test example set; Determining a target prompt word template for the first large language model according to the test performance index, the target prompt word template being a prompt word template with better test performance index between the initial prompt word template and the updated prompt word template.

11. An electronic device, comprising: Comprise: A memory, a processor, a communication interface; wherein the memory has stored executable codes, when the executable codes are executed by the processor, the processor executes the prompt word template updating method of the large language model according to any one of claims 1 to 10.

12. A non-transitory machine-readable storage medium, comprising: The non-transitory machine-readable storage medium stores executable code, which, when executed by a processor of an electronic device, causes the processor to perform the prompt word template updating method of the large language model according to any one of claims 1 to 10.

13. A computer program product, characterised in that, Comprise: A computer program, which, when executed by a processor of an electronic device, causes the processor to perform the prompt word template updating method of the large language model according to any one of claims 1 to 10.