Adaptive Template Reconstruction Method and System Based on a Small Number of Samples
By adopting the adaptive template reconstruction method in a small sample scenario, using a large language model and a multi-level similarity comparison module, the problem of low quality of template construction problems and answer generation in the existing technology is solved, and high-quality answer generation and task completion accuracy is achieved.
Patent Information
- Application Number
- CN202510297339.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-13
AI Technical Summary
The prior art is difficult to build high-quality templates in a small sample scenario, resulting in uneven text quality and inability to meet the accuracy required by the task.
Adaptive template reconstruction method based on a small number of samples is adopted, and the initial prompt template is constructed, and a large language model is input to generate a question pool and an answer pool. The multi-level similarity comparison module is used to filter and cluster analysis, and the template is dynamically adjusted to improve the accuracy of answer generation.
The task performance of large language models in small sample scenarios has been significantly improved, and the quality and diversity of the generated answers have been improved, which can more accurately meet the needs of complex tasks.
Smart Images

Figure CN119807421B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular, to an adaptive template reconstruction method and system based on a small number of samples. Background Art
[0002] With the wide application of large language models (such as GPT series, LLaMA, etc.) in natural language processing, they have demonstrated powerful capabilities in handling various tasks. However, current mainstream models usually rely on a large amount of training data and perform poorly in few-shot or zero-shot learning scenarios. When the input data is insufficient, the model may not be able to generate high-quality and relevant answers, thus affecting its performance on specific tasks. The existing few-shot learning techniques mainly face the following key problems:
[0003] 1. Insufficient template matching ability: Most existing generative models rely on manually constructed templates, but these templates are often difficult to adapt to complex task requirements in a small number of sample environments, and the generated text quality is uneven, affecting the output effect of the model.
[0004] 2. Inability to balance diversity and accuracy in generation: Generative models need to strike a balance between diversity and accuracy, often sacrificing one aspect to improve the performance of the other. This method is not applicable to fields with more precise task requirements, such as legal texts, bidding information, medical documents, etc.
[0005] 3. Imperfect answer screening and optimization process: For the generated answers, existing screening methods often lack a fine-grained similarity comparison mechanism and cannot fully utilize the semantic information of the expected answers, thus affecting the accuracy and consistency of the answers.
[0006] In recent years, few-shot learning and adaptive generation techniques have gradually attracted attention. These techniques can adjust template prompts to guide the model to output results more in line with task requirements without a large-scale dataset by combining existing pre-trained language models. However, there are still significant deficiencies in existing technologies when dealing with multi-dimensional tasks, generating diverse answers, and dynamically adjusting template prompts according to specific questions.
[0007] Therefore, finding a method that can both construct templates in a small number of sample scenarios and improve the accuracy of answer generation is a technical problem that needs to be urgently solved by those skilled in the art. Summary of the Invention
[0008] The present invention provides an adaptive template reconstruction method and system based on a small number of samples to solve the defect in the prior art that template reconstruction cannot be performed based on a small number of samples, and realizes improving the task performance in a small number of sample scenarios through adaptive template reconstruction and keyword supplementation.
[0009] The present invention provides an adaptive template reconstruction method based on a small number of samples, comprising the following steps:
[0010] S1. Construct an initial prompt template based on a small number of sample data, where the initial prompt template includes a task background description, context information, and a part to be completed;
[0011] S2. Input the initial prompt template into a large language model to generate a question pool through specific instructions, proxy instructions, and example instructions;
[0012] S3. Use the large language model to answer each question in the question pool to generate an answer pool;
[0013] S4. Screen and cluster-analyze the answers in the answer pool through a multi-level similarity comparison module to obtain a screened answer pool and screened answer groups, and dynamically adjust the weight parameters of similarity calculation in the multi-level similarity comparison module through a model self-feedback mechanism to obtain the overall score of the screened answers;
[0014] S5. Calculate the comprehensive similarity scores of each answer group based on the overall scores of the answers in each answer group, and conduct expert scoring on each answer group. Construct examples based on the comprehensive similarity scores and expert scoring, and reconstruct the initial prompt template based on the examples to obtain a reconstructed prompt template;
[0015] S6. Input the reconstructed prompt template into the large language model for iterative update until the final answer is obtained.
[0016] According to an adaptive template reconstruction method based on a small number of samples provided by the present invention, the specific instruction is an instruction designed according to the task objective, which is used to guide the large language model to generate answers related to the task. The proxy instruction is an instruction that guides the large language model to generate answers in an indirect manner. The example instruction is generated by providing example questions and answers.
[0017] According to an adaptive template reconstruction method based on a small number of samples provided by the present invention, the generation of the question pool specifically includes:
[0018] Input the initial prompt template into the large language model for entity expansion to obtain an entity set. Supplement keywords to the initial prompt template based on the entity set and adjust the weights of the keywords. The weight update formula is:
[0019] ;
[0020] where represents the weight value of the f-th keyword in the (t + 1)-th round, represents the weight value of the f-th keyword in the t-th round, Represents the adjustment parameter for relevance evaluation, Represents the keyword The relevance evaluation value of the keyword with the context Context, Represents the f-th keyword, Represents the current context environment, Represents the adjustment parameter for frequency evaluation, Represents the keyword The frequency of occurrence in the text;
[0021] The expanded keywords are divided into three categories: core keywords, secondary keywords, and background keywords, and are prioritized through differential initial weights to obtain a keyword sequence; where the initial weights are:
[0022] ;
[0023] Among them, Represents the initial weight value of the f-th keyword, where f represents the index of the keyword, Represents the weight adjustment coefficient, Represents the f-th keyword, Represents the priority scoring function;
[0024] The keyword sequence is combined with specific instructions, proxy instructions, and example instructions to generate multiple question variants, forming an initial question pool;
[0025] The initial question pool is de-duplicated and screened to obtain a screened question pool; where the screened question pool is represented as:
[0026] ;
[0027] Among them, Represents the screened question pool, Represents the current question to be evaluated, Represents except outside in other questions, where i represents the number of the current question, j represents the number of other questions except i, and Q represents the initial question pool, Represents the question and The similarity metric value between, Represents the similarity threshold.
[0028] According to an adaptive template reconstruction method based on a small number of samples provided by the present invention, the processing steps of the multi-level similarity comparison module are:
[0029] Perform primary similarity calculation on the answers in the answer pool to obtain a primary similarity score, and preliminarily screen the answer pool based on the primary similarity score to obtain a first answer pool;
[0030] Perform semantic similarity calculation on the answers in the first answer pool to obtain semantic similarity scores, and perform screening based on the semantic similarity scores to obtain a second answer pool;
[0031] Perform cluster analysis on the answers in the second answer pool based on the semantic similarity scores to form multiple answer groups.
[0032] According to an adaptive template reconstruction method based on a small number of samples provided by the present invention, the model self-feedback mechanism includes:
[0033] Perform weighted combination on the primary similarity score and the semantic similarity score to obtain an initial similarity score;
[0034] Dynamically adjust the initial similarity score based on coverage and specificity to obtain a total similarity score.
[0035] According to an adaptive template reconstruction method based on a small number of samples provided by the present invention, step S5 specifically includes:
[0036] Set a total score threshold. When the total score of the answers in the answer group is lower than the total score threshold, trigger the comprehensive scoring mechanism:
[0037] Calculate the comprehensive similarity scores of each answer group, and conduct expert scoring on each answer group. Sort the answer groups from largest to smallest according to the comprehensive similarity scores and expert scores to form an answer group sequence;
[0038] Use the first group of answers in the answer group sequence as an example, and reconstruct the initial prompt template based on the example to obtain a reconstructed prompt template.
[0039] According to an adaptive template reconstruction method based on a small number of samples provided by the present invention, the construction of the initial prompt template based on a small number of sample data specifically includes:
[0040] Preprocess the small number of sample data, and add identifiers for clarifying the context range at the beginning and end of the preprocessed small number of sample data;
[0041] Design an initial prompt template according to the task objective and context information;
[0042] Combined with the prompt engineering method of the large language model, understand the context through prompt words, and dynamically adjust the structure and vocabulary of the initial prompt template based on the context; wherein the prompt words include a clear description of the task objective and context prompts related to the task.
[0043] The present invention also provides an adaptive template reconstruction system based on a small number of samples, which implements the above-mentioned adaptive template reconstruction method, including:
[0044] A sample analysis module for constructing an initial prompt template based on a small amount of sample data, where the initial prompt template includes a task background description, context information, and a target information prompt to be completed;
[0045] A question generation module for inputting the initial prompt template into a large language model to generate a question pool through specific instructions, proxy instructions, and example instructions;
[0046] An answer generation module for using the large language model to answer each question in the question pool to generate an answer pool;
[0047] A similarity screening module for screening and clustering analysis of the answers in the answer pool through a multi-level similarity comparison module to obtain a screened answer pool and a screened answer group, and calculating the overall score of the answers according to the screened answer group;
[0048] A template reconstruction module for calculating the comprehensive similarity score and expert score of each answer group based on the overall score of the screened answer group, sorting the answer groups according to the comprehensive similarity score and expert score, selecting the answer group with the highest score as an example, and optimizing and reconstructing the initial prompt template to generate a reconstructed prompt template;
[0049] An iterative optimization module for inputting the reconstructed prompt template into the large language model for iterative update to continuously optimize the generated answers until the final answer is obtained.
[0050] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor implements the adaptive template reconstruction method as described in any one of the above when executing the program.
[0051] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the adaptive template reconstruction method as described in any one of the above.
[0052] An adaptive template reconstruction method and system based on a small amount of samples provided by the present invention, through adaptive template reconstruction and keyword supplementation techniques, combined with a multi-level similarity comparison module and a model self-feedback mechanism, effectively improve the quality of answer generation and the accuracy of task completion, significantly improve the task performance of the large language model in the small sample scenario, and enable the model to generate high-quality answers based on limited samples. Description of the Drawings
[0053] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the accompanying drawings required in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0054] Figure 1 is a flowchart of the adaptive template reconstruction method provided by the present invention;
[0055] Figure 2 is a block diagram of the adaptive template reconstruction method provided by the present invention;
[0056] Figure 3 is a schematic flowchart of an embodiment of the adaptive template reconstruction provided by the present invention;
[0057] Figure 4 is a schematic structural diagram of the electronic device provided by the present invention. Detailed Embodiments
[0058] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0059] As Figure 1 and Figure 2 shown, the present invention provides an adaptive template reconstruction method based on a small number of samples, including the following steps:
[0060] S1. Construct an initial prompt template based on a small amount of sample data, where the initial prompt template includes a task background description, context information, and a part to be completed;
[0061] S2. Input the initial prompt template into a large language model to generate a question pool through specific instructions, proxy instructions, and example instructions;
[0062] S3. Use the large language model to answer each question in the question pool to generate an answer pool;
[0063] S4. Through a multi-level similarity comparison module, screen and cluster analyze the answers in the answer pool to obtain a screened answer pool and a screened answer group, and dynamically adjust the weight parameters of the similarity calculation in the multi-level similarity comparison module through a model self-feedback mechanism to obtain the overall score of the screened answers;
[0064] S5. Calculate the comprehensive similarity scores of each answer group based on the overall scores of the answers in each answer group, and conduct expert scoring on each answer group. Construct examples based on the comprehensive similarity scores and expert scoring, and reconstruct the initial prompt template based on the examples to obtain the reconstructed prompt template;
[0065] S6. Input the reconstructed prompt template into the large language model for iterative update until the final answer is obtained.
[0066] Among them, the preprocessing includes removing HTML codes, page tags, escape symbols, etc., converting English punctuation to Chinese punctuation, deleting transfer symbols, converting line breaks to spaces, etc.
[0067] It can be understood that the task goal is to solve the task performance problem of the large language model (LLM) in the few-shot scenario. When constructing the initial prompt template, first embed the text to be processed into a unified prompt prefix structure, that is, place the specific activity text content between the "START" and "END" identifiers to form a standardized input format; then fuse the preset extraction target with the text to construct a prompt statement for entity recognition, requiring the model to output entity type information related to the extracted entity.
[0068] Such as Figure 2 As described above, the initial prompt template includes Context, Target, and Question. Among them, Context includes the text embedding prompt prefix " " that is, "START Current location: Home page Announcement information × Procurement winning bid (result) announcement of a certain unit in a certain city... Attachment: Bidding documents of a certain unit (final version).pdf; END", Target is: × Bidding Co., Ltd. of a certain province, and Question is: What are the three most relevant entity types corresponding to the entity in the Context of Target? Output in the form of a list;
[0069] Input the initial prompt template into the large language model (LLM) to generate a question pool, including Question 1, Question 2, and Question 3;
[0070] Input the question pool into the large language model to generate an answer pool, including Answer 1, Answer 2, and Answer 3;
[0071] Use the multi-level similarity comparison module to screen and cluster analyze the answers in the answer pool, and calculate the comprehensive similarity scores of each answer group, including Score 1, Score 2, and Score 3;
[0072] Select the answer group with the highest comprehensive similarity score to construct a new prompt template, and obtain the reconstructed prompt template, including examples, question 4, and answers;
[0073] Input the reconstructed prompt template into the large language model to obtain the final answer.
[0074] Through the adaptive template reconstruction and keyword supplementation technologies, combined with the multi-level similarity comparison module and the model self-feedback mechanism, the present invention effectively improves the quality of answer generation and the accuracy of task completion, significantly enhances the task performance of the large language model in the few-shot scenario, and enables the model to generate high-quality answers based on limited samples.
[0075] In an embodiment of the present invention, constructing the initial prompt template based on the few-shot data specifically includes:
[0076] Preprocess the few-shot data, and add identifiers for clarifying the context range at the beginning and end of the preprocessed few-shot data;
[0077] Design the initial prompt template according to the task objective and context information;
[0078] Combined with the prompt engineering method of the large language model, understand the context through the prompt words, and dynamically adjust the structure and vocabulary of the initial prompt template based on the context; wherein the prompt words include a clear description of the task objective and context prompts related to the task.
[0079] It can be understood that for the input bidding text, first construct the basic text template: "T(x) = Given the bidding activity text: START[specific text content]END", where the specific text content is the actual bidding text;
[0080] Integrate the preset extraction target into the template to form a complete prompt statement: "T'(x) = The corresponding entity result to be extracted is: Target. What are the entity types corresponding to the extracted entity values in the bidding activity? Output the three most relevant entity types in a list form.", where Target corresponds to the preset extraction target;
[0081] Input the complete prompt template into the large language model, and the model outputs the list of the three most relevant entity types:
[0082]
[0083] Among them, represents the three most relevant entity types output, and LM represents the processing of the large language model.
[0084] Illustrate with a specific embodiment:
[0085] When processing a bidding text containing information about a procurement agency, if the preset extraction target is "××× Bidding Co., Ltd. in a certain province", the constructed complete prompt template will guide the model to identify the possible entity types corresponding to this entity in the bidding activities, such as "procurement agency", "bidding agency", "company name", etc. This structured template design method not only ensures that the model can accurately understand the task requirements but also provides sufficient context information support, effectively improving the accuracy of entity type recognition.
[0086] In an embodiment of the present invention, the generation of the problem pool specifically includes:
[0087] Input the initial prompt template into a large language model for entity expansion to obtain an entity set, supplement keywords to the initial prompt template based on the entity set, and adjust the weights of the keywords. The weight update formula is:
[0088] ;
[0089] Where, represents the weight value of the f-th keyword in the (t + 1)-th round, represents the weight value of the f-th keyword in the t-th round, represents the adjustment parameter for relevance evaluation, represents the keyword and the relevance evaluation value with the context Context, represents the f-th keyword, represents the current context environment, represents the adjustment parameter for frequency evaluation, represents the keyword and the frequency of occurrence in the text;
[0090] Classify the expanded keywords into three categories: core keywords, secondary keywords, and background keywords, and perform priority sorting through differential initial weights to obtain a keyword sequence; where the initial weights are:
[0091] ;
[0092] Where, represents the initial weight value of the f-th keyword, f represents the index of the keyword, represents the weight adjustment coefficient, represents the f-th keyword, represents the priority scoring function;
[0093] Generate multiple question variants by combining keyword sequences with specific instructions, proxy instructions, and example instructions to form an initial question pool; where the specific instructions are instructions designed according to the task objective, used to guide the large language model to generate answers related to the task, the proxy instructions are instructions that guide the large language model to generate answers in an indirect way, and the example instructions are generated by providing example questions and answers;
[0094] Deduplicate and filter the initial question pool to obtain the filtered question pool; where the filtered question pool is represented as:
[0095] ;
[0096] Among them, represents the filtered question pool, represents the question to be evaluated currently, represents except outside in other questions, i represents the current question number, j represents the numbers of other questions except i, Q represents the initial question pool, represents the question and the similarity metric value between, represents the similarity threshold.
[0097] It can be understood that the specific instructions are used to guide the large language model to generate answers related to the task, such as a clear instruction like "Please extract the `Entity`-related entities from the following bidding text"; the proxy instructions are to let the model understand the task in a specific context through role-playing and other means, and identify key points in a more indirect way, such as "Suppose you are a model very proficient in information extraction. Now I will give you a passage about bidding. Please correctly extract the entities I want according to my requirements. Now I want to extract the contacts in the passage. The text is as follows: \nSTART…END ", and the example instructions are used to help the model understand the task requirements through examples and show how to perform entity extraction, such as "The entity type examples given in the example instructions help the model generate questions with similar structures:
[0098] Given the example START…END, extract the Entity as Target. Given the target extraction text START…END, now extract the Entity in it. Please help me extract".
[0099] The present invention introduces a dual evaluation of context relevance and frequency evaluation into the weight adjustment formula, enabling the adjustment of weights not only based on the global context but also taking into account the statistical characteristics of specific domains. By classifying keywords and designing priority ranking strategies respectively: Core keywords are directly associated with the domain task objectives; Secondary keywords are extended vocabulary closely associated with core keywords; Background keywords are low-priority vocabulary generated to supplement semantic integrity. By assigning different initial weights with different priorities to different keywords, the ranking effect of high-priority keywords is improved.
[0100] In an embodiment of the present invention, the processing steps of the multi-level similarity comparison module are as follows:
[0101] Perform primary similarity calculation on the answers in the answer pool to obtain primary similarity scores, and based on the primary similarity scores, preliminarily screen the answer pool to obtain the first answer pool;
[0102] Perform semantic similarity calculation based on deep learning on the answers in the first answer pool to obtain semantic similarity scores, and based on the semantic similarity scores, perform screening to obtain the second answer pool;
[0103] Perform clustering analysis on the answers in the second answer pool based on the semantic similarity scores to form multiple answer groups.
[0104] The present invention constructs a comprehensive similarity score model through weighted calculation of primary similarity scores and high-level similarity scores. The multi-level similarity comparison module can accurately measure and screen the similarity of the answers in the answer pool, avoiding the existence of redundant and duplicate answers. At the same time, by dynamically adjusting the weight parameters to adapt to different task requirements, the accuracy and robustness of the similarity scoring are further improved.
[0105] In an embodiment of the present invention, the calculation process of the primary similarity score is as follows:
[0106] Based on all the answers in the answer pool, construct a unified vocabulary. Select any two answers from the answer pool and convert them into word frequency vectors based on the same vocabulary and , where the vocabulary sizes of the word frequency vector A and the word frequency vector B are the same, represents the g-th word frequency value in the word frequency vector A, represents the h-th word frequency value in the word frequency vector B, and d represents the number of answers in the answer pool;
[0107] Calculate the primary similarity scores of the two answers based on the word frequency vector A and the word frequency vector B. The calculation formula is:
[0108]
[0109] Set a primary similarity threshold. If the cosine similarity is less than the threshold, the corresponding answer is filtered out.
[0110] Furthermore, the calculation process of the semantic similarity score is as follows:
[0111] Input any two answers in the first answer pool into a deep learning model (such as BERT, Bidirectional Encoder Representations from Transformers) to obtain the answer sentence vector representation and the answer sentence vector representation , similarly, use cosine similarity to calculate the similarity between sentence vectors to obtain the semantic similarity score. The calculation formula is:
[0112]
[0113] where q represents the number of answers in the first answer pool, r represents the index of the answer in the answer sentence vector representation and s represents the index of the answer in the answer sentence vector representation ;
[0114] Set a semantic similarity threshold. If the similarity between sentence vectors is less than the semantic similarity threshold, the corresponding answer is filtered out.
[0115] It can be understood that both the primary similarity threshold and the semantic similarity threshold can be set according to actual usage requirements, and the present invention does not make specific limitations on this.
[0116] Illustrate with a specific embodiment:
[0117] There are three answers in the answer pool: "The project budget amount is 5 million", "The total project budget is 5 million yuan", and "The tender budget is 5 million";
[0118] Construct a unified vocabulary as: "{"project", "budget", "amount", "engineering", "total", "yuan", "for", "tender", "million"}";
[0119] Select any two answers and convert them into word frequency vectors: Answer 1 is converted into vector A: (1, 1, 1, 0, 0, 0, 0, 0, 1), and Answer 2 is converted into vector B: (0, 1, 0, 1, 1, 1, 0, 0, 1);
[0120] Calculate the cosine similarity of these two word frequency vectors to obtain the primary similarity score.
[0121] Further, the two answers in the answer pool: "The project budget amount is 5 million" and "The total project budget is 5 million yuan" are respectively input into BERT to obtain two answer sentence vector representations and ;
[0122] Calculate the cosine similarity between the two answer sentence vector representations and to obtain the semantic similarity score.
[0123] In an embodiment of the present invention, the steps of clustering analysis are as follows:
[0124] Utilize the semantic similarity score to construct a similarity matrix , where represents the similarity between the answer corresponding to the semantic similarity score and . Secondly, the K-means clustering method is used to cluster the sentence vector representations of all answers into K clusters, and the cluster center is defined as . The clustering objective is:
[0125]
[0126] where represents the center point of the kth cluster, k represents the cluster number, K represents the total number of clusters, N represents the total number of answers in the second answer pool, represents the sentence vector representation of the ith answer in the second answer pool;
[0127] It can be understood that the cluster center with the highest similarity represents the finally selected answer set.
[0128] In an embodiment of the present invention, the top-k candidate examples with the highest comprehensive similarity score are selected, and according to their core structure and semantic characteristics, a prompt template is constructed. After reconstruction, the prompt template contains the core keywords and semantic framework of the task, which is convenient for guiding the large language model to generate consistent answers.
[0129] Specifically, extract the core keywords from each comprehensive similarity score, where represents the mth extracted core keyword, m represents the total number of core keywords, and combine with the answer structure to form the semantic framework of the template:
[0130]
[0131] where represents the example template, represents the operation of structuring the content, represents the candidate example with the highest comprehensive similarity score, and k represents the number of candidate examples selected;
[0132] Concatenate the example template with the target task to obtain the reconstructed prompt template :
[0133]
[0134] Among them, represents the concatenation operation;
[0135] Input the reconstructed prompt template into the large language model to obtain the initial result:
[0136]
[0137] Among them, represents the initial result;
[0138] Gradually optimize the template structure based on the initial result, and the formula is:
[0139]
[0140] Among them, t represents the current iteration round, represents the prompt template in the t-th round, is the keyword extracted from the result of the t-th round of iteration, p represents the number of newly extracted keywords, and input the updated template into the model again to obtain a new round of answers :
[0141]
[0142] Continue the iterative process until the final answer that meets the target task.
[0143] In an embodiment of the present invention, for the question , generate m different answers , among them, represents the m-th answer to the i-th question, i represents the question number, m represents the number of answers generated for each question, and the formula is:
[0144]
[0145] Among them, is the generation temperature parameter, which controls the diversity of the model output, represents the language model, j represents the number of answers generated, and the range of j is [1, m], represents The j-th value of the parameter; By setting different temperatures, various forms of answers can be generated;
[0146] Merge the answer sets of each question in to form an answer pool:
[0147]
[0148] where n represents the number of questions.
[0149] It can be understood that , , .
[0150] Furthermore, in order to avoid redundant and duplicate answers in the answer pool, a similarity metric is used to screen the answers. represents the similarity metric between two answers, and a similarity threshold is set. The finally screened answer pool is represented as:
[0151]
[0152] where represents the k-th answer generated by the i-th question, the range of i is [1, m], and the range of k is [1, m]. represents the l-th answer generated by the j-th question, the range of j is [1, m], and the range of l is [1, m]. represents the answer and the answer between the similarities.
[0153] The second answer pool calculates the similarity score between each answer and the target answer through a semantic similarity metric method:
[0154]
[0155] where represents the similarity score of the answer , represents the similarity score calculation function, and the answer set with the highest score is selected according to the similarity score: Answer set:
[0156]
[0157] Dynamically adjust the weight parameters of the similarity calculation through the model self-feedback mechanism to obtain the overall score of the screened answers.
[0158] In an embodiment of the present invention, the model self-feedback mechanism includes:
[0159] The primary similarity score and the semantic similarity score are weighted and combined to obtain an initial similarity score;
[0160] The initial similarity score is dynamically adjusted based on coverage and specificity to obtain a total similarity score; where the coverage is used to measure whether the answer covers the key topics or information, ensuring that the generated answer is complete and contains all the important information required by the question, and the specificity is used to ensure the specificity and accuracy of the answer, preventing the generated answer from being too general or ambiguous.
[0161] In an embodiment of the present invention, the calculation formula for the overall score of the screened answer is:
[0162]
[0163] Where, represents the comprehensive similarity score of the answer, represents the weight coefficient of the primary similarity, represents the weight coefficient of the semantic similarity, represents the primary similarity score, represents the semantic similarity score;
[0164] The optimized overall score is:
[0165]
[0166] Where, represents the final overall score, represents the weight coefficient of the coverage index, represents the weight coefficient of the specificity index, represents the degree of coverage of the answer to the key points of the question, represents the degree of specificity and accuracy of the answer;
[0167] Among them, the weight coefficient of the coverage index is obtained through iterative optimization by multiple rounds of experiments and manual evaluation, the weight coefficient of the specificity index is obtained through control experiments with different weight values and experimental optimization, the degree of coverage of the answer to the key points of the question is calculated by evaluating the inclusion of the key information points of the question in the answer, that is, for each answer A, check whether it contains the key information points, and the degree of specificity and accuracy of the answer is determined through information density evaluation, detail degree evaluation, accuracy evaluation and relevance evaluation.
[0168] It can be understood that based on the iterative generation and screening mechanism of the answer pool, the answer quality is continuously improved through multiple rounds of optimization, while maintaining the diversity of the answers, enabling the system to meet the multi-dimensional requirements of complex questions while ensuring the accuracy of the answers.
[0169] Through the model self - feedback mechanism, the present invention dynamically adjusts the weight parameters in the multi - level similarity comparison module, realizes the adaptive optimization of the similarity calculation method, gradually improves the accuracy and precision of answer selection, enhances the adaptability of the system to different types of tasks, and strengthens the robustness of the method.
[0170] In an embodiment of the present invention, step S5 specifically includes:
[0171] Set a total score threshold. When the total score of the answers in the answer group is lower than the total score threshold, trigger the comprehensive scoring mechanism:
[0172] Calculate the comprehensive similarity scores of each answer group, conduct expert scoring on each answer group, sort the answer groups from largest to smallest according to the comprehensive similarity scores and expert scores, and form an answer group sequence;
[0173] Take the first group of answers in the answer group sequence as an example, and reconstruct the initial prompt template based on the example to obtain the reconstructed prompt template.
[0174] The main evaluation indicators of expert scoring include accuracy, relevance, and coverage, etc., which measure the matching degree of answers respectively. Among them, accuracy is used to measure whether the answer meets the question requirements, relevance is used to measure the semantic relevance between the answer and the reference answer, and coverage is used to measure whether the answer covers key topics or information.
[0175] As Figure 3 shown, a specific embodiment is used to illustrate the adaptive template reconstruction method of the present invention:
[0176] S10. Perform de - coding processing on the obtained original tender document web page text to obtain web page text without code;
[0177] S20. Use regular expressions for preliminary entity recognition, use a large - language model for secondary screening, and conduct refined screening through manual verification to form an adaptive template prompt small - sample data set;
[0178] S30. Divide the data set into a training set and a test set, use the training set as the fine - tuning corpus of the large - language model, use the test set to test the accuracy of the model, and obtain the fine - tuned model;
[0179] S40. Use the fine - tuned model to extract the entity types and values in the tendering and bidding text, and complete the extraction of tendering and bidding text information.
[0180] Among them, after the web page file is downloaded, in addition to the text file displayed on the page, it also contains a large amount of information such as page code, page tags, and escape symbols. Therefore, it is necessary to preprocess it to initially remove the content other than the tender text in the downloaded file and improve the information density after preprocessing. Therefore, a web page parsing library is used to parse the web page text without code, remove the HTML tags in the parsed file to obtain the web page text without tags, and continue to process the punctuation marks of the target web page text, convert English punctuation to Chinese punctuation, delete transfer symbols, convert line breaks to spaces, and add text identifiers at the beginning and end of the text.
[0181] There is a large amount of entity-irrelevant information in the tender text. First, regular expressions are used to initially screen the entity types and values in the expression text. Since there are differences in the text formats of tender texts, regular expressions cannot cover all types of tender texts. Therefore, a large language model is used to conduct a secondary screening of the tender text. At this time, the obtained entity information is already relatively reliable. Finally, manual fine screening is performed on the tender text to obtain a small sample dataset for adaptive template prompts.
[0182] The present invention provides an adaptive template reconstruction system based on a small number of samples, which implements the adaptive template reconstruction method as described above, including:
[0183] A sample analysis module for constructing an initial prompt template based on a small amount of sample data, where the initial prompt template includes a task background description, context information, and a target information prompt to be completed;
[0184] A question generation module for inputting the initial prompt template into a large language model to generate a question pool through specific instructions, proxy instructions, and example instructions;
[0185] An answer generation module for using the large language model to answer each question in the question pool to generate an answer pool;
[0186] A similarity screening module for screening and clustering analysis of the answers in the answer pool through a multi-level similarity comparison module to obtain a screened answer pool and a screened answer group, and calculating the overall score of the answers according to the screened answer group;
[0187] A template reconstruction module for calculating the comprehensive similarity score of each answer group based on the overall score of the screened answer group, sorting the answer groups according to the comprehensive similarity score, selecting the answer group with the highest score as an example, and optimizing and reconstructing the initial prompt template to generate a reconstructed prompt template;
[0188] An iterative optimization module for inputting the reconstructed prompt template into a large language model for iterative update to continuously optimize the generated answers until the final answer is obtained.
[0189] The adaptive template reconstruction device provided by the present invention will be described below. The adaptive template reconstruction device described below can be correspondingly referred to the adaptive template reconstruction method described above.
[0190] Figure 4 The schematic physical structure diagram of an electronic device is exemplified, as Figure 4 shown. The electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communications interface 420, and the memory 430 complete the communication with each other through the communication bus 440. The processor 410 can call the logical instructions in the memory 430 to execute the adaptive template reconstruction method, which includes: constructing an initial prompt template based on a small amount of sample data, inputting the initial prompt template into a large language model, and generating a question pool through specific instructions, proxy instructions, and example instructions; using the large language model to answer each question in the question pool to generate an answer pool; screening and clustering analysis of the answers in the answer pool through a multi-level similarity comparison module to obtain a screened answer pool and a screened answer group, and dynamically adjusting the weight parameters of similarity calculation in the multi-level similarity comparison module through a model self-feedback mechanism to obtain the overall score of the screened answers; calculating the comprehensive similarity score of each answer group based on the overall score of the answers in each answer group, and conducting an expert evaluation on each answer group, constructing examples based on the comprehensive similarity score and the expert evaluation, and reconstructing the initial prompt template based on the examples to obtain a reconstructed prompt template; inputting the reconstructed prompt template into the large language model for iterative update until the final answer is obtained.
[0191] In addition, when the logical instructions in the above-mentioned memory 430 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0192] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the adaptive template reconstruction method provided by each of the above methods. The method includes: constructing an initial prompt template based on a small amount of sample data, inputting the initial prompt template into a large language model, and generating a question pool through specific instructions, proxy instructions, and example instructions; using the large language model to answer each question in the question pool to generate an answer pool; screening and clustering the answers in the answer pool through a multi-level similarity comparison module to obtain a screened answer pool and a screened answer group, and dynamically adjusting the weight parameters of similarity calculation in the multi-level similarity comparison module through a model self-feedback mechanism to obtain the overall score of the screened answers; calculating the comprehensive similarity score of each answer group based on the overall score of the answers in each answer group, and conducting an expert evaluation on each answer group. Constructing examples based on the comprehensive similarity score and the expert evaluation, reconstructing the initial prompt template based on the examples to obtain a reconstructed prompt template; inputting the reconstructed prompt template into the large language model for iterative update until a final answer is obtained.
[0193] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the adaptive template reconstruction method provided by each of the above methods. The method includes: constructing an initial prompt template based on a small amount of sample data, inputting the initial prompt template into a large language model, and generating a question pool through specific instructions, proxy instructions, and example instructions; using the large language model to answer each question in the question pool to generate an answer pool; screening and clustering the answers in the answer pool through a multi-level similarity comparison module to obtain a screened answer pool and a screened answer group, and dynamically adjusting the weight parameters of similarity calculation in the multi-level similarity comparison module through a model self-feedback mechanism to obtain the overall score of the screened answers; calculating the comprehensive similarity score of each answer group based on the overall score of the answers in each answer group, and conducting an expert evaluation on each answer group. Constructing examples based on the comprehensive similarity score and the expert evaluation, reconstructing the initial prompt template based on the examples to obtain a reconstructed prompt template; inputting the reconstructed prompt template into the large language model for iterative update until a final answer is obtained.
[0194] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.
[0195] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0196] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An adaptive template reconstruction method based on a small number of samples, characterized in that: The following steps are involved: S1. Construct an initial prompt template based on a small amount of sample data, wherein the initial prompt template includes a task background description, context information, and a part to be completed; S2, inputting the initial prompt template into a large language model, and generating a question pool through specific instructions, agent instructions and example instructions; The generation of the question pool specifically includes: inputting the initial prompt template into the large language model for entity expansion to obtain an entity set, supplementing the initial prompt template with keywords based on the entity set, and adjusting the weights of the keywords. The weight update formula is: ; in, represents the weight value of the fth keyword in round t+1, represents the weight value of the fth keyword in the tth round, represents the adjustment parameter of the correlation evaluation, Indicates keywords The relevance evaluation value with the context Context, represents the fth keyword, Represents the current context. represents the tuning parameter for frequency evaluation, Indicates keywords frequency of occurrence in the text; The expanded keywords are divided into three categories: core keywords, secondary keywords, and background keywords, and are prioritized by differentiated initial weights to obtain a keyword sequence; the initial weights are: ; in, represents the initial weight value of the fth keyword, f represents the index of the keyword, represents the weight adjustment coefficient, represents the fth keyword, represents the priority scoring function; Generate multiple question variants by combining keyword sequences with specific instructions, agent instructions, and example instructions to form an initial question pool; De-duplicate and filter the initial question pool to obtain a filtered question pool; S3. Using the large language model to answer each question in the question pool to generate an answer pool; S4. Screening and clustering the answers in the answer pool through a multi-level similarity comparison module to obtain a screened answer pool and a screened answer group, and dynamically adjusting the weight parameters of the similarity calculation in the multi-level similarity comparison module through a model self-feedback mechanism to obtain an overall score of the screened answers; S5. Calculate the comprehensive similarity score of each answer group based on the overall score of the answers in each answer group, and perform expert scoring on each answer group, construct examples based on the comprehensive similarity score and the expert score, and reconstruct the initial prompt template based on the example to obtain a reconstructed prompt template; S6. Input the reconstructed prompt template into the large language model for iterative updating until a final answer is obtained.
2. The method for adaptive template reconstruction based on a small number of samples according to claim 1, characterized in that: The specific instructions are instructions designed according to the task objectives and used to guide the large language model to generate answers related to the task. The proxy instructions are instructions that guide the large language model to generate answers in an indirect way. The example instructions are generated by providing example questions and answers.
3. The method for adaptive template reconstruction based on a small number of samples according to claim 1, characterized in that: The screened question pool is represented as: ; in, represents the filtered question pool, Indicates the current problem to be evaluated. Indicates except outside Other questions in the question pool, i represents the number of the current question, j represents the number of other questions except i, Q represents the initial question pool, Representation problem and The similarity measure between Represents the similarity threshold.
4. The method for adaptive template reconstruction based on a small number of samples according to claim 1, characterized in that: The processing steps of the multi-level similarity comparison module are: Performing primary similarity calculation on the answers in the answer pool to obtain a primary similarity score, and preliminarily screening the answer pool based on the primary similarity score to obtain a first answer pool; Performing semantic similarity calculation based on deep learning on the answers in the first answer pool to obtain a semantic similarity score, and screening based on the semantic similarity score to obtain a second answer pool; The answers in the second answer pool are clustered based on the semantic similarity scores to form multiple answer groups.
5. The method for adaptive template reconstruction based on a small number of samples according to claim 1, characterized in that: The model self-feedback mechanism includes: The primary similarity score and the semantic similarity score are weightedly combined to obtain an initial similarity score; The initial similarity score is dynamically adjusted based on coverage and specificity to obtain an overall similarity score.
6. The method for adaptive template reconstruction based on a small number of samples according to claim 1, characterized in that: Step S5 specifically includes: Set an overall score threshold. When the overall score of the answers in the answer group is lower than the overall score threshold, the comprehensive scoring mechanism is triggered: Calculate the comprehensive similarity score of each answer group, and perform expert scoring on each answer group, and sort the answer groups from large to small according to the comprehensive similarity score and expert scoring to form an answer group sequence; The first group of answers in the answer group sequence is taken as an example, and the initial prompt template is reconstructed based on the example to obtain a reconstructed prompt template.
7. The method for adaptive template reconstruction based on a small number of samples according to claim 1, characterized in that: The construction of the initial prompt template based on a small amount of sample data specifically includes: Preprocess a small amount of sample data, and add identifiers for clarifying the context scope at the beginning and end of the preprocessed small amount of sample data; Design the initial prompt template based on the task objectives and contextual information; The prompt engineering method combined with the large language model understands the context through prompt words, and dynamically adjusts the structure and vocabulary of the initial prompt template based on the context; wherein the prompt words include a clear description of the task goal and contextual prompts related to the task.
8. An adaptive template reconstruction system based on a small number of samples, characterized in that: Implementing the adaptive template reconstruction method according to any one of claims 1 to 7, comprising: A sample analysis module is used to construct an initial prompt template based on a small amount of sample data, wherein the initial prompt template includes a task background description, context information, and target information prompts that need to be completed; A question generation module, used to input the initial prompt template into a large language model, and generate a question pool through specific instructions, agent instructions and example instructions; An answer generation module, used to answer each question in the question pool using the large language model to generate an answer pool; A similarity screening module, used to screen and cluster the answers in the answer pool through a multi-level similarity comparison module to obtain a screened answer pool and a screened answer group, and calculate the overall score of the answer based on the screened answer group; A template reconstruction module is used to calculate the comprehensive similarity score and expert score of each answer group based on the overall score of the screened answer group, sort the answer groups according to the comprehensive similarity score and expert score, select the answer group with the highest score as an example, optimize and reconstruct the initial prompt template, and generate a reconstructed prompt template; The iterative optimization module is used to input the reconstructed prompt template into the large language model for iterative updating and continuous optimization of the generated answers until the final answer is obtained.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the adaptive template reconstruction method according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the adaptive template reconstruction method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Prompt word template generation method and device, electronic equipment and storage medium
CN119443094A