Soft prompt optimization method and system for large question-answer interaction model based on binary search tree
By adopting the Q&A interactive big model soft prompt optimization method based on binary tree search tree in large AI models, the problem that traditional soft prompt methods are difficult to accurately guide the model to generate text is solved, and more efficient matching search and stronger relevance content generation is achieved.
Patent Information
- Application Number
- CN202410321192.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-20
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-03-20
AI Technical Summary
When generating content, it is difficult for existing large AI models to correctly understand user intentions, and traditional soft prompt methods have problems such as fixed prefix length and low flexibility, making it difficult to provide sufficient information to accurately guide the model to generate text.
The soft prompt optimization method of Q&A interactive large model based on binary tree search tree is adopted. The token embed vector that matches the closest to the original question through depth-first search is added to the soft prompt, and the weight parameters of the soft prompt are updated to generate the soft prompt that is most similar to the input.
It improves matching search efficiency, provides the prompt that the original problem is closest to the embedding space, encourages the model to generate content that is more correlated with the original problem, and significantly improves the quality and accuracy of the generated content.
Smart Images

Figure CN118297151B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer technology and designs a method for optimizing soft prompts in a large model of question-answer interaction based on a binary search tree. The method can effectively select the most matching prompts during the large model training process, thereby achieving better results in content generation tasks. Background Art
[0002] The present invention relates to the technical field of large-scale artificial intelligence (AI) models, and in particular to the technology of using prompts to guide models to generate content. Large-scale AI models have achieved remarkable results in many fields, such as natural language processing, image generation, etc. However, when generating content, how the model correctly understands the user's intention and how to ensure the quality and accuracy of the generated content remains an important challenge.
[0003] Traditional prompt acquisition methods usually include manual design and manual definition. Specifically, designing prompts usually requires domain experts or developers to create them according to the characteristics and requirements of the task. This method may involve a lot of manual effort and expertise, and it is difficult to ensure that the generated prompts can fully guide the model to produce ideal results. Another way to obtain prompts is to select appropriate sentences from existing corpora, data sets or other sources as prompts, in order to guide the model to generate relevant content. However, there are some limitations in the traditional use and acquisition methods of prompts. Traditional fixed or manually designed prompts are time-consuming and labor-intensive, and the input of large models is usually limited by length. This limitation also easily leads to poor results of discrete text prompts. Compared with discrete text prompts, soft prompts can contain more qualitative information and can affect the generation results through changes in parameters, so as to better meet different needs and situations and obtain better results.
[0004] Existing soft hint methods also have some defects. For example, the Prefix Tuning method, which is often used when fine-tuning a large model after pre-training, can also be regarded as a form of soft hint. Its basic principle is to add a continuous and task-specific prefix vector sequence (prefix) before the model input, fix all parameters of the pre-trained model, and only update and optimize the prefix of the specific task. Although Prefix Tuning is an effective method that can be used to fine-tune pre-trained language models to adapt to specific tasks or fields, it also has some defects and limitations, including its fixed prefix length and low flexibility, which may not provide enough information to accurately guide the model to generate text. Summary of the invention
[0005] In order to overcome the above problems, the present invention proposes a soft prompt optimization method and system for a large question-answer interaction model based on a binary search tree, which improves the matching search efficiency.
[0006] The soft prompt optimization method for a large model of question-answer interaction based on a binary search tree designed by the present invention comprises the following steps:
[0007] Step 1: Define a full binary tree structure for storing soft prompts, and randomly initialize or create the embedding vector parameters of each layer in the tree structure based on the embedding vector of the vocabulary;
[0008] Step 2: Encode the questions input into the question-answering interaction model into embedding vectors;
[0009] Step 3: For the processed embedding vector, traverse the vector parameters of each layer in the binary tree, and generate soft prompts based on the vector closest to the sequence according to the depth-first search;
[0010] Step 4: The generated soft prompts are concatenated into the processed embedding vector. The model uses the concatenated input vector to generate content and calculate loss.
[0011] Step 5: Back propagation and optimization to update the parameters of the question-answer interaction model and the soft prompt weights;
[0012] Step 6: Use the trained question-answer interaction model for specific scenarios.
[0013] Preferably, the specific process of step 1 is as follows:
[0014] Step 1.1: Randomly initialize or cluster based on the embedding vectors of the vocabulary;
[0015] Step 1.2: For each type of embedded vector after clustering, calculate the vector average value as the parent node, and then divide it into two categories through clustering. The vector average values of the two categories are calculated as the left and right nodes respectively;
[0016] Step 1.3: Recursively construct left and right subtrees of the two types of embedding vectors after clustering according to step 1.2 until the leaf node is a single token embedding vector;
[0017] Step 1.4: Add these embedding vectors to the network as parameters.
[0018] Preferably, the encoding process into an embedding vector is specifically as follows: encoding the problem input into the model into an embedding vector, performing mean pooling on the input embedding vector, and averaging the embedding vector of each sample into a vector.
[0019] Preferably, the specific process of step 3 is as follows:
[0020] Step 3.1: For the input embedding vector obtained in step 2, traverse the root node of the binary tree, perform L2 regularization on the input and the vectors in the binary tree, then calculate the cosine similarity, and finally find the vector with the largest cosine similarity, record the corresponding vector index, and add the vector to the soft prompt;
[0021] Step 3.2: Based on the vector index of the binary tree recorded in step 3.1, recursively search for the child node with the greatest similarity to the input embedding vector according to the depth-first search, and add its vector to the soft hint until a leaf node is found.
[0022] Preferably, the specific process of step 4 is as follows:
[0023] Step 4.1: Add the soft prompt obtained in step 3 to the input embedding vector to obtain a vector matrix, where k represents the sum of the lengths of the input encoded embedding vector and the embedding vector of the soft prompt;
[0024] Step 4.2: Use the concatenated input vector and the generative part of the model to generate the corresponding content model, compare the generated content with the target content, and calculate the loss value according to the specific training task.
[0025] Preferably, the step 5 specifically comprises the following steps:
[0026] Step 5.1: Through back propagation, the loss value is back propagated to the parameters of the model, and the gradient of each parameter with respect to the loss is calculated;
[0027] Step 5.2: Using the calculated gradient information, use the optimization algorithm to update the parameters of the model, and at the same time update the soft prompt weight parameters added to the network as parameters.
[0028] Based on the same inventive concept, the present invention also designs an electronic device, including:
[0029] one or more processors;
[0030] A storage device for storing one or more programs;
[0031] When one or more programs are executed by the one or more processors, the one or more processors implement a soft prompt optimization method for a large model of question-answer interaction based on a binary tree search tree.
[0032] Based on the same inventive concept, the present invention also designs a computer-readable medium on which a computer program is stored, characterized in that when the program is executed by a processor, a soft prompt optimization method for a large model of question-and-answer interaction based on a binary tree search tree is implemented.
[0033] The advantages of the present invention are:
[0034] This paper proposes an efficient large-model soft prompt optimization method based on binary search tree. Through depth-first search, the token embedding sequence closest to the original question embedding is added to the soft prompt, providing the model with the prompt closest to the original question in the embedding space, encouraging the model to generate content that is more relevant to the original question. First, a binary search tree for searching matching soft prompts is generated using randomly initialized token embedding vectors. Then, during the training process, depth-first search is used to match the token embedding vector closest to the input text embedding vector and add it to the soft prompt. The weight parameters of the soft prompt are updated during the training process, thereby generating the soft prompt that is most similar to the input, encouraging the model to generate content that is more relevant to the original question. Finally, the trained model is used in a specific field.
[0035] The present invention provides the model with the prompt that is closest to the original question in the embedding space. Through the organization of the binary tree and the depth-first search, the matching problem of the prompt and the original question embedding is well solved. Compared with the discrete text prompt, the soft prompt can contain more dense information and obtain better results. Through the guidance of the prompt, the model can be trained for specific fields and can be well applied to various scenarios. The present invention uses multiple real data sets under different types of large models including language large models, visual large models, multimodal models and time series models to conduct experiments on different tasks, including wind prediction tasks under meteorological large models. The experimental results show that this method performs better than traditional prompt methods in these tasks, and has the advantages of fast training speed and short reasoning time, and has broad practical significance and commercial value. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 It is a schematic diagram of the principle of the present invention. DETAILED DESCRIPTION
[0037] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0038] Embodiment 1
[0039] The present invention proposes a method for optimizing soft prompts in a large model of question-answer interaction based on a binary search tree. First, a binary search tree for searching matching soft prompts is generated using randomly initialized token embedding vectors, and then a depth-first search is used during the training process to match the token embedding vector closest to the input text embedding vector and add it to the soft prompt, and the weight parameter of the soft prompt is updated, and finally the trained model is used in a specific field.
[0040] The technical solution adopted by the present invention is: a method for optimizing the high-efficiency large model soft prompt based on a binary search tree, characterized in that it comprises the following steps:
[0041] Step 1: Define a binary tree structure for storing soft prompts, randomly initialize or create the embedding vector parameters of each layer in the tree structure based on the embedding vector of the vocabulary, and these parameters will be optimized during model training; the specific implementation of step 1 includes the following sub-steps:
[0042] Step 1.1: Initialize the embedding vector e randomly or according to the vocabulary w ∈R d (where d represents the spatial size of the vector) for clustering;
[0043] Step 1.2: For each type of embedded vector after clustering, calculate the vector average value as the parent node, and then divide it into two categories through clustering. The vector average values of the two categories are calculated as the left and right nodes respectively;
[0044] Step 1.3: Recursively construct left and right subtrees of the two types of embedding vectors after clustering according to step 1.2 until the leaf node is a single token embedding vector.
[0045] Step 1.4: Add these embedding vectors to the network as parameters, which are optimized during the training of the model.
[0046] Step 2: Encode the question input into the model into an embedding vector, perform mean pooling on all the encoded embedding vectors, and obtain an embedding vector that contains the input information.
[0047] Step 2.1: Encode the question input to the model into an embedding vector matrix M u ∈R I·d , where l represents the length of the input text;
[0048] Step 2.2: Perform mean pooling on the input embedding vector matrix to average the vector matrix of each sample into a vector e q ∈R d .
[0049] Step 3: For the input embedding vector obtained in step 2, traverse the vector parameters of each layer in the binary tree, search the vector closest to its sequence according to the depth-first search, concatenate the embedding vectors stored in each matching intermediate node, and gradually generate soft prompts as the traversal from top to bottom; the specific process is as follows:
[0050] Step 3.1: Embedding vector e of the input obtained in step 2 q ∈R d, traverse the root node of the binary tree, perform L2 regularization on the input and the vectors in the binary tree, then calculate the cosine similarity, and finally find the vector with the largest cosine similarity, record the corresponding vector index, and add the vector to the soft prompt;
[0051] Step 3.2: According to the vector index of the binary tree recorded in step 3.1, recursively search for the child node with the greatest similarity to the input embedding vector based on depth-first search, and add its vector to the soft prompt until a leaf node is found.
[0052] Step 4: The soft hint obtained in step 3 is concatenated into the input embedding vector. The model uses the concatenated input vector to generate content and calculate loss.
[0053] Step 4.1: Add the soft hint obtained in step 3 to the input embedding vector to obtain the vector matrix M u ∈R K·d , where k represents the sum of the lengths of the encoded embedding vector of the input and the embedding vector of the soft prompt;
[0054] Step 4.2: Use the concatenated input vector to generate the corresponding content model using the generative part of the model, compare the generated content with the target content, and calculate the loss value based on the specific training task;
[0055] Step 5: Back propagation and optimization, update model parameters and soft prompt weights, including:
[0056] Step 5.1: Through back propagation, the loss value is back propagated to the parameters of the model, and the gradient of each parameter with respect to the loss is calculated;
[0057] Step 5.2: Using the calculated gradient information, select an adaptive optimization algorithm to update the model parameters according to different application scenarios and model structures, and update the soft prompt weight parameters added to the network as parameters.
[0058] Step 6: Use the trained model for a specific scenario; given an input, train according to steps 1-5, and the trained model can be used in a specific scenario. The soft prompts searched in step 3 contain information related to the input. The embedding vectors used to form the soft prompts are stored in a binary tree. The number of layers in the binary tree determines the search range of the soft prompts. The larger the number of layers, the more embedding vectors the vector library stores, the more information it has, and it can provide richer information for more types of inputs. It can be set according to the application scenario of the large model and its different parameter magnitudes.
[0059] The method for optimizing soft prompts for a large model of question-answer interaction based on a binary search tree proposed in the present invention has been tested under multiple large models (such as L l ama-2). In the large model question-answer scenario, the input is the user's question or chat content, and the output is the text generated by the model. The experimental results show that based on the soft prompt method, the perplexity of the large model is reduced by 19.78%, and the overall score is improved by 9.86%. In addition, in five scenarios of open question and answer, induction and summary, information extraction, mathematics, and code, the large model can output more accurate answers, achieving 12.05%, 8.06%, 7.50%, 14.28%, and 7.85% improvement respectively. The experimental results show that under the method based on soft prompts, the large model has achieved significant performance improvement in multiple field tasks such as question and answer, summary, mathematics, and code.
[0060] With the widespread application of large models, the method for optimizing soft prompts for large question-answer interaction models based on binary tree search tree proposed in the present invention can adapt to tasks in more fields, and as a general soft prompt optimization method, this method can be applied to more large models to help optimize the performance of large models.
[0061] Embodiment 2
[0062] Based on the same inventive concept, the present invention also provides an electronic device, including one or more processors; a storage device for storing one or more programs; when one or more programs are executed by the one or more processors, the one or more processors implement the method described in Example 1.
[0063] Since the device introduced in the third embodiment of the present invention is an electronic device used to implement the method for optimizing soft prompts of a large model of question-answer interaction based on a binary search tree in the first embodiment of the present invention, a person skilled in the art can understand the specific structure and deformation of the electronic device based on the method introduced in the first embodiment of the present invention, so it is not repeated here. All electronic devices used in a method of the embodiment of the present invention belong to the scope of protection of the present invention.
[0064] Embodiment 3
[0065] Based on the same inventive concept, the present invention further provides a computer-readable medium on which a computer program is stored, and when the program is executed by a processor, the method described in the first embodiment is implemented.
[0066] Since the device introduced in the third embodiment of the present invention is a computer-readable medium used to implement the method for optimizing soft prompts of a large model of question-answer interaction based on a binary search tree in the first embodiment of the present invention, a person skilled in the art can understand the specific structure and deformation of the electronic device based on the method introduced in the first embodiment of the present invention, so it is not repeated here. All electronic devices used in a method of the embodiment of the present invention belong to the scope of protection of the present invention.
[0067] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A soft prompt optimization method for a large model of question-answer interaction based on a binary search tree, characterized in that: The following steps are involved: Step 1: Define a full binary tree structure for storing soft prompts, and randomly initialize or create the embedding vector parameters of each layer in the tree structure based on the embedding vector of the vocabulary; Step 2: Encode the questions input into the question-answering interaction model into embedding vectors; Step 3: For the processed embedding vector, traverse the vector parameters of each layer in the binary tree, search the vector closest to the matching sequence according to the depth-first search, and generate soft prompts; Step 4: The generated soft prompts are concatenated into the processed embedding vector. The model uses the concatenated input vector to generate content and calculate loss. Step 5: Through back propagation, the loss value is back-propagated to the parameters of the model, and the gradient of each parameter with respect to the loss is calculated; using the calculated gradient information, an optimization algorithm is used to update the parameters of the model, and at the same time, the soft prompt weight parameters added as parameters to the network are updated; Step 6: Use the trained question-answering interaction model in question-answering interaction scenarios.
2. The soft prompt optimization method for a large model of question-answer interaction based on a binary search tree according to claim 1 is characterized in that: The specific process of step 1 is as follows: Step 1.1: Initialize embedding vectors randomly or from vocabulary e w ∈ R d Clustering is performed, where d Represents the spatial size of the vector; Step 1.2: For each type of embedded vector after clustering, calculate the vector average as the parent node, and then divide it into two categories through clustering. The vector averages of the two categories are calculated as the left and right nodes respectively; Step 1.3: Recursively construct left and right subtrees of the two types of embedding vectors after clustering according to step 1.2 until the leaf node is a single token embedding vector; Step 1.4: Add these embedding vectors to the network as parameters.
3. The soft prompt optimization method for a large model of question-answer interaction based on a binary search tree according to claim 1 is characterized in that: The encoding process into an embedding vector is specifically as follows: the problem input into the model is encoded into an embedding vector, the input embedding vector is mean pooled, and the embedding vector of each sample is averaged into one vector.
4. The soft prompt optimization method for a large model of question-answer interaction based on a binary search tree according to claim 1, characterized in that: The specific process of step 3 is as follows: Step 3.1: Embedding vector of the input obtained in step 2 e q ∈ R d , traverse the root node of the binary tree, perform L2 regularization on the input and the vectors in the binary tree, then calculate the cosine similarity, and finally find the vector with the largest cosine similarity, record the corresponding vector index, and add the vector index to the soft prompt; Step 3.2: Based on the vector index of the binary tree recorded in step 3.1, recursively search for the child node with the greatest similarity to the input embedding vector according to the depth-first search, and add its vector to the soft prompt until a leaf node is found.
5. The soft prompt optimization method for a large model of question-answer interaction based on a binary search tree according to claim 1, characterized in that: The specific process of step 4 is as follows: Step 4.1: Add the soft hint obtained in step 3 to the input embedding vector to get the vector matrix M u ∈R K·d , where k represents the sum of the lengths of the encoded embedding vector of the input and the embedding vector of the soft prompt; Step 4.2: Use the concatenated input vector and the generative part of the model to generate the corresponding content model, compare the generated content with the target content, and calculate the loss value according to the specific training task.
6. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs; When one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 5.
7. A computer readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Vector retrieval type dialogue method of common method question-answering system
CN113742471A
Data processing method and device, equipment and storage medium
CN117520523A