Large language model discrete cue word searching method and device

By introducing preset parameter matrix update and elite coordinate descent algorithms into the large language model, combined with the preset generative language model, the problem of large overhead of discrete prompt word search calculation is solved. The generated prompt words have better semantic rationality and interpretability, and the inference accuracy and efficiency of the model are improved.

CN120296148APending Publication Date: 2025-07-11SUN YAT SEN UNIV
View PDF 0 Cites 10 Cited by

Patent Information

Application Number
CN202510485327.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing discrete prompt word search technology leads to excessive computational overhead for prompt word optimization, and the generated prompt words lack semantic coherence and interpretability, making it difficult to efficiently optimize large language models in practical applications.

Method used

The preset parameter matrix update function and elite coordinate descent algorithm are used to combine the preset generative language model to generate the target strategy model through iterative optimization, and the target discrete prompt words are generated. The knowledge of the pre-trained language model is used to avoid retraining the policy network, reduce computing overhead and improve semantic generation capabilities.

Benefits of technology

It realizes efficient search of discrete prompt words, reduces the consumption of computing resources, and the generated prompt words have better semantic rationality and interpretability, improving the reasoning accuracy and generalization ability of large language models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296148A_ABST
    Figure CN120296148A_ABST
Patent Text Reader

Abstract

The invention discloses a large language model discrete cue word search method and device, and aims to solve the technical problem of overhigh cue word optimization calculation overhead caused by an existing discrete cue word search technology. The method comprises the steps of obtaining training prompt words and training sentences, updating an initial parameter matrix of an initial strategy model by adopting a preset parameter matrix updating function according to the training prompt words, and determining an intermediate strategy model; reasoning according to the training sentence by adopting an intermediate strategy model and a preset generative language model to generate a plurality of disturbance discrete cues and a plurality of reasoning results; based on a preset gradient calculation formula, performing iterative optimization on the intermediate parameter matrix of the intermediate strategy model by adopting an elite coordinate descent algorithm according to the plurality of disturbance discrete cues and the plurality of reasoning results, and determining a target strategy model; and generating a target discrete cue word based on the target strategy model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular to a method and device for searching discrete prompt words in a large language model. Background Art

[0002] In the actual application of Large Language Model (LLM), parameter fine-tuning is the mainstream means to improve the performance of downstream tasks. However, although the traditional model fine-tuning method is effective, its dependence on computing resources and gradient information makes it face great obstacles when applied to closed-source large models; at the same time, traditional model fine-tuning usually requires a large amount of memory resources and relies on the gradient information of model parameters, which greatly limits many practical application scenarios. However, prompt word search, especially the discrete prompt word method, has the advantages of low computational overhead and can work in a black box environment.

[0003] The prompt word search technology (also known as "prompt word fine-tuning") provides a lightweight method that does not require modifying model parameters. It guides the model to generate more accurate output text by inserting a specific text string (called a "prompt word") before the input expectation. Depending on whether the prompt text is continuous or discrete, the prompt word search is divided into continuous prompt word (or soft prompt word) search and discrete prompt word (or hard prompt word) search.

[0004] Compared with continuous prompt words, discrete prompt words are not only more interpretable, but also can achieve efficient prompt word search through optimization algorithms such as reinforcement learning, help freeze the parameters of large language models, and only optimize the prompt words, thus achieving an effect close to or even comparable to fine-tuning the full model. In addition, discrete prompt words have shown extensive potential in many fields such as few-shot learning and cross-task generalization, making them a hot research direction for large language model optimization.

[0005] Existing discrete prompt word search technology is mainly based on reinforcement learning-based prompt optimization (RLPrompt). The core idea of ​​RLPrompt is to transform the discrete prompt optimization problem into a reinforcement learning task, and use a parameter-efficient policy network to generate optimized prompt words. However, this method requires training the policy network from scratch for each optimization, and is highly dependent on the retraining of the policy network. At the same time, the complexity of the reinforcement learning mechanism itself (including the need for carefully designed reward functions, complex policy gradient calculations, and time-consuming trial-and-error training processes) requires intensive computing operations to be performed in each training cycle, resulting in excessive computational overhead for prompt word optimization. Summary of the invention

[0006] The present invention provides a method and device for discrete prompt search in a large language model, which is used to solve the technical problem that the existing discrete prompt search technology leads to excessive computational overhead in prompt optimization.

[0007] A method for discrete prompt search in a large language model provided by the first aspect of the present invention includes:

[0008] Obtain training prompts and training sentences, and update the initial parameter matrix of the initial policy model according to the training prompts by using a preset parameter matrix update function to determine an intermediate policy model;

[0009] Use the intermediate policy model and a preset generative language model to perform inference according to the training sentences to generate multiple perturbed discrete prompts and multiple inference results;

[0010] Based on a preset gradient calculation formula, use the elite coordinate descent algorithm to iteratively optimize the intermediate parameter matrix of the intermediate policy model according to the multiple perturbed discrete prompts and the multiple inference results to determine a target policy model;

[0011] Generate target discrete prompts based on the target policy model.

[0012] Optionally, the preset generative language model includes a noise model, a text-to-text transfer transformer, and a large language model; the using the intermediate policy model and the preset generative language model to perform inference according to the training sentences to generate multiple perturbed discrete prompts and multiple inference results includes:

[0013] Generate multiple intermediate discrete prompts based on the intermediate policy model;

[0014] Use the noise model to perturb each of the intermediate discrete prompts to generate a perturbed discrete prompt corresponding to each of the intermediate discrete prompts;

[0015] Respectively use each of the perturbed discrete prompts as the input of the text-to-text transfer transformer, and output a training semantic sentence corresponding to each of the perturbed discrete prompts;

[0016] Concatenate each of the training semantic sentences with the training sentence to generate a training concatenated sentence corresponding to each of the training semantic sentences;

[0017] Respectively input each of the training concatenated sentences into the large language model, and output a training inference result corresponding to each of the training concatenated sentences.

[0018] Optionally, the preset gradient calculation formula includes a parameter matrix gradient calculation formula and a policy model gradient calculation formula; based on the preset gradient calculation formula, using the elite coordinate descent algorithm to iteratively optimize the intermediate parameter matrix of the intermediate policy model according to multiple perturbation discrete prompts and multiple inference results to determine the target policy model, including:

[0019] Calculate the loss value corresponding to each inference result according to each inference result;

[0020] Based on each loss value, determine the objective function value corresponding to each loss value;

[0021] Sort each objective function value in ascending order to determine the ranking rank corresponding to each objective function value;

[0022] Normalize each ranking rank to determine the weight value corresponding to each objective function value;

[0023] Perform one-hot encoding on each perturbation discrete prompt to determine the discrete prompt encoding corresponding to each perturbation discrete prompt;

[0024] Use the policy model gradient calculation formula to calculate the policy model gradient corresponding to each perturbation discrete prompt according to each discrete prompt encoding and the intermediate parameter matrix of the intermediate policy model;

[0025] Use the parameter matrix gradient calculation formula to calculate the parameter matrix gradient according to each weight value and each policy model gradient;

[0026] Update the intermediate parameter matrix of the intermediate policy model according to the parameter matrix gradient to determine the updated policy model, and count the update times in real time;

[0027] Perform a mean operation on the loss values corresponding to each inference result to determine the loss mean;

[0028] Judge whether the update times reach the preset update threshold and whether the loss mean converges;

[0029] If the update times reach the preset update threshold or the loss mean converges, use the updated policy model as the target policy model.

[0030] Optionally, it further includes:

[0031] If the update times do not reach the preset update threshold and the loss mean does not converge, use the updated policy model as the new initial policy model;

[0032] Use the preset parameter matrix update function to update the initial parameter matrix of the new initial policy model according to the training prompt words, and determine the intermediate policy model;

[0033] Use the intermediate policy model and the preset generative language model to perform inference according to the training sentence, and generate multiple perturbed discrete prompt words and multiple inference results;

[0034] Based on the preset gradient calculation formula, use the elite coordinate descent algorithm to update the intermediate parameter matrix of the intermediate policy model according to multiple perturbed discrete prompt words and multiple inference results, determine the loss mean and the updated policy model, and count the update times in real time until the update times reach the preset update threshold or the loss mean converges;

[0035] Use the updated policy model determined when the update times reach the preset update threshold or the loss mean converges as the target policy model.

[0036] Optionally, after the step of generating the target discrete prompt word based on the target policy model, it includes:

[0037] When receiving the sentence to be tested, use the target discrete prompt word as the input of the text-to-text transfer transformer, and output the target semantic sentence;

[0038] Concatenate the target semantic sentence and the sentence to be tested to generate a target concatenated sentence;

[0039] Input the target concatenated sentence into the large language model, and output the target inference result.

[0040] Optionally, the parameter matrix gradient calculation formula in the preset gradient calculation formula is specifically:

[0041] ;

[0042] Where g is the parameter matrix gradient; N is the total number of perturbed discrete prompt words; is the i-th weight value; is the policy model gradient of the logarithm probability corresponding to the i-th perturbed discrete prompt word with respect to the k-th column of the parameter matrix of the policy model; is the i-th perturbed discrete prompt word.

[0043] A large language model discrete prompt word search device provided by the second aspect of the present invention includes:

[0044] An acquisition module, configured to acquire training prompt words and training sentences, and use a preset parameter matrix update function to update the initial parameter matrix of the initial policy model according to the training prompt words, and determine the intermediate policy model;

[0045] An inference module, configured to perform inference according to the training sentences by using the intermediate policy model and a preset generative language model, and generate a plurality of perturbed discrete prompt words and a plurality of inference results;

[0046] An optimization module, configured to iteratively optimize the intermediate parameter matrix of the intermediate policy model based on a preset gradient calculation formula by using an elite coordinate descent algorithm according to the plurality of perturbed discrete prompt words and the plurality of inference results, and determine a target policy model;

[0047] A generation module, configured to generate target discrete prompt words based on the target policy model.

[0048] A computer device provided in the third aspect of the present invention includes a memory and a processor. A computer program is stored in the memory. When the computer program is executed by the processor, the processor is caused to execute the steps of the large language model discrete prompt word search method as described in any one of the above.

[0049] A computer-readable storage medium provided in the fourth aspect of the present invention has a computer program stored thereon. When the computer program is executed, the steps of the large language model discrete prompt word search method as described in any one of the above are implemented.

[0050] A computer program product provided in the fifth aspect of the present invention includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer is caused to execute the steps of the large language model discrete prompt word search method as described in any one of the above.

[0051] It can be seen from the above technical solutions that the present invention has the following advantages:

[0052] The above technical solution of the present invention provides a method for searching discrete prompting words of a large language model. First, training prompting words and training sentences are obtained, and the initial parameter matrix of the initial policy model is updated according to the training prompting words by using a preset parameter matrix update function to determine an intermediate policy model. Then, the intermediate policy model and a preset generative language model are used to perform inference based on the training sentences to generate multiple perturbed discrete prompting words and multiple inference results. Based on a preset gradient calculation formula, the elite coordinate descent algorithm is used to iteratively optimize the intermediate parameter matrix of the intermediate policy model according to the multiple perturbed discrete prompting words and multiple inference results to determine a target policy model. Finally, based on the target policy model, target discrete prompting words are generated. Based on the above solution, the initial policy model is updated by using the preset parameter matrix update function to obtain an intermediate policy model, and then based on the elite coordinate descent algorithm, combined with the intermediate policy model, the preset generative language model, and the preset gradient calculation formula, iterative optimization of the parameter matrix is performed according to the training inference results to determine the target policy model. Finally, in the process of generating target discrete prompting words based on the target policy model, the present invention can achieve efficient search through the elite coordinate descent method and make full use of the existing knowledge of the pre-trained language model, without having to train the policy network from scratch every time for optimization. In practical applications, the target policy model can be directly used to generate target discrete prompting words, thereby reducing the computational overhead of prompting word optimization. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0054] Figure 1 It is a flowchart of the steps of a method for searching discrete prompting words of a large language model provided in Embodiment 1 of the present invention;

[0055] Figure 2 It is a schematic diagram of the working principle of a method for searching discrete prompting words of a large language model provided in Embodiment 2 of the present invention compared with the prior art;

[0056] Figure 3 It is a block diagram of the structure of a device for searching discrete prompting words of a large language model provided in Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] Embodiments of the present invention provide a method and device for searching discrete prompting words of a large language model, which are used to solve the technical problem that the existing discrete prompting word search technology leads to too large computational overhead for prompting word optimization.

[0058] To make the objectives, features, and advantages of the present invention more apparent and understandable, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the embodiments described below are only a part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0059] Please refer to Figure 1 , Figure 1 which is a flowchart of the steps of a discrete prompt word search method for a large language model provided in Embodiment 1 of the present invention.

[0060] A discrete prompt word search method for a large language model provided by the present invention includes:

[0061] Step 101: Obtain training prompt words and training sentences, and update the initial parameter matrix of the initial policy model according to the training prompt words using a preset parameter matrix update function to determine an intermediate policy model.

[0062] It should be noted that the training prompt words of the present invention are randomly generated, and a preset parameter matrix update function is introduced , whose role is to generate a new parameter matrix from the solution y, the coordinate k, and the parameter matrix , that is, taking the training prompt word a as the current solution y, randomly extracting an index k according to the uniform distribution corresponding to the current solution y using the preset parameter matrix update function, and this function will retain the k-th column of , and adjust the other columns according to the solution y for the parameter matrix to obtain a new parameter matrix , and further obtain an intermediate policy model.

[0063] Further, the preset parameter matrix update function is specifically:

[0064] ;

[0065] wherein, performs a random perturbation on the k-th row of the parameter matrix; is the preset parameter matrix update function; is the matrix element at the i-th row and j-th column of the parameter matrix; is the j-th token in the current solution; y is the current solution; k is the index; is the parameter matrix.

[0066] Exemplarily, please refer to Table 1. For a given coordinate k, the generated by the function R retains All values in the k-th column, while setting other columns according to the solution y. If , then the corresponding is set to , otherwise it is set to . The idea of this strategy is that when it can be determined that the k-th column (or coordinate) is the elite solution, the parameters of this part are retained, while other columns are adjusted according to the current solution y. After the above operations, the distribution degenerates at coordinates other than the k-th coordinate. That is, other coordinates are fixed, and only the k-th coordinate participates in the optimization.

[0067] Table 1 Algorithm framework for updating the parameter matrix using the preset parameter matrix update function

[0068]

[0069] Furthermore, the initial strategy model adopted by the present invention includes two optional strategy models, namely the Softmax strategy model and the hadamard strategy model; among them, the Softmax strategy model constructs a probability model using the classic Softmax method:

[0070] ;

[0071] Among them, is called the policy, n is the length of the discrete prompt word; x is the discrete prompt word; i is the i-th row in the parameter matrix; is the j-th token in the discrete prompt word; is the probability that the j-th token of x takes the value i under the given ; is the matrix element in the i-th row and j-th column of the parameter matrix of the strategy model; m is the number of the vocabulary of natural words.

[0072] The advantage of the Softmax strategy model is that it is relatively simple, has no additional hyperparameters, and has very good theoretical performance, which can guarantee the optimality of the search algorithm. Its possible performance bottleneck lies in the introduction of the exponential function, which will introduce pathological features in the function landscape, so it has high requirements for the ability of the search algorithm to handle irregular landscape features.

[0073] Furthermore, the hadamard strategy model is different from the Softmax strategy model. The hadamard strategy model has the following probability distribution:

[0074] ;

[0075] Among them, is a hyperparameter. This strategy is called the Hadamard strategy because for a given , the probability forms a vector and then there is , is the Hadamard product, and T is the transpose; are all elements of the j-th column in the parameter matrix. The Hadamard strategy model proposed by the present invention is significantly different from the existing Hadamard reparameterization method: the traditional reparameterization method still requires normalization changes, while the present invention does not require this step.

[0076] It is worth mentioning that for the probability modeling process, those skilled in the art can use Gumbel-Softmax reparameterization to replace Softmax or the Hadamard strategy, solve the discrete space optimization problem through differentiable sampling, and retain the gradient information. Or use a reinforcement learning policy network: use the Actor-Critic framework to replace the policy gradient method, evaluate the quality of the prompt words through the value function, reduce the variance and accelerate the convergence.

[0077] Step 102: Use the intermediate policy model and the pre-set generative language model to perform inference according to the training sentences, and generate multiple perturbed discrete prompt words and multiple inference results.

[0078] The pre-set generative language model includes a noise model, a text-to-text transfer transformer, and a large language model.

[0079] Specifically, step 102 may include the following sub-steps S21-S22:

[0080] Step S21: Generate multiple intermediate discrete prompt words based on the intermediate policy model;

[0081] Step S22: Use the noise model to perturb each intermediate discrete prompt word respectively to generate the perturbed discrete prompt words corresponding to each intermediate discrete prompt word;

[0082] It should be noted that in order to implement the noise model, the present invention adopts a structured perturbation mechanism, aiming to balance the depth search and breadth search in the discrete prompt optimization process. The noise model introduces controllable randomness in the discrete prompt word , while maintaining its structural integrity. Specifically, for the generated discrete prompt word x, the generated noise variable is as follows:

[0083] ;

[0084] where y is the perturbed discrete prompt word; is the noise model; is the i-th token in the perturbed discrete prompt; is the i-th token in the discrete prompt; is a randomly selected integer in each iteration, indicating that the k-th element is randomly selected to apply noise perturbation, , is a uniform distribution on; is the length of the discrete prompt; z is a value randomly selected from a uniform distribution in m dimensions. When i = k, = = , indicating that the k-th component of y changes from the original value to random noise, that is, , is a uniform distribution on, where m is the number of the vocabulary of natural words.

[0085] Furthermore, this noise injection mechanism ensures that the perturbation is local, allowing the algorithm to explore the changes of individual tokens without destroying the entire prompt structure. The design without hyperparameters simplifies the implementation, and the unified sampling of z guarantees unbiased exploration across the vocabulary.

[0086] It is worth mentioning that the noise model is applied in the candidate generation process of the self-elitist coordinate descent (ECD, Conclusion of Elitist Descent Framework) framework . For example, in ECD, the solution x (i.e., the discrete prompt) obtained from is injected with noise through to generate a perturbed candidate solution y (i.e., the perturbed discrete prompt).

[0087] This implementation ensures that the noise model systematically diversifies the search space while maintaining consistency with the policy-based optimization theoretical framework. By decoupling the depth search ( ) from the breadth search (elitist preservation and likelihood update), the algorithm achieves robust performance in the high-dimensional discrete domain.

[0088] Step S23: Respectively take each perturbed discrete prompt as the input of the text-to-text transfer transformer, and output the training semantic sentence corresponding to each perturbed discrete prompt;

[0089] Step S24: Respectively splice each training semantic sentence with the training sentence to generate the training spliced sentence corresponding to each training semantic sentence;

[0090] Step S25: Respectively input each training spliced sentence into the large language model, and output the training inference result corresponding to each training spliced sentence.

[0091] It should be noted that in order to introduce the semantic generation function into the algorithm framework of the present invention, the present invention first introduces the text-to-text transfer transformer, namely the T5 model (Text-to-Text Transfer Transformer). The T5 model was proposed by Colin Raffel et al. This model adopts a unified text-to-text framework, which converts all natural language processing tasks (such as text summarization, question answering, text classification, etc.) into the form of generating text. The T5 model achieves state-of-the-art performance on multiple benchmark tasks by pre-training on a large-scale dataset and then fine-tuning on specific tasks. This model utilizes the powerful ability of transfer learning and combines with the newly proposed large-scale dataset "Colossal CleanCrawled Corpus" to provide a powerful general model for the field of natural language processing.

[0092] Furthermore, in the framework of the present invention, the present invention introduces the T5 model to improve the semantic rationality and interpretability of the prompt words, especially to convert the discrete prompt words into fluent sentences that conform to natural language grammar. Through this process, the generated prompt words can not only improve the reasoning ability of the model but also increase the semantic consistency with the task objective.

[0093] The following is an example of introducing the T5 model for semantic generation as shown in Table 2:

[0094] Table 2 Example of Introducing the T5 Model for Semantic Generation

[0095]

[0096] From this example, it can be seen that the discrete prompt words generated in the framework of the present invention (such as "treeplant ground hole dig" in the example) can be directly used as the input of the T5 model. Through the generation process of the T5 model, the discrete prompt words are converted into natural language sentences with grammatical rationality and semantic clarity (such as "digging a hole in the ground to plant trees"). Therefore, the present invention only needs to make minor modifications in the model inference process, first taking the obtained discrete prompt word T as the input of the T5 model, after generating the semantic sentence, then splicing it with the obtained sentence S, and then sending it to the large language model for inference.

[0097] It is worth mentioning that for the semantic generation module, those skilled in the art can use a semantic correction based on BERT or GPT to replace the T5 model, and use other pre-trained language models (such as the masked filling mechanism of BERT or the generation ability of GPT) to convert discrete prompt words into natural language to enhance semantic coherence. Or adopt a rule-driven semantic template: perform structured correction on discrete prompt words through artificially designed grammar rules or template libraries (such as regular expressions) to avoid relying on pre-trained models.

[0098] Step 103: Based on the preset gradient calculation formula, use the elite coordinate descent algorithm to iteratively optimize the intermediate parameter matrix of the intermediate policy model according to multiple perturbed discrete prompt words and multiple inference results, and determine the target policy model.

[0099] The preset gradient calculation formula includes a parameter matrix gradient calculation formula and a policy model gradient calculation formula.

[0100] It should be noted that assuming the length of the discrete prompt word is n and it comes from a vocabulary with m natural words, the feasible space is . The original discrete prompt word search problem has the following definition:

[0101] ;

[0102] Among them, is the training sample; s is the input text; l is the target output; M is the large model; is to concatenate the input text with the prompt word; is the distribution of the training data; is the loss value between the output value of the large model at the given input and the target output l; is to find the mathematical expectation under the condition of sampling the training sample from D.

[0103] The present invention uses a discrete probability distribution to guide the search. Specifically, let be a real matrix, be a discrete probability distribution with parameter . At the same time, introduce a noise probability distribution to represent the fixed environmental noise applied to the new prompt word y generated after being perturbed by the noise model. Based on the above probability model, the objective function of the discrete prompt word problem is transformed into a surrogate problem:

[0104] ;

[0105] Among them, is the deterministic loss function obtained by calculating the average value of the original loss function F(y) at N sampling points, ; To find the mathematical expectation under the condition of sampling from ; To find the mathematical expectation under the condition of sampling from ; To find the mathematical expectation under the condition of sampling from ; Is the surrogate problem (surrogate function).

[0106] Thus, it can be seen that the present invention transforms the discrete search problem in the space into a continuous search problem in the continuous unconstrained real space .

[0107] Furthermore, the original stochastic optimization problem is defined as:

[0108] ;

[0109] Wherein, is to find the mathematical expectation under the condition of sampling from D ; represents the loss of the language model evaluated on the perturbed discrete prompt y and the data sample . To derive a deterministic objective, the present invention considers the special case where the distribution D is degenerate, i.e., it assigns probability 1 to a single fixed sample . This reduces the expectation to a deterministic evaluation:

[0110] ;

[0111] Wherein, is the loss of the language model evaluated on the perturbed discrete prompt y and the data sample .

[0112] In an actual scenario, for example, when the training data set is fixed, the deterministic objective can also be interpreted as the empirical risk of the entire data set:

[0113] ;

[0114] Wherein, are the pre-sampled and fixed data points.

[0115] Based on the above, the deterministic optimization surrogate function: In this deterministic setting, the surrogate function simplifies to:

[0116] ;

[0117] Wherein, is the original loss function, which calculates a loss for each sampling point and does not have stability, no longer depends on random samples This is consistent with the classical deterministic optimization framework, where the gradient can be directly calculated without the need for Monte Carlo approximation.

[0118] Furthermore, the impact on gradient estimation: The gradient of is:

[0119] ;

[0120] where is the log-likelihood gradient of policy ; is the gradient of ; this gradient eliminates the need for sampling . This simplification enhances computational stability and allows for deterministic policy updates in the ECD algorithm.

[0121] Furthermore, let e i be the vector with the i-th coordinate equal to 1 and other positions equal to 0. It follows that , is the mathematical expectation under the condition that k follows an n-dimensional uniform distribution , is the vector with the k-th coordinate equal to 1 and other positions equal to 0, is the n-dimensional all-ones vector. To use the gradient descent method, the present invention needs to discuss the gradient of the surrogate function . Therefore, the present invention can give the calculation process:

[0122] ;

[0123] ;

[0124] ;

[0125] ;

[0126] ;

[0127] ;

[0128] ;

[0129] ;

[0130] ;

[0131] where is the policy model; is the gradient of ; is the logarithmic likelihood gradient of the policy model; is the policy model after randomly perturbing the k-th row; is to calculate the mathematical expectation under the condition that x is sampled from the policy model after randomly perturbing the k-th row; is for the perturbed discrete prompt word y and the data sample is the language model loss evaluated on.

[0132] Furthermore, the present invention takes into account the pre-set parameter matrix update function That is:

[0133] ;

[0134] Furthermore, for a given metric and solution y, , keep the k-th column of, while all other columns are set according to y. Thus:

[0135] ;

[0136] ;

[0137] Among them, is the element in the i-th row and j-th column of the result of the sigma function based on the parameter , is the result of the sigma function based on the parameter ; is the result of the sigma function based on the parameter ; is the power, indicating to judge whether the j-th component value of the i-th prompt word x is exactly the row label l. If it holds, take 1, otherwise take 0; is the element in the j-th column selected from the result of the sigma function based on the parameter according to the value of the j-th component of the j-th prompt word x; is the element in the k-th column selected from the result of the sigma function based on the parameter according to the value of the k-th component of the k-th prompt word x; is the element in the j-th column selected from the result of the sigma function based on the parameter according to the value of the j-th component of the j-th prompt word x; is the element in the k-th column selected from the result of the sigma function based on the parameter according to the value of the k-th component of the k-th prompt word x.

[0138] Therefore, the gradient can be calculated as follows:

[0139] ;

[0140] However, the exact expression of the above gradient is difficult to directly implement in actual calculations. Therefore, the present invention adopts an approximation method based on multi-sample comparison, which includes two steps: The first step is sample generation, that is, sampling multiple prompt words from the probability distribution , and generating corresponding through the noise model ; The second step is gradient estimation, that is, based on these samples, by comparing the values of the loss , select samples with better performance for gradient estimation.

[0141] Therefore, the present invention has:

[0142] ;

[0143] where is the policy model after adding noise.

[0144] Furthermore, for the given parameter matrix , the following sampling is performed: , , , and then the gradient is defined as:

[0145] ;

[0146] where g is the gradient of the parameter matrix; N is the total number of perturbed discrete prompt words; is the probability value corresponding to the discrete prompt word based on the output of the policy model; is the probability value corresponding to the discrete prompt word based on the output of the policy model after adding noise; is the gradient of the policy model corresponding to the i-th perturbed discrete prompt word; is the objective function value corresponding to the i-th perturbed discrete prompt word; is a vector with the k-th coordinate being 1 and other positions being 0.

[0147] Based on the above, the present invention has , and the proof is as follows:

[0148] ;

[0149] where is to find the expectation.

[0150] Because , therefore:

[0151] ;

[0152] Expected value of the combined standard basis vectors There is . The above fact means that the gradient estimator g is an unbiased estimator of

[0153] In summary, the calculation formula for the parameter matrix gradient is specifically:

[0154] ;

[0155] where g is the parameter matrix gradient; N is the total number of perturbed discrete prompts; is the i-th weight value; is the policy model gradient of the logarithm probability corresponding to the i-th perturbed discrete prompt with respect to the k-th column of the parameter matrix of the policy model; is the i-th perturbed discrete prompt.

[0156] Furthermore, for the scoring function (measuring similarity) of the Softmax policy model, that is, the calculation formula for the policy model gradient of the Softmax policy model is defined as:

[0157] ;

[0158] ;

[0159] where, is the policy model gradient of the Softmax policy model corresponding to the perturbed discrete prompt y; is the discrete prompt encoding corresponding to the perturbed discrete prompt y, , which represents the one-hot encoding of the perturbed discrete prompt y satisfies , is the element in the i-th row and j-th column of the one-hot encoding. If y j = i, then take 1, otherwise take 0, is the j-th token in the discrete prompt, is the result of performing a Softmax transformation on the parameter matrix , that is, the transformation matrix; is the matrix element in the i-th row and j-th column of is the matrix element in the i-th row and j-th column of the parameter matrix; is the matrix element in the l-th row and j-th column of the parameter matrix; m is the number of the vocabulary of natural words.

[0160] Furthermore, the scoring function (measuring similarity) of the Hadamard policy model, that is, the calculation formula for the policy model gradient of the Hadamard policy model is defined as:

[0161] ;

[0162] ;

[0163] wherein, is the policy model gradient of the Hadamard policy model corresponding to the perturbation discrete prompt word y; is the first matrix; is the second matrix; is a very small positive number, whose function is to ensure that the denominator is not zero; is the matrix element at the i-th row and j-th column in ; is the matrix element at the i-th row and j-th column in .

[0164] Furthermore, for the proof of the scoring function (measuring similarity) of the Hadamard policy model, given a classification vector sampled from the probability distribution , wherein , it can be assumed that , wherein is a very small positive number. If is set to 0, then there is:

[0165] ;

[0166] wherein, is the j-th column of the parameter matrix ; is the matrix element at the m-th row and j-th column in ; is the Hadamard product; is the all-one vector of size m×1.

[0167] The Hadamard reparameterization technique has been studied in the literature of reinforcement learning, but the present invention is different from the existing methods because the present invention allows instead of directly constraining , thus avoiding projection on the probability simplex.

[0168] Specifically, the present invention has . Obviously, the term here is to prevent the denominator from being 0, so as to calculate the scoring function:

[0169] ;

[0170] ;

[0171] Furthermore, let For Taking the partial derivative, we have:

[0172] ;

[0173] ;

[0174] Thus, we have .

[0175] Wherein, is the matrix element at the u-th row and v-th column in the v-th token in the perturbed discrete prompt word.

[0176] Furthermore, next we calculate the second term:

[0177] ;

[0178] Putting the two results together, we have:

[0179] ;

[0180] ;

[0181] ;

[0182] ;

[0183] Wherein, , both are matrices, and we have

[0184] , ;

[0185] Therefore, we have:

[0186] ;

[0187] Wherein, is the matrix element at the l-th row and v-th column in the one-hot encoding of y.

[0188] It is worth mentioning that for the optimization algorithm, genetic algorithm or simulated annealing can replace the elite coordinate descent method (ECD) to explore the discrete prompt space through population evolution or temperature annealing mechanism. For example, regarding the prompt as a "chromosome", the combination is optimized through crossover, mutation and selection operations. Or, use the Bayesian optimization algorithm to replace the elite coordinate descent method (ECD), which uses a probability model to construct a surrogate model of the objective function and efficiently searches the discrete space through a sequential sampling strategy, reducing the number of API calls.

[0189] Specifically, step 103 may include the following sub-steps S31-S311:

[0190] Step S31: Calculate the loss value corresponding to each inference result according to each inference result;

[0191] It should be noted that the process of calculating the loss value corresponding to each inference result according to each inference result can refer to the prior art, and the present invention will not elaborate further.

[0192] Step S32: Determine the objective function value corresponding to each loss value based on each loss value;

[0193] It should be noted that the process of determining the objective function value is specifically:

[0194] ;

[0195] where, is the objective function value corresponding to the i-th perturbed discrete prompt; is the loss value corresponding to the i-th perturbed discrete prompt.

[0196] Step S33: Sort each objective function value in ascending order and determine the ranking rank corresponding to each objective function value;

[0197] Step S34: Normalize each ranking rank and determine the weight value corresponding to each objective function value;

[0198] It should be noted that since the policy gradient has a large noise, the present invention introduces an objective value shaping method (Objective Value Shaping Method) and tries to reduce the influence of the noise by adjusting the objective value, making the gradient update more stable and accurate. The present invention defines a negative gradient descent step size d:

[0199] ;

[0200] ;

[0201] where, the weight c i needs to be replaced with the set The ranking in. Specifically, the present invention replaces c i with , where represents the ranking rank of c i in the set . It is not difficult to note that is the ranking of the objective function value in the set . Assuming , , so it can be defined as:

[0202] ;

[0203] where is the element in the i-th row and j-th column of ; is the power, indicating to judge whether the j-th component of the i-th prompt word x takes the value exactly as the row label l. If it holds, take 1, otherwise take 0; is the element in the i-th row and j-th column of the updated parameter matrix; is the updated in the i-th row and j-th column of the element.

[0204] Based on the above basis, thus there is:

[0205] ;

[0206] where the second equal sign in this formula is due to and coincide in the k-th column, and the third equal sign is due to degenerates except for the k-th column. Since and have the same terms in the j (j≠k) -th coordinate, it is concluded that for all , this does not affect the ranking. Therefore, it is necessary to sort and calculate the normalized ranking to replace c i . On the other hand, since only the k-th column of g is non-zero, so can be updated column by column. This is the core idea of stochastic coordinate descent, which updates the gradient of only one coordinate each time, avoiding the high cost of calculating the gradient for all coordinates in high dimensions.

[0207] For example, assume that multiple objective function values correspond to [5, 2, 9, 1]. After sorting all the objective function values, it is [1, 2, 5, 9]. The ranking ranks corresponding to the multiple objective function values [5, 2, 9, 1] are [3, 2, 4, 1] (that is, the original value 1 ranks first, the original value 2 ranks second, and so on). Then, normalize the ranking rank of each objective function value to obtain multiple weight values ci 。

[0208] Step S35: Perform one-hot encoding on each perturbed discrete prompt word to determine the discrete prompt word encoding corresponding to each perturbed discrete prompt word;

[0209] Step S36: Using the policy model gradient calculation formula, calculate the policy model gradient corresponding to each perturbed discrete prompt word according to each discrete prompt word encoding and the intermediate parameter matrix of the intermediate policy model;

[0210] Step S37: Using the parameter matrix gradient calculation formula, calculate the parameter matrix gradient according to each weight value and each policy model gradient;

[0211] Step S38: Update the intermediate parameter matrix of the intermediate policy model according to the parameter matrix gradient, determine the updated policy model, and count the update times in real time;

[0212] Step S39: Perform a mean operation on the loss values corresponding to each inference result to determine the loss mean;

[0213] Step S310: Determine whether the update times reach the preset update threshold and whether the loss mean converges;

[0214] Step S311: If the update times reach the preset update threshold or the loss mean converges, use the updated policy model as the target policy model.

[0215] It should be noted that the elite coordinate descent algorithm integrates probability modeling, noise candidate generation, and elite preservation mechanisms to solve high-dimensional discrete optimization problems. Please refer to Table 3. First, perform initialization, randomly generate the training prompt word a , and initialize a -dimensional (i.e., the solution space dimension) parameter matrix as the initial parameter matrix. In the t-th iteration, the algorithm performs the following steps:

[0216] 1. Randomly draw an index: Randomly draw an index k from the uniform distribution u of the training prompt word a m to determine on which column to perform coordinate descent update.

[0217] 2. Generate a new parameter matrix : Use the preset parameter matrix update function to generate a new parameter matrix (intermediate parameter matrix), where the training prompt word a is used as the current solution y, and k is the randomly selected coordinate index. The preset parameter matrix update function will retain the k-th column of

[0218] 3. From the intermediate policy model Medium sampling: According to the generated sampling x i , that is, generate multiple discrete prompt words, indicating the generation of new candidate solutions.

[0219] 4. Calculate the objective function value: For the sampled x i , through the noise model generate the corresponding perturbed discrete prompt word y i , and then calculate the objective function value f i =f(y i ).

[0220] 5. Calculate the weight value c i : For each f i , calculate its rank r in the set of objective function values i for subsequent gradient updates. A better f i corresponds to a higher rank and thus occupies a greater weight in subsequent updates.

[0221] 6. Gradient calculation: The calculation of the gradient depends on the weight value and the k-th column of the current parameter matrix , and the gradient g is calculated using the following formula:

[0222] ;

[0223] ;

[0224] where g is the parameter matrix gradient; N is the total number of perturbed discrete prompt words; is the weight value corresponding to the i-th perturbed discrete prompt word, is the rank corresponding to the i-th perturbed discrete prompt word; is the policy model gradient of the logarithm probability corresponding to the i-th perturbed discrete prompt word with respect to the k-th column of the parameter matrix of the policy model; is the i-th perturbed discrete prompt word.

[0225] 7. Update the intermediate parameter matrix : According to the calculated gradient g, use the learning rate to update the k-th column of the parameter matrix :

[0226] ;

[0227] where, is the k-th column of the updated parameter matrix of the policy model; is the k-th column of the intermediate parameter matrix.

[0228] This column-by-column update method is a typical coordinate descent strategy, which updates the gradient of a specific coordinate each time instead of a global update. The algorithm will repeat the above steps continuously until the preset termination condition is reached. Specifically:

[0229] S41. If the number of updates has not reached the preset update threshold and the loss mean has not converged, then use the updated policy model as the new initial policy model;

[0230] S42. Use the preset parameter matrix update function to update the initial parameter matrix of the new initial policy model according to the training prompt words to determine the intermediate policy model;

[0231] S43. Use the intermediate policy model and the preset generative language model to perform inference based on the training sentences to generate multiple perturbed discrete prompt words and multiple inference results;

[0232] S44. Based on the preset gradient calculation formula, use the elite coordinate descent algorithm to update the intermediate parameter matrix of the intermediate policy model according to multiple perturbed discrete prompt words and multiple inference results to determine the loss mean and the updated policy model, and count the number of updates in real time until the number of updates reaches the preset update threshold or the loss mean converges;

[0233] S45. Use the updated policy model determined when the number of updates reaches the preset update threshold or the loss mean converges as the target policy model.

[0234] It should be noted that if the number of updates has not reached the preset update threshold and the loss mean has not converged, then use the updated policy model as the new initial policy model and jump to execute step 101 until the number of updates reaches the preset update threshold or the loss mean converges; use the updated policy model determined when the number of updates reaches the preset update threshold or the loss mean converges as the target policy model.

[0235] Table 3 Elite Coordinate Descent (ECD) Algorithm Framework

[0236]

[0237] It is worth mentioning that the pseudocode encapsulates the core ideas of the elite coordinate descent framework: (1) a parameterized probability search strategy, (2) generating new candidate objects through a noise model, (3) a ranking-based objective function value shaping method for stabilizing gradient estimation, (4) selective coordinate updates for reducing computational overhead. These features ultimately form an effective and scalable high-dimensional discrete optimization method that guides the parameter model to high-quality solutions in a principled and adaptable manner.

[0238] ​Furthermore, for the phased search process, those skilled in the art can adopt a combination of rough screening and fine-tuning: first, narrow the search scope through heuristic algorithms (such as keyword matching), and then perform local optimization on the candidate prompt words to reduce the computational complexity. Or, adopt multi-objective optimization: simultaneously optimize the performance metrics of the prompt words (such as accuracy) and semantic quality (such as interpretability), and select the optimal solution through the Pareto frontier.

[0239] For the hybrid architecture design, those skilled in the art can adopt a hybrid optimization of black box and white box: in the scenario where some model parameters are accessible, combine gradient optimization (white box) and derivative-free optimization (black box) to improve the search efficiency. Or adopt manual intervention enhancement: introduce an expert knowledge base or an interactive feedback mechanism, and guide the generation of prompt words through manual annotation to reduce the algorithm complexity.

[0240] Step 104: Generate a target discrete prompt word based on the target policy model.

[0241] It should be noted that through the trained target policy model, a target discrete prompt word is generated, and the target discrete prompt word is used to achieve accurate reasoning. Specifically:

[0242] S51: When receiving the sentence to be tested, use the target discrete prompt word as the input of the text-to-text transfer transformer, and output the target semantic sentence;

[0243] S52: Concatenate the target semantic sentence and the sentence to be tested to generate a target concatenated sentence;

[0244] S53: Input the target concatenated sentence into the large language model, and output the target reasoning result.

[0245] In this embodiment, the present invention addresses the deficiencies of RLPrompt (Reinforcement Learning-based Prompt Optimization) that rely on artificial reward functions and repeated training of the policy network. The present invention introduces a semantic generation module based on the T5 model to convert discrete prompt words into natural language sentences. This design not only solves the problems of inconsistent semantics and poor interpretability of traditional discrete prompt words, but also enhances the semantic consistency between the prompt words and the task objectives by leveraging the semantic generation ability of the pre-trained language model, thereby improving the reasoning accuracy and generalization ability of the large language model. At the same time, since there is no need to retrain the policy network, the present invention significantly reduces the consumption of computing resources and realizes a more efficient optimization process.

[0246] As a comparison of technical effects, it can be referenced in combination with the existing technology. In the existing discrete prompt search methods, there are several mainstream technologies: ① Manual construction: Tom et al. proposed manually designing prompts, relying on experts to convert common downstream tasks into multi-classification tasks based on mask filling (fill-mask); at the same time, a classification label generation method based on statistical optimization was proposed. ② White-box discrete prompt search: AutoPrompt is an algorithm that uses white-box discrete prompt search technology. It optimizes prompts by analyzing the internal structure of the model to improve performance. RLPrompt converts the optimization problem of discrete prompts into a reinforcement learning problem and uses a continuous policy network to explore the prompt space. Plum adopts a prompt search based on meta-heuristic methods, including various typical optimization methods such as hill climbing, simulated annealing, genetic algorithms (with or without crossover), tabu search, and harmony search. These methods perform well in both discrete and continuous prompt searches and can be used to generate more human-interpretable prompts, making them effective tools for optimizing prompts in a white-box environment. GAP3 uses a genetic algorithm to optimize prompts. The model treats prompts as "chromosomes", generates different prompt combinations through crossover and mutation, and selects the optimal prompt through fitness scoring. ③ Black-box discrete prompt search: BBT (Black-Box Tuning) is a continuous prompt search technology for optimizing the inference effect of language models. Black-box discrete prompt search technologies (such as BBT) improve the inference performance of language models by optimizing continuous prompts, especially suitable for closed-source commercial models that cannot directly access internal parameters. BBT adopts a derivative-free optimization method and only adjusts the soft prompts in the input through the API interface. Experiments on models such as RoBERTa show that its effect is better than manual design and gradient-based prompt adjustment methods. BBTv2 further improves performance through hierarchical prompt optimization and can be comparable to full-model fine-tuning with a small amount of data. In addition, similar technologies such as Proxy-Tuning achieve black-box optimization by fine-tuning small models and using their prediction differences from large models, and also perform well in the adaptation of specific tasks for closed-source models such as GPT-3.5. These methods provide effective solutions for the optimization of large models in a commercial environment. ④ Methods based on continuous rounding (such as Hard Prompts Made Easy) optimize model performance through the rounding or approximation of prompt content, combining the efficiency of gradient optimization and the interpretability of hard prompts when dealing with complex prompts.⑤Context learning techniques significantly improve model performance by combining prompt words with input context: in terms of task execution, enabling large language models (LLMs) to complete multi-dimensional evaluations with only a few examples; in terms of generalization ability, helping the model adapt to new tasks through meta-in-context learning; in terms of domain applications, successfully applied to scenarios such as automatic metadata annotation, despite the limitations in obtaining subject rules. These two methods together have promoted the development of prompt engineering in terms of efficiency, adaptability, and practicality.

[0247] Based on the above foundation, existing discrete prompt word search methods face three core challenges: First, it is difficult to efficiently optimize the discrete search space. The failure of traditional gradient methods leads to reliance on inefficient random / heuristic searches and makes it difficult to obtain the global optimal solution; Second, the generated discrete prompt words lack natural semantic associations, damaging the interpretability of the results and affecting model performance; Finally, the search cost is too high - the solution space that grows exponentially with the length of the prompt words makes reinforcement learning methods (such as RLPrompt) computationally intensive, and although black-box optimization methods (such as BDPL) adopt variance reduction policy gradients, they still require a large number of API calls, significantly increasing the time and economic costs. These limitations together restrict the practical application effectiveness of discrete prompt word search technology.

[0248] The existing discrete prompt word optimization schemes closest to the present invention mainly include two types of methods: discrete prompt word learning based on direct probability modeling (BDPL) and prompt word optimization based on reinforcement learning (RLPrompt):

[0249] (1) Discrete prompt word learning based on direct probability modeling (BDPL): BDPL realizes prompt word optimization in a black-box environment through the variance reduction policy gradient algorithm (VR-PGE). The specific process is as follows: Independent categorical distribution sampling: Establish a categorical distribution for each prompt word position and independently sample to generate a prompt sequence; Black-box inference and loss calculation: Concatenate the prompt word with the input and perform inference through a pre-trained model, and calculate the prediction loss; Gradient estimation and distribution update: Use the VR-PGE algorithm to estimate the gradient and update the categorical distribution parameters through projected gradient descent to ensure the legality of the probability space. This method only requires API calls and is applicable to closed-source models, but it has the defects of high search cost (requiring a large number of samplings to calculate the gradient) and semantic incoherence (independent distribution sampling results in the lack of context association of prompt words).

[0250] (2)Reinforcement Learning-based Prompt Optimization (RLPrompt): RLPrompt transforms discrete prompt optimization into a reinforcement learning problem: Policy network design: Insert a trainable MLP layer into a frozen distillation model (such as distilGPT-2) as the policy network to gradually generate prompt words; Reward engineering: Input normalization reward: Eliminate the reward scale differences of different inputs through z-score; Piecewise reward function: Combine the continuous label probability signal with the sparse classification correct signal to prevent adversarial prompts; Policy optimization: Use the Soft Q-Learning (SQL) algorithm to maximize the expected reward. Although this method performs excellently in few-shot classification tasks, its policy network needs to be trained from scratch, fails to reuse pre-trained knowledge, and the generated prompt words are mostly meaningless and disordered texts, with poor interpretability.

[0251] In summary, as an emerging black-box prompt learning method, BDPL has successfully solved the problem of discrete prompt word optimization in an environment where the internal parameters and gradient information of the pre-trained language model cannot be accessed and has shown excellent performance in terms of security, computational efficiency, etc. However, it still faces a series of challenges and drawbacks. In the BDPL framework, although the projection processing strategy is adopted to map the discrete prompt gradient to the classification probability space, the curse of dimensionality in the high-dimensional space leads to a sharp increase in computational complexity, seriously affecting the search accuracy and convergence efficiency. In the RLPrompt method, although systematic space exploration is achieved through the reinforcement learning policy network, there are three limitations: 1) The inherent computational complexity of reinforcement learning significantly increases the time cost; 2) The artificial design of the reward function introduces subjective biases and reduces the optimization stability; 3) The design that requires re-training the policy network fails to reuse the knowledge of the pre-trained model, resulting in waste of computational resources. These structural defects jointly restrict the practical process of discrete prompt word optimization technology.

[0252] To address the above problems, the present invention proposes a discrete prompt word search method for large language models, aiming to solve the above problems by adding a semantic generation process during the discrete prompt word search. By integrating the advantages of BDPL and RLPrompt, a new method is introduced to overcome these limitations. Specifically, the present invention introduces a semantic generation tool, enabling the generated prompt words to not only have better interpretability but also utilize the information of the previously solved prompt words as prior knowledge to guide the subsequent optimization process. This approach can effectively reduce the search difficulty of prompt words, overcome the curse of dimensionality, and improve the optimization accuracy and efficiency of prompt words. At the same time, by leveraging the existing knowledge of the pre-trained language model, the cumbersome process of re-training the neural network in RLPrompt is avoided, thereby improving the overall robustness of the algorithm, reducing the training time and computational overhead, and ultimately achieving a faster convergence speed and higher optimization efficiency.

[0253] In an embodiment of the present invention, the present invention provides a method for searching discrete prompt words of a large language model. The initial policy model is updated by using a preset parameter matrix update function to obtain an intermediate policy model. Then, based on the elite coordinate descent algorithm, combined with the intermediate policy model, a preset generative language model, and a preset gradient calculation formula, iterative optimization of the parameter matrix is performed according to the training inference results to determine the target policy model. Finally, based on the target policy model, the process of generating the target discrete prompt words is carried out. The present invention can achieve efficient search through the elite coordinate descent method and make full use of the existing knowledge of the pre-trained language model, without having to train the policy network from scratch every time for optimization. In practical applications, the target discrete prompt words can be directly generated by using the target policy model, thereby reducing the computational overhead of prompt word optimization.

[0254] For better illustration, refer to Figure 2 , which shows a schematic diagram of the working principle of a method for searching discrete prompt words of a large language model provided in the second embodiment of the present invention compared with the prior art, including:

[0255] The working principle of the present invention is as shown in Figure 2 the right sub-figure (b): Assume that the number of tokens to be learned by the model is 2, represented by one blue and one yellow. T is the discrete prompt word learned by the model. Each token corresponds to an individual distribution function, and the token at the current position is determined by the distribution function. S is the input sentence. T is concatenated before S and then input into the large model for inference to obtain a prediction result, and then the Loss value is calculated together with the known label y. By minimizing the Loss, the training parameters in the distribution function are updated, and thus the content of the discrete prompt word is updated. Compared with the prior art, the present invention has three main points:

[0256] 1) Continuous modeling method for discrete probability distribution: By introducing a parametric classification model, the discrete prompt word is represented as discrete sampling points of the classification model, thereby transforming the original discrete search problem into an unconstrained problem-solving in a continuous real space. This model can effectively reduce the difficulty of problem-solving and at the same time does not require an additional constraint processing process during the solution process, improving the solution speed.

[0257] 2) Efficient solution method based on stochastic coordinate descent: A stochastic coordinate gradient descent method for efficient search in the continuous parameter space of the classification model is designed. Compared with the standard gradient descent method often used in existing similar work, the high-quality prompt words that appear during the algorithm operation are more likely to be retained, thereby improving the convergence efficiency of the search process.

[0258] 3) Semantic enhancement strategy based on the seq2seq model: By introducing seq2seq models such as T5 between discrete prompt words and the large model interface, convert prompt words that do not conform to the semantic of natural language writing into sentences that conform to daily language, thereby improving the interpretability of the final results of discrete prompt word search problems and being able to provide human users with more understandable prompt word options. By using the seq2seq model, the generated prompt words have better understandability, making the model behavior more transparent and facilitating user understanding and application. At the same time, the generated natural language prompt words have a higher semantic matching degree with the task objective, thereby improving the accuracy and effectiveness of task execution. In addition, through the T5 model, the generated prompt words are more grammatically standard, conform to the expression habits of natural language, and improve the input quality of the large language model. This semantic generation method based on the seq2seq model can effectively solve the problems of unsmooth grammar and unclear semantics that may occur during the generation of discrete prompt words, provide more accurate and natural prompts for the large language model, and thus improve its performance in various natural language processing tasks.

[0259] In this embodiment, by converting discrete search into continuous space optimization, the dimension disaster problem of high-dimensional combinations is effectively solved by combining the Softmax and Hadamard strategies. This method uses column-by-column update and ranking gradient estimation techniques to reduce the computational cost, and uses a pre-trained seq2seq model to improve the semantic quality of prompt words. Through the co-design of probability modeling, efficient search, and semantic generation, black-box optimization without repeated training is achieved, significantly improving the semantic rationality of prompt words and the model inference performance.

[0260] Please refer to Figure 3 , Figure 3 which is the structural block diagram of a discrete prompt word search device for a large language model provided in Embodiment III of the present invention

[0261] A discrete prompt word search device for a large language model provided by the present invention includes:

[0262] An acquisition module 301, configured to acquire training prompt words and training sentences, and update the initial parameter matrix of the initial policy model according to the training prompt words by using a preset parameter matrix update function to determine an intermediate policy model;

[0263] An inference module 302, configured to perform inference according to the training sentences by using the intermediate policy model and a preset generative language model to generate multiple perturbed discrete prompt words and multiple inference results;

[0264] An optimization module 303, configured to iteratively optimize the intermediate parameter matrix of the intermediate policy model based on a preset gradient calculation formula by using an elite coordinate descent algorithm according to the multiple perturbed discrete prompt words and the multiple inference results to determine a target policy model;

[0265] A generation module 304, configured to generate a target discrete prompt based on a target policy model.

[0266] The pre - set generative language model includes a noise model, a text - to - text transfer transformer, and a large language model; further, the inference module 302 is specifically configured to:

[0267] Generate multiple intermediate discrete prompts based on an intermediate policy model;

[0268] Use the noise model to perturb each intermediate discrete prompt respectively to generate a perturbed discrete prompt corresponding to each intermediate discrete prompt;

[0269] Respectively use each perturbed discrete prompt as the input of the text - to - text transfer transformer, and output a training semantic sentence corresponding to each perturbed discrete prompt;

[0270] Concatenate each training semantic sentence with a training sentence respectively to generate a training concatenated sentence corresponding to each training semantic sentence;

[0271] Respectively input each training concatenated sentence into the large language model, and output a training inference result corresponding to each training concatenated sentence.

[0272] Further, the pre - set gradient calculation formula includes a parameter matrix gradient calculation formula and a policy model gradient calculation formula; the optimization module 303 is specifically configured to:

[0273] Calculate a loss value corresponding to each inference result according to each inference result;

[0274] Determine an objective function value corresponding to each loss value based on each loss value;

[0275] Sort each objective function value in ascending order to determine a ranking rank corresponding to each objective function value;

[0276] Normalize each ranking rank to determine a weight value corresponding to each objective function value;

[0277] Perform one - hot encoding on each perturbed discrete prompt to determine a discrete prompt encoding corresponding to each perturbed discrete prompt;

[0278] Use the policy model gradient calculation formula to calculate a policy model gradient corresponding to each perturbed discrete prompt according to each discrete prompt encoding and the intermediate parameter matrix of the intermediate policy model;

[0279] Use the parameter matrix gradient calculation formula to calculate a parameter matrix gradient according to each weight value and each policy model gradient;

[0280] Update the intermediate parameter matrix of the intermediate policy model according to the parameter matrix gradient, determine the updated policy model, and count the update times in real time;

[0281] Perform a mean operation on the loss values corresponding to each inference result to determine the mean loss;

[0282] Judge whether the number of updates reaches the preset update threshold and whether the mean loss converges;

[0283] If the number of updates reaches the preset update threshold or the mean loss converges, use the updated policy model as the target policy model.

[0284] In an alternative device embodiment, it further includes:

[0285] The first module is used to, if the number of updates does not reach the preset update threshold and the mean loss does not converge, use the updated policy model as the new initial policy model;

[0286] The second module is used to update the initial parameter matrix of the new initial policy model according to the training prompt words by using the preset parameter matrix update function to determine the intermediate policy model;

[0287] The third module is used to perform inferences according to the training sentences by using the intermediate policy model and the preset generative language model to generate multiple perturbed discrete prompt words and multiple inference results;

[0288] The fourth module is used to, based on the preset gradient calculation formula, use the elite coordinate descent algorithm to update the intermediate parameter matrix of the intermediate policy model according to multiple perturbed discrete prompt words and multiple inference results, determine the mean loss and the updated policy model, and count the update times in real time until the number of updates reaches the preset update threshold or the mean loss converges;

[0289] The fifth module is used to use the updated policy model determined when the number of updates reaches the preset update threshold or the mean loss converges as the target policy model.

[0290] In an alternative device embodiment, it further includes:

[0291] The sixth module is used to, when receiving the sentence to be tested, use the target discrete prompt word as the input of the text-to-text transfer transformer and output the target semantic sentence;

[0292] The seventh module is used to splice the target semantic sentence and the sentence to be tested to generate the target spliced sentence;

[0293] The eighth module is used to input the target spliced sentence into the large language model and output the target inference result.

[0294] Furthermore, the parameter matrix gradient calculation formula in the preset gradient calculation formula is specifically:

[0295] ;

[0296] where g is the gradient of the parameter matrix; N is the total number of perturbed discrete prompts; is the i-th weight value; is the policy model gradient of the k-th column of the parameter matrix of the policy model with respect to the log probability corresponding to the i-th perturbed discrete prompt; is the i-th perturbed discrete prompt.

[0297] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0298] The embodiment of the present invention also provides a computer device, including a memory and a processor, where a computer program is stored in the memory; when the computer program is executed by the processor, the processor executes the steps of the large language model discrete prompt search method in any of the foregoing embodiments.

[0299] The embodiment of the present invention also provides a computer-readable storage medium, on which a computer program / instruction is stored, and when the computer program / instruction is executed by the processor, the steps of the large language model discrete prompt search method in any of the foregoing embodiments are implemented.

[0300] The embodiment of the present invention also provides a computer program product, including a computer program / instruction, and when the computer program / instruction is executed by the processor, the steps of the large language model discrete prompt search method in any of the foregoing embodiments are implemented.

[0301] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A discrete prompt search method for large language models, characterized in that, Including: Obtain training prompt words and training sentences, and update the initial parameter matrix of the initial policy model according to the training prompt words by using a preset parameter matrix update function to determine an intermediate policy model; Perform inference according to the training sentences by using the intermediate policy model and a preset generative language model to generate a plurality of perturbed discrete prompt words and a plurality of inference results; Based on a preset gradient calculation formula, use the elite coordinate descent algorithm to iteratively optimize the intermediate parameter matrix of the intermediate policy model according to the plurality of perturbed discrete prompt words and the plurality of inference results to determine a target policy model; Generate a target discrete prompt word based on the target policy model.

2. The discrete prompt word search method for large language models according to claim 1, characterized in that The preset generative language model includes a noise model, a text-to-text transfer transformer, and a large language model; the performing inference according to the training sentences by using the intermediate policy model and the preset generative language model to generate a plurality of perturbed discrete prompt words and a plurality of inference results includes: Generate a plurality of intermediate discrete prompt words based on the intermediate policy model; Use the noise model to perturb each of the intermediate discrete prompt words to generate a perturbed discrete prompt word corresponding to each of the intermediate discrete prompt words; Respectively use each of the perturbed discrete prompt words as an input to the text-to-text transfer transformer to output a training semantic sentence corresponding to each of the perturbed discrete prompt words; Concatenate each of the training semantic sentences with the training sentence to generate a training concatenated sentence corresponding to each of the training semantic sentences; Respectively input each of the training concatenated sentences into the large language model to output a training inference result corresponding to each of the training concatenated sentences.

3. The discrete prompt word search method for large language models according to claim 1, characterized in that, The preset gradient calculation formula includes a parameter matrix gradient calculation formula and a policy model gradient calculation formula; the using the elite coordinate descent algorithm to iteratively optimize the intermediate parameter matrix of the intermediate policy model according to the plurality of perturbed discrete prompt words and the plurality of inference results based on the preset gradient calculation formula to determine a target policy model includes: Calculate a loss value corresponding to each of the inference results according to each of the inference results; Determine an objective function value corresponding to each of the loss values based on each of the loss values; Sort each of the objective function values in ascending order to determine a ranking rank corresponding to each of the objective function values; Normalize each of the ranking ranks to determine a weight value corresponding to each of the objective function values; Perform one-hot encoding on each of the perturbed discrete prompt words to determine a discrete prompt word encoding corresponding to each of the perturbed discrete prompt words; Use the policy model gradient calculation formula to calculate a policy model gradient corresponding to each of the perturbed discrete prompt words according to each of the discrete prompt word encodings and the intermediate parameter matrix of the intermediate policy model; Use the parameter matrix gradient calculation formula to calculate a parameter matrix gradient according to each of the weight values and each of the policy model gradients; Update the intermediate parameter matrix of the intermediate policy model according to the parameter matrix gradient to determine an updated policy model, and count the update times in real time; Perform a mean operation on the loss values corresponding to each of the inference results to determine a loss mean; Judge whether the update times reach a preset update threshold and whether the loss mean converges; If the number of updates reaches a preset update threshold or the mean loss converges, the updated policy model is taken as the target policy model.

4. The method for searching discrete prompt words of the large language model according to claim 3, wherein, It further includes: If the number of updates does not reach the preset update threshold and the mean loss does not converge, the updated policy model is taken as the new initial policy model; The initial parameter matrix of the new initial policy model is updated according to the training prompt words by using a preset parameter matrix update function to determine an intermediate policy model; The intermediate policy model and a preset generative language model are used to perform inference according to the training sentences to generate a plurality of perturbed discrete prompt words and a plurality of inference results; Based on a preset gradient calculation formula, the elite coordinate descent algorithm is used to update the intermediate parameter matrix of the intermediate policy model according to the plurality of perturbed discrete prompt words and the plurality of inference results to determine the mean loss and the updated policy model, and the number of updates is statistically counted in real time until the number of updates reaches the preset update threshold or the mean loss converges; The updated policy model determined when the number of updates reaches the preset update threshold or the mean loss converges is taken as the target policy model.

5. The method for searching discrete prompt words of a large language model according to claim 2, wherein After the step of generating the target discrete prompt words based on the target policy model, it includes: When a sentence to be measured is received, the target discrete prompt words are used as the input of the text-to-text transfer transformer to output a target semantic sentence; The target semantic sentence and the sentence to be measured are concatenated to generate a target concatenated sentence; The target concatenated sentence is input into a large language model to output a target inference result.

6. The method for searching discrete prompt words of the large language model according to claim 1, wherein The parameter matrix gradient calculation formula in the preset gradient calculation formula is specifically: ; Among them, g is the parameter matrix gradient; N is the total number of perturbed discrete prompts; is the i-th weight value; is the policy model gradient of the k-th column of the parameter matrix of the policy model with respect to the log probability corresponding to the i-th perturbed discrete prompt; is the i-th perturbed discrete prompt.

7. A discrete prompt search device for large language models, characterized in that, It includes: An acquisition module, configured to acquire training prompt words and training sentences, and update the initial parameter matrix of the initial policy model according to the training prompt words by using a preset parameter matrix update function to determine an intermediate policy model; An inference module, configured to perform inference according to the training sentences by using the intermediate policy model and a preset generative language model to generate a plurality of perturbed discrete prompt words and a plurality of inference results; An optimization module, configured to iteratively optimize the intermediate parameter matrix of the intermediate policy model based on a preset gradient calculation formula by using the elite coordinate descent algorithm according to the plurality of perturbed discrete prompt words and the plurality of inference results to determine a target policy model; A generation module, configured to generate target discrete prompt words based on the target policy model.

8. A computer device, characterized in that, It includes a memory and a processor. A computer program is stored in the memory. When the computer program is executed by the processor, the processor is caused to execute the steps of the method for searching for discrete prompt words of a large language model according to any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed, it implements the method for searching for discrete prompt words of a large language model according to any one of claims 1-6.

10. A computer program product, characterized in that, The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer is caused to execute the method for searching for discrete prompt words of a large language model according to any one of claims 1-6.

Citation Information

Cited By

  • Large language model cue word automatic optimization method

    CN120542583A

  • Prompt word optimization method and system based on approximate submodule function and continuous learning

    CN120597895A

  • Prompt word optimization method and system based on approximate submodular function and continuous learning

    CN120597895B

  • Network protocol intelligent extraction method based on large language model and application

    CN120725152A

  • Large model security vulnerability detection method based on multi-agent reinforcement learning

    CN120805146A