Knowledge-guided reinforcement learning discrete cue word optimization method

By optimizing prompt word editing using a structured knowledge base and a deep Q-network, the efficiency and controllability issues of prompt word optimization in black-box large language models are resolved, thereby improving the quality of generated prompt words and the model's performance in downstream tasks.

CN120832891APending Publication Date: 2025-10-24SHANGHAI SECOND POLYTECHNIC UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510905316.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-28
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently and controllably optimize prompt words in black-box large language models, and lack structured knowledge constraints and diverse reward mechanisms, resulting in poor performance of the generated prompt words.

Method used

We employ a structured phrase-level knowledge base and a deep Q-network, combined with semantic embedding and task statistical features, to design a diverse reward mechanism. Through reinforcement learning, we optimize the prompt word editing operation, ensuring the compliance of editing and global exploration capabilities.

Benefits of technology

It enables efficient and controllable optimization of prompt words in a black-box environment, improves the model's performance in downstream tasks, enhances the quality and interpretability of optimization results, and adapts to different tasks and API budget conditions.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention discloses a black box large language model-oriented knowledge-guided reinforcement learning discrete cue word optimization method, and belongs to the technical field of natural language processing and artificial intelligence optimization. According to the method, a KPE-RL (Known-gued Prompt Evolution with Reforming Learning) algorithm is put forward, editing operation is constrained by using a structured phrase-level knowledge base, cue words are optimized and modeled into a Markov decision process, and a discrete editing strategy is learned through a deep Q network. According to the method, a mixed state coding mode fusing semantic embedding and task statistical characteristics is designed, and a diversity regular reward is designed, so that efficient exploration is encouraged, and the optimization capability of the model in a limited API calling scene is improved. The method is suitable for automatic prompt optimization of the large language model in the API black box scene, and has generalization, robustness and practical application value.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of artificial intelligence optimization and natural language processing, and specifically relates to a reinforcement learning driven discrete prompt optimization method for a large language model black box API scenario, which can improve the performance of the model in various natural language processing (NLP) downstream tasks. BACKGROUND

[0002] With the increasingly wide application of large language models (LLMs) in natural language processing, automatic question answering, text generation, knowledge extraction and other fields, guiding the model to generate high-quality outputs related to the target has become an important research direction. Prompt learning, as a convenient model adaptation technology, can enable the model to complete various downstream tasks without modifying parameters by constructing specific input templates. Compared with traditional fine-tuning methods, prompt learning has the characteristics of high efficiency and flexibility, and therefore is valued in practical applications. Currently, automatic prompt optimization techniques mainly include gradient-based, heuristic search-based and reinforcement learning-based methods. Gradient-based optimization methods automatically adjust templates or trigger words by obtaining model gradient information. Such methods usually require access to internal parameters of the model and are suitable for open models, but are difficult to apply to black box models. Heuristic search-based methods use greedy search, evolutionary algorithms and other methods to optimize discrete editing operations. Such methods can be applied to black box models without gradient information, but have large search spaces, are prone to local optimization, have limited exploration range, and optimization efficiency needs to be improved. Reinforcement learning-based methods treat prompt optimization as a sequence decision process and use a reinforcement learning framework for optimization. Such methods have global search capabilities in theory, but current methods focus on a single state or reward signal and do not fully combine semantic and task statistical information, resulting in limited actual performance and unstable training processes.

[0003] In addition, existing methods are still insufficient in terms of structuring and controllability of editing operation space, lack effective constraints of domain knowledge and grammar rules, and result in poor performance of generated prompts. At the same time, for black box large language models, the number of API calls is limited, and prompt optimization methods need to achieve more efficient performance improvement under limited feedback. In summary, current automatic prompt optimization methods face the following challenges:

[0004] (1) Difficult to adapt to large language models that can only be called through API, unable to use gradient information, resulting in limited optimization efficiency;

[0005] (2) Single state representation and reward mechanism design, difficult to fully reflect the semantic expression and actual task performance of prompts;

[0006] (3) Lack of structured knowledge base support, lack of controllability and explainability of editing operation.

[0007] Therefore, it is urgent to develop a reinforcement learning prompt optimization method that integrates structured knowledge constraints, deep expression of semantic and task state, and introduces a diversity reward mechanism, to realize efficient and controllable automatic prompt optimization in the scene of black box large language model. SUMMARY

[0008] In view of the deficiencies of the prior art in the field of prompt optimization, the present application proposes a structured knowledge guided reinforcement learning discrete prompt optimization method, aiming to improve the automation, controllability and optimization efficiency of prompt editing.

[0009] To achieve the above-mentioned goal, the present application designs a structured short phrase knowledge base to constrain the editing operation of the prompt and ensure the semantic consistency and domain relevance of the editing result. A hybrid state encoding method combining semantic embedding and task statistical features is proposed to comprehensively reflect the expression ability and actual task performance of the prompt. A diversity reward mechanism is designed to encourage diverse editing paths during optimization, prevent falling into local optimum, and improve exploration efficiency and convergence stability. A deep Q network reinforcement learning optimizer is constructed to realize global optimal search of discrete editing operation of the prompt.

[0010] Further, the method of the present application specifically comprises the following steps:

[0011] Step 1, construct a structured short phrase knowledge base, classify phrases according to application domain and syntactic type, and attach form degree score, context label and other meta information to each phrase. Segment the initial prompt, combine the segmented prompt with semantic embedding and statistical features of recent task performance, and these statistical features include mean, standard deviation and trend of accuracy, and finally splice to form a hybrid state vector.

[0012] Step 2, define the action space including four types of discrete operations of short phrase level addition, replacement, deletion and order exchange, all operations are subject to knowledge base structured constraints to ensure the compliance and explainability of editing;

[0013] Step 3, based on deep Q network, input the current state and output the Q value of each editing action; realize the balance between exploration and utilization through greedy strategy; use experience replay and target network update in training process.

[0014] Step 4, the reward signal comprehensively considers the accuracy improvement and editing diversity brought by the current editing operation, and the diversity reward guides the optimizer to explore more possibilities under the limited API budget, preventing single editing behavior.

[0015] Step 5: The editing operation candidate phrase needs to meet multiple conditions such as grammar category, context relevance and similarity threshold, etc., to improve the actual effect.

[0016] Step 6: After multiple rounds of sequence decision, the reinforcement learning optimizer gradually optimizes the prompt word in the black box environment and outputs the final optimization result.

[0017] Further, the above method has the following advantages:

[0018] 1. It can automatically and efficiently optimize the prompt word in a black box environment without model gradient information, and improve the performance of the model in downstream tasks;

[0019] 2. The structured knowledge base ensures the controllability of the editing operation and the effective constraint of semantics and syntax, improving the quality and explainability of the optimization result;

[0020] 3. The mixed state encoding and diversity reward mechanism effectively improve the global exploration ability and convergence speed of the optimization process;

[0021] 4. The reinforcement learning optimizer can adapt to different tasks and API budget conditions, and has strong universality and practicality. DETAILED DESCRIPTION

[0022] The present application relates to a structured knowledge guided reinforcement learning prompt word optimization method for a black box large language model. The method models the discrete prompt word optimization problem as a Markov decision process, uses a deep Q network to learn the editing strategy, and automatically optimizes the prompt word under the constraint of the knowledge base. The typical implementation steps of the present application are as follows:

[0023] Step 1: Construct a structured phrase-level knowledge base K to constrain and filter the candidate phrases during prompt word editing. The knowledge base classifies phrases according to application domain and syntax category, and assigns each phrase form score, context label and other meta information. Application domains include general, medical and financial, and syntax categories such as noun phrase, verb phrase and conditional clause.

[0024] Step 2: For each round of prompt word editing, define a hybrid state representation s t , composed of semantic embedding and statistical features:

[0025] S t =[embeddingt,μt,σ t, Δ t ]

[0026] Where embedding t is the semantic embedding of the current prompt word, μ t is the mean of the accuracy of the last t rounds of tasks, σ t is the standard deviation, and Δt For accuracy rate trend. All state features are normalized by running mean and standard deviation during training process:

[0027]

[0028] Step 3: Define four phrase-level editing operations: addition, replacement, deletion and exchange. All operations are subject to the knowledge base K, and the edited candidate phrase needs to meet the consistency of grammatical type and context relevance. The filtering threshold is generally set to 0.65.

[0029] Step 4: The reward function is designed as follows:

[0030] r t =ΔAcc t +H(P t )

[0031] Where ΔAcc t is the change in task performance after editing, H(P t ) is the entropy of the current action distribution, and the diversity reward can prevent the editing from falling into a single operation mode and promote global exploration.

[0032] Step 5: The editing strategy is learned by a deep Q network Q(s, a; θ), with the network input being the current state and the output being the Q value of the four types of actions. A greedy strategy is used to balance exploration and utilization, and experience replay and target network are used to update regularly during training to improve convergence stability. The Q network loss is:

[0033]

[0034] Where γ is the discount factor and θ is the target network parameter.

[0035] Step 6: At each decision-making round, the agent selects and executes the editing operation, and the system updates the Q network according to the reward function feedback until the set number of rounds or the convergence condition is reached, and outputs the final optimized prompt word.

[0036] Step 7: Each time the "add" or "replace" action is performed, the system retrieves candidate phrases from the knowledge base K according to the phrase syntax category and context label, and filters them according to the similarity threshold to ensure that the edited prompt word is contextually coherent and grammatically correct.

[0037] The preferred embodiments of the present application are described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and changes without creative labor based on the concept of the present application. Therefore, any technical solutions obtained by logical analysis, reasoning or limited experiments based on the prior art according to the concept of the present application shall be within the scope of protection determined by the claims.

Claims

1. A knowledge-guided reinforcement learning discrete prompt optimization method, characterized in that, The method comprises the following steps: Step 1, constructing a structured phrase-level knowledge base, which groups candidate phrases by domain and syntactic type, and annotates each phrase with style score, grammatical category, context label, and additional semantic metadata; Step 2, inputting the initial prompt into a hybrid state, which is composed of the semantic embedding of the prompt and the statistical characteristics of the task performance; these statistical characteristics include the mean, standard deviation, and trend of the accuracy rate, which are spliced with the semantic embedding to serve as the input of the reinforcement learning agent; Step 3, defining discrete editing actions, including phrase-level addition, replacement, deletion, and exchange operations, and the action space is constrained by the structured knowledge base; Step 4, learning the action-value function based on the deep Q network and the experience replay mechanism, selecting actions through the epsilon-greedy strategy, and gradually optimizing the prompt; Step 5, designing a composite reward function, including the improvement of task performance and the entropy regularization term of action diversity, guiding the agent to obtain an optimization strategy that improves accuracy and edits diversity; Step 6, outputting the optimized prompt when the maximum number of steps or the convergence criterion is met.

2. The method of claim 1, wherein, The hybrid state encoding uses the pre-trained sentence vector model Sentence-BERT to obtain semantic embedding, and dynamically normalizes statistical features to achieve the expression ability and stability of state representation.

3. The method of claim 1, wherein, In addition to task performance, the reward function includes an entropy regularization term based on action distribution, encouraging strategy diversity and preventing falling into local optimum.

4. The method of claim 1, wherein, Each phrase in the knowledge base contains domain label, syntactic type, style score, applicable scenario label, and optional metadata such as reference ID and applicable jurisdiction, which are used to constrain and filter editing actions.

5. The method of claim 1, wherein, It is suitable for black box API calling scenarios where internal parameters and gradients of large language models cannot be obtained.

Citation Information

Cited By

  • Pumping well production condition analysis method based on cue word self-adaptive generation

    CN121882220A

  • A pumping unit well production condition analysis method based on prompt word adaptive generation

    CN121882220B