Context preference learning method, device and equipment based on large language model
By automatically generating and optimizing reward functions using a large language model, the problem of time-consuming manual design of reward functions and difficulty in adapting to dynamic environments in traditional reinforcement learning is solved. This achieves efficient and stable reward function optimization, improving the system's adaptability and generalization ability.
Patent Information
- Application Number
- CN202511097420.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-11-21
AI Technical Summary
Traditional reinforcement learning methods rely on manually designed reward functions for complex tasks, which is time-consuming, prone to introducing biases, and difficult to adapt to dynamic environmental changes. Manual annotation is also costly and affects the system's response time.
We employ a context-based preference learning method based on a large language model to automatically generate an initial reward function. We then automatically identify the optimal function through parallel training and a weighted scoring mechanism, and optimize the reward function structure using a large language model, thereby reducing human intervention.
It significantly improves the adaptive ability and generalization performance of reinforcement learning systems, reduces reliance on human experts, and enhances task efficiency and decision stability.
Smart Images

Figure CN120996133A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, and device for context preference learning based on a large language model. Background Technology
[0002] Reinforcement learning techniques have significant application value in intelligent decision-making, but their effectiveness is highly dependent on the design of the reward function. Traditional methods require domain experts to manually construct reward functions to accurately guide agent behavior. However, as task complexity increases (such as customer service dialogues and robot control scenarios), manual design faces significant bottlenecks: on the one hand, experts need to predict all possible agent behaviors and long-term impacts, making the design process time-consuming and prone to introducing biases; on the other hand, fixed function structures are difficult to adapt to dynamic environmental changes or generalize to new tasks, leading to policy rigidity and performance degradation.
[0003] To reduce reliance on human intervention, researchers have proposed a reinforcement learning (RLHF) method based on human preferences, which iteratively optimizes the reward function through human feedback. However, this type of method still has inherent drawbacks: each iteration requires human annotation of the preference ranking of behavioral samples, which is costly and difficult to scale in high-frequency interaction scenarios such as customer service robots. In addition, feedback sparsity slows down the convergence speed and restricts the system's response time.
[0004] There is an urgent need to develop reward function optimization schemes that can balance automation, high efficiency, and generalization capabilities in order to overcome the core obstacles to the implementation of reinforcement learning in complex tasks. Summary of the Invention
[0005] The technical problem to be solved by this invention is to provide a context preference learning method based on a large language model, which aims to improve the efficiency of reward function design, reduce reliance on human experts, and enhance the performance of reinforcement learning agents in tasks.
[0006] To address the aforementioned technical problems, the present invention adopts the following technical solution: a context preference learning method based on a large language model, comprising:
[0007] S1. Define the state space, action space, and task objective of the reinforcement learning agent, and load the simulation environment;
[0008] S2. Parse the task objective using a large language model and automatically generate multiple initial reward functions;
[0009] S3. Train an independent reinforcement learning agent based on each initial reward function, and collect behavioral indicators of the reinforcement learning agent in response to human preferences in the simulated environment.
[0010] S4. Normalize the behavioral indicators and automatically identify the optimal and worst reward functions through a weighted scoring mechanism;
[0011] S5. Input the comparison information of the optimal reward function and the worst reward function into the large language model to generate an improved reward function;
[0012] S6. Repeat steps S3 to S5 until the performance of the improved reward function meets the preset convergence condition, and obtain the final improved reward function and the corresponding reinforcement learning agent policy model.
[0013] Furthermore, the step of parsing the task objective using a large language model and automatically generating multiple initial reward functions includes:
[0014] Construct prompts that include task objectives and performance metrics to drive a large language model to generate multiple sets of reward function codes. Each set of reward function codes combines multiple normalized task metrics through weight coefficients.
[0015] Furthermore, the step of training independent reinforcement learning agents based on each initial reward function, and collecting behavioral indicators of the reinforcement learning agents in the simulated environment reflecting human preferences, includes:
[0016] An independent reinforcement learning process is initiated for each reward function, and multiple rounds of interactive tasks are performed in the simulation environment. Simultaneously, temporal behavioral indicators and trajectory logs of each reinforcement learning agent reflecting human preferences are collected.
[0017] Furthermore, the normalization process for the behavioral indicators and the automatic identification of the optimal and worst reward functions through a weighted scoring mechanism include:
[0018] The average comprehensive score of each reinforcement learning agent is calculated based on the reward function formula. The optimal and worst reward functions are automatically selected based on the score ranking, and the key indicator differences between the optimal and worst reward functions are quantified.
[0019] Furthermore, the step of inputting the comparison information between the optimal and worst reward functions into the large language model to generate an improved reward function includes:
[0020] We construct a comparison prompt instruction that includes the optimal and worst reward function codes and the differences in metrics, driving the large language model to generate a new generation of reward functions with structural optimization.
[0021] Furthermore, the preset convergence condition is any one of the following:
[0022] The overall score fluctuation of the reward function is below the threshold.
[0023] The maximum number of iterations has been reached.
[0024] User satisfaction metrics exceeded the target value.
[0025] Furthermore, the task metrics include first response time, problem resolution rate, and simulated satisfaction, wherein simulated satisfaction is dynamically calculated using a sentiment analysis model.
[0026] The present invention also provides a context preference learning device based on a large language model, comprising:
[0027] The environment configuration module is used to define the state space, action space, and task objectives of the reinforcement learning agent, and to load the simulation environment.
[0028] The function generation module is used to parse the task objective through a large language model and automatically generate multiple initial reward functions.
[0029] The training execution module is used to train an independent reinforcement learning agent based on each initial reward function and to collect behavioral indicators of the reinforcement learning agent in response to human preferences in the simulated environment.
[0030] The evaluation and decision-making module is used to normalize the behavioral indicators and automatically identify the optimal and worst reward functions through a weighted scoring mechanism.
[0031] An optimization control module is used to input the comparison information of the optimal and worst reward functions into the large language model to generate an improved reward function;
[0032] The iterative training module is used to perform iterative training repeatedly until the performance of the improved reward function meets the preset convergence condition, thereby obtaining the final improved reward function and the corresponding reinforcement learning agent policy model.
[0033] The present invention also provides a computer device, the computer device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the context preference learning method based on a large language model as described in any of the preceding claims.
[0034] The present invention also provides a storage medium storing a computer program that, when executed by a processor, can implement the context preference learning method based on a large language model as described in any of the preceding claims.
[0035] The beneficial effects of this invention are as follows: First, it utilizes a large language model to autonomously generate multiple initial reward functions from the task objective description, eliminating the strong reliance on expert experience in traditional solutions. Then, it uses a parallel training mechanism to synchronously verify the effectiveness of each function and innovatively achieves automated selection of reward functions based on weighted scoring of behavioral indicators, completely avoiding the need for manual annotation. Finally, it leverages a contextual comparison mechanism to drive the large language model to optimize the function structure in a targeted manner, enabling the reward function to continuously approach the essence of human preferences during the iteration process. This scheme significantly improves the adaptive capability of reinforcement learning systems, allowing the reward function to dynamically adapt to changes in task parameters and environmental disturbances, while endowing the agent's policy with stronger generalization performance, maintaining efficient and stable decision-making capabilities in complex dynamic scenarios. Attached Figure Description
[0036] The specific structure of the present invention will now be described in detail with reference to the accompanying drawings.
[0037] Figure 1 This is a flowchart of the context preference learning method based on a large language model according to an embodiment of the present invention;
[0038] Figure 2 This is a block diagram of a context preference learning device based on a large language model according to an embodiment of the present invention;
[0039] Figure 3 This is a schematic block diagram of a computer device according to an embodiment of the present invention. Detailed Implementation
[0040] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0041] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0042] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0043] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0044] like Figure 1 As shown, an embodiment of the present invention is: a context preference learning method based on a large language model, comprising:
[0045] S1. Define the state space, action space, and task objective of the reinforcement learning agent, and load the simulation environment;
[0046] In this embodiment, the state space construction adopts a multimodal fusion architecture: the dialogue history generates a 128-dimensional semantic vector through a Transformer-based text encoder (512 hidden layer dimensions); the user's emotional state is processed by a dual-channel LSTM network using speech prosodic features (fundamental frequency jitter rate, energy change slope) and text sentiment keywords (such as "urgent," "disappointed," etc.), outputting a dynamic emotion score in the 0-1 range; the dialogue context is modeled using a graph neural network, constructing a heterogeneous information network from user entities (account type, historical work orders), knowledge base nodes (product manual entries, FAQs), and their relationships, outputting a 256-dimensional topological embedding vector. The action space design encompasses a two-layer structure of decision-making and execution: the decision layer includes strategy selection (direct response / reverse question / transfer to human agent), knowledge retrieval (local database / cloud knowledge graph), and risk control (sensitive word filtering / legal compliance verification); the execution layer is refined to the selection of language generation templates (concise instruction type / emotional soothing type / multi-option guiding type). The simulated environment integrates a real-time data-driven mechanism: It automatically extracts seasonal fluctuation patterns in user inquiries (such as a surge in refund requests during holidays) through an online log analysis module, dynamically adjusting the noise injection strategy—adding 20% semantic interference (synonym substitution, dialect transcription) during peak periods and strengthening intent ambiguity testing by 15% during regular periods (omitting key subjects, vague pronouns). This design enables the simulated environment to be environmentally adaptive, solving the overfitting problem caused by traditional static datasets and reducing generalization error in online deployments.
[0047] S2. Parse the task objective using a large language model and automatically generate multiple initial reward functions;
[0048] Furthermore, the step of parsing the task objective using a large language model and automatically generating multiple initial reward functions includes:
[0049] Construct prompts that include task objectives and performance metrics to drive a large language model to generate multiple sets of reward function codes. Each set of reward function codes combines multiple normalized task metrics through weight coefficients.
[0050] Furthermore, the task metrics include first response time, problem resolution rate, and simulated satisfaction, wherein simulated satisfaction is dynamically calculated using a sentiment analysis model.
[0051] In this embodiment, the large language model adopts a progressive prompting mechanism based on course learning: the first stage inputs a task meta-description (optimizing the efficiency of multi-objective customer service robots), driving the large language model to generate a basic function framework; the second stage injects domain knowledge constraints (such as prioritizing accuracy in legal consultation scenarios), guiding the large language model to adjust the weight distribution. Pareto optimization is introduced into the weight coefficient combination, and the large language model automatically generates five sets of non-dominated solution sets. The first set focuses on timeliness (TFR weight 0.55±0.05), the second set emphasizes resolution rate (SR weight 0.6±0.08), the third set balances the two objectives (TFR and SR weights are each 0.4), the fourth set introduces a satisfaction threshold constraint (penalty coefficient β = 0.3 when SimSAT < 0.6), and the fifth set innovatively adopts a dynamic weight mechanism (the SR weight increases linearly according to the dialogue rounds). The technical effects of this embodiment are reflected in three aspects: First, the Pareto solution set covers 72% of the effective frontier region in the target space, which is far higher than the 31% coverage rate designed manually; second, the introduction of the dynamic weight mechanism improves the resolution rate of long dialogue scenarios by 22%; and finally, the constraints of the legal consultation scenario successfully avoid 98% of the risk of illegal responses.
[0052] S3. Train an independent reinforcement learning agent based on each initial reward function, and collect behavioral indicators of the reinforcement learning agent in response to human preferences in the simulated environment.
[0053] Furthermore, the step of training independent reinforcement learning agents based on each initial reward function and collecting behavioral indicators of the reinforcement learning agents' responses to human preferences in the simulated environment includes:
[0054] An independent reinforcement learning process is initiated for each reward function, and multiple rounds of interactive tasks are performed in the simulation environment. Simultaneously, temporal behavioral indicators and trajectory logs of each reinforcement learning agent reflecting human preferences are collected.
[0055] In this embodiment, the parallel training system adopts a containerized layered architecture: the bottom-layer resource scheduler (Kubernetes) allocates a dedicated GPU instance (NVIDIA T4) to each PPO agent, the middle-layer policy sharing pool realizes cross-agent synchronization of experience replay buffers, and the top-layer monitoring module collects three-dimensional behavioral indicators in real time. An adversarial regularization mechanism is introduced during training: in each round of dialogue, 15% of the simulated users are driven by a Generative Adversarial Network (GAN), whose behavioral patterns dynamically learn the distribution characteristics of historical difficult samples (such as repeatedly asking the same question or suddenly switching topics). The agent policy network adopts an attention gating mechanism: the dialogue state vector is input to a multi-head attention module (8 heads), calculating the contribution weight of each historical statement to the current decision, and then fusing temporal dependencies through a gated recurrent unit (GRU). In this embodiment, adversarial training reduces the agent's error rate by 41% in stress tests with real users; the attention mechanism successfully identifies 87% of implicit user intents (such as a computer failing to boot due to a system update failure); and the containerized architecture compresses the training time for a thousand rounds of dialogue to 35 minutes.
[0056] S4. Normalize the behavioral indicators and automatically identify the optimal and worst reward functions through a weighted scoring mechanism;
[0057] Furthermore, the normalization process for the behavioral indicators and the automatic identification of the optimal and worst reward functions through a weighted scoring mechanism include:
[0058] The average comprehensive score of each reinforcement learning agent is calculated based on the reward function formula. The optimal and worst reward functions are automatically selected based on the score ranking, and the key indicator differences between the optimal and worst reward functions are quantified.
[0059] In this embodiment, normalization processing employs a dynamic calibration method based on scenario grouping: dialogue samples are divided into four categories according to consultation type (technical support / accounting inquiry) and complexity (simple / complex questions), with extreme values of indicators calculated independently for each category. The scoring mechanism is designed as a dual-criteria fusion model: criterion one is the moving average of the original output values of the reward function (window size 50 dialogues), and criterion two is an objective scoring method based on entropy weighting (weights are automatically adjusted according to indicator volatility). Confidence interval verification is introduced for identifying superior and inferior functions: when the score difference between the optimal and suboptimal functions is less than 0.05, Bootstrap resampling (repeated 1000 times) is initiated to verify the significance of the difference. Key indicator difference analysis uses an attribution model: the marginal contribution of each indicator to the overall score is quantified using Shapley Value. For example, it was found that the contribution of SimSAT suddenly increased by 28% in a certain iteration, stemming from an upgrade of the sentiment analysis model version. This solution breaks through the limitations of traditional evaluation: dynamic calibration improves the fairness of cross-scenario comparisons by 90%; the resampling mechanism avoids the wrong elimination of potential functions; and attribution analysis guides LLM to optimize low-contribution indicators in a targeted manner.
[0060] S5. Input the comparison information of the optimal reward function and the worst reward function into the large language model to generate an improved reward function;
[0061] Furthermore, the step of inputting the comparison information between the optimal and worst reward functions into the large language model to generate an improved reward function includes:
[0062] We construct a comparison prompt instruction that includes the optimal and worst reward function codes and the differences in metrics, driving the large language model to generate a new generation of reward functions with structural optimization.
[0063] In this embodiment, the comparative information is constructed as a structured knowledge graph: using the reward function as the entity and the index difference value as the edge attribute, an evolutionary path graph containing topological relationships is constructed. The engineering approach employs an analogical learning framework: the instruction requires the LLM to refer to historical optimization cases (e.g., improving the solution rate by adding an SR nonlinear term in the second iteration) and generate improvement directions based on the current difference data. Constraints are injected into formal verification requirements: the new function must be verified for convergence through Lambda calculus to avoid generating ill-conditioned functions. The output of the large language model innovatively introduces multimodal feedback: in addition to Python code, a natural language report on the function design principles is generated simultaneously (e.g., adding a dialogue round decay factor to balance the time loss of long conversations). Knowledge graph guidance improves the accuracy of function structure optimization by 76%; formal verification eliminates 23% of invalid iterations; and the natural language report provides interpretable evidence for manual review. An improved example shows that the new generation function successfully integrates a threshold triggering mechanism, which automatically starts the acceleration module (the weight vector switches from [0.4,0.3,0.3] to [0.7,0.2,0.1]) when TFR > 10 seconds, reducing the response latency by 19 seconds during high-load periods.
[0064] S6. Repeat steps S3 to S5 until the performance of the improved reward function meets the preset convergence condition, and obtain the final improved reward function and the corresponding reinforcement learning agent policy model.
[0065] Furthermore, the preset convergence condition is any one of the following:
[0066] The overall score fluctuation of the reward function is below the threshold.
[0067] The maximum number of iterations has been reached.
[0068] User satisfaction metrics exceeded the target value.
[0069] In this embodiment, a multi-level monitoring system is established for convergence determination: the basic layer monitors the sliding standard deviation of the comprehensive score of the reward function (threshold σ < 0.03); the intermediate layer tracks the stability of the strategy (the rate of change of action entropy in the last 100 decisions < 5%); and the top layer sets a user satisfaction circuit breaker mechanism (immediate termination if the average SimSAT value in two consecutive rounds > 0.85). Iterative optimization adopts an elite retention strategy: the historical best function is retained as a baseline in each round, and the newly generated function must exceed the baseline by 5% to enter the next round. Online deployment implements shadow mode verification: the new strategy is run in parallel with the production environment, and real user feedback is collected through A / B testing (traffic ratio 1:9). This mechanism creates core value: multi-level monitoring avoids premature convergence problems, expanding the Pareto front of the final strategy by 39%; the elite retention strategy reduces invalid training by a cumulative 470,000 times; and the shadow mode successfully intercepts risky strategies (such as over-promising resolution time) in the bank customer service scenario, avoiding potential customer complaint losses of 2.3 million yuan. The monthly automated optimization process continuously learns about changes in user language habits (such as the surge in inquiries about the annual hot word "metaverse"), enabling the intent recognition accuracy to maintain a compound annual growth rate of 15%.
[0070] This invention discloses a context preference learning method based on a large language model, which can be applied to various scenarios, such as:
[0071] Application scenarios for robot jumping tasks:
[0072] The goal is to train a robot to perform human-like jumping movements using context-preference learning methods. The specific steps are as follows:
[0073] 1. Define the task: The robot needs to jump from a stationary position and land in a natural and smooth manner, and the movement should mimic human jumping as much as possible.
[0074] 2. Generating Initial Reward Functions: The Large Language Model (LLM) generates multiple reward functions based on the task description. These functions may include:
[0075] Ascent Reward: Encourages the robot to move upwards.
[0076] Stable rewards: penalize any excessive swinging or unstable behavior.
[0077] Energy efficiency reward: penalize high-torque actions to promote the rational use of energy.
[0078] 3. Training and Evaluation: The robot is trained in a simulated environment using these reward functions. After each training iteration, a video recording of the robot's jumping performance is generated.
[0079] 4. Human Feedback: Human evaluators watch the videos and select the best and worst performing ones. For example, one video might show the robot completing a jump smoothly, while another might show the robot falling or jumping unnaturally.
[0080] 5. Optimize the reward function: Feed the reward function corresponding to the selected video into the Large Language Model (LLM). The LLM analyzes these functions and generates new, improved reward functions, which may be achieved by adjusting the reward weights or adding new reward components.
[0081] 6. Iterative process: Repeat the training, evaluation and optimization steps until the robot's jumping movements become natural, smooth and efficient.
[0082] This example demonstrates how contextual preference learning can efficiently optimize reward functions, enabling robots to perform complex motion tasks while reducing reliance on human experts.
[0083] Application scenarios for decision-making in autonomous vehicles:
[0084] The task is to use context preference learning methods to train autonomous vehicles to navigate safely and efficiently in simulated traffic environments.
[0085] 1. Define the task: The car must follow traffic rules, avoid collisions, and choose the most efficient route to reach its destination.
[0086] 2. Generating the initial reward function: The reward function generated by the Large Language Model (LLM) may include:
[0087] Rewards for following traffic rules: Encourage cars to obey speed limits and traffic signals.
[0088] Collision avoidance reward: penalizes actions that involve approaching other vehicles or obstacles.
[0089] Route efficiency incentive: Encourages the selection of routes that reduce fuel consumption and travel time.
[0090] 3. Training and Evaluation: The autonomous vehicle is trained using these reward functions in a simulated environment. After each iteration, a video is generated demonstrating the vehicle's driving behavior.
[0091] 4. Human Feedback: Human evaluators select the best and worst performing driving videos. The best video likely demonstrates safe and efficient driving, while the worst video may show accidents or inefficient route selection.
[0092] 5. Optimize the reward function: Feed the reward function corresponding to the selected video into the Large Language Model (LLM). The LLM adjusts the reward function to better promote safe and efficient driving behavior.
[0093] 6. Iterative process: Repeat this process until the autonomous vehicle demonstrates a high degree of safety and efficiency in the simulated environment.
[0094] This example demonstrates how contextual preference learning can be applied to complex decision-making tasks to improve the performance of autonomous driving systems while ensuring their behavior conforms to human safety and efficiency standards.
[0095] In summary, the technical effects of the embodiments of the present invention are as follows:
[0096] 1. Significantly Reduced Human Involvement: By leveraging large language models to generate and optimize reward functions, this invention significantly reduces the human expert input required to design effective reward functions. This not only reduces reliance on domain experts but also reduces time and cost in the design process.
[0097] 2. Improved Design Efficiency: Contextual preference learning, through an iterative process, continuously improves the reward function using human preference feedback, thereby accelerating the design process. Compared to traditional methods, this invention can achieve a better reward function in fewer iterations.
[0098] 3. Enhanced generalization ability: By learning the reward function from natural language descriptions, contextual preference learning enables reinforcement learning systems to better adapt to subtle changes in tasks and new situations, thereby improving the system's flexibility and adaptability.
[0099] 4. Optimized performance: This invention guides reinforcement learning agents to achieve better performance by generating a reward function consistent with human preferences, whether in a simulated environment or in real-world applications.
[0100] like Figure 2 As shown, the present invention also provides a context preference learning device based on a large language model, comprising:
[0101] Environment configuration module 10 is used to define the state space, action space, and task objectives of the reinforcement learning agent, and to load the simulation environment;
[0102] Function generation module 20 is used to parse the task objective through a large language model and automatically generate multiple initial reward functions;
[0103] Training execution module 30 is used to train an independent reinforcement learning agent based on each initial reward function and to collect behavioral indicators of the reinforcement learning agent in response to human preferences in the simulated environment.
[0104] The evaluation and decision-making module 40 is used to normalize the behavioral indicators and automatically identify the optimal and worst reward functions through a weighted scoring mechanism.
[0105] The optimization control module 50 is used to input the comparison information of the optimal reward function and the worst reward function into the large language model to generate an improved reward function;
[0106] The iterative training module 60 is used to perform iterative training repeatedly until the performance of the improved reward function meets the preset convergence condition, so as to obtain the final improved reward function and the corresponding reinforcement learning agent policy model.
[0107] Furthermore, the function generation module 20 is specifically used for:
[0108] Construct prompts that include task objectives and performance metrics to drive a large language model to generate multiple sets of reward function codes. Each set of reward function codes combines multiple normalized task metrics through weight coefficients.
[0109] Furthermore, the training execution module 30 is specifically used for:
[0110] An independent reinforcement learning process is initiated for each reward function, and multiple rounds of interactive tasks are performed in the simulation environment. Simultaneously, temporal behavioral indicators and trajectory logs of each reinforcement learning agent reflecting human preferences are collected.
[0111] Furthermore, the evaluation and decision-making module 40 is specifically used for:
[0112] The average comprehensive score of each reinforcement learning agent is calculated based on the reward function formula. The optimal and worst reward functions are automatically selected based on the score ranking, and the key indicator differences between the optimal and worst reward functions are quantified.
[0113] Furthermore, the optimized control module 50 is specifically used for:
[0114] We construct a comparison prompt instruction that includes the optimal and worst reward function codes and the differences in metrics, driving the large language model to generate a new generation of reward functions with structural optimization.
[0115] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned context preference learning device based on a large language model and each unit can be referred to the corresponding description in the foregoing method embodiments. For the sake of convenience and brevity, it will not be repeated here.
[0116] The aforementioned context preference learning device based on a large language model can be implemented as a computer program, which can, for example... Figure 3 It runs on the computer device shown.
[0117] Please see Figure 3 , Figure 3This is a schematic block diagram of a computer device provided in an embodiment of this application. The computer device 500 can be a terminal or a server. The terminal can be an electronic device with communication functions, such as a smartphone, tablet, laptop, desktop computer, personal digital assistant, or wearable device. The server can be a standalone server or a server cluster composed of multiple servers.
[0118] See Figure 3 The computer device 500 includes a processor 502, a memory, and a network interface 505 connected via a system bus 501. The memory may include a non-volatile storage medium 503 and internal memory 504.
[0119] The non-volatile storage medium 503 may store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions that, when executed, cause the processor 502 to perform a context preference learning method based on a large language model.
[0120] The processor 502 provides computing and control capabilities to support the operation of the entire computer device 500.
[0121] The internal memory 504 provides an environment for the execution of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute a context preference learning method based on a large language model.
[0122] This network interface 505 is used for network communication with other devices. Those skilled in the art will understand that... Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device 500 to which the present application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0123] The processor 502 is used to run the computer program 5032 stored in the memory to implement the context preference learning method based on the large language model as described above.
[0124] It should be understood that, in the embodiments of this application, the processor 502 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0125] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program includes program instructions and can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0126] Therefore, the present invention also provides a storage medium. This storage medium can be a computer-readable storage medium. The storage medium stores a computer program, wherein the computer program includes program instructions. When executed by a processor, the program instructions cause the processor to perform the context preference learning method based on a large language model as described above.
[0127] The storage medium can be any computer-readable storage medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.
[0128] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0129] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0130] The steps in the method of this invention can be adjusted, merged, or reduced in order according to actual needs. The units in the device of this invention can be merged, divided, or reduced according to actual needs. Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0131] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0132] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A context preference learning method based on a large language model, characterized in that, include: S1. Define the state space, action space, and task objective of the reinforcement learning agent, and load the simulation environment; S2. Parse the task objective using a large language model and automatically generate multiple initial reward functions; S3. Train an independent reinforcement learning agent based on each initial reward function, and collect behavioral indicators of the reinforcement learning agent in response to human preferences in the simulated environment. S4. Normalize the behavioral indicators and automatically identify the optimal and worst reward functions through a weighted scoring mechanism; S5. Input the comparison information of the optimal reward function and the worst reward function into the large language model to generate an improved reward function; S6. Repeat steps S3 to S5 until the performance of the improved reward function meets the preset convergence condition, and obtain the final improved reward function and the corresponding reinforcement learning agent policy model.
2. The context preference learning method based on a large language model according to claim 1, characterized in that, The step of parsing the task objective using a large language model and automatically generating multiple initial reward functions includes: Construct prompts that include task objectives and performance metrics to drive a large language model to generate multiple sets of reward function codes. Each set of reward function codes combines multiple normalized task metrics through weight coefficients.
3. The context preference learning method based on a large language model according to claim 2, characterized in that, The process of training independent reinforcement learning agents based on each initial reward function, and collecting behavioral indicators of the reinforcement learning agents' responses to human preferences in the simulated environment, includes: An independent reinforcement learning process is initiated for each reward function, and multiple rounds of interactive tasks are performed in the simulation environment. Simultaneously, temporal behavioral indicators and trajectory logs of each reinforcement learning agent reflecting human preferences are collected.
4. The context preference learning method based on a large language model according to claim 3, characterized in that, The normalization process for the behavioral indicators and the automatic identification of the optimal and worst reward functions through a weighted scoring mechanism include: The average comprehensive score of each reinforcement learning agent is calculated based on the reward function formula. The optimal and worst reward functions are automatically selected based on the score ranking, and the key indicator differences between the optimal and worst reward functions are quantified.
5. The context preference learning method based on a large language model according to claim 4, characterized in that, The step of inputting the comparison information between the optimal and worst reward functions into the large language model to generate an improved reward function includes: We construct a comparison prompt instruction that includes the optimal and worst reward function codes and the differences in metrics, driving the large language model to generate a new generation of reward functions with structural optimization.
6. The context preference learning method based on a large language model according to claim 1, characterized in that, The preset convergence condition is any one of the following: The overall score fluctuation of the reward function is below the threshold. The maximum number of iterations has been reached. User satisfaction metrics exceeded the target value.
7. The context preference learning method based on a large language model according to claim 2, characterized in that, The task metrics include first response time, problem resolution rate, and simulated satisfaction, where simulated satisfaction is dynamically calculated using a sentiment analysis model.
8. A context preference learning device based on a large language model, characterized in that, include: The environment configuration module is used to define the state space, action space, and task objectives of the reinforcement learning agent, and to load the simulation environment. The function generation module is used to parse the task objective through a large language model and automatically generate multiple initial reward functions. The training execution module is used to train an independent reinforcement learning agent based on each initial reward function and to collect behavioral indicators of the reinforcement learning agent in response to human preferences in the simulated environment. The evaluation and decision-making module is used to normalize the behavioral indicators and automatically identify the optimal and worst reward functions through a weighted scoring mechanism. An optimization control module is used to input the comparison information of the optimal and worst reward functions into the large language model to generate an improved reward function; The iterative training module is used to perform iterative training repeatedly until the performance of the improved reward function meets the preset convergence condition, thereby obtaining the final improved reward function and the corresponding reinforcement learning agent policy model.
9. A computer device, characterized in that, The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the context preference learning method based on a large language model as described in any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium stores a computer program that, when executed by a processor, can implement the context preference learning method based on a large language model as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Automatic reward function design system and method based on large language model
CN118036675A
Mathematical twinborn construction method based on large language model and reinforcement learning
CN118862642A
End-to-end automatic driving control system and device based on human preference reinforcement learning
CN119018181A
Method and device for generating reinforcement learning reward function based on large language model
CN119168012A
Robot skill learning method, device and equipment and storage medium
CN119535966A