Recommendation strategy generation method based on large language model enhancement and related equipment

By combining the large language model with reinforcement learning agents and using the large language model to predict user behavior, the problem that traditional methods are difficult to quickly provide accurate recommendation strategies is solved, and efficient and personalized recommendation effects are achieved.

CN119513292BActive Publication Date: 2025-05-13INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411250877.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-06
Publication Date
2025-05-13
Estimated Expiration
2044-09-06

AI Technical Summary

Technical Problem

Traditional recommendation strategies based on reinforcement learning are difficult to quickly provide accurate recommendations, especially in cold start scenarios or rapid changes in user interests.

Method used

Combining large language model with reinforcement learning agents, predict user behavior of sample recommendation tasks through large language model, providing prior knowledge to agents, thereby accelerating the learning and strategy optimization of agents.

Benefits of technology

It significantly improves the accuracy and personalization of the recommendation strategy, shortens the convergence time of the strategy, improves the efficiency and adaptability of the recommendation system, and can better cope with changes in user interests.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119513292B_ABST
    Figure CN119513292B_ABST
Patent Text Reader

Abstract

The present invention provides a recommendation strategy generation method based on large language model enhancement and related equipment. The method includes obtaining a recommendation task; inputting the recommendation task into an intelligent agent, and obtaining a recommendation strategy output by the intelligent agent; the intelligent agent is trained based on sample recommendation tasks and a large language model, and the large language model is used to predict user behavior based on the sample recommendation task. By combining a large language model with a reinforcement learning intelligent agent, the present invention effectively solves the problem that traditional methods are difficult to quickly provide accurate recommendation strategies, and achieves the improvement of the accuracy and personalization of the recommendation strategy while also improving the efficiency and adaptability of the recommendation system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of learning algorithm technology, and in particular to a recommendation strategy generation method based on large language model enhancement and related equipment. Background Art

[0002] With the rapid development of Internet technology and the explosive growth of digital content, recommendation systems play an increasingly important role in various online platforms. Recommendation systems aim to provide users with personalized content recommendations by analyzing user behavior and content features. Among the many recommendation methods, recommendation strategies based on reinforcement learning have attracted widespread attention because they can dynamically adapt to changes in user interests.

[0003] Existing reinforcement learning-based recommendation strategies usually use traditional reinforcement learning algorithms to continuously optimize strategies through continuous interaction with users' recommendation tasks to maximize long-term cumulative rewards. During the interaction process, the reinforcement learning-based agent learns through trial and error and gradually builds an understanding of users' interests and behavior patterns.

[0004] However, traditional reinforcement learning methods often require a large amount of interaction data to learn effective strategies, which may lead to a long convergence time in practical applications. Especially in cold start scenarios or when user interests change rapidly, traditional methods are difficult to quickly provide accurate recommendation strategies. Summary of the invention

[0005] The present invention provides a recommendation strategy generation method based on large language model enhancement and related equipment. By combining the large language model with a reinforcement learning agent, the problem that traditional methods are difficult to quickly provide accurate recommendation strategies is effectively solved. While improving the accuracy and personalization of the recommendation strategy, the efficiency and adaptability of the recommendation system are also improved.

[0006] The present invention provides a recommendation strategy generation method based on large language model enhancement, comprising:

[0007] Get recommended tasks;

[0008] Inputting the recommendation task into an intelligent agent, and obtaining a recommendation strategy output by the intelligent agent;

[0009] The intelligent agent is trained based on a sample recommendation task and a large language model, and the large language model is used to predict user behavior based on the sample recommendation task.

[0010] Optionally, inputting the recommendation task into an intelligent agent and obtaining a recommendation strategy output by the intelligent agent includes:

[0011] Obtaining the status of the user corresponding to the recommended task through the agent;

[0012] Obtaining the user's action based on the user's state through the agent;

[0013] A recommendation strategy is generated by the agent based on the user's actions.

[0014] Optionally, the training process of the agent includes:

[0015] Inputting a sample recommendation task into the intelligent agent, and obtaining multiple groups of sample interaction data output by the intelligent agent;

[0016] Obtaining a predicted behavior output by the large language model based on the sample recommendation task;

[0017] For each set of the sample interaction data, obtaining a target expected value based on the interaction data and the corresponding predicted behavior;

[0018] Based on the loss function, obtaining a difference value of the target expected value;

[0019] adjusting parameters of the agent based on the difference value;

[0020] The large language model is trained based on the sample recommendation task.

[0021] Optionally, the interaction data includes a sample state, a sample action, a sample user behavior, and a sample reward at a current moment, and a sample state at a next moment, and obtaining a target expected value based on the interaction data and the corresponding predicted behavior includes:

[0022] Based on the sample state, sample action, and sample user behavior at the current moment, obtaining a first expected value of the agent at the current moment;

[0023] Based on the sample state and predicted behavior at the next moment, obtaining the maximum expected value of the agent at the next moment;

[0024] The first expected value is adjusted based on the sample reward and the maximum expected value to obtain a target expected value.

[0025] Optionally, the loss function includes a first loss function and a second loss function, wherein:

[0026] ;

[0027] In the formula, represents the first loss value, B represents multiple groups of sample interaction data, It means taking the average value of multiple groups of sample interaction data. Respectively represent the current sample state, sample action, and sample user behavior, represents the target expected value under the recommended strategy, represents the sample reward at the current moment, Indicates the sample state at the next moment;

[0028] The second loss function is used to characterize the difference of the value evaluation network in the agent, including:

[0029] ;

[0030] In the formula, represents the second loss value, represents the discount factor, which is used to balance the impact of sample rewards and future rewards. represents the expected value of the sample at the next moment, They represent the sample state, sample action, and predicted behavior at the next moment respectively.

[0031] Optionally, the training process of the large language model includes:

[0032] Inputting the sample recommendation task into the large language model to obtain the predicted behavior output by the large language model;

[0033] Obtaining a low-rank matrix corresponding to the predicted behavior;

[0034] The weights of the large language model are adjusted based on the low-rank matrix.

[0035] Optionally, inputting the sample recommendation task into the large language model to obtain a predicted behavior output by the large language model includes:

[0036] generating a task instruction and a binary interaction sequence based on the recommendation task, wherein the binary interaction sequence is used to represent the user's preference for the recommended data in the recommendation task;

[0037] The task instruction and the binary historical interaction sequence are input into the large language model to obtain a predicted behavior output by the large language model, where the predicted behavior is a binary classification result corresponding to the task instruction.

[0038] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for generating a recommendation strategy based on large language model enhancement as described above is implemented.

[0039] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described methods for generating a recommendation strategy based on large language model enhancement.

[0040] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the above-mentioned methods for generating a recommendation strategy based on large language model enhancement.

[0041] In summary, one or more technical solutions provided in the present invention have at least the following technical effects or advantages:

[0042] The present invention effectively solves the problem that traditional methods are difficult to quickly provide accurate recommendation strategies by combining a large language model with a reinforcement learning agent. This improves the accuracy and personalization of the recommendation strategy while also improving the efficiency and adaptability of the recommendation system.

[0043] The present invention first uses a large language model to predict user behavior for sample recommendation tasks, providing the intelligent agent with rich prior knowledge. Subsequently, through reinforcement learning training, the intelligent agent can make full use of this prior knowledge to quickly learn and adapt to different recommendation tasks.

[0044] The above combination significantly improves the efficiency and accuracy of the agent learning recommendation strategies. When faced with new recommendation tasks, the trained agent can quickly output an adaptable recommendation strategy and achieve high-quality recommendations without a large amount of interactive data.

[0045] Compared with traditional methods, this invention greatly shortens the strategy convergence time, improves the recommendation effect in cold start scenarios, and can better cope with the situation where user interests change rapidly. In addition, since the agent has internalized the knowledge of the large language model through reinforcement learning training, there is no need to deploy a large language model in practical applications, thereby reducing system complexity and computing resource requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0047] Figure 1 It is a flowchart of a recommendation strategy generation method based on large language model enhancement provided by the present invention.

[0048] Figure 2 It is a schematic diagram of a training architecture of an intelligent agent provided by the present invention.

[0049] Figure 3 It is a flow chart of an intelligent agent training method provided by the present invention.

[0050] Figure 4 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0051] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0052] Before introducing the embodiments of the present invention, some terms involved in the embodiments of the present invention are first defined and explained:

[0053] (1) Intelligent agent.

[0054] An agent refers to an entity that can perceive the environment, make decisions, and learn from environmental feedback in a reinforcement learning framework. It optimizes its behavior strategy through continuous interaction with the environment to maximize long-term cumulative rewards. In the embodiment of the present invention, it can be understood as a complex decision-making system composed of a deep neural network, which includes core components such as a state encoder, a policy network, a value assessment network, an experience replay pool, and an exploration mechanism.

[0055] The agent plays a core role in the recommendation system. It first perceives the recommendation environment through the state encoder, converting user information, historical interaction data, and candidate item pools into vector representations that can be processed by the neural network. Based on these encoded state information, the policy network generates a recommendation strategy and outputs the probability distribution of recommended actions, such as selecting specific items or item sequences for recommendation. At the same time, the value evaluation network evaluates the value of each possible action in the current state to assist the policy network in making better decisions.

[0056] After performing the recommendation action, the agent will observe the user's interactive behavior on the recommended item (such as clicking, purchasing, pulling down, etc.) and calculate the immediate reward based on these feedbacks. These interactive experiences are stored in the experience replay pool for subsequent offline learning. By continuously interacting and learning with the environment, the agent can continuously optimize its internal models, including the policy network and the value evaluation network.

[0057] A key feature of the intelligent agent is its ability to adapt to dynamic environments. As user interests change and new items are added, it can adjust its recommendation strategy in real time. In addition, the intelligent agent needs to balance multiple goals, such as improving user satisfaction while considering multiple factors such as recommendation diversity and system efficiency.

[0058] In the present invention, the combination of reinforcement learning agent and large language model is used. The large language model helps the agent to better understand the semantics of content and user interests, so as to gradually learn and optimize the recommendation strategy in a complex and changing recommendation environment. The above combination not only improves the content understanding ability of the agent, but also enhances its ability to handle cold start problems and provide personalized recommendations.

[0059] (2) Large Language Model

[0060] A large language model refers to a large-scale natural language processing model trained based on deep learning technology, which has powerful semantic understanding, contextual reasoning and knowledge extraction capabilities. This type of model is usually based on the Transformer architecture and uses massive text data for pre-training to obtain extensive language knowledge and world knowledge. In an embodiment of the present invention, the large language model can be understood as a powerful semantic analysis and user behavior prediction tool. As an auxiliary system for reinforcement learning agents, it enhances the semantic understanding ability and prediction accuracy of the entire recommendation system by deeply understanding and analyzing the text information in the recommendation task.

[0061] Based on the above, the embodiment of the present invention provides a recommendation strategy generation method based on large language model enhancement, please refer to Figure 1 , Figure 1 This is a flowchart of a method for generating a recommendation strategy based on a large language model enhancement provided by an embodiment of the present invention. The method can be implemented by a computer program, which can be integrated into an application or run as an independent tool application. The method can also be implemented by a single-chip microcomputer or run in a recommendation strategy generation system based on a von Neumann system. Specifically, the method can include the following steps:

[0062] Step 101: Get recommended tasks.

[0063] The recommendation task refers to the process of selecting and presenting the most relevant and valuable content or items to users based on user information and system status in a specific scenario. It aims to filter out the options that best meet the user's interests and needs from a large number of candidate items by analyzing user preferences and behavior patterns. In the embodiment of the present invention, the recommendation task can be understood as a structured data set that contains all the key information required to perform the recommendation.

[0064] The above data set mainly consists of user information u, historical interaction sequence = , ... , , candidate item pool I, recommendation requirements and scenario information. User information u includes the user's unique identifier and related attributes; historical interaction sequence It reflects the user's interest changes and behavior patterns; the candidate item pool I contains independent items that can be recommended; the recommendation requirements specify constraints such as the length, timeliness, and diversity of the recommendation list; and the scene information describes the specific environment in which the recommendation occurs.

[0065] The above structured recommendation task provides the reinforcement learning agent with an initial state s, enabling the agent to understand the current recommendation environment and user needs, and thus make appropriate recommendation decisions. At the same time, it also guides the large language model to understand content and predict user behavior, enabling the model to more accurately analyze user interests and item features.

[0066] In addition, the definition of the recommendation task also directly affects the design of the action space and reward function of reinforcement learning. Based on the recommendation requirements specified in the task, the system can clearly define the scope of the recommendation action and the criteria for evaluating the recommendation effect. This not only provides an evaluation benchmark for the recommendation system, but also supports personalized and context-aware recommendations, enabling the system to consider the user's personal characteristics and current scenario to generate more targeted recommendations.

[0067] In summary, the recommendation task in the present invention serves as both input data and a guiding framework for the recommendation process, ensuring that the recommendation system can accurately understand user needs and make optimal recommendation decisions in a complex and changing environment.

[0068] Step 102: input the recommendation task into the intelligent agent, and obtain the recommendation strategy output by the intelligent agent, wherein the intelligent agent is trained based on the sample recommendation task and the large language model, and the large language model is used to predict user behavior based on the sample recommendation task.

[0069] Among them, recommendation strategy refers to a series of decisions and action plans taken by the recommendation system in a specific scenario to achieve the established goals. It is the specific method and rules for the recommendation system to select and sort the recommended content from the candidate items based on factors such as user characteristics, historical behavior, current environment, etc. A good recommendation strategy can balance the user's short-term interests and long-term preferences while considering the overall goals of the system, such as user satisfaction, click-through rate, conversion rate, etc.

[0070] In the embodiment of the present invention, the recommendation strategy can be understood as the intelligent agent being able to output the corresponding recommended action a based on the current state s. This action a represents a specific scheme for selecting K items from the candidate item pool I as the recommendation list. More specifically, it can be represented as a K-dimensional vector, where each element corresponds to an item index in the candidate item pool I. The above vector not only contains the selected items, but also implies the recommended order of these items.

[0071] Specifically, because traditional content encoders often have difficulty accurately obtaining the semantic representation of project content and user dynamic interests, the stability of reinforcement learning recommendation strategies is poor, and the intelligent agent can only learn suboptimal strategies. The present invention cleverly solves this technical problem by introducing a large language model. With its powerful semantic understanding ability, the large language model can deeply analyze project content and user behavior, providing the intelligent agent with a more accurate and rich semantic representation.

[0072] Specifically, when the recommendation task is input to the agent, the large language model first analyzes the user information u in the task, the historical interaction sequence x 1:t The deep semantic analysis is performed on the candidate item pool I. The above analysis is not just a simple feature extraction, but a comprehensive semantic understanding and context association of the content. Based on the above understanding, the large language model can predict the user's possible behavior and provide important prior knowledge for the decision-making of the intelligent agent.

[0073] When generating a recommendation strategy, the agent not only considers its own state representation, but also combines the user behavior prediction provided by the large language model. This enables the agent to consider historical data and current state when making decisions, and to foresee possible user reactions. This improves the accuracy and adaptability of the recommendation strategy.

[0074] Furthermore, the introduction of a large language model enables the system to capture more subtle and complex semantic information, thereby generating more accurate recommendations. Secondly, the above method significantly enhances the stability of the reinforcement learning recommendation strategy. Through the high-quality semantic representation and behavior prediction provided by the large language model, the agent can converge to a high-quality strategy more quickly, avoiding the problem of falling into local optimality in traditional methods.

[0075] In addition, the present invention also greatly improves the system's ability to handle cold start problems. For new users or new items, the large language model can make reasonable inferences based on its rich prior knowledge and provide valuable decision-making basis for the intelligent agent. This enables the system to maintain good recommendation quality when facing unknown situations.

[0076] Based on the above embodiment, as an optional embodiment, step 102, inputting the recommendation task to the intelligent agent and obtaining the recommendation strategy output by the intelligent agent, may further include the following steps:

[0077] Step 201: Obtain the status of the user corresponding to the recommended task through the agent.

[0078] Specifically, the system extracts the user's basic information and historical interaction sequences from the recommendation task. Then, the interaction data is preprocessed, for example, encoding the category features and normalizing the numerical features. Next, the system uses the pre-trained embedding layer to convert these discrete and continuous features into dense vector representations.

[0079] Step 202: Obtain the user's actions based on the user's state through the agent.

[0080] Specifically, the agent inputs the user state into its policy network. The policy network is a carefully designed and trained deep neural network that can capture the complex nonlinear relationship between the state and the optimal action. The policy network outputs a probability distribution that represents the tendency to choose each possible action in the current state.

[0081] Step 203: Generate a recommendation strategy based on the user's actions through the intelligent agent.

[0082] Specifically, the process of generating a recommendation strategy first involves decoding user actions into a specific list of recommended items. The user action is a K-dimensional vector, where each element corresponds to an item index in the candidate item pool. Based on this vector, the system retrieves the corresponding K items from the candidate item pool to form a preliminary recommendation list. However, the generation of the recommendation strategy does not stop there. The system will further optimize and adjust this preliminary list.

[0083] For example, taking the movie recommendation scenario as an example, the specific implementation of steps 201 to 203 is described in detail. Assume that the user is a 25-year-old female software engineer named Xiaohong, who has recently watched several science fiction action movies and gave them high ratings. When Xiaohong opens the movie recommendation application, the system starts to generate personalized recommendations for her.

[0084] The system constructs Xiaohong's user state s. This state vector combines Xiaohong's static features (such as age, gender, occupation), historical interaction information (such as recently watched movies such as "Interstellar", "The Matrix", "Blade Runner 2049" and their ratings), and current context (Friday night, using smart TV). This information is encoded into the user state vector s, providing a comprehensive and accurate user portrait for subsequent recommendations.

[0085] Next, the agent policy network trained with the large language model receives this user state vector s as input. The policy network has learned rich semantic understanding capabilities from the large language model and can better interpret the interest tendencies contained in the user state. The policy network outputs a probability distribution representing the tendency to recommend different movies. Assuming that three movies need to be recommended, the policy network may output the recommendation probabilities of "Blade Runner" (original version), "Ready Player One", "Arrival", "iPartment" and "Titanic". The system finally selected the first three movies as the recommendation list, and the action a is represented by the index vectors corresponding to these three movies.

[0086] Finally, the system generates the final recommendation strategy based on user action a. The system first retrieves the three movies corresponding to action a: Blade Runner, Ready Player One, and Arrival. Since the policy network has acquired strong semantic understanding capabilities through the training of a large language model, it can effectively capture the thematic associations and style differences between these movies. Based on the above internal understanding, the system optimizes the recommendation order: Ready Player One is placed first, Arrival is second, and Blade Runner is the last. The above order not only takes into account Xiaohong's recent viewing preferences, but also introduces a certain degree of diversity.

[0087] Through the recommendation strategy generated by the agent policy network enhanced by the large language model training, the system can more accurately meet Xiaohong's movie-watching needs. This recommendation list not only reflects Xiaohong's preference for science fiction movies, but also provides subtle changes in themes and styles, which may inspire her to explore different types of science fiction movies. The semantic understanding ability learned by the policy network makes the recommendation results both personalized and insightful, which can bring higher user acceptance rates and more positive feedback.

[0088] The above embodiment describes the application process of the decision network of the intelligent agent. Based on the above embodiment, the training process of the intelligent agent will be described below.

[0089] Specifically, we first build an intelligent agent, including: perception layer, decision-making layer, execution layer, feedback and adjustment layer, learning layer, and deployment and monitoring layer.

[0090] The perception layer is the entrance of the entire architecture and is responsible for receiving and processing raw input data. It mainly includes two key functions:

[0091] Data input processing: Receive user status information, such as historical behavior, personal information, and current context, and perform preprocessing, including standardization, feature extraction, and dimensionality reduction, to prepare for subsequent processing.

[0092] User behavior modeling: Use large language models or other complex models to model user behavior and predict the user's potential behavior in different states. These prediction results will be passed to the decision-making layer as important input.

[0093] The decision layer is the core of the agent, responsible for generating recommendation strategies and evaluating their quality. It contains the following components:

[0094] Actor Network: Generates recommendation strategies based on a given state and selects the best recommended content to optimize the user experience.

[0095] Critic Network: Evaluates the quality of the recommended strategies generated by the strategy network and measures the possible future returns by estimating the Q value.

[0096] Behavioral decision making: Based on the output of the policy network, the recommended actions are executed in the actual environment.

[0097] Among them, the execution layer is responsible for applying the output of the decision layer to the actual system: it can provide recommended content to users based on the current status, process real-time feedback from users, and pass the new status and feedback information back to the perception layer and learning layer.

[0098] Among them, the feedback and adjustment layer is mainly used to process user feedback and adjust the strategy accordingly, such as click, browse or purchase behavior, and to adjust the reward value, as well as dynamically adjust the strategy network according to user feedback and environmental changes to optimize the recommendation effect.

[0099] Among them, the learning layer is responsible for the continuous learning and optimization of the agent, including:

[0100] Experience replay pool: stores the interaction data between the agent and the environment, including state, action, user behavior, reward, and next state.

[0101] Network update and optimization: Sample data from the experience replay pool and update the parameters of the Actor and Critic networks to improve the recommendation strategy and evaluation accuracy.

[0102] Among them, the deployment and monitoring layer is responsible for applying the trained model to the actual system, continuously monitoring the system performance and recommendation effect, collecting user feedback, and feeding back information to the learning layer or perception layer so as to adjust the strategy in time.

[0103] Please refer to Figure 2 , Figure 2 A schematic diagram of a training framework of an intelligent agent provided by an embodiment of the present invention is shown. The training framework of the intelligent agent is described below.

[0104] This architecture cleverly combines reinforcement learning and large language models to achieve an efficient and accurate personalized recommendation system. The core of this architecture consists of four main components: state input (State), decision network (Actor), value assessment network (Critic) and large language model (LLM). The interaction between them forms a complex and efficient decision-making and learning system.

[0105] First, the state input is the starting point of the entire system, which captures the user's current state, historical behavior, and context information. This state information is simultaneously passed to three key components: Actor, Critic, and LLM.

[0106] The Actor network, as the core of decision-making, generates the corresponding action a after receiving state information. This action represents the recommended decision of the system in the current state. The design goal of the Actor network is to learn a strategy that can maximize the long-term cumulative reward.

[0107] The Large Language Model (LLM) receives state information and outputs predicted user behavior It is worth noting that LLM uses LORA (Low-Rank Adaptation) technology for training, which can achieve fine-tuning for specific tasks while keeping most of the model parameters unchanged. During the agent training process, LLM's parameters are frozen, which means that LLM acts as a stable predictor of user behavior and provides reliable behavior prediction for the entire system.

[0108] The Critic network is the evaluation center of the entire architecture. It receives the state s, the action a generated by the Actor, and the user behavior predicted by the LLM. Based on these three inputs, Critic outputs Q(s,a, ), this Q value represents the expected user behavior when performing action a in a given state s. Long-term value estimation.

[0109] After the Critic network outputs the value estimate, the training process of the agent continues. The system first stores the current state, action, reward, next state, and user behavior predicted by LLM into the experience replay pool, and then randomly samples a batch of data from it for training. The Critic network updates its parameters by minimizing the error between the predicted Q value and the target Q value, while the Actor network updates its parameters based on the Q value output by the Critic using the policy gradient method to optimize the recommendation strategy for higher long-term returns.

[0110] After training is completed, the system deploys the trained Actor network to the production environment for actual recommendation tasks. At the same time, the system maintains a continuous learning mechanism, collects new user feedback during actual use, and regularly updates the model to ensure that the recommendation system can continue to adapt to changes in user preferences and environment.

[0111] Please refer to Figure 3 , Figure 3 A flow chart of an agent training method is shown, which may specifically include the following steps:

[0112] Step 301: Input the sample recommendation task into the intelligent agent, and obtain multiple groups of sample interaction data output by the intelligent agent.

[0113] Specifically, the system first constructs a series of sample recommendation tasks. These sample tasks are generated based on real user data and historical interaction records, covering a variety of possible recommendation scenarios and user types. Each sample recommendation task contains user information u, historical interaction sequence , candidate item pool I and corresponding scene information.

[0114] After these sample recommendation tasks are input into the agent, the agent will make recommendation decisions for each task based on the current policy network parameters. In this process, the agent will output a series of actions a, each of which corresponds to a set of K-dimensional vectors, representing the K recommended items selected from the candidate item pool I.

[0115] Next, the system simulates the user's feedback on these recommendations and generates sample interaction data. Each set of sample interaction data includes the sample state s at the current moment t , Sample user behavior b t , sample action a t And the sample reward r t , and the sample state s at the next moment t+1 .

[0116] Through the above process, the system can obtain multiple sets of high-quality sample interaction data. These data not only reflect the current decision-making ability of the agent, but also contain rich state transition information and reward signals. They will be stored in the experience replay pool to provide key learning materials for subsequent strategy optimization and value evaluation.

[0117] Step 302: Obtain the predicted behavior output by the large language model based on the sample recommendation task, wherein the large language model is trained based on the sample recommendation task.

[0118] Specifically, the system first needs to conduct targeted training on the large language model to ensure that the large language model can fully understand the semantic features and contextual information of the specific recommendation scenario. After the training is completed, the system will input each sample recommendation task into the fine-tuned large language model. These inputs include user information, historical interaction sequences, candidate item pools, and related scene information. The large language model will perform deep semantic analysis on this information, not only understanding the literal meaning of each element, but also capturing the potential associations and contextual information between them.

[0119] Based on semantic understanding, the large language model outputs a series of predicted behaviors. These predicted behaviors are essentially the large language model's speculation on items that the user may like. When generating these predicted behaviors, the large language model not only considers the user's historical behavior and preferences, but also incorporates a deep understanding of the content of the items. For example, it may identify potential topics in the user's historical interactions and look for semantically related but not identical items among the candidate items, thereby maintaining the relevance of the recommendations while improving novelty. In addition, the large language model is also able to take into account a wider range of contextual information, such as season, time, social trends, etc., which may affect the user's interests and behaviors.

[0120] Step 303: For each set of sample interaction data, obtain a target expected value based on the interaction data and the corresponding predicted behavior.

[0121] Among them, the target expected value is an estimate of the long-term cumulative reward of a specific state-action pair in the reinforcement learning framework, reflecting the agent's expectation of the future benefits that can be obtained by taking a certain action in a given state. In the embodiment of the present invention, the target expected value can be understood as a comprehensive evaluation indicator that combines actual interaction experience and semantic understanding of a large language model. It not only takes into account the current immediate reward, but also incorporates the prediction of possible rewards in the future.

[0122] Furthermore, the target expected value is mainly used to guide the learning process of the agent. By comparing the actual Q value and the target expected value, the agent can adjust its policy network so that its decision is closer to the choice that can obtain the maximum long-term benefit. At the same time, it provides the agent with a benchmark for evaluating the quality of actions, which is used to judge the relative value of an action in a specific state. The introduction of the prediction of the large language model enables the target expected value to better capture the long-term benefits and prevent the agent from falling into short-sighted decisions, thereby achieving a balance between short-term and long-term benefits. Accurate target expected values ​​can accelerate the learning process of the agent, allowing it to converge to the optimal strategy faster and improve learning efficiency.

[0123] In addition, since the target expectation combines the semantic understanding of the large language model, it can also help the agent better handle unseen states and actions, enhancing the generalization ability of the system.

[0124] Based on the above embodiment, as an optional embodiment, the calculation process of the target calculated value may further include the following steps:

[0125] Step 401: Based on the sample state, sample user behavior, and sample action at the current moment, obtain the first expected value of the agent at the current moment.

[0126] The first expected value refers to the agent's immediate value estimate of the current state-action pair. It reflects the agent's expectation of the cumulative rewards that may be obtained in the future when taking a specific action in a given state and observing a specific user behavior. In the embodiment of the present invention, the first expected value can be understood as the preliminary evaluation made by the agent's value evaluation network based on the currently observed information.

[0127] Specifically, the system uses the agent’s value evaluation network to calculate the first expected value. The value evaluation network is a trained deep neural network whose input includes the sample state s at the current moment. t , Sample user behavior b t and sample action a t .

[0128] The value evaluation network will integrate these input information and perform deep processing, and finally output the first expected value Q(s t , a t , b t ). It represents the agent's response to being in state s t Take action a and observe user behavior b t An estimate of the long-term cumulative reward that can be obtained when

[0129] Step 402: Based on the sample state and predicted behavior at the next moment, obtain the maximum expected value of the agent at the next moment.

[0130] The maximum expected value refers to the estimated maximum value that the agent may obtain at the next moment. It reflects the best expected benefit that can be achieved by taking the optimal action in a given next state. In the embodiment of the present invention, the maximum expected value can be understood as the agent's prediction of the best possible situation in the future.

[0131] Specifically, the system first obtains the sample state at the next moment, which reflects the new situation of the recommendation system after executing a recommendation. Next, the system considers all possible candidate actions in this new state and introduces the predicted behavior output by the large language model. At the same time, the system also considers the actions that may be selected based on the current strategy. For these two actions, for all possible candidate actions and their predicted behaviors, the system will use the value evaluation network to calculate the corresponding expected values, and select the action with the largest expected value among all candidate actions, and accordingly, obtain the maximum expected value. Among them, one is the user behavior predicted based on the large language model, and the other is the user behavior predicted based on the current strategy. Take the larger of the two expected values ​​to get the maximum expected value.

[0132] Step 403: Adjust the first expected value based on the sample reward and the maximum expected value to obtain a target expected value.

[0133] Specifically, the system first obtains the sample reward at the current moment, which reflects the user's immediate feedback on the agent's current recommendation. Then, the system uses a discount factor to balance the importance of immediate rewards and future benefits. The discount factor reflects the principle that recent benefits are more important than long-term benefits, while also avoiding the problem that cumulative rewards may tend to infinity.

[0134] Exemplarily, the calculation formula of the target expected value can be expressed as: target expected value = sample reward + discount factor × maximum expected value.

[0135] The target expected value obtained by the above method can enable the agent to take into account both short-term and long-term interests when making decisions. By adding the immediate reward and the discounted future benefits, the agent can make more balanced and strategic decisions, avoiding the problem of pursuing only short-term benefits and ignoring long-term development. Secondly, this method enhances the stability and robustness of the system. Even though the immediate rewards may be inaccurate or noisy in some cases, the introduction of the maximum expected value can smooth these fluctuations to a certain extent, making the system's learning process more stable.

[0136] In addition, the above method also improves the adaptability of the system. By continuously adjusting the target expected value, the system can quickly respond to changes in user interests and the dynamic characteristics of the environment. If the user's preference changes, this change will be reflected in the sample reward, which in turn affects the target expected value and ultimately leads to adjustments in the recommendation strategy.

[0137] In another feasible embodiment, the target expected value can also be calculated using the following formula:

[0138] ;

[0139] In the formula, Q π(∙) represents the expected value estimate under strategy π, Respectively represent the current sample state, sample action, and sample user behavior, represents the target expected value, represents the sample reward at the current moment, They represent the sample state, sample action, and predicted behavior at the next moment, respectively; α represents the learning rate, which is used to control the update step size and determines the weight of new information in the update; γ represents the discount factor, which ranges from [0,1) and is used to balance the weight of current rewards and future returns. A larger γ value means that the agent attaches more importance to future rewards.

[0140] Step 304: Based on the loss function, obtain the difference value of the target expected value, and adjust the parameters of the agent based on the difference value.

[0141] Specifically, the system first defines a suitable loss function, which is usually the difference between the expected value of the agent's output and the target expected value. The system calculates the value of the loss function, that is, the difference value, which reflects the distance between the agent's current output and the ideal output. Based on the difference value, the system can use optimization algorithms such as gradient descent to adjust the agent's parameters, with the goal of minimizing the loss function so that the agent's output is as close to the target expected value as possible.

[0142] After each parameter adjustment, the system will determine whether the preset training termination condition has been reached. This condition is usually whether the number of interaction steps between the agent and the recommendation system has reached the upper limit. The purpose of setting the upper limit of the number of interaction steps is to balance the training time and model performance and prevent overfitting problems caused by overtraining. If the termination condition is not reached, the system will continue the next round of interaction and learning process.

[0143] When the training termination condition is reached, the training process ends. The system obtains the last updated policy network parameters as the final training results. These parameters contain the knowledge and experience gained by the agent through a large number of interactions and learning. Finally, the system deploys these trained parameters to the policy network so that it can be applied in actual recommendation scenarios.

[0144] Based on the above embodiment, as an optional embodiment, the embodiment of the present invention designs two loss functions, including a first loss function and a second loss function, wherein:

[0145] The first loss function is used to characterize the differences in decision networks among agents, including:

[0146] ;

[0147] In the formula, represents the first loss value, B represents multiple groups of sample interaction data, It means taking the average value of multiple groups of sample interaction data. Respectively represent the current sample state, sample action, and sample user behavior, represents the target expected value under the recommended strategy, represents the sample reward at the current moment, Indicates the sample state at the next moment.

[0148] Among them, the first loss function is mainly used to characterize the difference in decision networks in the intelligent agent. This difference reflects the distance between the output of the current decision network and the ideal decision, and is the core driving force for the continuous optimization and learning of the intelligent agent.

[0149] Furthermore, the first loss function quantifies this difference by calculating the average of the negative target expected value. The target expected value represents the system's best estimate of the future cumulative reward when a certain action is taken and a specific user behavior is observed in a given state. By minimizing the first loss function, the system is actually trying to make the output of the decision network closer to this best estimate.

[0150] The second loss function is used to characterize the differences in the value evaluation network in the agent, including:

[0151] ;

[0152] In the formula, represents the second loss value, represents the discount factor, which is used to balance the impact of sample rewards and future rewards. represents the expected value of the sample at the next moment, They represent the sample state, sample action, and predicted behavior at the next moment respectively.

[0153] Among them, the second loss function is mainly used to characterize the difference in the value evaluation network in the intelligent agent. This difference reflects the distance between the current prediction of the value evaluation network and the actual long-term return, and is the core driving force for the continuous optimization and learning of the value evaluation network.

[0154] Furthermore, the second loss function quantifies this difference by calculating the mean squared error between the current prediction value and the target value. The target value consists of the immediate reward and the discounted next-moment prediction value, which represents the system's best estimate of the long-term cumulative reward. By minimizing the second loss function, the system is actually trying to make the prediction of the value evaluation network closer to this best estimate.

[0155] The above embodiment describes the training process of the agent. Based on the above embodiment, the training process of the large language model will be described below. Specifically, the process may include the following steps:

[0156] Step 501: Input the sample recommendation task into the large language model to obtain the predicted behavior output by the large language model.

[0157] Specifically, the system first constructs a series of representative sample recommendation tasks, which cover various possible recommendation scenarios and user types. Each sample recommendation task usually contains user information u, historical interaction sequence, and , candidate item pool I and related scene information. This information is formatted into structured input that can be understood by the large language model.

[0158] After these formatted sample recommendation tasks are input into the large language model, the model performs a deep semantic analysis of the input. The large language model not only understands the literal meaning of each element, but also captures the potential associations and contextual information between them. For example, it may identify potential topic preferences in a user's historical interactions, or understand the inherent logic of certain item combinations. Based on this deep understanding, the large language model outputs predicted behaviors, which are essentially the model's speculation on the actions that the user may take in a given situation.

[0159] Based on the above embodiment, as an optional embodiment, step 501 may further include the following steps:

[0160] Step 601: Generate a task instruction and a binary interaction sequence based on a recommendation task, where the binary interaction sequence is used to represent the user's preference for the recommended data in the recommendation task.

[0161] Specifically, the system will first design a task instruction based on the characteristics of the recommendation task. This task instruction is a natural language description that clearly tells the large language model the specific task to be completed. For example, the task instruction may be "based on the user's historical interactions and personal information, predict whether the user will like a given recommended item." This clear instruction can guide the large language model to focus its powerful language understanding and reasoning capabilities on the recommendation task, improving the pertinence and accuracy of the prediction.

[0162] Next, the system generates a binary interaction sequence based on the user's historical interaction data in the recommendation task. This sequence is the user's feedback on the historical recommended items, expressed in two states: "like" and "dislike". For example, if the user gives a high score to the recommended movie A or watches it for a long time, this will be marked as "like" in the binary sequence; conversely, if the user quickly skips or gives a low score to movie B, it will be marked as "dislike". This binary representation not only simplifies the expression of user preferences, but also helps the large language model understand the user's preference pattern more clearly.

[0163] By combining task instructions and binary interaction sequences, a structured input can be formed. For example, "The historical interaction sequence of user Xiao Ming is: Movie A - like, Movie B - dislike, Movie C - like. Based on this sequence, predict whether the user will like Movie D." This input method not only provides a clear task goal, but also gives a specific expression of user preferences, which can give full play to the semantic understanding and reasoning capabilities of the large language model.

[0164] Step 602: Input the task instruction and the binary historical interaction sequence into the large language model, and obtain the predicted behavior output by the large language model, where the predicted behavior is the binary classification result corresponding to the task instruction.

[0165] Specifically, the system first combines the generated task instructions and binary historical interaction sequences into a structured input. For example, the input may be: "Task: Predict whether the user will like the movie "Interstellar". User historical interactions: "Inception" - like, "Avatar" - like, "Titanic" - dislike. Please make predictions based on this information." This input method not only clarifies the task goal, but also provides the user's historical preference information, providing sufficient context for the prediction of the large language model.

[0166] After passing this structured input to the large language model, the model performs deep semantic analysis and reasoning. The large language model not only understands the specific requirements of the task instructions, but also captures the user's preference patterns from the binary historical interaction sequence. For example, the model may recognize that the user prefers science fiction movies, but is not very interested in romance movies. Based on this deep understanding, the large language model outputs a binary classification result, which predicts whether the user "likes" or "dislikes" the target recommendation item.

[0167] Step 502: Obtain a low-rank matrix corresponding to the predicted behavior.

[0168] In specific implementation, the system first constructs a low-rank matrix based on the predicted behavior output by the large language model. This low-rank matrix is ​​essentially a low-rank update of the weights between certain layers of the large language model. The construction process of the low-rank matrix involves the analysis of the predicted behavior and the optimization of the weights. The system will adjust the values ​​of the A and B matrices according to the characteristics of the predicted behavior to minimize the prediction error. This process can be achieved through optimization algorithms such as gradient descent.

[0169] Step 503: Adjust the weight of the large language model based on the low-rank matrix.

[0170] Specifically, the system first applies the obtained low-rank matrix to specific layers of the large language model, and performs incremental updates based on the original weights. During the adjustment process, the system carefully selects the layers that need to be updated. Typically, these layers are the parts of the model that are most sensitive to recommendation tasks, such as layers related to semantic understanding and decision-making. For each selected layer, the system applies the corresponding low-rank update. This selective update ensures the accuracy and efficiency of the adjustment.

[0171] It is worth noting that during the adjustment process, only the low-rank matrix is ​​trainable, while the original weight matrix remains unchanged. This can reduce the number of parameters that need to be updated, thereby reducing computational and storage costs; secondly, it retains most of the knowledge of the original model, which helps maintain the generalization ability of the model; finally, it makes the adjustment process more stable and reduces the risk of overfitting.

[0172] The above method of adjusting the weights of the large language model based on the low-rank matrix can improve the model's adaptability to specific recommendation tasks. Through precise low-rank updates, the model can be quickly adjusted to adapt to different user groups or recommendation scenarios without retraining the entire model. This high degree of adaptability is particularly important for dynamically changing recommendation environments.

[0173] Figure 4 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 4 As shown, the electronic device may include: a processor 410, a communication interface 420, a memory 430 and a communication bus 440, wherein the processor 410, the communication interface 420 and the memory 430 communicate with each other through the communication bus 440. The processor 410 may call the logic instructions in the memory 430 to execute the recommendation strategy generation method based on the large language model enhancement, the method comprising: obtaining a recommendation task; inputting the recommendation task into an intelligent agent, and obtaining a recommendation strategy output by the intelligent agent; wherein the intelligent agent is trained based on sample recommendation tasks and a large language model, and the large language model is used to predict user behavior based on sample recommendation tasks.

[0174] In addition, the logic instructions in the above-mentioned memory 430 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0175] On the other hand, the present invention also provides a computer program product, which includes a computer program, and the computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the recommendation strategy generation method based on large language model enhancement provided by the above methods, and the method includes: obtaining a recommendation task; inputting the recommendation task to an intelligent agent, and obtaining a recommendation strategy output by the intelligent agent; wherein the intelligent agent is trained based on sample recommendation tasks and a large language model, and the large language model is used to predict user behavior based on sample recommendation tasks.

[0176] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the recommendation strategy generation method based on large language model enhancement provided by the above-mentioned methods, the method comprising: obtaining a recommendation task; inputting the recommendation task to an intelligent agent, and obtaining a recommendation strategy output by the intelligent agent; wherein the intelligent agent is trained based on sample recommendation tasks and a large language model, and the large language model is used to predict user behavior based on the sample recommendation tasks.

[0177] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0178] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0179] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A recommendation strategy generation method based on large language model enhancement, characterized in that: include: Get recommended tasks; Inputting the recommendation task into an intelligent agent, and obtaining a recommendation strategy output by the intelligent agent; The agent is trained based on a sample recommendation task and a large language model, and the large language model is used to predict user behavior based on the sample recommendation task; The training process of the agent includes: Inputting a sample recommendation task into the intelligent agent, and obtaining multiple groups of sample interaction data output by the intelligent agent; Obtaining a predicted behavior output by the large language model based on the sample recommendation task; For each set of the sample interaction data, obtaining a target expected value based on the interaction data and the corresponding predicted behavior; Based on the loss function, obtaining a difference value of the target expected value; adjusting parameters of the agent based on the difference value; Wherein, the large language model is trained based on the sample recommendation task; The interaction data includes a sample state, a sample user behavior, a sample action, and a sample reward at a current moment, and a sample state at a next moment. The step of obtaining a target expected value based on the interaction data and the corresponding predicted behavior includes: Based on the sample state, sample user behavior, and sample action at the current moment, obtaining a first expected value of the agent at the current moment; Based on the sample state and predicted behavior at the next moment, obtaining the maximum expected value of the agent at the next moment; The first expected value is adjusted based on the sample reward and the maximum expected value to obtain a target expected value.

2. The method for generating a recommendation strategy based on large language model enhancement according to claim 1, characterized in that: The step of inputting the recommendation task into an intelligent agent and obtaining a recommendation strategy output by the intelligent agent includes: Obtaining the status of the user corresponding to the recommended task through the agent; Obtaining the user's action based on the user's state through the agent; A recommendation strategy is generated by the agent based on the user's actions.

3. The method for generating a recommendation strategy based on large language model enhancement according to claim 1 or 2, characterized in that: The loss function includes a first loss function and a second loss function, wherein: The first loss function is used to characterize the difference of the decision network in the agent, including: Where, L Actor represents the first loss value, B represents multiple groups of sample interaction data, |B| represents the average of multiple groups of sample interaction data, s t ,a t ,b t Respectively represent the current sample state, sample action, and sample user behavior, represents the target expected value under the recommended strategy, r t represents the sample reward at the current moment, s t+1 Indicates the sample state at the next moment; The second loss function is used to characterize the difference of the value evaluation network in the agent, including: Where, L Critic represents the second loss value, γ represents the discount factor, which is used to balance the impact of sample rewards and future rewards. represents the expected value of the sample at the next moment, s t+1 ,a t+1 , They represent the sample state, sample action, and predicted behavior at the next moment respectively.

4. The method for generating a recommendation strategy based on large language model enhancement according to claim 1 or 2, characterized in that: The training process of the large language model includes: Inputting the sample recommendation task into the large language model to obtain the predicted behavior output by the large language model; Obtaining a low-rank matrix corresponding to the predicted behavior; The weights of the large language model are adjusted based on the low-rank matrix.

5. The method for generating a recommendation strategy based on large language model enhancement according to claim 4, characterized in that: The step of inputting the sample recommendation task into the large language model and obtaining the predicted behavior output by the large language model includes: generating a task instruction and a binary interaction sequence based on the recommendation task, wherein the binary interaction sequence is used to represent the user's preference for the recommended data in the recommendation task; The task instruction and the binary historical interaction sequence are input into the large language model to obtain a predicted behavior output by the large language model, where the predicted behavior is a binary classification result corresponding to the task instruction.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method for generating a recommendation strategy based on large language model enhancement as described in any one of claims 1 to 5 is implemented.

7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for generating a recommendation strategy based on large language model enhancement as described in any one of claims 1 to 5 is implemented.

8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for generating a recommendation strategy based on large language model enhancement as described in any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • User multi-behavior recommendation method and device based on large language model

    CN117851685A

  • Using large language models for dialogue management and recommendations in a conversational recommender system

    WO2024167498A1