A recommendation method and device based on multi-agent
By adopting a multi-agent-based recommendation method in the social network platform, combining the policy network and evaluation network, the recommendation value is generated from the individual and group levels, and the recommendation inaccuracy problem caused by ignoring user relationships in the existing technology is solved, and more accurate user recommendations are achieved.
Patent Information
- Application Number
- CN202210540560.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-17
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-05-17
AI Technical Summary
When the prior art recommends relevant users in social network platforms, the relationship between users is ignored, resulting in inaccurate recommendation results.
The recommendation method based on multi-agents is adopted, and the final recommendation value is generated from the individual and group levels through the policy network and evaluation network, and the recommended value agent set is determined.
It realizes more objective and accurate user recommendations, taking into account the relationship between users and individual historical status information.
Smart Images

Figure CN114817744B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a recommendation method and device based on multi-agents. Background Art
[0002] In some social networking platforms, users often mention other related users (indicated by the "@" symbol) when posting messages. If the platform can automatically recommend related users that may be mentioned as candidates in these scenarios, it can effectively improve the user experience. In the prior art, users' preferences and relationships with other users are usually inferred through the user's most recent messages or a few randomly selected historical messages. However, the relationship between users is ignored in this inference process, resulting in inaccurate recommendation results. Summary of the invention
[0003] The present disclosure provides a multi-agent based recommendation method and device, which realizes obtaining the final recommendation value from the individual (agent) and group (agent set) levels, thereby making recommendations more objective and accurate.
[0004] In a first aspect, the present disclosure provides a multi-agent based recommendation method, comprising:
[0005] When determining the current input information of the current agent, obtaining the historical state information of the current agent, wherein the historical state information includes the historical input information of the current agent and the list information related to other agents;
[0006] Input the historical state information and current input information of the current agent into the policy network to generate the first true value of the current agent and the first true values of other agents;
[0007] Processing is performed based on the first true value of the current agent and the first true values of other agents to obtain a feedback value;
[0008] Input the feedback value of the current agent, the feedback values of the other agents, and a preset discount factor into the evaluation network, and output the evaluation value vectors corresponding to the current agent and the other agents;
[0009] The evaluation value vector is input into the policy network, the final recommendation values of other agents relative to the current agent are output, and the recommended value agent set is determined.
[0010] According to the multi-agent recommendation method provided by the present disclosure, the first real value of the current agent and the first real values of other agents are processed to obtain the feedback value, including:
[0011] Based on the historical state information of the current agent and the current input information, the initial recommendation value of the other agents relative to the current agent corresponding to each input information is obtained through the strategy network;
[0012] Label each input information with a sample label, wherein the label value of the sample label is 1 or 0;
[0013] Sampling the first true value of the current agent to obtain a sampled value;
[0014] The feedback value is updated based on the sampled value and the tag value.
[0015] According to the multi-agent recommendation method provided by the present disclosure, the feedback value of the current agent, the feedback value of the other agents and the preset discount factor are input into the evaluation network, and the evaluation value vectors corresponding to the current agent and the other agents are output, including:
[0016] Input the feedback value of the current agent, the feedback values of the other agents, and a preset discount factor into the evaluation network, and output the evaluation values corresponding to the current agent and the other agents;
[0017] The evaluation value is evaluated through an evaluation network, and the evaluation value vectors corresponding to the current agent and the other agents are output.
[0018] According to the multi-agent based recommendation method provided by the present disclosure, the evaluation value vector is input into the strategy network, and the final recommendation value of other agents relative to the current agent is output, including:
[0019] Inputting the evaluation value vector into the policy network to obtain an updated policy network;
[0020] Generate a second true value of the current agent and second true values of other agents based on the updated policy network, historical state information of the current agent, and current input information;
[0021] Determine the weight values of other agents relative to the current agent based on a preset agent utility matrix; wherein the agent utility matrix includes the weight value of each agent;
[0022] Based on the second true value of the current agent, the second true values of other agents and the weight value, the final recommended values of other agents relative to the current agent are output.
[0023] According to the multi-agent based recommendation method provided by the present disclosure, the step of determining a set of recommendation value agents includes:
[0024] Comparing the recommendation values of the other agents relative to the current agent with a preset threshold;
[0025] If the recommended value is greater than or equal to a preset threshold, a recommended value agent set is recommended for the current agent, wherein the recommended value agent set is the sum of agents related to the input information and historical state information of the current agent;
[0026] If the recommendation value is less than the preset threshold, an agent is randomly recommended for the current agent.
[0027] According to the multi-agent based recommendation method provided by the present disclosure, based on the sampled value and the label value, updating the feedback value is achieved by the following formula:
[0028]
[0029] Among them, p t Represents the number of historical positive samples in the historical state information, n t represents the number of historical negative samples in the historical state information, G t represents the label value of the tth time, Represents the recommendation value of agent i at the tth time.
[0030] According to the multi-agent based recommendation method provided by the present disclosure, the method further includes:
[0031] Based on the first loss function and the second loss function, the weight value in the agent utility matrix is updated; wherein the first loss function is:
[0032]
[0033] The second loss function is:
[0034]
[0035] Among them, p t Represents the number of historical positive samples in the historical state information, n t represents the number of historical negative samples in the historical state information, G t represents the label value of the tth time, o t represents the recommended value, μ represents the strategy of each agent, i and j represent agents, s ij Represents the similarity between the state of agent i and the state of agent j. The agent utility matrix is decomposed into two small matrices represented by A and B, with sizes N×d and d×N respectively, and d is less than N, a i is the i-th row of matrix A, a j is the jth row of matrix A, b iis the i-th column of matrix B, b j is the j-th column of matrix B.
[0036] In a second aspect, the present disclosure provides a multi-agent based recommendation device, comprising:
[0037] An acquisition module, configured to acquire historical state information of the current agent when current input information of the current agent is determined, wherein the historical state information includes historical input information of the current agent and list information related to other agents;
[0038] A generation module, used to input the historical state information of the current agent and the current input information into the policy network, and generate the first true value of the current agent and the first true values of other agents;
[0039] A processing module, configured to process the first true value of the current agent and the first true values of other agents to obtain a feedback value;
[0040] An input module, used to input the feedback value of the current agent, the feedback values of the other agents and a preset discount factor into the evaluation network, and output the evaluation value vectors corresponding to the current agent and the other agents;
[0041] The determination module is used to input the evaluation value vector into the policy network, output the final recommendation value of other agents relative to the current agent, and determine the recommended value agent set.
[0042] In a third aspect, the present disclosure provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the multi-agent-based recommendation method as described in any one of the above items are implemented.
[0043] In a fourth aspect, the present disclosure provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the multi-agent-based recommendation method as described in any one of the above items.
[0044] In a fifth aspect, the present disclosure provides a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the multi-agent-based recommendation method as described in any one of the above items.
[0045] The present disclosure provides a multi-agent based recommendation method and device, which determines the current input information and historical status information of the current agent, wherein the historical status information includes the historical input information of the current agent and list information involving other agents; inputs the historical status information and current input information of the current agent into a strategy network to generate a first true value of the current agent and the first true values of other agents; performs processing based on the first true value of the current agent and the first true values of other agents to obtain a feedback value; inputs the feedback value of the current agent, the feedback value of other agents and a preset discount factor into an evaluation network, and outputs evaluation value vectors corresponding to the current agent and other agents; inputs the evaluation value vector into a strategy network, outputs final recommendation values of other agents relative to the current agent, and determines a set of recommended value agents. Not only does the feedback value get obtained through the first true value of the current agent and the first true value of other agents, but also the evaluation value vector is input into the policy network to output the final recommendation value of other agents relative to the current agent, and the agent recommended by the current agent is determined among other agents, thus achieving the final recommendation value from the individual (agent) and group (agent set) levels, thereby enabling more objective and accurate recommendations for the current agent. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the present disclosure or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0047] Figure 1 is a flowchart of a multi-agent based recommendation method provided by an embodiment of the present disclosure;
[0048] Figure 2 is a schematic diagram of a process for obtaining a feedback value provided by an embodiment of the present disclosure;
[0049] Figure 3 is a schematic diagram of a process for obtaining a final recommendation value provided by an embodiment of the present disclosure;
[0050] Figure 4 is a block diagram of a multi-agent based recommendation method provided by an embodiment of the present disclosure;
[0051] Figure 5 is a specific flow chart of the multi-agent based recommendation method provided by an embodiment of the present disclosure;
[0052] Figure 6 is a schematic diagram of the structure of a multi-agent based recommendation device provided by an embodiment of the present disclosure;
[0053] Figure 7 It is a schematic diagram of the structure of an electronic device provided by the present disclosure. DETAILED DESCRIPTION
[0054] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all of the embodiments. Based on the embodiments in the embodiments of the present disclosure, every other embodiment obtained by ordinary technicians in this field without creative work belongs to the scope of protection of the embodiments of the present disclosure.
[0055] The multi-agent based recommendation method provided by the disclosed embodiment is applied to a multi-agent delayed aggregate graph neural network architecture (MA-DGNN). The multi-agent is composed of a series of interacting agents, and the internal agents complete complex and large-scale tasks that a single agent cannot complete through mutual communication, cooperation, competition, etc.
[0056] Below, the aggregate graph neural network (GNN) in the prior art is first introduced, and then the improved delayed aggregate neural network (DGNN) disclosed in the present invention is described.
[0057] Aggregate Graph Neural Network (GNN) is a class of information processing frameworks that operate on network data in a decentralized manner, obtain useful information through repeated communication with neighbors, process ordinary aggregate GNNs on fixed graphs, and support processing fixed signals on fixed graphs.
[0058] In the prior art, there is also a process of extending the application of GNN on fixed graphs to process time-varying graphs on time-varying graph support, specifically:
[0059] Introduce (i, j)∈ε n Indicates that j may send data i at time n, ε n Represents the set of edges on a time-varying graph using the graph shift operator Description Graph Support G n , G n Indicates that [S n ] ij May be non-zero if and only if (j, i)∈ε n Or i = j, which shows the sparsity of the graph. More specifically, the state of node j at time n-1 is represented as Then through [S n ] ij May be non-zero if and only if (j, i)∈ε n or the property i=j, represents the set of all elements that send data to i at time n, and we get formula (1):
[0060]
[0061]
[0062] Based on the locality of (1), the delayed aggregation GNN constructs a recursive k-hop neighborhood aggregation sequence, as shown in formulas (2)(3)(4). Specifically, define a signal sequence Yes 0n =x n and:
[0063] y kn =S n y (k-1)(n-1) (3)
[0064] Get y kn =(S n S n-1 ...S n-k+1 )x n-k , therefore, (3) characterizes the time-varying network sequence S n-k+1 To S n Medium n-k Next, aggregate y kn (k∈{0, 1, ..., K-1}) forms a nested state along the multi-hop neighborhood:
[0065] z in =[[y 0n ] i ; [y 1n ] i ;…;[y (K-1)n ] i ] (4)
[0066] Among them, z in The k+1th element of kn ] i =[S n S n-1 ...S n-k+1 x n-k ] i , is x j(n-k) The average value of It is the k of nk moment - Hop neighborhood.
[0067] Specify z in Due to its nested aggregation properties, it has a regular temporal structure and is trained by a convolutional neural network of length L for z in Modeling, that is, for l∈[L], let:
[0068]
[0069] Among them, σ (l) is a point-wise nonlinear function, H (l) is a class of small support filters with learnable parameters shared by every node.
[0070] In summary, the delayed aggregation neural network described by formulas (3)-(5) includes the strategy The local parameterization of captures the dynamics and sparse interactions of the network and allows long-range communication via multi-hop information diffusion.
[0071] Reference Figure 1 As shown, it is a flowchart of a multi-agent based recommendation method provided by an embodiment of the present disclosure, including:
[0072] 110. When the current input information of the current agent is determined, historical state information of the current agent is obtained, wherein the historical state information includes historical input information of the current agent and list information related to other agents.
[0073] In this step, it refers to a hardware, software or other entity with adaptive and autonomous capabilities, such as a robot, a drone and a game character, etc. In the embodiment of the present disclosure, the intelligent agent can be understood as a user.
[0074] The current input information can be understood as the information when the user publishes a message, including the content information of the published message or the information of the user related to the published message.
[0075] The list information of other intelligent agents can be understood as the list information of tweets and the list information of persons mentioned by the user in the past.
[0076] 120, inputting the historical state information of the current agent and the current input information into the strategy network to generate the first true value of the current agent and the first true values of other agents.
[0077] In this step, the policy network can be understood as a delayed aggregation neural network, and the true value can be understood as a real number.
[0078] Specifically, the historical state information and current input information of the current user are input into the delayed aggregation neural network to generate the first true value of the current user and the first true values of other users.
[0079] 130, performing processing based on the first true value of the current intelligent agent and the first true values of other intelligent agents to obtain a feedback value.
[0080] In this step, processing refers to modeling processing, which is specifically achieved through a Markov Decision Process (MDP).
[0081] Markov decision process is a mathematical model of sequential decision making, which is used to simulate the random strategies and rewards that can be achieved by agents in an environment where the system state has Markov properties. MDP is constructed based on a set of interacting objects, namely agents and environments, and its elements include state, action, feedback value and reward. In the simulation of MDP, the agent will perceive the current system state, take actions on the environment according to the feedback value, thereby changing the state of the environment and receiving rewards. The accumulation of rewards over time is called rewards.
[0082] The feedback value is obtained based on a Markov decision process, and can also be understood as a feedback value calculated through a model to improve system performance.
[0083] 140, input the feedback value of the current agent, the feedback value of the other agents and the preset discount factor into the evaluation network, and output the evaluation value vectors corresponding to the current agent and the other agents.
[0084] In this step, the preset discount factor γ is set to 0.99.
[0085] The evaluation network can be understood as a centralized evaluation of each agent. The evaluation body in the evaluation network absorbs the cascade of each agent's state and action. Then, the evaluation body evaluates the state-action pair of each agent and outputs an evaluation value vector to update the policy network.
[0086] 150, input the evaluation value vector into the policy network, output the final recommendation values of other agents relative to the current agent, and determine the recommended value agent set.
[0087] In this step, the recommended value agent set refers to one or more agents mentioned by the current agent in the list information of other agents.
[0088] The present disclosure provides a multi-agent based recommendation method, which determines the current input information and historical status information of the current agent, wherein the historical status information includes the historical input information of the current agent and list information involving other agents; inputs the historical status information and current input information of the current agent into a strategy network to generate a first true value of the current agent and the first true values of other agents; performs processing based on the first true value of the current agent and the first true values of other agents to obtain a feedback value; inputs the feedback value of the current agent, the feedback value of other agents and a preset discount factor into an evaluation network, and outputs evaluation value vectors corresponding to the current agent and other agents; inputs the evaluation value vector into a strategy network, outputs final recommended values of other agents relative to the current agent, and determines a set of recommended value agents. Not only does the feedback value get obtained through the first true value of the current agent and the first true value of other agents, but also the evaluation value vector is input into the policy network to output the final recommendation value of other agents relative to the current agent, and the agent recommended by the current agent is determined among other agents, thus achieving the final recommendation value from the individual (agent) and group (agent set) levels, thereby enabling more objective and accurate recommendations for the current agent.
[0089] Based on any of the above embodiments, refer to Figure 2 FIG. 1 is a flow chart of obtaining a feedback value provided by an embodiment of the present disclosure, including:
[0090] 210. Based on the historical state information of the current agent and the current input information, the initial recommendation value of the other agents relative to the current agent corresponding to each input information is obtained through the strategy network.
[0091] 220, labeling a sample label for each input information, wherein the label value of the sample label is 1 or 0.
[0092] In this step, when an agent j mentioned by agent i in the tweet list information belongs to the agents in the mentioned list information, agent j is marked as a positive sample with a label value of 1, otherwise it is marked as a negative sample with a label value of 0.
[0093] 230, sampling the first true value of the current agent to obtain a sampled value.
[0094] In this step, since each agent will output a real value, each agent will correspond to a sample value. The first real value of the current agent is adopted to obtain a binary value, and this binary value follows the Bernoulli distribution, which can be expressed as:
[0095] The current agent i∈[N] outputs a first true value Thus sampling a binary value
[0096] 240. Update the feedback value based on the sampled value and the tag value.
[0097] In this step, since the number of negative samples in the real data set far exceeds the number of positive samples, it is necessary to adjust the items in the feedback value to solve the problem of unbalanced data sets. Therefore, the feedback value is updated by obtaining the sampling value and label value to make the data set relatively balanced.
[0098] Based on any of the above embodiments, step 140 specifically includes the following steps 141 to 142:
[0099] 141, input the feedback value of the current agent, the feedback value of the other agents and the preset discount factor into the evaluation network, and output the evaluation values corresponding to the current agent and the other agents.
[0100] In this step, it is specifically implemented by the following formula:
[0101] Customize a minimax objective for each agent to update the evaluation network at the tth time:
[0102]
[0103] in, represents the minimax goal of each agent, represents the policy loss, π is in state s t and strategy μ to obtain a t,i The probability of t is the concatenation of the states of each agent at the tth time, a t,i (i∈[N]) is the action of agent i, r t,i represents the reward value, γ represents the discount factor, a′ i represents the action of agent i in the target policy network, μ is the strategy of each agent, and μ i (s t ) is μ(s t ), μ′ is the target policy network of M3DDPG, ★ represents the optimal solution, and M3DDPG represents minimax multi-agent deep deterministic policy gradient.
[0104] In order to derive for each agent i∈[N] An end-to-end solution is adopted to replace the inner loop minimization process in formula (1) with a single step gradient descent. Specifically, the y in (1) is retrieved i for:
[0105]
[0106] in, represents the gradient symbol, a′ k represents the action of agent k in the target policy network, ∈ j Represents the minimum value of the strategy loss.
[0107] Get yj and the corresponding Then, for each i∈[N], we will Perform linear combination, the evaluation value in this step is used express.
[0108] This step achieves the goal of customizing a minimax objective for each agent in the face of multiple non-stationarities from the environment and other agents, so that the multi-agent delayed aggregation graph neural network architecture can learn robust strategies during training.
[0109] 142. Evaluate the evaluation value through an evaluation network, and output evaluation value vectors corresponding to the current agent and the other agents.
[0110] In this step, the evaluation network evaluates the state-action pair of each agent, outputs the evaluation value vectors corresponding to the current agent and the other agents, and thus updates the policy network.
[0111] Reference Figure 3 FIG. 1 is a flow chart of obtaining a final recommendation value provided by an embodiment of the present disclosure, including:
[0112] 310, input the evaluation value vector into the policy network to obtain an updated policy network.
[0113] In this step, a delayed aggregation neural network is used as the policy network.
[0114] 320. Generate a second true value of the current agent and second true values of other agents based on the updated policy network, the historical state information of the current agent, and the current input information.
[0115] 330. Determine the weight values of other agents relative to the current agent based on a preset agent utility matrix; wherein the agent utility matrix includes the weight value of each agent.
[0116] In this step, the preset agent utility matrix models the dynamic relationship between each pair of users and is updated by a neural network matrix decomposition method. Its weight values are based on the similarity between users and are optimized independently of the DGNN, rather than jointly updating them slowly with the DGNN through feedback one at a time.
[0117] 340. Based on the second true value of the current agent, the second true values of other agents and the weight value, output the final recommendation values of other agents relative to the current agent.
[0118] In this step, the preset agent utility matrix aggregates the second true value of each agent and outputs a linear weighted value o t , to make a final decision, the final recommended value is o t express.
[0119] Based on any of the above embodiments, step 150 specifically includes the following steps 151 to 153:
[0120] 151, comparing the recommendation values of the other agents relative to the current agent with a preset threshold.
[0121] In this step, the preset threshold is thre t Indicates that o t With thre t Make a comparison.
[0122] 152. If the recommendation value is greater than or equal to a preset threshold, a recommendation value agent set is recommended for the current agent, wherein the recommendation value agent set is the sum of agents related to the input information and historical state information of the current agent.
[0123] In this step, if o t ≥thre t , then it is the current agent i t Recommend the recently mentioned K 0 Agents, K 0 can be one or more;
[0124] 153. If the recommendation value is less than the preset threshold, a random agent is recommended for the current agent.
[0125] In this step, if o t <thre t , otherwise, it is the current agent i t Randomly recommend an agent.
[0126] Based on any of the above embodiments, based on the sampled value and the tag value, updating the feedback value is achieved by the following formula:
[0127]
[0128] Among them, p t Represents the number of historical positive samples in the historical state information, n t represents the number of historical negative samples in the historical state information, G trepresents the label value of the tth time, Represents the recommendation value of agent i at the tth time.
[0129] In this step, when an agent j mentioned by agent i in the tweet list information belongs to the agents in the mentioned list information, agent j is marked as a positive sample with a label value of 1, otherwise it is marked as a negative sample with a label value of 0.
[0130] The number of positive samples is the number of samples with a label value of 1 in the entire sample data, and the number of negative samples is the number of samples with a label value of 0 in the entire sample data.
[0131] Based on any of the above embodiments, the method further includes:
[0132] Based on the first loss function and the second loss function, the weight value in the agent utility matrix is updated; wherein the first loss function is:
[0133]
[0134] The second loss function is:
[0135]
[0136] Among them, p t Represents the number of historical positive samples in the historical state information, n t represents the number of historical negative samples in the historical state information, G t represents the label value of the tth time, o t represents the recommended value, μ represents the strategy of each agent, i and j represent agents, s ij Represents the similarity between the state of agent i and the state of agent j. The agent utility matrix is decomposed into two small matrices represented by A and B, with sizes N×d and d×N respectively, and d is less than N, a i is the i-th row of matrix A, a j is the jth row of matrix A, b i is the i-th column of matrix B, b j is the j-th column of matrix B.
[0137] In this step, the weight values are updated based on the preset agent utility matrix, which has N 2 parameters, decomposing the preset agent utility matrix into two small matrices A and B with sizes N×d and d×N, respectively, where d is less than N. To estimate these two matrices, we apply a neural network matrix factorization method, whose loss is composed of the above two loss functions.
[0138] Based on any of the above embodiments, before executing the multi-agent based recommendation method, the delayed aggregation graph neural network architecture (MA-DGNN) provided by the embodiment of the present disclosure needs to be trained. In order to obtain other agents with greater similarity and better decision-making quality, a customized experience replay mechanism method is adopted when collecting sample agents in the replay buffer in the delayed aggregation graph neural network architecture. Specific factors that need to be considered include: (i) time; (ii) whether the user feedback is positive or negative; (iii) the number of times a user mentions a certain user; (iv) the number of times a user is mentioned by other users; (v) the average performance of users in assisting other users in making decisions; and (vi) the average similarity between the user state and the states of other users.
[0139] For each sample agent i∈|length buffer |Calculate the above six factors For each factor j∈[6], we Execute the softmax function on it and get Then for each sampling process, the sample i in the playback buffer (i∈[length buffer ]), the corresponding sampling probability is
[0140] In order to improve the diversity of experience playback, the above sampling method is adopted to obtain a batch of training samples. in addition Batches are obtained by state clustering. Specifically, the samples in the buffer are divided into 10 groups according to their states, and the number of samples is randomly and uniformly selected from each group. In addition, during the sampling process, we measure the standard error of the true label of each group of samples and obtain the average std of the 10 standard errors. avg,t If std avg,t Greater than history {std avg,i} i∈[t-1] 80%, then it is judged that there may be obvious non-stationarity, so the oldest percent min{100, std avg,t / 100) to keep track of the user's latest preferences.
[0141] Further, the implementation of this disclosure is further explained with reference to Figure 4 As shown, it is a block diagram of a multi-agent based recommendation method provided by an embodiment of the present disclosure, using Figure 4 The MDP, policy network and evaluation network in Figure 5 The process in Figure 5 As shown, it is a specific flow chart of the multi-agent based recommendation method provided by the embodiment of the present disclosure, including the following steps 510 to 590:
[0142] 510. Obtain the current input information and historical status information of the current agent through MDP, wherein the historical status information includes the historical input information of the current agent and list information involving other agents, and the list information of other agents includes tweet list information and mentioned person list information.
[0143] 520, input the historical state information of the current agent and the current input information into the policy network to generate the first true value of the current agent and the first true values of other agents.
[0144] 530, based on the historical state information of the current agent and the current input information, the initial recommendation value of other agents corresponding to each input information relative to the current agent is obtained through the strategy network.
[0145] 540, label each input information with a sample label, the label value of the sample label is 1 or 0, and sample the first true value of the current agent to obtain a sampled value, and update the feedback value based on the sampled value and the label value.
[0146] 550, input the feedback value of the current agent, the feedback value of other agents and the preset discount factor into the evaluation network, output the evaluation values corresponding to the current agent and other agents, evaluate the evaluation values through the evaluation network, and output the evaluation value vectors corresponding to the current agent and other agents.
[0147] 560, input the evaluation value vector into the policy network, and output the final recommendation value o of other agents relative to the current agent t , determine the set of recommended value agents.
[0148] 570, will o t With the preset threshold thre t Make a comparison.
[0149] 580, if o t ≥thre t , then it is the current agent i t Recommend the recently mentioned K 0 Agents, K 0 The number can be one or more.
[0150] 590, if o t <thre t , otherwise, it is the current agent i t Randomly recommend an agent.
[0151] The following is a description of a multi-agent based recommendation device provided in an embodiment of the present disclosure. The multi-agent based recommendation device described below and the multi-agent based recommendation method described above can refer to each other.
[0152] Specific reference Figure 6 FIG. 1 is a schematic diagram of a multi-agent-based recommendation device provided in an embodiment of the present disclosure, and the device includes:
[0153] The acquisition module 610 is used to acquire the historical state information of the current agent when the current input information of the current agent is determined, wherein the historical state information includes the historical input information of the current agent and the list information involving other agents.
[0154] The generation module 620 is used to input the historical state information of the current agent and the current input information into the strategy network to generate the first true value of the current agent and the first true values of other agents.
[0155] The processing module 630 is used to process based on the first true value of the current agent and the first true values of other agents to obtain a feedback value.
[0156] The input module 640 is used to input the feedback value of the current agent, the feedback value of the other agents and the preset discount factor into the evaluation network, and output the evaluation value vectors corresponding to the current agent and the other agents.
[0157] The determination module 650 is used to input the evaluation value vector into the policy network, output the final recommendation value of other agents relative to the current agent, and determine the recommended value agent set.
[0158] The present disclosure provides a multi-agent based recommendation device, which determines the current input information and historical status information of the current agent, wherein the historical status information includes the historical input information of the current agent and list information involving other agents; inputs the historical status information and current input information of the current agent into a strategy network to generate a first true value of the current agent and the first true values of other agents; performs processing based on the first true value of the current agent and the first true values of other agents to obtain a feedback value; inputs the feedback value of the current agent, the feedback values of other agents and a preset discount factor into an evaluation network, and outputs evaluation value vectors corresponding to the current agent and other agents; inputs the evaluation value vectors into a strategy network, outputs final recommendation values of other agents relative to the current agent, and determines a set of recommended value agents. Not only does the feedback value get obtained through the first true value of the current agent and the first true value of other agents, but also the evaluation value vector is input into the policy network to output the final recommendation value of other agents relative to the current agent, and the agent recommended by the current agent is determined among other agents, thus achieving the final recommendation value from the individual (agent) and group (agent set) levels, thereby enabling more objective and accurate recommendations for the current agent.
[0159] Based on any of the above embodiments, the processing module 630 is specifically used for:
[0160] Based on the historical state information of the current agent and the current input information, the initial recommendation value of the other agents relative to the current agent corresponding to each input information is obtained through the strategy network;
[0161] Label each input information with a sample label, wherein the label value of the sample label is 1 or 0;
[0162] Sampling the first true value of the current agent to obtain a sampled value;
[0163] The feedback value is updated based on the sampled value and the tag value.
[0164] Based on any of the above embodiments, the input module 640 is specifically used for:
[0165] Input the feedback value of the current agent, the feedback values of the other agents, and a preset discount factor into the evaluation network, and output the evaluation values corresponding to the current agent and the other agents;
[0166] The evaluation value is evaluated through an evaluation network, and the evaluation value vectors corresponding to the current agent and the other agents are output.
[0167] Based on any of the above embodiments, the determining module 650 is specifically configured to:
[0168] Inputting the evaluation value vector into the policy network to obtain an updated policy network;
[0169] Generate a second true value of the current agent and second true values of other agents based on the updated policy network, historical state information of the current agent, and current input information;
[0170] Determine the weight values of other agents relative to the current agent based on a preset agent utility matrix; wherein the agent utility matrix includes the weight value of each agent;
[0171] Based on the second true value of the current agent, the second true values of other agents and the weight value, the final recommended values of other agents relative to the current agent are output.
[0172] Based on any of the above embodiments, the determining module 650 is further configured to:
[0173] Comparing the recommendation values of the other agents relative to the current agent with a preset threshold;
[0174] If the recommended value is greater than or equal to a preset threshold, a recommended value agent set is recommended for the current agent, wherein the recommended value agent set is the sum of agents related to the input information and historical state information of the current agent;
[0175] If the recommendation value is less than the preset threshold, an agent is randomly recommended for the current agent.
[0176] Based on any of the above embodiments, based on the sampled value and the tag value, updating the feedback value is achieved by the following formula:
[0177]
[0178] Among them, p t Represents the number of historical positive samples in the historical state information, n t represents the number of historical negative samples in the historical state information, G t represents the label value of the tth time, Represents the recommendation value of agent i at the tth time.
[0179] The multi-agent-based recommendation device provided in the embodiment of the present disclosure further includes an updating module, which is specifically used to:
[0180] Based on the first loss function and the second loss function, the weight value in the agent utility matrix is updated; wherein the first loss function is:
[0181]
[0182] The second loss function is:
[0183]
[0184] Among them, p t Represents the number of historical positive samples in the historical state information, n t represents the number of historical negative samples in the historical state information, G t represents the label value of the tth time, o t represents the recommended value, μ represents the strategy of each agent, i and j represent agents, s ij Represents the similarity between the state of agent i and the state of agent j. The agent utility matrix is decomposed into two small matrices represented by A and B, with sizes N×d and d×N respectively, and d is less than N, a i is the i-th row of matrix A, a j is the jth row of matrix A, b i is the i-th column of matrix B, b j is the j-th column of matrix B.
[0185] Figure 7 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 7 As shown, the electronic device may include: a processor (processor) 710, a communication interface (Communications Interface) 720, a memory (memory) 730 and a communication bus 740, wherein the processor 710, the communication interface 720, and the memory 730 communicate with each other through the communication bus 740. The processor 710 can call the logic instructions in the memory 730 to execute a multi-agent based recommendation method, including: obtaining the historical status information of the current agent while determining the current input information of the current agent, wherein the historical status information includes the historical input information of the current agent and the list information involving other agents; inputting the historical status information and the current input information of the current agent into the strategy network to generate the first true value of the current agent and the first true values of other agents; processing based on the first true value of the current agent and the first true values of other agents to obtain a feedback value; inputting the feedback value of the current agent, the feedback value of the other agents and a preset discount factor into the evaluation network, and outputting the evaluation value vectors corresponding to the current agent and the other agents; inputting the evaluation value vector into the strategy network, outputting the final recommendation values of the other agents relative to the current agent, and determining the recommended value agent set.
[0186] In addition, the logic instructions in the above-mentioned memory 730 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the embodiment of the present disclosure is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present disclosure. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0187] On the other hand, the present disclosure also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the multi-agent based recommendation method provided by the above methods, including: when determining the current input information of the current agent, obtaining the historical state information of the current agent, wherein the historical state information includes the historical input information of the current agent and the list information involving other agents; inputting the historical state information and the current input information of the current agent into the strategy network to generate the first true value of the current agent and the first true values of other agents; processing based on the first true value of the current agent and the first true values of other agents to obtain a feedback value; inputting the feedback value of the current agent, the feedback value of the other agents and a preset discount factor into the evaluation network, and outputting the evaluation value vectors corresponding to the current agent and the other agents; inputting the evaluation value vector into the strategy network, outputting the final recommendation values of the other agents relative to the current agent, and determining the recommended value agent set.
[0188] On the other hand, the present disclosure also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the above-mentioned multi-agent-based recommendation methods, including: obtaining historical status information of the current agent under the condition of determining the current input information of the current agent, wherein the historical status information includes the historical input information of the current agent and list information involving other agents; inputting the historical status information and the current input information of the current agent into a strategy network to generate a first true value of the current agent and the first true values of other agents; processing based on the first true value of the current agent and the first true values of other agents to obtain a feedback value; inputting the feedback value of the current agent, the feedback value of the other agents and a preset discount factor into an evaluation network, and outputting evaluation value vectors corresponding to the current agent and the other agents; inputting the evaluation value vector into a strategy network, outputting the final recommendation values of the other agents relative to the current agent, and determining a set of recommended value agents.
[0189] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0190] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0191] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them. Although the present disclosure has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present disclosure.
Claims
1. A recommendation method based on multi-agent, It is characterized in that include: When the current input information of the current agent is determined, the historical state information of the current agent is obtained, wherein the historical state information includes the historical input information of the current agent and the list information related to other agents; the current input information is the information of the user when publishing a message, including the content information of the published message or the information of the user related to the published message; the list information of other agents is the list information of tweets and the list information of persons mentioned that were previously published by the user; Input the historical state information and current input information of the current agent into the policy network to generate the first true value of the current agent and the first true values of other agents; Based on the first true value of the current agent and the first true values of other agents, a feedback value of the current agent is obtained; Input the feedback value of the current agent, the feedback values of the other agents, and a preset discount factor into the evaluation network, and output the evaluation value vectors corresponding to the current agent and the other agents; Input the evaluation value vector into the policy network, output the final recommended values of other agents relative to the current agent, and determine the recommended value agent set; The processing based on the first true value of the current agent and the first true values of other agents to obtain the feedback value of the current agent includes: Based on the historical state information of the current agent and the current input information, the initial recommendation value of the other agents relative to the current agent corresponding to each input information is obtained through the strategy network; Label each input information with a sample label, wherein the label value of the sample label is 1 or 0; Sampling the first true value of the current agent to obtain a sampled value; Based on the sampled value and the label value, update the reward value of the current agent; The step of determining a recommended value agent set includes: Comparing the recommendation values of the other agents relative to the current agent with a preset threshold; If the recommended value is greater than or equal to a preset threshold, a recommended value agent set is recommended for the current agent, wherein the recommended value agent set is the sum of agents related to the input information and historical state information of the current agent; If the recommendation value is less than the preset threshold, an agent is randomly recommended for the current agent.
2. The multi-agent based recommendation method according to claim 1, It is characterized in that The step of inputting the feedback value of the current agent, the feedback values of the other agents, and a preset discount factor into the evaluation network and outputting the evaluation value vectors corresponding to the current agent and the other agents includes: Input the feedback value of the current agent, the feedback values of the other agents, and a preset discount factor into the evaluation network, and output the evaluation values corresponding to the current agent and the other agents; The evaluation value is evaluated through an evaluation network, and the evaluation value vectors corresponding to the current agent and the other agents are output.
3. The multi-agent based recommendation method according to claim 1, It is characterized in that Inputting the evaluation value vector into the policy network to output the final recommendation value of other agents relative to the current agent, including: Inputting the evaluation value vector into the policy network to obtain an updated policy network; Generating a second true value of the current agent and a second true value of other agents based on the updated policy network, the historical state information of the current agent, and the current input information; Determining the weight value of other agents relative to the current agent based on a preset agent utility matrix; wherein, the agent utility matrix includes the weight value of each agent; Outputting the final recommendation value of other agents relative to the current agent based on the second true value of the current agent, the second true value of other agents, and the weight value.
4. The multi-agent based recommendation method according to claim 1, wherein, Updating the feedback value based on the sampled value and the tag value is achieved through the following formula: ; in, Represents the number of historical positive samples in the historical state information, Represents the number of historical negative samples in the historical state information, represents the label value of the tth time, Representing an Agent The recommended value at time t.
5. The multi-agent based recommendation method according to claim 3, wherein, The method further includes: Updating the weight value in the agent utility matrix based on a first loss function and a second loss function; wherein, the first loss function is: ; The second loss function is: ; in, Represents the number of historical positive samples in the historical state information, Represents the number of historical negative samples in the historical state information, represents the label value of the tth time, Indicates the recommended value. represents the strategy of each agent, and represents an intelligent agent, Representing an Agent The state and agent The similarity between the states of the agent is decomposed into two small matrices using and Indicates that they have dimensions and ,and Less than , is a matrix No. OK, is a matrix No. OK, is a matrix No. List, is a matrix No. List.
6. A multi-agent based recommendation device, wherein, Comprising: An acquisition module, configured to acquire the historical state information of the current agent when determining the current input information of the current agent, wherein the historical state information includes the historical input information of the current agent and the list information related to other agents; the current input information is the information when the user publishes a message, including the content information of the published message or the information of the user related to the published message; the list information of other agents is the list information of the tweets published by the user in the past and the list information of the mentioned persons; the current input information is the information when the user publishes a message, including the content information of the published message or the information of the user related to the published message; the list information of other agents is the list information of the tweets published by the user in the past and the list information of the mentioned persons; A generation module, configured to input the historical state information of the current agent and the current input information into the policy network to generate a first true value of the current agent and a first true value of other agents; A processing module, configured to process based on the first true value of the current agent and the first true value of other agents to obtain the feedback value of the current agent; An input module, configured to input the feedback value of the current agent, the feedback value of other agents, and a preset discount factor into the evaluation network to output the evaluation value vector corresponding to the current agent and other agents; A determination module, configured to input the evaluation value vector into the policy network to output the final recommendation value of other agents relative to the current agent, and determine the recommended value agent set; The processing module is specifically configured to: Based on the historical state information of the current agent and the current input information, the initial recommendation value of the other agents relative to the current agent corresponding to each input information is obtained through the strategy network; a sample label is labeled for each input information, wherein the label value of the sample label is 1 or 0; the first true value of the current agent is sampled to obtain a sample value; based on the sample value and the label value, the feedback value is updated; The determining module is further used for: The recommendation values of the other agents relative to the current agent are compared with a preset threshold; if the recommendation value is greater than or equal to the preset threshold, a recommended value agent set is recommended for the current agent, wherein the recommended value agent set is the sum of agents related to the input information and historical status information of the current agent; if the recommendation value is less than the preset threshold, an agent is randomly recommended for the current agent.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the program, the steps of the multi-agent based recommendation method as described in any one of claims 1 to 5 are implemented.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the steps of the multi-agent based recommendation method as described in any one of claims 1 to 5 are implemented.
9. A computer program product comprising a computer program, It is characterized in that When the computer program is executed by a processor, the steps of the multi-agent based recommendation method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Dialogue recommendation method and device, electronic equipment and storage medium
CN112925892A
Interactive recommendation method and system based on offline user environment and dynamic reward
CN113449183A