Diversified recommendation method based on graph neural network and reinforcement learning

By using the facility site selection method and A3C algorithm in the diversified recommendation methods of graph neural networks and reinforcement learning, the diversity problems of graph neural networks in diversified recommendations and the time-variable user interest problems are solved, and higher recommendation accuracy and diversity are achieved.

CN120086441APending Publication Date: 2025-06-03TIANJIN UNIV
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510213541.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

When existing graph neural network methods achieve diversified recommendations, it is difficult to effectively aggregate neighbor information, resulting in popular things covering up long-tail things, and user interests are time-changing, making it difficult to capture and adapt.

Method used

Diversified recommendation methods based on graph neural network and reinforcement learning are adopted, and neighbor nodes are diversified sampled through facility site selection method and A3C algorithm, recommendation strategies are dynamically adjusted, changes in user interests are captured, and users' prediction scores for each thing are calculated through graph neural network model.

Benefits of technology

It significantly improves the accuracy and diversity of the recommendation system, can better represent the basic set, alleviate the impact of popular things on long-tail things, and improves the diversity and adaptability of recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086441A_ABST
    Figure CN120086441A_ABST
Patent Text Reader

Abstract

The invention discloses a diversified recommendation method based on a graph neural network and reinforcement learning. The method comprises the following steps of: preprocessing collected data containing users and things to form a data set, and dividing the data set into a training set, a test set and a verification set: mapping the users and things in the data of the training set to a dense vector space through a lookup table of an embedded layer, and selecting the data of the users and things through a facility site selection method; iteratively screening an object with the highest dependency from the vector representation embedded by the embedding layer, and constructing a diversified neighborhood subset of each user; on the basis of the constructed neighborhood subset, performing secondary neighbor selection by using an A3C algorithm, dynamically adjusting a recommendation strategy through continuous interaction between the user and the environment, capturing and adapting to the change of user interest, and performing secondary neighbor selection; inputting a selection result of the secondary neighbor selection module into a graph neural network model for processing, and training the model; and verifying the trained model by using the verification set, and performing model performance evaluation through the test set. The diversity and adaptability of recommendation can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of personalized recommendation systems, and particularly to a diversified recommendation method based on graph neural networks and reinforcement learning. Background Art

[0002] Recommendation systems aim to provide personalized information to users to enhance the user experience. They are widely used in various industries and are one of the typical representatives of machine learning in practical applications. To maximize the utility of recommendation systems, accuracy is usually the only criterion for measuring the likelihood of user interaction with relevant data. However, a well-designed recommendation system should be evaluated from multiple perspectives, such as diversity. Accuracy only reflects correctness, and a purely accuracy-oriented approach may lead to the echo chamber and filter bubble effects, which limit users to a small set of similar data without exploring the vast majority of other data. Diversified recommendations provide a rich and diverse perspective recommendation list for users, thus meeting the diverse interest needs of users.

[0003] In the process of data recommendation, graph neural networks can easily access high-order connections and effectively capture the complex relationships between users and relevant data as things (such as authors, institutions, news, policy documents, patents, papers). The user, the thing, and their interaction relationships are represented by nodes and edges in the graph structure, so as to better model user interests and thing characteristics. However, there are also the following challenges in using graph neural networks to achieve diversified recommendations:

[0004] (1) The problem of diversity in neighbor aggregation. Traditional graph neural network methods integrate neighbor information by direct aggregation or random sampling, resulting in popular things covering up long-tail things and making it difficult to achieve diversified recommendations.

[0005] (2) The time-varying nature of user interests. User interests are not static but change over time. The recommendation system needs to be able to capture and adapt to the dynamic changes of user interests to provide more accurate and personalized recommendations.

[0006] In particular, when making diversified recommendations for literature databases such as ocean databases, the above challenges exist. Therefore, how to effectively use graph neural networks to aggregate neighbor nodes and capture the dynamic changes of user interests, and make diversified recommendations for relevant literature things based on the literature database, has become a technical problem to be solved. Summary of the Invention

[0007] The object of the present invention is to overcome the deficiencies and defects of the prior art, and provide a diversified recommendation method based on graph neural network and reinforcement learning. By using the facility location method and reinforcement learning for diversified sampling of neighbor nodes, and using the graph neural network to capture the potential interaction relationship between users and things, the accuracy and diversity of the recommendation system used in the literature database are significantly improved

[0008] A diversified recommendation method based on graph neural network and reinforcement learning, comprising:

[0009] Preprocessing the collected data containing users and things to form a data set, and dividing the data set into a training set, a test set and a validation set:

[0010] Mapping the users and things in the data of the training set to a dense vector space through the lookup table of the embedding layer, and iteratively screening the things with the highest dependence from the vector representations embedded by the embedding layer through the facility location method to construct a diversified neighborhood subset for each user;

[0011] Based on the constructed neighborhood subset, use the A3C algorithm for secondary neighbor selection, and dynamically adjust the recommendation strategy through the continuous interaction between the user and the environment to capture and adapt to the changes in the user's interests, and perform secondary selection of neighbors;

[0012] Input the selection result of the secondary neighbor selection module into the graph neural network model for processing, train the model, calculate the predicted score of each thing by the user, sort the things in descending order according to the predicted score, the higher the score, the more forward the item, and present the sorted list of things to the user as the recommendation result;

[0013] Use the validation set to verify the trained model, and evaluate the model performance through the test set.

[0014] Among them, the mapping of the users and things in the data to a dense vector space includes:

[0015] For a set of users U = {u1, u2,..., u|U|}, a set of things I = {i1, i2,..., i|I|}, each thing is mapped to its category through the mapping function C(·);

[0016] The user-thing interaction is represented as an interaction matrix R ∈ R|U|×|I|. If the user u interacts with the thing i, then R u,i = 1, otherwise R u,i = 0; The historical interaction is represented by the user-thing bipartite graph G = (V, E), where V = U ∪ I. If R u,i = 1, then there is an edge e u,i ∈ E;

[0017] Finally, the user-thing is mapped to a dense vector space, expressed as follows:

[0018]

[0019] where, e (0) ∈R d is the d-dimensional dense vector of the user-thing.

[0020] Among them, the method of iteratively screening the thing with the highest dependence degree through the facility location method to construct a diverse domain subset for each user includes:

[0021] First, identify the neighborhood subset S u selected by user u and the most similar thing i' in each thing i in the basic set , and then sum the similarity values; the higher the dependence value, the higher the similarity between the two subsets; N u represents all neighbors of a user node u; the facility location method is defined as follows:

[0022]

[0023] The similarity sim(i, i') between thing i and i' is measured by a Gaussian kernel parameterized by the kernel width σ 2 :

[0024]

[0025] S u is restricted to no more than k items, that is, |S u | ≤ k. Under this constraint, maximizing the submodular function is NP-hard, and the greedy algorithm approximately solves it with 1 - e -1 ; the greedy algorithm starts from an empty set and then adds a thing i ∈ I\S u with the largest marginal benefit to S u in each step:

[0026] S u ← S u ∪ i * ,

[0027]

[0028] After k-step greedy neighborhood selection, a diverse neighborhood subset for each user is obtained.

[0029] Among them, the secondary neighbor selection using the A3C algorithm is to learn a strategy through the interaction between the user and the environment to maximize the long-term cumulative reward. The learning process follows the framework of the Markov decision process MDP: In the recommendation scenario, MDP is a five-tuple <S, A, P, R, γ>. The state space S describes the environmental state in a fixed-length historical trajectory, including all possible state sets, and each state S t describes the current environmental situation at time step t; the action space A defines all possible behaviors that the recommendation system takes when interacting with the user. At time step t, the system selects and executes action A t , that is, recommends a thing to the user; the state transition probability P defines the probability p(s′, r|s, a) that the system transfers to a new state s′ and the agent receives user feedback given the current state s and the action a taken, reflecting the possible changes in the state of the user after responding to the recommendation; the reward function R defines the immediate reward r(s, a) obtained by the recommendation system when transferring from state s to a new state after taking action a; the discount factor γ is a real number between 0 and 1, used to balance the importance of immediate rewards and future rewards. When calculating the cumulative reward, future rewards are multiplied by the power of γ and gradually decay as the time step increases; the closer γ is to 1, the more the system tends to pursue long-term high returns.

[0030] Among them, when using the A3C algorithm for secondary neighbor selection, the actor-critic method is adopted, combining the policy-based method and the value-based method to learn the policy and the value function together. As an actor, the policy-based method trains the policy according to the value function feedback by the value-based method as a critic, and the critic trains the value function and updates it step by step using the temporal difference method.

[0031] Among them, the formula for replacing the full return with a single-step return in the single-step actor-critic method is:

[0032]

[0033] Among them, S t represents the screening state of the diverse neighborhood subset S u of the user; A t represents selecting things from the set S u for recommendation; the reward R t+1 includes user feedback rewards and diversity rewards; the user feedback reward means giving a positive reward when the user clicks on the recommended thing; if the user does not click, a negative reward is given; the diversity reward means increasing the diversity of recommendations by calculating the similarity between the recommended things; the lower the similarity between the recommended things, the higher the diversity of the recommendations and the greater the reward.

[0034] Among them, the graph neural network model is combined with a layer attention mechanism to dynamically adjust the weights of different-order neighbors, increase the diversity of high-order neighbors, and selectively aggregate information; each user-thing has L embedding representations generated by L GNN layers, and the readout function on [e (0) ,e (1) ,...,e (L) is used to learn different weights of the GNN layer to optimize the loss function:

[0035]

[0036] Among them, a (l) is the attention weight of the l-th layer, and W Att ∈R d is the parameter for attention calculation.

[0037] Among them, during model training, the training sample weights are adjusted according to the number of valid things of the category of the thing, so that the training of long-tail categories receives more attention and the learning ability of the model for rare categories is improved; borrowing the idea of category balance loss, the samples are weighted according to the number of valid things of the category; the calculation method of the weight wC(i) is:

[0038]

[0039] Among them, β is a hyperparameter that determines the weight. The larger β is, the smaller the weight of the popular category will be.

[0040] Among them, the data in the dataset comes from the ASEAN Marine Science and Technology Platform, covering behaviors such as clicks and skips. The dataset is split according to the ratio of 60% training set, 20% validation set, and 20% test set; the training set is used for model training, the validation set is used for model hyperparameter tuning and early stopping strategies to optimize model performance, and the results of the test set are used as the final evaluation basis for model performance.

[0041] Among them, the data preprocessing includes data filtering. An N-core setting is adopted to filter out sparse data, so that the collected data only retains users or things with at least N interactions.

[0042] The present invention selects a diverse subset of neighbors through a facility location method, enabling the user embedding to more evenly represent the base set, thereby alleviating the impact of popular things on long-tail things and improving the diversity of the recommendation system. The sampling based on A3C dynamically adjusts the recommendation strategy through asynchronous interaction and policy update, captures the changes in user interests while optimizing the long-term cumulative reward, and enhances the diversity and adaptability of the recommendation.

[0043] The present invention uses a lightweight graph convolutional neural network (LGC) as the backbone GNN layer, abandons feature transformation and non-linear activation, and directly aggregates the embeddings of neighbors selected by facility location and A3C sampling. Layer attention is used in the GNN to improve the diversity of higher-order neighbors by dynamically adjusting the influence of different-order neighbors, while alleviating the over-smoothing problem caused by the increase in the number of layers.

[0044] The present invention adjusts the loss function, increases the weights of long-tail categories, and combines Bayesian personalized ranking loss to improve the diversity and personalization effect of recommendations. Brief Description of the Drawings

[0045] Figure 1 is a schematic flowchart of the diversified recommendation method based on graph neural network and reinforcement learning of the present invention;

[0046] Figure 2 is a schematic overall architecture diagram of the model of the diversified recommendation method based on graph neural network and reinforcement learning of the present invention.

[0047] Figure 3 is a schematic diagram of the diversified recommendation effect in the marine science and technology platform of the embodiment of the present invention. Detailed Description of the Invention

[0048] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0049] See Figure 1 As shown, a diversified recommendation method based on graph neural network and reinforcement learning includes:

[0050] Preprocess the collected literature data containing users and things to form a data set, and divide the data set into a training set, a test set, and a validation set:

[0051] Map the users and things in the data of the training set to a dense vector space through the lookup table of the embedding layer, and iteratively screen the things with the highest dependence from the vector representations embedded by the embedding layer through the facility location method to construct a diversified neighborhood subset for each user;

[0052] Based on the constructed neighborhood subset, use the A3C algorithm for secondary neighbor selection, and dynamically adjust the recommendation strategy through the continuous interaction between the user and the environment to capture and adapt to the changes in the user's interests, and perform secondary selection of neighbors;

[0053] Input the selection result of the secondary neighbor selection module into the graph neural network model for processing, train the model, calculate the prediction score of the user for each thing, sort the things in descending order according to the prediction score, with the item having a higher score ranked higher, and present the sorted list of things to the user as the recommendation result;

[0054] Use the validation set to verify the trained model and perform data preprocessing for model performance evaluation through the test set.

[0055] The method of the embodiment of the present application is mainly based on Figure 1 the model architecture of diversified recommendation based on graph neural network and reinforcement learning to achieve diversified recommendation of the data in the literature database to users when the users browse the database.

[0056] Specifically, the data collection can be the Marine dataset that collects the user behavior of the marine science and technology platform through log data or a similar literature database. The data covers click behaviors, and the data transactions include authors, institutions, news, policy documents, patents, papers, etc.

[0057] To ensure the quality of the dataset, it is necessary to filter the data. For example, set 10 cores for the dataset, that is, only retain users or things with at least 10 interactions to filter out sparse data, improve the stability and recommendation effect of the model, reduce data noise at the same time, and make the training more efficient.

[0058] In some embodiments of the present application, the formed dataset is split according to the ratio of 60% training set, 20% validation set, and 20% test set. The validation set is used for hyperparameter tuning and early stopping strategies to optimize the model performance, and the result of the final test set is used as the final evaluation basis for the model performance.

[0059] In the embodiment of the present application, the embedding layer represents the feature information of users and things, maps users and things to a dense vector space through a lookup table, and provides input for subsequent sampling and the graph neural network model.

[0060] For the diversified recommendation task, there is a set of users U = {u 1 ,u 2 ,...,u |U|}, a set of things I = {i 1 ,i 2 ,...,i |I|}, and a mapping function C(·). The mapping function C(·) maps each thing to its category, that is, each transaction corresponds to a similarity. For example, taking marine data as an example, the category includes authors, institutions, news, policy documents, patents, papers, and there are multiple things with the same attributes in the category.

[0061] Represent the user-thing interaction as an interaction matrix \(R\in\mathbb{R}\) |U|×|I| . If user \(u\) interacts with thing \(i\), then \(R\) u,i = 1, otherwise \(R\) u,i = 0. For the graph-based recommendation model, the historical interaction is represented by the user-thing bipartite graph \(G=(V, E)\), where \(V = U\cup I\). If \(R\) u,i = 1, then there is an edge \(e\) u,i \(\in E\) between \(u\) and \(i\).

[0062] The recommendation system learned from the user-thing bipartite graph \(G\) aims to recommend the \(k\) most interesting things \(i\) 1 , \(i\) 2 ,..., \(i\) k for each user \(u\). The diverse recommendation task requires the top \(k\) recommended things to be different from each other. The diversity of the recommendation list is usually measured by the coverage of the recommended categories .

[0063] The embedding layer mentioned above is a lookup table that maps the user-thing to a dense vector \(E\) (0) , expressed as follows:

[0064]

[0065] where \(e\) (0) \(\in\mathbb{R}\) d is the \(d\)-dimensional dense vector of the user-thing. Then, the embeddings indexed from the embedding table are input into the graph neural network \(GNN\) for information aggregation. Thus, it is called the output of the "zero" layer

[0066] In the embodiments of the present application, the subset is constructed by iteratively screening the things with the highest similarity through the facility location method, greatly improving the diversity of the candidate nodes.

[0067] The facility location method is a widely used sub-module method for evaluating diversity. This method first identifies the thing in each thing \(i\) in the selected subset \(S\) u that is most similar to the basic set , and then sums the similarity values, where \(N\) u represents all the neighbors of a user node \(u\). The higher the value, the higher the similarity between the two subsets. In other words, the selected subset is very diverse and can better represent the basic set. The definition of the facility location method is as follows:

[0068]

[0069] where \(S\) u is the neighborhood subset selected by user \(u\), and the similarity of things is measured by the Gaussian kernel parameterized by the kernel width \(\sigma\) 2 :

[0070]

[0071] Among them, S u is restricted to no more than k items, that is, |Su| ≤ k. Maximizing the submodular function under this constraint is NP-hard, but the greedy algorithm can use 1 - e -1 to approximately solve it. The greedy algorithm starts from an empty set and then at each step adds an item i ∈ I\S with the largest marginal benefit to S u : u :

[0072] S u ← S u ∪ i * ,

[0073]

[0074] After k-step greedy neighborhood selection, a diverse neighborhood subset for each user can be obtained, and then this subset is aggregated.

[0075] In the embodiments of the present application, the A3C algorithm is used for secondary neighbor selection, and by continuously interacting with the environment, the recommendation strategy is dynamically adjusted, so as to effectively capture and adapt to the changes in user interests.

[0076] In the reinforcement learning using the A3C algorithm, especially in the context of applying it to the recommendation system task, the problem is redefined as an optimization problem, the core of which is to learn a strategy through the user's continuous interaction with the environment to maximize the long-term cumulative reward. This learning process follows the framework of the Markov Decision Process (MDP). More precisely, the MDP in the recommendation scenario is a five-tuple <S, A, P, R, γ>. The state space S describes the environmental state in a fixed-length historical trajectory, including all possible state sets, and each state S t describes the current environmental situation at time step t; the action space A defines all possible behaviors that the recommendation system takes when interacting with the user. At time step t, the system selects and executes action A t, that is, recommending a thing to the user; the state transition probability P defines the probability p(s′, r|s, a) that the system transfers to a new state s′ and the agent receives user feedback given the current state s and the action a taken, reflecting the possible changes in the user's state after responding to the recommendation; the reward function R defines the immediate reward r(s, a) obtained by the recommendation system when transferring from state s to a new state after taking action a; the discount factor γ is a real number between 0 and 1, used to balance the importance of immediate rewards and future rewards; when calculating the cumulative reward, future rewards are multiplied by the power of γ and gradually decay as the time step increases; the closer γ is to 1, the more the system tends to pursue long-term high rewards.

[0077] In the MDP in the recommendation scenario, the value function method can be used to achieve the global optimal return, that is, to enable the agent to achieve the global optimal return through the best action in state s, that is, the maximum value is determined by the optimal policy π * The best action a * after that. The optimal state value function is defined by the Bellman equation as follows:

[0078]

[0079] where is the expected value under the optimal policy π * The action value function represents the expected value of the cumulative reward that the agent may obtain in the future after taking action a * in state s and then acting according to the policy π * . The Bellman equation of the action value function is defined as:

[0080]

[0081] Different from the optimal state value function , it evaluates the long-term benefits of taking a certain action in a certain state under a certain policy, helping the agent select a better action.

[0082] In addition to using the value function, the policy search method can also be used to optimize the policy. Compared with the value function method, the policy search method directly optimizes the policy, and these policies are parameterized by a set of policy parameters θ t These policy parameters can be updated through gradient-free or gradient-based optimization methods to maximize the expected return. To find the optimal policy, the gradient-based optimization method uses the gradient of a certain performance measure J(θ) to learn the policy parameters. Formally, the approximate gradient ascent function for updating J(θ) is:

[0083]

[0084] in, Indicates that in θ t Approximately J(θ t ) is a random estimate of the gradient, and α is the step size that affects the learning rate. The policy gradient method gives an analytical expression for the gradient of J(θ), and uses an equation proportional to the policy gradient and Monte Carlo sampling to approximate the expected value, thereby estimating the gradient. The formula is as follows:

[0085]

[0086] Preferably, the embodiment of the present application uses the actor-critic method to implement the Markov decision process. The actor-critic method combines the advantages of the value function method and the policy search method. It attempts to estimate a value function while using policy gradients to search in the policy space. The actor-critic is one of the most representative algorithms. It combines the policy-based method (i.e., actor) with the value-based method (i.e., critic) to learn the policy and value function together.

[0087] Furthermore, in this application, in order to improve the high variance problem in the policy gradient method, the advantage function is introduced, which is an important concept in reinforcement learning. It is used to measure the average performance of an action relative to the current state. Specifically, the advantage function measures the degree of advantage of taking a specific action in a certain state compared to the average level. The advantage function is defined as follows:

[0088]

[0089] in is the action value function, is the state value function. The advantage function Adv(s,a) measures the quality of choosing an action a compared to the average action under the current strategy in a given state. If Adv(s,a)>0, it means that choosing action a is worse than the average action. On the contrary, if Adv(s,a)<0, it means that action a is not as good as the average action.

[0090] In the embodiment of the present application, when the actor-critic method is used to implement the Markov decision process, the actor trains the strategy according to the value function fed back by the critic, and the critic trains the value function and updates it in a single step using the temporal difference method (TD). The formula for replacing the complete return with a single step return in the single-step actor-critic algorithm is:

[0091]

[0092] Among them, S tIndicates the screening status of the diverse neighborhood subset S for the user u ; A t Indicates selecting things from the set S u for recommendation; Reward R t+1 Includes two parts: user feedback reward and diversity reward. ω represents the state value weight vector learned by the actor-critic algorithm

[0093] The user feedback reward means that when the user clicks on the recommended thing, a positive reward is given; if the user does not click, a negative reward is given. The diversity reward means that by calculating the similarity between the recommended things, the diversity of the recommendation is increased. The lower the similarity between the recommended things, the higher the diversity of the recommendation, and the greater the reward should be

[0094] In the embodiments of this application, the A3C (Asynchronous Advantage Actor-Critic) algorithm adopts an asynchronous training method. Each thread independently interacts with the environment and achieves gradient update through parameter sharing. This asynchronous training method can improve the efficiency and stability of training and can learn better policies and value functions

[0095] In the embodiments of this application, layer attention is used to increase the diversity of high-order neighbors while alleviating the over-smoothing problem. In a graph neural network, different GNN layers generate embeddings based on information from different subsets of nodes: the l-th layer aggregates from the l-th hop neighbors. Diversified embeddings can be achieved by aggregating from high-order neighbors. However, the direct stacking of several GNN layers will lead to the over-smoothing problem. Therefore, layer attention is designed in the GRRec embodiment of this application to increase the diversity of high-order neighbors while alleviating the over-smoothing problem

[0096] For each user-thing, there are L embedding representations generated by L GNN layers. Layer attention is a readout function learned through the attention mechanism on the embedding representations [e (0) , e (1) ,..., e (L) :

[0097]

[0098] where a(l) is the attention weight of the l-th layer. The calculation method is:

[0099]

[0100] W Att ∈R d is the parameter for attention calculation, e (l) , e (l`)Denote the embedding representations of nodes at \(l\) and \(l'\). Thus, the attention mechanism can learn different weights \(a(l)\) of the GNN layer to optimize the loss function, thereby effectively alleviating the over-smoothing problem.

[0101] In the embodiments of the present application, by adjusting the sample weights according to the number of valid things in each category, the training of long-tail categories receives more attention, thereby improving the model's learning ability for rare categories.

[0102] In reality, the number of things within each category is highly unbalanced. A small number of categories contain the most things, while most categories have only a limited number of things. Training the model by directly optimizing the average loss of all samples makes the training of long-tail categories hardly noticeable. Therefore, in the embodiments of the present application, GRRec re-weights the sample losses during the training process according to their categories. If a thing belongs to a popular category, GRRec will relatively reduce the weight, and if a thing belongs to a long-tail category, GRRec will relatively increase the weight.

[0103] In practice, borrowing the idea of category balance loss, re-weight the sample \((u, i)\) according to the number of valid things in the category. The weight \(w\) C(i) is calculated as:

[0104]

[0105] where \(\beta\) is a hyperparameter that determines the weight. The larger \(\beta\) is, the smaller the weight of the popular category will be.

[0106] Experimental verification

[0107] 1. Experimental dataset

[0108] Marine dataset: Users click to browse marine science and technology resource information.

[0109] The detailed list of the dataset is shown in the table:

[0110] Table 1 Dataset description table

[0111]

[0112] 2. Evaluation metrics

[0113] Accuracy metrics: such as Rec, HR.

[0114] Diversity metrics: such as Cov.

[0115] 3. Experimental settings

[0116] The comparison models include traditional collaborative filtering, recommendation models based on graph neural networks, and diversified recommendation models without integrated reinforcement learning. Parameter tuning and ablation experiments are conducted to analyze the effects of the number of graph neural network layers, the number of aggregated neighbors, neighbor similarity, and loss reweighting in the models on accuracy and diversity.

[0117] 4. Experimental Results

[0118] The model of the present invention is superior to the comparison models in terms of diversity evaluation metrics and achieves the optimal or sub-optimal results in terms of accuracy evaluation metrics. In particular, it shows significant advantages in neighbor selection based on reinforcement learning. The following table shows the comparison of experimental results with mainstream recommendation models based on graph neural networks. The best and sub-optimal results are indicated in bold and underlined respectively. Among them, GRRec is the model proposed by the method of the present invention. The comparative evaluation results of the algorithms are shown in the following figure:

[0119] Table 2 Comparative Analysis Table of Performance Evaluation of Marine Dataset

[0120]

[0121] From the description of the embodiments of the present application, it can be seen that the present application has the following beneficial effects:

[0122] 1. Improve recommendation accuracy

[0123] Adopt a lightweight graph convolutional neural network to directly aggregate neighbor information, enabling more accurate modeling of the complex relationships between users and things and improving the matching degree of recommendations. Optimize the recommendation strategy through reinforcement learning, and dynamically adjust the recommended content in combination with the user's historical behavior, making the recommendation results more in line with the user's interest changes and improving the personalization degree of recommendations.

[0124] 2. Enhance recommendation diversity

[0125] Neighbor selection is carried out through the facility location method to ensure that the recommended things cover more categories and avoid the problem of popular things dominating the recommendation list. The layer attention mechanism is adopted to optimize the contribution of different-order neighbors to the recommendation results, reduce the over-reliance on low-order neighbors, and improve the diversity of the recommendation list. Through reinforcement learning combined with the diversity reward mechanism, while optimizing the user click-through rate, the exploration of the recommended content is increased, making the recommendation list more diverse.

[0126] 3. Have dynamic adaptability

[0127] The reinforcement learning strategy allows the system to continuously learn the user's feedback and dynamically adjust the recommendation strategy to adapt to the changes in the user's interests. The reward mechanism not only focuses on the user's click behavior but also combines the diversity reward, enabling the recommendation system to continuously optimize the recommendation effect during the long-term interaction process.

[0128] The technology of the embodiments of the present application can better explore users' interest points, provide more diverse recommendations, be applicable to professional fields (such as the Marine dataset), and can provide accurate and diverse recommendations in relatively professional fields (such as marine technology platforms), improving the user experience.

[0129] The foregoing has shown and described the basic principles, main features and advantages of the present invention. For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and without departing from the spirit or basic features of the present invention, the present invention can be implemented in other specific forms.

[0130] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to embrace all changes falling within the meaning and scope of the equivalent elements of the claims in the present invention.

[0131] In addition, it should be understood that although this specification is described in accordance with the embodiments, not every embodiment only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A diversified recommendation method based on graph neural network and reinforcement learning, characterized in that: include: Preprocess the collected data including users and things to form a data set, and divide the data set into a training set, a test set, and a validation set: The users and things in the training set data are mapped to a dense vector space through the lookup table of the embedding layer. Through the facility location selection method, the things with the highest degree of dependence are iteratively selected from the vector representation embedded in the embedding layer to construct a diverse neighborhood subset for each user. Based on the constructed neighborhood subset, the A3C algorithm is used to perform secondary neighbor selection. Through the continuous interaction between the user and the environment, the recommendation strategy is dynamically adjusted to capture and adapt to the changes in user interests and perform secondary neighbor selection. Input the selection results of the secondary neighbor selection module into the graph neural network model for processing, train the model, calculate the user's predicted score for each item, sort the items in descending order according to the predicted score, and present the sorted item list to the user as the recommendation result; The trained model is verified using the validation set, and the model performance is evaluated using the test set.

2. The diversified recommendation method based on graph neural network and reinforcement learning according to claim 1, characterized in that: The mapping of users and things in the data into a dense vector space includes: For a set of users U = {u1,u2,...,u|U|} and a set of things I = {i1,i2,...,i|I|}, each thing is mapped to its category through the mapping function C(·); The user-thing interaction is represented as an interaction matrix R∈R|U|×|I|. If user u interacts with thing i, then R u,i =1, otherwise R u,i = 0; historical interactions are represented by the user-thing bipartite graph G = (V, E), where V = U ∪ I. If R u,i = 1, then there is an edge e between u and i u,i ∈E; Map users-things into a dense vector space: Among them, e (0) ∈R d is a d-dimensional dense vector of user-things.

3. The diversified recommendation method based on graph neural network and reinforcement learning according to claim 1, characterized in that: The facility location selection method is used to iteratively select the things with the highest degree of dependency to construct a diverse domain subset for each user, including: First, identify the neighborhood subset S selected by user u u With the basic collection The most similar thing i′ in each thing i in N, and then sum up the similarity values; the higher the dependency value, the higher the similarity of the two subsets; N u represents all neighbors of a user node u; the facility location method is defined as follows: The similarity sim(i, i′) between objects i and i′ is given by the kernel width σ 2 Parameterized Gaussian kernel metric: S u is restricted to no more than k items, i.e. |S u |≤k, maximizing the submodular function under this constraint is NP-hard, and the greedy algorithm uses 1-e -1 Approximate solution; Greedy algorithm from an empty set Start, then each step towards S u Add a thing with the largest marginal benefit i∈I\S u : S u ←S u ∪i * , After k steps of greedy neighborhood selection, a diverse neighborhood subset for each user is obtained.

4. The diversified recommendation method based on graph neural network and reinforcement learning according to claim 1, characterized in that: The use of the A3C algorithm for secondary neighbor selection is to learn a strategy through interaction between the user and the environment to maximize the long-term cumulative reward; The learning process follows the framework of Markov decision process MDP: In the recommendation scenario, MDP is a five-tuple<S、A、P、R、γ> , the state space S describes the state of the environment in a fixed-length history trajectory, which contains all possible state sets. Each state S t It describes the current environmental conditions at time step t; The action space A defines all possible actions that the recommendation system can take when interacting with a user. At time step t, the system selects and executes action A. t , that is, recommending something to the user; the state transition probability P defines the probability p(s′,r|s,a) that the system transfers to a new state s′ and the agent receives user feedback after a given current state s and action a, reflecting the possible changes in the user's state after responding to the recommendation; the reward function R defines the immediate reward r(s,a) obtained by the recommendation system when it transfers from state s to a new state after taking action a; the discount factor γ is a real number between 0 and 1, which is used to balance the importance of immediate rewards and future rewards; when calculating the cumulative reward, the future reward is multiplied by the power of γ, which gradually decays with the increase of time steps; the closer γ is to 1, the more the system tends to pursue long-term high returns.

5. The diversified recommendation method based on graph neural network and reinforcement learning according to claim 4 is characterized in that: When the A3C algorithm is used for secondary neighbor selection, the actor-critic method is adopted, and the policy-based method is combined with the value-based method to learn the policy and value function together. The actor-based policy method trains the policy according to the value function fed back by the critic-value method, and the critic trains the value function and updates it in a single step using the temporal difference method.

6. The diversified recommendation method based on graph neural network and reinforcement learning according to claim 5, characterized in that: The formula for the single-step actor-critic method to replace a full return with a single-step return is: Among them, S t Represents the diverse neighborhood subset S of the user u The screening status of t Indicates selecting things from the set Su for recommendation; reward R t+1 It includes user feedback reward and diversity reward; user feedback reward refers to giving positive rewards when users click on recommended items; if users do not click, negative rewards are given; diversity reward refers to increasing the diversity of recommendations by calculating the similarities between recommended items; the lower the similarity between recommended items, the higher the diversity of recommendations and the greater the reward.

7. The diversified recommendation method based on graph neural network and reinforcement learning according to claim 1, characterized in that: The graph neural network model is combined with the layer attention mechanism to dynamically adjust the weights of neighbors of different orders, increase the diversity of high-order neighbors, and selectively aggregate information; each user-thing has L embedding representations generated by L GNN layers, and the attention mechanism is used to learn [e (0) ,e (1) ,...,e (L) ], learning different weights of the GNN layer to optimize the loss function: Among them, a (l) is the attention weight of the lth layer, W Att ∈R d are the parameters for attention calculation.

8. The diversified recommendation method based on graph neural network and reinforcement learning according to claim 7, characterized in that: During model training, the weight of training samples is adjusted according to the number of valid objects in the object category, so that the training of long-tail categories gets more attention and the model's learning ability for rare categories is improved; the idea of ​​category balance loss is borrowed to weight samples according to the number of valid objects in the category; the weight wC(i) is calculated as: Among them, β is a hyperparameter that determines the weight. The larger β is, the smaller the weight of the popular category will be.

9. The diversified recommendation method based on graph neural network and reinforcement learning according to claim 1, characterized in that: The data in the dataset comes from the ASEAN Marine Science and Technology Platform, covering behaviors such as clicks and skips. The dataset is split into 60% training set, 20% validation set, and 20% test set. The training set is used for model training, and the validation set is used for model hyperparameter tuning and early stopping strategy to optimize model performance. The results of the test set are used as the final evaluation basis for model performance.

10. The diversified recommendation method based on graph neural network and reinforcement learning according to claim 1, characterized in that: The data preprocessing includes data filtering, adopting an N-core setting to filter out sparse data so that the collected data only retains users or things with at least N interactions.

Citation Information

Cited By

  • Event deduction and early warning method based on causal inference

    CN120258155A

  • Pumped storage station selection method, system and equipment based on reinforcement learning

    CN121303538A

  • A method, system and device for pumped storage site selection based on reinforcement learning

    CN121303538B

  • Method and system for intelligently judging electricity stealing behaviors of low-voltage resident users based on large language model

    CN121765487A

  • A Smart Method and System for Identifying Electricity Theft Behavior of Low-Voltage Residential Users Based on Large Language Models

    CN121765487B