Cold Start Method for Recommendation Systems Based on Meta-Learning and Long Short-Term Memory Networks
By introducing meta-learning and long-term memory networks into the recommendation system, high-order meta-models and neural task schedulers are built, the problem of sparse interaction information in cold start of the recommendation system is solved, and higher recommendation accuracy and more effective task sampling and preference prediction are achieved.
Patent Information
- Application Number
- CN202211296829.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-21
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-10-21
AI Technical Summary
In the cold start scenario of the recommendation system, the interactive information between the user and the project is very sparse, making it difficult to achieve good results based on collaborative filtering.
A cold start method for recommendation system based on meta-learning and long-term short-term memory networks is proposed. By constructing a high-order meta-model and neural task scheduler, integrating task sampling and preference prediction, a more accurate node embedding is generated using the potential semantic relationship between node attributes and network topology.
It effectively solves the problem of sparse interaction information in cold start of recommendation system, improves the accuracy of recommendations, and improves the effect of task sampling and preference prediction through a two-way promotion mechanism.
Smart Images

Figure CN115600648B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of meta - learning and graph neural networks, and proposes a meta - learning model for a sparse network of structured data to improve the accuracy of recommendations. Specifically, it relates to a cold - start method for a recommendation system based on meta - learning and long - short - term memory networks. Background Art
[0002] With the development of network technology, recommendation systems have attracted great attention in the past decade to solve the problem of users' information overload (e.g., e - commerce platforms and online food ordering platforms). Collaborative - filtering - based methods have played a key role in solving recommendation system problems. They usually estimate the likelihood of purchasing the current item based on the user's historical interaction information (such as past purchase records). However, in cold - start scenarios, the interaction information between users and items is very sparse, and it is difficult for collaborative - filtering - based methods to achieve good results. Recently, due to the success of meta - learning technology in the field of computer vision, many researchers have started to focus on its feasibility in the field of graph neural networks. The core idea is to learn prior knowledge in similar tasks so as to quickly adapt to new tasks with a small amount of training data. Current meta - learning technologies are mainly divided into three types: metric - based meta - learning, model - based meta - learning, and optimization - based meta - learning. Metric - based meta - learning learns a good similarity kernel function to measure the relationship between samples. Model - based meta - learning relies on a model that can learn quickly, and its effect depends on the internal structure of the model. Optimization - based meta - learning algorithms assign two data sets to each task, a support set and a query set. The support set and the query set are used to calculate the training loss and test loss of each task respectively. The optimization - based method is simpler and more effective than other methods and has become one of the mainstream methods in recent years. Most existing meta - learning methods randomly sample tasks with a uniform probability. This assumption means that all tasks are equally important for new tasks, which is obviously inconsistent with reality. Recently, some studies have used manually defined or fixed sample sampling strategies (such as defining a reward function in reinforcement learning) to solve this problem. These methods are greatly affected by human factors and have great limitations in improving the effect. Summary of the Invention
[0003] The purpose of the present invention is to overcome the problem of sparse interaction information in the cold - start of the recommendation system, and proposes a cold - start method for the recommendation system based on meta - learning and long - short - term memory networks. The cold - start method integrates a model of task sampling and preference prediction, and at the same time utilizes the potential semantic relationship between node attributes and network topology to generate more accurate node embeddings, thereby improving the accuracy of recommendations. Moreover, preference prediction also guides the process of task sampling, realizing two - way promotion.
[0004] To achieve the above object, the present invention provides a cold start method for a recommendation system based on meta-learning and long short-term memory networks, including:
[0005] Step 1: Construct a high-order meta-model, where the high-order meta-model includes a feature aggregator and a task predictor; the feature aggregator uses a heterogeneous neural network to capture high-order semantic relationships between nodes, supplement the situation where there is less current user interaction information, uses a graph convolutional neural network as an encoder to encode node attributes and network topologies, and generates node embeddings; use a fully connected neural network to construct the task predictor, which takes the user node embedding and the node embedding of the evaluated item as inputs and outputs the predicted value of the evaluation score;
[0006] Step 2: Take the true value and predicted value of the item evaluation score as inputs and calculate the square loss function;
[0007] Step 3: Construct a neural task scheduler with a long short-term memory network, and take the square loss function of the query set and the gradient similarity of the square loss functions of the support set and the query set as inputs to obtain the sampling probability of the current task, and finally obtain the sampling weight of the current task relative to the new task:
[0008]
[0009]
[0010] In the formula, Ms represents the gradient similarity of the loss functions of the support set and the query set, respectively represent the gradients of the support set and the query set, represents the long short-term memory network with parameter ω u represents the sampling weight of the current user u;
[0011] Step 4: Use the method of gradient descent to minimize the gap between the true value and the predicted value, construct the joint optimization of the neural task scheduler and the high-order meta-model, and obtain the optimal parameters of the high-order meta-model and the task scheduler for direct application to new user tasks that have not been trained.
[0012] Furthermore, use the loss function to train the high-order meta-model and the long short-term memory network, and train the high-order meta-model and the neural task scheduler through the loss function and the gradient similarity of the loss function; the training process of the cold start method uses the Adam optimizer to train the neural network parameters.
[0013] Furthermore, Step 1 specifically includes the following steps:
[0014] Sample the user-item interactions to obtain the dataset of meta-learning tasks, where the dataset includes a support set and a query set; assume that the task of a user u is Tu =(S u , Q u ), according to the idea of meta - learning, the items interacting with user u are divided into a support set S u and a query set Q u ; the support set is used for training data, and the query set is used for test data;
[0015] Use a graph convolutional neural network to perform feature aggregation on the high - order context item information of users in the heterogeneous neural network, and generate a preliminary node embedding of the user, where the user node embedding captures the topological relationship between context item nodes:
[0016] x u = g φ (u, b)=σ(MEAN({Wb + b: b ∈ B}))
[0017] In the formula, b represents the item attributes interacting with the current user, σ is the activation function, φ=(W, b) represents the parameters of the graph convolutional neural network g, and g is named the feature aggregator;
[0018] Step 2: Use a fully - connected neural network to construct a task predictor. Utilize the generated preliminary node embedding of the user and the items rated by the user, and use the task predictor to predict the score of the user for the item; then, based on the predicted score, measure the gap between the predicted value and the true value obtained in Step 1, and represent it with a squared loss function:
[0019]
[0020]
[0021] In the formula, h w represents a predictor with parameters w, MLP represents a fully - connected neural network, x u , x b respectively represent the embeddings of the user and the item, L us represents the loss function of the support set of user u, |B s | represents the number of items in the support set, r u,b , respectively represent the true value and the predicted value of user u's score for item b;
[0022] Finally, the feature aggregator and the task predictor are unified into a high - order meta - model together, and its parameters θ=(φ, w).
[0023] Furthermore, Step 4 specifically includes the following steps:
[0024] a) Taking the loss function of the high - order meta - model and the corresponding gradient similarity as inputs, model two aspects of the training results and training process of the high - order meta - model respectively to update the parameters of the neural task scheduler.
[0025]
[0026] b) Taking the sampling weights calculated by the neural task scheduler as a reference, select the tasks with higher sampling weights for the training of the high - order meta - model to update its parameters; this process can ensure that tasks providing more information will be used for training more preferentially, alleviating the limitations of unified sampling of all tasks to a certain extent:
[0027]
[0028]
[0029] c) The high - order meta - model based on meta - learning technology models the information of the topological structure, and the neural task scheduler based on the long - short - term memory network breaks through the limitations of the previous fixed sampling strategy, realizes its automated learning, and obtains the optimal parameters of the high - order meta - model and the task scheduler:
[0030]
[0031] The advantages and beneficial effects of the present invention are as follows:
[0032] 1. By designing an integrated model of task sampling and preference prediction, effectively solve the problem of missing interaction information in real recommendation cold - start scenarios, can restore the real interaction relationship by using the implicit relationship between high - order contexts, and promote the preference prediction process. At the same time, the task selection process also contributes to the correctness of preference prediction;
[0033] 2. Utilize the high - order context of nodes in space to provide more interaction information, predict preference information based on the relationships and laws existing in the data itself, which is entirely driven by the data itself without introducing additional prior knowledge;
[0034] 3. By using graph deep learning and meta - learning technologies, the model can be efficiently trained, has strong scalability and the ability to process big data, and can be applied to large - scale recommendation scenarios with less interaction information; Experiments on multiple real - world datasets show that even when most node interaction information is scarce, the proposed method can still achieve high accuracy, demonstrating the robustness and effectiveness of the method;
[0035] 4. The present invention has broad application prospects in the fields of social network analysis, recommendation systems, information retrieval, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1It is the flowchart of the cold start method described in the present invention;
[0037] Figure 2 It is the workflow framework diagram of the embodiment of the present invention; the upper part is the flowchart of the present invention, and the lower part is the structural framework diagram of the present invention. Detailed implementation manners
[0038] The following is by combining Figure 1 and specific embodiments to further illustrate the present invention. The embodiments of the present invention are to enable those skilled in the art to better understand the present invention and do not impose any limitations on the present invention.
[0039] The following takes a preferred embodiment of the present invention to further illustrate the workflow and working principle of the cold start method described in the present invention. As Figure 2 shown in the lower part of, the meta-learning of the cold start method includes a high-order meta-model and a neural task scheduler. Among them, the high-order meta-model includes a feature aggregator and a task predictor; the feature aggregator aggregates structural information from the context of the heterogeneous information network and generates node embeddings using a convolutional neural network; the task predictor is constructed using a fully connected neural network, takes the user node embedding and the evaluated item embedding as inputs, and outputs the predicted value of the evaluation score; constructs a squared loss function, takes the true value and the predicted value of the item evaluation score as inputs, and calculates the loss function; the neural task scheduler takes the loss function and the gradient similarity as inputs respectively to obtain the sampling probability of the current task, where the loss function is a summary of the results of the high-order meta-model, and the gradient similarity is a summary of the process of the high-order meta-model; the iterative promotion relationship between the two helps the two to learn better and faster.
[0040] As Figure 1 shown, a cold start method for a recommendation system based on meta-learning and long short-term memory network includes the following steps:
[0041] Step 1: Use the meta-learning framework to formalize the user and their historical interaction data into a learning task. Assume that the task of a user u is T u =(S u , Q u ). According to the idea of meta-learning, the books interacted with user u are divided into a support set S u and a query set Q u ; the high-order meta-model includes a feature aggregator and a task predictor; the support set is used for training data, and the query set is used for test data.
[0042] Obtain the node types and the corresponding number of nodes, edge types and the number of edges in the DBook dataset from the Internet, where the number of edges represents the number of interactions that occur between users-books and books-authors; Figure 2On the left of the upper part is the interaction diagram of user - book - author, and on the right is the flow chart of this embodiment. First, the neural task scheduler calculates the initial sample weights, and the training task optimizes the high - order meta - model; second, the high - order meta - model initially optimizes the neural task scheduler based on the validation task; finally, the initialized high - order meta - model and neural task scheduler are used for loss function calculation and model optimization. Table 1 shows the specific description of a selected test data:
[0043] Table 1: Test Data
[0044]
[0045] Use the graph convolutional neural network to perform feature aggregation on the high - order context book information of users in the heterogeneous neural network to generate the initial node embedding of the user, where the user node embedding captures the topological relationship between the context book nodes.
[0046] x u = g φ (u, b) = σ(MEAN({Wb + b: b ∈ B}))
[0047] In the formula, b represents the book attributes that interact with the current user, σ is the activation function, φ=(W, b) represents the parameters of the graph convolutional neural network g, and g is named the feature aggregator.
[0048] Step 2: Use a fully - connected neural network to construct a task predictor. Use the generated initial node embedding of the user and the books rated by the user (i.e., the books in the dataset obtained in Step 1), and use the established task predictor to predict the user's rating of the book; then, based on the predicted rating, use the squared loss function to measure the gap between the predicted value and the true value obtained in Step 1. Among them, the Adam optimizer is used to train the neural network parameters during the training process of this step.
[0049]
[0050]
[0051] In the formula, h w represents the predictor with parameters w, MLP represents the fully - connected neural network, x u , x b represent the embeddings of the user and the book respectively, L us represents the loss function of the support set of user u, |B s | represents the number of books in the support set, r u,b , represent the true value and the predicted value of user u's rating of book b respectively. Finally, the feature aggregator and the task predictor are unified into the high - order meta - model together, and its parameters θ=(φ, w).
[0052] Step 3: According to the calculation methods in Step 1 and Step 2, similarly obtain the loss function of the query set.
[0053]
[0054] In the formula, L uq represents the loss function of the query set of user u, and |B q | represents the number of books in the query set.
[0055] Step 4: Construct a neural task scheduler with a long short-term memory network (LSTM, Long Short-Term Memory). Use the loss function of the query set and the gradient similarity between the support set and the query set as inputs to obtain the sampling probability of the current task. Model from two aspects of the results and processes of the high-order meta-model training to fit the relationship between the current task and the new task; among them, the loss function is a summary of the results of the high-order meta-model, and the gradient similarity is a summary of the process of the high-order meta-model; finally, obtain the sampling weight of the current task relative to the new task:
[0056]
[0057]
[0058] In the formula, M s represents the gradient similarity between the support set and the query set loss function, represent the gradients of the support set and the query set respectively, represents the long short-term memory network with parameter , and ω u represents the sampling weight of the current user u. The training process in this step uses the Adam optimizer to train the neural network parameters.
[0059] Step 5: Since the high-order meta-model and the neural task scheduler are interdependent in the meta-training stage, considering that the high-order meta-model minimizes the average loss requires the parameters of the neural task scheduler as the basis, where the parameters of the neural task scheduler are obtained by optimizing the training loss of the high-order meta-model. Therefore, construct a joint optimization process for the neural task scheduler and the high-order meta-model, and transform it into a two-layer optimization process to obtain the optimal parameters of the high-order meta-model and the task scheduler. The specific steps include:
[0060] a) Use the loss function of the high-order meta-model and the corresponding gradient similarity as inputs, and model the training results and training processes of the high-order meta-model respectively to update the parameters of the neural task scheduler.
[0061]
[0062] b) Referring to the sampling weights calculated by the neural task scheduler, sample the top K tasks for the training of the high-order meta-model to update its parameters. This process can ensure that tasks with more information will be preferentially used for training, alleviating the limitations of unified sampling for all tasks to a certain extent:
[0063]
[0064]
[0065] c) The high-order meta-model based on meta-learning technology models the information of the topological structure. The neural task scheduler based on the long short-term memory network breaks through the limitations of the previous fixed sampling strategy, realizes its automated learning, and obtains the optimal parameters of the high-order meta-model and the task scheduler:
[0066]
[0067] Directly apply the trained optimal parameters to new users who have not been trained. In the joint optimization process of the above steps, on the one hand, the neural task scheduler can better select the importance of tasks, and on the other hand, it also helps the high-order meta-model to more accurately predict user preferences.
[0068] In the present invention, the high-order meta-model and the neural task scheduler are not optimized using training tasks, but the high-order meta-model is optimized using training tasks, and the neural task scheduler is optimized using validation tasks. Among them, the loss function of the validation task is the same as that of the training task, and its performance can be regarded as a certain reward. We can obtain the parameters of the neural task scheduler by minimizing the average loss on the validation task , and the best parameters of the high-order meta-model are obtained by optimizing the average loss of the training tasks.
[0069] Through the integrated training of task sampling and preference prediction, the prediction results can be made more accurate. That is, in multiple cold start scenarios, by comparing with the latest various methods, it can be seen that the method of the present invention can obtain high-quality recommendation results.
[0070] Table 2 shows the comparison of the experimental results of the present invention with those of other recommendation methods in the user cold start scenario, where the recommended metrics are the mean absolute error and the root mean square error.
[0071] Table 2: Comparison of mean absolute error and root mean square error
[0072] Method Mean Absolute Error Root Mean Square Error FM 0.702 0.915 NeuMF 0.654 0.805 GC-MC 0.906 0.976 mp2vec 0.666 0.839 HERec 0.651 0.819 DropoutNet 0.831 0.901 MeteEmb 0.678 0.855 MetaHIN 0.625 0.756 The method of the present invention 0.592 0.713
[0073] The embodiments described above are only used to illustrate the technical idea and features of the present invention. The purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The patent scope of the present invention cannot be limited only by these embodiments. That is, any equivalent changes or modifications made in accordance with the spirit disclosed by the present invention still fall within the patent scope of the present invention.
Claims
1. A cold start method for a recommendation system based on meta - learning and long - short - term memory network, including: Step 1: Construct a high - order meta - model, where the high - order meta - model includes a feature aggregator and a task predictor; The feature aggregator uses a heterogeneous neural network to capture the high - order semantic relationships between nodes, supplement the situation where there is less current user interaction information, uses a graph convolutional neural network as an encoder to encode node attributes and network topologies, and generates node embeddings; Use a fully - connected neural network to construct the task predictor, which takes the user node embedding and the node embedding of the evaluated item as inputs and outputs the predicted value of the evaluation score; Step 2: Take the true value and predicted value of the item evaluation score as inputs and calculate the squared loss function; Step 3: Construct a neural task scheduler with a long - short - term memory network, and take the squared loss function of the query set and the gradient similarity between the squared loss functions of the support set and the query set as inputs to obtain the sampling probability of the current task, and finally obtain the sampling weight of the current task relative to the new task: where M s represents the gradient similarity of the support set and the query set loss functions, respectively represent the gradients of the support set and the query set, represents a long short-term memory network with parameters ω, u represents the sampling weight of the current user u; Step 4: Use the gradient descent method to minimize the gap between the true value and the predicted value, construct the joint optimization of the neural task scheduler and the high - order meta - model, and obtain the optimal parameters of the high - order meta - model and the task scheduler for direct application to new user tasks that have not been trained.
2. The cold start method for a recommendation system according to claim 1, characterized in that, Use a loss function to train the high - order meta - model and the long - short - term memory network, and train the high - order meta - model and the neural task scheduler through the loss function and the gradient similarity of the loss function; the training process of the cold start method uses an Adam optimizer to train the neural network parameters.
3. The cold start method for a recommendation system according to claim 1, characterized in that, Step 1 specifically includes the following steps: Sample the user-item interactions to obtain a dataset for the meta-learning task, where the dataset includes a support set and a query set; assume that the task of a user u is T u =(S u , Q u ). According to the idea of meta-learning, divide the items interacted with user u into a support set S u and a query set Q u ; the support set is used for training data, and the query set is used for testing data; Use a graph convolutional neural network to perform feature aggregation on the high - order context item information of users in the heterogeneous neural network to generate the initial node embedding of the user, where the user node embedding captures the topological relationship between context item nodes: x u = g φ (u, b) =σ(MEAN({Wb + b:b∈B})) In the formula, b represents the item attributes interacting with the current user, σ is the activation function, φ=(W, b) represents the parameters of the graph convolutional neural network g, and g is named the feature aggregator; Step 2: Use a fully - connected neural network to construct a task predictor, use the generated initial node embedding of the user and the items evaluated by the user, and use the task predictor to predict the user's score for the item; then based on the predicted score, measure the gap between the predicted value and the true value obtained in Step 1, and represent it with a squared loss function: where h w denotes a predictor with parameter w, MLP denotes a fully connected neural network, and x u , x b represent the embeddings of the user and the item respectively, and L us denotes the loss function of the support set of user u, and |B s | represents the number of items in the support set, and r u,b , represent the true value and the predicted value of user u's rating for item b respectively; Finally, the feature aggregator and the task predictor are unified into a high - order meta - model together, and its parameters θ=(φ, w).
4. The cold start method for a recommendation system according to claim 1, characterized in that, Step 4 specifically includes the following steps: a) Take the loss function of the high - order meta - model and the corresponding gradient similarity as inputs, model the training results and training process of the high - order meta - model respectively, and realize the update of the parameters of the neural task scheduler: b) With the sampling weights calculated by the neural task scheduler as a reference, tasks with higher sampling weights are selected for training the high-order meta-model to update its parameters; this process can ensure that tasks providing more information will be given higher priority for training, alleviating to some extent the limitations of unified sampling for all tasks: c) The high-order meta-model based on meta-learning technology models the information of the topological structure. The neural task scheduler based on the long short-term memory network breaks through the limitations of the previous fixed sampling strategy, realizes its automated learning, and obtains the optimal parameters of the high-order meta-model and the task scheduler:
Citation Information
Patent Citations
Information-enhanced meta-learning method for relieving cold start problem of recommended user
CN113343094A
Device and method for training meta learning network
JP2020144849A