Robust Recommendation System Representation Learning Method Based on Cross-Distillation

By adopting multi-view learning and cross-distillation methods in the recommendation system, combined with variational fitting approximate mutual information, the problem of insufficient robustness of the existing recommendation system representation learning method in noise data processing is solved, and a more efficient and robust representation learning effect is achieved.

CN118154277BActive Publication Date: 2025-06-27DALIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410338921.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-25
Publication Date
2025-06-27
Estimated Expiration
2044-03-25

AI Technical Summary

Technical Problem

The existing recommendation system based on graph structure learning indicates that learning methods have insufficient robustness and difficulty in processing noise data and training target calculation.

Method used

Multi-view learning and cross-distillation are used to combine variational fitting approximate mutual information to improve the robustness and training efficiency of the model.

Benefits of technology

Through multi-view learning and cross-distillation, redundant information between a single view and multiple views is eliminated, improving the robustness and efficiency of recommendation system representation learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118154277B_ABST
    Figure CN118154277B_ABST
Patent Text Reader

Abstract

The present invention provides a method for robust recommendation system representation learning based on cross-distillation. It uses the principle of cross-distillation and a multi-view learning method to learn the representations of users and items in the recommendation system. The graph structure learning based on multi-view learning is introduced into the representation learning of the recommendation system. The method of graph data augmentation is used to augment the original subgraph, generating new views. Then, using the ideas of information bottleneck and cross-distillation, the redundant information within a single view and between views is removed respectively, improving the robustness of the recommendation system representation learning from multiple perspectives. This enables graph structure learning to better overcome the problem that there is a lot of noisy data in the recommendation system dataset in real life, which affects the robustness of representation learning. When training, the variational fitting method is used to approximate the mutual information instead of directly calculating the mutual information, improving the training efficiency of the model while enhancing the robustness of the recommendation system representation learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The data of the present invention belongs to the field of computer big data systems, and in particular relates to a representation learning method of a recommendation system based on the interaction relationship between users and commodities in the e-commerce scenario. Background Art

[0002] With the improvement of people's living efficiency in recent years, people's demand for the accuracy of recommendation systems during online shopping has become increasingly strong, and scholars have long used deep learning technologies in the construction of recommendation systems. In the recommendation scenario, the interaction records between users and commodities, as well as the social network among users, can be abstracted into a graph structure. Therefore, the construction of a recommendation system is inseparable from graph structure learning. In recent years, scholars have conducted a lot of research on graph structure learning. Graph structure learning is to learn the representation of a graph based on the information of the nodes and edges of the graph data, that is, the structure of the graph. Most of the research on the representation learning method of recommendation systems based on graph structure learning focuses on the improvement of the neural network structure in the graph structure learning problem, and optimizes the structure of the graph neural network to obtain better representations of users or commodities. In addition, due to the large amount of noise data in the recommendation system, there are also many studies on the robustness of the model, starting from the perspective of information theory, aiming to learn more robust representations of the graph, so as to obtain more stable results. In order to make the robustness of the recommendation results stronger, we also introduce the methods of multi-view learning and cross-distillation to learn the representations of users and commodities from multiple perspectives, remove the redundant information that will affect the recommendation results, and further improve the robustness of the recommendation system.

[0003] Most of the current representation learning methods for recommendation systems based on graph structure learning obtain better representations of users or items by improving the structure of graph neural networks. For example, on the basis of the basic graph neural network (GNN) that aggregates information of adjacent nodes in a graph, the graph convolutional neural network (GCN) that introduces convolutional methods, the graph attention network (GAT) that introduces attention mechanisms, etc. However, there is a lot of noise in the real-world recommendation system datasets, such as user accidental touches, the influence of active users / hot-selling items on other users or items. To address this problem, many studies have used information-theoretic methods and introduced information bottlenecks to learn more robust representations of users or items. The information bottleneck aims to make the finally learned representation contain as much predictive information as possible while containing as little information in the input data as possible, so that the learned representation contains as little noise data irrelevant to the task as possible, thereby obtaining a more robust representation. This part of the work includes, for example, the graph information bottleneck (GIB) that combines information bottleneck with graph structure learning tasks, the subgraph information bottleneck (SIB) that simplifies the complexity of the graph by extracting key subgraphs, and the VIB-GSL that first generates a more simplified IB graph using variational information bottleneck (VIB) and then uses the IB graph for downstream prediction tasks. However, these methods face difficulties in calculating the training objective and high costs during actual calculations. Therefore, we use the method of variational fitting to approximately approximate the training objective instead of explicitly calculating it, which can greatly improve the training efficiency of the model. In addition, to further improve the robustness of the model, we use multi-view learning and cross-distillation methods to learn the structure of the graph more comprehensively from multiple perspectives, and finally obtain user and item representations that take into account both effectiveness and robustness for recommendation tasks in e-commerce scenarios.

[0004] The key point of the present invention is to use multi-view learning and cross-distillation methods to enhance the robustness of the existing representation learning work of recommendation systems; at the same time, in the training of the model, the method of variational fitting is used to approximate the mutual information instead of explicitly calculating it, improving the efficiency of the representation learning of the recommendation system and laying a solid foundation for specific downstream recommendation tasks, etc.

[0005] What the present invention intends to protect is to introduce the ideas of data augmentation, multi-view learning, information bottleneck, and cross-distillation into the representation learning task of recommendation systems, ensure the effectiveness and robustness of the representations of users and items learned in the recommendation system, and use the method of variational fitting to improve the efficiency of model training. Summary of the Invention

[0006] Aiming at the problem of a large amount of noise data in recommendation systems, the present invention proposes a robust recommendation system representation learning method based on cross-distillation.

[0007] To solve the above problems, the present invention adopts the following specific steps:

[0008] 1. A robust recommendation system representation learning method based on cross distillation, characterized by comprising the following steps:

[0009] S1: Construction of the recommended system graph structure;

[0010] S1-1: Based on the interaction records between users and products and the social network between users, the entire scenario is abstracted into a graph structure. The nodes of the graph are users and products, and the edges of the graph are the relationships between users and products and between users.

[0011] The recommended system graph structure is constructed and recorded as G0. The node feature matrix of the graph is X0. The initial features of users and products are randomly initialized using the standard normal distribution. The adjacency matrix of the graph is A0. Different weights are assigned according to whether there are edges in the graph and the specific relationship type.

[0012] The number of rows of the node feature matrix X0 is the number of nodes in the graph. Each row of the node feature matrix X0 represents the initial node feature of a node in the recommendation system structure graph G0. The adjacency matrix A0 is constructed according to the relationship between the nodes in the recommendation system structure graph G0 and the type of relationship. The adjacency matrix A0 is an n-order square matrix, where n is the number of nodes in the graph, that is, the total number of users and products.

[0013] S1-2: Construct a subgraph of each node in the recommendation system structure graph G0;

[0014] Taking each node in the recommendation system structure graph G0 as the starting point, the subgraph with a hop count of K is used as the input of downstream representation learning, K∈[1,3]. Taking all nodes in the recommendation system structure graph G0 as the center, the corresponding subgraph is obtained by limiting the hop count, and the corresponding features are learned using the graph structure learning method. represents the subgraph generated with the i-th node of the recommendation system structure graph G0 as the center, i∈[1,n], Representing a subgraph The adjacency matrix of Representing a subgraph The node feature matrix of is used to learn the representation of all nodes and generate a subgraph built around the corresponding node, i.e., the original view G1 = (A1, X1);

[0015] S2: graph data enhancement;

[0016] Traverse all nodes in the recommendation system structure graph G0, and perform Perform data augmentation to generate new subgraphs In the following, G1 and G2 are used to represent the original view and the new view generated by data augmentation, respectively;

[0017] S2-1: The original view G1 = (A1, X1). Aggregate the information of adjacent nodes in the original graph using a graph neural network to obtain new node features, i.e.:

[0018] x 1,i = FC(Conv(A1, X1)) i (1)

[0019] Where Conv represents a graph neural network, FC represents a common linear neural network in deep learning, and x 1,i represents the feature of the i-th node in graph G1, X1 represents the node feature matrix of view G1, and A1 represents the adjacency matrix of view G1;

[0020] S2-2: Use a Bernoulli distribution with parameter ω 1,ij to sample the edges connecting node i and node j in the original view G1, i.e., A 1,ij ~ Ber(ω 1,ij ). Connect the two node features and use a neural network to obtain the parameters of the distribution to obtain the parameter ω 1,ij of the Bernoulli distribution. The calculation formula is:

[0021] ω 1,ij = sigmoid(FC(e 1,ij )) (2)

[0022] e 1,ij = x 1,i || x 1,j (3)

[0023] Where ω 1,ij represents the parameter of the Bernoulli distribution, x 1,j represents the feature of the j-th node in G1, || represents the concatenation of two vectors, and sigmoid represents a common activation function in deep learning to ensure that ω 1,ij ∈ [0, 1];

[0024] In actual use, for the convenience of backpropagation during training, the present invention uses the Gumbel reparameterization method to replace the ordinary Bernoulli distribution, i.e.:

[0025]

[0026] Where A 2,ij represents the value of the edge connecting node i and node j in the adjacency matrix of G2, and represent the values sampled from the Gumbel distribution Gumbel(0, 1), and s represents the weight of the Bernoulli distribution in the final adjacency matrix;

[0027] Obtain the subgraph of the original recommended network subgraph structure, that is, the original view G1=(A1, X1), and the newly generated view G2=(A2, X2) after resampling the edges using the Bernoulli distribution, where A2 and X2 represent the adjacency matrix and node feature matrix of G2 respectively. At initialization, X2 = X1;

[0028] S3: Graph structure learning for a single view;

[0029] Traverse each node in the recommended system structure graph G0. For the two views G1 and G2 of the node, the training objective for a single view is defined as follows:

[0030]

[0031] In the formula, y represents the class label of the current user or commodity, z1 and z2 represent the representations of views G1 and G2, P represents the probability distribution of random variables, f represents the graph encoder, θ and represent the parameters of the graph encoder f for the two views, D KL represents the KL divergence between the two distributions;

[0032] S4: Graph structure learning for the robustness between views;

[0033] S4-1: Traverse each node in the recommended system structure graph G0. For the two views G1 and G2 of the node, eliminate the specific redundant information between the two views, and the training objective is as follows:

[0034]

[0035] In the formula, D SKL represents the average value of two symmetric KL divergences, S represents the shared parameters of the graph encoders of the two views, and represent the representations of the two views obtained through mutual learning and cross-distillation between the views;

[0036] S4-2: Use the cross-distillation method to remove the common redundant information between views, and the training objective is as follows:

[0037]

[0038] In the formula, D KL represents the KL divergence;

[0039] S5: Training of the overall model;

[0040] S5-1: For the redundant information between a single view and multiple views, obtain the training objective of the model as follows:

[0041]

[0042] In the formula, CE represents the cross-entropy loss function commonly used in deep learning, y and represent the true category of the user or item and the predicted category output by the model respectively, and β represents the hyperparameter;

[0043] S5-2: Calculate the representation z of the user or item in the recommendation system as:

[0044]

[0045] In the formula, MLP represents the multi-layer perceptron commonly seen in deep learning, and || represents the concatenation of two vectors.

[0046] The relationship types between users and items include viewing, collecting, adding to cart, and purchasing, and the relationship types between users include viewing, following, chatting, and being friends.

[0047] The present invention can achieve the following beneficial effects: Aiming at the problems of poor robustness and difficult model training in the existing recommendation system representation learning work, we propose a method to learn the representations of users and items in the recommendation system using the principle of cross-distillation and the method of multi-view learning. We introduce graph structure learning based on multi-view learning into the representation learning of the recommendation system. First, aiming at the problem of scarce multi-view data in the real-world recommendation system representation learning task, we use the method of graph data augmentation to augment the original subgraph and generate new views. Next, we use the ideas of information bottleneck and cross-distillation to eliminate the redundant information within a single view and between views respectively, improving the robustness of the recommendation system representation learning from multiple perspectives, enabling the graph structure learning to better overcome the problem that there are many noisy data in the real-world recommendation system dataset, which affects the robustness of the representation learning. In addition, when training, we use the method of variational fitting to approximate the mutual information instead of directly calculating the mutual information, improving the training efficiency of the model while enhancing the robustness of the recommendation system representation learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is the original graph structure of the recommendation system of the present invention;

[0049] Figure 2 is the subgraph generated with user 2 as the center in the present invention;

[0050] Figure 3 is the architecture diagram of the robust recommendation system representation learning model based on cross-distillation in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] The product consists of five steps. Step 1: Construction of the recommendation system graph structure. In the e-commerce scenario, first, the entire scenario needs to be abstracted into a graph structure based on the interaction records between users and products and the social network relationships between users. Specifically, the nodes of the graph are individual users and products, and the edges of the graph are the relationships between users and products and the relationships between users. The types of relationships between users and products include viewing, favoriting, adding to the shopping cart, and purchasing. The types of relationships between users include viewing, following, chatting, and being friends. If the constructed recommendation system graph structure is denoted as G0, then the node feature matrix X0 of the graph, that is, the initial features of users and products, is randomly initialized using the standard normal distribution. For the adjacency matrix A0 of the graph, different weights are assigned according to whether there is an edge in the graph and the specific type of relationship. Specifically, in the relationships between users and products, the weights of viewing, favoriting, adding to the shopping cart, and purchasing are 1 to 4 respectively. In the relationships between users, the weights of viewing, following, chatting, and being friends are 1 to 4 respectively.

[0052] For example, for Figure 1 the recommendation system graph structure G0 in

[0053]

[0054]

[0055] where n represents the number of nodes in the graph, that is, the total number of users and products in the recommendation system. In Figure 1 the G0 represented, n = 9. dim represents the custom feature dimension. Each row of G0 represents the initial node features of a node in the graph G0. The adjacency matrix A0 is defined as follows according to the relationships between the nodes in the graph G0 and the types of relationships:

[0056]

[0057] The obtained adjacency matrix is an n-order square matrix, that is, the total number of users and products. In Figure 1 the G0 represented, n = 9. We define the elements on the diagonal of the adjacency matrix as 1. Next, a subgraph is constructed for each node in the graph G0 to be used as the input for the representation learning of each subsequent user or product node. Specifically, for each node in the graph G0, the subgraph with a restricted hop count of K starting from it is used as the input for the downstream representation learning, where K ∈ [1, 3]. A larger hop count means that the model can capture more complex correlations between users and products. For example, on the graph G0, with user 2 as the center and a restricted hop count of 2, that is, in the graph G0, starting from the node represented by user 2, the nodes and the edges passed through by walking at most 2 steps on the graph form a new set of nodes and a set of edges, thus generating a subgraph as Figure 2As shown, denoted as where the subscript 1 represents the initial view, and the superscript 2 represents the second node in graph G0. Similarly, after restricting the hop count and obtaining the corresponding subgraphs centered around all the nodes in graph G0 in sequence, we can use the graph structure learning method to learn their corresponding features, which can then be used in downstream recommendation tasks. We use to represent the subgraph generated centered around the i-th node of graph X0, where i ∈ [1, n], and represents the adjacency matrix of the subgraph , and represents the node feature matrix of the subgraph . Since we will perform representation learning on all the nodes next, for convenience of representation in the following text, we uniformly use G1 = (A1, X1) to represent the subgraph generated centered around the corresponding node.

[0058] Step 2: Graph data augmentation. After completing the graph structure modeling for the recommendation scenario, we can use the graph structure learning method to learn the representations of users and items, thus laying a foundation for downstream specific recommendation tasks. In the field of representation learning, multi-view learning can not only improve the prediction performance of the model but also enhance the robustness of the representations of various objects in the recommendation system. Therefore, in order to use the multi-view learning method, the present invention first uses the graph data augmentation method to generate new views, so as to learn the graph structure from different perspectives respectively and enhance the robustness of the representation learning. Here, we traverse all the nodes in the recommendation system structure graph G0 and perform data augmentation on their respective subgraphs to generate new views . For convenience of representation, we will omit the superscript i next and use G1 and G2 to represent the two views respectively. The present invention uses the method of constructing a new graph adjacency matrix to achieve the purpose of graph data augmentation. For the original view G1 = (A1, X1), we perform the following operations:

[0059] First, use the graph convolution operation to aggregate the information of adjacent nodes in the original graph to obtain new node features, that is:

[0060] x 1,i = FC(Conv(A1, X1)) i (1)

[0061] where Conv represents the graph neural network, FC represents the common linear neural network in deep learning, and x 1,iRepresent the features of the \(i\)-th node in graph \(G1\). Among them, for graph neural network references [Scarselli F, Gori M, Tsoi A C, et al. The Graph Neural Network Model [J]. IEEE Transactions on Neural Networks, 2009, 20(1): 61. DOI: 10.1109 / TNN.2008.2005605. ALEMI Alexander A, FISCHER I, DILLON Joshua V, et al. Deep Variational Information Bottleneck [J]. Learning, Learning, 2016.]. Secondly, use a Bernoulli distribution with parameter \(\omega\) 1,ij to sample the edges in the original graph \(G1\), that is, \(A\) 1,ij \(\sim Ber(\omega\) 1,ij ). Specifically, for the Bernoulli distribution for each edge in graph \(G1\), assuming it connects the \(i\)-th node and the \(j\)-th node in the graph, the probability of retaining this edge in the new view is \(\omega\) 1,ij , and the probability of deleting this edge is \(1 - \omega\) 1,ij . That is, a new graph adjacency matrix is generated based on the original graph adjacency matrix by sampling with the Bernoulli distribution, so as to achieve the purpose of graph data augmentation. Among them, the Bernoulli distribution is also called the binomial distribution, which can be referred to [Sheng Zhou, Xie Shiqian, Pan Chengyi. Probability Theory and Mathematical Statistics (Fourth Edition) [M]. Beijing: Higher Education Press, 2008.6: pp. 33 - 35]. In order to obtain the parameter \(\omega\) 1,ij of the Bernoulli distribution, the present invention connects the features of two nodes and uses a neural network to obtain the parameters of the distribution, that is:

[0062] \(\omega\) 1,ij = sigmoid(FC(e 1,ij )) (2)

[0063] where, \(e\) 1,ij = \(x\) 1,i || \(x\) 1,j , \(c\) 1,i represents the features of the \(i\)-th node in graph \(G1\), \(x\) 1,j represents the features of the \(j\)-th node in graph \(G1\), and || represents the concatenation of two vectors. sigmoid is a common activation function in deep learning, which can ensure that \(\omega\) 1,ij ​​​​​​​∈[0,1]. In actual use, for the convenience of backpropagation during the training process, the present invention uses the Gumbel reparameterization method introduced in the literature [JANG E, GUS, POOLE B. Categorical Reparameterization with Gumbel-Softmax[J]. International Conference on Learning Representations, International Conference on Learning Representations, 2016.] to replace the ordinary Bernoulli distribution, that is:

[0064]

[0065] where ω 1,ij represents the parameter of the Bernoulli distribution, a2 represents the adjacency matrix of G2, i and j respectively represent the nodes i and j of G2, and A 2,ij represents the value of the edge connecting nodes i and j in the adjacency matrix of G2, that is and are both values sampled from the Gumbel distribution Gumbel(0,1), and s is used to adjust the weight of the Bernoulli distribution in the final adjacency matrix. Finally, the original view G1=(A1,X1) is obtained, where A1 and X1 respectively represent the adjacency matrix and node feature matrix of G1, and the newly generated view G2=(A2,X2) after resampling the edges using the Bernoulli distribution, where A2 and X2 respectively represent the adjacency matrix and node feature matrix of G2. At initialization, X2 = X1. Next, the present invention uses multi-view learning and cross-distillation methods to remove redundant information irrelevant to the downstream task within a single view and between multiple views, so as to learn effective and more robust representations of users and commodities and improve the performance of the downstream recommendation task.

[0066] Step 3: Graph structure learning for a single view. In the representation learning problem, the representation z of the input data x obtained should at least satisfy sufficiency: the representation should contain all the prediction information required for downstream prediction tasks. The representation z of the input data x is sufficient for predicting the label y if and only if I(x; y|z) = 0, where I represents mutual information, which can be regarded as the amount of information about another random variable contained in a random variable. For example, I(x; y) represents the amount of information related to the prediction information y contained in the input data x, and I(x; y|z) is the conditional mutual information, indicating the mutual information between x and y given z. However, even when the sufficiency constraint is satisfied, the representation z may still contain redundant information unrelated to the downstream task, and this redundant information may make the learned representation extremely unstable and less robust. To improve the robustness of the learned representation z, the present invention introduces the idea of information bottleneck to eliminate redundant information, and the loss function of the information bottleneck is as follows:

[0067] L IB = I(z; x) - βI(z; y) (12)

[0068] where x is the input information, z is the representation of x, y is the prediction label, and β is the Lagrange multiplier used to balance the amount of information of x contained in the learned representation z. The ideal representation z should contain all the prediction information required for downstream prediction, that is, maximize I(z; y), while containing as little information of the input x as possible, that is, minimize I(z; x), so as to contain as little redundant information in the input as possible to improve the robustness of the model. However, according to the literature [Federici, M.; Dutta, A.; Forré, P.; Kushman, N.; Akata, Z. Learning Robust Representations via Multi-View Information Bottleneck 2020, [arXiv:cs.LG / 2002.07017]], since β needs to be optimized by balancing high compression and high mutual information, it is impossible to achieve these two optimization goals simultaneously in the above formula. In addition, the calculation of mutual information, especially in the high-dimensional case, is very difficult. Therefore, the present invention does not directly use formula (6) as the training objective, but uses the variational fitting method to approximate the mutual information and obtain an analytical solution of the mutual information instead of explicitly calculating it. In this step, we traverse each node in the recommendation system structure graph G0. For the two views G1 and G2 of the node, we first eliminate the redundant information unrelated to the prediction information within a single view: specifically, for the first view G1, according to the chain rule of mutual information, the mutual information between the encoding f θ (G1) of graph G1 and the representation z1 of graph G1 can be decomposed into the following two terms:

[0069] I(f θ (G1); z1) = I(z1; y) + I(f θ (G1); z1|y) (13)

[0070] where f is the graph encoder and θ is the parameter of the encoder f. In the above formula, I(z1; y) represents all the predictive information contained in z1, while I(f θ (G1); z1|y) represents the information irrelevant to the downstream prediction task, that is, redundant information. According to the idea of the information bottleneck, in order to make the representation z1 contain all the predictive information while containing as little redundant information as possible, it is necessary to maximize I(z1; y) while minimizing I(f θ (G1); z1|y). Maximizing I(z1; y) means making the representation z1 contain as much predictive information as possible, and this part can be completed by minimizing the prediction loss of the downstream task. For I(f θ (G1); z1|y), obviously, according to the definition of redundant information in the prediction task, minimizing I(f θ (G1); z1|y) is equivalent to minimizing the difference between the following mutual information: I(f θ (G1); y) - I(z1; y). According to the symmetry of mutual information, that is, I(x; y) = I(y; x) and the definition of mutual information, that is, I(x; y) = H(x) - H(x|y), it can be known that minimizing the difference between the above mutual information is equivalent to minimizing the difference between the following conditional entropies: H(y|z) - H(y|f θ (G1)), that is:

[0071]

[0072] According to the definition of conditional entropy under continuous conditions, the following upper bound can be obtained:

[0073]

[0074] where D KL represents the KL divergence between two distributions. The smaller the KL divergence, the more similar the two distributions are. Obviously, minimizing the KL divergence between the distribution P(y|f θ (G1)) and the distribution P(y|z1), that is, making P(y|z1) fit P(y|f θ (G1)), can minimize both terms of the upper bound (9). Therefore, the KL divergence between the distribution P(y|f θ (G1)) and the distribution P(y|z1) being 0, that is, these two distributions being exactly the same, is a sufficient condition for the difference between the above conditional entropies to be 0, that is:

[0075]

[0076] Therefore, the purpose of removing the redundant information unrelated to the downstream task in the representation can be achieved by minimizing the KL divergence between the two distributions, that is, the training objective for the first view is defined as follows:

[0077] Lz1 = D KL (P(y|f θ (G1))||P(y|z1))

[0078] The above process is the same for the second view G2. Therefore, the training objective for a single view is defined as follows:

[0079]

[0080] where θ is the parameter of the graph encoder f of the first view, is the parameter of the graph encoder f of the second view. The representation of the user or commodity learned in this way eliminates the noise information unrelated to the downstream recommendation task within a single view, thereby improving the robustness of the model. And through this method, the mutual information in formula (6) can be obtained without directly calculating it, but by using variational fitting to approximate the mutual information, finding the variational upper bound of the mutual information, and obtaining an analytical solution of the mutual information during the calculation of the mutual information, thereby greatly improving the optimization efficiency of the model.

[0081] Step 4: Robust graph structure learning among multiple views. In multi-view data, each view describes the same object from a different perspective, thereby improving the robustness of the model. To effectively utilize the rich information of multi-view data, the present invention introduces the idea of cross-distillation to eliminate the redundant information in the multi-view scenario. In this step, we traverse each user or commodity node in the recommendation system structure graph G0, and perform the following operations on the two views G1 and G2 of the node: First, the specific redundant information between the two views needs to be eliminated. Specifically, it is known that G1 and G2 are respectively two views of the original input subgraph. Assume that the input of the first view G1 is f S (G1) and its representation is The input of the second view is f S (G2) and its representation is where, f S is the graph encoder, and S is the shared parameter of the graph encoders of the two views. Here, in order to make the two views in the multi-view learning learn representations with consistent modalities, we use a graph encoder with parameter sharing and an information bottleneck. Assume that the inputs of the two views f S (G1) and f S(G2) both contain all the information about the predicted label y. At the same time, according to the literature [Federici, M.; Dutta, A.; Forré, P.; Kushman, N.; Akata, Z. Learning Robust Representations via Multi-View Information Bottleneck 2020, [arXiv:cs.LG / 2002.07017]], if the representation of the first view is sufficient for the input of the second view i.e., then i.e., also contains all the prediction information of y, and the same is true for In addition, if and only capture the common information of f S (G1) and f S (G2), while ignoring the inconsistent information between views, that is, the specific redundant information between views, then it can improve the robustness of graph structure learning. According to the chain rule of mutual information, decomposing the mutual information between the input f S (G1) of the first view and its representation gives:

[0082]

[0083] In the above formula, represents the mutual information between the input f S (G2) of the second view and the representation of the first view, that is, the mutual information shared between the two views, while represents the information in the representation of the first view that is only related to the input f S (G1) of the first view and has nothing to do with the second view, that is, the view-specific information. Since both views contain all the information required for the downstream task, the view-specific information in a single view is redundant information for the prediction task and needs to be eliminated. To eliminate the view-specific redundant information, it is necessary to minimize It can be derived from the definition of conditional mutual information that the upper bound of is as follows:

[0084]

[0085] Therefore, minimizing in the training objective can be transformed into minimizing the following KL divergence:

[0086]

[0087] Similarly for z2, the following training objective can be obtained:

[0088]

[0089] Combining formula (15) and formula (16), the final training objective can be obtained as follows:

[0090]

[0091] Among them, D SKL represents the average of two symmetric KL divergences, f represents the graph encoder, S represents the shared parameters of the graph encoders of the two views, and respectively represent the representations of the two views obtained through mutual learning and cross-distillation between the views in this step. Minimizing the training objective L2 can eliminate the specific redundant information between the two views, thereby improving the robustness of the model. In addition to the specific redundant information between the views, there may still be information irrelevant to the downstream task in the information shared between the views. To further improve the robustness of the model, the present invention also uses the method of cross-distillation to remove this part of redundant information.

[0092] Specifically, for the mutual information between the input f S (G1) of the first view and the representation z2 of the second view As can be seen from the previous section, this mutual information represents the information of the first view input contained in the representation of the second view, that is, the shared information between the two views. First, decompose this mutual information to obtain the following equation:

[0093]

[0094] Obviously, the second term in the decomposition result is the prediction information contained in the representation z1 of the first view, which is also what we hope to maximize. And represents the redundant information shared between the two views and irrelevant to the downstream task. Therefore, it is also necessary to minimize Similar to the derivation process of the upper bound (9), minimizing can be transformed into the following training objective:

[0095]

[0096] Similarly, for the representation of the second view, it is necessary to minimize the following training objective:

[0097]

[0098] Combining the training objectives (19) and (20), we can obtain the following training objective:

[0099]

[0100] where D KL represents the KL divergence, y represents the predicted label of the current user or item, f represents the graph encoder, S represents the shared parameters of the graph encoders of the two views, and respectively represent the representations of the two views obtained through mutual learning and cross-distillation between the views in this step.

[0101] This part of the training objective uses the method of cross-distillation to eliminate the redundant information shared between the views that is irrelevant to the downstream task, further enhancing the robustness of the model.

[0102] Step 5: The framework of the entire model is as Figure 3 shown. The left half of the figure shows the graph data augmentation module, which realizes graph data augmentation by sampling the edges in the original subgraph. The right half of the figure represents the graph structure learning module, where GSL represents graph structure learning, E represents the graph encoder, that is, the graph neural network, IB represents the information bottleneck, and the internal structure is the multi-layer perceptron (MLP) commonly used in deep learning. The subscripts 1 and 2 represent the encoders and information bottlenecks of each view. The network structure within each view can minimize the loss within each view, so that the representation obtained for each view contains as little redundant information as possible; the subscript S represents the encoder and information bottleneck calculated using the shared parameters of the two views. This part of the network minimizes the loss between the two views, making the two views as consistent as possible and reducing redundant information. The final representation z of the user or item is obtained by encoding with the MLP: Combining the redundant information between a single view and multiple views described above, the final training objective of the model is as follows:

[0103]

[0104] The training objective consists of two parts: The first term is the prediction loss, which is the prediction loss for the user or item category classification task here. Among them, CE is the cross-entropy loss function commonly used in deep learning, and y and respectively represent the true category of the user or item and the predicted category output by the model. This term enables the representation z to contain as much prediction information as possible, thereby improving the effectiveness of the representation in the downstream recommendation task. The second term considers the differences between a single view and multiple views, eliminating the redundant information in the representation, thereby improving the robustness of the representation in the recommendation system. β is a hyperparameter used to balance the relationship between the effectiveness and robustness of the model. The final representation z of the user or item in the recommendation system is:

[0105]

[0106] Among them, MLP represents the multi-layer perceptron commonly used in deep learning, || represents the concatenation of two vectors, z1 and z2 are the representations of the two views respectively, and and represent the representations of the two views obtained through mutual learning and cross-distillation between the views respectively. Next, the representation z can be used in downstream recommendation tasks to improve the effectiveness and robustness of the recommendation results.

[0107] The above is the preferred embodiment of the present invention, but the present invention is not limited to the above embodiments. It should be noted that other improvements and changes directly derived or associated by those skilled in the art of this technology without departing from the spirit and concept of this application should be considered to be included in the protection scope of the present invention.

Claims

1. A robust recommendation system representation learning method based on cross distillation, characterized in that: The steps include: S1: Construction of the recommended system graph structure; S1-1: Based on the interaction records between users and products and the social network between users, the entire scenario is abstracted into a graph structure. The nodes of the graph are users and products, and the edges of the graph are the relationships between users and products and between users. The recommended system graph structure is constructed and recorded as G0. The node feature matrix of the graph is X0. The initial features of users and products are randomly initialized using the standard normal distribution. The adjacency matrix of the graph is A0. Different weights are assigned according to whether there are edges in the graph and the specific relationship type. The number of rows of the node feature matrix X0 is the number of nodes in the graph. Each row of the node feature matrix X0 represents the initial node feature of a node in the recommendation system structure graph G0. The adjacency matrix A0 is constructed according to the relationship between the nodes in the recommendation system structure graph G0 and the type of relationship. The adjacency matrix A0 is an n-order square matrix, where n is the number of nodes in the graph, that is, the total number of users and products. S1-2: Construct a subgraph of each node in the recommendation system structure graph G0; Taking each node in the recommendation system structure graph G0 as the starting point, a subgraph with a hop count of K is used as the input for downstream representation learning, K∈[1,3]. All nodes in the recommendation system structure graph G0 are taken as the center, and the hop count is limited to obtain the corresponding subgraph. The corresponding features are learned using the graph structure learning method. use represents the subgraph generated with the i-th node of the recommendation system structure graph G0 as the center, i∈[1,n], Representing a subgraph The adjacency matrix of Representing a subgraph The node feature matrix of is used to learn the representation of all nodes and generate a subgraph built around the corresponding node, i.e., the original view G1 = (A1, X1); S2: graph data enhancement; Traverse all nodes in the recommendation system structure graph G0, and perform Perform data augmentation to generate new subgraphs In the following, G1 and G2 are used to represent the original view and the new view generated by data augmentation, respectively; S2-1: The original view G1 = (A1, X1), uses the graph neural network to aggregate the information of adjacent nodes in the original graph to obtain new node features, namely: x 1,i =FC(Conv(A1,X1)) i (1) In the formula, Conv represents graph neural network, FC represents linear neural network commonly used in deep learning, and x 1,i represents the feature of the i-th node in the graph G1, X1 represents the node feature matrix of view G1, and A1 represents the adjacency matrix of view G1; S2-2: Use parameter ω 1,ij The Bernoulli distribution is used to sample the edges connecting nodes i and j in the original view G1, that is, A 1,ij ~Ber(ω 1,ij ), connect the two node features, use the neural network to get the distribution parameters, and obtain the Bernoulli distribution parameters ω 1,ij , the calculation formula is: ω 1,ij =sigmoid(FC(e 1,ij )) (2) e 1,ij =x 1,i ||x 1,j (3) In the formula, ω 1,ij represents the parameters of the Bernoulli distribution, x 1,j represents the feature of the j-th node in G1, || represents the concatenation of two vectors, sigmoid represents a common activation function in deep learning, and ensures ω 1,ij ∈[0,1]; In actual use, in order to facilitate back propagation during training, the Gumbel reparameterization method is used to replace the ordinary Bernoulli distribution, that is: In the formula, A 2,ij represents the value of the edge connecting node i and node j in the adjacency matrix of G2, and represents the value sampled from the Gumbel distribution Gumbel(0,1), and s represents the weight of the Bernoulli distribution in the final adjacency matrix; Obtain the subgraph of the original recommendation network subgraph structure, that is, the original view G1 = (A1, X1), and the newly generated view G2 = (A2, X2) after edge resampling using Bernoulli distribution, where A2 and X2 represent the adjacency matrix and node feature matrix of G2 respectively, and when initialized, X2 = X1; S3: Graph structure learning for a single view; Traverse each node in the recommendation system structure graph G0. For the two views G1 and G2 of the node, the training objective of a single view is defined as follows: In the formula, y represents the category label of the current user or product, z1 and z2 represent the representation of views G1 and G2, P represents the probability distribution of random variables, f represents the graph encoder, θ and Denote the parameters of the image encoder f for two views, D KL Represents the KL divergence between two distributions; S4: Graph structure learning for robustness between views; S4-1: Traverse each node in the recommendation system structure graph G0, and for the two views G1 and G2 of the node, eliminate the specific redundant information between the two views. The training objectives are as follows: Where D SKL represents the average of two symmetric KL divergences, S represents the shared parameters of the graph encoders of the two views, and Representation of two views obtained by mutual learning and cross-distillation between representation views; S4-2: Use the cross distillation method to remove the common redundant information between views. The training objectives are as follows: Where D KL represents KL divergence; S5: training of the overall model; S5-1: For the redundant information between a single view and multiple views, the training objectives of the model are as follows: In the formula, CE represents the cross entropy loss function commonly used in deep learning, y and They represent the true category of the user or product and the predicted category of the model output, and β represents a hyperparameter; S5-2: Calculate the representation z of users or products in the recommendation system as: In the formula, MLP represents the multi-layer perceptron commonly used in deep learning, and || represents the concatenation of two vectors.

2. The cross-distillation-based robust recommendation system representation learning method according to claim 1, characterized in that: The types of relationships between users and products include view, favorite, add to cart, and purchase; the types of relationships between users include view, follow, chat, and friend.

Citation Information

Patent Citations

  • Recommendation method and recommendation device for learning user preferences by utilizing multiple views

    CN117473167A

  • Cross-lingual unsupervised classification with multi-view transfer learning

    EP3879429A2