An out-of-distribution generalization deep recommendation method based on double test-time training
By adding self-supervised methods of self-distillation and comparison tasks during the test-time training phase, and updating the recommendation model, the problem of external distribution generalization caused by changes in user and item characteristics is solved, and the adaptability and accuracy of the recommendation system are improved.
Patent Information
- Application Number
- CN202410680721.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-29
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-05-29
AI Technical Summary
The existing recommendation algorithm assumes that the distribution of users and items remains unchanged during the training and testing stages, violating the independent and homogeneous distribution assumptions, resulting in the inability to effectively handle changes in user characteristics and item characteristics in the real world.
Using a dual-test training method, the pre-trained model is updated through self-supervised tasks, including self-distillation tasks and comparison tasks. The self-distillation tasks minimize the distance between user interest centers, and the comparison tasks capture similar preferences between users to form an update model.
It enhances the robustness of the model for distribution changes, can better understand user preferences and adapt to data offsets, and improves the generalization ability of the recommendation system.
Smart Images

Figure CN118520170B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data recommendation, and in particular to an out-of-distribution generalization deep recommendation method based on training during double testing. Background Art
[0002] Recommendation systems play a crucial role in alleviating information overload on social media platforms by providing personalized information filtering. In recent years, a large number of recommendation algorithms have emerged, including collaborative filtering, graph-based recommendations, retraining, etc.
[0003] Although guaranteed performance has been achieved, most existing recommendation algorithms assume that the user and item distributions remain unchanged between the training phase and the testing phase, following the independent and identically distributed (I.I.D.) assumption. However, in the real world, these distributions often change, which violates this assumption. This shift is mainly manifested in two aspects: the shift of user characteristics and the shift of item characteristics. The shift of user characteristics specifically refers to the change of user preferences over time as user characteristics evolve. For example, as a boy grows older, his interests may change from comic books to sports, such as basketball. On the other hand, the shift of item characteristics is related to the change of preference relationships caused by the replacement or update of specific types of items. For example, when items are renewed or updated, they may have new characteristics, leading to different preferences and recommendation relationships among users.
[0004] The out-of-distribution (OOD) generalization problem has been widely studied in various fields, including image classification and detection. However, it has received limited attention in the context of recommendation systems. To address this issue, disentangled recommendation methods aim to learn a factored representation of user preferences, which can enhance the model's robustness to distribution changes. Causality-based methods use causal learning to solve the OOD problem. However, these methods often require intervention during the training process, and they are less practical when only a pre-trained model is available. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide an out-of-distribution generalization deep recommendation method based on training during double testing.
[0006] The technical solution adopted by the present invention to solve the above technical problems is: an out-of-distribution generalization depth recommendation method based on dual testing-time training, including the steps of: 1) in the pre-training stage, training a recommendation system model to form a pre-trained model; 2) in the testing-time training stage, updating the pre-trained model through a self-supervised task to form an updated model; 3) in the inference stage, using the updated model for inference to obtain a recommendation result; the self-supervised task in step 2) includes a self-distillation task, and the distance between user interest centers with the same preference is minimized through the self-distillation task.
[0007] The method for updating the pre-trained model through the self-distillation task is as follows:
[0008] 211) First, extract user features, item features, and historical interactions through the pre-trained model;
[0009] E u = f θ (X u , D);
[0010] E i = f θ (X i , D);
[0011] where X u and X i represent user features and item features respectively, D is the interaction history, E u and E i are embedding features, and f θ is the pre-trained model;
[0012] 212) Perform fusion: where is the concatenation operation, and g() represents a multi-layer perceptron;
[0013] 213) Obtain the clustering distribution result: where represents the K-means algorithm;
[0014] 214) Calculate the target distribution P:
[0015] where S() represents the sharpening function and T is the temperature parameter;
[0016] The objective function of the self-distillation task:
[0017] where q i represents the user preference center obtained by clustering, and p i represents the sharpened user preference distribution.
[0018] Further, the self-supervised task in step 2) further includes a contrast task: the contrast task is to capture the correlation between users with similar preferences. The method for updating the pre-trained model through the contrast task is as follows:
[0019] 221) Select multiple high-confidence samples based on the clustering distribution results in the self-distillation task to form a high-confidence sample set, perform clustering operations to obtain corresponding pseudo-labels, and divide them into multiple non-overlapping clusters R i , and regard the samples in the same cluster as positive samples;
[0020] 222) Calculate the center of the high-confidence samples: C i = avg(R i ), i = 1, 2,..., K, where avg() represents the average function; regard different high-confidence centers as negative samples;
[0021] 223) The total loss function of the contrast task:
[0022] Loss of positive sample pairs:
[0023] Loss of negative sample pairs:
[0024] where N represents the number of samples, K represents the number of clustering centers, i represents the index of the feature, n and m represent variables, R im and R in are different feature vectors in the same clustering center, and C i and C j represent different preference centers respectively.
[0025] Further, the total loss function of the self-supervised task:
[0026] where α represents the balance parameter.
[0027] The beneficial effects of the present invention are as follows: The present invention proposes an out-of-distribution generalization deep recommendation method based on dual test-time training. This method adds a self-distillation task and a contrast task during the test-time training stage to effectively update the recommendation model, thereby allowing the model to adaptively shift the data to better understand the invariant preferences between users and the variable user features or item features with shifted data. Brief Description of the Drawings
[0028] Figure 1 is the framework schematic diagram of the out-of-distribution generalization deep recommendation method of the present invention;
[0029] Figure 2 is the schematic diagram of the self-distillation task in the recommendation method of the present invention;
[0030] Figure 3 This is the schematic diagram of the contrast task in the recommended method of the present invention. Detailed implementation manners
[0031] Now, the present invention will be further described in detail with reference to the accompanying drawings and preferred embodiments. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0032] A distribution - out - of - domain generalization depth recommendation method based on double - testing training includes the steps: 1) In the pre - training stage, training a recommendation system model to form a pre - trained model; 2) In the testing - time training stage, updating the pre - trained model through self - supervised tasks to form an updated model; 3) In the inference stage, using the updated model to perform inference to obtain a recommendation result. The self - supervised tasks in step 2) include a self - distillation task and a contrast task. The self - distillation task is to minimize the distance between user interest centers with the same preference; the contrast task is to capture the correlation between users with similar preferences. As Figure 1 shown, in the pre - training stage, skirts are recommended to girls, watches are recommended to boys, and books are recommended to boys. After the testing - time training stage and updating the model, in the inference stage, houses are recommended to boys, baby bottles are recommended to pregnant girls, and basketballs are recommended to boys.
[0033] The self - supervised tasks in step 2) include a self - distillation task. The self - distillation task aims to minimize the distance between user interest centers with the same preference, thereby improving the consistency of user interest representation in the latent space. By encouraging user interest centers to cluster more closely, the model becomes better at capturing the latent patterns in user preferences.
[0034] Figure 2 This is the schematic diagram of the self - distillation task in the recommended method of the present invention. As Figure 2 shown, the method for updating the pre - trained model through the self - distillation task is as follows:
[0035] 211) First, extract user features, item features, and historical interactions through the pre - trained model;
[0036] E u = f θ (X u , D);
[0037] E i = f θ (X i , D);
[0038] where X u and X irespectively represent user features and item features, D is the interaction history, and E u and E i are the embedding features, and f θ is the pre-trained model;
[0039] In a recommendation system, the interaction history and user features are indicators of user preferences. The interaction history, such as a user's purchase behavior, can reflect their preferences for specific types of items. Similarly, the items with which a user interacts can indicate their preferences.
[0040] are fused: where is the concatenation operation, and g() represents a multi-layer perceptron (MLP).
[0041] 213) Obtain the clustering distribution result: where represents the K-means algorithm; the variable q is the probability of assigning a sample to a user preference, indicating a soft assignment. In addition, the variables q corresponding to all samples form a matrix Q.
[0042] 214) Calculate the target distribution P:
[0043] where S() represents the sharpening function and T is the temperature parameter;
[0044] The objective function of the self-distillation task:
[0045] q i represents the user preference center obtained by clustering, and p i represents the user preference distribution after sharpening; here, through this loss function, the target distribution P and p i are pulled closer to each other to achieve self-distillation.
[0046] After obtaining the clustering distribution result, the present invention applies a sharpening function to reduce the entropy of the clustering distribution result and uses a lower entropy assignment to optimize the network. When the temperature T approaches zero, the output of S(q, T) tends to approximate a Dirac distribution (a one-hot encoding). A lower temperature helps to produce a lower entropy distribution. Minimizing the KL divergence loss between the target distribution and the clustering distribution result, the target distribution can help the model learn more and better embeddings because the target distribution has low entropy, and this process is regarded as a self-distillation mechanism. Since P is derived from Q; in addition, the distribution of P sequentially supervises the update of the distribution of Q.
[0047] By clustering the fused embeddings of users and items, the present invention obtains interest centers representing different types of user preferences, where each type of user shows a preference for a specific type of item. The self-distillation task of the present invention enhances the saliency of these interest centers, allowing for a better understanding and modeling of user preferences.
[0048] Furthermore, the self-supervised task in step 2) also includes a contrastive task. Contrastive learning is a self-supervised paradigm that has achieved good results in various tasks. However, inappropriate data augmentation can easily lead to semantic drift, resulting in low-quality positive samples. In addition, existing negative sample construction strategies directly regard all other non-positive samples as negative samples, which easily brings false negative samples.
[0049] Figure 3 is the schematic diagram of the contrastive task in the recommendation method of the present invention, as Figure 3 shown, the method for updating the pre-trained model through the contrastive task is as follows:
[0050] 221) Select multiple high-confidence samples from the clustering distribution results in the self-distillation task to form a high-confidence sample set, perform clustering operations to obtain corresponding pseudo-labels, and divide them into multiple non-overlapping clusters R i , and regard the samples in the same cluster as positive samples;
[0051] On the basis of the high-confidence sample set, construct positive sample pairs, and regard the samples in the same cluster as positive samples. That is, the same preference relationship (items of similar users) is considered a positive sample.
[0052] 222) Calculate the center of the high-confidence samples: C i = avg(R i ), i = 1, 2,..., K, where avg() represents the average function; regard different high-confidence centers as negative samples;
[0053] The present invention regards different high-confidence centers as negative samples because they correspond to different interest centers. In this way, the present invention can reduce the possibility of false negative samples.
[0054] Generally speaking, using high-confidence clustering information as supervision information can improve the quality of positive and negative samples, thereby improving the performance of the contrastive task.
[0055] 223) The total loss function of the contrastive task:
[0056] Loss of positive sample pairs:
[0057] Loss of negative sample pairs:
[0058] Among them, N represents the number of samples, K represents the number of clustering centers, i represents the index of features, n and m represent variables, R im and R in are different feature vectors in the same clustering center, C i and C j respectively represent different preference centers.
[0059] The total loss function of the self-supervised task:
[0060] Among them, α represents the balance parameter.
[0061] Test-Time Training (TTT) is a widely adopted strategy to solve the Out-of-Distribution (OOD) problem. The test-time training stage uses self-supervised learning tasks to deeply understand the basic features of test samples. To solve the OOD recommendation problem, the present invention proposes a depth recommendation method for out-of-distribution generalization based on dual test-time training (DT3OR). This method adds a self-distillation task and a contrast task in the test-time training stage to effectively update the recommendation model, thereby allowing the model to adaptively shift the data to better understand the invariant preferences between users and the variable user features or item features with shifted data.
[0062] What is described in the above specification is only the specific implementation manners of the present invention. Various illustrative examples do not constitute limitations to the essence of the present invention. Those of ordinary skill in the art can make modifications or deformations to the previously described specific implementation manners after reading the specification without departing from the essence and scope of the invention.
Claims
1. An out-of-distribution generalization depth recommendation method based on double-test-time training, comprising the steps of: 1) In the pre-training stage, training a recommendation system model to form a pre-trained model; 2) During the test training phase, the pre-trained model is updated through self-supervised tasks to form an updated model. 3) During the inference phase, the updated model is used for inference to obtain the recommendation results; characterized in that: The self-supervised task in step 2) includes a self-distillation task, and the distance between user interest centers with the same preference is minimized through the self-distillation task; The method for updating the pre-trained model through the self-distillation task is: 211) First, user features, item features, and historical interactions are extracted through the pre-trained model; E u = f θ (X u , D); E i = f θ (X i , D); where X u and X i represent user features and item features respectively, D is the interaction history, E u and E i are embedding features, and f θ is a pre-trained model; Perform fusion: where is a join operation, and g() represents a multi-layer perceptron; 213) Obtain the clustering distribution result: where represents the K-means algorithm; 214) Calculate the target distribution P: where T is the temperature parameter; Objective function of self-distillation task: where q i represents the user preference center obtained by clustering, and p i represents the user preference distribution after sharpening; The self-supervised task in step 2) also includes a contrast task: the contrast task is to capture the correlation between users with similar preferences. The method for updating the pre-trained model through the contrast task is: 221) Select multiple high-confidence samples from the clustering distribution results in the self-distillation task to form a high-confidence sample set, perform clustering operations to obtain corresponding pseudo-labels, and divide them into multiple non-overlapping clusters R i , and regard the samples in the same cluster as positive samples; 222) Calculate the center of the high-confidence samples: C i = avg(R i ), i = 1, 2,..., K, where avg() represents the average function; regard different high-confidence centers as negative samples; 223) Total loss function of the comparison task: Loss of positive sample pairs: Loss of negative sample pairs: Among them, N represents the number of samples, K represents the number of clustering centers, i represents the index of features, n and m represent variables, R im and R in are different feature vectors in the same clustering center, C i and C j represent different preference centers respectively.
2. The out-of-distribution generalization depth recommendation method based on dual-test training according to claim 1, wherein: The total loss function of the self-supervised task: Among them, α represents the balance parameter.
Citation Information
Patent Citations
Knowledge perception-based graph self-supervised learning recommendation method
CN117634600A
Recommendation method and system based on comparison learning of double social views
CN117688255A