Self-attention sequence recommendation method fusing time-assisted features and contrast optimization
By fusing temporal auxiliary features with contrast-optimized self-attention sequence recommendation methods, the problems of data sparsity and insufficient temporal feature modeling are solved, and the robustness and accuracy of the recommendation system are improved, especially maintaining efficient recommendation performance under sparse data conditions.
Patent Information
- Application Number
- CN202510733046.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-12
AI Technical Summary
Existing sequence recommendation methods are unable to effectively capture changes in user interests when faced with data sparsity and insufficient temporal feature modeling, resulting in decreased recommendation performance, especially in cold start or interest transfer scenarios.
A self-attention sequence recommendation method with time-assisted features and contrastive optimization is introduced. Through the self-attention mechanism, time-aware information and contrastive learning, time-aware sequences are constructed and multi-view sequences are generated. The sequence encoder is trained using the contrastive learning loss function, and multiple rounds of fast adaptive training are performed through the meta-optimization module.
It improves the robustness and recommendation accuracy of the model in sparse interaction scenarios, alleviates the performance degradation caused by insufficient user data, and improves the model's generalization ability and recommendation performance.
Smart Images

Figure CN120632213A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of recommendation systems, and more specifically, to a self-attention sequence recommendation method that integrates time-assisted features with contrast optimization. Background Art
[0002] In the context of big data, sequential recommender systems have become an important means of addressing the information explosion and achieving personalized services. Their core goal is to predict users' likely future interests by analyzing their continuous interactions over time (such as browsing history, purchases, and viewing history), thereby enabling precise recommendations. These methods effectively capture the dynamics of user interests and therefore hold great potential for application in scenarios such as e-commerce, social networking platforms, and content recommendation.
[0003] As the scale of information continues to expand, mining valuable user behavior sequences from massive amounts of interaction data and providing content of interest to users has become a key research direction in the recommendation field. Current mainstream sequential recommendation methods typically divide user interaction trajectories into long-term and short-term behaviors, respectively modeling their stable interests and immediate preferences, thereby improving the personalization of recommendations.
[0004] However, this modeling process relies heavily on historical user interaction data, such as clicks, views, ratings, and favorites. In reality, most users have limited behavioral records, especially new and long-tail users, whose behavioral samples are scarce. This results in a highly sparse user-item interaction matrix. This data sparsity directly restricts the model's ability to learn users' personalized preferences and underlying intentions. It also affects the model's ability to capture item attributes and user behavior patterns, thereby weakening recommendation performance. Furthermore, during training, too few interaction samples can make it difficult for the model to achieve effective generalization, making it prone to overfitting. This is particularly effective in practical application scenarios such as cold start or interest shifts.
[0005] At the same time, traditional collaborative filtering methods or deep learning-based recommendation models often struggle to construct a complete profile of user interests and behavior trajectories when faced with sparse data, causing recommendation results to deviate from the user's true preferences. Therefore, there is an urgent need to design a novel recommendation method that not only enhances user behavior modeling capabilities under data-sparse conditions, but also maintains high recommendation accuracy and computational efficiency, thereby improving the system's robustness and generalization capabilities, and meeting the dual requirements of recommendation quality and efficiency in real-world applications. Summary of the Invention
[0006] To address the challenges of existing recommendation systems in dealing with data sparsity and inadequate temporal feature modeling, this paper proposes a self-attention sequence recommendation method that integrates temporal auxiliary features with contrastive optimization. This method effectively integrates temporal information from user behavior sequences to enhance the model's ability to capture user preferences. Conversely, contrastive learning is introduced to bring similar samples closer together and dissimilar samples further apart, thereby capturing user interests. A self-attention mechanism is also introduced to apply different weighting strategies to different parameters, thereby improving the model's sensitivity to temporal variations and the accuracy of recommendations. This method demonstrates good robustness in sparse interaction scenarios, helping to alleviate performance degradation caused by insufficient user data.
[0007] The technical solution provided by the present invention is: a self-attention sequence recommendation method integrating time-assisted features and contrast optimization, comprising the following steps:
[0008] S1. Based on the self-attention mechanism, we introduce the construction of time-aware information and contrastive learning to obtain the TCMSRec framework;
[0009] S2. Data processing and pre-training: aligning user behavior sequences with timestamps, constructing time-aware sequences, and generating multi-view sequences to pre-train the dataset.
[0010] S3. Construct positive and negative sample pairs based on the enhanced view pairs and use contrastive learning loss function to train the sequence encoder to improve the recognition of user representations;
[0011] S4. Introduce the meta-optimization module to perform multiple rounds of fast adaptive training on the model to obtain the trained TCMSRec framework.
[0012] S5. Use the trained TCMSRec to predict the recommendation results and obtain the final recommendation results.
[0013] Furthermore, the TCMSRec framework includes a time-aware enhancement module, a contrastive learning meta-optimization module, and a model optimization training module, which specifically include:
[0014] A time-aware sequence enhancement module built based on timestamps;
[0015] A contrastive learning meta-optimization module based on multi-view sequences;
[0016] A model optimization training module that builds dependencies between users and items based on self-attention mechanism learning.
[0017] Furthermore, the data processing and pre-training steps refer to organizing historical user interaction data into time-stamped user behavior sequences, reconstructing the original sequences based on time information, and pre-training sample user data using time-aware enhancement and contrastive learning mechanisms, which specifically include:
[0018] Obtain user historical data, including user ID number, project ID number and corresponding interaction timestamp, classify and process by user ID, and organize the data of user u's interaction within a certain time range into user subsequences S of length n u ={(v1,t1),(v2,t2),…,(v n ,t n )},v∈V,t∈N * , and according to the timestamp t i Sort in ascending order to ensure that the user behavior sequence meets the time dependency;
[0019] Divide all user sequences into training set, validation set and test set according to the set ratio;
[0020] Sample data is selected from the training set to form the input data of the pre-training stage. A time-aware data augmentation strategy is applied to the selected samples to construct multi-view user sequences with different temporal structures. Users with more uniform time interval changes are prioritized for pre-training data construction.
[0021] Select the sample data set as the training data for the pre-training module.
[0022] Furthermore, the pre-processed data is sorted by time-homogenized user behavior and then input into the model to train the TCMSRec framework, ultimately obtaining a trained TCMSRec model. The specific process includes:
[0023] By sorting the pre-processed user-item interaction time intervals, we select user interaction sequences corresponding to items with uniform interaction time as the training set for TCMSRec.
[0024] When training the training set, the sequential recommendation model TCMSRec introduces a time-aware data augmentation strategy. Different augmented data are generated based on the length of the user behavior sequence. A contrastive learning loss function is introduced by constructing contrasting positive and negative sample pairs. This makes positive samples closer and negative samples further apart, resulting in an augmented sequence closer to the original sequence. This generates new training data while maintaining the integrity of the original user behavior sequence.
[0025] The optimizer is used to iteratively optimize the joint loss function consisting of the contrast loss and the weighted recommendation loss for the original sequence. The trained TCMSRec model has stronger sequence modeling and time perception capabilities.
[0026] The optimal parameter combination of the model is determined through the validation set, and its recommendation performance is evaluated on the test set, thus obtaining a sequence recommendation model that integrates contrastive learning and temporal information modeling, with good convergence and excellent results.
[0027] Furthermore, the recommendation task is performed based on the trained TCMSRec framework to generate the final recommendation results. The specific process includes:
[0028] Collect the historical interaction data of the target user u, analyze the historical interaction time series of the user according to the data preprocessing process defined in step 2), calculate the variance value of the time interval of the behavior, and perform structural modeling on the user by combining the variance sorting strategy to construct the user behavior sequence S u , S u The data is input into the sequential recommendation model TCMSRec trained and optimized by contrastive learning in step 3), and the sequence encoder is used to perform temporal feature modeling and semantic enhancement on the user sequence, and output the user's preference scores for all candidate products.
[0029] The preference scores are represented by the sequence h u The dot product with the product embedding vector matrix E is calculated as:
[0030]
[0031] Among them, e i represents the embedding vector of the i-th item, represents the user's predicted interest score for product i;
[0032] The candidate products are sorted in descending order according to the predicted scores of all products, and the top N products with the highest scores are selected to form the recommendation result set, completing the Top-N product recommendations for user u.
[0033] Furthermore, the contrastive learning module is used to train the sequence encoder, which specifically includes:
[0034] Suppose there is a function f:X→R d , used to map the input user behavior sequence into a dense representation vector. The goal of contrastive learning training is to minimize the distance between enhanced views and maximize the discrimination between different user representations.
[0035] Let S u Represents the historical behavior sequence of user u, and uses the time-aware data enhancement strategy to generate two different views of it and Constructing positive sample pairs and using other users' views or shuffled views as negative sample pairs includes the following steps:
[0036] 1) Input the enhanced view into the shared sequence encoder to obtain the representation vector and
[0037] 2) Each pair of positive samples With all negative samples Compute the contrastive loss function.
[0038] The above process is iterated continuously until the training loss reaches the convergence condition, thereby obtaining the user behavior sequence representation optimized by contrastive learning, which is used in the training phase of the TCMSRec model.
[0039] In addition, a time-aware inter-item dependency modeling module is constructed based on the self-attention mechanism to mine the potential correlation and temporal features between items at different time points in the sequence. The specific form of its self-attention mechanism is as follows:
[0040] Let the input sequence be matrix X∈R n×d , where n is the sequence length and d is the embedding dimension of each element. First, the input X is mapped into three vectors: query, key, and value. The expression is as follows:
[0041] Q=XW Q ,K=XW K ,V=XW V
[0042] in, is the learnable parameter matrix,
[0043] Calculate the similarity between the query and the key, and perform scaling and normalization. The attention weight calculation expression is as follows:
[0044]
[0045] Among them, QK T ∈R n×n Indicates the attention score of each position to other positions, To scale the factor and prevent the inner product from being too large, softmax normalizes the scores into weights.
[0046] In order to enhance the expressiveness of the model, multi-head attention is introduced, and the expression is as follows:
[0047] MultiHead(Q,K,V)=Concat(head1,head2,…,head h )W O
[0048]
[0049] Concat represents the concatenation operation of the representation vector output by each attention head, h is the number of attention heads, and W O Used to fuse the multi-head attention output into a unified dimension, A separate projection matrix for each attention head.
[0050] The beneficial effects of the method and system of the present invention are as follows: the present invention introduces time-auxiliary information based on the self-attention mechanism and proposes a TCMSRec framework, aligns user behavior sequences with timestamps to construct time-aware sequences, and better learns user interest preferences; constructs positive and negative sample pairs based on enhanced views, and uses a contrastive learning loss function to train the sequence encoder to maximize the distance between positive samples and push negative samples away, so that the generated new samples are closer to the original users and reduce noise interference; introduces a meta-optimization module to construct a minimum task unit for model training, and performs multiple rounds of rapid adaptive training on the model to obtain a trained model framework; while performing data enhancement operations, the framework can retain the key features of user behavior and effectively improve the performance of sequence modeling. At the same time, it adopts a dynamic optimization enhancement strategy to improve the recommendation performance and generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 This is a flowchart of the steps of a self-attention sequence recommendation method that integrates time-assisted features and contrast optimization in the present invention;
[0052] Figure 2 This is a structural block diagram of a self-attention sequence recommendation method that integrates time-assisted features and contrast optimization in the present invention;
[0053] Figure 3 This is a model training and result prediction flow chart of a self-attention sequence recommendation method that integrates time-assisted features and contrast optimization. DETAILED DESCRIPTION
[0054] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments. The step numbers in the following embodiments are for ease of explanation only and do not limit the order of the steps. The specific execution process can be flexibly adjusted by those skilled in the art according to actual needs.
[0055] See also Figure 1 The present invention proposes a self-attention sequence recommendation method that integrates time-assisted features and contrast optimization, which specifically includes the following steps:
[0056] S1 introduces contrastive learning and temporal position encoding based on the self-attention mechanism and proposes the TCMSRec framework;
[0057] S1.1 Time-aware sequence enhancement module based on timestamp construction;
[0058] Specifically, we extract user IDs, interaction items, and time information and convert them into binary pairs, forming a data structure containing user and item, and user and time. We then enhance the data based on user behavior and time information, transforming uneven user sequences into uniform sequences through enhancement strategies, thereby preserving the key characteristics of user behavior at the data level.
[0059] S1.2 Contrastive learning meta-optimization module based on multi-view sequences;
[0060] Specifically, an augmentation module is introduced based on the existing contrastive learning architecture to increase sample diversity by generating richer sequence views. This method constructs multiple differentiated views, calculates a contrastive loss for each view, and optimizes this loss function together with the original sequence loss function. During training, the loss function is iteratively updated to dynamically adjust the augmentation module parameters, thereby implementing an adaptive augmentation strategy and further improving the effectiveness and flexibility of contrastive learning.
[0061] S1.3 is a model optimization training module that builds dependencies between users and items based on self-attention mechanism learning.
[0062] Specifically, the self-attention mechanism effectively models multiple dependencies between users and items, including user-item, item-time, and item-item interactions. For each user behavior sequence, sequence enhancement is first performed according to S1.1. Then, based on the multi-view representation constructed in S1.2, the self-attention mechanism is used to mine item features that are highly correlated with user preferences. This allows users to predict items they might be interested in next, thereby capturing potential connections within the sequence and significantly improving the accuracy and generalization capabilities of the recommendation system.
[0063] S2 data processing and pre-training: sorting the variance of the user's historical behavior sequence, selecting user data with more uniform interaction intervals, and training them;
[0064] S2.1 Data processing and pre-training: converting historical user interaction data in the dataset into user behavior sequences;
[0065] Specifically, datasets for recommender systems were obtained from Yelp and the Amazon Product Dataset. The Yelp dataset, a public dataset released by Yelp, covers various business categories, such as restaurants, haircuts, and shopping. It can be used for categorical information fusion modeling and multi-task learning. The data includes user reviews, ratings, and timestamps of businesses. In recommender systems, only positive feedback interactions with a rating greater than 3 are retained. Timestamps are extracted to form an ordered sequence of user interactions, and k-core filtering is applied, for example, retaining users or businesses with at least five interactions. The Beauty and Sports datasets are from the Amazon Product Dataset. The Beauty dataset records user reviews of products in the "Beauty" category and is suitable for studying personalized representation learning and interest transfer tasks. The Sports dataset records user purchase ratings of sports and outdoor products. Its characteristics include fewer user behaviors, a low average number of interactions, a large variety of items, and low repetition rates, making it suitable for tasks such as data sparsity processing and temporal structure learning. Processing the dataset yields user IDs, interaction timestamps, and the IDs of the items interacted with.
[0066] First, for the above dataset, we can read the raw data to obtain user-item-time triples. Next, we sort by user-time and user-item pairs, construct a sequence, and save it in the format of the input model. Specifically, each user ID corresponds to the timestamps of all their historical interactions. Based on the user's historical interaction sequence and the degree of change in interaction time, we use variance to process and sort the user behavior sequence. Users with high rankings have small timestamp changes, that is, the user behavior intervals are relatively uniform. Users with low rankings have large timestamp changes, that is, the behavior distribution is uneven or not concentrated. Similarly, each user ID corresponds to the ID sequence of their historical interaction items. Based on the above operations, we sort the items that the user has interacted with, that is, obtain data on user-behavior timestamps and user-interaction items.
[0067] Reference Figure 3 ,The preprocessed dataset is divided by first evenly dividing the dataset by the ,interaction amount, generating the first half of low variance user interactions and the second half of high variance user interactions. ,Secondly, evenly dividing the dataset by the number of users, generating the first 50% of users and the last 50% of users, ,dynamically segmenting the data.
[0068] S2.2 Select a sample data set from the training set as training data for the pre-training module;
[0069] Specifically, to improve pre-training efficiency and avoid wasting resources, we don't directly use all user interaction data in the training set as pre-training input. This is because real-world data often exhibits a long-tail distribution, with a large number of users experiencing extreme low or high interaction frequencies. These samples contribute little to the pre-training objective. In contrast, pre-training relies more on representative user data with relatively stable behavior patterns to achieve more effective model initialization and optimization.
[0070] S2.3 For the user sequences selected during the pre-training phase, we calculated the variance of their interaction time intervals and sorted the user behavior sequences accordingly. We prioritized user samples with relatively stable time intervals and more representative interaction distributions to construct a pre-training data subset, which served as the input for the pre-training module. This strategy effectively improved the model's ability to learn user behavior patterns, enhanced the generalization of parameter initialization, and provided a better convergence foundation for subsequent training.
[0071] In specific implementation, we expand the user's interaction records into a one-dimensional behavior sequence in chronological order, which is recorded as:
[0072] S u =[v1,v1,…,v n ].
[0073] Among them, S u represents the interaction sequence of user u, and n is the total length of its historical interaction behaviors.
[0074] Sort each user's interaction timestamp sequence in ascending order of time, and calculate the variance of their continuous time differences, which is recorded as:
[0075]
[0076] Var u is the variance of the continuous time difference, t i+1 -t i is the continuous time difference, is the average time interval for users.
[0077] Then according to Var u Sort users in ascending order, prioritizing users with more uniform time intervals to construct a temporally stable training subset for pre-training or view construction. Align user-item sequences and user-time interval sequences one-to-one to generate a standardized temporal behavior sequence, which serves as input for subsequent time-aware modeling and contrastive learning training.
[0078] S3 uses the sorted sequences obtained after pre-training to train the TCMSRec framework using contrastive learning to obtain a trained TCMSRec framework;
[0079] Specifically, data augmentation is performed based on user interaction time intervals to generate diverse comparative learning samples. The variance of the time series is calculated to measure the volatility of user interaction times. A time series is input and the variance value is obtained. A larger value indicates greater time interval fluctuation. Data augmentation is then performed, with different augmentation strategies being selected, such as generating multi-view comparative samples or randomly selecting an augmentation method. Similarity calculations are performed for the insertion or replacement of similar items in the data augmentation to generate an item similarity matrix.
[0080] The processed training set is fed into the sequential recommendation model TCMSRec for training. The Transformer Encoder is used to encode the dataset, and then multi-head attention is used to assign different weights to each parameter. The recommendation task and contrastive learning task are integrated, and the optimizer is used to iteratively optimize the TCMSRec loss function until TCMSRec converges.
[0081] The validation set is used to screen the different recommendation results and obtain the optimal model parameters. Finally, the test set is used to evaluate the model effect using the optimal parameters, and the optimal performance of the sequence recommendation model based on the self-attention mechanism and the introduction of time-assisted information is obtained.
[0082] S3.1 To enhance the temporal consistency and discriminability of user behavior representation, the TCMSRec model introduces a contrastive learning module. During training, two data-augmented views based on temporal perturbations are constructed for each user and fed into a shared encoder to generate representation vectors. The InfoNCE loss is used to minimize the distance between two views of the same user while maximizing the distance to views of other users. This contrastive loss, together with the original sequence loss, forms a joint training objective:
[0083] L total =L rec +λ·L cl
[0084] Among them L total represents the combined traditional recommendation loss and contrast loss, L rec is the traditional recommendation loss, L cl is the contrast loss, and λ is the weight hyperparameter that balances the two losses.
[0085] Specifically, we use cross entropy loss to calculate the original recommendation loss, that is, for the user sequence S u =(v1,v2,…,v n ) The goal is to predict the last item v n , the cross entropy loss expression is as follows:
[0086] Lrec =-logP(v n |v1,v2,…,v n-1 )
[0087] The goal of contrastive learning is to make the two augmented views of the same user closer together and the representations of different users farther apart. First, the above augmentation operation is applied to a given user behavior sequence to generate two different augmented views. The specific expression is as follows:
[0088]
[0089] in For two different enhanced views, is the vector output by the encoder for the corresponding view, which is used as a positive sample pair.
[0090] For the loss between enhanced views, we use the InfoNCE loss function, that is, assume there are N users in the batch, each user generates two different enhanced views, a total of 2N representations, and constructs positive and negative sample pairs.
[0091] Specifically, the contrast loss expression is as follows:
[0092]
[0093] Among them, L cl represents contrast loss, v represents item, u represents user, is the cosine similarity, and τ is the temperature parameter used to control the smoothness of the distribution.
[0094] S3.2 Multi-head self-attention mechanism inter-item dependency learning, its expression is:
[0095] Let the input sequence be matrix X∈R n×d , where n is the sequence length and d is the embedding dimension of each element. First, the input X is mapped into three vectors: query, key, and value. The expression is as follows:
[0096] Q=XW Q ,K=XW K ,V=XW V
[0097] in, is the learnable parameter matrix,
[0098] Calculate the similarity between the query and the key, and perform scaling and normalization. The attention weight calculation expression is as follows:
[0099]
[0100] Among them, QK T ∈R n×n Indicates the attention score of each position to other positions, To scale the factor and prevent the inner product from being too large, softmax normalizes the scores into weights.
[0101] In order to enhance the expressiveness of the model, multi-head attention is introduced, and the expression is as follows:
[0102] MultiHead(Q,K,V)=Concat(head1,head2,…,head h )W O
[0103]
[0104] Concat represents the concatenation operation of the representation vector output by each attention head, h is the number of attention heads, and W O Used to fuse the multi-head attention output into a unified dimension, A separate projection matrix for each attention head.
[0105] S3.3 further verifies and optimizes the parameter configuration of the TCMSRec framework by evaluating it on the validation set, and finally obtains the optimal parameter set for deployment.
[0106] Specifically, if Figure 3 As shown in Figure 2, the operation process of the verification phase is consistent with the training process in step S3. Its main goal is to evaluate the generalization ability of the model parameters learned by the TCMSRec framework during training on unseen data, ensuring that it still has good recommendation accuracy and time series modeling capabilities on the verification set.
[0107] The validation process uses the same input format and model structure as the training one, including the multi-head self-attention module, the time-aware enhancement mechanism, and the random inheritance strategy, while keeping the contrastive learning module active to ensure the consistency of user representations in the enhanced view.
[0108] Based on the evaluation results on the validation set, the system automatically analyzes hyperparameter configurations that may have performance bottlenecks and fine-tunes and retrains them based on performance metrics. Ultimately, the parameter combination that achieves the best performance on the validation set is retained as the final parameter set for the TCMSRec framework for subsequent testing, evaluation, and online deployment.
[0109] S3.4 evaluates the final parameter set of the TCMSRec framework on the test set to verify the performance of the model in actual recommendation scenarios, and ultimately obtains a trained and deployable TCMSRec framework.
[0110] Specifically, the testing phase follows the training process in step S3, including the input format, network structure, time-aware enhancement module and multi-head self-attention mechanism, as well as the joint optimization strategy to ensure the consistency of the test evaluation process and training.
[0111] The core goal of the test is to verify whether the TCMSRec framework, which has undergone pre-training, contrastive learning optimization, and hyperparameter screening, can maintain stable and generalized recommendation accuracy, temporal consistency, and user preference expression capabilities on unseen user behavior data. In particular, the model can robustly output Top-N recommendation results in complex behavioral scenarios such as sparse interactions, temporal disturbances, or interest jumps.
[0112] During the evaluation process, the performance indicators NDCG@K (Normalized Discounted Cumulative Gain) and Hit Rate@K (Click-through Rate) will be used to quantify the recommendation performance of TCMSRec, ultimately confirming the effectiveness of the current model structure and parameter set in real-world application environments, and completing the final training and performance verification of the TCMSRec framework.
[0113] S4 uses the trained TCMSRec framework to make recommendation predictions for target users and generate the final recommendation results.
[0114] Specifically, the system first collects the historical interaction data of the target user u, and through the preprocessing process defined in step 2), organizes the user's historical behavior into a user behavior sequence S in chronological order. u .
[0115] Then, the user sequence S u Input into the sequence recommendation model TCMSRec with optimal parameter configuration based on the self-attention mechanism and the introduction of time auxiliary information obtained in step 3) for prediction. The model generates the user's behavior representation vector h through the encoding module u , and matches it with the embedding vector matrix E of all products to calculate the user's preference score for each candidate product. The expression is as follows:
[0116]
[0117] in, represents the predicted interest score of user u for product i, h u is the user representation vector based on time perception enhancement and behavior dependency modeling, e i is the embedding vector of product i.
[0118] Based on the scores of all candidate products Sort from largest to smallest and output the top-N products with the highest scores to form a top-N recommendation result set for user u.
[0119] like Figure 2 As shown in the figure, a self-attention sequence recommendation method that integrates time-assisted features and contrast optimization includes the following modules:
[0120] The time-aware sequence module built based on timestamps is used to perform time-aware enhancement processing on the user's historical behavior sequence to enhance the model's ability to model the temporal changes in behavior and improve the temporal consistency and generalization ability of user representation.
[0121] A contrastive learning meta-optimization module built based on multi-view sequences is used to perform contrastive learning optimization on user behavior sequences under different time views in a shared representation space. The module introduces a meta-optimization strategy to automatically search for the optimal combination configuration under the structure of multiple views, further improving the contrastive learning effect.
[0122] A model optimization training module that builds dependencies between users and items based on self-attention mechanism learning is used to model high-order dependencies between items in user behavior sequences during the training phase, capture the evolutionary pattern of user interests, and achieve dual optimization of recommendation accuracy and representation consistency.
[0123] The prediction module outputs the user's preference score for each product based on the similarity relationship between the user encoding vector and the product embedding matrix, and generates the final recommendation result.
[0124] The above is a schematic description of the present invention and its embodiments, which is not restrictive. The drawings show only one embodiment of the present invention, and the actual structure is not limited thereto. Therefore, if a person skilled in the art is inspired by this and, without departing from the purpose of the present invention, designs a structure and embodiment similar to this technical solution without inventiveness, they shall fall within the scope of protection of the present invention.
Claims
1. A self-attention sequence recommendation method that integrates temporal auxiliary features and contrast optimization, characterized in that: The following steps are involved: S1. Based on the self-attention mechanism, we introduce the construction of time-aware information and contrastive learning to obtain the TCMSRec framework; S2. Data processing and pre-training: aligning user behavior sequences with timestamps, constructing time-aware sequences, and generating multi-view sequences to pre-train the dataset. S3. Construct positive and negative sample pairs based on the enhanced view pairs and use contrastive learning loss function to train the sequence encoder to improve the recognition of user representations; S4. Introduce the meta-optimization module to perform multiple rounds of fast adaptive training on the model to obtain the trained TCMSRec framework. S5. Use the trained TCMSRec to predict the recommendation results and obtain the final recommendation results.
2. The self-attention sequence recommendation method integrating temporal auxiliary features and contrast optimization according to claim 1, characterized in that: The TCMSRec framework includes a time-aware enhancement module, a contrastive learning optimization module, and a model optimization training module, which specifically include: A time-aware sequence enhancement module built based on timestamps; A contrastive learning meta-optimization module based on multi-view sequences; A model optimization training module that builds dependencies between users and items based on self-attention mechanism learning.
3. The self-attention sequence recommendation method integrating temporal auxiliary features and contrast optimization according to claim 1, characterized in that: The data processing and pre-training steps involve organizing historical user interaction data into time-stamped user behavior sequences, reconstructing the original sequences based on time information, and pre-training sample user data using time-aware enhancement and contrastive learning mechanisms. Specifically, they include: Obtain user historical data, including user ID number, project ID number and corresponding interaction timestamp, classify and process by user ID, and organize the data of user u's interaction within a certain time range into user subsequences S of length n u ={(v1,t1),(v2,t2),…,(v n ,t n )},v∈V,t∈N * , and according to the timestamp t i Sort in ascending order to ensure that the user behavior sequence meets the time dependency; Divide all user sequences into training set, validation set and test set according to the set ratio; Sample data is selected from the training set to form the input data of the pre-training stage. A time-aware data augmentation strategy is applied to the selected samples to construct multi-view user sequences with different temporal structures. Users with more uniform time interval changes are prioritized for pre-training data construction. Select the sample data set as the training data for the pre-training module.
4. The self-attention sequence recommendation method integrating temporal auxiliary features and contrast optimization according to claim 1, characterized in that: The data obtained by the preprocessing is used to sort the users through the time uniformization operation and then input into the model to train the TCMSRec framework to obtain the trained TCMSRec framework, which specifically includes: By sorting the pre-processed user-item interaction time intervals, we select user interaction sequences corresponding to items with uniform interaction time as the training set for TCMSRec. When training the training set, the sequential recommendation model TCMSRec introduces a time-aware data augmentation strategy. Different augmented data are generated based on the length of the user behavior sequence. A contrastive learning loss function is introduced by constructing contrasting positive and negative sample pairs. This makes positive samples closer and negative samples further apart, resulting in an augmented sequence closer to the original sequence. This generates new training data while maintaining the integrity of the original user behavior sequence. The optimizer is used to iteratively optimize the joint loss function consisting of the contrast loss and the weighted recommendation loss for the original sequence. The trained TCMSRec model has stronger sequence modeling and time perception capabilities. The validation set is used to screen the optimal parameter configuration of the model, and finally its recommendation performance is evaluated on the test set to obtain a sequence recommendation model based on contrastive learning and time-assisted information modeling with stable convergence and optimal effect.
5. The self-attention sequence recommendation method integrating temporal auxiliary features and contrast optimization according to claim 1, characterized in that: The method of using the trained TCMSRec framework to predict recommendation results and obtain recommendation results specifically includes: Collect the historical interaction data of the target user u, analyze the historical interaction time series of the user according to the data preprocessing process defined in step 2), calculate the variance value of the time interval of the behavior, and perform structural modeling on the user by combining the variance sorting strategy to construct the user behavior sequence S u , S u The data is input into the sequential recommendation model TCMSRec trained and optimized by contrastive learning in step 3), and the sequence encoder is used to perform temporal feature modeling and semantic enhancement on the user sequence, and output the user's preference scores for all candidate products. The preference scores are represented by the sequence h u The dot product with the product embedding vector matrix E is calculated as: Among them, e i represents the embedding vector of the i-th item, represents the user's predicted interest score for product i; The candidate products are sorted in descending order according to the predicted scores of all products, and the top N products with the highest scores are selected to form the recommendation result set, completing the Top-N product recommendations for user u.
6. The self-attention sequence recommendation method integrating temporal auxiliary features and contrast optimization according to claim 2, characterized in that: The method of using a contrastive learning module to train a sequence encoder specifically includes: Suppose there is a function f:X→R d , which is used to map the input user behavior sequence into a dense representation vector. The goal of contrastive learning training is to minimize the distance between enhanced views and maximize the discrimination between different user representations. Let S u Represents the historical behavior sequence of user u, and uses the time-aware data enhancement strategy to generate two different views of it and Constructing positive sample pairs and using other users' views or shuffled views as negative sample pairs includes the following steps: 1) Input the enhanced view into the shared sequence encoder to obtain the representation vector and 2) For each pair of positive samples With all negative samples Compute the contrastive loss function. The above process is repeated until the training loss converges, and finally the user behavior sequence encoding optimized by contrastive learning is obtained for training the TCMSRec model.
7. The self-attention sequence recommendation method integrating temporal auxiliary features and contrast optimization according to claim 3, characterized in that: The time-aware item dependency learning module based on the self-attention module is used to capture the potential dependencies and temporal patterns between items at different moments in the sequence. The self-attention mechanism is expressed as follows: Let the input sequence be matrix X∈R n×d , where n is the sequence length and d is the embedding dimension of each element. First, the input X is mapped into three vectors: query, key, and value. The expression is as follows: Q=XW Q ,K=XW K ,V=XW V in, is the learnable parameter matrix, Calculate the similarity between the query and the key, and perform scaling and normalization. The attention weight calculation expression is as follows: Among them, QK T ∈R n×n Indicates the attention score of each position to other positions, To scale the factor and prevent the inner product from being too large, softmax normalizes the scores into weights. In order to enhance the expressiveness of the model, multi-head attention is introduced, and the expression is as follows: MultiHead(Q,K,V)=Concat(head1,head2,…,head h )W O head i =Attention(QW i Q ,KW i K ,VW i V ) Concat represents the concatenation operation of the representation vector output by each attention head, h is the number of attention heads, and W O Used to fuse the multi-head attention output into a unified dimension, QW i Q ,KW i K ,VW i V A separate projection matrix for each attention head.
Citation Information
Patent Citations
Authentication of subscriber station
WO2001030104A1
Cited By
Cross-sequence mixed sequence recommendation method based on user correlation guidance
CN121705519A
A cross-sequence hybrid sequence recommendation method based on user relevance guidance
CN121705519B
Encrypted traffic classification method and system based on cross-modal comparative learning, and medium
CN121727866A