Timing sequence enhancement and multi-granularity intention guidance adversarial generation recommendation method

By generating timing enhancement sequences and building multiple loss functions, the problems of sparse data and insufficient user intention capture in sequence recommendations are solved, and the accuracy and reliability of the recommendation system are improved.

CN120336632APending Publication Date: 2025-07-18CHONGQING UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510426307.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing sequence recommendation methods have shortcomings in data sparseness and user intention capture, and it is difficult to effectively use user historical interactive data for accurate recommendations.

Method used

By obtaining user product interaction sequences and time data, using multiple data enhancement methods to generate timing enhancement sequences, and building multiple cross entropy and comparison loss functions, jointly training feature coding modules and prediction modules to improve the model's understanding and generalization ability of user intentions.

Benefits of technology

It improves the accuracy and reliability of the recommendation system, can better capture the user's true intentions and behavior patterns, and provide more accurate recommendation services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336632A_ABST
    Figure CN120336632A_ABST
Patent Text Reader

Abstract

The invention relates to a time sequence enhanced and multi-granularity intention guided adversarial generation recommendation method, which comprises the following steps of: firstly, obtaining an original user commodity interaction sequence and interaction time data, and obtaining a time sequence enhanced sequence through a plurality of data enhancement modes; then, by constructing a plurality of cross entropy loss functions and comparison loss functions, model training is constrained from different angles, and user behavior patterns and potential correlation are fully mined; and finally, performing joint training on the feature coding module, and performing fine adjustment on the prediction module to obtain a recommendation model. The method focuses on solving the problems of data sparseness and insufficient model generalization ability of sequence recommendation (SR) under information overload. Under the background that information technology development causes serious information overload, although the SR is concerned, the SR faces many challenges; a traditional method depends on project prediction task optimization parameters, is easily influenced by data sparsity, and is difficult to capture real intentions of users and correlation between sequences; the method effectively improves the accuracy and reliability of a recommendation system, provides more accurate recommendation services for users, and has significant application value in the field of sequence recommendation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer science and technology, and particularly relates to an adversarial generation recommendation method with temporal enhancement and multi-granularity intention guidance. Background Art

[0002] With the rapid development of information technology, the problem of information overload has become increasingly severe. To alleviate this problem, personalized recommendation systems have been widely applied. Among them, sequential recommendation (SR) has attracted much attention because it can predict the next item of interest to the user based on the time series characteristics of the user's historical items. The core of the sequential recommendation task lies in inferring high-quality user behavior representations from the user's historical interaction data to accurately recommend items. In the early stage, Markov chains were used to learn the pairwise transition relationships between items, and RNN-based models further explored sequential correlations. Recently, the Transformer architecture has been introduced into the SR field due to its powerful sequence encoding ability, and the self-attention mechanism is used to capture the importance of items, improving the recommendation effect. However, most of these methods only optimize a large number of parameters depending on the item prediction task, and are vulnerable to the problem of data sparsity. When the training data is limited, it is difficult to accurately infer the user representation.

[0003] For modeling the SR problem, it is crucial to capture the relationships between items in the sequence. However, the interaction between users and platform items is limited, and affected by cross-platform restrictions and privacy protection, a large amount of historical data cannot be collected or used for model training, severely restricting the recommendation performance of sequential models. To address this challenge, contrastive self-supervised learning has been introduced, creating fine-grained views through enhancement operators to solve problems such as data sparsity and noisy data. However, the time intervals in the sequence vary greatly, the user preferences drift over time, and the feature encoders of existing methods may overly separate similar but different categories of sequences, hindering the capture of the consistency between similar sequences.

[0004] In a real shopping environment, sequences with similar purchase intentions have important reference value for improving the accuracy of commodity prediction. For example, ICLRec uses the expectation maximization (EM) to maximize the consistency between the sequence view and a single intention, constructing a user intention distribution function to understand the diverse preferences of users. However, such methods usually only model the interaction sequences of a single user, ignoring the potential correlations between users with similar subsequence patterns, and not fully utilizing the shared behavior patterns between different users, restricting the model's understanding and generalization ability of complex user intentions. Summary of the Invention

[0005] To solve the problems in the background art, one aspect of the present invention provides an adversarial generation recommendation method with temporal enhancement and multi-granularity intention guidance, including:

[0006] S1: Obtain the original user-item interaction sequences for training, as well as the time data of the interaction between the user and the item;

[0007] S2: Use multiple data augmentation methods according to the interaction time between the user and the commodity to perform data augmentation on the original user-commodity interaction sequence to obtain multiple time-series augmented sequences;

[0008] S3: Input the time-series augmented sequences into the feature encoding module to obtain the feature representations of the time-series augmented sequences;

[0009] S4: Determine the data augmentation methods adopted by the time-series augmented sequences according to the feature representations of the time-series augmented sequences, and construct the first cross-entropy loss function based on the discrimination results;

[0010] S5: Determine whether two different time-series augmented sequences adopt the same data augmentation method according to the feature representations of the two different time-series augmented sequences, and construct the second cross-entropy loss function based on the discrimination results;

[0011] S6: Divide the original user-commodity interaction sequence into multiple different first subsequences, and use the first subsequences of the same target object as positive sample pairs and the first subsequences of different target objects as negative sample pairs; wherein, the target object of the first subsequence represents the last commodity interacted by the user in the first subsequence;

[0012] S7: Input the first subsequences into the feature encoding module to obtain the feature representations of the first subsequences, and construct the first contrast loss function according to the feature representations of the positive sample pairs and negative sample pairs;

[0013] S8: Use the first subsequences of the same target object as the second subsequences, perform clustering on the second subsequences by the kmean algorithm according to the feature representations of the second subsequences, and use the second subsequences under the same cluster as positive sample pairs and the second subsequences under different clusters as negative sample pairs to construct the second contrast loss function;

[0014] S9: Use the augmented sequences under the same user as positive sample pairs and the augmented sequences under different users as negative sample pairs to construct the third contrast loss function;

[0015] S10: Construct the total loss function according to the first cross-entropy loss function, the second cross-entropy loss function, the first contrast loss function, the second contrast loss function and the third contrast loss function to jointly train the feature encoding module. After the training is completed, fix the parameters of the feature encoding module and fine-tune the prediction module on the labels of the original user-commodity interaction sequence to obtain the trained recommendation model.

[0016] Another aspect of the present invention provides a time-series augmentation and multi-granularity intention-guided adversarial generation recommendation system, the system includes a memory and a processor; the memory is used to store application programs; the processor is used to run the application programs and execute the described time-series augmentation and multi-granularity intention-guided adversarial generation recommendation method.

[0017] Another aspect of the present invention provides a computer storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned method for adversarial generation recommendation with time series enhancement and multi-granularity intention guidance is implemented.

[0018] The present invention has at least the following beneficial effects

[0019] In terms of data processing, the present invention obtains the original user-item interaction sequence and interaction time data, and uses a variety of data enhancement methods to generate a variety of time series enhanced sequences, alleviating the data sparsity problem and enriching the training data. During the model training process, multiple loss functions are constructed, including a first cross-entropy loss function and a second cross-entropy loss function constructed according to the feature representations of the time series enhanced sequences, as well as multiple contrastive loss functions constructed based on subsequences and enhanced sequences. These loss functions constrain the model from different perspectives, comprehensively improving the model performance. By discriminating the data enhancement methods used for the time series enhanced sequences and whether different time series enhanced sequences use the same enhancement method to construct the loss function, it helps the model better learn the features under different enhancement methods and improve the model's understanding ability of data. Dividing the original user-item interaction sequence into subsequences and constructing corresponding contrastive loss functions can guide the model to learn the user behavior pattern from the multi-granularity intention level and capture the user's true intention. At the same time, by constructing contrastive loss functions for the enhanced sequences of the same user and different users, the potential correlations between users are fully explored, and the model's understanding and generalization ability for complex user intentions are improved. Finally, by constructing a total loss function to jointly train the feature encoding module and fine-tuning the prediction module, a trained recommendation model is obtained, effectively improving the accuracy and reliability of the recommendation system and providing users with higher-quality and more accurate recommendation services. Description of the Drawings

[0020] Figure 1 It is a schematic flowchart of the method of the present invention. Detailed Embodiments

[0021] The following specific examples illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0022] Please refer to Figure 1, the present invention provides an adversarial generation recommendation method with temporal enhancement and multi-granularity intention guidance, including inputting the user-item interaction sequence into a trained recommendation model for prediction to obtain a recommendation result. Among them, the recommendation model includes a feature encoding module and a prediction module, and the training process of the recommendation model includes:

[0023] Preferably, in this embodiment, the feature extraction module includes a Transformer encoder, and the prediction module includes a fully connected neural network + softmax.

[0024] Using a Transformer encoder as the feature extraction module in this embodiment can effectively capture the long-term and short-term dependencies and complex patterns in the user interaction sequence, and comprehensively consider the associations between items in the sequence. The prediction module of the fully connected neural network + softmax can convert the encoded features into accurate recommendation probabilities. Through such model training and prediction methods, compared with traditional recommendation algorithms, it can capture user intentions more accurately and improve the accuracy and relevance of recommendations.

[0025] S1: Obtain the original user-item interaction sequence for training, as well as the time data of the user-item interaction;

[0026] In this embodiment, the original user-item interaction sequence for training can be selected from Sports and Beauty in the Amazon dataset, or the Yelp dataset can also be used, etc.

[0027] S2: Use a variety of data augmentation methods to perform data augmentation on the original user-item interaction sequence according to the interaction time between the user and the item to obtain a variety of temporally enhanced sequences;

[0028] Preferably, the data augmentation methods include: insertion operation, cropping operation, masking operation, and rearrangement operation;

[0029] The insertion operation includes: calculating the interaction time difference between adjacent items according to the interaction time between the user and the item in the original user-item interaction sequence, inserting the item with the largest sum of similarities to the two adjacent items between the two adjacent items with the largest interaction time difference, and the inserted item is not in the original user-item interaction sequence, to obtain a temporally enhanced sequence;

[0030] The cropping operation includes: setting a cropping length l, based on the principle of a sliding window, cropping a sequence with a length of l from the original user-item interaction sequence with a step size of 1, and calculating the time standard deviation of this sequence, and selecting the sequence with the smallest time standard deviation among the cropped sequences as the temporally enhanced sequence;

[0031] The masking operation includes: calculating the interaction time difference between adjacent items according to the interaction time between the user and the item in the original user-item interaction sequence, and selecting one item to delete between the two adjacent items with the largest interaction time difference to obtain a time-series enhanced sequence;

[0032] The rearrangement operation includes: screening out the subsequence with the largest time standard deviation from all subsequences of the original user-item interaction sequence and shuffling it to obtain a time-series enhanced sequence.

[0033] In this embodiment, it is assumed that the original user-item interaction sequence is S = [A, B, C, D], and the corresponding interaction times are t A = 1, t B = 3, t c = 7, and t D = 10;

[0034] In the insertion operation, calculate the interaction time difference between adjacent items, t B -t A = 2, t c -t B = 4, t D -t c = 3, and it is found that t c -t B is the largest. Assume that through item similarity calculation, the sum of the similarities between item X and B, C is the largest and it is not in the original sequence. Insert X between B and C to obtain the time-series enhanced sequence S 插入 = [A, B, C, X, D]. This helps to supplement information at positions with large time differences, making the sequence more uniform and able to capture possible interest changes of the user over a long time interval. For example, when there is a long time interval between the purchase of a mobile phone and a mobile phone case, the user may also purchase a mobile phone film, and the insertion operation can simulate this potential behavior.

[0035] In the cropping operation, assume the cropping length l = 2, and crop sequences of length 2 from the original sequence with a step size of 1 to obtain [A, B], [B, C], [C, D]. Calculate their time standard deviations. Assume that the time standard deviation of [B, C] is the smallest, and select [B, C] for cropping S 裁剪 = [B, C]. The cropping operation can highlight the part of the sequence with relatively stable time intervals, retain key user interest patterns, remove possible noise information, and focus on the core interests of the user.

[0036] In the masking operation, from the previous calculation, it is known that t c -t B is the largest. Select one item to delete between B and C. Assume that B is deleted to obtain the masked S 掩码= [A, C, D]. The masking operation can simulate the situation of missing data, enabling the model to learn to capture the user's interest trend even when some information is missing, thus improving the robustness of the model.

[0037] In the rearrangement operation, calculate the time standard deviation of all subsequences of the original sequence. Suppose it is found that the time standard deviation of the subsequence [A, D] is the largest. Shuffle [A, D] to get [D, A], and obtain the rearrangement S 重排 = [D, A, B, C]. The rearrangement operation changes the order of elements while keeping the sequence elements unchanged, enabling the model to learn the features of the sequence under different orders, enhancing the model's understanding of the diversity of user interests, and discovering potential interest combinations.

[0038] S3: Input the time series enhanced sequence into the feature encoding module to obtain the feature representation of the time series enhanced sequence;

[0039] S4: Determine the data augmentation method adopted by the time series enhanced sequence according to the feature representation of the time series enhanced sequence, and construct the first cross-entropy loss function based on the discrimination result;

[0040] Preferably, construct a multi-layer fully connected neural network as the time series enhancement discriminator. The input of the time series enhancement discriminator is the feature representation of the time series enhanced sequence after data augmentation; the output of the time series enhancement discriminator is the data augmentation method adopted by the time series enhanced sequence.

[0041] In this embodiment, a multi-layer fully connected neural network is constructed as the time series enhancement discriminator. This discriminator has 3 hidden layers, containing 10, 8, and 6 neurons respectively, and the ReLU activation function is used. The output layer has 4 neurons (corresponding to the 4 data augmentation methods of insertion, cropping, masking, and rearrangement), and the Softmax activation function is used to output the probability distribution.

[0042] By constructing the time series enhancement discriminator and the first cross-entropy loss function, the feature extraction module can learn the feature differences of the sequence under different data augmentation methods. During the training process, the feature extraction module and the discriminator will continuously adjust the parameters according to this loss function, enabling the discriminator to more accurately judge the data augmentation method adopted by the time series enhanced sequence. At the same time, the feature extraction module can better understand the effect of data augmentation, thereby improving the adaptability of the feature extraction module to different change patterns and the ability to generate high-quality feature representations.

[0043] S5: Determine whether two different time series enhanced sequences adopt the same data augmentation method according to the feature representations of the two different time series enhanced sequences, and construct the second cross-entropy loss function based on the discrimination result;

[0044] Preferably, a multi-layer fully-connected neural network is constructed as a temporal stability discriminator, and the input of the temporal stability discriminator is the feature representations of two temporally augmented sequences after data augmentation; the output of the temporal stability discriminator is whether the two temporally augmented sequences adopt the same data augmentation method.

[0045] In this embodiment, a multi-layer fully-connected neural network is constructed as a temporal stability discriminator. This discriminator has 3 hidden layers, which contain 10, 8, and 6 neurons respectively, and the ReLU activation function is used. The output layer has 2 neurons (corresponding to two results of "yes" and "no"), and the Sigmoid activation function is used to output probability values. Input h1 and h2 into the discriminator. After calculation, the output probability value is p = 0.1, indicating that the discriminator believes that the probability that these two temporally augmented sequences adopt the same data augmentation method is 0.1, that is, it is more inclined to believe that they adopt different data augmentation methods; h1 and h2 represent the feature representations of two different temporally augmented sequences.

[0046] By constructing a temporal stability discriminator and a second cross-entropy loss function, the feature extraction module can learn the difference degree of sequence features under different data augmentation methods. During the training process, the feature extraction module and the temporal stability discriminator adjust parameters according to this loss function, so that the temporal stability discriminator can more accurately judge whether two temporally augmented sequences adopt the same data augmentation method. This helps to ensure that the generated embeddings are more consistent and stable under the same augmentation operation. For example, if multiple sequences are generated by the same insertion operation, the temporal stability discriminator should be able to accurately judge that they come from the same augmentation method, so that the feature extraction module can better capture this consistency during the learning process, improving the model's understanding and application ability of data augmentation operations, thereby improving the stability and accuracy of the recommendation system.

[0047] S6: Divide the original user-item interaction sequence into multiple different first subsequences, and use the first subsequences with the same target item as positive sample pairs, and the first subsequences with different target items as negative sample pairs; wherein, the target item of the first subsequence represents the last item interacted by the user in the first subsequence.

[0048] In this embodiment, assume that the original user-item interaction sequence is [Item A, Item B, Item C, Item D, Item E]. The following is the subsequence division and construction of positive and negative sample pairs according to the requirements. Set the subsequence length to 2 (which can be adjusted according to the actual situation), and divide it in the form of a sliding window with a step size of 1. In this way, the following first subsequences can be obtained: [Item A, Item B], [Item B, Item C], [Item C, Item D], [Item D, Item E].

[0049] Based on the target object, the first subsequences with the same target object are used as positive sample pairs. For example, the target objects of [Product A, Product B] and [Product B, Product C] are both Product C, and they can form a positive sample pair; the target objects of [Product C, Product D] and [Product D, Product E] are both Product E, and these two subsequences can also form a positive sample pair.

[0050] The first subsequences with different target objects are used as negative sample pairs. For example, the target object of [Product A, Product B] is Product B, and the target object of [Product C, Product D] is Product D, and these two subsequences can form a negative sample pair; [Product B, Product C] and [Product D, Product E] can also form a negative sample pair.

[0051] By constructing positive and negative sample pairs in this way, the model can learn the relationships between different subsequences during the training process, especially the differences between subsequences around the same target object and different target objects. This helps the model better understand the changes in the user's interests at different stages, as well as the differences and connections between different interests, thereby improving the accuracy and effectiveness of the recommendation system. For example, in a recommendation system, if the model finds many positive sample pairs around a certain type of product (such as electronic products) in the user's historical interaction sequences, then it can more accurately recommend other electronic products of the same type during the recommendation; at the same time, through the learning of negative sample pairs, the model can avoid recommending product categories that the user is clearly not interested in (such as food), improving the quality of the recommendation.

[0052] S7: Input the first subsequence into the feature encoding module to obtain the feature representation of the first subsequence, and construct the first contrastive loss function according to the feature representations of the positive and negative sample pairs;

[0053] In this embodiment, the first contrastive loss function uses the InfoNCE loss function.

[0054] S8: Use the first subsequences with the same target object as the second subsequences, cluster the second subsequences according to the feature representations of the second subsequences by the kmean algorithm, use the second subsequences under the same cluster as positive sample pairs, and construct the second contrastive loss function with the second subsequences under different clusters as negative sample pairs;

[0055] In this embodiment, it is assumed that the original user-item interaction sequence is [Item A, Item B, Item C, Item B, Item E, Item B, Item E]. The first subsequences [Item A, Item B], [Item B, Item C], [Item C, Item B], [Item B, Item E], [Item E, Item B], [Item D, Item E] have been obtained, and their feature representations have been obtained through the feature encoding module. Assume that the first subsequences with the same target object form the second subsequences [Item A, Item B], [Item C, Item B], [Item E, Item B]; cluster [Item A, Item B], [Item C, Item B], [Item E, Item B] using the kmean algorithm, set the number of clusters K to 2, and first randomly initialize two cluster centers c1 and c2; calculate the Euclidean distance between the feature representation of each second subsequence and the cluster center, divide the second subsequence into the cluster where the nearest cluster center is located, and update the cluster center. After multiple iterations, the clustering result is obtained. For example, Cluster 1: [Item A, Item B], [Item C, Item B]; Cluster 2 [Item E, Item B]. The second subsequences under the same cluster are used as positive sample pairs, and [Item A, Item B] and [Item C, Item B] are used as positive sample pairs; [Item A, Item B] and [Item E, Item B], [Item C, Item B] and [Item E, Item B] are used as two negative sample pairs; in this embodiment, the second contrast loss function uses the InfoNCE loss function.

[0056] Through K-Means clustering, the model can discover that even if the target objects are the same, there are differences in the user's behavior patterns. For example, in the above example, although the target objects of some second subsequences are all Item B or Item C, they can be divided into different groups through clustering, indicating that there are different interest preferences or behavior patterns when users purchase these items. The constructed positive and negative sample pairs and the second contrast loss function enable the model to learn the differences between different clusters and the similarities within the same cluster. The negative sample pairs under different clusters enable the model to learn the boundaries between different behavior patterns, prevent the model from only focusing on a certain behavior pattern, and thus enhance the generalization ability of the model when facing various user behavior data and improve the recommendation performance of the model in different scenarios.

[0057] S9: Use the enhanced sequences under the same user as positive sample pairs and the enhanced sequences under different users as negative sample pairs to construct the third contrast loss function;

[0058] In this embodiment, the third contrast loss function uses the InfoNCE loss function. By constructing the third contrast loss function, the model can learn the similarities between different enhanced sequences of the same user and the differences between enhanced sequences of different users. In the recommendation system, it helps the model capture the personalized behavior patterns of users and enhance the understanding of user interests.

[0059] S10: Construct a total loss function based on the first cross-entropy loss function, the second cross-entropy loss function, the first contrastive loss function, the second contrastive loss function, and the third contrastive loss function to jointly train the feature encoding module. After the training is completed, fix the parameters of the feature encoding module, and fine-tune the prediction module on the labels of the original user-item interaction sequence to obtain a trained recommendation model.

[0060] Preferably, the total loss function includes:

[0061] L = λ1·L c1 + λ2·L c2 + β1·L T1 + β2·L T2 + β3·L T3

[0062] where L represents the total loss function; λ1 and λ2 represent adjustable weight parameters; β1, β2, and β3 represent adjustable weight parameters; L c1 represents the first cross-entropy loss function; L c2 represents the second cross-entropy loss function; L T1 represents the first contrastive loss function; L T2 represents the second contrastive loss function; L T3 represents the third contrastive loss function. Use the Adam optimizer and set the learning rate to 0.001. In each training iteration, update the parameters of the feature encoding module according to the calculated gradient, and continuously adjust the weights of the feature encoding module to minimize the total loss function. When the training of the feature encoding module is completed, fix its parameters. Then fine-tune the prediction module (fully connected neural network + softmax) on the labels of the original user-item interaction sequence. By constructing the total loss function to jointly train the feature encoding module, various factors (such as the discrimination of data augmentation methods, the contrast between subsequences, etc.) can be comprehensively considered, enabling the feature encoding module to learn richer and more accurate feature representations. After fixing the parameters of the feature encoding module and fine-tuning the prediction module, the prediction module can better adapt to specific tasks and data labels, improving the prediction accuracy.

[0063] In this embodiment, the calculated total loss function L is used to jointly train the feature encoding module. During the training process, calculate the gradient of L with respect to the parameters of the feature encoding module through the backpropagation algorithm, and then use an optimizer (such as stochastic gradient descent, Adam, etc.) to update the parameters of the feature encoding module, making L gradually decrease.

[0064] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0065] In summary, in terms of data processing, the present invention obtains the original user-item interaction sequence and interaction time data, and uses a variety of data augmentation methods to generate multiple time-series augmented sequences, alleviating the problem of data sparsity and enriching the training data. During the model training process, multiple loss functions are constructed, including a first cross-entropy loss function and a second cross-entropy loss function constructed based on the feature representations of the time-series augmented sequences, and multiple contrastive loss functions constructed based on subsequences and augmented sequences. These loss functions constrain the model from different perspectives, comprehensively improving the model performance. By discriminating the data augmentation methods used in the time-series augmented sequences and whether different time-series augmented sequences use the same augmentation method to construct the loss function, it helps the model better learn the features under different augmentation methods and improve the model's understanding ability of the data. Dividing the original user-item interaction sequence into subsequences and constructing corresponding contrastive loss functions can guide the model to learn the user behavior patterns from the multi-granularity intention level and capture the true intentions of the users. At the same time, by constructing contrastive loss functions for the augmented sequences under the same user and different users, the potential correlations between users are fully explored, and the model's understanding and generalization ability of complex user intentions are improved. Finally, by constructing a total loss function to jointly train the feature encoding module and fine-tuning the prediction module, a trained recommendation model is obtained, effectively improving the accuracy and reliability of the recommendation system and providing users with higher-quality and more accurate recommendation services.

[0066] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A method for adversarial generation recommendation with temporal enhancement and multi-granularity intention guidance, characterized in that Including: S1: Obtain the original user-item interaction sequence for training and the time data of the user's interaction with the item; S2: Use various data augmentation methods to augment the original user-item interaction sequence according to the interaction time between the user and the item to obtain multiple time-series augmented sequences; S3: Input the time-series augmented sequences into the feature encoding module to obtain the feature representations of the time-series augmented sequences; S4: Determine the data augmentation method used for the time-series augmented sequence according to the feature representation of the time-series augmented sequence, and construct the first cross-entropy loss function based on the discrimination result; S5: Determine whether two different time-series augmented sequences use the same data augmentation method according to the feature representations of the two different time-series augmented sequences, and construct the second cross-entropy loss function based on the discrimination result; S6: Divide the original user-item interaction sequence into multiple different first subsequences, and use the first subsequences of the same target object as positive sample pairs and the first subsequences of different target objects as negative sample pairs; wherein, the target object of the first subsequence represents the last item interacted by the user in the first subsequence; S7: Input the first subsequences into the feature encoding module to obtain the feature representations of the first subsequences, and construct the first contrastive loss function according to the feature representations of the positive sample pairs and the negative sample pairs; S8: Use the first subsequences of the same target object as the second subsequences, cluster the second subsequences by the kmean algorithm according to the feature representations of the second subsequences, and use the second subsequences under the same cluster as positive sample pairs and the second subsequences under different clusters as negative sample pairs to construct the second contrastive loss function; S9: Use the augmented sequences under the same user as positive sample pairs and the augmented sequences under different users as negative sample pairs to construct the third contrastive loss function; S10: Construct a total loss function according to the first cross-entropy loss function, the second cross-entropy loss function, the first contrastive loss function, the second contrastive loss function and the third contrastive loss function to jointly train the feature encoding module. After the training is completed, fix the parameters of the feature encoding module and fine-tune the prediction module on the labels of the original user-item interaction sequence to obtain the trained recommendation model.

2. The adversarial generation recommendation method with time series enhancement and multi-granularity intention guidance according to claim 1, characterized in that The feature extraction module includes a transformer encoder, and the prediction module includes a fully connected neural network + softmax.

3. A method for adversarial generation recommendation with temporal enhancement and multi-granularity intention guidance according to claim 1, wherein The data augmentation methods include: insertion operation, cropping operation, masking operation and rearrangement operation; The insertion operation includes: calculating the interaction time difference between adjacent items according to the interaction time between the user and the item in the original user-item interaction sequence, inserting the item with the largest sum of similarities to the two adjacent items between the two adjacent items with the largest interaction time difference, and the inserted item is not in the original user-item interaction sequence to obtain a time-series augmented sequence; The cropping operation includes: setting the cropping length l, based on the principle of a sliding window, cropping a sequence of length l from the original user-item interaction sequence with a step size of 1, and calculating the time standard deviation of the sequence, and selecting the sequence with the smallest time standard deviation among the cropped sequences as the time-series augmented sequence; The masking operation includes: calculating the interaction time difference between adjacent items according to the interaction time between the user and the item in the original user-item interaction sequence, and selecting one item to delete between the two adjacent items with the largest interaction time difference to obtain a time-series enhanced sequence; The rearrangement operation includes: screening out the subsequence with the largest time standard deviation from all subsequences of the original user-item interaction sequence and shuffling the subsequence to obtain a time-series enhanced sequence.

4. A method for adversarial generation recommendation with temporal enhancement and multi-granularity intention guidance according to claim 1, characterized in that, Construct a multi-layer fully connected neural network as a time-series enhancement discriminator, and the input of the time-series enhancement discriminator is the feature representation of the time-series enhanced sequence after data enhancement; The output of the time-series enhancement discriminator is the data enhancement method adopted by the time-series enhanced sequence.

5. A method for adversarial generation recommendation with temporal enhancement and multi-granularity intention guidance according to claim 1, characterized in that Construct a multi-layer fully connected neural network as a time-series stability discriminator, and the input of the time-series stability discriminator is the feature representations of two time-series enhanced sequences after data enhancement; The output of the time-series stability discriminator is whether the two time-series enhanced sequences adopt the same data enhancement method.

6. The adversarial generation recommendation method with time series enhancement and multi-granularity intention guidance according to claim 1, wherein The total loss function includes: L = λ1·L c1 + λ2·L c2 + β1·L T1 + β2·L T2 + β3·L T3 Among them, \(L\) represents the total loss function; \(\lambda_1\) and \(\lambda_2\) represent adjustable weight parameters; \(\beta_1\), \(\beta_2\) and \(\beta_3\) represent adjustable weight parameters; \(L\) c1 represents the first cross-entropy loss function; \(L\) c2 represents the second cross-entropy loss function; \(L\) T1 represents the first contrastive loss function; \(L\) T2 represents the second contrastive loss function; \(L\) T3 represents the third contrastive loss function.

7. A time-series enhanced and multi-granularity intention-guided adversarial generation recommendation system, characterized in that, The system includes a memory and a processor; the memory is used to store the application program; the processor is used to run the application program and execute a time-series enhancement and multi-granularity intention-guided adversarial generation recommendation method according to any one of claims 1 to 6.

8. A computer storage medium, characterized in that, A remote monitoring program is stored on the computer storage medium, and when the remote monitoring program is executed by the processor, it implements a time-series enhancement and multi-granularity intention-guided adversarial generation recommendation method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Cross-sequence mixed sequence recommendation method based on user correlation guidance

    CN121705519A