Sequence recommendation method based on Siamese generative adversarial network and uncertainty adversarial memory
By introducing the Siamese network and uncertainty adversarial memory mechanism into the generative adversarial network, the "hard boundary" classification and noise sensitivity problems are solved, more accurate sequence recommendations are achieved, and the long sequence dependency modeling capabilities and recommendation performance are improved.
Patent Information
- Application Number
- CN202510703737.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-05-29
AI Technical Summary
Existing generative adversarial networks have feature space constraints caused by "hard boundary" classification in sequential recommendation systems, and Transformer-based sequence modeling methods are sensitive to uncertainty and noise and lack the ability to model long sequence dependencies.
The adversarial network framework of the Siamese network and the uncertainty adversarial memory mechanism are integrated to optimize the learning process of the generator and discriminator through similarity loss, introduce penalty correction terms and Gaussian noise into the attention mechanism, and combine with the external memory mechanism to model long-range dependencies.
It improves the learning coordination between the generator and the discriminator, suppresses high uncertainty and noise interference, enhances the ability to model long-term dependencies on user behavior, and improves the accuracy and recall of recommendation results.
Smart Images

Figure BDA0005425197840000041 
Figure BDA0005425197840000046 
Figure BDA0005425197840000049
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning recommendation systems, and relates to the technical field of adversarial learning optimization frameworks and deep sequence modeling. Background Art
[0002] As a core technology for information filtering and personalized services, recommendation systems have been widely used in e-commerce, social media, and streaming media. Traditional methods are based on static user preference modeling and have difficulty capturing dynamic behavioral characteristics. This is especially true in scenarios where user interests drift in real time (such as sudden scenario demands) or behaviors contain implicit temporal logic (such as the strongly associated consumption chain of "mobile phone → screen protector → earphones"). This can easily lead to misalignment between recommendation results and true intentions. For this reason, sequential recommendation systems, as an important branch of the recommendation field, have received increasing attention. Unlike traditional recommendation methods, sequential recommendation captures the temporal sequence of user interaction behaviors, explores potential needs from users' continuous operations, and can more accurately predict users' future behaviors. This method models based on the user's historical behavior sequence. It not only focuses on single behaviors or static preferences, but also can identify dynamic changes in user interests, providing users with more contextually relevant recommendations.
[0003] Early research relied on Markov chains and matrix factorization to characterize short-range dependencies and global associations. Subsequently, deep learning models based on CNNs and RNNs significantly improved feature extraction capabilities by capturing local patterns and modeling long-range dependencies. The Transformer further optimized behavioral association modeling by dynamically weighting key behavioral nodes through a self-attention mechanism (e.g., identifying short-term peaks in demand for flowers during Valentine's Day). However, existing scaled dot-product attention mechanisms are sensitive to uncertainty and noisy data, resulting in insufficient robustness in representation. Furthermore, engineered window truncation strategies limit their ability to model long-sequence dependencies. Specifically, the high uncertainty of user behavior sequences causes the traditional dot-product attention mechanism to confuse occasional noise with true preferences due to deterministic vector compression (e.g., misclassifying holiday purchases as daily needs). Furthermore, attention's oversensitivity to noisy interactions can easily amplify low-quality feature associations (e.g., a bread lover mistakenly clicking on the refrigerator triggers irrelevant product recommendations). While window truncation improves computational efficiency, it sacrifices the periodic recurrence and gradual evolution of user behavior, resulting in inadequate modeling of long-range behavioral patterns.
[0004] In recent years, generative adversarial networks (GANs) have been used to generate high-quality interaction sequences through adversarial training. Previous research has restructured sequential recommendation into a discriminator adversarial training framework. However, the discriminator in traditional GANs typically employs a "hard boundary" classification strategy, assigning label 1 to real samples and label 0 to generated samples. However, this binary classification strategy forces the discriminator to construct a rigid classification boundary in feature space, resulting in samples near the boundary being misclassified due to insufficient separability. The generator relies on feedback from the discriminator for parameter updates, but this "hard boundary" classification only measures the degree of match between generated samples and real samples, failing to provide fine-grained supervision of local patterns within the sequence, such as item associations and temporal dependencies.
[0005] To address the above problems and challenges, this paper proposes a method to enhance sequential recommendation by integrating an adversarial network framework with a Siamese network and an uncertainty adversarial memory mechanism. Summary of the Invention
[0006] The purpose of this invention is to solve the problem of feature space constraints caused by the "hard boundary" classification of existing generative adversarial networks in sequence recommendation systems, as well as the shortcomings of Transformer-based sequence modeling methods such as sensitivity to uncertainty and noise and insufficient ability to model long sequence dependencies.
[0007] To achieve the above objectives, this application provides the following solutions:
[0008] S1. Obtain a historical dataset of interactions between users and items, and analyze the application of generative adversarial networks to sequential recommendation problems.
[0009] S2. Build an enhanced adversarial network framework that integrates a Siamese network. The generator receives the masked user interaction sequence to generate candidate sequences, and the discriminator achieves accurate classification.
[0010] S3. Input the real samples and generated samples into the Siamese network and map them into the same feature space;
[0011] S4. Calculate the similarity loss between real samples and generated samples, and optimize the learning process of the generator and discriminator through the similarity loss;
[0012] S5. After the real sequence is encoded by the encoder, it is combined with dynamic random enhancement and contrastive learning techniques to model sequence features from multiple perspectives.
[0013] S6. An uncertainty adversarial memory mechanism is proposed, which integrates query and key penalty correction terms and adversarial Gaussian noise into the attention mechanism to dynamically suppress the interference of high uncertainty and noise features, and uses the memory slots of the external memory mechanism to model the long-term dependencies of user behavior.
[0014] S7. A multi-task training strategy is adopted to jointly optimize the contrastive learning task, generation task, discrimination task, and next item prediction task.
[0015] S8. Generate a list of recommended items for each user based on the user-item interaction probability finally output by the model.
[0016] S9. This paper selects three datasets from different real-world scenarios to test and evaluate the recommendation performance of the model. This paper uses the recall rate (Recall@20) and the normalized discounted cumulative gain (NDCG@20) of the top 20 items in the results as quantitative indicators of recommendation performance.
[0017] In summary, compared with the existing technology, the present invention has the following advantages:
[0018] (1) This paper refines the mapping relationship between the real sequence and the generated sequence in the feature space by integrating the Siamese network into the generative adversarial network, and optimizes the learning process of the generator and discriminator through similarity loss. This is an improved way to gradually transition from "hard classification" to "soft classification".
[0019] (2) Compared with the existing methods, the uncertainty adversarial memory mechanism designed in this paper makes two improvements based on Transformer: the uncertainty-aware adversarial attention mechanism incorporates query and key penalty correction terms and Gaussian noise to optimize the attention weight distribution; the external memory mechanism realizes continuous modeling of users' long-term behavior patterns through a learnable memory slot matrix. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 This is a recommended model structure diagram of the present invention;
[0021] Figure 2 Schematic diagram of the discriminator's "hard boundary" classification;
[0022] Figure 3 For Siamese network architecture;
[0023] Figure 4 To provide uncertainty adversarial memory mechanism, the uncertainty-aware adversarial attention mechanism and external memory mechanism are integrated;
[0024] Figure 5 Statistics of three real-scene data sets obtained by the present invention;
[0025] Figure 6 The figure shows the comparison of the overall performance of the present invention with different baselines in three real scenarios; DETAILED DESCRIPTION
[0026] The present invention will be further described below with reference to the accompanying drawings and examples, which are for illustrative purposes only and are not to be construed as limiting the present invention.
[0027] The present invention discloses a method for enhancing adversarial sequence recommendation by integrating Siamese network, and its model structure is shown in the figure below: Figure 1 shown.
[0028] S1. Obtain a historical dataset of interactions between users and items, and analyze the application of generative adversarial networks to sequential recommendation problems.
[0029] In specific implementation, S1 includes the following steps:
[0030] S101. The present invention collects three datasets in different real-world scenarios, namely the movie recommendation dataset ML-1M, the commercial venue recommendation dataset Yelp, and the sports event recommendation dataset Sports.
[0031] S102. Figure 2 As shown, the present invention visualizes a schematic diagram of the discriminator's "hard boundary" classification strategy, thereby showing that the generative adversarial network has a significant rigid classification boundary in sequence recommendation data, which causes samples near the boundary to be misclassified due to insufficient separability.
[0032] S2. Construct an enhanced adversarial network framework that integrates the Siamese network. The generator receives the masked user interaction sequence to generate candidate sequences, and the discriminator achieves accurate classification.
[0033] In specific implementation, S2 includes the following steps:
[0034] S201. The present invention adopts the BERT4Rec model based on bidirectional Transformer as the generator G. Given a user interaction sequence S u , we first hide some items in these sequences through masking operations. Specifically, we randomly mask several items according to the predefined mask ratio γ to obtain the partially masked interaction sequence S' u The generator G is based on S' u To restore the original sequence, we get the generated sequence S' u '=G(S' u ), the loss function of the generator is:
[0035]
[0036] in, is the set of mask items of user u, is the real item label that masks the location, Represents the generator given a mask sequence S' u When the mask position v is predictedm is a real term probability.
[0037] S202. The present invention uses the SASRec model based on the self-attention mechanism as the discriminator D. The discriminator accepts the real sequence and the generated sequence processed by the aggregation function as input and distinguishes them. Its loss function is:
[0038]
[0039] Among them, T is the maximum time step of the sequence, is the item that generates the sequence, 1(·) is the indicator function, σ is the Sigmoid function; w is the learnable parameter matrix; f D represents the discriminator network; Represents the historical information of user u at time t.
[0040] S3. Input the real samples and generated samples into the Siamese network and map them into the same feature space;
[0041] The present invention innovatively embeds the Siamese network into the GAN framework, and achieves joint optimization of generation and identification by measuring the feature similarity between the generated sequence and the real sequence in the projection space. The Siamese network involved in the present invention adopts a two-layer fully connected structure to achieve dimensionality reduction mapping of the feature space. The schematic diagram of the architecture is shown in the figure. Figure 3 Specifically, the generated sequence G(S u ) and the true sequence S u Input them into the two-branch network with shared weights to obtain the corresponding feature representation:
[0042] f gen =Sn(G(S u ))
[0043] f real =Sn(S u )
[0044] Here, Sn(·) represents a Siamese network.
[0045] S4. Calculate the similarity loss between real samples and generated samples, and optimize the learning process of the generator and discriminator through the similarity loss;
[0046] In specific implementation, S4 includes the following steps:
[0047] S401. In order to quantify the feature similarity between two types of sequences, the cosine similarity metric is used:
[0048] Sim_p=cos(f real ,f gen )
[0049] S402. Finally, the generator’s loss function consists of the standard cross entropy loss and similarity loss:
[0050]
[0051] The loss function of the discriminator consists of the standard cross entropy loss and the similarity loss:
[0052]
[0053] S5. After the real sequence is encoded by the encoder, it is combined with dynamic random enhancement and contrastive learning techniques to model sequence features from multiple perspectives.
[0054] In specific implementation, S5 includes the following steps:
[0055] S501. To enhance the robustness of sequence representation, the present invention proposes a joint optimization framework that integrates dynamic data enhancement and contrastive learning. Map it to the latent space through the encoder to obtain the potential representation of each time step:
[0056]
[0057] Among them, f θ is the encoder function and θ is the model parameter.
[0058] S502. Obtaining the potential representation of each time step After that, the potential representation of the entire sequence can be aggregated to obtain the final potential representation h of the user behavior sequence u Here, the average value of the last k time steps is used, as follows:
[0059]
[0060] Different from the traditional single data augmentation method, we adopt a dynamic random data augmentation strategy for the latent representation. During each training, the enhanced samples are generated by this mechanism. and compare it with the original sample h u The goal of contrastive learning is to enhance the similarity between the sample and the original sample by comparison, so that the representations of the two are as close as possible, while being far away from other negative samples. To this end, we introduce the contrastive loss function:
[0061]
[0062] where sim(·,·) represents the cosine similarity function, τ is the temperature coefficient that controls the smoothness of the similarity, and M(u) represents the set of users in the mini-batch containing user u.
[0063] S6. Incorporate query and key penalty correction terms and adversarial Gaussian noise into the attention mechanism to dynamically suppress the interference of high uncertainty and noise features, and use the memory slots of the external memory mechanism to model the long-range dependencies of user behavior.
[0064] This step is the uncertainty-versus-memory mechanism, such as Figure 4 In specific implementation, S6 includes the following steps:
[0065] S601. Uncertainty-aware adversarial memory mechanism. The traditional attention mechanism only calculates the similarity between the basic query and the key, while the present invention introduces a penalty correction term in the attention score calculation to reduce the attention weight of high uncertainty features. Specifically, given an input sequence tensor X∈R B×L×d (B is the batch size, L is the sequence length, and d is the hidden layer dimension). The raw attention score is calculated as the base relevance score minus the uncertainty penalty:
[0066]
[0067] To further improve the robustness of the model, we add Gaussian noise ∈ ~N(0,σ 2 ), generate the perturbed attention score:
[0068]
[0069] in, and are the query and key after random perturbation respectively.
[0070] S602. External memory mechanism. The model extracts long-term interest representations related to the current behavior through the attention interaction between the query vector and the memory matrix. Specifically, given the current query vector Calculate its memory matrix (N is the number of memory slots, d m is the associated weight of the memory vector dimension), and then use this weight to associate all memory slots M in the memory matrix i Perform a weighted sum:
[0071] α=softmax(query·W r )
[0072]
[0073] In order to dynamically store new features in the sequence, an update strategy based on the gating mechanism is designed. Concatenate it with the old memory to generate a candidate memory:
[0074]
[0075] By updating the gate weight g u ∈[0,1] controls the mixing ratio of new and old memories:
[0076]
[0077] Finally, smooth updates are achieved through gated weighting to avoid memory conflicts:
[0078]
[0079] S7. A multi-task training strategy is adopted to jointly optimize the contrastive learning task, generation task, discrimination task, and next item prediction task.
[0080] To optimize the sequential recommendation model, we adopt a multi-task training strategy by jointly optimizing the contrastive learning task, the generation task, the discrimination task, and the next item prediction task. We define the following total loss function:
[0081] L=L rec +λ3L cl +λ4L gen +λ5L disc
[0082] By properly adjusting the weight coefficients λ3, λ4, and λ5, the contribution of each task to the total loss can be flexibly controlled. Ultimately, this multi-task learning strategy will effectively promote the coordinated optimization of the model across multiple objectives.
[0083] S8. Generate a list of recommended items for each user based on the user-item interaction probability finally output by the model.
[0084] In a specific implementation, the goal of the top-k recommendation task studied in the present invention is to find k uninteracted items that are most likely to be of interest to the user and generate a ranked recommendation list.
[0085] S9. This paper selects three datasets of different real scenarios to test and evaluate the recommendation performance of the model. Figure 5 The statistical data of the dataset are listed. This paper uses the recall rate (Recall@20) and normalized discounted cumulative gain (NDCG@20) of the top 20 items in the results as quantitative indicators of recommendation performance.
[0086] The present invention evaluates the overall performance of the recommendation model. Figure 6As shown, the overall performance of the present invention is compared with different baselines in three different real scenarios. The comparison results show that the present invention is always better than other baselines in different data sets and evaluation indicators. Specifically, compared with the suboptimal model, the present invention improves the Recall@10 index by 11.3% and 5.8% on the Yelp dataset and ML-1M dataset respectively; and improves the Recall@20 index by 4.9% on the Sports dataset. These performance improvements are due to the following core innovations: First, by introducing the twin network architecture and the similarity loss function into the generative adversarial network framework, the "hard boundary" decision bias problem caused by the traditional discriminator binary classification is innovatively solved, and the collaborative optimization of the generator and the discriminator is achieved. Secondly, the present invention applies the uncertainty adversarial memory mechanism to the encoder, generator and discriminator, which effectively suppresses high uncertainty and noise interference and captures long-term dependency patterns.
[0087] This paper proposes a method for enhanced adversarial sequence recommendation based on Siamese networks, aiming to address the limitations of the "hard boundary" classification strategy in traditional generative adversarial networks, as well as the shortcomings of Transformer in uncertainty and noise robustness and long sequence modeling. Specifically, by introducing the Siamese network into the generative adversarial network, the similarity mapping between the real sequence and the generated sequence in the feature space is refined, and the learning process of the generator and discriminator is optimized with the help of similarity loss, thereby generating sequence recommendation results that are more in line with user preferences. In addition, the design of the uncertainty adversarial memory mechanism can effectively suppress highly uncertain sequence data, reduce noise interference, and achieve long-range dependency modeling. In summary, the present invention not only makes up for the shortcomings of existing generative adversarial networks, but also provides new ideas for improving the sequence modeling capabilities of the Transformer architecture. In the future, we will focus on enhancing the interpretability of the present invention to provide better transparency and insight into the recommendation generation process, which will further enhance the credibility and acceptability of the system.
[0088] Finally, the details of the above embodiments of the present invention are merely examples for explaining the present invention. For those skilled in the art, any modifications, improvements and replacements of the above embodiments should be included in the scope of protection of the claims of the present invention.
Claims
1. A sequence recommendation method based on Siamese generative adversarial networks and uncertainty adversarial memory, the method comprising the following steps: S1. Obtain a historical dataset of interactions between users and items, and analyze the application of generative adversarial networks to sequential recommendation problems. S2. Build an enhanced adversarial network framework that integrates a Siamese network. The generator receives the masked user interaction sequence to generate candidate sequences, and the discriminator achieves accurate classification. S3. Input the real samples and generated samples into the Siamese network and map them into the same feature space; S4. Calculate the similarity loss between real samples and generated samples, and optimize the learning process of the generator and discriminator through the similarity loss; S5. After the real sequence is encoded by the encoder, it is combined with dynamic random enhancement and contrastive learning techniques to model sequence features from multiple perspectives. S6. An uncertainty adversarial memory mechanism is proposed, which integrates query and key penalty correction terms and adversarial Gaussian noise into the attention mechanism to dynamically suppress the interference of high uncertainty and noise features, and uses the memory slots of the external memory mechanism to model the long-term dependencies of user behavior. S7. A multi-task training strategy is adopted to jointly optimize the contrastive learning task, generation task, discrimination task, and next item prediction task. S8. Generate a list of recommended items for each user based on the user-item interaction probability finally output by the model. S9. This paper selects three datasets from different real-world scenarios to test and evaluate the recommendation performance of the model. This paper uses the recall rate (Recall@20) and the normalized discounted cumulative gain (NDCG@20) of the top 20 items in the results as quantitative indicators of recommendation performance.
2. According to the construction of the enhanced adversarial network framework of the fused Siamese network in claim 1, the specific process of S2 is: S201. The present invention adopts the BERT4Rec model based on bidirectional Transformer as the generator G. Given a user interaction sequence S u ,We first hide some items in these sequences through masking operations. Specifically, several items are randomly masked according to the predefined mask ratio γ to obtain the partially masked interaction sequence S ' u The generator G is based on S ' u To restore the original sequence, we get the generated sequence S ' u ' =G(S ' u ), the loss function of the generator is: in, is the set of mask items of user u, is the real item label that masks the location, Represents the generator given a mask sequence S' u When the mask position v is predicted m is a real term probability. S202. The present invention uses the SASRec model based on the self-attention mechanism as the discriminator D. The discriminator accepts the real sequence and the generated sequence processed by the aggregation function as input and distinguishes them. Its loss function is: Among them, T is the maximum time step of the sequence, is the item that generates the sequence, 1(·) is the indicator function, σ is the Sigmoid function; w is the learnable parameter matrix; f D represents the discriminator network; Represents the historical information of user u at time t.
3. The construction of the Siamese network according to claim 1, wherein the specific process of S3 is as follows: The present invention innovatively embeds the Siamese network into the GAN framework, and realizes the joint optimization of generation and identification by measuring the feature similarity between the generated sequence and the real sequence in the projection space. The Siamese network involved in the present invention adopts a two-layer fully connected structure to realize the dimensionality reduction mapping of the feature space, and the schematic diagram of the architecture is shown in Figure 3. Specifically, the generated sequence G(S u ) and the true sequence S u Input them into the two-branch network with shared weights to obtain the corresponding feature representation: f gen =Sn(G(S u )) f real =Sn(S u ) in, Sn(·) represents a Siamese network.
4. The construction of contrastive learning according to claim 1, wherein the specific process of S5 is: S501. To enhance the robustness of sequence representation, the present invention proposes a joint optimization framework that integrates dynamic data enhancement and contrastive learning. Map it to the latent space through the encoder to obtain the potential representation of each time step: in, f θ is the encoder function and θ is the model parameter. S502. Obtaining the potential representation of each time step After that, the potential representation of the entire sequence can be aggregated to obtain the final potential representation h of the user behavior sequence u Here, the average value of the last k time steps is used, as follows: Different from the traditional single data augmentation method, we adopt a dynamic random data augmentation strategy for the latent representation. During each training, the enhanced samples are generated by this mechanism. and compare it with the original sample h u The goal of contrastive learning is to enhance the similarity between the sample and the original sample by comparison, so that the representations of the two are as close as possible, while being far away from other negative samples. To this end, we introduce the contrastive loss function: where sim(·,·) represents the cosine similarity function, τ is the temperature coefficient that controls the smoothness of the similarity, and M(u) represents the set of users in the mini-batch containing user u.
5. According to the construction of the uncertainty counter-memory mechanism of claim 1, the specific process of S6 is: S601. Uncertainty-aware adversarial memory mechanism. The traditional attention mechanism only calculates the similarity between the basic query and the key, while the present invention introduces a penalty correction term in the attention score calculation to reduce the attention weight of high uncertainty features. Specifically, given an input sequence tensor X∈R B×L×d (B is the batch size, L is the sequence length, and d is the hidden layer dimension). The raw attention score is calculated as the base relevance score minus the uncertainty penalty: To further improve the robustness of the model, we add Gaussian noise ∈ ~N(0,σ 2 ), generate the perturbed attention score: in, and are the query and key after random perturbation respectively. S602. External memory mechanism. The model extracts long-term interest representations related to the current behavior through the attention interaction between the query vector and the memory matrix. Specifically, given the current query vector Calculate its memory matrix (N is the number of memory slots, d m is the associated weight of the memory vector dimension), and then use this weight to associate all memory slots M in the memory matrix i Perform a weighted sum: α=softmax(query·W r ) In order to dynamically store new features in the sequence, an update strategy based on the gating mechanism is designed. Concatenate it with the old memory to generate a candidate memory: By updating the gate weight g u ∈[0,1] controls the mixing ratio of new and old memories: Finally, smooth updates are achieved through gated weighting to avoid memory conflicts:
6. The multi-task training strategy according to claim 1, wherein the specific process of S7 is as follows: To optimize the sequential recommendation model, we adopt a multi-task training strategy by jointly optimizing the contrastive learning task, the generation task, the discrimination task, and the next item prediction task. We define the following total loss function: L=L rec +λ3L cl +λ4L gen +λ5L disc By properly adjusting the weight coefficients λ3, λ4, and λ5, the contribution of each task to the total loss can be flexibly controlled. Ultimately, this multi-task learning strategy will effectively promote the coordinated optimization of the model across multiple objectives.
Citation Information
Patent Citations
Recommendation algorithm based on adversarial learning and bidirectional long-short-term memory network
CN112035745A
Fair and controllable image generation method based on structured network priori knowledge
CN116109719A
System and method for detecting anomalies in images
US20220076053A1
Using generative adversarial networks to construct realistic counterfactual explanations for machine learning models
US20220188645A1
Generative adversarial multi-head attention neural network self-learning method for aero-engine data reconstruction
WO2024087129A1