Frequency domain enhanced filtering method and system for sequence recommendation
By using frequency domain enhanced filtering methods in sequence recommendation, including frequency occlusion, mixing and frequency domain regularization, the problem of comparative learning dependence overweight and low negative sample quality in the prior art is solved, and the robustness and recommendation performance of the model are improved.
Patent Information
- Application Number
- CN202510256737.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-20
AI Technical Summary
The existing sequence recommendation methods rely too much on contrast learning, resulting in limited frequency domain information mining, low negative sample quality, weakening model discrimination ability, and the mixed attention mechanism increases model complexity.
A frequency domain enhanced filtering method for sequence recommendation is proposed. By acquiring the dynamic preferences of user historical interactions, frequency masking and mixing strategies are processed, and the encoder model is constructed using multiple learnable filter blocks stacks, combining frequency trapezoidal structure and frequency domain regularization, reducing dependence on negative samples.
The model's fine-grained characterization ability of user behavior characteristics is improved, the ability to capture different frequency modes is enhanced, noise interference is reduced, and the model's robustness and discrimination ability is improved.
Smart Images

Figure CN120179901A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of sequence recommendation prediction, and in particular relates to a frequency domain enhanced filtering method and system for sequence recommendation. Background Art
[0002] In today's era of the Internet of Everything, as the scale of data in real life continues to grow, recommender systems (RS) have become one of the effective ways to deal with the problem of information overload. Among them, the classic sequential recommendation (SRS) algorithm has been widely used. In the real world, users' shopping behaviors usually occur continuously in sequence rather than in isolation, and users' preferences and the popularity of goods also change over time. Therefore, it is necessary to use previous sequential interactions as context to predict items that may be interacted with in the future.
[0003] Current methods rely too much on contrastive learning to extract frequency domain information, and the goal of contrastive learning is to maximize the similarity between positive samples and minimize the similarity between negative samples. Experiments have found that the frequency domain information mined using the similarity between self-supervised signals is relatively limited, and low-quality negative samples may weaken the model's discriminative ability. Secondly, the hybrid attention mechanism also significantly increases the complexity of the model. Summary of the invention
[0004] In order to solve the above technical problems, the present invention provides a frequency domain enhanced filtering method for sequence recommendation, comprising:
[0005] Acquire dynamic preferences in historical user interactions, encode the dynamic preferences in historical user interactions, and obtain a user behavior sequence;
[0006] Performing frequency masking and mixing strategy processing on the user behavior sequence to obtain processed sequence data;
[0007] Building an encoder model based on a stack of multiple learnable filter blocks, and training the encoder model to obtain a prediction model;
[0008] The processed sequence data is input into the prediction model for calculation to obtain the prediction result.
[0009] Preferably, the process of obtaining the user behavior sequence includes:
[0010] Construct an item embedding matrix based on the dynamic preferences in the user's historical interactions, and project the high-dimensional one-hot representation of the item embedding matrix into a low-dimensional dense representation.
[0011] For the user's historical interaction sequence, the embedding matrix is used to find the embedding vector of each item to form the input embedding matrix;
[0012] Introduce a learnable position embedding matrix, add the input embedding matrix to the position embedding matrix to obtain the user behavior sequence.
[0013] Preferably, the process of obtaining the processed sequence data includes:
[0014] Perform a fast Fourier transform on each feature dimension of the user behavior sequence to obtain a frequency-domain sequence;
[0015] Randomly shuffle the order of user interaction items to obtain a new embedding matrix;
[0016] Perform Fourier transforms on the frequency-domain sequence and the new embedding matrix respectively, retain the high-amplitude frequency components, and at the same time exchange the positions of the low-amplitude components to obtain a processed frequency-domain signal;
[0017] Convert the processed frequency-domain signal back to the time domain through a fast Fourier transform, and optimize the signals in different frequency bands using a frequency trapezoidal structure to obtain the processed sequence data.
[0018] Preferably, the process of obtaining the prediction model includes:
[0019] Stack L filter blocks and construct a pointwise feed-forward network;
[0020] Construct an encoder model based on the stacked L filter blocks and the pointwise feed-forward network;
[0021] Train the encoder model based on a training set to obtain the prediction model.
[0022] Preferably, the operation process of the filter block includes:
[0023] Perform a filtering operation on the features of each dimension in the frequency domain, modulate the spectrum through a learnable low-pass filter, and then convert the signal back to the time domain through a fast Fourier transform.
[0024] Preferably, the pointwise feed-forward network further includes: combining an MLP and a ReLU activation function to further capture non-linear characteristics, and generating an output through residual connection and layer normalization.
[0025] Preferably, the model parameters of the encoder model include: the maximum sequence length is set to 50, the training batch size is set to 256, the embedding size is set to 64, the number of encoder blocks is set to 2, the optimizer is Adam, and the learning rate is 0.001.
[0026] On the other hand, the present invention also provides a frequency-domain enhanced filtering system for sequence recommendation, including:
[0027] A data acquisition module, configured to acquire dynamic preferences in a user's historical interaction, encode the dynamic preferences in the user's historical interaction, and obtain a user behavior sequence;
[0028] A data processing module, configured to perform frequency masking and mixing strategy processing on the user behavior sequence to obtain processed sequence data;
[0029] A model construction module, configured to construct an encoder model based on a stack of multiple learnable filter blocks, and train the encoder model to obtain a prediction model;
[0030] A prediction module, configured to input the processed sequence data into the prediction model for calculation to obtain a prediction result.
[0031] On the other hand, the present invention also provides an electronic device, including a memory, a processor, and a computing program stored in the memory and executable on the processor, where the processor implements the method when executing the computer program.
[0032] On the other hand, the present invention also provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and the computer program implements the method when executed by a processor.
[0033] Compared with the prior art, the present invention has the following advantages and technical effects:
[0034] The present invention proposes a novel frequency-enhanced filter (FARec) for sequence recommendation. The model uses a learnable filter as an encoder to capture the preference features of users, and introduces a data augmentation module and a frequency trapezoidal structure to improve the model's ability to capture different frequency features. The data augmentation module draws on the idea of the amplitude spectrum in Fourier analysis and introduces two data augmentation methods to adapt to events with different frequency magnitudes: frequency masking and mixing. In addition, the present invention uses frequency domain regularization alignment to enhance the view and the original view. The frequency trapezoidal structure divides the original sequence spectrum into multiple frequency bands, enabling the encoder to focus on different spectra to capture different frequency patterns. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0036] Figure 1 It is a schematic diagram for explaining the trend in the model of the embodiment of the present invention;
[0037] Figure 2 It is a visualization schematic diagram of the learnable filter in the Beauty dataset in the embodiment of the present invention;
[0038] Figure 3 This is the overall framework diagram of the model FARec according to the embodiment of the present invention;
[0039] Figure 4 This is the schematic diagram of the frequency trapezoidal structure according to the embodiment of the present invention;
[0040] Figure 5 This is the schematic diagram of the noise verification result of the present invention on the Beauty dataset according to the embodiment of the present invention. Detailed implementation manners
[0041] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments may be combined with each other. The following will describe this application in detail with reference to the drawings and in combination with the embodiments.
[0042] It should be noted that the steps shown in the flowchart of the drawings may be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.
[0043] Embodiment 1
[0044] As Figure 1-2 shown, in this embodiment, a frequency domain enhanced filtering method for sequence recommendation is provided, including:
[0045] S1. Embedding encoding of the user behavior sequence:
[0046] The sequence recommendation system aims to model the dynamic preferences in the user's historical interactions and predict the subsequent items that the user may interact with in the future. In this embodiment, an item embedding matrix MI ∈ R|I|xd is maintained to project the high-dimensional one-hot representation of the item into a low-dimensional dense representation. Given an item sequence of length n, this embodiment applies a lookup operation from MI to form the input embedding matrix E ∈ Rn×d. In addition, in order to retain the time order of the sequence, this embodiment adds a learnable position encoding matrix P ∈ Rn×d to enhance the input representation of the item sequence. In this way, adding the two embedding matrices can obtain the sequence representation EI ∈ Rn×d.
[0047] S2. Two frequency domain enhancement methods:
[0048] Frequency masking creates an extended interaction sequence for users by randomly masking low - amplitude components. This interaction sequence removes the local detailed features of historical interactions, retains the main features of the sequence, and focuses on learning the long - term interest trends of users. Its idea is similar to compression theory. The long - term interest trends may only depend on some important attributes. Therefore, frequency masking can filter out noise information and improve the discrimination between item embeddings. The mixing frequency strategy generates enhanced samples with different feature distributions from the original samples by swapping the positions of low - amplitude components, and designs a frequency trapezoidal structure to enhance the model's ability to capture information in different frequency bands. This strategy explicitly models high - frequency information, which is often ignored in traditional methods, thus improving the model's fine - grained characterization ability of user behavior features and providing a new direction for the performance improvement of sequential recommendation.
[0049] S3. Construction of the encoder model:
[0050] In this embodiment, a project encoder is developed by stacking multiple learnable filter blocks. A learnable filter block usually includes two sub - layers, namely a filter layer and a point - wise feed - forward network. In the filter layer, this embodiment performs a filtering operation on the features of each dimension in the frequency domain, then performs a residual connection and layer normalization. In the point - wise feed - forward network, this embodiment combines an MLP and a ReLU activation function to further capture non - linear characteristics.
[0051] S4. Prediction of user - interacted items:
[0052] The item sequence interacted by the user is input into the trained prediction model. Using the interest change trend of the user learned by the model, the items that the user may interact with in the future are predicted. Four public benchmark datasets are used. The Amazon dataset is widely used to evaluate various sequential recommendation algorithms, and three subsets are selected from it: beauty, sports, and toys. The fourth dataset, MovieLens - 1M (ML - 1M), contains user rating information for movies. Unpopular items and inactive users with less than five interaction records are removed. The top - k metrics HR@k and NDCG@k are used to evaluate the recommendation list. HR@k indicates how many true items (items that the user is actually interested in) can be successfully hit by the recommendation system. NDCG@k further considers the position ranking information. Where K ∈ {5, 10, 20}. The last item of each user interaction sequence is used as a test sample. The maximum sequence length is set to 50, the training batch size is set to 256, the embedding size is set to 64, and the number of encoder blocks is set to 2. For the optimizer, Adam is used. For the learning rate, it is set to 0.001.
[0053] Furthermore, the embedding encoding of the user behavior sequence specifically includes the following content:
[0054] First, create a project embedding matrix Project the high-dimensional sparse features of the project into low-dimensional dense vectors. The historical interaction sequence of each user is transformed from discrete ID information into a matrix where is the embedding of the project , and n represents the maximum sequence length processed by the model. Given any interaction sequence of length n, this embodiment can use M I to generate the embedding matrix To preserve the time order of the sequence, this embodiment combines a learnable position encoding matrix In addition, this embodiment uses dropout and layer normalization to alleviate the problems of overfitting and unstable training E u = Dropout(LayerNorm(E u + P)).
[0055] Furthermore, in the step S2, it includes: S2-1 discrete Fourier transform; S2-2 frequency masking; S2-3 frequency mixing; S2-4 frequency trapezoidal structure.
[0056] S2-1 discrete Fourier transform: The one-dimensional discrete Fourier transform (1D DFT) is a mathematical tool that transforms a discrete signal in the time domain into a frequency domain representation. Through this transform, this embodiment can obtain the amplitude spectrum of the signal to analyze the importance of each frequency component in the sequence, thereby revealing the change trend of the sequence. The fast Fourier transform is a more efficient implementation method of the discrete Fourier transform. This embodiment first performs the fast Fourier transform on each feature dimension of the embedding matrix to transform it into the frequency domain. This embodiment sorts each frequency component according to the magnitude of the amplitude, and retains the top c frequency components with the largest amplitude without any processing represents a percentage.
[0057] S2-2 frequency masking: This embodiment uses the amplitude spectrum of each dimension to create a mask. Frequency masking retains the high-amplitude frequency components, and this part of the frequency components represents the difference between attributes, that is, the frequency components stacked around ω Figure 1 in i . For the low-amplitude frequency components stacked at ω k , the mask is set to zero, and finally it is inverse-transformed back to the time domain to obtain the enhanced sample.
[0058] S2-3 Mixing Frequency: In many research fields, low-amplitude frequency components are usually regarded as noise, especially the part related to high-frequency information. However, these components often carry local detailed information and have potential modeling value. If these information can be fully exploited, it may be transformed into a key factor to improve the model performance. Based on this, in this embodiment, enhanced samples with different feature distributions from the original samples are generated through a mixing frequency strategy, and a frequency trapezoidal structure is designed to enhance the model's ability to capture information in different frequency bands. This strategy explicitly models high-frequency information, which is often ignored in traditional methods, thereby improving the model's fine-grained characterization ability of user behavior features and providing a new direction for the performance improvement of sequence recommendation.
[0059] S2-4 Frequency Trapezoidal Structure: The mixing frequency strategy uses a frequency trapezoidal structure to optimize the spectrum of the original sequence. This structure divides the entire spectrum of the sequence into multiple frequency bands, and each frequency band corresponds to an encoder of a different layer. In this way, the signal processing of different frequency bands can be balanced, and the model's ability to capture multi-scale user interests can be improved by gradually emphasizing, thereby alleviating the problem of poor performance in high-frequency tasks.
[0060] Furthermore, the construction of the encoder model specifically includes the following content: Stack L filter blocks as the sequence encoder, and the filter block consists of a filtering layer and a feed-forward network. Given the input item representation matrix H of layer l l , in this embodiment, the FFT is performed on each dimension of the feature to transform it into the frequency domain. Then, in this embodiment, the spectrum can be modulated by multiplying with a learnable low-pass filter W ∈ C n ×d . The multiplication in the frequency domain is equivalent to the circular convolution in the time domain, so the filter can also capture sequential features. Finally, in this embodiment, the inverse FFT is used to transform back to the time domain and update the sequence representation. To prevent overfitting, this embodiment introduces residual connections and layer normalization. Finally, a pointwise feed-forward network is used to further capture non-linear features.
[0061] Furthermore, in the step S4-1: Cross-entropy loss is used to optimize the model parameters; S4-2: Frequency domain regularization is used for joint optimization.
[0062] S4-1: After the user's behavior pattern information is extracted through L layers of encoders, the output where the last hidden vector in represents the user representation of this sequence. The user representation is multiplied by the embedding matrix to obtain the relevance of the candidate item, and the softmax is used to convert it into a recommendation probability. Then, the cross-entropy loss is used to optimize the model parameters.
[0063] S4-2: This embodiment adopts a new regularization strategy. By enhancing the high-quality association between positive samples, a more robust feature representation is constructed. This method naturally enhances the ability to capture high-frequency detailed features and low-frequency global patterns, thereby more accurately representing the dynamic interests of users and improving the discriminative ability of the model.
[0064] Embodiment 2
[0065] This embodiment provides a frequency-domain enhanced filtering method for sequence recommendation, including:
[0066] S1. Embedding encoding of the user behavior sequence.
[0067] The basic idea of the embedding layer is to map the original high-dimensional sparse representation (one_hot) into a low-dimensional dense vector space, so that this low-dimensional vector can be used to represent the item (movie). The property of this embedding vector is that vectors corresponding to objects with similar distances have similar meanings, and then the similarity between two items can be measured by calculating the similarity between two low-dimensional vectors.
[0068] Step1: The sequence recommendation system aims to model the dynamic preferences in the user's historical interactions and predict the subsequent items that the user may interact with in the future. We maintain an item embedding matrix MI ∈ R|I|xd, which projects the high-dimensional one-hot representation of the item into a low-dimensional dense representation. Given a sequence of n items, we apply a lookup operation from MI to form the input embedding matrix E ∈ Rn×d. In addition, to preserve the temporal order of the sequence, we add a learnable position encoding matrix P ∈ Rn×d to enhance the input representation of the item sequence. In this way, adding the two embedding matrices together gives the sequence representation EI ∈ Rn×d.
[0069] Step2: However, when the network goes deeper, several problems become worse: the increased model capacity leads to overfitting; the training process becomes unstable (due to vanishing gradients, etc.); models with more parameters usually require more training time. To alleviate these problems, this embodiment combines residual connections, layer normalization, and dropout operations.
[0070] S2. Two frequency-domain enhancement methods:
[0071] Inspired by Fourier analysis in signal processing, this embodiment conducts visual analysis (such as Figure 2As shown, it significantly enhances the filtering effect, enabling more accurate identification and differentiation of different frequency components, thereby more effectively adapting to the ever-changing interest trends of users. Specifically, the embedded representations of historical interactions are regarded as multiple "discrete signals", and the features of each dimension of the embedding matrix are decomposed into sine and cosine components of different frequencies using the discrete Fourier transform, and self-supervised signals are constructed using the amplitude spectrum. The amplitude of each frequency component represents the contribution or "energy" size of that component in the original signal. Among them, the high-amplitude components represent the parts with higher energy in the "signal", usually corresponding to the main features or the main contours in the "signal". The low-amplitude components represent the parts with lower energy in the "signal", corresponding to the secondary features, local details or noise information in the "signal". In view of this, two novel data augmentation strategies reflecting interest trends are proposed: frequency masking and frequency mixing. Frequency masking discards local detail features and retains the main features to learn the long-term basic trends of user interests.
[0072] Frequency mixing differentiates local detail features by swapping the positions of low-amplitude components. For example, a user's frequent consumption behavior in the short term may be only due to shopping festivals or promotional activities. In this case, the user's purchase behavior for a certain item is not affected by certain attributes, but these attributes are very likely to determine the user's purchase behavior for another item. And a frequency trapezoidal structure is applied to adapt to different frequency patterns of user behavior, thereby alleviating the problem of poor performance in high-frequency tasks. Finally, in this embodiment, the augmented sequence is directly applied to the recommendation task itself, and a custom frequency domain regularization is used to maximize the similarity between the augmented sequence and the original sequence of the same user, rather than maximizing the similarity between different augmented sequences of the same user. In this way, the dependence on negative sample pairs is abandoned, thereby improving the inference ability of the model. The following are the detailed steps of the two frequency domain augmentation methods:
[0073] Step1: Discrete Fourier transform.
[0074] The one-dimensional discrete Fourier transform (1DDFT) is a mathematical tool that converts a discrete signal in the time domain to a frequency-domain representation. Through this transformation, in this embodiment, the amplitude spectrum of the signal can be obtained to analyze the importance of each frequency component in the sequence, thereby revealing the changing trend of the sequence. In addition, the non-periodic historical behavior sequence may be difficult to accurately reflect the user's intention in the time domain, while these intentions are easier to identify and understand in the frequency domain. In view of this, this embodiment analyzes the user's behavior pattern with the help of the amplitude spectrum to better capture the user's interests and behavior trends. Given any discrete-time signal x[n], it can be transformed to the frequency domain using 1DDFT. The fast Fourier transform is a more efficient implementation method of the discrete Fourier transform. First, the fast Fourier transform is performed on each feature dimension of the embedding matrix to transform it to the frequency domain. The amplitude reflects the energy and intensity of the "signal" at the corresponding frequency. The fast Fourier transform decomposes the attributes of each dimension in the embedding matrix into a linear combination of different-frequency sine and cosine components, and the superposition of multiple components at the same frequency will produce a higher amplitude.
[0075] Step2: Frequency masking.
[0076] Use the amplitude spectrum of each dimension to create a mask. Frequency masking retains the high-amplitude frequency components, which represent the differences between attributes, that is, the frequency components stacked around ω Figure 1 in i , and the mask for the low-amplitude frequency components stacked at ω k is zero. Finally, it is inverse-transformed back to the time domain to obtain the enhanced sample. Frequency masking creates an extended interaction sequence for the user by randomly masking the low-amplitude components. This interaction sequence removes the local detailed features of historical interactions, retains the main features of the sequence, and focuses on learning the user's long-term interest trends. Its idea is similar to the compression theory. The long-term interest trend may only depend on some important attributes. Therefore, frequency masking can filter out noise information and improve the discrimination between item embeddings. As Figure 5 , the robustness of FARec in noisy data was evaluated on the Beauty dataset. To simulate the scenario of noisy data, a certain proportion of negative items was added to the test dataset while keeping the training dataset unchanged. The specific proportions were 10%, 20%, 30%, and 40%. The results are as Figure 5 shown. After adding negative items, the performance of all models decreased. However, FARec outperformed all baselines at all set noise levels, which further verified the effectiveness of this embodiment. This performance improvement can be attributed to the combined use of the data augmentation method and the frequency-domain regularization technique in this embodiment, thus enhancing the robustness of the model.
[0077] Step3: Frequency mixing.
[0078] In many research fields, low-amplitude frequency components are usually regarded as noise, especially the part related to high-frequency information. However, these components often carry local detailed information and have potential modeling value. If these information can be fully explored, it may be transformed into a key factor to improve the model performance. Based on this, in this embodiment, enhanced samples with different feature distributions from the original samples are generated through a frequency mixing strategy, and a frequency trapezoidal structure is designed to enhance the model's ability to capture information in different frequency bands. This strategy explicitly models high-frequency information, which is often ignored in traditional methods, thus improving the model's fine-grained characterization ability of user behavior features and providing a new direction for the performance improvement of sequential recommendation. In a specific example, frequency mixing can be regarded as an exchange feature between commodities. For example, after a user purchases a commodity of a certain brand, it is very likely that the user will continue to purchase other commodities of this brand. Therefore, the enhanced sequence largely retains the semantic consistency of the prediction. Finally, the frequency mixing strategy uses a frequency trapezoidal structure to optimize the spectrum of the original sequence. As Figure 4 , this structure subdivides the entire spectrum of the sequence into multiple frequency bands, and each frequency band corresponds to an encoder of a different layer. In this way, the signal processing of different frequency bands can be balanced, and the model's ability to capture multi-scale user interests can be improved by gradually emphasizing, thus alleviating the problem of poor performance in high-frequency tasks.
[0079] S3. Construction of the encoder model.
[0080] Deep neural networks (such as RNN, CNN, and Transformer) have been applied to sequential recommendation tasks, aiming to capture dynamic preference features from logged user behavior data for accurate recommendation. However, in an online platform, the recorded user behavior data inevitably contains noise, and deep recommendation models are prone to overfitting on these recorded data. In this embodiment, to address these problems, a full MLP model with a learnable filter for sequential recommendation tasks is proposed. The full MLP architecture gives it a low time complexity, and the learnable filter can adaptively attenuate noise information in the frequency domain. The learnable filter block usually includes two sub-layers, namely the filter layer and the pointwise feed-forward network.
[0081] Step1: Stack L filter blocks.
[0082] Based on the embedding layer, an item encoder is developed by stacking multiple learnable filtering blocks. In the filter layer, a filtering operation is performed on the features of each dimension in the frequency domain. Given the input item representation matrix of the layer, the FFT is performed on the features of each dimension to transform them into the frequency domain. Then, the spectrum is modulated by multiplying with a learnable low-pass filter W∈C×. Finally, the inverse FFT is used to transform the modulated spectrum back to the time domain and update the sequence representation.
[0083] Step2: Pointwise feed-forward network.
[0084] In the pointwise feed-forward network, the MLP and ReLU activation functions are combined to further capture non-linear characteristics. Finally, residual connection and layer normalization operations are performed to generate the output of the \(l\)th layer.
[0085] S4. Prediction of user interaction items.
[0086] The sequence of items interacted by the user is input into the trained prediction model. Using the trend of the user's interest learned by the model, the items that the user may interact with in the future are predicted. Four public benchmark datasets are used. The Amazon dataset is widely used to evaluate various sequential recommendation algorithms, and three subsets are selected from it: beauty, sports, and toys. The fourth dataset MovieLens-1M (ML-1M) contains user rating information for movies. Unpopular items and inactive users with less than five interaction records are removed. The top-k metrics HR@k and NDCG@k are used to evaluate the recommendation list. HR@k indicates how many true items (items that the user is actually interested in) can be successfully hit by the recommendation system. NDCG@k further considers the position ranking information. Among them, \(K\in\{5, 10, 20\}\). The last item of each user interaction sequence is used as the test sample. The maximum sequence length is set to 50, the training batch size is set to 256, the embedding size is set to 64, and the number of encoder blocks is set to 2. For the optimizer, Adam is used. For the learning rate, it is set to 0.001.
[0087] Its main steps include:
[0088] Step1: Optimize the model parameters with cross-entropy loss:
[0089] After the user's behavior pattern information is extracted by the \(L\) layer encoder, the output is obtained where the last hidden vector in represents the user representation of the sequence. The user representation is multiplied by the embedding matrix Obtain the relevance of candidate items and convert it into a recommendation probability using softmax. Then, use cross-entropy loss to optimize the model parameters. SASRec captures long-term dependencies in the user behavior sequence through the self-attention mechanism, thus more precisely understanding user preferences. FMLP-Rec uses learnable filters to attenuate high-frequency noise information in the frequency domain, improving performance. The effectiveness of contrastive learning can be seen in CL4SRec, CoSeRec, and DuoRec. CFIT4SRec constructs self-supervised signals by introducing second-order Fourier transform, uses contrastive learning to enhance the model's adaptability to different frequency components, and achieves good performance. FEARec designs a time-frequency domain hybrid attention mechanism to capture the periodic characteristics of users, and combines contrastive learning to achieve the best performance among the baselines. However, the hybrid attention mechanism significantly increases the model complexity. At the same time, the time-domain data augmentation strategy it adopts has limitations in capturing the global patterns of user behavior and is prone to amplifying the influence of noise. In contrast, FARec improves the performance on all datasets, further verifying the effectiveness of this embodiment. This improvement is mainly attributed to the proposed data augmentation strategy and frequency domain regularization. Compared with CFIT4SRec, which also constructs augmented samples from the frequency domain perspective, as Figure 3 shown, this embodiment abandons the contrastive loss, thus effectively reducing the negative impact of low-quality negative samples on the model. In addition, the frequency masking strategy uses the idea of signal filtering to significantly alleviate the interference of noise in the embedding representation; while the frequency mixing strategy pays more attention to information reconstruction and further enriches the expressive power of the embedding representation through frequency domain interaction. In addition, there may be significant differences in the frequency characteristics of user behavior sequences in different datasets. For example, in the Toys and ML-1M datasets, user behavior is mainly dominated by long-term preferences; while in the Beauty dataset, users' short-term behavior is more prominent. Therefore, the performance of FARec-mask and FARec-mix on different datasets also shows certain differences. This indicates that the frequency characteristics of the dataset have an important impact on the model's enhancement strategy. The experimental results are shown in Table 1;
[0090] Table 1
[0091]
[0092] Step 2: Joint optimization is performed by frequency domain regularization.
[0093] The contrastive loss is a training strategy. Its goal is to bring similar samples closer in the feature space while pushing dissimilar samples farther apart. The original samples and the samples obtained by data augmentation are regarded as positive sample pairs, and the contrastive loss will reduce the distance between them. Negative samples are usually generated by random sampling, that is, the interaction sequences of different users are randomly paired and labeled as negative samples. However, this method of generating negative samples may have quality problems. For example, two users may show a common interest in the same type of items (such as "science fiction movies"), but in contrastive learning, these sequences are regarded as negative sample pairs, which may cause the model to wrongly learn the differences between common interests while ignoring their potential sequence feature correlations. This mislabeling forces the model to overly separate the negative samples in the feature space, thus weakening the model's ability to capture sequence features and ultimately leading to a decline in recommendation performance. This problem reflects the limitations of the current negative sample generation strategy in contrastive learning, highlighting the need for a more refined and semantically consistent construction method when designing high-quality negative sample pairs. The Noise-Contrastive Estimation (NCE) loss is usually applied to train the encoder. Assuming the vectors are normalized, it can be regarded as alignment and consistency. To reduce the negative impact of low-quality negative samples on the contrastive loss, this embodiment adopts a new regularization strategy, namely the Mean Squared Error (MSE). By enhancing the high-quality associations between positive samples to construct more robust feature representations, it no longer relies on the selection of negative samples. This method naturally strengthens the ability to capture high-frequency detail features and low-frequency global patterns, thus more accurately representing the user's interest dynamics and improving the discriminative ability of the model.
[0094] On the other hand, this embodiment also provides a frequency-domain enhanced filtering system for sequence recommendation, including:
[0095] A data acquisition module, configured to acquire the dynamic preferences in the user's historical interactions, encode the dynamic preferences in the user's historical interactions, and obtain a user behavior sequence;
[0096] A data processing module, configured to perform frequency masking and mixing strategy processing on the user behavior sequence to obtain processed sequence data;
[0097] A model construction module, configured to construct an encoder model based on the stacking of multiple learnable filtering blocks, train the encoder model, and obtain a prediction model;
[0098] A prediction module, configured to input the processed sequence data into the prediction model for calculation to obtain a prediction result.
[0099] On the other hand, this embodiment also provides an electronic device, including a memory, a processor, and a computing program stored in the memory and executable on the processor. When the processor executes the computer program, the method is implemented.
[0100] On the other hand, this embodiment also provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the method.
[0101] The above are only the preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A frequency domain enhanced filtering method for sequence recommendation, characterized in that: include: Acquire dynamic preferences in historical user interactions, encode the dynamic preferences in historical user interactions, and obtain a user behavior sequence; Performing frequency masking and mixing strategy processing on the user behavior sequence to obtain processed sequence data; Building an encoder model based on a stack of multiple learnable filter blocks, and training the encoder model to obtain a prediction model; The processed sequence data is input into the prediction model for calculation to obtain the prediction result.
2. The method according to claim 1, characterized in that: The process of obtaining the user behavior sequence includes: Construct an item embedding matrix based on the dynamic preferences in the user's historical interactions, and project the high-dimensional one-hot representation of the item embedding matrix into a low-dimensional dense representation. For the user's historical interaction sequence, the embedding matrix is used to find the embedding vector of each item to form the input embedding matrix; A learnable position embedding matrix is introduced, and the input embedding matrix is added to the position embedding matrix to obtain the user behavior sequence.
3. The method according to claim 1, characterized in that: The process of obtaining the processed sequence data comprises: Performing a fast Fourier transform on each feature dimension of the user behavior sequence to obtain a frequency domain sequence; Randomly shuffle the order of user interaction items to obtain a new embedding matrix; Performing Fourier transform on the frequency domain sequence and the new embedding matrix respectively, retaining the high-amplitude frequency components and exchanging the positions of the low-amplitude components, to obtain a processed frequency domain signal; The processed frequency domain signal is converted back to the time domain by fast Fourier transform, and the signals of different frequency bands are optimized by using a frequency ladder structure to obtain processed sequence data.
4. The method according to claim 1, characterized in that: The process of obtaining the prediction model includes: Stack L layers of filter blocks and build a point-by-point feedforward network; Constructing an encoder model based on the stacked L-layer filter blocks and the point-by-point feed-forward network; The encoder model is trained based on the training set to obtain the prediction model.
5. The method according to claim 4, characterized in that The operation process of the filter block includes: A filtering operation is performed on the features of each dimension in the frequency domain, the spectrum is modulated by a learnable low-pass filter, and then the signal is converted back to the time domain through a fast Fourier transform.
6. The method according to claim 4, characterized in that The point-by-point feedforward network also includes: combining MLP and ReLU activation functions to further capture nonlinear characteristics, and generating outputs through residual connections and layer normalization.
7. The method according to claim 1, characterized in that The model parameters of the encoder model include: the maximum sequence length is set to 50, the training batch size is set to 256, the embedding size is set to 64, the number of encoder blocks is set to 2, the optimizer is Adam, and the learning rate is 0.
001.
8. A frequency domain enhanced filtering system for sequence recommendation, characterized in that: include: A data acquisition module, used to acquire dynamic preferences in user historical interactions, encode the dynamic preferences in the user historical interactions, and obtain a user behavior sequence; A data processing module, used for performing frequency masking and mixing strategy processing on the user behavior sequence to obtain processed sequence data; A model building module, used for building an encoder model based on a stack of multiple learnable filter blocks, and training the encoder model to obtain a prediction model; The prediction module is used to input the processed sequence data into the prediction model to calculate and obtain the prediction result.
9. An electronic device comprising a memory, a processor, and a computing program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method described in any one of claims 1 to 7 is implemented.