Sequence recommendation model based on time-frequency domain conversion and semantic fusion

By introducing time-frequency domain transformation and semantic fusion techniques into the sequence recommendation model, text embedding and feature fusion are optimized, combined with multi-head attention mechanism and residual connection, the problem of insufficient understanding of semantic information and context in the existing model is solved, and the recommendation performance is significantly improved.

CN120144856APending Publication Date: 2025-06-13HUZHOU UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510110280.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing sequence recommendation model lacks in-depth understanding of semantic information and context in user behavior modeling, resulting in limited improvement in recommendation performance.

Method used

A sequence recommendation model based on time-frequency domain transformation and semantic fusion is proposed. Text embedding is optimized through a hybrid expert model, text embedding and ID embedding are used to frequency domain transformation using short-time Fourier transform, and feature fusion is performed in the frequency domain space, combining multi-head attention mechanism and residual connection for sequence modeling.

Benefits of technology

It significantly improves recommendation performance, can more effectively capture the deep-level characteristics of user interests and changes in user dynamic behavior, and enhances the model's adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144856A_ABST
    Figure CN120144856A_ABST
Patent Text Reader

Abstract

The invention discloses a sequence recommendation model based on time-frequency domain conversion and semantic fusion, which comprises the following steps: firstly, introducing a hybrid expert model to enhance the expression ability of text embedding, and then carrying out frequency domain conversion on text embedding and ID embedding through short-time Fourier transform to extract key features; and then fusing the two embedded vectors by using a weight adaptive method, and finally modeling the sequence by using a multi-head attention mechanism with residual connection. Experimental studies on three public data sets show that compared with a current most advanced sequence recommendation model, the method has the advantage that the recommendation performance is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of recommendation models, in particular to the technical field of sequential recommendation models based on time-frequency domain conversion and semantic fusion. Background Art

[0002] With the rapid growth of network information, recommendation systems play an important role in helping users discover relevant and personalized items; sequential recommendation systems identify user preferences based on the user's historical interaction behaviors, and predict the user's next-hop interaction by calculating the correlation between the user preferences and item features; to achieve accurate prediction of the user's next-hop interaction, various sequential recommendation models have been proposed successively, including CNN, RNN, and Transformer; currently, most sequential recommendation models mainly rely on item IDs for user behavior modeling; although such methods perform well in certain scenarios, the modeling method that simply relies on IDs lacks in-depth understanding of semantic information and context, thus limiting the potential for improving recommendation performance.

[0003] To solve the above problems, researchers have begun to explore integrating auxiliary information into the recommendation system to more comprehensively characterize the user's behavior preferences; item-related text information (such as item titles, descriptions, and category labels) has been widely used to improve the performance of sequential recommendation systems; however, existing methods usually use relatively simple fusion techniques, such as simply weighting the ID embedding and text embedding of items; this fusion method is difficult to fully mine the deep features of user interests; in addition, existing methods often ignore the frequency characteristics and time variations of signals, limiting the adaptability of the model to the dynamic behavior changes of users, which has become a bottleneck for the further development of recommendation systems. Summary of the Invention

[0004] The purpose of the present invention is to solve the problems in the prior art, and propose a sequential recommendation model based on time-frequency domain conversion and semantic fusion, which can solve the above problems.

[0005] To achieve the above purpose, the present invention proposes a sequential recommendation model based on time-frequency domain conversion and semantic fusion, including a signal conversion and feature fusion module and a sequential representation learning module;

[0006] In the signal transformation and feature fusion module, first, a mixture of experts model is used to optimize the embedding of item text, then the text embedding and ID embedding are subjected to frequency domain conversion through short-time Fourier transform, the two embeddings are fused in the frequency domain space using a weight adaptive learning method, and finally the feature signal is converted back to the time domain through inverse short-time Fourier transform;

[0007] In the sequence representation learning module, the multi-head attention mechanism is used to model the long-range dependencies in the interaction sequence, and the residual connection is combined to retain the key information in the context. Finally, the Transformer architecture is used to achieve sequence modeling and next-hop recommendation.

[0008] Preferably, the working steps of the sequence recommendation model based on time-frequency domain conversion and semantic fusion are as follows:

[0009] S1. Text embedding encoding and optimization;

[0010] S2. Feature fusion based on time-frequency domain signal conversion;

[0011] S3. Sequence representation learning.

[0012] Preferably, the specific steps of S1 are as follows:

[0013] S1.1. BERT-based text embedding encoding:

[0014] Use BERT to encode the text information of the item: Add the marker "CLS" at the starting position of the text description t i of the item v i ={w 1 , w 2 ,…, w c}, and input the new sequence into BERT for text encoding. The formula is as follows:

[0015] w_t v =PLM([CLS]; w 1 , w 2 ,…, w n ) (1)

[0016] where, ";" represents the text concatenation operation, w_t v represents the hidden vector corresponding to the marker CLS, that is, the initial embedding representation of the item text; BERT is a pre-trained model. During the training process of the model in this article, w_t v always remains unchanged;

[0017] S1.2. Text embedding optimization based on the Mixture of Experts (MoE) model:

[0018] During the embedding optimization process, the text embedding is fused with the position embedding. Specifically, the initial embedding passes through an adapter composed of multiple experts, and a learnable gating mechanism is used to adjust the output of each expert. The formula is as follows:

[0019]

[0020] g = Softmax(t j ·WG +δ) (3)

[0021] Among them, G represents the number of experts, and s j,k represents the embedding of the j-th sequence position when modulating the k-th expert, represents the feature transformation matrix of the k-th expert, g k represents the combined weight of the gating routing, W G represents the feature transformation matrix corresponding to the expert model, and δ represents the introduced random Gaussian noise.

[0022] Preferably, the specific steps of S2 are as follows:

[0023] S2.1. Feature extraction based on short-time Fourier transform:

[0024] First, the input signal x ∈ R n×k is dimensionally reduced through a compression layer to x' ∈ R n×d , where d << k. This process aims to reduce the complexity of subsequent calculations; then, the STFT method is used to extract the frequency-domain representation of the input signal in each channel, and the formula is as follows:

[0025]

[0026] Among them, x(t) represents the input signal in the time domain space; w(t - τ) represents the window selection, which restricts the time range of the signal, and usually a Hanning window is selected; e -j2πfτ represents the complex exponential function, indicating the oscillation of different frequency components; X(t, f) ∈ C T×F represents the spectral representation of the input signal at time t and frequency f, which is in the form of a complex number;

[0027] S2.2. Feature fusion based on dynamic adaptive weights:

[0028] A learnable complex weight matrix w is defined for each feature m in each channel k m,k , and an adaptive feature weight adjustment method is used for feature fusion, and the formula is as follows:

[0029]

[0030]

[0031] Among them, X E,k and X T,k respectively represent the ID embedding E = {e 1 , e 2 ,..., e n} and the text embedding T = {t 1 , t 2 ,..., tn The characteristic representation of}, W E,k and W T,k respectively represent the complex weight matrices corresponding to the ID embedding and the text embedding, and ⊙ represents the element-wise multiplication operation of vectors;

[0032] S2.3. Signal recovery based on inverse short-time Fourier transform:

[0033] After completing the feature fusion in the frequency domain space, it is necessary to restore the fused feature representation to the time domain in order to complete the subsequent sequence modeling; this process involves the reconstruction of each frequency component, and we use the inverse short-time Fourier transform method to achieve it. The formula is as follows:

[0034]

[0035] where x(t) represents the feature representation restored to the time domain.

[0036] Preferably, the specific steps of S3 are as follows:

[0037] S3.1. Sequence item representation learning based on multi-head attention (MHA):

[0038] Taking the reconstructed time series signal obtained by ISTFT as the input, use the MHA method to model the complex relationships between sequence items. The formula is as follows:

[0039] Z = Concat(head 1 , head 2 , …, head h )W O (8)

[0040] where Z represents the output sequence item representation, and W O represents the weight matrix; head i represents the output of the i-th attention module, which is implemented by the self-attention method here. The calculation formula is as follows:

[0041]

[0042] where d k represents the dimension of the input vector, represents the scaling factor, which is used to prevent the calculation result from being too large due to too large dimensions;

[0043] S3.2. Model optimization based on residual connection:

[0044] After the multi-head self-attention model, we introduce a residual connection to further optimize the sequence modeling process. The formula is as follows:

[0045] Y = LN(X + Z) (10)

[0046] Among them, X represents the output vector of the ISTFT, Z represents the output of the multi-head attention (MHA) module, LN represents the layer normalization operation, and Y represents the output sequence representation, that is, the user preference vector;

[0047] S3.3. Sequence Representation Learning and Score Prediction:

[0048] After the above process is iterated L times, the output vector of the last layer is [y 1 , y 2 , …, y n ; First, take the output vector y n at the last position in the sequence as the embedding representation of the sequence, that is, the user preference vector;

[0049] Then, the score prediction of the user for the candidate item is realized by using the vector multiplication method, and the formula is as follows:

[0050] z i = y n T e i (11)

[0051] Among them, e i represents the embedding representation of the i-th candidate item.

[0052] Preferably, a cross-entropy loss function with a temperature coefficient is used for model training, and the formula is as follows:

[0053]

[0054] Among them, z i represents the output of the model for the positive sample, z j represents the output of the model for all candidate items, and T represents the temperature coefficient.

[0055] Advantages of the present invention: The present invention proposes a feature fusion method based on time-frequency domain conversion; the short-time Fourier transform is used to convert the text embedding and ID embedding into the frequency domain, and feature fusion is realized through complex weight adaptive adjustment;

[0056] The present invention proposes a text embedding optimization method based on a mixture of experts model; MoE is introduced to dynamically adjust and optimize the text embedding output by the preprocessing text encoder, improving the discriminability of the project text embedding; a sequence modeling method integrating the multi-head attention mechanism and residual connection is proposed;

[0057] The present invention uses MHA to model the long-range dependencies in the interaction sequence, extracts feature information from multiple perspectives, and retains key contexts through residual connections;

[0058] The present invention has been extensively experimentally verified on public datasets; the experimental results show that compared with the current most advanced sequence recommendation model, the method in this paper has achieved significant improvement in recommendation performance, verifying the effectiveness of the method.

[0059] The features and advantages of the present invention will be described in detail through embodiments in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 This is the overall architecture diagram of the sequence recommendation model based on time-frequency domain conversion and semantic fusion;

[0061] Figure 2 It is a schematic diagram of the text encoding and optimization framework;

[0062] Figure 3 It is a diagram of the feature fusion process based on short-time Fourier transform;

[0063] Figure 4 This is a comparison chart of Recall@10 of each data set;

[0064] Figure 5 It is a schematic diagram of the efficiency of each data set under different numbers of layers. DETAILED DESCRIPTION

[0065] Problem Definition

[0066] In the recommendation system, the user set is U = {u 1 ,u 2 ,…,u m}, the item set is V = {v 1 ,v 2 ,…,v r}, each item v i There is a text description sequence t associated with it i ={w 1 ,w 2 ,…,w c}, where c is the length of the text sequence; each user u has a historical interaction sequence s u = {v 1 ,v 2 ,…,v n}, the items in the sequence are arranged in the order of interaction time, n is the number of items in the sequence; sequence recommendation is based on the user's historical interaction sequence s u , predict the interactive item v of user u at the next moment n+1 , that is, P(v n+1 ∣s u ).

[0067] Text Model

[0068] We propose a novel sequence recommendation model SF-Rec based on time-frequency domain conversion and semantic fusion. This model consists of a signal conversion and feature fusion module and a sequence representation learning module, as Figure 1 shown; in the signal transformation and feature fusion module, first, the mixture of experts model (MoE) is used to optimize the embedding of item texts. Then, the text embedding and ID embedding are converted to the frequency domain through the short-time Fourier transform (STFT). In the frequency domain space, a weight adaptive learning method is used to fuse the two embeddings. Finally, the feature signal is converted back to the time domain through the inverse short-time Fourier transform (ISTFT); in the sequence representation learning module, the multi-head attention (MHA) mechanism is used to model the long-range dependencies in the interactive sequence, and the residual connection is combined to retain the key information in the context. Finally, the Transformer architecture is used to achieve sequence modeling and next-hop recommendation;

[0069] Text Embedding Encoding and Optimization

[0070] The processing flow of encoding and optimizing the item text in this paper is as Figure 2 shown: First, the pre-trained language model BERT is used to obtain the initial embedding of the item text; then, the mixture of experts model (MoE) is introduced to optimize the text embedding to obtain a more distinguishable item text embedding representation.

[0071] Text Embedding Encoding Based on BERT

[0072] BERT is a pre-trained language model based on the Transformer architecture, mainly used for natural language processing (NLP) tasks; here we use BERT to encode the text information of the item: Add the marker "CLS" at the starting position of the text description t i of the item v i ={w 1 ,w 2 ,…,w c}, and input the new sequence into BERT for text encoding. The formula is as follows:

[0073] wt v =PLM([CLS];w 1 ,w 2 ,…,w n ) (1)

[0074] where, ";" represents the text concatenation operation, and wt v represents the hidden vector corresponding to the marker CLS, that is, the initial embedding representation of the item text; BERT is a pre-trained model, and during the training process of the model in this paper, wt v always remains unchanged.

[0075] Optimization of Text Embedding Based on Mixture of Experts (MoE)

[0076] To improve the discriminability of text representation, we further optimize the project text embedding representation using the Mixture of Experts (MoE); during the embedding optimization process, text embedding and positional embedding are fused to prepare for subsequent sequence modeling; specifically, the initial embedding passes through an adapter composed of multiple experts, and a learnable gating mechanism is used to adjust the output of each expert. The formula is as follows:

[0077]

[0078] g = Softmax(t j ·W G + δ) (3)

[0079] where G represents the number of experts, s j,k represents the embedding of the j-th sequence position when modulating the k-th expert, represents the feature transformation matrix of the k-th expert, g k represents the combined weight of the gating routing, W G represents the feature transformation matrix corresponding to the expert model, and δ represents the introduced random Gaussian noise.

[0080] By combining multiple modulated experts, the model can dynamically capture the context information of the interaction sequence and generate a discriminative text representation for each item in the sequence; this flexible structure not only improves the diversity of text representation but also enhances the model's adaptive ability in the face of different tasks.

[0081] Feature Fusion Based on Time-Frequency Domain Signal Transformation

[0082] To achieve semantic fusion in sequence recommendation, it is necessary to explore an effective fusion strategy to fuse ID embedding and text embedding, capturing the changes in user interests while using text information to deeply understand user behavior; this paper proposes a feature fusion method based on the Short-Time Fourier Transform (STFT). First, ID embedding and text embedding are transformed into the frequency domain to capture the instantaneous changes of the signal, and then the two features are fused in the frequency domain space. The detailed process is as Figure 3 shown.

[0083] Feature Extraction Based on Short-Time Fourier Transform (STFT)

[0084] The Short-Time Fourier Transform (STFT) is a method for converting a time-domain signal to the frequency domain. It can provide the frequency components of the signal at different time periods, thereby observing the changes in the time-series signal at the microscopic level; specifically, STFT analyzes the frequency characteristics of the signal by dividing the signal into several short-time segments and performing Fourier transforms on each segment.

[0085] In this paper, the STFT method is used to extract the user's behavioral features from the user interaction sequence, and the process is as follows: First, the input signal x ∈ R n×k is dimensionally reduced through a compression layer to x' ∈ R n×d , where d << k. This process aims to reduce the complexity of subsequent calculations; then, the STFT method is used to extract the frequency-domain representation of the input signal in each channel, and the formula is as follows:

[0086]

[0087] where x(t) represents the input signal in the time domain space; w(t - τ) represents the window selection, which restricts the time range of the signal, and usually a Hanning window is selected; e -j2πfτ represents the complex exponential function, indicating the oscillation of different frequency components; X(t, f) ∈ C T×F represents the spectral representation of the input signal at time t and frequency f, which is in the form of a complex number; the calculation result in the complex number form retains the amplitude information and phase information, which is very important for comprehensively understanding the signal changes.

[0088] Feature Fusion Based on Dynamic Adaptive Weights

[0089] To enhance the interaction between text features and ID features, we propose a feature fusion method based on dynamic adaptation of complex weights; specifically, a learnable complex weight matrix w is defined for each feature m on each channel k m,k , and an adaptive feature weight adjustment method is used for feature fusion, and the formula is as follows:

[0090]

[0091]

[0092] where X E,k and X T,k represent the feature representations of the ID embedding E = {e 1 , e 2 ,..., e n} and the text embedding T = {t 1 , t 2 ,..., t n} respectively, W E,k and W T,k represent the complex weight matrices corresponding to the ID embedding and the text embedding respectively, and ⊙ represents the element-wise multiplication operation of vectors.

[0093] By introducing a learnable complex weight vector, the model can effectively fuse context information, enhance the model's sensitivity to different frequency information, and then generate richer and more meaningful feature representations. Signal recovery based on the inverse short-time Fourier transform (ISTFT)

[0094] After completing feature fusion in the frequency domain space, it is necessary to restore the fused feature representation to the time domain in order to complete subsequent sequence modeling; this process involves the reconstruction of each frequency component, and we use the inverse short-time Fourier transform (ISTFT) method to achieve it. The formula is as follows:

[0095]

[0096] Among them, x(t) represents the feature representation restored to the time domain; the above transformation process effectively suppresses the edge effect in the reconstructed signal by weighting each frequency component using the Hann window function, making the reconstruction process smoother and more natural; the length of the reconstructed signal after the inverse transformation is the same as that of the original signal, ensuring the integrity of information.

[0097] Sequence representation learning

[0098] Sequence item representation learning based on multi-head attention (MHA)

[0099] The core idea of the multi-head attention (MHA) mechanism is to introduce multiple parallel attention structures, enabling the model to learn different aspects of feature representations in different subspaces; taking the reconstructed time series signal obtained by ISTFT as the input, the MHA method is used to model the complex relationships between sequence items. The formula is as follows:

[0100] Z = Concat(head 1 , head 2 , …, head h )W O (8)

[0101] Among them, Z represents the output sequence item representation, and W O represents the weight matrix; head i represents the output of the i-th attention module, which is implemented using the self-attention method here. The calculation formula is as follows:

[0102]

[0103] Among them, d k represents the dimension of the input vector, represents the scaling factor, which is used to prevent the calculation result from being too large due to too large a dimension.

[0104] Model optimization based on residual connection

[0105] The residual connection constructs a skip connection to retain key information in the context and reduce the risk of gradient vanishing. After the multi-head self-attention model, we introduce a residual connection to further optimize the sequence modeling process. The formula is as follows:

[0106] Y = LN(X + Z) (10)

[0107] Among them, X represents the output vector of ISTFT, Z represents the output of the multi-head attention (MHA) module, LN represents the layer normalization operation, and Y represents the output sequence representation, that is, the user preference vector.

[0108] The residual connection fuses the output of ISTFT and the output of MHA, enabling the model to comprehensively consider time-frequency domain features and context information, and providing a richer feature representation for the model.

[0109] Sequence Representation Learning and Score Prediction

[0110] After the above process is iterated L times, the output vector of the last layer is [y 1 , y 2 , …, y n ; First, take the output vector y n at the last position in the sequence as the embedding representation of the sequence, that is, the user preference vector; then, use the vector multiplication method to implement the score prediction of the user for the candidate items. The formula is as follows:

[0111] z i = y n T e i (11)

[0112] Among them, e i represents the embedding representation of the i-th candidate item.

[0113] Model Training

[0114] In this paper, a cross-entropy loss function with a temperature coefficient is used for model training. The formula is as follows:

[0115]

[0116] Among them, z i represents the output of the model for the positive sample, z j represents the output of the model for all candidate items, and T represents the temperature coefficient; the introduction of the temperature coefficient aims to control the smoothness of the probability distribution, so that the model pays different attentions to different items, thereby enhancing the accuracy of sequence recommendation.

[0117] Experiments

[0118] Datasets

[0119] The experiments in this paper are conducted on three public datasets. The MovieLens-1M (ml-1m) dataset contains 1 million movie rating records from the MovieLens website. We merge the movie title, type, and year fields as item text information; the OnlineRetail (OR) dataset records the cross-border transaction information of a British e-commerce platform, in which the product description field is used as item text information; Office is a dataset of product reviews from the Amazon website, in which inactive users and unpopular items with less than 5 interactions are filtered out; the statistical details of the dataset are shown in Table 1:

[0120] Table 1 Statistics of the dataset

[0121]

[0122] Comparison Models

[0123] This article selects 8 representative sequence recommendation models as comparison models, which are introduced as follows: SASRec is the first sequence recommendation model based on the Transformer architecture, which shows good performance in sequence modeling.

[0124] FEARec proposes a frequency-enhanced attention network, combined with contrastive learning alignment representation, which improves the model's ability to capture user preferences.

[0125] SASRecF is an extended version of SASRec, which enhances the expressiveness of the model by combining item IDs and item attributes.

[0126] FDSA uses two different Transformer encoders to encode the item ID and item attributes respectively, and uses the attention method to fuse the output encodings, thereby improving the expressiveness of the model.

[0127] DIF-SR is a Transformer-based sequential recommendation model that improves the fusion method of auxiliary information through a decoupled self-attention mechanism and enhances the flexibility of the model.

[0128] UniSRec leverages item text descriptions to learn transferable item representations, aiming to improve the accuracy and generalization ability of recommendations.

[0129] SMLP4Rec achieves three-way auxiliary information fusion by capturing sequential, cross-channel, and cross-feature dependencies.

[0130] TedRec transforms text embedding and ID embedding from time domain to frequency domain, achieving sequence-level semantic fusion.

[0131] Experimental setup

[0132] In this paper, the leave-one-out method is adopted to evaluate the recommendation performance. That is, given a user interaction sequence, the last sequence item is used as the test item. The evaluation metrics use the common ranking metrics Recall@K (R@K) and NDCG@K (N@K), and K is set to 10 and 20 in the experiment. We implemented the proposed method on Recbole, and the comparison models were directly implemented using the models provided by Recbole. To ensure a fair comparison, we uniformly adopted the Adam optimizer and the cross-entropy loss function to optimize the models, and set the training batch size of all models to 2048. During the experiment, when the N@10 metric on the validation set did not improve within 10 epochs, an early stopping strategy was adopted to avoid overfitting.

[0133] Experimental Results

[0134] Performance Comparison

[0135] The performance of the proposed model and the comparison models on three datasets is shown in Table 2, where the best performance is shown in bold and the second-best performance is underlined. It can be seen from the table that the SMLP4Rec model performs the worst, perhaps because the optimal ratio is not achieved in the three-way auxiliary information fusion process. The performance of FDSA and UniSRec has improved, indicating that the introduction of auxiliary information has played a positive role. Next, SASRec performs well, indicating that the Transformer architecture adopted in it plays a key role in sequence modeling. SASRecF further improves the recommendation performance by incorporating item attributes on the basis of SASRec, once again demonstrating the role of auxiliary information. DIF-SR enhances the role of auxiliary information due to the adoption of the decoupled self-attention mechanism. FEARec uses the frequency domain transformation method, further improving the recommendation performance, indicating that the frequency domain transformation helps to capture the dynamic preferences of users. TedRec adopts a method similar to this paper, performing frequency domain transformation on both ID information and text information, and it is the best-performing comparison model, further demonstrating the dual role of frequency domain transformation and auxiliary information. Finally, on all datasets, the recommendation performance of the proposed model is significantly better than all comparison models, and this result fully proves the effectiveness and generalization ability of the proposed model.

[0136] Table 2 Recommendation Performance of Different Models

[0137]

[0138] Ablation Experiment

[0139] In this section, we will conduct an in-depth analysis of the key components included in the proposed model to evaluate the role of different components in the recommendation performance. Specifically, we implemented the following four model variants:

[0140] w / o MOE: Without using the Mixture of Experts (MOE) for text embedding optimization, directly use the text representation generated by the BERT encoder as the text embedding of the project.

[0141] w / o STFT: Remove the time-frequency domain conversion module based on the Short-Time Fourier Transform (STFT) in the model, and keep the rest unchanged.

[0142] w / o MHA: Remove the feature optimization module based on the Multi-Head Attention (MHA) mechanism and residual connection, and directly use the Transformer architecture to implement sequence modeling.

[0143] w / o FUSION: In the model of this paper, instead of using the dynamic adaptive feature fusion method, choose to simply add the ID embedding and the text embedding.

[0144] The performance comparison between the method of this paper and four variants on three datasets is as Figure 4 shown. Here we only show the Recall@10 metric, and the results of the NDCG@K metric are exactly the same as those of the Recall@10 metric; it can be seen from the figure that all components in the model of this paper have different degrees of influence on the recommendation performance; comparatively speaking, the Short-Time Fourier Transform (STFT) and dynamic adaptive feature fusion (FUSION) have a greater impact on the recommendation performance than the Mixture of Experts (MOE) and the Multi-Head Attention mechanism (MHA); on the OR dataset, the role of the Short-Time Fourier Transform (STFT) is the most obvious, because the OR dataset is highly sparse, so the frequency features of the time series provide important information for recommendations; on the other two datasets, the role of the dynamic adaptive feature fusion (FUSION) is relatively obvious, and this phenomenon indicates that the Multi-Head Attention mechanism and residual connection play a key role in maintaining the stability of the model.

[0145] Parameter Sensitivity

[0146] In this section, we conduct a detailed analysis of the impact of the key hyperparameters used in the model, that is, the number of compression layers in the Short-Time Fourier Transform (STFT), on the recommendation performance; we respectively choose the number of layers as 1, 4, and 10 for testing, and the experimental results are as Figure 5 shown; it can be seen from the figure that the effect of 1-layer compression is the best, and this finding prompts us to give priority to the retention of effective information in model design, because the most critical information is extracted by 1-layer compression; theoretically, more layers of compression can increase the expressive power of the model, however, more layers of compression may also lead to the loss of effective information, thus affecting the overall performance of the model.

[0147] Furthermore, we recorded the average training time required by models with different numbers of compression layers in each epoch, and the results are shown in Table 3. It can be seen from the table that as the number of compression layers increases, the training time of the model rises rapidly. This phenomenon indicates that deeper network structures require more computing resources for parameter learning and model optimization. In summary, choosing 1-layer compression not only has advantages in performance but also has an absolute advantage in training efficiency, providing guarantee for the efficient operation of the model. Based on the above experimental results, the model in this paper uses 1-layer compression.

[0148] Table 3 Training Time (seconds) Required for Different Compression Layers

[0149]

[0150] The present invention proposes a novel sequence recommendation framework SF-Rec, which introduces a mixture of experts model (MOE) to optimize the text embeddings of items, uses the short-time Fourier transform (STFT) to transform the text embeddings and ID embeddings from the time domain to the frequency domain, adopts a dynamic adaptive weight adjustment method in the frequency domain space to achieve effective fusion of text features and ID features, uses the multi-head attention (MHA) mechanism to learn long-range dependencies in the sequence, and combines residual connections to improve the stability of the model. The next research will consider applying the frequency-domain-based feature extraction method to more recommendation tasks, such as multi-modal recommendation and cross-domain recommendation.

[0151] The above embodiments are illustrative of the present invention and not restrictive thereof. Any simple transformation of the present invention falls within the protection scope of the present invention.

Claims

1. A sequence recommendation model based on time-frequency domain conversion and semantic fusion, characterized by: It includes signal conversion and feature fusion module and sequence representation learning module; In the signal transformation and feature fusion module, the hybrid expert model is first used to optimize the embedding of the project text, and then the text embedding and ID embedding are transformed into the frequency domain through short-time Fourier transform. The two embeddings are fused in the frequency domain space using the weight adaptive learning method, and finally the feature signal is converted back to the time domain through the inverse short-time Fourier transform. In the sequence representation learning module, the multi-head attention mechanism is used to model the long-distance dependency relationships in the interaction sequence, and the residual connection is combined to retain the key information in the context. Finally, the Transformer architecture is used to achieve sequence modeling and next-hop recommendation.

2. The sequence recommendation model based on time-frequency domain conversion and semantic fusion according to claim 1, characterized in that: The working steps of the sequence recommendation model based on time-frequency domain conversion and semantic fusion are as follows: S1, text embedding encoding and optimization; S2, feature fusion based on time-frequency domain signal transformation; S3. Sequence representation learning.

3. The sequence recommendation model based on time-frequency domain conversion and semantic fusion as claimed in claim 2, characterized in that: The specific steps of S1 are as follows: S1.

1. Text embedding encoding based on BERT: Use BERT to encode the text information of the project: In project v i Text description of i ={w1,w2,…,w c }, add the tag "CLS" at the beginning of the sequence, and input the new sequence into BERT for text encoding. The formula is as follows: wt v =PLM([CLS];w1,w2,…,w n ) (1) Among them, ";" indicates the text connection operation, wt v represents the hidden vector corresponding to the tag CLS, that is, the initial embedding representation of the project text; BERT is a pre-trained model. In the training process of this model, wt v Always remain the same; S1.

2. Text embedding optimization based on hybrid expert model: In the process of embedding optimization, the text embedding and the position embedding are fused. Specifically, the initial embedding is passed through an adapter composed of multiple experts, and the output of each expert is adjusted using a learnable gating mechanism. The formula is as follows: g=Softmax(t j ·W G +δ) (3) Among them, G represents the number of experts, s j,k represents the embedding of the jth sequence position when modulating the kth expert, represents the feature transformation matrix of the kth expert, g k represents the combined weight of the gating routing, W G represents the feature transformation matrix corresponding to the expert model, and δ represents the introduced random Gaussian noise.

4. The sequence recommendation model based on time-frequency domain conversion and semantic fusion as claimed in claim 2, characterized in that: The specific steps of S2 are as follows: S2.1, Feature extraction based on short-time Fourier transform: First, the input signal \(x\in\mathbb{R}\) n×k is dimensionally reduced to \(x'\in\mathbb{R}\) through a compression layer n×d , where \(d\ll k\). This process aims to reduce the complexity of subsequent calculations. Then, the STFT method is used to extract the frequency-domain representation of the input signal in each channel, and the formula is as follows: Among them, x(t) represents the input signal in the time domain; w(t-τ) represents the window selection, which limits the time range of the signal, and usually the Hanning window is selected; e -j2πfτ represents a complex exponential function, which indicates the oscillation of different frequency components; X(t, f)∈C T×F Represents the spectrum of the input signal at time t and frequency f, which is in the form of a complex number; S2.

2. Feature fusion based on dynamic adaptive weights: For each feature m, a learnable complex weight matrix w is defined on each channel k m,k , the feature fusion is performed using the adaptive feature weight adjustment method, the formula is as follows: Among them, X E,k and X T,k They represent ID embedding E = {e1, e2, ..., e n } and text embedding T = {t1,t2,...,t n }, W E,k and W T,k They represent the complex weight matrices corresponding to ID embedding and text embedding respectively, and ⊙ represents the vector element-by-element multiplication operation; S2.3, signal recovery based on inverse short-time Fourier transform: After completing feature fusion in the frequency domain, the fused feature representation needs to be restored to the time domain in order to complete subsequent sequence modeling; this process involves the reconstruction of each frequency component, which we achieve using the inverse short-time Fourier transform method, as shown in the following formula: Where x(t) represents the feature representation restored to the time domain.

5. The sequence recommendation model based on time-frequency domain conversion and semantic fusion as claimed in claim 2, characterized in that: The specific steps of S3 are as follows: S3.

1. Sequence Item Representation Learning Based on Multi-Head Attention: The reconstructed time series signal obtained by ISTFT is used as input, and the MHA method is used to model the complex relationship between sequence items. The formula is as follows: Z=Concat(head1,head2,…,head h )W O (8) Among them, Z represents the output sequence term representation, W O Represents the weight matrix; head i Represents the output of the i-th attention module, which is implemented using the self-attention method. The calculation formula is as follows: Among them, d k represents the dimension of the input vector, Represents the scaling factor, which is used to prevent the dimension from being too large and causing the calculation result to be too large; S3.

2. Model optimization based on residual connection: We introduce residual connections after the multi-head self-attention model to further optimize the sequence modeling process. The formula is as follows: Y=LN(X+Z) (10) Among them, X represents the output vector of ISTFT, Z represents the output of the multi-head attention module, LN represents the layer normalization operation, and Y represents the sequence representation of the output, that is, the user preference vector; S3.3, Sequence Representation Learning and Rating Prediction: After L layers of iterations of the above process, the final output vector is [y1,y2,…,y n ]; First, take the output vector y at the last position in the sequence n As an embedding representation of the sequence, i.e., the user’s preference vector; Then, the vector multiplication method is used to predict the user's rating of the candidate items. The formula is as follows: z i =y n T yes i (11) Among them, e i represents the embedding representation of the i-th candidate.

6. The sequence recommendation model based on time-frequency domain conversion and semantic fusion as claimed in claim 2, characterized in that: The cross entropy loss function with a temperature coefficient is used to train the model. The formula is as follows: Among them, z i Represents the output of the model for positive samples, z j represents the output of the model for all candidates, and T represents the temperature coefficient.

Citation Information

Cited By

  • Multi-modal sequence recommendation method based on double-gating hybrid expert model and Fourier noise reduction

    CN120804383A

  • Multimodal sequential recommendation method based on double-gated hybrid expert model and fourier denoising

    CN120804383B