Sequence recommendation method and system based on temporal course learning and medium

Through the temporal course learning method, the recommendation system is trained in stages to learn short-term and long-term preferences, the problems of user interest changes and forgetting in the sequence recommendation system are solved, and the accuracy and adaptability of recommendations are improved.

CN120296245APending Publication Date: 2025-07-11NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510347553.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Existing sequence recommendation systems have challenges in capturing short-term and long-term preferences, especially how to effectively integrate changes in user interests and avoid catastrophic forgetting.

Method used

Using a time-based course learning method, the user interaction data is sorted by timestamp and divided into multiple time stages. Embedding representations containing temporal sequential features are generated through the embedding layer and the position embedding layer. Implicit features are extracted using the sequence feature extraction layer, and model parameters are optimized through backpropagation and loss functions to gradually learn short-term and long-term preferences.

Benefits of technology

It effectively solves the problem of balance between short-term and long-term preferences, alleviates catastrophic forgetting, improves recommendation accuracy and user satisfaction, and has better time-dependent modeling capabilities and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296245A_ABST
    Figure CN120296245A_ABST
Patent Text Reader

Abstract

The invention relates to a sequence recommendation method and system based on temporal course learning and a medium. The method comprises the steps that user interaction data are sorted according to timestamps and divided into a plurality of time stages; converting the interaction data into a sequence with a fixed length; mapping items in the sequence into embedded vectors in a potential space; time sequence information is injected, and embedded representation with time sequence features is generated; extracting hidden features to generate hidden representation; inputting the hidden representation into a recommendation model, performing preliminary training, and updating model parameter learning short-term preference; gradually introducing historical interaction data according to a time reverse order, and updating model parameter learning long-term preference; the weight of the loss function is dynamically adjusted, so that the short-term preference has a relatively high weight, and the long-term preference has a relatively low weight; carrying out overall training to obtain final model parameters; and outputting a recommendation item of the next interaction of the user according to the final model. The method effectively captures time dependence, adapts to user interest changes, and improves the generalization ability of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information retrieval and personalized services, and more particularly to a sequential recommendation method, system, and medium based on time-based course learning. Background Art

[0002] Recommendation systems play a crucial role in the fields of information retrieval and personalized services. With the widespread popularity of Internet applications and the increasing richness of user behavior data, recommendation systems have become the core tools for various online platforms (such as e-commerce, social media, and news websites) to provide personalized content and services. By analyzing the historical behavior data of users, recommendation systems can predict the products or content that users may be interested in, thereby improving the user experience and the overall efficiency of the platform.

[0003] As an important branch of recommendation systems, sequential recommendation systems mainly focus on predicting the items or content that users may be interested in in the future based on their past interaction sequences. Compared with traditional recommendation systems, sequential recommendation systems not only need to capture the immediate interests of users but also must consider the time dynamics and be able to adapt to the changes in users' short-term and long-term preferences. Short-term preferences are usually determined by users' recent behaviors or interactions, reflecting their immediate needs at a certain moment, while long-term preferences represent the stable interests that users maintain over a longer period. To improve the accuracy and personalization of recommendations, accurately modeling short-term and long-term preferences becomes particularly important.

[0004] However, current sequential recommendation systems still face many challenges in capturing short-term and long-term preferences. First, users' interests change over time. Especially in a short period, users' preferences may fluctuate violently due to newly occurred interaction behaviors. How to effectively capture these changes and reflect them in recommendations is an urgent problem to be solved. Second, existing recommendation systems often face the problem of catastrophic forgetting, that is, when the model processes new user interaction data, it is easy to forget the historical knowledge learned before.

[0005] In addition, although some advanced sequential recommendation methods (such as deep learning models like LSTM, GRU, and Transformer) can capture the time dependence of user behaviors to a certain extent, these methods still face the problems of how to effectively integrate short-term and long-term interests and how to avoid catastrophic forgetting. Summary of the Invention

[0006] The present invention provides a sequential recommendation method, system, and medium based on time-based course learning, aiming to solve the deficiencies in the balance between short-term and long-term preferences, the problem of catastrophic forgetting, and the processing of time-series data in current recommendation systems.

[0007] To achieve the above object, the first aspect of the present invention provides a sequential recommendation method based on temporal curriculum learning, including the following steps:

[0008] Sort the user interaction data by timestamp and divide it into multiple time stages according to time periods, where each time stage contains a set of user interaction data arranged in chronological order;

[0009] Construct a training set for the user interaction data of each time stage and convert the user interaction data of each time stage into an interaction sequence of a fixed length;

[0010] Map the items in the interaction sequence of fixed length into an embedding vector representation in the latent space through an embedding layer;

[0011] Inject chronological information, encode the temporal positions of the embedding vector representations in the latent space through a positional embedding layer to form an embedding representation containing chronological features;

[0012] Use a sequence feature extraction layer to extract implicit features from the embedding representation with chronological features to generate a hidden representation;

[0013] Input the generated hidden representation into a recommendation model for preliminary training, and update the model parameters through backpropagation to learn short-term preferences;

[0014] In subsequent time stages, introduce historical interaction data step by step in reverse chronological order, gradually add the interaction data of earlier time stages to the training set, and update the model parameters using the trained model to learn long-term preferences;

[0015] During the training process of each time stage, dynamically adjust the weights of the loss function so that the loss term of short-term preferences has a higher weight and the loss term of long-term preferences has a lower weight;

[0016] In the final stage, use the user interaction data of all time stages for overall training to obtain the finally trained model parameters;

[0017] According to the finally trained model parameters, output the recommended items for the user's next interaction.

[0018] Further, the method for converting the user interaction data of each time stage into an interaction sequence of a fixed length includes:

[0019] Process the user interaction data of each time stage and arrange it in chronological order;

[0020] According to a preset fixed sequence length, intercept the user interaction data of each time stage into several interaction sequences of a fixed length;

[0021] If the length of the user interaction data in a certain time period exceeds the preset fixed sequence length, the most recent interaction records are retained;

[0022] If the length of the user interaction data in a certain time period is less than the preset fixed sequence length, the sequence is padded with placeholder tokens to reach the fixed length.

[0023] Furthermore, the method of mapping the items in the fixed-length interaction sequence to the embedded vector representation in the latent space includes:

[0024] Define an embedding matrix for each item, where each row in the embedding matrix corresponds to the latent space embedding vector of an item;

[0025] Input the item IDs in each fixed-length interaction sequence into the embedding layer, and map the item IDs to the corresponding embedded vector representations through this embedding matrix.

[0026] Furthermore, the method of forming the embedded representation containing the chronological features includes:

[0027] Inject chronological information into the embedded vectors in each interaction sequence through the positional embedding layer;

[0028] The positional embedding layer uses learnable positional embedding vectors to encode each item's embedded vector according to the item's chronological position in the interaction sequence, forming an embedded representation containing chronological features;

[0029] The positional embedding vector is added to or concatenated with the item's embedded vector to generate the final embedded representation with chronological features.

[0030] Furthermore, the method of generating the hidden representation includes:

[0031] Input the embedded representation containing the chronological features into the sequence feature extraction layer;

[0032] The sequence feature extraction layer extracts features from the embedded representation through at least one temporal model;

[0033] The temporal model captures the long-term dependencies and short-term dynamic features in the user interaction sequence according to the chronological order, generating the corresponding hidden feature representation.

[0034] Furthermore, the method of learning the short-term preferences includes:

[0035] Input the generated hidden representation into the recommendation model as the initial input of the recommendation model;

[0036] The recommendation model processes the input hidden representation through forward propagation, generating the prediction result of the next interaction that the user may perform at the current time stage;

[0037] According to the error between the user's real interaction data and the interaction result predicted by the recommendation model, use the backpropagation algorithm to calculate the error and update the model parameters;

[0038] During the process of updating the model parameters through backpropagation, optimize the model to learn short-term preferences and capture the user's immediate interests within the current time period;

[0039] Repeat the above steps for iterative training until the recommendation model can accurately predict the user's short-term preferences.

[0040] Furthermore, the method for learning long-term preferences includes:

[0041] In subsequent time stages, gradually introduce the interaction data from earlier time stages in reverse chronological order, add the data of each stage to the training set, and gradually expand the data in chronological order;

[0042] Use the recommendation model trained in the current time stage as the base model, load the model weights of the previous stage, and merge the new time stage data with the data of the previous stage for incremental training;

[0043] In each time stage, keep the training results of the previous stage unchanged and only update the newly added historical interaction data part;

[0044] Update the model parameters through the backpropagation algorithm to enable the model to learn long-term preferences and effectively capture the stable interest changes of users over a long time span;

[0045] Repeat the above steps for iterative training until the recommendation model can accurately predict the user's long-term preferences.

[0046] Furthermore, the method also includes using the cross-entropy loss function and introducing a loss term for contrastive learning during the training process, specifically including:

[0047] Within each time stage, use the cross-entropy loss function to measure the error between the user interaction data and the model prediction results, and calculate the correlation score of each candidate item;

[0048] Introduce a loss term for contrastive learning, and optimize the model parameters by comparing the differences between positive samples and negative samples;

[0049] Use the following formula to calculate the loss function:

[0050]

[0051] Where, represents the total loss function of the recommendation system, represents the set of users Sum over all users u, where j represents the index of an item or event. And Represent positive and negative samples respectively. Denotes the interaction sequence of user u before the j-th step, used to describe the user's historical behavior. σ represents the Sigmoid activation function, which is used to map the prediction score to the range (0, 1). Denotes the probability that the user interacts with the positive sample Given the interaction sequence of user u before the j-th step Of the user with the positive sample. Denotes the probability that the user interacts with the negative sample Given the interaction sequence of user u before the j-th step Of the user with the negative sample;

[0052] Update the model parameters through the backpropagation of this loss function.

[0053] To achieve the above object, the first aspect of the present invention provides a sequential recommendation system based on temporal curriculum learning, including the following modules:

[0054] User interaction data receiving module, used to sort the user interaction data by timestamp;

[0055] Time period division module, which divides into multiple time stages according to time periods, where each time stage contains a set of user interaction data arranged in chronological order;

[0056] Training set construction module, used to construct a training set for the user interaction data of each time stage, and convert the user interaction data of each time stage into an interaction sequence of fixed length;

[0057] Embedding layer module, used to map the items in the interaction sequence of fixed length into an embedding vector representation in the latent space through the embedding layer;

[0058] Position embedding layer module, used to encode the time position of the embedding vector representation in the latent space through the position embedding layer to form an embedding representation containing chronological order features;

[0059] Sequence feature extraction layer module, used to extract implicit features from the embedding representation with chronological order features by using the sequence feature extraction layer to generate a hidden representation;

[0060] Recommendation model module, used to input the generated hidden representation into the recommendation model for preliminary training, and update the model parameters through backpropagation to learn short-term preferences;

[0061] A historical data introduction module, which is used to introduce historical interaction data step by step in reverse chronological order in subsequent time phases, gradually add interaction data in earlier time phases to the training set, and update model parameters using the trained model to learn long-term preferences;

[0062] A dynamic adjustment module, which is used to dynamically adjust the weights of the loss function during the training process of each time phase, so that the loss term of short-term preference has a higher weight and the loss term of long-term preference has a lower weight;

[0063] An overall training module, which is used to perform overall training using user interaction data from all time phases in the final stage to obtain the model parameters of the final training;

[0064] A recommendation output module, which is used to output recommended items for the user's next interaction according to the model parameters of the final training.

[0065] To achieve the above object, the first aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is run by a processor, it executes the steps of the sequential recommendation method based on temporal curriculum learning.

[0066] Advantages of the present invention:

[0067] Compared with the prior art, a sequential recommendation method, system and medium based on temporal curriculum learning provided by the present invention effectively solves the problem of balancing short-term and long-term preferences and alleviates the problem of catastrophic forgetting by introducing a time series curriculum learning strategy and a reverse training strategy. Specifically, through a phased training process, the method first allows the model to learn short-term preferences, focusing on the most recent interaction data, and then gradually introduces historical interaction data, thereby gradually capturing long-term preferences. During this process, the reverse training strategy processes the data in reverse chronological order to ensure that the model does not lose historical information when learning new data and maintains the memory of early user preferences. In this way, the model can adapt to new user behaviors while retaining long-term interest trends, avoid catastrophic forgetting, and improve the accuracy of recommendations and user satisfaction. Different from traditional recommendation systems, this method has better time-dependent modeling ability, can effectively handle dynamically changing user interests, and improves the generalization ability and adaptability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments.

[0069] Figure 1 It is a flowchart of a sequential recommendation method based on temporal curriculum learning disclosed in an embodiment of the present invention.

[0070] Figure 2 It is a schematic diagram of data partitioning disclosed in an embodiment of the present invention.

[0071] Figure 3 It is an effect diagram of ablation experiment disclosed in an embodiment of the present invention. Detailed implementation manners

[0072] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0073] According to the embodiments of the present invention, it should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the following methods, in some cases, the steps shown or described can be executed in a different order than here.

[0074] As Figure 1 、 Figure 2 shown, the present invention provides a sequential recommendation method based on temporal curriculum learning, including the following steps:

[0075] Step S100: Sort the user interaction data according to the timestamp and divide it into multiple time stages according to the time period, where each time stage contains a set of user interaction data arranged in chronological order;

[0076] User interaction data refers to the data generated by various interaction behaviors of a user with the system when using a certain application or platform. These interaction behaviors can include actions such as user clicks, browsing, purchases, comments, likes, and shares, reflecting the interaction between the user and content, goods, or services. By collecting and analyzing this interaction data, the platform can understand the user's interests, needs, and behavior patterns, thereby providing a basis for the recommendation system, predicting the items or content that the user may be interested in in the future, and further optimizing the personalized recommendation effect. User interaction data usually contains information such as timestamps, user identifiers, item identifiers, and interaction types.

[0077] Sort the user interaction data according to the timestamp to ensure that the chronological order of each interaction record is preserved. Divide these interaction data into multiple time phases according to time periods, for example, it can be divided by time units such as days, weeks, or months. Each time phase contains a set of user interaction data arranged in chronological order, representing the user's behavior trajectory within that time period. The purpose of this time period division is to help the model capture the temporal evolution of user preferences and enable the data of each time phase to effectively reflect the user's interests and behavior changes within that time period, thereby providing ordered data input for subsequent training.

[0078] Step S200: Construct a training set for the user interaction data of each time phase, and convert the user interaction data of each time phase into interaction sequences of a fixed length;

[0079] Extract the interaction data arranged in chronological order within each time phase as the input data set for training the model. Then, convert the user interaction data of each time phase into interaction sequences of a fixed length.

[0080] A time phase refers to dividing the user interaction data into different time periods according to the time axis in a recommendation system. For example, it can be divided by time units such as days, weeks, months, etc., and each time period is an independent time phase. Within each time phase, the user's interaction behaviors are arranged in chronological order and represent the user's interest changes and behavior trajectory within that time period.

[0081] An interaction sequence refers to a sequence formed by arranging all the behaviors of a user interacting with the system within a certain time phase in chronological order. Each interaction sequence usually consists of multiple interaction events (such as clicks, views, purchases, etc.) and reflects the user's activity pattern within that time period.

[0082] The purpose of this step is to ensure that the data input into the model each time has the same length, which is convenient for the model to process. If the data length of a certain time phase exceeds the preset fixed length, only the latest interaction records are retained; if the data length is insufficient, it is supplemented by filling special markers (such as padding characters, placeholders) to ensure the unity and integrity of the sequence.

[0083] Preferably, the method for converting the user interaction data of each time phase into interaction sequences of a fixed length includes:

[0084] Step S201: Process the user interaction data of each time phase and arrange it in chronological order;

[0085] Step S202: According to the preset fixed sequence length, intercept the user interaction data of each time phase into several interaction sequences of a fixed length;

[0086] Step S203: If the length of the user interaction data in a certain time period exceeds the preset fixed sequence length, retain the most recent interaction records;

[0087] Step S204: If the length of the user interaction data in a certain time period is less than the preset fixed sequence length, pad the sequence to the fixed length by filling placeholders.

[0088] Step S300: Map the items in the fixed-length interaction sequence into embedding vector representations in the latent space through an embedding layer;

[0089] The embedding layer is a component in a neural network designed to convert discrete categorical data (such as item IDs) into continuous vector representations. In this step, the role of the embedding layer is to map each item (such as a product or content) in the interaction sequence into a low-dimensional continuous vector space. This process can be understood as learning the "semantic" representation of each item, so that the similarity and relationship between items can be reflected in the vector space.

[0090] In this step, each item in the interaction sequence refers to the specific item or content with which the user interacts in a certain time period, and each item has a unique identifier (such as an ID). To convert these items into vector representations, an embedding matrix is first defined for each item, where each row of the embedding matrix corresponds to the latent space embedding vector of an item. Next, the item IDs in each fixed-length interaction sequence are input into the embedding layer, and the item IDs are mapped into the corresponding embedding vector representations through this embedding matrix.

[0091] This mapping method helps the model capture the similarity between items by converting the original discrete item representations (such as product IDs) into continuous embedding vectors. In the latent space, the distance and direction of vectors can represent the correlation between items. For example, products that are semantically similar may be mapped to nearby positions in the latent space. Through this process, the embedding vectors provide rich semantic information for subsequent feature extraction and learning, thus helping the model better understand the user's preferences and behaviors.

[0092] Step S400: Inject chronological information, encode the temporal positions of the embedding vector representations in the latent space through a positional embedding layer, and form an embedding representation containing chronological features;

[0093] Temporal order information refers to the order of occurrence of each item in the user interaction sequence, that is, the temporal relationship between items. Since users' interests and behaviors usually change over time, by introducing temporal order information, the model can understand the evolution process of user preferences, thereby more accurately predicting users' future interests. For example, the items that a user has recently paid attention to are usually more in line with the current interests, while earlier interactions may represent the user's long-term preferences.

[0094] In this step, the role of the position embedding layer is to provide a "temporal position" for each item, that is, the relative position of the item in the interaction sequence. Different from the traditional embedding layer that converts item IDs into embedding vectors, the position embedding layer helps the model understand the temporal relationship of items in the interaction sequence by adding temporal order features to the embedding vectors of items.

[0095] Temporal position encoding combines the temporal order information with the embedding vector of the item according to the temporal position of the item in the interaction sequence (such as whether the item is the first, second, third interaction, etc.), forming an embedding representation containing temporal order features. The position embedding layer usually assigns a learnable vector to each temporal position, and these vectors are encoded according to the order of items in the sequence. By adding or concatenating these position embedding vectors with the embedding vector of the item, an embedding representation containing temporal order information is finally generated. This enables the model to better understand the position of each item in the interaction sequence, thereby capturing temporal dependencies and improving the accuracy and personalization effect of recommendations. The specific steps are as follows:

[0096] Step S401: Inject temporal order information into the embedding vectors in each interaction sequence through the position embedding layer;

[0097] Step S402: The position embedding layer uses learnable position embedding vectors to encode the embedding vector of each item according to the temporal position of the item in the interaction sequence, forming an embedding representation containing temporal order features;

[0098] Step S403: Add or concatenate the position embedding vector with the embedding vector of the item to generate a final embedding representation with temporal order features.

[0099] Step S500: Use the sequence feature extraction layer to extract implicit features from the embedding representation with temporal order features and generate a hidden representation;

[0100] The sequence feature extraction layer is the layer in the neural network responsible for extracting important features from the input interaction sequence. In this step, the main task of the sequence feature extraction layer is to process the embedded representation with time-order features, so as to capture the potential patterns in the user behavior sequence. To achieve this goal, the sequence feature extraction layer adopts models such as recurrent neural network (RNN), long short-term memory network (LSTM), gated recurrent unit (GRU), Transformer, etc. These models can effectively learn the temporal dependencies in the sequence, identify the interest changes of users at different time points, and the dynamic evolution of short-term and long-term preferences.

[0101] Latent features refer to the features that are automatically learned through the training process and are difficult to directly observe. These features contain the potential laws or patterns in the data. In sequence recommendation, latent features help the model understand the deep motivations behind user behaviors, such as why a user is interested in a specific product or content at a certain moment. Through the sequence feature extraction layer, the model can extract these latent features from the embedded representation containing time-order features, generating a set of vectors representing user behavior patterns and interest changes, and these vectors are called hidden representations. The hidden representation not only contains the user's current interests but also can reflect the evolution of the user's preferences throughout the interaction sequence.

[0102] The process of generating the hidden representation can be achieved through the following steps:

[0103] Step S501: Input the embedded representation containing time-order features into the sequence feature extraction layer;

[0104] Step S502: The sequence feature extraction layer extracts features from the embedded representation through at least one temporal model (such as RNN, LSTM, GRU, Transformer, etc.);

[0105] Step S503: The temporal model captures the long-term dependencies and short-term dynamic features in the user interaction sequence according to the time order, generating the corresponding latent feature representation.

[0106] Step S600: Input the generated hidden representation into the recommendation model for preliminary training, and update the model parameters through backpropagation to learn short-term preferences;

[0107] The hidden representation is passed to the recommendation model, and the recommendation model starts to learn short-term preferences and make predictions based on the user's latest behavior. During the preliminary training process, the model updates its parameters through the backpropagation algorithm. Specifically, by calculating the gradient of the loss function with respect to each parameter, the model's parameters are gradually adjusted to make the prediction results closer to the actual data.

[0108] The methods for learning short-term preferences include:

[0109] Step S601: Input the generated hidden representation into the recommendation model as the initial input of the recommendation model;

[0110] Step S602: The recommendation model processes the input hidden representation through forward propagation to generate the prediction result of the next interaction that the user may perform at the current time stage;

[0111] Step S603: According to the error between the user's real interaction data and the interaction result predicted by the recommendation model, use the backpropagation algorithm to calculate the error and update the model parameters;

[0112] Step S604: During the process of updating the model parameters through backpropagation, optimize the model to learn short-term preferences and capture the user's immediate interests within the current time period;

[0113] Step S605: Repeat the above steps for iterative training until the recommendation model can accurately predict the user's short-term preferences.

[0114] Step 700: In subsequent time stages, introduce historical interaction data step by step in reverse chronological order, gradually add the interaction data of earlier time stages to the training set, and use the trained model to update the model parameters to learn long-term preferences;

[0115] In subsequent time stages, the model will gradually introduce the interaction data of earlier time stages in reverse chronological order. First, the model uses the data of the latest time stage for training to learn short-term preferences. Then, during the training process, the model gradually introduces the previous interaction data. This process ensures that while learning new interaction data, the model gradually captures the long-term preferences hidden in the historical data. At each new time stage, the model does not need to train from scratch but uses the model parameters obtained from the previous training stage. In this way, the model can inherit the previous knowledge and, by introducing new historical data, gradually adjust and optimize its parameters. Each time new data is introduced, the model will fine-tune its parameters according to these data to learn the changes in the user's long-term preferences.

[0116] Long-term preferences refer to the stable interests and behavior patterns of users over a long period of time. By gradually introducing historical interaction data, the model can understand the user's long-term behavior trends and capture the deep-seated interests beyond short-term changes. During the model training process, earlier interaction data helps the model maintain attention to these long-term interests, avoid catastrophic forgetting, and effectively learn the user's long-term preferences.

[0117] Specifically, the methods for learning long-term preferences include:

[0118] Step S701: In subsequent time phases, gradually introduce the interaction data of earlier time phases in reverse chronological order, add the data of each phase to the training set, and gradually expand the data in chronological order;

[0119] Step S702: Use the recommendation model trained in the current time phase as the base model, load the model weights of the previous phase, and merge the data of the new time phase with the data of the previous phase for incremental training;

[0120] Step S703: In each time phase, keep the training results of the previous phase unchanged and only update the newly added part of the historical interaction data;

[0121] Step S704: Update the model parameters through the backpropagation algorithm to enable the model to learn long-term preferences and effectively capture the stable interest changes of users over a long time span;

[0122] Step S705: Repeat the above steps for iterative training until the recommendation model can accurately predict the long-term preferences of users.

[0123] Step S800: During the training process of each time phase, dynamically adjust the weights of the loss function so that the loss term of short-term preferences has a higher weight and the loss term of long-term preferences has a lower weight;

[0124] During the training process, the loss function of the model is divided into two parts for short-term and long-term preferences. The loss term of short-term preferences reflects the immediate interests of users in recent interactions, while the loss term of long-term preferences represents the stable interests of users over a longer period. To consider both in training, the system dynamically adjusts the loss weights of these two parts.

[0125] At the beginning of training, the model mainly relies on the latest user behavior data to learn short-term preferences. Therefore, a higher weight is given to the loss term of short-term preferences so that the model can quickly adapt to the current interests and needs of users. This weighting strategy ensures that the model can effectively learn and optimize immediate interests when facing new interaction data.

[0126] During the training process of each time phase, since long-term preferences represent the long-term stable interests of users, their changes are usually relatively slow. Therefore, a lower weight is given to the loss term of long-term preferences. This strategy helps the model avoid over-relying on historical data, enabling the model to gradually learn the long-term preferences of users while paying attention to current interests without being overly interfered by historical interactions.

[0127] Dynamically adjusting the weights of the loss function enables the model to flexibly adjust its learning focus according to the characteristics of the current data. Giving more attention to short-term preferences can quickly respond to the current needs of users, while giving appropriate attention to long-term preferences can ensure that the model captures the long-term interest evolution of users.

[0128] Step S900: In the final stage, use the user interaction data of all time stages for overall training to obtain the model parameters of the final training.

[0129] In the final stage, the model will review and integrate the user interaction data of all previous time stages. Different from the previous stages that only focus on the interaction data of the current time period, the training in the final stage uses the complete time series data set, which includes all user interaction records from the earliest to the most recent, covering all aspects of short-term and long-term preferences.

[0130] When using the data of all time stages for overall training, the model will learn from the entire interaction sequence. At this time, the model no longer gradually introduces historical data, but uses the knowledge learned in all stages to further train on the full amount of data. The goal of overall training is to optimize the model on the interaction data of multiple stages in order to accurately capture the preference evolution and interest changes of users. During the overall training process, the model will update and optimize its parameters according to the data of all time stages. Through the backpropagation and gradient descent algorithms, the parameters of the model will be adjusted according to the error of the entire data set to reduce the overall loss. Finally, after sufficient training, the model can extract more comprehensive features from all historical interaction data and improve its recommendation accuracy through parameter optimization. After the complete training process, the final model parameters will contain the learning results of the user preferences at all time stages.

[0131] Step S1000: According to the model parameters of the final training, output the recommended items for the user's next interaction.

[0132] After calculating the items that the user may be interested in, the model will output the recommended items according to its predicted results. The recommended items can be products, content, services, etc., depending on the application scenario of the recommendation system.

[0133] In step S600, calculate the gradient of the loss function with respect to each parameter through the backpropagation algorithm, and gradually adjust the parameters of the model to optimize the prediction ability of the model. The following is the specific process of backpropagation:

[0134] In the forward propagation process, the input data (such as the generated hidden representation) will be processed through the layers of the model, and finally the prediction results of the model will be output. These prediction results are calculated according to the current model parameters.

[0135] The difference between the output of the model and the actual user behavior data (such as the items that the user actually interacts with) is used to calculate the loss function. The loss function is a criterion for measuring the difference between the model prediction and the true label. Commonly used loss functions include mean squared error (MSE) and cross-entropy loss, etc.

[0136] In this embodiment, binary cross-entropy loss is used to measure whether a user will interact with a certain product or content. Assuming that the output of the model is the predicted interaction probability, the actual interaction is 1, and the non-interaction is 0, the loss function can be calculated according to the prediction result and the actual behavior.

[0137] Once the loss is calculated, the backpropagation algorithm starts to calculate the gradient of each parameter with respect to the loss function through the chain rule. This means that backpropagation will calculate the contribution of each model parameter (such as embedding vectors, weights, etc.) to the loss function. Specifically, backpropagation first calculates the gradient of the loss function with respect to the output layer, and then passes these gradients layer by layer to the previous layers. The gradient value of each layer indicates how the parameters of that layer affect the loss function. After calculating the gradient of each parameter, the model updates the parameters according to the gradient descent algorithm. Gradient descent adjusts the parameters in the opposite direction of the gradient to reduce the loss.

[0138] Specifically, each parameter will be updated according to the following formula:

[0139]

[0140] where θ represents the parameters of the model, η is the learning rate, is the gradient of the loss function with respect to this parameter.

[0141] In the method of the present invention, the preliminary training learns short-term preferences through this backpropagation process. Since short-term preferences are usually related to the most recent interaction data, the model adjusts the parameters according to the loss of these latest interaction data to quickly adapt to the user's current interests and behaviors.

[0142] By continuously iterating this process, the model gradually optimizes its parameters during the training process, making the prediction results more accurate and better able to capture the user's short-term preferences.

[0143] Preferably, during the training process, a cross-entropy loss function is used and a loss term for contrastive learning is introduced, specifically including:

[0144] Within each time stage, the cross-entropy loss function is used to measure the error between the user interaction data and the model prediction result, and the correlation score of each candidate item is calculated;

[0145] A loss term for contrastive learning is introduced to optimize the model parameters by comparing the differences between positive samples and negative samples;

[0146] Use the following formula to calculate the loss function:

[0147]

[0148] where, represents the total loss function of the recommendation system, represents the summation over all users u in the user set j represents the index of the item or event, and represent the positive sample and the negative sample respectively. Among them, the positive sample is the item that the user actually clicks, and the negative sample is the non-clicked item randomly sampled. represents the interaction sequence of user u before the j-th step, which is used to describe the user's historical behavior. σ represents the Sigmoid activation function, which is used to map the prediction score to the range (0, 1). represents the probability that the user interacts with the positive sample under the condition of the given interaction sequence of user u before the j-th step, represents the probability that the user interacts with the negative sample under the condition of the given interaction sequence of user u before the j-th step;

[0149] Update the model parameters through the backpropagation of this loss function.

[0150] To verify the effectiveness of the method of the present invention, two main evaluation methods are adopted in the present invention: leave-one-out evaluation and global timeline evaluation.

[0151] Leave-one-out evaluation: This method uses the last interaction of each user as the test set, and all the previous interaction data of the user is used for training and verification. This method ensures that the model can rely on the user's historical behavior during training and make accurate predictions for the last interaction.

[0152] Global timeline evaluation: In this method, the training set, validation set, and test set are divided according to the time order to ensure that the data selection of the test set comes from the later interaction records in time. This way avoids the problem of data leakage and thus improves the reliability of the evaluation.

[0153] For evaluation, four real-world data sets (including "Toys", "Sports", "Beauty", and "Games") in the Amazon review data set and the movie review data set MI-1M are selected. In addition, SASRec, LinRec, and Mamba4Rec are also used as the baseline models, and the NDCG and HIT metrics are used for performance comparison, as shown in Tables 1 and 2:

[0154] Table 1 shows the evaluation results of the leave-one-out method. "w / TCL4Rec" represents the method of the present invention

[0155]

[0156] Table 2 shows the evaluation results of the global timeline. "w / TCL4Rec" represents the method of the present invention

[0157]

[0158] The experimental results show that the method of the present invention (w / TCL4Rec) significantly improves the recommendation performance of the basic model on these two metrics, verifying its advantages in recommendation accuracy and personalization.

[0159] To further verify the effectiveness of the method of the present invention, ablation experiments were also conducted. The specific experiments include the following variants:

[0160] 1. Reverse order within groups: In this experiment, the items are arranged in reverse order according to the timestamp, but the training order between groups remains unchanged.

[0161] 2. No grouping (k = 1): The items are sorted from earliest to latest according to the timestamp, and no grouped training is performed.

[0162] 3. Unordered training (common training): The interaction order of users is randomly shuffled for training.

[0163] 4. Only current training (1 / k): Only 1 / k of the data is used for training each time, and no overlapping training is performed, which may lead to the problem of catastrophic forgetting.

[0164] As Figure 3 shown, the experimental results show that CCLSR performs best in all variants, demonstrating its excellent ability to accurately model temporal dependencies and the evolution of user preferences. These results further prove the effectiveness and superiority of the method of the present invention, which can significantly improve the performance of the recommendation system and user satisfaction.

[0165] According to another aspect of the embodiments of the present application, a sequential recommendation system based on temporal curriculum learning is also provided, including the following modules:

[0166] A user interaction data receiving module, configured to sort the user interaction data according to the timestamp;

[0167] A time period division module, which divides the time period into multiple time stages, where each time stage contains a set of user interaction data arranged in chronological order;

[0168] A training set construction module, which is used to construct a training set for the user interaction data of each time stage, and convert the user interaction data of each time stage into an interaction sequence of a fixed length;

[0169] An embedding layer module, which is used to map the items in the interaction sequence of a fixed length into an embedding vector representation in the latent space through the embedding layer;

[0170] A positional embedding layer module, which is used to encode the time positions of the embedding vector representations in the latent space through the positional embedding layer to form an embedding representation containing temporal order features;

[0171] A sequence feature extraction layer module, which is used to extract implicit features from the embedding representation with temporal order features by using the sequence feature extraction layer to generate a hidden representation;

[0172] A recommendation model module, which is used to input the generated hidden representation into the recommendation model for preliminary training, and update the model parameters through backpropagation to learn short-term preferences;

[0173] A historical data introduction module, which is used to gradually introduce historical interaction data in reverse chronological order in subsequent time stages, gradually add the interaction data of earlier time stages to the training set, and update the model parameters using the trained model to learn long-term preferences;

[0174] A dynamic adjustment module, which is used to dynamically adjust the weights of the loss function during the training process of each time stage, so that the loss term of short-term preferences has a higher weight and the loss term of long-term preferences has a lower weight;

[0175] An overall training module, which is used to perform overall training using the user interaction data of all time stages in the final stage to obtain the model parameters of the final training;

[0176] A recommendation output module, which is used to output the recommended items for the user's next interaction according to the model parameters of the final training.

[0177] According to another aspect of the embodiments of the present application, an electronic device is further provided, including a processor and a memory, and the processor is used to implement the steps of the method when executing the computer program stored in the memory.

[0178] In the above embodiments of the present invention, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0179] In several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of units or modules can be in an electrical or other form.

[0180] In addition, each functional unit in various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0181] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0182] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A sequential recommendation method based on temporal curriculum learning, characterized in that, It includes the following steps: Sort the user interaction data by timestamp and divide it into multiple time phases according to time periods, where each time phase contains a set of user interaction data arranged in chronological order; Construct a training set for the user interaction data of each time phase and convert the user interaction data of each time phase into an interaction sequence of a fixed length; Map the items in the interaction sequence of a fixed length into an embedded vector representation in the latent space through an embedding layer; Inject chronological information and encode the time positions of the embedded vector representations in the latent space through a positional embedding layer to form an embedded representation containing chronological features; Use a sequence feature extraction layer to extract implicit features from the embedded representation with chronological features to generate a hidden representation; Input the generated hidden representation into a recommendation model for preliminary training and update the model parameters through backpropagation to learn short-term preferences; In subsequent time phases, introduce historical interaction data step by step in reverse chronological order, gradually add the interaction data of earlier time phases to the training set, and update the model parameters using the trained model to learn long-term preferences; During the training process of each time phase, dynamically adjust the weights of the loss function so that the loss term of short-term preferences has a higher weight and the loss term of long-term preferences has a lower weight; In the final stage, perform overall training using the user interaction data of all time phases to obtain the model parameters of the final training; According to the model parameters of the final training, output the recommended items for the user's next interaction.

2. The sequential recommendation method based on temporal curriculum learning according to claim 1, wherein The method of converting the user interaction data of each time phase into an interaction sequence of a fixed length includes: Process the user interaction data of each time phase and arrange it in chronological order; According to the preset fixed sequence length, intercept the user interaction data of each time phase into several interaction sequences of a fixed length; If the length of the user interaction data in a certain time phase exceeds the preset fixed sequence length, retain the most recent interaction records; If the length of the user interaction data in a certain time phase is less than the preset fixed sequence length, fill in placeholders to complete the sequence to the fixed length.

3. The sequential recommendation method based on temporal curriculum learning according to claim 1, wherein The method of mapping the items in the interaction sequence of a fixed length into an embedded vector representation in the latent space includes: Define an embedding matrix for each item, where each row in the embedding matrix corresponds to the embedded vector of an item in the latent space; Input the item IDs in each interaction sequence of a fixed length into the embedding layer, and map the item IDs into the corresponding embedded vector representations through this embedding matrix.

4. The sequential recommendation method based on temporal curriculum learning according to claim 1, wherein The method of forming an embedded representation containing chronological features includes: Inject chronological information into the embedded vectors in each interaction sequence through a positional embedding layer; The positional embedding layer uses learnable positional embedding vectors to encode each item's embedded vector according to the item's time position in the interaction sequence to form an embedded representation containing chronological features; Add or concatenate the positional embedding vectors with the item's embedded vectors to generate the final embedded representation with chronological features.

5. The sequential recommendation method based on temporal curriculum learning according to claim 1, characterized in that, The method of generating a hidden representation includes: Input the embedded representation containing chronological features into a sequence feature extraction layer; The sequence feature extraction layer extracts features from the embedded representation through at least one temporal model; The temporal model captures long-term dependencies and short-term dynamic features in the user interaction sequence according to the time order, and generates corresponding hidden feature representations.

6. The sequential recommendation method based on temporal curriculum learning according to claim 1, wherein Methods for learning short-term preferences include: Input the generated hidden representation into the recommendation model as the initial input of the recommendation model; The recommendation model processes the input hidden representation through forward propagation to generate a prediction result of the next interaction that the user may perform at the current time stage; According to the error between the user's real interaction data and the interaction result predicted by the recommendation model, use the backpropagation algorithm to calculate the error and update the model parameters; During the process of updating the model parameters through backpropagation, optimize the model to learn short-term preferences and capture the user's immediate interests within the current time period; Repeat the above steps for iterative training until the recommendation model can accurately predict the user's short-term preferences.

7. The sequential recommendation method based on temporal curriculum learning according to claim 1, wherein Methods for learning long-term preferences include: In subsequent time stages, gradually introduce the interaction data of earlier time stages in reverse chronological order, add the data of each stage to the training set, and gradually expand the data in chronological order; Use the recommendation model that has been trained at the current time stage as the base model, load the model weights of the previous stage, and merge the new time stage data with the data of the previous stage for incremental training; At each time stage, keep the training results of the previous stage unchanged and only update the newly added historical interaction data part; Update the model parameters through the backpropagation algorithm to enable the model to learn long-term preferences and effectively capture the stable interest changes of users over a long time span; Repeat the above steps for iterative training until the recommendation model can accurately predict the user's long-term preferences.

8. The sequential recommendation method based on temporal curriculum learning according to claim 1, wherein The method also includes using a cross-entropy loss function and introducing a loss term for contrastive learning during the training process, specifically including: Within each time stage, use the cross-entropy loss function to measure the error between the user interaction data and the model prediction result, and calculate the correlation score of each candidate item; Introduce a loss term for contrastive learning, and optimize the model parameters by comparing the differences between positive and negative samples; Use the following formula to calculate the loss function: Among them, represents the total loss function of the recommendation system, represents the summation over all users u in the user set where j represents the index of an item or event, and represent the positive sample and the negative sample respectively, represents the interaction sequence of user u before the j-th step, which is used to describe the historical behavior of the user. σ represents the Sigmoid activation function, which is used to map the predicted score to the range (0, 1), represents the probability that the user interacts with the positive sample given the interaction sequence of user u before the j-th step, represents the probability that the user interacts with the negative sample given the interaction sequence of user u before the j-th step; Update the model parameters through the backpropagation of this loss function.

9. A sequential recommendation system based on temporal curriculum learning, characterized in that, It includes the following modules: A user interaction data receiving module for sorting the user interaction data according to timestamps; A time period division module that divides according to time periods into multiple time stages, where each time stage contains a set of user interaction data arranged in chronological order; A training set construction module for constructing a training set for the user interaction data of each time stage and converting the user interaction data of each time stage into an interaction sequence of a fixed length; An embedding layer module for mapping the items in the fixed-length interaction sequence into an embedding vector representation in the latent space through the embedding layer; A positional embedding layer module for encoding the time positions of the embedding vector representations in the latent space through the positional embedding layer to form an embedding representation containing chronological order features; The sequence feature extraction layer module is used to extract implicit features from the embedded representation with chronological features by using the sequence feature extraction layer, and generate a hidden representation; The recommendation model module is used to input the generated hidden representation into the recommendation model for preliminary training, and update the model parameters through backpropagation to learn short-term preferences; The historical data introduction module is used to introduce historical interaction data step by step in reverse chronological order in subsequent time stages, gradually add the interaction data in earlier time stages to the training set, and update the model parameters using the trained model to learn long-term preferences; The dynamic adjustment module is used to dynamically adjust the weights of the loss function during the training process of each time stage, so that the loss term of short-term preferences has a higher weight and the loss term of long-term preferences has a lower weight; The overall training module is used to perform overall training using the user interaction data of all time stages in the final stage to obtain the model parameters of the final training; The recommendation output module is used to output the recommended items for the user's next interaction according to the model parameters of the final training.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is run by a processor, it executes the steps of the sequential recommendation method based on temporal curriculum learning according to any one of claims 1 to 8.