A Contrastive Learning Sequence Recommendation Method Based on Large Language Model View Enhancement

The method addresses dynamic user interest changes in sequence recommendations by using large language models for adaptive sequence trimming and dual-view contrastive learning, enhancing accuracy and adaptability in diverse scenarios.

CN120123600BActive Publication Date: 2025-07-15NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510625192.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-07-15
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

The prior art is difficult to effectively capture the dynamic evolution characteristics and timing dependencies of user interests, and is not robust in sparse data and cold start scenarios, and traditional methods are difficult to adapt to the complexity of user behavior.

Method used

By establishing a dynamic importance scoring mechanism and adaptive sequence cropping method, a large language model is used to generate enhanced views, and a multimodal dual-view enhanced contrast learning framework is constructed, and user interest representations are extracted in combination with linear recursive units to optimize contrast loss to enhance the model's perception ability.

Benefits of technology

The characterization accuracy and recommendation accuracy of user behavior sequences are improved, especially in sparse data and cold start scenarios, which significantly enhances the adaptability and robustness of the model, reduces computational overhead and improves training efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123600B_ABST
    Figure CN120123600B_ABST
Patent Text Reader

Abstract

The present invention provides a contrastive learning sequence recommendation method based on large language model view enhancement, including: Step 1, establishing a dynamic importance scoring mechanism, using a large language model to calculate importance and generate a dynamic adjustment factor according to the hidden layer states of user and item historical interaction sequences; Step 2, generating an enhanced view containing negative sample pairs; Step 3, constructing a sequence and item dual-granularity contrastive learning method under multi-modal dual-view enhancement; Step 4, inputting the original user and item historical interaction sequences into a linear recurrent unit; Step 5, dynamically fusing the contrastive learning method and the conventional sequence recommendation method, optimizing through a weighted combination loss function, and jointly training the model through backpropagation. The method of the present invention improves resource utilization while ensuring the robustness of the model in cold start item recommendation and long-tail distribution scenarios, providing an efficient and stable solution for dynamic recommendation systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of recommendation systems, and particularly relates to a contrastive learning sequential recommendation method based on large language model view enhancement. Background Art

[0002] The sequential recommendation model is a deep learning method widely used in the field of personalized recommendation, and has made remarkable progress in multiple scenarios such as user purchase prediction, web content recommendation, and point of interest navigation, becoming a core part of modern recommendation system research.

[0003] The execution paradigm of traditional recommendation systems is mainly based on static feature modeling. Its core logic is to achieve recommendations through the association and matching of user portraits and product features. Such methods rely on two stages: offline training and online inference of static data. In the training stage, fixed patterns of user interests are learned through historical interaction data, and in the inference stage, recommendation results are generated based on real-time inputs. However, this static modeling method fails to fully consider the dynamic evolution characteristics of user interests and the temporal dependence relationships in the behavior sequence, resulting in limited recommendation effects in actual scenarios such as interest drift and scenario migration. For example, in the e-commerce platform scenario, users may exhibit completely different shopping patterns on weekdays and weekends, and traditional methods are difficult to capture such short-term interest fluctuations.

[0004] In recent years, large language pre-trained models can capture potential semantic associations across domains through pre-training learning of massive behavioral data. Its multi-level attention mechanism can model both long-term interest preferences and short-term behavioral motivations, providing a new paradigm for refined user portraits. However, directly applying large language models to sequential recommendation still faces significant challenges: First, there are essential differences between the spatio-temporal sparsity of user behavior sequences and the dense semantics of natural language texts; second, the contradiction between the model parameter scale and real-time inference requirements is particularly prominent in mobile scenarios; third, how to effectively fuse context features and real-time behavioral signals still needs to be explored in depth.

[0005] Recommendation methods based on contrastive learning have effectively improved the representation quality of user behavior sequences and alleviated the data sparsity problem. Although existing research has optimized the data distribution to a certain extent, it is difficult to dynamically adapt to complex and changing recommendation scenarios. Traditional methods are limited by the manually designed generation logic, and have insufficient ability to capture potential deep semantic associations and cross-scenario generalization features in user behavior, resulting in limited robustness of the model in sparse data and cold start scenarios.

[0006] Therefore, how to construct an adaptive contrastive view generation mechanism based on large language models, break through the constraints of preset rules, and achieve semantic-driven dynamic enhancement has become the core breakthrough point for improving sequential recommendation performance. Summary of the Invention

[0007] Purpose of the Invention: The technical problem to be solved by the present invention is to provide a contrastive learning sequence recommendation method based on large language model view enhancement in view of the deficiencies of the prior art, including the following steps:

[0008] Step 1, establish a dynamic importance scoring mechanism, use a large language model, calculate the importance and generate a dynamic adjustment factor according to the hidden layer states of the user and item historical interaction sequences;

[0009] Step 2, establish an adaptive sequence cropping method, use the attention mechanism for sequence data enhancement, dynamically adjust the cropping position through semantic understanding, and generate an enhanced view containing negative sample pairs according to the dynamic adjustment factor;

[0010] Step 3, construct a sequence and item dual-granularity contrastive learning method under multi-modal dual-view enhancement, build a contrastive learning task based on the enhanced view generated in Step 2, and optimize the contrastive loss to strengthen the contrastive learning method's perception ability of key behavior differences by narrowing the representation distance of positive sample pairs and pushing away the similarity of negative sample pairs;

[0011] Step 4, input the original user and item historical interaction sequences into a linear recurrent unit LRU (Linear Recurrent Units), extract the user's dynamic interest representation and predict the probability distribution of the next interaction item, and synchronously calculate the loss function of conventional sequence recommendation;

[0012] Step 5, dynamically fuse the contrastive learning method and the conventional sequence recommendation method, optimize through weighted combination of loss functions, and jointly train the model through backpropagation.

[0013] Step 1 includes: constructing the user and item historical interaction sequences , where represents the i-th element in the sequence, is the sequence length;

[0014] Process the user and item historical interaction sequence S through the pre-trained large language model BERT to establish a dynamic importance scoring mechanism, and the dynamic importance scoring mechanism performs the following operation process:

[0015] Step 1-1, convert the user and item historical interaction sequence S into word vectors, including:

[0016] Input data: Input the user's historical interaction records;

[0017] Clean data: Remove invalid interactions (such as duplicates, outliers);

[0018] Sequence segmentation: Split the long sequence into fixed-length segments of 256;

[0019] Build a global item pool: Count all the items that have appeared and generate unique IDs;

[0020] Vocabulary mapping: Assign an index to each item ID;

[0021] Initialize the embedding matrix: Create a random matrix of shape (B1, B2), where B1 represents the vocabulary size and B2 represents the embedding dimension;

[0022] Sequence vectorization: Convert the user and item historical interaction sequence S into a list of indices, and then map it to a vector sequence through the embedding matrix to obtain a standardized vector;

[0023] Input to the embedding layer: Input the standardized vector sequence into the embedding layer;

[0024] Step 1-2, Extract hidden layer states through multiple layers of Transformer,

[0025] Step 1-3, Calculate the position importance score along the feature dimension for subsequent information pruning; A large variance of the hidden layer state across different layers indicates a high importance score, while a small fluctuation of the hidden layer state indicates a low importance score;

[0026] Step 1-4, Generate a dynamic adjustment factor through the aggregation of multiple layers of hidden layer states, which essentially fuses features at different abstraction levels and indirectly reflects the position importance. The formula is:

[0027] ,

[0028] where, represents the dynamic adjustment factor of the i-th sample, represents the hidden layer state matrix output by the k-th layer of the pre-trained large language model BERT, and d is the number of layers of the BERT model.

[0029] Step 2 includes:

[0030] Apply the adjustment factor directly to the position encoding to obtain the pruned dynamic start position and the elastic pruning length :

[0031] ,

[0032] ,

[0033] ,

[0034] where, represents the pruning ratio parameter, represents the original sequence length; Denotes uniformly distributed random sampling, Denotes the position encoding matrix of the enhanced sequence; Denotes a sequence of integers from 1 to ;

[0035] Then, through the calculated cropped dynamic start position and the elastic cropping length the enhanced sequence is calculated, and the formula is:

[0036] ,

[0037] where Denotes the original user and item historical interaction sequence; Denotes the intercepted sequence part starting from with a length of .

[0038] Step 3 includes:

[0039] Using the original sequence S and the enhanced sequence as positive and negative samples respectively for contrastive learning;

[0040] The sequence-level alignment method captures the global temporal pattern of user behavior (such as the interest evolution path) by extracting the end position embedding of the enhanced sequence. The loss function of sequence-level alignment is as follows:

[0041] ,

[0042] where N represents the total number of samples in a training batch, is the sequence-level embedding vector, denotes the positive example-level embedding of the i-th sample, denotes the j-th negative example embedding vector, is an adjustable temperature coefficient, is the cosine similarity, exp represents the natural exponential function, and K represents the total number of negative examples to be compared for each anchor point;

[0043] Using the SASRec model as the base model, with the position encoding matrix and the enhanced sequence as parameters, input into the SASRec model to obtain the vector embedding of the enhanced sequence;

[0044] SASRec (Self-Attention-based Sequential Recommendation) model: This model uses the self-attention mechanism to comprehensively consider the information of the entire sequence and can effectively capture the long-term dependencies in user behavior.

[0045] Using the Transformer encoding layer in the SASRec model, learn the relationships between items implicitly, achieve item-level alignment, construct a local item similarity matrix through attention weights, and the loss function for item-level alignment is:

[0046] ,

[0047] where represents the length of the interaction sequence between the user and the item, is the hidden state output by the Transformer of the q-th item, is the local contrast representation of the q-th item; is the hidden state after encoding the same augmented variant generated by sequence shuffling through the Transformer; Attn(x,y) is the attention function used to calculate the similarity between x and y; represents and similarity;

[0048] Finally, adopt the dual-granularity contrast loss for joint optimization to obtain the loss of contrastive learning :

[0049] ,

[0050] where, , represent the weight coefficients used to balance the contributions of the global and local losses.

[0051] Step 4 includes: Initialize the hidden state with the original sequence S as the input, set the initial state , and update the hidden state:

[0052] ,

[0053] where, is the input vector at time step t, has a dimension of , is the hidden layer state vector at time step t, has a dimension of , is the weight matrix of the hidden state, is the input weight matrix, is the bias; represents the real number space;

[0054] obtain the interest at each time step , where T represents the total number of time steps of the input interaction sequence, and take the final state as the user's final interest for predicting the next interaction;

[0055] Map the hidden state to the item space:

[0056] ,

[0057] where is the mapping result in the item space corresponding to time step t, is the weight matrix, is the bias term;

[0058] Generate the predicted probability of the user's next interaction with each item through the Softmax function :

[0059] .

[0060] Step 4 also includes: calculating the loss function of conventional sequence recommendation :

[0061] ,

[0062] where is the true label, is the number of candidate items, represents the indicator value corresponding to item c in the true label, is the output predicted probability, represents the weight coefficient of the c-th item.

[0063] Step 5 includes:

[0064] Dynamically fuse the contrastive learning method and the conventional sequence recommendation method, and the formula for the total loss function is:

[0065] ,

[0066] where is the weight hyperparameter of the contrastive loss.

[0067] Step 5 also includes: updating the model parameters by gradient descent , minimizing the total loss:

[0068] ,

[0069] where, denotes the optimal model parameters obtained after optimization, denotes finding the parameters that minimize the value of the expression in .

[0070] The present invention also provides an electronic device, including a processor and a memory. The memory stores program code, and when the program code is executed by the processor, the processor is caused to execute the steps of the method described above.

[0071] The present invention also provides a storage medium storing a computer program or instruction, and when the computer program or instruction runs on a computer, the steps of the method described above are executed.

[0072] The present invention has the following beneficial effects: (1) The contrastive learning sequence recommendation method based on large language model view enhancement proposed by the present invention combines the semantic understanding ability of the large language model through a dynamic importance scoring mechanism, quantifies the key differences in user historical behavior in real time, and effectively filters out noise interference. Compared with the traditional static weight allocation method, this method can adaptively identify the significant change nodes of user interests, improve the representation accuracy of the behavior sequence, and is especially suitable for scenarios where user behavior is sparse or there is implicit feedback.

[0073] (2) The adaptive sequence cropping and dual-view contrastive learning framework designed by the present invention generates diverse enhanced views through a semantics-driven dynamic cropping strategy, and constructs sequence-level and item-level contrast tasks. Compared with the traditional uniform sampling method, this framework significantly enhances the model's ability to capture long-term and short-term user interests, and at the same time solves the data sparsity problem through contrastive loss optimization, improving the recommendation accuracy.

[0074] (3) The linear recurrence unit (LRU) and temporal hidden state extraction technology introduced by the present invention accurately models the dynamic evolution process of user interests by fusing time-sensitive feature encoding and semantic enhancement representation of the large language model. Compared with the traditional RNN or Transformer architecture, this method reduces the computational overhead by 30% in training efficiency and improves the adaptability to user interest drift scenarios by more than 40%.

[0075] (4) The dynamic multi-task fusion mechanism implemented by the present invention adaptively balances the self-supervised signal and the supervised signal by weighted joint optimization of the contrastive learning loss and the recommendation task loss. Compared with the fixed weight strategy, this mechanism improves the resource utilization rate, and at the same time ensures the robustness of the model in cold start item recommendation and long-tail distribution scenarios, providing an efficient and stable solution for the dynamic recommendation system. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] Figure 1 is a flowchart of the method of the present invention. Detailed implementation manners

[0077] The following further specifically describes the present invention in conjunction with the accompanying drawings and specific implementation manners, and the above and / or other advantages of the present invention will become clearer.

[0078] An embodiment of the present invention provides a contrast learning sequence recommendation method based on large language model view enhancement, as Figure 1 shown, specifically including the following steps:

[0079] Step 1, establish a dynamic importance scoring mechanism, and use a large language model to calculate importance and generate a dynamic adjustment factor according to the hidden layer states of the user and item historical interaction sequences.

[0080] Step 2, establish an adaptive sequence cropping method, use the attention mechanism for sequence data enhancement, dynamically adjust the cropping position through semantic understanding, and generate an enhanced view containing negative sample pairs according to the dynamic adjustment factor.

[0081] Step 3, construct a sequence and item dual-granularity contrast learning method under multi-modal dual-view enhancement, build a contrast learning task based on the enhanced view generated in Step 2, and optimize the contrast loss to strengthen the contrast learning method's perception ability of key behavior differences by narrowing the representation distance of positive sample pairs and pushing away the similarity of negative sample pairs.

[0082] Step 4, input the original user and item historical interaction sequences into a linear recurrent unit LRU (Linear Recurrent Units), extract the user's dynamic interest representation and predict the probability distribution of the next interaction item, and synchronously calculate the loss function of conventional sequence recommendation.

[0083] Step 5, dynamically fuse the contrast learning method and the conventional sequence recommendation method, optimize through weighted combination of loss functions, and jointly train the model through backpropagation.

[0084] Step 1 includes:

[0085] Construct the user and item historical interaction sequences , where represents the i-th element in the sequence, is the sequence length;

[0086] Process the user and item historical interaction sequence S through a pre-trained large language model BERT to establish a dynamic importance scoring mechanism, and the dynamic importance scoring mechanism performs the following operation process:

[0087] Step 1-1, convert the user and item historical interaction sequence S into word vectors, including:

[0088] Input data: Input the historical interaction records of the user;

[0089] Data cleaning: Remove invalid interactions (such as duplicates, outliers);

[0090] Sequence segmentation: Split long sequences into fixed-length segments of 256;

[0091] Construct the global item pool: Count all the items that have appeared and generate unique IDs;

[0092] Vocabulary mapping: Assign an index to each item ID;

[0093] Initialize the embedding matrix: Create a random matrix of shape (B1, B2), where B1 represents the vocabulary size and B2 represents the embedding dimension;

[0094] Sequence vectorization: Convert the user and item historical interaction sequence S into a list of indices, and then map it to a vector sequence through the embedding matrix to obtain the standardized vector;

[0095] Input to the embedding layer: Input the standardized vector sequence into the embedding layer;

[0096] Steps 1-2, Extract the hidden layer states through multiple layers of Transformer,

[0097] Steps 1-3, Calculate the position importance scores along the feature dimension for subsequent information pruning; A large variance of the hidden layer states across different layers indicates a high importance score, while a small fluctuation of the hidden layer states indicates a low importance score;

[0098] Steps 1-4, Generate a dynamic adjustment factor through the aggregation of multiple layers of hidden layer states, which essentially fuses features at different abstraction levels and indirectly reflects the position importance. The formula is:

[0099] ,

[0100] where, represents the dynamic adjustment factor of the i-th sample, represents the hidden layer state matrix output by the k-th layer of the pre-trained large language model BERT, and d is the number of layers of the BERT model.

[0101] Step 2 includes:

[0102] Apply the adjustment factor directly to the position encoding to obtain the cropped dynamic start position and the elastic cropping length :

[0103] ,

[0104] ,

[0105] ,

[0106] Among them, represents the cropping ratio parameter, represents the length of the original sequence; represents uniform distribution random sampling, represents the position encoding matrix of the augmented sequence; represents an integer sequence from 1 to ;

[0107] Then, through the calculated dynamic start position of cropping and the elastic cropping length the augmented sequence is calculated, and the formula is:

[0108] ,

[0109] where represents the original user and item historical interaction sequence.

[0110] Step 3 includes:

[0111] Using the original sequence S and the augmented sequence as positive and negative samples respectively for contrastive learning;

[0112] The sequence-level alignment method captures the global temporal pattern of user behavior (such as the interest evolution path) by extracting the end position embedding of the augmented sequence. The loss function of sequence-level alignment has the formula:

[0113] ,

[0114] where N represents the total number of samples included in a training batch, is the sequence-level embedding vector, represents the positive example-level embedding of the i-th sample, represents the j-th negative example embedding vector, is the adjustable temperature coefficient, is the cosine similarity, exp represents the natural exponential function, and K represents the total number of negative examples to be contrasted for each anchor ;

[0115] Using the SASRec model as the basic model, with the position encoding matrix and the augmented sequence as parameters, input into the SASRec model to obtain the vector embedding of the augmented sequence;

[0116] SASRec (Self-Attention-based Sequential Recommendation) model: This model uses the self-attention mechanism to comprehensively consider the entire sequence information and can effectively capture the long-term dependencies in user behavior.

[0117] Using the Transformer encoding layer in the SASRec model, learn the relationships between items implicitly, achieve item-level alignment, construct a local item similarity matrix through attention weights, and the loss function for item-level alignment is:

[0118] ,

[0119] where represents the length of the interaction sequence between the user and the item, is the hidden state output by the Transformer for the q-th item, is the local contrast representation of the q-th item; is the hidden state after Transformer encoding of the same augmented variant generated by sequence shuffling; Attn(x,y) is the attention function used to calculate the similarity between x and y; represents and similarity;

[0120] Finally, adopt dual-granularity contrast loss for joint optimization to obtain the loss of contrastive learning :

[0121] ,

[0122] where, , represent the weight coefficients used to balance the contributions of the global and local losses.

[0123] Step 4 includes: Initialize the hidden state with the original sequence S as the input, set the initial state , and update the hidden state:

[0124] ,

[0125] where, is the input vector at time step t, has a dimension of , is the hidden layer state vector at time step t, has a dimension of , is the weight matrix of the hidden state, is the input weight matrix, is the bias; represents the real number space;

[0126] obtain the interest at each time step , where T represents the total number of time steps of the input interaction sequence, and take the final state as the user's final interest for predicting the next interaction;

[0127] Map the hidden state to the item space:

[0128] ,

[0129] where is the mapping result in the item space corresponding to time step t, is the weight matrix, is the bias term;

[0130] Generate the predicted probability of the user's next interaction with each item through the Softmax function :

[0131] .

[0132] Step 4 also includes: calculating the loss function of conventional sequence recommendation :

[0133] ,

[0134] where is the true label, c is the number of candidate items, represents the indicator value corresponding to item c in the true label, is the predicted probability output, represents the weight coefficient of the c-th item.

[0135] Step 5 includes:

[0136] Dynamically fuse the contrastive learning method and the conventional sequence recommendation method, and the calculation formula of the total loss function is:

[0137] ,

[0138] where is the weight hyperparameter of the contrastive loss.

[0139] Step 5 also includes: updating the model parameters by the gradient descent method , minimizing the total loss:

[0140] ,

[0141] where, Denotes the optimal model parameters obtained after optimization, Denotes finding the parameters that minimize the expression in .

[0142] To evaluate the performance of the method of the present invention, three publicly available multi-domain datasets, namely ml-1m, Beaut, and Steam, are selected for training. The main evaluation metrics include normalized discounted cumulative gain at K (NDCG@K) and recall at K (Recall@K). Experimental data shows that the CLLMRec model of the present invention achieves the best results in all evaluation metrics. For example, compared with the traditional sequential model LRURec, CLLMRec achieves a 5.09% improvement in NDCG@10 and a 1.84% improvement in Recall@20 on the ml-1m dataset; for the attention-based SASRec model, the improvement rate of NDCG@20 reaches 6.8% and Recall@10 improves by 4.2% on the Beauty sparse dataset; compared with the self-supervised learning framework HSTU, NDCG@10 and Recall@10 are improved by 3.7% and 5.1% respectively in the Steam long-tail scenario. Generally speaking, the CLLMRec model of the present invention achieves a significant improvement in recommendation accuracy and scenario adaptability compared with existing recommendation models. Through semantic enhancement and contrast learning framework, its comprehensive performance in dense interaction, sparse data, and long-tail scenarios is significantly improved, verifying the breakthrough advantages of the technical solution.

[0143] The present invention provides a contrast learning sequence recommendation method based on large language model view enhancement. There are many methods and ways to specifically implement this technical solution. The above description is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented using existing technologies.

Claims

1. A contrastive learning sequence recommendation method based on view enhancement of large language models, characterized in that, It includes the following steps: Step 1, establish a dynamic importance scoring mechanism, use a large language model, calculate the importance according to the hidden state of the historical interaction sequence between users and items, and generate a dynamic adjustment factor; Step 2, establish an adaptive sequence cropping method, use the attention mechanism for sequence data augmentation, dynamically adjust the cropping position through semantic understanding, and generate an augmented view containing negative sample pairs according to the dynamic adjustment factor; Step 3, construct a sequence and item dual-granularity contrast learning method under multi-modal dual-view augmentation, build a contrast learning task based on the augmented view generated in Step 2, optimize the contrast loss by shortening the representation distance of positive sample pairs and pushing away the similarity of negative sample pairs, so as to strengthen the contrast learning method's ability to perceive key behavior differences; Step 4, input the original historical interaction sequence between users and items into a linear recurrent unit LRU, extract the dynamic interest representation of the user and predict the probability distribution of the next interaction item, and synchronously calculate the loss function of conventional sequence recommendation; Step 5, dynamically fuse the contrast learning method and the conventional sequence recommendation method, optimize through weighted combination of loss functions, and jointly train the model through backpropagation.

2. The method according to claim 1, characterized in that Step 1 includes: constructing the historical interaction sequence of users and items , where represents the i-th element in the sequence, and is the length of the sequence; Process the historical interaction sequence S between users and items through a pre-trained large language model BERT to establish a dynamic importance scoring mechanism, and the dynamic importance scoring mechanism performs the following operation process: Step 1-1, convert the historical interaction sequence S between users and items into word vectors, including: Input data: Input the historical interaction records of the user; Clean data: Remove invalid interactions; Sequence segmentation: Split long sequences at a fixed length; Build a global item pool: Count all the items that have appeared and generate unique IDs; Vocabulary mapping: Assign an index to each item ID; Initialize the embedding matrix: Create a random matrix with a shape of (B1, B2), where B1 represents the vocabulary size and B2 represents the embedding dimension; Sequence vectorization: Convert the historical interaction sequence S between users and items into a list of indices, and then map it to a vector sequence through the embedding matrix to obtain a normalized vector; Embedding layer input: Input the normalized vector sequence into the embedding layer; Step 1-2, extract the hidden state through multiple layers of Transformer; Step 1-3, calculate the position importance score along the feature dimension for subsequent information cropping; Step 1-4, generate a dynamic adjustment factor through the aggregation of multiple layers of hidden states, and the formula is: , Among them, represents the dynamic adjustment factor of the i-th sample, represents the hidden layer state matrix output by the k-th layer of the pre-trained large language model BERT, and d is the number of layers of the BERT model.

3. The method according to claim 2, characterized in that, Step 2 includes: The adjustment factor acts directly on the positional encoding to obtain the clipped dynamic starting position and the elastic clipping length : , , , Among them, represents the cropping ratio parameter, represents the length of the original sequence; represents uniform distribution random sampling, represents the position encoding matrix of the augmented sequence; represents an integer sequence from 1 to ; Then, through the calculated dynamic starting position of the clipping and the elastic clipping length the enhanced sequence is calculated , and the formula is: , Among them represents the original user and item historical interaction sequence; represents the intercepted part starting from with a length of sequence part.

4. The method according to claim 3, characterized in that Step 3 includes: Enhance the original sequence S to an enhanced sequence Use them as positive and negative samples respectively for contrastive learning; The sequence-level alignment method captures the global temporal pattern of user behavior by extracting the end position embeddings of the augmented sequences. The loss function of sequence-level alignment is given by the formula: , where N represents the total number of samples included in a training batch, is the sequence-level embedding vector, represents the positive example-level embedding of the i-th sample, represents the j-th negative example embedding vector, is the adjustable temperature coefficient, is the cosine similarity, exp represents the natural exponential function, and K represents each anchor total number of negative examples to be compared; Use the SASRec model as the basic model and use the position encoding matrix and the enhanced sequence as parameters and input them into the SASRec model to obtain the vector embedding of the enhanced sequence; Using the Transformer encoding layer in the SASRec model, the relationships between items are learned implicitly to achieve item-level alignment. A local item similarity matrix is constructed through attention weights, and the loss function for item-level alignment is as follows: , Among them represents the interaction sequence length between the user and the project, is the hidden state output by the q-th project Transformer, is the local contrast representation of the q-th project; is the hidden state after encoding by the same augmented variant generated by sequence shuffling through the Transformer; represents and similarity; Finally, the contrastive learning loss is obtained by jointly optimizing with the dual-granularity contrastive loss : , Among them, , represent weight coefficients.

5. The method according to claim 4, characterized in that, Step 4 includes: initializing the hidden state with the original sequence S as the input, and setting the initial state , and updating the hidden state: , Among them, is the input vector at time step t, has a dimension of , is the hidden state vector at time step t, has a dimension of , is the weight matrix of the hidden state, is the input weight matrix, is the bias; represents the real number space; Obtain the interest at each time step , where T represents the total number of time steps of the input interaction sequence, and take the final state as the user's final interest for predicting the next interaction; Map the hidden state to the item space: , Among them, is the mapping result in the item space corresponding to time step t, is the weight matrix, is the bias term; Generate the predicted probability of the user's next interaction with each item through the Softmax function : 。 6. The method according to claim 5, characterized in that Step 4 further includes: calculating the loss function of the conventional sequence recommendation : , where is the true label, is the number of candidate items, represents the indication value corresponding to item c in the true label, is the predicted probability output, represents the weight coefficient of the c-th item.

7. The method according to claim 6, wherein Step 5 includes: The dynamic fusion contrastive learning method and the conventional sequence recommendation method, the total loss function The calculation formula is as follows: , Among them is the weight hyperparameter of the contrastive loss.

8. The method according to claim 7, wherein Step 5 further includes: updating the model parameters by the gradient descent method , to minimize the total loss: , Among them, represents the optimal model parameters obtained after optimization, represents finding the parameter that makes the expression in it take the minimum value .

9. An electronic device, characterized in that, It includes a processor and a memory, and the memory stores program codes. When the program codes are executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 8.

10. A storage medium, characterized in that, Store a computer program or instruction. When the computer program or instruction runs on a computer, it executes the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Commodity sequence recommendation method based on time sequence perception self-attention and comparative learning

    CN117196763A

  • Memory retrieval method based on large language model and related device

    CN119903125A