Generative recommendation pre-training method based on multi-word segmentation device enhancement

By using a pre-training method with multi-partitioner enhancement in the generative recommendation model, data augmentation and dynamically adjusting the training data distribution is solved by using the RQ-VAE model checkpoints of adjacent epochs, which is difficult for the model to adapt to new items and long-tail items in the prior art and is easy to overfit, achieving higher effectiveness and generalization of the model.

CN120030350APending Publication Date: 2025-05-23RENMIN UNIVERSITY OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510111475.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing generative recommendation model is trained on a fixed item set and a single item word segmenter, making it difficult to adapt to new items and long-tail items with low interaction frequency. Moreover, due to data sparseness, it is easy to overfit and lacks generalization.

Method used

The generative recommended pre-training method is adopted with multi-word participle enhancement. The RQ-VAE model checkpoint of adjacent epoch is used as semantic-related multiple item word participle, and the training data distribution is dynamically adjusted, and the gradient influence score is used for course learning.

Benefits of technology

Through data augmentation and dynamic adjustment of the training data distribution, the effectiveness and generalization of the model are improved, overfitting is avoided, and large-scale parameter expansion of the generative recommendation model is possible.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030350A_ABST
    Figure CN120030350A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of generative recommendation, in particular to a generative recommendation pre-training method based on multi-word segmentation device enhancement, which comprises the following steps: S1, inputting training data; s2, selecting an article word segmentation device according to the sampling probability; s3, performing word segmentation on the interaction sequence and the target article; s4, calculating the negative logarithm likelihood loss of the generated target and performing model optimization; s5, periodically filtering the word segmentation device and updating the sampling probability; and S6, circulating the steps from S1 to S5 until the model converges. According to the method, the limitation is solved in a targeted mode through data enhancement of the multi-item word segmentation device and a data course technology based on gradient influence scores, the potential of the generative recommendation model is further stimulated through model pre-training, and generalization and recommendation performance of the generative recommendation model are fully improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of generative recommendation technology, and in particular to a generative recommendation pre-training method based on multi-word segmenter enhancement. Background Art

[0002] In order to help users find items of interest, modern recommendation systems often model users' sequential behavior patterns to infer their personalized preferences and predict the next interactive item. Generative recommendation methods convert this recommendation task into a sequence-to-sequence generation task. Specifically, this type of method mainly includes two key modules: item tokenizer and generative recommendation model. The item tokenizer is used to map each item into a set of tokens with specific semantics. The shared tokens between different items represent the semantic relevance between them. For the item sequence that the user has interacted with, the tokenizer tokenizes the items separately and finally splices them together as the input of the generative recommendation model. In existing methods, the generative recommendation model is often a Transformer model with an encoder-decoder structure, which generates tokens corresponding to the target item in an autoregressive manner. In actual deployment, the model uses a beam search algorithm to generate N predicted items as recommendation results.

[0003] There are two main limitations of existing technologies when training models. First, existing generative recommendation models are trained on a fixed set of items and a single item segmenter, which makes the model tend to generate seen or more popular items, but difficult to adapt to new items or long-tail items with low interaction frequency. On the other hand, there is a data sparsity problem in the recommendation system, which results in a limited scale of word sequences involved in the training process and a lack of diverse supervisory signals, which makes the generative recommendation model prone to overfitting and lack of generalization. More importantly, as the model size increases, this phenomenon becomes more serious, which limits the full use of the model's capabilities.

[0004] The information disclosed in this background technology section is only intended to deepen the understanding of the overall background technology of the present invention, and should not be regarded as acknowledging or suggesting in any form that the information constitutes the prior art already known to those skilled in the art. Summary of the invention

[0005] The purpose of the present invention is to provide a generative recommendation pre-training method based on multi-word segmenter enhancement to solve the technical problems existing in the prior art.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] The present invention provides a generative recommendation pre-training method based on multi-word segmenter enhancement, which comprises:

[0008] S1, input training data;

[0009] S2, select an item segmenter based on the sampling probability;

[0010] S3, segment the interaction sequence and target items;

[0011] S4, calculate the negative log-likelihood loss of the generated target and perform model optimization;

[0012] S5, periodically filter the word segmenter and update the sampling probability;

[0013] S6. Repeat the above steps S1 to S5 until the model converges.

[0014] Preferably, step S1 comprises:

[0015] Input a batch of training data, each data instance contains a user's item interaction sequence S = [v 1 ,v 2 ,…,v t ] and the next target item v t+1 .

[0016] Preferably, step S2 comprises:

[0017] Sample one tokenizer from multiple candidate items for subsequent data tokenization based on probability:

[0018] T~{T 1 ,T 2 ,...,T n}

[0019]

[0020] Among them, T i represents the i-th item segmenter, T represents the final item segmenter sampled from multiple candidate item segmenters, and p i It represents the sampling probability of each item segmenter. Initially, the probabilities of each segmenter are equal, and then they are dynamically adjusted according to the gradient influence score calculated in S5.

[0021] Preferably, step S3 comprises:

[0022] Based on the sampled item segmenter, the input interaction sequence S and the target item v t+1 To perform word segmentation:

[0023]

[0024] in, represents the jth token of the ith item. Each item has a total of H tokens. X is the user interaction sequence after word segmentation, and Y is the generation target of the generative recommendation model.

[0025] Preferably, step S4 comprises:

[0026] The sequence recommendation task is converted into a sequence-to-sequence generation task. Taking the user interaction sequence X as a condition, the negative log-likelihood of the generated target Y is calculated as the model loss for model optimization:

[0027]

[0028] in, It is the autoregressive generation loss of the model, and then the generative recommendation model is optimized based on the back propagation and gradient descent algorithms.

[0029] Preferably, step S5 comprises:

[0030] Based on the initial state in step S2, there are n word segmenters, and the sampling probability of each word segmenter is equal. The word segmenters with lower quality are filtered out periodically, and the sampling probability of the remaining word segmenters is dynamically updated. The quality of the word segmenter is evaluated based on the validation set effect of each word segmenter. Different word segmenters are used for evaluation on the same validation set, and the common indicator NDCG@10 in the recommendation system is calculated. Finally, half of the word segmenters with higher quality are retained, and the minimum number of word segmenters is limited to m.

[0031] The sampling probability of the word segmenter is obtained according to the gradient influence score of each word segmenter, including:

[0032] The contribution of the word segmenter to model optimization is defined as: the impact of model optimization based on the current word segmenter on the validation set loss. According to the first-order Taylor expansion formula, the validation set loss is written as:

[0033]

[0034] in, is the validation data, θ t represents the model parameters of the tth iteration. According to the above Taylor expansion, the change of the validation set loss is written as When using the Adam optimizer (θ t+1 -θ t ) is: k = 1

[0035] in, is the training data, β 1 and β 2 They are the hyperparameters of the Adam optimizer, ∈ is a small constant, and η t is the learning rate of the current model. Based on the above two formulas, the gradient influence scores of each tokenizer required are:

[0036] in, Represents the word segmenter T i The training data for word segmentation is used. The gradient influence scores are calculated based on K model checkpoints, but the previous K-1 calculation results can be reused each time. Finally, the sampling probability of each item segmenter is:

[0037]

[0038] Where τ is the temperature coefficient used to control the smoothness of the distribution;

[0039] By controlling the sampling probability of different word segmenters and adjusting the proportion of data from different sources in training, dynamic adjustment of data courses in the pre-training process is achieved.

[0040] By adopting the above technical solution, the present invention has the following beneficial effects:

[0041] The present invention uses the RQ-VAE model checkpoints of adjacent epochs as semantically related multiple item segmenters for data augmentation. An item interaction sequence is augmented into multiple item word sequences to obtain larger and more diverse training data. The present invention also proposes a curriculum learning technology based on gradient influence scores to dynamically adjust the distribution of training data. Data augmentation and dynamic adjustment of data distribution improve the effectiveness of the model while ensuring the generalization and robustness of the model, and make large-scale parameter expansion of the generative recommendation model possible. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0043] Figure 1 A flowchart of the generative recommendation pre-training method based on multi-word segmenter enhancement provided by the present invention.

[0044] Figure 2 This is a diagram of the overall architecture of the multi-item word segmenter-enhanced generative recommendation model technology provided by the present invention. DETAILED DESCRIPTION

[0045] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0046] The following is combined with Figure 1 and Figure 2 The specific embodiments of the present invention are described in detail. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0047] The present invention provides a generative recommendation pre-training method based on multi-word segmenter enhancement, which comprises:

[0048] S1, input training data;

[0049] Input a batch of training data, each data instance contains a user's item interaction sequence S = [v 1 ,v 2 ,…,v t ] and the next target item v t+1 .

[0050] S2, select an item segmenter based on the sampling probability;

[0051] Different from the previous method that only uses one item segmenter, the present invention uses multiple RQ-VAE (Residual-Quantized Variational AutoEncoder) model checkpoints of adjacent epochs in a training process as candidate item segmenters, which have associated semantic information. Specifically, the present invention probabilistically samples one from multiple candidate item segmenters for subsequent data segmentation:

[0052] T~{T 1 ,T 2 ,...,T n}

[0053]

[0054] Where T i represents the i-th item segmenter, and T represents the final item segmenter sampled from multiple candidate item segmenters. i It represents the sampling probability of each item segmenter. Initially, the probabilities of each segmenter are equal, and then they are dynamically adjusted according to the gradient influence score calculated in step 5.

[0055] S3, segment the interaction sequence and target items;

[0056] Based on the sampled item segmenter, the present invention performs a segmentation on the input interaction sequence S and the target item v t+1 To perform word segmentation:

[0057]

[0058] in represents the jth token of the i-th item, and each item has a total of H tokens. X is the user interaction sequence after token segmentation, and Y is the generation target of the generative recommendation model. The specific RQ-VAE training and token segmentation process can be referred to as follows:

[0059] RQ-VAE (Residual-Quantized Variational AutoEncoder) is the most commonly used item segmenter in generative recommendation, thanks to

[0060]

[0061] It has the advantage of modeling prior semantic knowledge. In addition, RQ-VAE can obtain equal-length item word sequences to avoid length bias. RQ-VAE takes the item semantic embedding x (e.g., text embedding encoded by a pre-trained language model) as input and encodes it into a latent semantic representation z. Then, z is quantized from coarse to fine into a set of word units by nearest neighbor matching with the H-level codebook vectors. Specifically, each level of the codebook consists of Indicates that is the learnable cluster center, and K is the codebook size. Based on this, the residual quantization process can be written as:

[0062] where c j represents the jth word of the item, r j is the residual vector of level j, and r 1 = z. Subsequently, the present invention can obtain a quantitative representation of the object Afterwards, Input to the decoder to reconstruct the semantic embedding of the item:

[0063] in is the reconstructed item embedding, sg[] represents the gradient stop operation, and β is the balanced encoder and

[0064]

[0065] The coefficient for optimization between codebooks is usually set to 0.25.

[0066] S4, calculate the negative log-likelihood loss of the generated target and perform model optimization;

[0067] The generative recommendation method converts the sequence recommendation task into a sequence-to-sequence generation task. It takes the user interaction sequence X as a condition and calculates the negative log-likelihood of the generated target Y as the model loss for model optimization:

[0068]

[0069] in, is the autoregressive generation loss of the model. Then, the present invention optimizes the generative recommendation model based on back propagation and gradient descent algorithms.

[0070] S5, periodically filter the word segmenter and update the sampling probability;

[0071] As described in step 2, the present invention has n word segmenters in the initial state, and the sampling probability of each word segmenter is equal. In order to further optimize the technology of the present invention, the present invention is inspired by the curriculum learning technology, periodically filters out the word segmenters with lower quality, and dynamically updates the sampling probability of the remaining word segmenters.

[0072] The quality of the word segmenters is evaluated based on the validation set effect of each word segmenter. The present invention uses different word segmenters for evaluation on the same validation set, and calculates the commonly used indicator NDCG@10 (Normalized Discounted Cumulative Gain) in the recommendation system. Finally, the present invention retains half of the word segmenters with higher quality and limits the minimum number of word segmenters to m (hyperparameter).

[0073] The sampling probability of the word segmenter is obtained according to the gradient influence score of each word segmenter. Specifically, inspired by the LESS algorithm in the field of natural language, the contribution of the word segmenter to model optimization is defined as: the impact of model optimization based on the current word segmenter on the validation set loss. According to the first-order Taylor expansion formula, the validation set loss can be written as:

[0074]

[0075] in is the validation data, θ t represents the model parameters of the tth iteration. According to the above Taylor expansion, the change in the validation set loss can be written as In use

[0076]

[0077] In the case of Adam optimizer (θ t+1 -θ t ) is: k = 1

[0078] in, is the training data, β 1 and β 2 They are the hyperparameters of the Adam optimizer, ∈ is a small constant, and η t is the learning rate of the current model. Based on the above two formulas, the gradient influence score of each word segmenter required by the present invention is:

[0079] in, Represents the word segmenter T i Training data for word segmentation. It is worth noting that the present invention calculates the gradient influence score based on K model checkpoints, but each calculation can reuse the previous K-1 calculation results. Finally, the sampling probability of each item segmenter is:

[0080]

[0081] where τ is the temperature coefficient used to control the smoothness of the distribution.

[0082] Based on the above sampling probability updating method, the present invention adjusts the proportion of data from different sources in training by controlling the sampling probabilities of different word segmenters, thereby realizing dynamic adjustment of data courses in the pre-training process.

[0083] S6. Repeat the above steps S1 to S5 until the model converges.

[0084] Existing methods are trained on the basis of a single and fixed item tokenizer. The model is limited by the long-tail distribution problem of item interaction frequency in the recommendation scenario, and tends to be more popular items. On the other hand, due to data sparsity, the model repeatedly iteratively optimizes the limited item token metadata during the training process, which makes the generative recommendation model prone to overfitting and lack of generalization. Compared with these methods, the scheme proposed in the present invention uses the RQ-VAE model checkpoints of adjacent epochs as multiple semantically related item tokenizers for data enhancement. An item interaction sequence is augmented into multiple item token sequences to obtain larger and more diverse training data. The present invention also proposes a curriculum learning technology based on gradient influence scores to dynamically adjust the distribution of training data. Data enhancement and dynamic adjustment of data distribution ensure the generalization and robustness of the model while improving the effectiveness of the model, and make large-scale parameter expansion of the generative recommendation model possible.

[0085] like Figure 2 As shown, the overall architecture of this embodiment mainly includes two key contributions.

[0086] First, the present invention proposes to use the RQ-VAE model checkpoints of adjacent epochs as multiple semantically related item segmenters for data augmentation. This technical innovation can augment an item interaction sequence into multiple item word-meta sequences for pre-training of generative recommendation models, thereby improving the effectiveness of the model. Secondly, in order to improve the data distribution of the training process, the present invention proposes to filter low-quality item segmenters based on the validation set performance and dynamically update the sampling probability of multiple item segmenters based on the gradient impact score. This data course based on dynamic distribution further enhances the generalization and robustness of the model.

[0087] Specifically, under this framework, a data instance containing the user's item interaction sequence and the next target item is first input, and then an item segmenter is sampled probabilistically and item segmentation is completed as the input of the generative recommendation model. Model optimization is based on the negative log-likelihood of the generated target. In addition, during the model training process, the present invention periodically filters low-quality item segmenters and dynamically updates their sampling probabilities to achieve model course learning. This process is iterative until the model converges.

[0088] It is worth mentioning that compared with the existing methods, the technical solution of the present invention can provide more and more diverse data for model training, which makes large-scale parameter expansion of the generative model possible, rather than falling into overfitting and making it difficult to stimulate the model's potential. After completing pre-training on a multi-item tokenizer, a basic model with both generalization and effectiveness was obtained. Since items and their tokens in actual recommendation scenarios need to satisfy a one-to-one mapping to achieve item recognition, the present invention then fine-tunes the item tokenizer with the best verification effect for final deployment. In actual deployment, the present invention uses a beam search algorithm to generate N predicted items as recommendation results.

[0089] In summary, as an emerging personalized recommendation technology, generative recommendation technology converts traditional item recommendations based on nearest neighbor search into an autoregressive generation method, effectively unifying the multi-stage recall process into a single generation, while improving the recommendation effect. However, the existing technical solutions are based on a single item segmenter and limited word sequence data for training, which makes the model also limited by the data sparsity and the long-tail distribution of the number of item interactions in the recommendation scenario. It is difficult to adapt to items with low interaction frequency and is prone to overfitting, lacking generalization. These problems are more serious when the model parameter scale is expanded, so it is impossible to improve performance by expanding the parameter scale like a large language model. The present invention solves the above limitations in a targeted manner through data enhancement of multi-item segmenters and data course technology based on gradient influence scores, further stimulates the potential of the generative recommendation model through model pre-training, and fully improves the generalization and recommendation performance of the generative recommendation model.

[0090] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A generative recommendation pre-training method based on multi-segmenter enhancement, characterized in that: include: S1, input training data; S2, select an item segmenter based on the sampling probability; S3, segment the interaction sequence and target items; S4, calculate the negative log-likelihood loss of the generated target and perform model optimization; S5, periodically filter the word segmenter and update the sampling probability; S6. Repeat the above steps S1 to S5 until the model converges.

2. The generative recommendation pre-training method based on multi-word segmenter enhancement according to claim 1 is characterized in that: Step S1 includes: Input a batch of training data, each data instance contains a user's item interaction sequence S = [v1, v2, ..., v t ] and the next target item v t+1 .

3. The generative recommendation pre-training method based on multi-word segmenter enhancement according to claim 1 is characterized in that: Step S2 includes: Sample one tokenizer from multiple candidate items for subsequent data tokenization based on probability: T~{T1,T2,...,T n } where P(T i )=p i , Among them, T i represents the i-th item segmenter, T represents the final item segmenter sampled from multiple candidate item segmenters, and p i It represents the sampling probability of each item segmenter. Initially, the probabilities of each segmenter are equal, and then they are dynamically adjusted according to the gradient influence score calculated in S5.

4. The generative recommendation pre-training method based on multi-word segmenter enhancement according to claim 1 is characterized in that: Step S3 includes: Based on the sampled item segmenter, the input interaction sequence S and the target item v t+1 To perform word segmentation: in, represents the jth token of the ith item. Each item has a total of H tokens. X is the user interaction sequence after word segmentation, and Y is the generation target of the generative recommendation model.

5. The generative recommendation pre-training method based on multi-word segmenter enhancement according to claim 4 is characterized in that: Step S4 includes: The sequence recommendation task is converted into a sequence-to-sequence generation task. Taking the user interaction sequence X as a condition, the negative log-likelihood of the generated target Y is calculated as the model loss for model optimization: in, It is the autoregressive generation loss of the model, and then the generative recommendation model is optimized based on the back propagation and gradient descent algorithms.

6. The generative recommendation pre-training method based on multi-word segmenter enhancement according to claim 4 is characterized in that: Step S5 includes: Based on the initial state in step S2, there are n word segmenters, and the sampling probability of each word segmenter is equal. The word segmenters with lower quality are filtered out periodically, and the sampling probability of the remaining word segmenters is dynamically updated. The quality of the word segmenter is evaluated based on the validation set effect of each word segmenter. Different word segmenters are used for evaluation on the same validation set, and the common indicator NDCG@10 in the recommendation system is calculated. Finally, half of the word segmenters with higher quality are retained, and the minimum number of word segmenters is limited to m. The sampling probability of the word segmenter is obtained according to the gradient influence score of each word segmenter, including: The contribution of the word segmenter to model optimization is defined as: the impact of model optimization based on the current word segmenter on the validation set loss. According to the first-order Taylor expansion formula, the validation set loss is written as: in, is the validation data, θ t represents the model parameters of the tth iteration. According to the above Taylor expansion, the change of the validation set loss is written as When using the Adam optimizer (θ t+1 -θ t ) is: k = 1 in, is the training data, β1 and β2 are the hyperparameters of the Adam optimizer, ∈ is a small constant, and η t is the learning rate of the current model. Based on the above two formulas, the gradient influence scores of each tokenizer required are: in, Represents the word segmenter T i The training data for word segmentation is used. The gradient influence scores are calculated based on K model checkpoints, but the previous K-1 calculation results can be reused each time. Finally, the sampling probability of each item segmenter is: Where τ is the temperature coefficient used to control the smoothness of the distribution; By controlling the sampling probability of different word segmenters and adjusting the proportion of data from different sources in training, dynamic adjustment of data courses in the pre-training process is achieved.