Model training method and recommendation method, device, equipment, storage medium, computer program product
Patent Information
- Application Number
- CN202610912311.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-23
- Publication Date
- 2026-09-29
AI Technical Summary
然而,评估器存在价值评估不准确的问题,从而降低了推荐的准确性
[0021]本申请实施例中,确定第一样本数据,第一样本数据包括多个第一物品信息序列,第一物品信息序列中的每个物品信息用于描述一个物品的物品侧特征以及用户侧特征;然后,基于多个第一物品信息序列,利用第一损失函数与第二损失函数对第一评估模型进行多次训练迭代,直至达到设定的第一收敛条件,得到第二评估模型,其中,第一评估模型包括第一评估网络与第一辅助网络;第一评估网络用于基于输入的物品信息序列生成中间特征,并基于中间特征对物品信息序列进行点级价值评估,第一辅助网络用于基于第一评估网络生成的中间特征对物品信息序列进行序列级价值评估,第一损失函数用于衡量关于价值的点级损失;第二损失函数用于衡量关于价值的序列级损失。上述方案中,将第一损失函数与第二损失函数联合用于第一评估模型的训练,其中,第一损失函数用于衡量关于价值的点级损失,相当于为训练过程引入了点级训练目标,第二损失函数用于衡量关于价值的序列级损失,相当于为训练过程引入了序列级训练目标,因此,本申请方案中的训练过程相当于采用了基于点级训练目标与序列级训练目标构成的混合训练目标(HTO,HybridTraining Objective),在此基础上,训练完成的第一评估模型能够有效捕捉物品序列的整体价值,且仍然可以保证点级价值评估的准确性,稳定了训练过程,避免了绝对分数的退化,从而相比于相关技术,提升了价值评估的准确性,进而提升了推荐的准确性。
Smart Images

Figure CN122840148A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a model training method and recommendation method, apparatus, device, storage medium, and computer program product. Background Technology
[0002] In related technologies, during the reordering phase of the recommendation process, a trained generator produces item sequences, and item recommendations are made to users based on these sequences. During the generator's training phase, an evaluator assesses the value of the generated item sequences, and the evaluated value guides the generator's training, enabling it to produce high-quality item sequences. However, the evaluator suffers from inaccurate value assessments, thus reducing the accuracy of the recommendations. Summary of the Invention
[0003] To address the related technical issues, embodiments of this application provide a model training method and recommendation method, apparatus, device, storage medium, and computer program product.
[0004] The technical solution of this application embodiment is implemented as follows: This application provides a model training method, the method comprising: First sample data is determined; the first sample data includes multiple first item information sequences; each item information in the first item information sequence is used to describe the item-side features and user-side features of an item; Based on the multiple first item information sequences, the first evaluation model is trained and iterated multiple times using the first loss function and the second loss function until the set first convergence condition is met, and the second evaluation model is obtained. The first evaluation model includes a first evaluation network and a first auxiliary network. The first evaluation network is used to generate intermediate features based on the input item information sequence and to perform point-level value evaluation on the item information sequence based on the intermediate features. The first auxiliary network is used to perform sequence-level value evaluation on the item information sequence based on the intermediate features generated by the first evaluation network. The first loss function is used to measure point-level loss with respect to value. The second loss function is used to measure sequence-level loss with respect to value.
[0005] In the above scheme, based on each of the plurality of first item information sequences, the first evaluation model is trained and iterated multiple times using a first loss function and a second loss function, including: The first evaluation network is invoked to process the first item information sequence to obtain a first evaluation result and a first intermediate feature; and the first auxiliary network is invoked to process the first intermediate feature to obtain a second evaluation result. The first evaluation result is processed based on a first loss function to obtain a first loss value, and the first evaluation network is updated based on the first loss value; and the second evaluation result is processed based on a second loss function to obtain a second loss value, and the first auxiliary network and the first evaluation network are updated sequentially based on the second loss value.
[0006] In the above scheme, the step of processing the second evaluation result based on the second loss function to obtain a second loss value, and then sequentially updating the first auxiliary network and the first evaluation network based on the second loss value, includes: In the case of the first training iteration in the multiple training iterations of the first evaluation model, the second evaluation result is processed based on the second loss function to obtain the second loss value, and the first auxiliary network and the first evaluation network are updated sequentially based on the second loss value. The step of processing the first evaluation result based on the first loss function to obtain a first loss value, and updating the first evaluation network based on the first loss value, includes: In the case of the second training iteration in the multiple training iterations of the first evaluation model, the first evaluation result is processed based on the first loss function to obtain the first loss value, and the first evaluation network is updated based on the first loss value. Wherein, the first training iteration represents one training iteration performed every n training iterations; the second training iteration represents other training iterations in the multiple training iterations besides the first training iteration; and n is a positive integer.
[0007] This application embodiment also provides a model training method, the method comprising: Determine the second sample data; the second sample data includes multiple sets of first item information; each item information in the first item information set is used to describe the item-side features and user-side features of an item; Based on the multiple sets of first item information, the first generation model is trained and iterated multiple times using the second evaluation model until the set second convergence condition is met, thus obtaining the second generation model; the second evaluation model is obtained by training using the model training method described above.
[0008] In the above scheme, based on each of the plurality of first item information sets, the first generation model is trained and iterated multiple times using the second evaluation model, including: The first generation model is invoked to process the first item information set to obtain a second item information sequence; each item information in the second item information sequence is used to identify an item. The second evaluation model is invoked to process the second item information sequence to obtain a third evaluation result; the third evaluation result represents the evaluation result output by the evaluation network in the second evaluation model. Based on the third evaluation result, determine the first value score corresponding to the second item information sequence; Update the first generative model based on the first value score.
[0009] In the above scheme, determining the first value score corresponding to the second item information sequence based on the third evaluation result includes: Based on the third evaluation result, the item information in the second item information sequence is reordered to obtain the third item information sequence; For each item information in the third item information sequence, the difference between the first position index value corresponding to the item information in the third item information sequence and the second index value corresponding to the item information in the second item information sequence is calculated to obtain the first difference value corresponding to the item information; A weighted summation is performed on multiple first differences corresponding to the third item information sequence to obtain a first value score; the first weight corresponding to the first difference in the first value score is related to the first position index value corresponding to the first difference.
[0010] In the above scheme, the third evaluation result includes a first score sequence composed of multiple second value scores; The step of reordering the item information in the second item information sequence based on the third evaluation result to obtain the third item information sequence includes: Based on the second value score corresponding to each item in the second item information sequence in the first score sequence, the item information in the second item information sequence is sorted to obtain the third item information sequence; each second value score in the first score sequence is used to describe the value corresponding to an item in the second item information sequence.
[0011] The above scheme, based on each of the plurality of first item information sets, uses a second evaluation model to perform multiple training iterations on the first generation model, and further includes: The second evaluation model is invoked to process the fourth item information sequence corresponding to the first item information set to obtain a fourth evaluation result; the fourth item information sequence represents a preset item information sequence for the first item information set, and each item information in the fourth item information sequence is used to identify an item; the fourth evaluation result represents the evaluation result output by the evaluation network in the second evaluation model. Based on the fourth evaluation result, determine the third value score corresponding to the fourth item information sequence; Correspondingly, updating the first generation model based on the first value score includes: The first generative model is updated based on the second difference between the third value score and the first value score.
[0012] This application embodiment also provides a recommended method, the method comprising: The second generation model is invoked to process the second item information set to obtain the fifth item information sequence; the second item information set represents the item information set determined through recall and sorting processes. Based on the fifth item information sequence, item recommendations are made to the user; The second generative model is obtained by training using the model training method described above.
[0013] This application embodiment also provides a model training apparatus, the apparatus comprising: The first determining unit is used to determine the first sample data; the first sample data includes multiple first item information sequences; each item information in the first item information sequence is used to describe the item-side features and user-side features of an item. The first training unit is used to train the first evaluation model multiple times based on the multiple first item information sequences, using a first loss function and a second loss function, until a set first convergence condition is reached, and thus obtain the second evaluation model. The first evaluation model includes a first evaluation network and a first auxiliary network. The first evaluation network is used to generate intermediate features based on the input item information sequence and to perform point-level value evaluation on the item information sequence based on the intermediate features. The first auxiliary network is used to perform sequence-level value evaluation on the item information sequence based on the intermediate features generated by the first evaluation network. The first loss function is used to measure point-level loss with respect to value. The second loss function is used to measure sequence-level loss with respect to value.
[0014] This application embodiment also provides a model training apparatus, the apparatus comprising: The second determining unit is used to determine the second sample data; the second sample data includes multiple sets of first item information; each item information in the first item information set is used to describe the item-side features and user-side features of an item; The second training unit is used to train the first generation model multiple times based on the multiple sets of first item information and using the second evaluation model until the set second convergence condition is reached, so as to obtain the second generation model; the second evaluation model is obtained by training using the above-mentioned model training method.
[0015] This application embodiment also provides a model training apparatus, the apparatus comprising: The calling unit is used to call the second generation model to process the second item information set to obtain the fifth item information sequence; the second item information set represents the item information set determined through recall processing and sorting processing. The recommendation unit is used to recommend items to the user based on the fifth item information sequence; The second generative model is obtained by training using the model training method described above.
[0016] This application also provides a model training device, the device comprising: a first processor, a first memory, and a first communication bus; The first communication bus is used to establish a communication connection between the first processor and the first memory; The first processor is configured to execute the model training program in the first memory to implement the steps of the model training method as described in any of the preceding claims.
[0017] This application also provides a model training device, the device comprising: a second processor, a second memory, and a second communication bus; The second communication bus is used to establish a communication connection between the second processor and the second memory; The second processor is used to execute the model training program in the second memory to implement the steps of the model training method as described in any of the preceding claims.
[0018] This application also provides a model training device, the device comprising: a third processor, a third memory, and a third communication bus; The third communication bus is used to realize the communication connection between the third processor and the third memory; The third processor is used to execute the recommendation program in the third memory to implement the steps of the recommendation method as described above.
[0019] This application also provides a computer-readable storage medium storing one or more programs that can be executed by one or more processors, following the steps of the method provided in this application.
[0020] This application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the method provided in this application.
[0021] In this embodiment, first sample data is determined, which includes multiple first item information sequences. Each item information sequence describes the item-side features and user-side features of an item. Then, based on the multiple first item information sequences, a first evaluation model is trained and iterated multiple times using a first loss function and a second loss function until a set first convergence condition is reached to obtain a second evaluation model. The first evaluation model includes a first evaluation network and a first auxiliary network. The first evaluation network is used to generate intermediate features based on the input item information sequences and to perform point-level value evaluation on the item information sequences based on the intermediate features. The first auxiliary network is used to perform sequence-level value evaluation on the item information sequences based on the intermediate features generated by the first evaluation network. The first loss function is used to measure point-level loss related to value, and the second loss function is used to measure sequence-level loss related to value. In the above scheme, the first loss function and the second loss function are used together to train the first evaluation model. The first loss function measures the point-level loss related to value, which is equivalent to introducing a point-level training objective into the training process. The second loss function measures the sequence-level loss related to value, which is equivalent to introducing a sequence-level training objective into the training process. Therefore, the training process in this application scheme is equivalent to adopting a hybrid training objective (HTO) based on a point-level training objective and a sequence-level training objective. On this basis, the trained first evaluation model can effectively capture the overall value of the item sequence, while still ensuring the accuracy of point-level value evaluation. This stabilizes the training process, avoids the degradation of absolute scores, and thus improves the accuracy of value evaluation compared to related technologies, thereby improving the accuracy of recommendations. Attached Figure Description
[0022] Figure 1 A schematic diagram illustrating the implementation process of a model training method provided in this application embodiment; Figure 2 A schematic diagram illustrating the implementation process of another model training method provided in this application embodiment; Figure 3 A schematic diagram illustrating the implementation flow of a recommended method provided for an application embodiment of this application; Figure 4 A training diagram provided for an application embodiment of this application; Figure 5 A schematic diagram of sequence value assessment provided for an application embodiment of this application; Figure 6A schematic diagram of the structure of a model training device provided in an application embodiment of this application; Figure 7 A schematic diagram of another model training device provided in the application embodiments of this application; Figure 8 A schematic diagram of the structure of a recommended device provided for an application embodiment of this application; Figure 9 A schematic diagram of the hardware composition structure of a model training device provided for an application embodiment of this application; Figure 10 A schematic diagram of the hardware composition structure of another model training device provided for an application embodiment of this application; Figure 11 This is a schematic diagram of the hardware structure of a recommended device provided for an application embodiment of this application. Detailed Implementation
[0023] A recommendation system is a system used to recommend content to users. Specifically, a recommendation system can select several items from a set of candidate items based on a user's interests and preferences, and present the selected items to the user in a specific order. Items are the content units used to recommend content to the user; for example, items can include goods, services, or news. In practical applications, recommendation systems are widely used in various internet applications across different fields, such as e-commerce, social media, and news recommendations.
[0024] In related technologies, the recommendation process of recommendation systems typically follows a multi-stage pipeline structure, mainly including three stages: recall, ranking, and re-ranking. The goal of the recall stage is to filter out items that may be relevant to the user from a massive number of items. The goal of the ranking stage is to finely score and rank the items selected in the recall stage, thereby further filtering out... The goal of the rearrangement phase is to select items from the sorting phase. Further filtering from the items One item, and Items are organized in a specific order to form a sequence with the highest overall value. Based on this sequence, item recommendations are made to the user; for example, the user is presented with this item sequence. as well as All are positive integers. It can be less than or equal to The overall value of an item sequence can be used to describe the quality of a recommendation after recommending the item sequence to a user, such as user satisfaction or the user's purchase rate of the item. The higher the overall value of an item sequence, the better the recommendation effect corresponding to that item sequence, and the higher the quality of the item sequence can be considered.
[0025] In related technologies, a generator-evaluator framework is used to model the problem of generating high-quality item sequences during the reordering stage. Specifically, in the reordering stage of the recommendation process, a trained generator generates item sequences, and item recommendations are made to users based on these sequences. During the generator's training phase, the generator generates one or more candidate item sequences, and an evaluator assesses the overall value of these candidate item sequences. Furthermore, a reinforcement learning method is used to establish an interaction mechanism between the generator and the evaluator. The value evaluated by the evaluator guides the generator's training, enabling the trained generator to generate high-quality item sequences. The value evaluated by the evaluator can be quantified into a specific numerical value, also known as a value score or score.
[0026] In practical applications, the accuracy of the values evaluated by the evaluator can affect the training performance of the generator, and consequently, the accuracy of the recommendations. The more accurate the values evaluated by the evaluator, the more effective it can be considered. Related technologies primarily address the problem of training an effective evaluator using the following two methods.
[0027] Method 1: Classification modeling.
[0028] Classification modeling, or modeling the problem as a classification task, specifically models the training process of the evaluator as a supervised learning problem based on discrete category labels. Multiple categories are preset, and the evaluator is trained to independently classify each item in a sequence into one of these categories, thus determining its corresponding value. For example, with two preset categories, clicked and unclicked, the evaluator is trained to determine the probability that each item in the sequence belongs to the clicked category, and uses this probability as its value score. If an item has a high probability of belonging to the clicked category, it can be considered that the item is more likely to be clicked by the user after being recommended; if the probability is low, it can be considered that the item is more likely not to be clicked by the user after being recommended. Based on this, the value scores of each item in the sequence can be aggregated into an overall value score for the entire sequence.
[0029] In Method 1, the evaluator independently assesses the value of each item in the same item sequence, thus ensuring that the values of different items are independent of each other. The value determined by the evaluator for each item can be understood as an item-level value, treating each item as a point in the item sequence; item-level value can also be understood as point-wise value. Accordingly, the evaluator in Method 1 is equivalent to performing item-level or point-level value assessment, and the training objective corresponding to the training process in Method 1 can be understood as an item-level training objective or a point-level training objective.
[0030] In practical applications, Method 1 can achieve high accuracy in point-level value assessment. However, the overall value score determined by Method 1 for a sequence of items cannot reflect the overall impact of the sequence structure on the value. For example, for a sequence of items, if the order of the items changes, the overall value score determined by Method 1 remains unchanged. Therefore, the evaluator trained based on Method 1 struggles to capture the overall value of a sequence of items, thus reducing the accuracy of value assessment.
[0031] Method 2: Regression modeling.
[0032] Regression modeling, in other words, models the problem as a regression task. Specifically, the training process for the evaluator is constructed as a supervised learning problem based on continuous value output. Through training, the evaluator is enabled to directly assess the overall value of a sequence of items, obtaining an overall value score corresponding to that sequence.
[0033] In Method 2, the overall value score determined by the evaluator can be used to reflect the impact of sequence structure on the overall value of the sequence. For example, for a sequence of items, if the order of items in the sequence changes, the overall value score determined based on Method 2 can change accordingly. Based on this, the corresponding value can be understood as a sequence-level value; if the item sequence is viewed as a list, the value can also be understood as a list-wise value. Accordingly, the evaluator in Method 2 is equivalent to performing a sequence-level value evaluation or a list-level value evaluation, and the training objective corresponding to the training process in Method 2 can be understood as a sequence-level training objective or a list-level training objective.
[0034] However, for Method 2, due to the complexity of the sequence combination space, the supervision signals during training are often sparse and have high variance, making the optimization process difficult. Furthermore, this training process easily leads the evaluator to focus too much on the relative order relationships within the item sequence, lacking supervision of the point-level value corresponding to individual items. Therefore, the evaluator struggles to learn fine-grained ranking patterns. Moreover, the determined value scores are difficult to use to explain the actual physical meaning of the items, such as whether the click-through rate is low; that is, absolute score degradation occurs. Based on these factors, Method 2 easily makes the training process of the evaluator unstable, reducing the accuracy of value assessment.
[0035] It can be seen that the evaluators trained in the relevant technologies have the problem of inaccurate value assessment, which reduces the accuracy of recommendations.
[0036] Based on this, in this embodiment of the application, a first sample data is determined, which includes multiple first item information sequences. Each item information in the first item information sequence is used to describe the item-side features and user-side features of an item. Then, based on the multiple first item information sequences, the first evaluation model is trained and iterated multiple times using a first loss function and a second loss function until a set first convergence condition is reached to obtain a second evaluation model. The first evaluation model includes a first evaluation network and a first auxiliary network. The first evaluation network is used to generate intermediate features based on the input item information sequences and to perform point-level value evaluation on the item information sequences based on the intermediate features. The first auxiliary network is used to perform sequence-level value evaluation on the item information sequences based on the intermediate features generated by the first evaluation network. The first loss function is used to measure point-level loss regarding value. The second loss function is used to measure sequence-level loss regarding value. In the above scheme, the first loss function and the second loss function are used together to train the first evaluation model. The first loss function measures the point-level loss related to value, which is equivalent to introducing a point-level training objective into the training process. The second loss function measures the sequence-level loss related to value, which is equivalent to introducing a sequence-level training objective into the training process. Therefore, the training process in this application scheme is equivalent to using a hybrid training objective based on point-level and sequence-level training objectives. On this basis, the trained first evaluation model can effectively capture the overall value of the item sequence and still ensure the accuracy of point-level value evaluation, stabilize the training process, and avoid the degradation of absolute scores. Thus, compared with related technologies, the accuracy of value evaluation is improved, thereby improving the accuracy of recommendation.
[0037] The present application will now be described in further detail with reference to the accompanying drawings and embodiments.
[0038] This application provides a model training method. In practical applications, this model training method can be used to support the relevant processing logic of a recommendation system in the recommendation process. Through this model training method, a trained evaluation model can be obtained. Based on this model, a generator can be trained, and then, during the reordering stage of the recommendation process, an item sequence can be generated using the trained generator to recommend items to the user based on this item sequence.
[0039] See Figure 1 The method includes: Step 101: Determine the first sample data.
[0040] The first sample data includes multiple sequences of first item information; each item information in the first item information sequence is used to describe the item-side features and user-side features of an item.
[0041] Here, the first sample data refers to the dataset used for training the model for evaluating the model. Each first item sequence included in the first sample data can be understood as a training sample.
[0042] In practical applications, in data processing related to items, an item sequence can be represented as an item information sequence. Each item information in the item information sequence corresponds to one item in the item sequence, and the order of the items in the item sequence is the same as the order of their corresponding information in the item information sequence. Processing an item sequence can be seen as processing an item information sequence, or in other words, processing an item information sequence can be understood as processing an item sequence. Similarly, processing items within an item sequence can be seen as processing item information within an item information sequence, or in other words, processing item information within an item information sequence can be understood as processing items within an item sequence.
[0043] Here, each item information in the first item information sequence is used to describe the item-side features and user-side features of an item.
[0044] In practical applications, item-side features can include the item's attribute features. Attribute features can be used to describe the item's own attributes, such as describing one or more of the following: item identifier, item brand, item price, and item sales volume.
[0045] User-side features corresponding to an item can be understood as the relevant characteristics of the user when an item is recommended to them. User-side features can include: user profile features and / or user behavioral features. User profile features can be used to describe a user persona, for example, one or more of the following: user ID, location, number of days registered, and membership level. User behavioral features can be used to describe the user's behavior in the user interface of the recommendation system, for example, one or more of the following: frequently viewed item categories, purchase rate, and click-through rate.
[0046] In practical applications, recommendation data generated during the operation of the recommendation system can be collected to determine the first sample data. For example, recommendation data may include one or more of the following: item attributes, user profiles, and user behavior data. It should be noted that the collection involved in this application is conducted with the user's awareness of the information collection and their authorization and consent.
[0047] Step 102: Based on multiple first item information sequences, the first evaluation model is trained and iterated multiple times using the first loss function and the second loss function until the set first convergence condition is met, thus obtaining the second evaluation model; The first evaluation model includes a first evaluation network and a first auxiliary network. The first evaluation network is used to generate intermediate features based on the input item information sequence and to perform point-level value evaluation on the item information sequence based on the intermediate features. The first auxiliary network is used to perform sequence-level value evaluation on the item information sequence based on the intermediate features generated by the first evaluation network. A first loss function is used to measure point-level loss related to value. A second loss function is used to measure sequence-level loss related to value.
[0048] Here, the first evaluation model refers to the evaluation model that has not been fully trained, and the second evaluation model refers to the evaluation model that has been fully trained. The evaluation model consists of two core networks: an evaluation network and an auxiliary network. The first evaluation network refers to the evaluation network that has not been fully trained, and the first auxiliary network refers to the auxiliary network that has not been fully trained.
[0049] Here, the first evaluation network is used to perform point-level value evaluation on the input item information sequence. In other words, the first evaluation network can independently evaluate the value of each item in the item information sequence, or in other words, the first evaluation network can independently evaluate the value of each item in the item sequence.
[0050] In practical applications, the first evaluation network can determine a value score for each item in the item information sequence, which is equivalent to having point-level scoring capabilities. The multiple value scores determined by the first evaluator for the item information sequence can be organized into a sequence, which can be called a score sequence. This score sequence can be used to form the value evaluation result determined by the first evaluation network.
[0051] In practical applications, the evaluation result output by the first evaluation network based on the input item information sequence may include one or more score sequences. For example, the evaluation result output by the first evaluation network may include two score sequences; one score sequence can be obtained by the first evaluation network processing the item-side features and corresponding user profile features in the input item information sequence, and the other score sequence can be obtained by the first evaluation network processing the item-side features, corresponding user profile features, and corresponding user behavioral features in the input item information sequence. In subsequent processing of the evaluation result, all or part of the score sequences can be selected for processing.
[0052] The first evaluation network can be understood as an evaluator. For example, the network architecture of the first evaluation network can be set up with reference to models such as Generative Rerank Network (GRN) or Personalized Re-ranking Model (PRM).
[0053] Here, the first evaluation network generates corresponding intermediate features during the processing of the input item sequence, and then performs point-level value evaluation based on these intermediate features. In practical applications, these intermediate features can be understood as the embedded features generated by the first evaluation network based on the input. The intermediate features generated by the first evaluation network can be input into the first auxiliary network.
[0054] Here, the first auxiliary network is used to perform sequence-level value assessment on the item information sequence based on the intermediate features generated by the first evaluation network. In other words, after the first evaluation network determines the intermediate features based on the input item information sequence and inputs these intermediate features into the first auxiliary network, the first auxiliary network can process these intermediate features to achieve sequence-level value assessment of the item information sequence input into the first evaluation network. Sequence-level value assessment of the item information sequence can also be understood as sequence-level value assessment of the corresponding item sequence.
[0055] Here, the value assessment performed by the first auxiliary network is a sequence-level value assessment, meaning that the value assessment result determined by the first auxiliary network reflects the overall value of the item information sequence. In practical applications, the first auxiliary network can be used to assist the training of the first evaluation network during the training of the first evaluation model. For example, the network architecture of the first auxiliary network can be set up with reference to a multilayer perceptron (MLP).
[0056] In practical applications, the first auxiliary network can determine and output a value assessment result for each item in the item information sequence. This value assessment result can be used to reflect the value score corresponding to that item. The value scores corresponding to the value assessment results determined by the first auxiliary network for each item are not independent of each other, but rather combine the overall characteristics of the item information sequence. These value assessment results can still be understood as sequence-level value assessment results.
[0057] Here, based on multiple sequences of first item information, the first evaluation model is trained and iterated multiple times using the first loss function and the second loss function until the set first convergence condition is reached.
[0058] In practical applications, supervised learning can be used to train the first evaluation model.
[0059] In multiple training iterations of the first evaluation model, the first evaluation model can be updated after each training iteration. Updating the first evaluation model may include updating the first evaluation network and / or the first auxiliary network. Updating the model may also include updating the model parameters. Model parameters can be understood as configuration variables within the model, such as the weights and biases of each node in the model. Model parameters in the first evaluation network and / or the first auxiliary network can also be understood as network parameters.
[0060] The first convergence condition can include a calculated loss value being less than a set threshold, whereby the loss value can be calculated during each training iteration. The loss value can be used to describe the degree of inconsistency between the output obtained based on the model input and the target output obtained based on the model input.
[0061] Here, during the training of the first evaluation model, both the first and second loss functions are utilized. In other words, in multiple training iterations of the first evaluation model, both the first and second loss functions are used, effectively combining them.
[0062] In practical applications, in each training iteration, the loss value can be calculated based on the first loss function and / or the second loss function, and then the first evaluation model can be updated based on the loss value.
[0063] Here, the first loss function is used to measure point-level loss regarding value, which is equivalent to introducing a point-level training objective into the training process. Based on this, for a model or network updated through the first loss function, the model or network can learn the ability to evaluate point-level value during training, thus ensuring the accuracy of point-level value evaluation after training is completed.
[0064] Here, the second loss function measures the sequence-level loss regarding value, essentially introducing a sequence-level training objective into the training process. In practical applications, the second loss function can model the relative utility relationships between items. For example, these relationships can include the impact of different item rankings on the number of purchases, clicks, or impressions for each item. Based on this, the model or network updated using the second loss function can learn the ability to evaluate value at the sequence level during training, thus accurately capturing the overall value of the item sequence after training is complete.
[0065] In this embodiment, a first loss function and a second loss function are jointly used to train the first evaluation model. The first loss function measures point-level loss related to value, essentially introducing a point-level training objective into the training process. The second loss function measures sequence-level loss related to value, essentially introducing a sequence-level training objective into the training process. Therefore, the training process in this application is equivalent to using a hybrid training objective based on point-level and sequence-level training objectives. In practical applications, the point-level training objective can be understood as the training objective corresponding to classification modeling in related technologies, and the sequence-level training objective can be understood as the training objective corresponding to regression modeling in related technologies. Therefore, the embodiment of this application establishes a correlation between classification modeling and regression modeling.
[0066] Based on this, the first evaluation model trained by the embodiment of this application can effectively capture the overall value of the item sequence, while still ensuring the accuracy of point-level value evaluation, stabilizing the training process, and avoiding the degradation of absolute scores. Thus, compared with related technologies, it improves the accuracy of the evaluator's value evaluation.
[0067] In this embodiment, a second evaluation model is obtained by training a first evaluation model. In practical applications, after obtaining the second evaluation model, the generator can be trained based on it, thereby enhancing the training effect of the generator and improving the quality of the item sequences generated by the trained generator. After obtaining the trained generator, item sequences can be generated based on the generator during the reordering stage of the recommendation process, and item recommendations can be made to users based on the item sequences generated by the generator, thereby improving the accuracy of the recommendations.
[0068] The training process based on the first loss function and the second loss function will be further explained below.
[0069] In one embodiment, based on each of a plurality of first item information sequences, a first evaluation model is trained iteratively multiple times using a first loss function and a second loss function, including: The first evaluation network is invoked to process the first item information sequence to obtain a first evaluation result and a first intermediate feature; and the first auxiliary network is invoked to process the first intermediate feature to obtain a second evaluation result. The first evaluation result is processed based on the first loss function to obtain the first loss value, and the first evaluation network is updated based on the first loss value; and the second evaluation result is processed based on the second loss function to obtain the second loss value, and the first auxiliary network and the first evaluation network are updated sequentially based on the second loss value.
[0070] In practical applications, each of the multiple first item information sequences can be understood as a training sample. When training the first evaluation model multiple times based on each first item information sequence, the first evaluation network can be invoked to process the first item information sequence to obtain a first evaluation result. Then, a first loss function can be invoked to process the first evaluation result and the labels of the first item information sequence to obtain a first loss value. Furthermore, a first auxiliary network can be invoked to process the first intermediate features generated by the first evaluation network during processing to obtain a second evaluation result. Then, a second loss function can be invoked to process the second evaluation result and the labels of the first item information sequence to obtain a second loss value.
[0071] In practical applications, the initial recommendation data generated by the recommendation system during operation can be collected to determine the relevant features of items. Then, based on these features, a first item sequence can be constructed, thus creating the first training sample. The initial recommendation data may include: relevant recommendation data of item sequences exposed online by the recommendation system.
[0072] In practical applications, the label of the first item information sequence can be understood as the label of the training samples. The label of the first item information sequence can be used to describe the user's interaction with the items in the item sequence after the first item information sequence is recommended to the user.
[0073] For example, the tags of the first item information sequence may be used to describe one or more of the following: user purchases of items, user clicks on items, and item exposure.
[0074] For example, the tags of the first item information sequence may include one or more of the following: a purchase tag, a click tag, and an exposure tag. The purchase tag can be used to describe whether a user has purchased an item; for example, if the user has purchased an item, the corresponding value in the purchase tag can be 1, and if the user has not purchased an item, the corresponding value in the purchase tag can be 0. The click tag can be used to describe whether a user has clicked an item; for example, if the user has clicked an item, the corresponding value in the click tag can be 1, and if the user has not clicked an item, the corresponding value in the click tag can be 0. The exposure tag can be used to describe whether the item has been exposed; for example, if the item has been exposed, the corresponding value in the exposure tag can be 1, and if the item has not been exposed, the corresponding value in the exposure tag can be 0.
[0075] In practical applications, updating the first auxiliary network and the first evaluation network sequentially based on the second loss value can include: updating the first auxiliary network based on the second loss value, and then backpropagating the gradient generated by the first auxiliary network during the update process to the first evaluation network, thereby updating the first evaluation network based on this gradient. The gradient generated by the first auxiliary network during the update process carries information related to the second loss value; therefore, after backpropagating this gradient to the first evaluation network, the update performed on the first evaluation network can be considered as an update based on the second loss value. In practical applications, this update method can also be regarded as a collaborative update of the first auxiliary network and the first evaluation network. Through collaborative updating, the relevant gradient information of the first auxiliary network used for sequence-level value evaluation can be passed to the first evaluation network, thereby enhancing the ability of the trained first evaluation network to capture the overall value of the item sequence.
[0076] In this embodiment, both the first loss value and the second loss value are used to update the first evaluation network. This is equivalent to introducing a hybrid training objective based on point-level training objectives and sequence-level training objectives into the first evaluation network during the training process. In this way, the trained first evaluation network can effectively capture the overall value of the item sequence while still ensuring the accuracy of point-level value evaluation. This stabilizes the training process and avoids the degradation of absolute scores, thereby improving the accuracy of value evaluation compared to related technologies, and thus improving the accuracy of recommendations.
[0077] In one embodiment, the second evaluation result is processed based on a second loss function to obtain a second loss value, and the first auxiliary network and the first evaluation network are updated sequentially based on the second loss value, including: In the first training iteration of multiple training iterations of the first evaluation model, the second evaluation result is processed based on the second loss function to obtain the second loss value, and the first auxiliary network and the first evaluation network are updated sequentially based on the second loss value. The first evaluation result is processed based on the first loss function to obtain a first loss value, and the first evaluation network is updated based on the first loss value, including: In the second training iteration of multiple training iterations of the first evaluation model, the first evaluation result is processed based on the first loss function to obtain the first loss value, and the first evaluation network is updated based on the first loss value. Wherein, the first training iteration represents one training iteration performed every n training iterations; the second training iteration represents other training iterations besides the first training iteration in multiple training iterations; n is a positive integer.
[0078] In practical applications, the first training iteration and the second training iteration can together form multiple training iterations for the first evaluation model.
[0079] Here, the first evaluation network was updated in both the first and second training iterations. In other words, the first evaluation network was updated in each training iteration across multiple training iterations.
[0080] Here, the first auxiliary network and the first evaluation network are updated sequentially based on the second loss value only during the first training iteration. That is, the first auxiliary network and the first evaluation network are updated sequentially based on the second loss value only every n training iterations, which is equivalent to interval learning of the sequence-level loss, thus improving the stability of the evaluator learning.
[0081] In practical applications, the first loss function can be constructed based on the cross-entropy loss function and the list-wise maximum likelihood estimation (ListMLE) loss function.
[0082] The first loss function can be a linear combination of the cross-entropy loss function and the ListMLE loss function. The loss value corresponding to the first loss function can be a weighted sum of the loss values corresponding to the cross-entropy loss function and the ListMLE loss function.
[0083] For example, the first evaluation result may include two score sequences, which can be denoted as output_logits and pv_logits, respectively. Output_logits can be obtained by the first evaluation network processing the item-side features and corresponding user profile features in the first item information sequence. Pv_logits can be obtained by the first evaluation network processing the item-side features, corresponding user profile features, and corresponding user behavior features in the input first item information sequence. The first loss value calculated based on the first loss function can consist of multiple losses, including base_loss, pv_loss, order_loss, and clk_loss. Base_loss can be determined based on the cross-entropy loss function processing of the click labels of the first item information sequence and output_logits in the first evaluation result. Pv_loss can be determined based on the cross-entropy loss function processing of the click labels of the first item information sequence and pv_logits in the first evaluation result. The order_loss can be determined based on the ListMLE loss function by processing the purchase tags of the first item information sequence and the output_logits in the first evaluation result. The clk_loss can be determined based on the ListMLE loss function by processing the click tags of the first item information sequence and the output_logits in the first evaluation result.
[0084] For example, the loss value corresponding to the cross-entropy loss function can be expressed as: , in, The loss value corresponding to the cross-entropy loss function. The number of items in the first item information sequence. It represents the number of item information in the first item information sequence. A label representing the sequence of information for the first item. The first evaluation network represents the first item in the first item information sequence. The estimated probability of the output of the item information, which represents the probability of the first item. The probability that an item belongs to a corresponding tag. It was determined based on the results of the first assessment.
[0085] For example, the loss value corresponding to the ListMLE loss function can be expressed as: , in, The loss value corresponding to the ListMLE loss function. The number of items in the first item information sequence. Represents the number of items in the first item sequence. Characterizes a monotonically increasing function. The first evaluation network represents the first item in the reordered first item information sequence. The estimated probability of each item's output. The first evaluation network represents the first item in the reordered first item information sequence. The estimated probability of each item information output. The reordered first item information sequence can be obtained based on the following processing: Based on the first evaluation result, determine the value score corresponding to the item information in the first item information sequence, and sort the item information in the first item information sequence in descending order according to the value score to obtain the reordered first item information sequence.
[0086] For example, the loss value corresponding to the first loss function can be expressed as: , in, The loss value corresponding to the first loss function is represented. The characterization is based on the loss value corresponding to the cross-entropy loss function. Characterization The corresponding weights Characterizes the loss value corresponding to the ListMLE loss function. Characterization The corresponding weights.
[0087] In practical applications, the second loss function can include the Classification-Restoration framework with Error-Adaptive Discretization (CREAD) loss function. CREAD can be understood as a processing framework, or simply the CREAD framework; the CREAD loss function is the loss function used within this framework. The CREAD framework can be applied to the second model, and the relevant loss values can be calculated using the CREAD loss function.
[0088] In practical applications, the CREAD framework's processing flow mainly includes the following stages: discretization, classification, and restoration. Discretization divides value scores with continuous value ranges into several ordered intervals, ensuring the comparability of values for each item within the item information sequence. Classification allows the first auxiliary network to estimate the probability of a value score exceeding the upper limit of an interval, rather than directly estimating specific value scores. Restoration converts the probabilities estimated in the separation stage back into specific value scores. Thus, the calculated value scores can comprehensively measure the value of the item information sequence.
[0089] The loss value corresponding to the CREAD loss function can be composed of the classification loss value in the classification stage and the recovery loss value in the recovery stage.
[0090] In practical applications, a first target score can be determined for each item based on the user's interaction with the items in the first item information sequence. Then, for each item's first target score, a first vector is determined based on the comparison between the first target score and the upper limit of each value score interval. After that, a label for the first item information sequence is formed based on the multiple first vectors obtained. This label can be used for loss value calculation under the CREAD framework.
[0091] The first target score corresponding to the item information can be represented as the weighted sum of the second target scores corresponding to each item information related to that item information. The weight of the item information in this weighted sum can be related to the position of the item information in the first item information sequence. For example, the first... i The relevant item information for an item may include: the first i Information on the first item, and the first... i Before the item information i -1 item information; iIt is a positive integer; the weight of the item information in the weighted sum can be represented as the reciprocal of the location index value.
[0092] The first vector can be determined by comparing the first target score corresponding to the first item information sequence with the upper limit of each value score interval.
[0093] Each element in the first vector corresponds to a comparison result for a score interval. For example, if the first target score for an item is greater than or equal to the upper limit of a score interval, the element corresponding to that score interval in the first vector can have a value of 1; if the first target score for an item is less than the upper limit of a score interval, the element corresponding to that score interval in the first vector can have a value of 0.
[0094] For example, suppose the first item information sequence contains 20 item information items. The user clicked and purchased the item corresponding to the first item information item, clicked the item corresponding to the third item information item, and did not interact with any other items besides those corresponding to the first and third item information items. Based on this, the second target score corresponding to the first item information item can be 2, the second target score corresponding to the third item information item can be 1, and the second target score corresponding to the other item information items can be 0. Based on this, the first target score corresponding to each item information item in the first item information sequence can be calculated. For example, the first target score corresponding to the third item information item can be 2.333 (=2+(1 / 2)*1). Further suppose the value score range is divided into three intervals, with upper limits of 0.25, 0.5, and 1 respectively. Based on this, the first vector can be determined by comparing the first target score corresponding to each item information item in the first item information sequence with the upper limits of each interval. For example, the first vector corresponding to the third item information item can be set to [1,1,1]. Then, the first vector can be included in the labels of the first training samples to support the calculation of the CREAD loss function.
[0095] Based on the model training methods in the above embodiments, this application also provides a model training method. In practical applications, the model training method provided here can be used to support the relevant processing logic of the recommendation system in the recommendation process. This model training method can obtain a trained generator. Based on this, in the reordering stage of the recommendation process, the trained generator can generate item sequences to recommend items to users based on these item sequences.
[0096] The method includes: Step 201: Determine the second sample data.
[0097] The second sample data includes multiple sets of first item information; each item information in the first item information set is used to describe the item-side features and user-side features of an item.
[0098] Here, the second sample data refers to the data set used for training the generative model, and each set of first information included in the second sample data can be understood as a training sample.
[0099] In practical applications, in data processing related to items, an item set can be represented as a set of item information. Each piece of item information in the item information set corresponds to an item in the item set. Processing an item set can be seen as processing the item information set, or in other words, processing the item information set can be understood as processing the item set. Similarly, processing an item within an item set can be seen as processing the item information within the item information set, or in other words, processing the item information within the item information set can be understood as processing the items within the item set.
[0100] In practical applications, the first item information set can be constructed based on the recommendation data generated during the operation of the recommendation system. For example, the first item information set may include: the item information set generated by the recommendation system in the ranking stage of the recommendation process, which can be regarded as the item information set generated through recall and ranking processes.
[0101] Step 202: Based on multiple sets of first item information, use the second evaluation model to train and iterate the first generation model multiple times until the set second convergence condition is met, and obtain the second generation model.
[0102] The second evaluation model is trained using the model training method described in any of the preceding items.
[0103] Here, the first generative model refers to the generative model that has not been trained, and the second generative model refers to the generative model that has been trained.
[0104] In practical applications, a generative model can be understood as a generator. A generator rearranges the input set of item information to output a sequence of item information. The items corresponding to the item information in the generator's output sequence can be understood as items selected from the items corresponding to each item in the input set of item information, and these selected items are arranged in a specific sorting order to form the item information sequence.
[0105] In practical applications, during multiple training iterations of the first generative model, it can be updated after each iteration. Updating the first generative model can include updating its model parameters. Model parameters can be understood as configuration variables within the model, such as the weights and biases of each node.
[0106] The second convergence condition can include a calculated loss value being less than a set threshold, whereby the loss value can be calculated during each training iteration. The loss value can be used to describe the degree of inconsistency between the output obtained based on the model input and the target output obtained based on the model input.
[0107] Here, a second generative model is obtained by training the first generative model. In practical applications, after obtaining the second generative model, item information sequences can be generated based on the second generative model during the reordering stage of the recommendation process, and then item recommendations can be made to users based on the item information sequences generated by the second generator.
[0108] In the model training method provided in the foregoing embodiments, the present application employs a hybrid training objective based on point-level and sequence-level training objectives to train the first evaluation model, resulting in a second evaluation model. Thus, the second evaluation model can effectively capture the overall value of the item sequence while still ensuring the accuracy of point-level value assessment. Based on this, the second evaluation model can provide accurate training guidance during the training of the first generation model, thereby enhancing the training effect of the first generation model. Furthermore, the trained first generator can generate high-quality item information sequences, or in other words, high-quality item sequences, thereby improving the accuracy of recommendations.
[0109] In one embodiment, based on each of a plurality of first item information sets, a second evaluation model is used to perform multiple training iterations on a first generative model, including: The first generation model is invoked to process the first set of item information to obtain a second sequence of item information; each item information in the second sequence of item information is used to identify an item. The second evaluation model is invoked to process the second item information sequence to obtain the third evaluation result; the third evaluation result represents the evaluation result output by the evaluation network in the second evaluation model. The first value score corresponding to the second item information sequence is determined based on the third evaluation results; Update the first generative model based on the first value score.
[0110] In practical applications, calling the second evaluation model to process the second item information sequence and obtain the third evaluation result can include: calling the evaluation network in the second evaluation model to process the second item information sequence and obtain the third evaluation result.
[0111] In the model training method provided in the foregoing embodiments, the present application adopts a hybrid training objective based on point-level training objectives and sequence-level training objectives to train the first evaluation model, thereby obtaining a second evaluation model. Based on this, the evaluation network in the second evaluation model can effectively capture the overall value of the item sequence while still ensuring the accuracy of point-level value evaluation. Therefore, the third evaluation result output by the evaluation network in the second evaluation model can be used to describe both the sequence-level value and the point-level value of the second item information sequence, thus accurately reflecting the value of the second item information sequence.
[0112] Here, the first value score is used to update the first generative model. In practical applications, the first value score can guide the update of the first generative model during multiple training iterations. Since the third evaluation result used to calculate the first value score can accurately reflect the value of the second item information sequence, the first generative model can obtain accurate training guidance, thereby enhancing the training effect of the first generative model. Based on this, the trained first generative model can generate high-quality item sequences, improving the accuracy of recommendations.
[0113] In one embodiment, determining the first value score corresponding to the second item information sequence based on the third evaluation result includes: Based on the third evaluation results, the item information in the second item information sequence is reordered to obtain the third item information sequence; For each item information in the third item information sequence, the difference between the first position index value corresponding to the item information in the third item information sequence and the second index value corresponding to the item information in the second item information sequence is calculated to obtain the first difference value corresponding to the item information. The first value score is obtained by weighted summation of multiple first differences corresponding to the third item information sequence; the first weight of the first difference in the first value score is related to the first position index value corresponding to the first difference.
[0114] Here, by reordering the item information in the second item information sequence, a third item information sequence is obtained. That is, the item information in the third item information sequence is the same as that in the second item information sequence, but the arrangement of the item information in the third item information sequence is different from that in the second item information sequence. In practical applications, the third item information sequence can be a higher quality item information sequence than the second item information sequence; in other words, the item sequence represented by the third item information sequence can have a higher quality, or better recommendation effect, compared to the item sequence represented by the second item information sequence.
[0115] In practical applications, the position index value can be used to describe the position of item information within an item information sequence. The first difference can be used to describe the change in the position of item information in the third item information sequence compared to its position in the second item information sequence. Therefore, the first value score determined based on the first difference can be used to describe the sequence structure difference between the second and third item information sequences, or in other words, the structural alignment, thus reflecting the structural difference between the item information sequence generated by the first generative model and the higher-quality item information sequence provided by the second evaluation model. Based on this, using the first value score to guide the training of the first generative model enables it to learn higher-quality item information sequence generation methods, enhancing the training effect of the first generative model.
[0116] Here, the first value score represents the weighted sum of multiple first differences corresponding to the third item information sequence. The first weight of each first difference in the first value score is related to the position of the corresponding item information in the third item sequence. Thus, the first value score can reflect the influence of the position of the item information in the sequence on the sequence quality, or in other words, on the recommendation effect, improving the accuracy of the first value score in describing the sequence, thereby enhancing the training effect of the first generative model.
[0117] In practical applications, after recommending items to users based on item information sequences, users typically prioritize items that appear earlier in the sequence, meaning they focus their attention on those items first. Based on this, a first weight can be set, increasing the weight of items appearing earlier in the third item information sequence. This explicitly incorporates location-related user attention during training, guiding the first generation model to prioritize the location where user attention is focused through a first value score, thus enhancing the training effect of the first generation model.
[0118] In practical applications, when the position index value of an item is negatively correlated with the degree to which the item is positioned, the first weight can be negatively correlated with the position index value of the item in the third item information sequence. The negative correlation between the position index value of an item and the degree to which the item is positioned can be understood as the smaller the position index value of the item in the item information sequence, the earlier the item is positioned in the item information sequence.
[0119] When the position index value of an item is positively correlated with the degree to which the item is located in the order of its position, the first weight can be set to be positively correlated with the position index value of the item in the third item information sequence. The positive correlation between the position index value of an item and the degree to which the item is located in the order of its position can be understood as the larger the position index value of the item in the item information sequence, the later the item is located in the item information sequence.
[0120] For example, the first value score can be represented as: , in, Characterizing the first value score, Represents the index value of the first position. Represents the index value of the second position. Characterizes the first weight, It represents the number of item information in the third item information sequence.
[0121] For example, the first weight can be expressed as: , in, Characterizes the first weight, The first in the sequence of information representing items The location index corresponding to each item's information.
[0122] In one embodiment, the third evaluation result includes a first score sequence consisting of a plurality of second value scores; Based on the third evaluation results, the item information in the second item information sequence is reordered to obtain the third item information sequence, which includes: Based on the second value score corresponding to each item in the second item information sequence in the first score sequence, the item information in the second item information sequence is sorted to obtain the third item information sequence; each second value score in the first score sequence is used to describe the value corresponding to an item in the second item information sequence.
[0123] In practical applications, sorting the item information in the second item information sequence can include sorting the item information in the second item information sequence in descending order.
[0124] Here, based on the second value score, the items in the second item information sequence are sorted. In this way, the items at the beginning of the third item information sequence can have a higher second value score, making the third item information sequence have a higher sequence quality than the second item information sequence. This ensures that the first generative model can learn a higher quality item information sequence generation method during training, thus enhancing the training effect.
[0125] In one embodiment, based on each of a plurality of first item information sets, the first generative model is trained and iterated multiple times using a second evaluation model, and the method further includes: The second evaluation model is invoked to process the fourth item information sequence corresponding to the first item information set, and the fourth evaluation result is obtained. The fourth item information sequence represents the item information sequence preset for the first item information set, and each item information in the fourth item information sequence is used to identify an item. The fourth evaluation result represents the evaluation result output by the evaluation network in the second evaluation model. The third value score corresponding to the fourth item information sequence is determined based on the fourth assessment results; Correspondingly, based on the first value score, the first generative model is updated, including: The first generative model is updated based on the second difference between the third value score and the first value score.
[0126] In practical applications, the fourth item information sequence can be understood as the initial (base) item information sequence of the first item set, which can be used as the adjustment benchmark for the first generative model during training.
[0127] Here, the second evaluation model is invoked to process the fourth item information sequence corresponding to the first item information set, resulting in a fourth evaluation result. Based on this fourth evaluation result, a third value score is determined for the fourth item information sequence. In practical applications, the generation method of the fourth evaluation result can be understood in accordance with the generation method of the third evaluation result, and the generation method of the third value score can be understood in accordance with the generation method of the second value score.
[0128] In practical applications, determining the third value score corresponding to the fourth item sequence based on the fourth evaluation result can include: Based on the results of the fourth evaluation, the item information in the fourth item information sequence is reordered to obtain the sixth item information sequence; For each item in the sixth item information sequence, the difference between the third position index value of the item information in the sixth item information sequence and the fourth position index value of the item information in the fourth item information sequence is calculated to obtain the third difference value of the item information. The third value score is obtained by weighted summation of multiple third differences corresponding to the information sequence of the sixth item; the second weight of the third difference in the third value score is related to the third position index value corresponding to the third difference.
[0129] In practical applications, the fourth evaluation result may include a second score sequence consisting of multiple third value scores; each third value score in the second score sequence is used to describe the value of an item information in the fourth item information sequence; Based on the results of the fourth evaluation, the item information in the fourth item information sequence may be reordered, which may include: Based on the third value score corresponding to each item in the fourth item information sequence in the second score sequence, the item information in the fourth item information sequence is sorted to obtain the sixth item information sequence.
[0130] Here, the first generative model is updated based on the second difference between the third value score and the first value score. In practical applications, during the training of the first generative model, the fourth item information sequence can be used as the learning benchmark for the first generative model. The second difference is used to measure the improvement in the sequence quality of the item information sequence generated by the first generative model compared to the fourth item information sequence, thereby guiding the training of the first generative model more efficiently and accurately and enhancing the training effect.
[0131] Based on the model training methods described in the above embodiments, this application also provides a recommendation method. In practical applications, this recommendation method can be used to support the relevant processing logic of the recommendation system in the recommendation process. Through this recommendation method, a high-quality item sequence can be generated during the reordering stage of the recommendation process, and then item recommendations can be made to users based on this item sequence.
[0132] See Figure 3 The method includes: Step 301: Call the second generation model to process the second item information set to obtain the fifth item information sequence.
[0133] The second item information set represents the item information set determined through recall and sorting processes.
[0134] In practical applications, recall processing can include the processing performed by the recommendation system during the recall phase of the recommendation process. Ranking processing can include the processing performed by the recommendation system during the ranking phase of the recommendation process.
[0135] Step 302: Recommend items to the user based on the fifth item information sequence.
[0136] The second generative model is obtained by training using the model training method described in any of the preceding items.
[0137] In practical applications, recommending items to users based on the fifth item information sequence can include recommending a first item sequence corresponding to the fifth item information sequence. The item information in the fifth item information sequence can be used to indicate each item in the first item sequence, and the order of the item information in the fifth item information sequence is the same as the order of the items in the first item sequence. The generation of the fifth information sequence by the second generative model can also be considered as the generation of the first item sequence. Recommending items to users based on the fifth information sequence can also be considered as recommending items to users based on the first item sequence.
[0138] Recommending a first item sequence to a user may include presenting the first item sequence in the user interface of the recommendation system. For example, the display controls corresponding to each item in the first item sequence may be presented sequentially in the user interface according to the order in which the items are arranged.
[0139] In this embodiment, the second generation model is trained under the guidance of the second evaluation model. In the model training method of this embodiment, a hybrid training objective consisting of point-level and sequence-level training objectives is used to train the first evaluation model, resulting in the second evaluation model. Therefore, the second evaluation model can effectively capture the overall value of the item sequence while still ensuring the accuracy of point-level value evaluation. Based on this, the accurate value evaluation of the second evaluation model enhances the training effect of the first generation model. Furthermore, during the training process of the first generation model, this solution guides the first generation model to prioritize the location where the user's attention is focused, further enhancing the training effect of the first generation model. Based on this, the quality of the item sequences generated by the second generation model is improved, thereby increasing the accuracy of the recommendations.
[0140] The present application will be further described in detail below with reference to application embodiments.
[0141] This application provides a unified sequence value assessment framework (USVA). The USVA framework can be used for recommendation-related model training and application. The corresponding processing logic can be understood by referring to the foregoing method embodiments.
[0142] In practical applications, the USVA framework can be used for evaluator-related model training.
[0143] For example, the model training process related to the evaluator may include the following steps: Step 1: Data collection.
[0144] In practical applications, recommendation data can be collected. For example, recommendation data may include one or more of the following: item attributes, user profiles, and user behavior data. This collected data can be used to construct training samples.
[0145] Step 2: Model building.
[0146] In practical applications, an evaluator and an auxiliary network can be constructed, and these two networks can be combined to form an evaluation model. The evaluator can be trained by training the evaluation model.
[0147] The evaluator can have point-level scoring capabilities, equivalent to the evaluation network in the embodiments of this application. The auxiliary network can be used to predict sequence values, thereby assisting in the training of the evaluator. Predicting sequence values can also be referred to as evaluating sequence values.
[0148] Step 3: Model training.
[0149] In practical applications, supervised learning can be used to train the evaluation model.
[0150] During the training of the evaluation model, multiple training iterations can be performed based on multiple training samples until the set convergence condition is met. Each training sample may include a sequence of item information.
[0151] In multiple training iterations of the evaluation model, two types of loss can be used: point-level prediction loss and sequence-level prediction loss.
[0152] The point-level prediction loss is equivalent to the loss corresponding to the first loss function in the embodiments of this application, and can also be called the point-level evaluation loss; the sequence-level prediction loss is equivalent to the loss corresponding to the second loss function in the embodiments of this application, and can also be called the sequence-level evaluation loss.
[0153] During training, point-level prediction loss can be used solely to update the evaluator, while sequence-level prediction loss can be used to update both the evaluator and the auxiliary network. Evaluator updates can be performed in each training iteration across multiple training iterations. Updates to the auxiliary network and the learning of sequence-level prediction loss can be performed intermittently across multiple training iterations.
[0154] For example, see Figure 4During the forward propagation phase of training, the evaluator can be invoked to process the sequence to be evaluated, and the intermediate features determined during the processing can be input into the auxiliary network. The sequence to be evaluated is equivalent to the training samples. The auxiliary network can also be invoked to process the intermediate features determined by the evaluator. During the backpropagation phase of training, the auxiliary network can be updated based on loss 3 every n training iterations, and the gradient of the auxiliary network with respect to loss 3 can be backpropagated to the evaluator to update it. In other training iterations, the evaluator can be updated based only on loss 1 and loss 2. Figure 4 In this context, loss 3 can be understood as sequence-level prediction loss. For example, loss 3 can include the loss corresponding to the CREAD loss function. Loss 1 and loss 2 can be used together to construct point-level prediction loss. For example, loss 1 can include the loss corresponding to the cross-entropy loss function, and loss 2 can include the loss corresponding to the ListMLE loss function.
[0155] Here, point-level prediction loss is equivalent to introducing a point-level training objective, sequence-level prediction loss is equivalent to introducing a sequence-level training objective, and combining point-level and sequence-level prediction losses for the relevant training of the evaluator is equivalent to introducing a hybrid training objective. In other words, the USVA framework in this application embodiment introduces a hybrid training objective mechanism. Thus, by combining point-level stability and sequence-level relationship modeling, training stability and value evaluation accuracy are ensured.
[0156] In practical applications, the USVA framework can be used for evaluator-related model applications. Specifically, it can guide the training of the generator based on the evaluator in the trained evaluation model.
[0157] During the training process of the generator, it can be trained iteratively multiple times based on multiple training samples until the set convergence condition is reached. Each training sample can include a set of item information, and an initial item information sequence can be preset for each set of item information.
[0158] In multiple training iterations of the generator, the generator can be invoked to process the item information set in the training samples to obtain the item information sequence output by the generator. For ease of description, the item information sequence output by the generator will be referred to as the generated sequence, and the initial item information sequence corresponding to the item information set input to the generator will be referred to as the candidate sequence. The generated sequence is equivalent to the second item information sequence in this embodiment, and the candidate sequence is equivalent to the fourth item information sequence in this embodiment. Then, the generated sequence and the candidate sequence can be processed respectively based on the evaluator in the trained evaluation model to obtain the sequence value score corresponding to the generated sequence and the sequence value score corresponding to the candidate sequence. The generator is then updated based on the difference between these two sequence value scores. Here, the sequence value score corresponding to the generated sequence is equivalent to the first value score in this embodiment, and the sequence value score corresponding to the candidate sequence is equivalent to the third value score in this embodiment.
[0159] For sequence value scores, see Figure 5 It can be calculated in the following way: The original sequence is processed by the evaluator in the trained evaluation model to obtain the value evaluation result. Then, the item information in the original sequence is reordered based on the value evaluation result to obtain a new sequence. After that, for each item information in the new sequence, the difference between the position index value of the item information in the new sequence and the position index value of the item information in the original sequence is calculated to obtain the corresponding difference value. Then, multiple differences corresponding to the new sequence are weighted and summed to obtain the value evaluation result. The weight of the difference value in the value evaluation result is related to the position index of the corresponding item information in the new sequence.
[0160] In the process of calculating the sequence value score, the original sequence can be understood as the generated sequence or candidate sequence processed by the evaluator, and the new sequence can be understood as the item information sequence obtained after reordering the original sequence.
[0161] In the process of calculating the sequence value score, the application embodiment of this application measures the structural alignment between the original sequence and the new sequence and explicitly incorporates position-related user attention. In this way, the generator can be guided to pay priority to the position where the user's attention is focused, which is equivalent to integrating the position-aware reward (PAR) mechanism into the USVA framework, thereby enhancing the training effect of the generator.
[0162] In practical applications, based on the trained generator, a high-quality item sequence can be generated during the reordering stage of the recommendation process in the recommendation system, and then the item sequence can be recommended to the user.
[0163] In the application embodiments of this application, the USVA framework achieves a balance between sequence modeling and optimization efficiency by combining a training objective mechanism with a location-aware reward mechanism, thereby improving the accuracy of recommendations.
[0164] Based on the foregoing embodiments, this application provides a model training device that can be applied to... Figure 1 In the model training method provided in the corresponding embodiment, refer to Figure 6 As shown, the model training device 60 may include: a first determining unit 601 and a first training unit 602, wherein: The first determining unit 601 is used to determine the first sample data; the first sample data includes a plurality of first item information sequences; each item information in the first item information sequence is used to describe the item-side features and user-side features of an item; The first training unit 602 is used to train the first evaluation model multiple times based on the multiple first item information sequences, using a first loss function and a second loss function, until a set first convergence condition is reached, to obtain the second evaluation model. The first evaluation model includes a first evaluation network and a first auxiliary network. The first evaluation network is used to generate intermediate features based on the input item information sequence and to perform point-level value evaluation on the item information sequence based on the intermediate features. The first auxiliary network is used to perform sequence-level value evaluation on the item information sequence based on the intermediate features generated by the first evaluation network. The first loss function is used to measure point-level loss with respect to value. The second loss function is used to measure sequence-level loss with respect to value.
[0165] In one embodiment, the first training unit 602, based on each of the plurality of first item information sequences, performs multiple training iterations on the first evaluation model using a first loss function and a second loss function, including: The first evaluation network is invoked to process the first item information sequence to obtain a first evaluation result and a first intermediate feature; and the first auxiliary network is invoked to process the first intermediate feature to obtain a second evaluation result. The first evaluation result is processed based on a first loss function to obtain a first loss value, and the first evaluation network is updated based on the first loss value; and the second evaluation result is processed based on a second loss function to obtain a second loss value, and the first auxiliary network and the first evaluation network are updated sequentially based on the second loss value.
[0166] In one embodiment, the first training unit 602 processes the second evaluation result based on a second loss function to obtain a second loss value, and updates the first auxiliary network and the first evaluation network sequentially based on the second loss value, including: In the case of the first training iteration in the multiple training iterations of the first evaluation model, the second evaluation result is processed based on the second loss function to obtain the second loss value, and the first auxiliary network and the first evaluation network are updated sequentially based on the second loss value. The first training unit 602 processes the first evaluation result based on a first loss function to obtain a first loss value, and updates the first evaluation network based on the first loss value, including: In the case of the second training iteration in the multiple training iterations of the first evaluation model, the first evaluation result is processed based on the first loss function to obtain the first loss value, and the first evaluation network is updated based on the first loss value. Wherein, the first training iteration represents one training iteration performed every n training iterations; the second training iteration represents other training iterations in the multiple training iterations besides the first training iteration; and n is a positive integer.
[0167] It should be noted that a detailed explanation of the steps performed by each unit can be found in [reference needed]. Figure 1 The model training method provided in the corresponding embodiments will not be described in detail here.
[0168] Based on the foregoing embodiments, this application provides a model training device that can be applied to... Figure 2 In the model training method provided in the corresponding embodiment, refer to Figure 7 As shown, the model training device 70 may include: a second determining unit 701 and a second training unit 702, wherein: The second determining unit 701 is used to determine the second sample data; the second sample data includes multiple first item information sets; each item information in the first item information set is used to describe the item-side features and user-side features of an item; The second training unit 702 is used to train the first generative model multiple times based on the plurality of first item information sets and using a second evaluation model, until a set second convergence condition is met, to obtain the second generative model; the second evaluation model is obtained through... Figure 1 The model was trained using the model training method provided in the corresponding embodiment.
[0169] In one embodiment, the second training unit 702, based on each of the plurality of first item information sets, uses a second evaluation model to perform multiple training iterations on the first generative model, including: The first generation model is invoked to process the first item information set to obtain a second item information sequence; each item information in the second item information sequence is used to identify an item. The second evaluation model is invoked to process the second item information sequence to obtain a third evaluation result; the third evaluation result represents the evaluation result output by the evaluation network in the second evaluation model. Based on the third evaluation result, determine the first value score corresponding to the second item information sequence; Update the first generative model based on the first value score.
[0170] In one embodiment, the second training unit 702 determines a first value score corresponding to the second item information sequence based on the third evaluation result, including: Based on the third evaluation result, the item information in the second item information sequence is reordered to obtain the third item information sequence; For each item information in the third item information sequence, the difference between the first position index value corresponding to the item information in the third item information sequence and the second index value corresponding to the item information in the second item information sequence is calculated to obtain the first difference value corresponding to the item information; A weighted summation is performed on multiple first differences corresponding to the third item information sequence to obtain a first value score; the first weight corresponding to the first difference in the first value score is related to the first position index value corresponding to the first difference.
[0171] In one embodiment, the third evaluation result includes a first score sequence consisting of a plurality of second value scores; The second training unit 702 reorders the item information in the second item information sequence based on the third evaluation result to obtain a third item information sequence, including: Based on the second value score corresponding to each item in the second item information sequence in the first score sequence, the item information in the second item information sequence is sorted to obtain the third item information sequence; each second value score in the first score sequence is used to describe the value corresponding to an item in the second item information sequence.
[0172] In one embodiment, the second training unit 702, based on each of the plurality of first item information sets, uses a second evaluation model to perform multiple training iterations on the first generative model, and further includes: The second evaluation model is invoked to process the fourth item information sequence corresponding to the first item information set to obtain a fourth evaluation result; the fourth item information sequence represents a preset item information sequence for the first item information set, and each item information in the fourth item information sequence is used to identify an item; the fourth evaluation result represents the evaluation result output by the evaluation network in the second evaluation model. Based on the fourth evaluation result, determine the third value score corresponding to the fourth item information sequence; Correspondingly, the second training unit 702 updates the first generative model based on the first value score, including: The first generative model is updated based on the second difference between the third value score and the first value score.
[0173] It should be noted that a detailed explanation of the steps performed by each unit can be found in [reference needed]. Figure 2 The model training method provided in the corresponding embodiments will not be described in detail here.
[0174] Based on the foregoing embodiments, this application provides a recommendation device that can be applied to... Figure 3 In the recommended method provided in the corresponding embodiments, refer to Figure 8 As shown, the recommendation device 80 may include: a first calling unit 801 and a recommendation unit 803, wherein: Calling unit 801 is used to call the second generation model to process the second item information set to obtain the fifth item information sequence; the second item information set represents the item information set determined by recall processing and sorting processing; Recommendation unit 802 is used to recommend items to the user based on the fifth item information sequence; The second generative model is achieved through, for example, Figure 2 The model was trained using the model training method provided in the corresponding embodiment.
[0175] It should be noted that a detailed explanation of the steps performed by each unit can be found in [reference needed]. Figure 3 The recommended methods provided in the corresponding embodiments will not be repeated here.
[0176] Based on the foregoing embodiments, embodiments of this application provide a model training device that can be applied to... Figure 1 In the model training method provided in the corresponding embodiment, refer to Figure 9As shown, the model training device 90 may include: a first processor 901, a first memory 902, and a first communication bus 903, wherein: The first communication bus 903 is used to realize the communication connection between the first processor 901 and the first memory 902; The first processor 901 is used to execute the model training program in the first memory 902 to perform the following steps: First sample data is determined; the first sample data includes multiple first item information sequences; each item information in the first item information sequence is used to describe the item-side features and user-side features of an item; Based on the multiple first item information sequences, the first evaluation model is trained and iterated multiple times using the first loss function and the second loss function until the set first convergence condition is met, and the second evaluation model is obtained. The first evaluation model includes a first evaluation network and a first auxiliary network. The first evaluation network is used to generate intermediate features based on the input item information sequence and to perform point-level value evaluation on the item information sequence based on the intermediate features. The first auxiliary network is used to perform sequence-level value evaluation on the item information sequence based on the intermediate features generated by the first evaluation network. The first loss function is used to measure point-level loss with respect to value. The second loss function is used to measure sequence-level loss with respect to value.
[0177] In one embodiment, the first processor 901 is used to execute the model training program in the first memory 902 based on each of the plurality of first item information sequences, and to perform multiple training iterations on the first evaluation model using a first loss function and a second loss function to achieve the following steps: The first evaluation network is invoked to process the first item information sequence to obtain a first evaluation result and a first intermediate feature; and the first auxiliary network is invoked to process the first intermediate feature to obtain a second evaluation result. The first evaluation result is processed based on a first loss function to obtain a first loss value, and the first evaluation network is updated based on the first loss value; and the second evaluation result is processed based on a second loss function to obtain a second loss value, and the first auxiliary network and the first evaluation network are updated sequentially based on the second loss value.
[0178] In one embodiment, the first processor 901 is used to execute the model training program in the first memory 902 to process the second evaluation result based on the second loss function to obtain a second loss value, and to update the first auxiliary network and the first evaluation network sequentially based on the second loss value, so as to achieve the following steps: In the case of the first training iteration in the multiple training iterations of the first evaluation model, the second evaluation result is processed based on the second loss function to obtain the second loss value, and the first auxiliary network and the first evaluation network are updated sequentially based on the second loss value. The first processor 901 is used to execute the model training program in the first memory 902 to process the first evaluation result based on the first loss function, obtain a first loss value, and update the first evaluation network based on the first loss value, so as to achieve the following steps: In the case of the second training iteration in the multiple training iterations of the first evaluation model, the first evaluation result is processed based on the first loss function to obtain the first loss value, and the first evaluation network is updated based on the first loss value. Wherein, the first training iteration represents one training iteration performed every n training iterations; the second training iteration represents other training iterations in the multiple training iterations besides the first training iteration; and n is a positive integer.
[0179] It should be noted that a detailed description of the steps performed by the first processor can be found in [reference needed]. Figure 1 The model training method provided in the corresponding embodiments will not be described in detail here.
[0180] Based on the foregoing embodiments, embodiments of this application provide a model training device that can be applied to... Figure 2 In the model training method provided in the corresponding embodiment, refer to Figure 10 As shown, the model training device 100 may include: a second processor 1001, a second memory 1002, and a second communication bus 1003, wherein: The second communication bus 1003 is used to realize the communication connection between the second processor 1001 and the second memory 1002; The second processor 1001 is used to execute the model training program in the second memory 1002 to perform the following steps: Determine the second sample data; the second sample data includes multiple sets of first item information; each item information in the first item information set is used to describe the item-side features and user-side features of an item; Based on the multiple sets of first item information, the first generation model is trained and iterated multiple times using the second evaluation model until a set second convergence condition is met, thus obtaining the second generation model; the second evaluation model is obtained through, as follows: Figure 1 It was obtained by training the corresponding model using the corresponding model training method.
[0181] In one embodiment, the second processor 1001 is used to execute the model training program in the second memory 1002, based on each of the plurality of first item information sets, and using a second evaluation model to perform multiple training iterations on the first generative model to achieve the following steps: The first generation model is invoked to process the first item information set to obtain a second item information sequence; each item information in the second item information sequence is used to identify an item. The second evaluation model is invoked to process the second item information sequence to obtain a third evaluation result; the third evaluation result represents the evaluation result output by the evaluation network in the second evaluation model. Based on the third evaluation result, determine the first value score corresponding to the second item information sequence; Update the first generative model based on the first value score.
[0182] In one embodiment, the second processor 1001 is used to execute the model training program in the second memory 1002 to determine the first value score corresponding to the second item information sequence based on the third evaluation result, in order to achieve the following steps: Based on the third evaluation result, the item information in the second item information sequence is reordered to obtain the third item information sequence; For each item information in the third item information sequence, the difference between the first position index value corresponding to the item information in the third item information sequence and the second index value corresponding to the item information in the second item information sequence is calculated to obtain the first difference value corresponding to the item information; A weighted summation is performed on multiple first differences corresponding to the third item information sequence to obtain a first value score; the first weight corresponding to the first difference in the first value score is related to the first position index value corresponding to the first difference.
[0183] In one embodiment, the third evaluation result includes a first score sequence consisting of a plurality of second value scores; The second processor 1001 is used to execute the model training program in the second memory 1002 to reorder the item information in the second item information sequence based on the third evaluation result, so as to obtain a third item information sequence, in order to achieve the following steps: Based on the second value score corresponding to each item in the second item information sequence in the first score sequence, the item information in the second item information sequence is sorted to obtain the third item information sequence; each second value score in the first score sequence is used to describe the value corresponding to an item in the second item information sequence.
[0184] In one embodiment, the second processor 1001 is further configured to execute a model training program in the second memory 1002 to perform the following steps: The second evaluation model is invoked to process the fourth item information sequence corresponding to the first item information set to obtain a fourth evaluation result; the fourth item information sequence represents a preset item information sequence for the first item information set, and each item information in the fourth item information sequence is used to identify an item; the fourth evaluation result represents the evaluation result output by the evaluation network in the second evaluation model. Based on the fourth evaluation result, determine the third value score corresponding to the fourth item information sequence; The second processor 1001 is used to execute the model training program in the second memory 1002, updating the first generative model based on the first value score, to achieve the following steps: The first generative model is updated based on the second difference between the third value score and the first value score.
[0185] It should be noted that a detailed description of the steps performed by the second processor can be found in [reference needed]. Figure 2 The model training method provided in the corresponding embodiments will not be described in detail here.
[0186] Based on the foregoing embodiments, embodiments of this application provide a recommended device that can be applied to... Figure 3 In the recommended method provided in the corresponding embodiments, refer to Figure 11 As shown, the recommended device 110 may include: a third processor 1101, a third memory 1102, and a third communication bus 1103, wherein: The third communication bus 1103 is used to realize the communication connection between the third processor 1101 and the third memory 1102; The third processor 1101 is used to execute the model training program in the third memory 1102 to perform the following steps: The second generation model is invoked to process the second item information set to obtain the fifth item information sequence; the second item information set represents the item information set determined through recall and sorting processes. Based on the fifth item information sequence, item recommendations are made to the user; The second generative model is achieved through, for example, Figure 2 The model was trained using the model training method provided in the corresponding embodiment.
[0187] It should be noted that a detailed explanation of the steps performed by the third processor can be found in [reference needed]. Figure 3The recommended methods provided in the corresponding embodiments will not be repeated here.
[0188] Based on the foregoing embodiments, embodiments of this application provide a computer-readable storage medium storing one or more programs, which can be executed by one or more processors to implement... Figure 1 The corresponding implementation provides the model training method, or Figure 2 The corresponding implementation provides the model training method, or Figure 3 The steps of the recommended method provided in the corresponding embodiment.
[0189] Based on the foregoing embodiments, embodiments of this application provide a computer program product, which includes a computer program that, when executed by a processor, implements... Figure 1 The corresponding implementation provides the model training method, or Figure 2 The corresponding implementation provides the model training method, or Figure 3 The steps of the recommended method provided in the corresponding embodiment.
[0190] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.
[0191] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0192] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0193] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0194] It should be noted that terms such as "first" and "second" are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0195] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Additionally, the term "one or more" in this document refers to any combination of at least two of any one or more elements from a set of A, B, and C. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set of A, B, and C. Furthermore, the term "one or more" in this document is an exemplary expression and can be replaced with any possible expression, such as one or more, at least one, or at least one of, etc.
[0196] Furthermore, the technical solutions described in the embodiments of this application can be combined arbitrarily without conflict.
[0197] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. A model training method, characterized in that, The method includes: First sample data is determined; the first sample data includes multiple first item information sequences; each item information in the first item information sequence is used to describe the item-side features and user-side features of an item; Based on the multiple first item information sequences, the first evaluation model is trained and iterated multiple times using the first loss function and the second loss function until the set first convergence condition is met, and the second evaluation model is obtained. The first evaluation model includes a first evaluation network and a first auxiliary network. The first evaluation network is used to generate intermediate features based on the input item information sequence and to perform point-level value evaluation on the item information sequence based on the intermediate features. The first auxiliary network is used to perform sequence-level value evaluation on the item information sequence based on the intermediate features generated by the first evaluation network. The first loss function is used to measure point-level loss with respect to value. The second loss function is used to measure sequence-level loss with respect to value.
2. The method according to claim 1, characterized in that, Based on each of the plurality of first item information sequences, the first evaluation model is trained and iterated multiple times using a first loss function and a second loss function, including: The first evaluation network is invoked to process the first item information sequence to obtain a first evaluation result and a first intermediate feature; and the first auxiliary network is invoked to process the first intermediate feature to obtain a second evaluation result. The first evaluation result is processed based on a first loss function to obtain a first loss value, and the first evaluation network is updated based on the first loss value; and the second evaluation result is processed based on a second loss function to obtain a second loss value, and the first auxiliary network and the first evaluation network are updated sequentially based on the second loss value.
3. The method according to claim 2, characterized in that, The step of processing the second evaluation result based on the second loss function to obtain a second loss value, and then sequentially updating the first auxiliary network and the first evaluation network based on the second loss value, includes: In the case of the first training iteration in the multiple training iterations of the first evaluation model, the second evaluation result is processed based on the second loss function to obtain the second loss value, and the first auxiliary network and the first evaluation network are updated sequentially based on the second loss value. The step of processing the first evaluation result based on the first loss function to obtain a first loss value, and updating the first evaluation network based on the first loss value, includes: In the case of the second training iteration in the multiple training iterations of the first evaluation model, the first evaluation result is processed based on the first loss function to obtain the first loss value, and the first evaluation network is updated based on the first loss value. Wherein, the first training iteration represents one training iteration performed every n training iterations; the second training iteration represents other training iterations in the multiple training iterations besides the first training iteration; and n is a positive integer.
4. A model training method, characterized in that, The method includes: Determine the second sample data; the second sample data includes multiple sets of first item information; each item information in the first item information set is used to describe the item-side features and user-side features of an item; Based on the multiple sets of first item information, the first generation model is trained and iterated multiple times using the second evaluation model until the set second convergence condition is reached, thereby obtaining the second generation model; the second evaluation model is obtained by training the model training method as described in any one of claims 1 to 3.
5. The method according to claim 4, characterized in that, Based on each of the plurality of first item information sets, the first generation model is trained and iterated multiple times using the second evaluation model, including: The first generation model is invoked to process the first item information set to obtain a second item information sequence; each item information in the second item information sequence is used to identify an item. The second evaluation model is invoked to process the second item information sequence to obtain a third evaluation result; the third evaluation result represents the evaluation result output by the evaluation network in the second evaluation model. Based on the third evaluation result, determine the first value score corresponding to the second item information sequence; Update the first generative model based on the first value score.
6. The method according to claim 5, characterized in that, The determination of the first value score corresponding to the second item information sequence based on the third evaluation result includes: Based on the third evaluation result, the item information in the second item information sequence is reordered to obtain the third item information sequence; For each item information in the third item information sequence, the difference between the first position index value corresponding to the item information in the third item information sequence and the second index value corresponding to the item information in the second item information sequence is calculated to obtain the first difference value corresponding to the item information; A weighted summation is performed on multiple first differences corresponding to the third item information sequence to obtain a first value score; the first weight of the first difference in the first value score is related to the first position index value corresponding to the first difference.
7. The method according to claim 6, characterized in that, The third evaluation result includes a first score sequence consisting of multiple second value scores; The step of reordering the item information in the second item information sequence based on the third evaluation result to obtain the third item information sequence includes: Based on the second value score corresponding to each item in the second item information sequence in the first score sequence, the item information in the second item information sequence is sorted to obtain the third item information sequence; each second value score in the first score sequence is used to describe the value corresponding to an item in the second item information sequence.
8. The method according to claim 5, characterized in that, Based on each of the plurality of first item information sets, and using a second evaluation model, the first generation model is trained and iterated multiple times, further comprising: The second evaluation model is invoked to process the fourth item information sequence corresponding to the first item information set to obtain a fourth evaluation result; the fourth item information sequence represents a preset item information sequence for the first item information set, and each item information in the fourth item information sequence is used to identify an item; the fourth evaluation result represents the evaluation result output by the evaluation network in the second evaluation model. Based on the fourth evaluation result, determine the third value score corresponding to the fourth item information sequence; Correspondingly, updating the first generation model based on the first value score includes: The first generative model is updated based on the second difference between the third value score and the first value score.
9. A recommended method, characterized in that, The method includes: The second generation model is invoked to process the second item information set to obtain the fifth item information sequence; the second item information set represents the item information set determined through recall and sorting processes. Based on the fifth item information sequence, item recommendations are made to the user; The second generative model is obtained by training using the model training method described in any one of claims 4 to 8.
10. A model training device, characterized in that, The device includes: The first determining unit is used to determine the first sample data; the first sample data includes multiple first item information sequences; each item information in the first item information sequence is used to describe the item-side features and user-side features of an item. The first training unit is used to train the first evaluation model multiple times based on the multiple first item information sequences, using a first loss function and a second loss function, until a set first convergence condition is reached, and thus obtain the second evaluation model. The first evaluation model includes a first evaluation network and a first auxiliary network. The first evaluation network is used to generate intermediate features based on the input item information sequence and to perform point-level value evaluation on the item information sequence based on the intermediate features. The first auxiliary network is used to perform sequence-level value evaluation on the item information sequence based on the intermediate features generated by the first evaluation network. The first loss function is used to measure point-level loss with respect to value. The second loss function is used to measure sequence-level loss with respect to value.
11. A model training device, characterized in that, The device includes: The second determining unit is used to determine the second sample data; the second sample data includes multiple sets of first item information; each item information in the first item information set is used to describe the item-side features and user-side features of an item; The second training unit is used to perform multiple training iterations on the first generation model based on the plurality of first item information sets and using the second evaluation model until a set second convergence condition is reached to obtain the second generation model; the second evaluation model is obtained by training the model training method as described in any one of claims 1 to 3.
12. A recommendation device, characterized in that, The device includes: The calling unit is used to call the second generation model to process the second item information set to obtain the fifth item information sequence; the second item information set represents the item information set determined through recall processing and sorting processing. The recommendation unit is used to recommend items to the user based on the fifth item information sequence; The second generative model is obtained by training using the model training method described in any one of claims 4 to 8.
13. A model training device, characterized in that, The device includes: a first processor, a first memory, and a first communication bus; The first communication bus is used to establish a communication connection between the first processor and the first memory; The first processor is configured to execute the model training program in the first memory to implement the steps of the model training method as described in any one of claims 1 to 3.
14. A model training device, characterized in that, The device includes: a second processor, a second memory, and a second communication bus; The second communication bus is used to establish a communication connection between the second processor and the second memory; The second processor is used to execute the model training program in the second memory to implement the steps of the model training method as described in any one of claims 4 to 8.
15. A recommended device, characterized in that, The device includes: a third processor, a third memory, and a third communication bus; The third communication bus is used to realize the communication connection between the third processor and the third memory; The third processor is used to execute the recommendation program in the third memory to implement the steps of the recommendation method as described in claim 9.
16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs that can be executed by one or more processors to implement the steps of the model training method as described in any one of claims 1 to 3, or the model training method as described in any one of claims 4 to 8, or the recommended method as described in claim 9.
17. A computer program product, the computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the model training method according to any one of claims 1 to 3, or the model training method according to any one of claims 4 to 8, or the recommended method according to claim 9.