Model training method and device, auction method and device, electronic equipment and medium

By preprocessing the user's historical behavior data and item data, and using the list recommendation model to train the auction model, the problem that the existing auction model fails to effectively consider the actual e-commerce scenarios is solved, and more effective auction results are achieved.

CN120070017APending Publication Date: 2025-05-30BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311634092.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-30
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing auction models tend to be more inclined to algorithmic theoretical allocation, and fail to effectively consider the actual e-commerce scenarios, resulting in the limitation of the effectiveness of the auction results.

Method used

By obtaining the historical behavior data and item data of multiple users, after preprocessing, the initial list recommendation model is trained to obtain the list recommendation model, and then it is used as a teacher model, and a loss function is built in combination with the initial auction model, and training is carried out to obtain the target auction model.

Benefits of technology

Ensure that the output of the auction model takes into account the user's interest in the item and improves the effectiveness and accuracy of the auction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070017A_ABST
    Figure CN120070017A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a model training method and device, an auction method and device, electronic equipment and a computer storage medium, and the method comprises the steps: obtaining behavior data and article data of a plurality of users in a first historical time period, and the behavior data comprises data generated when the users carry out behavior operation on articles; preprocessing the behavior data and the article data to obtain a preprocessed data set; training an initial list recommendation model according to the preprocessed data set to obtain a list recommendation model; constructing a loss function by taking the list recommendation model as a teacher model and taking the initial auction model as a student model; and training the initial auction model according to the loss function to obtain a target auction model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of Internet technologies, and in particular, to a model training method, an auction method, an apparatus, an electronic device, and a computer storage medium. Background Art

[0002] Currently, online auctions have become an important field. The auction process involves auctioneers, bidders, and auction items, and its goal is to find a mapping from bids to allocation and payment; from the perspective of the auctioneer, the goal of auction mechanism design is to find an optimal solution to optimize the overall revenue. However, existing auction models tend to focus on algorithmic theoretical allocation and do not consider the actual scenario; in order to utilize them in practice and better fit the e-commerce scenario, an auction model in the e-commerce scenario needs to be provided. Summary of the Invention

[0003] The present application provides a model training method, an auction method, an apparatus, an electronic device, and a computer storage medium.

[0004] The technical solution of the present application is implemented as follows:

[0005] An embodiment of the present application provides a model training method, and the method includes:

[0006] Obtain the behavior data and item data of multiple users in the first historical time period, where the behavior data includes the data generated by the users' behavior operations on the items;

[0007] Preprocess the behavior data and item data to obtain a preprocessed data set; train an initial list recommendation model according to the preprocessed data set to obtain a list recommendation model;

[0008] Use the list recommendation model as a teacher model and an initial auction model as a student model to construct a loss function; train the initial auction model according to the loss function to obtain a target auction model.

[0009] An embodiment of the present application provides an auction method, and the method includes:

[0010] Obtain the behavior data of multiple users in the current time period and a set of candidate items; the set of candidate items includes at least two candidate items;

[0011] Input the behavior data of the multiple users in the current time period and the set of candidate items into the target auction model to obtain a predicted item list; the predicted item list includes the sorting results of the at least two candidate items;

[0012] Based on the sorting result, conduct an auction for the display positions corresponding to the at least two candidate items; wherein, the target auction model is obtained according to the model training method provided by the foregoing one or more technical solutions.

[0013] An embodiment of the present application further provides a model training device, which includes a first acquisition module, a first training module, and a second training module, wherein,

[0014] The first acquisition module is used to acquire the behavior data and item data of multiple users in the first historical time period, and the behavior data includes the data generated by the users' behavior operations on the items;

[0015] The first training module is used to preprocess the behavior data and item data to obtain a preprocessed data set; and train an initial list recommendation model according to the preprocessed data set to obtain a list recommendation model;

[0016] The second training module is used to construct a loss function with the list recommendation model as the teacher model and the initial auction model as the student model; and train the initial auction model according to the loss function to obtain a target auction model.

[0017] An embodiment of the present application provides an auction device, which includes a second acquisition module, a prediction module, and a position auction module, wherein,

[0018] The second acquisition module is used to acquire the behavior data of multiple users in the current time period and a set of candidate items; the set of candidate items includes at least two candidate items;

[0019] The prediction module is used to input the behavior data of the multiple users in the current time period and the set of candidate items into the target auction model to obtain a predicted item list; the predicted item list includes the sorting results of the at least two candidate items;

[0020] The position auction module is used to conduct an auction for the display positions corresponding to the at least two candidate items according to the sorting result; wherein, the target auction model is obtained according to the model training method provided by the foregoing one or more technical solutions.

[0021] An embodiment of the present application provides an electronic device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the model training method or the auction method provided by the foregoing one or more technical solutions.

[0022] An embodiment of the present application provides a computer storage medium storing a computer program, which, when executed, can implement the model training method or auction method provided by one or more of the foregoing technical solutions.

[0023] An embodiment of the present application provides a model training method, an auction method, a device, an electronic device, and a computer storage medium. The method includes: obtaining behavior data and item data of multiple users in a first historical time period, where the behavior data includes data generated by the users' behavior operations on the items; preprocessing the behavior data and item data to obtain a preprocessed data set; training an initial list recommendation model according to the preprocessed data set to obtain a list recommendation model; using the list recommendation model as a teacher model and an initial auction model as a student model to construct a loss function; and training the initial auction model according to the loss function to obtain a target auction model.

[0024] It can be seen that, by preprocessing the obtained historical behavior data and item data of the users, the embodiment of the present application can ensure the accuracy and reliability of the data for subsequent model training. After obtaining the preprocessed data set, the initial list recommendation model is first trained. Since this model can take into account the users' behavior operations on the items during the training process, it can help identify the users' interest preferences for the items. Therefore, using the list recommendation model as a teacher model to train the initial auction model, which is relatively small in scale and easy to deploy, can ensure that the output of the initial auction model also takes into account the users' interest preferences for the items. In this way, when using the auction model to auction the display positions corresponding to each item, such as advertising positions, in the subsequent process, the effectiveness of the auction results can be ensured. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 is a flowchart of a model training method in an embodiment of the present application;

[0026] Figure 2 is a structural diagram of an initial list recommendation model in an embodiment of the present application;

[0027] Figure 3 is a network structure diagram of an initial auction model in an embodiment of the present application;

[0028] Figure 4 is a flowchart of an auction method in an embodiment of the present application;

[0029] Figure 5 is a structural diagram of a model training framework in an embodiment of the present application;

[0030] Figure 6Schematic diagram of the composition structure of a model training device according to an embodiment of the present application;

[0031] Figure 7 Schematic diagram of the composition structure of an auction device according to an embodiment of the present application;

[0032] Figure 8 Schematic diagram of the structure of an electronic device according to an embodiment of the present application. Detailed implementation manners

[0033] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the embodiments provided herein are only used to explain the present application and are not used to limit the present application. In addition, the embodiments provided below are partial embodiments for implementing the present application, rather than all embodiments for implementing the present application. Without conflict, the technical solutions described in the embodiments of the present application can be implemented in any combined manner.

[0034] It should be noted that, in the embodiments of the present application, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a method or device including a series of elements not only includes the elements specifically recited, but also includes other elements not expressly listed, or also includes elements inherent to the implementation of the method or device. Without further limitation, the element defined by the statement "including one..." does not exclude the existence of other relevant elements in the method or device including the element (such as steps in the method or units in the device, for example, the unit may be a part of a circuit, a part of a processor, a part of a program or software, etc.).

[0035] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, I and / or J can represent: I exists alone, I and J exist simultaneously, and J exists alone. In addition, the term "at least one" in this article means any one or any combination of at least two of a plurality. For example, including at least one of I, J, and R can represent including any one or more elements selected from the set composed of I, J, and R.

[0036] For example, the model training method provided in the embodiments of the present application includes a series of steps, but the model training method provided in the embodiments of the present application is not limited to the recited steps. Similarly, the model training device provided in the embodiments of the present application includes a series of modules, but the model training device provided in the embodiments of the present application is not limited to including the expressly recited modules, and may also include modules required for obtaining relevant timing data or processing based on timing data.

[0037] In some embodiments of the present application, the model training method can be implemented by a processor in the model training device. The above-mentioned processor can be at least one of an Application Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), a Digital Signal Processing Device (DSPD), a Programmable Logic Device (PLD), a Field Programmable Gate Array (FPGA), a Central Processing Unit (CPU), a controller, a microcontroller, and a microprocessor.

[0038] It should be noted that in the technical solution of the present application, the collection, use, storage, sharing, transfer, and other processing of the user's personal information involved all comply with the provisions of relevant laws and regulations, and the user needs to be informed and obtain the consent or authorization of the user. When applicable, technical processing such as de-identification and / or anonymization and / or encryption is performed on the user's personal information.

[0039] Figure 1 It is a schematic flowchart of a model training method in an embodiment of the present application, as Figure 1 shown. The method includes the following steps:

[0040] Step 100: Obtain the behavior data and item data of multiple users in the first historical period.

[0041] In the embodiments of the present application, the model training method can be applied to the e-commerce scenario of multi-item auctions; here, the multi-item auction can be an auction for the display positions corresponding to each item, such as an advertisement slot.

[0042] Exemplarily, the behavior data includes the data generated by the user's behavior operations on the item. Here, the type of the behavior operation is not limited. For example, the behavior operation can include operations such as the user searching, browsing, purchasing, and collecting the item.

[0043] Exemplarily, the behavior data and item data of each user in the first historical period can be collected from the e-commerce platform and recorded in the data system related to the user. In this way, when it is necessary to obtain the behavior data and item data of multiple users in the first historical period, it can be directly obtained from the e-commerce platform or from the data system related to the user.

[0044] It should be noted that the first historical time period can be a time period with a set duration before the current moment. Here, the duration of the historical time period is not specifically limited. For example, it can be the past six months, or the past three months, etc.

[0045] Exemplarily, the item data may include attribute data such as the category, name, description, price, etc. of each item that the user has operated on. The behavior data may also include user data, where the user data may include attribute data such as the user's age, gender, occupation, etc.

[0046] It can be understood that before building the model, by obtaining the behavior data and item data of multiple users in the first historical time period, it is convenient for subsequent mining or processing.

[0047] Step 101: Preprocess the behavior data and item data to obtain a preprocessed data set; train the initial list recommendation model according to the preprocessed data set to obtain a list recommendation model.

[0048] In the embodiments of the present application, after obtaining the behavior data and item data of multiple users in the first historical time period, these data can be preprocessed first to obtain preprocessed data; then, the constructed initial list recommendation model is trained according to the preprocessed data set to obtain a list recommendation model.

[0049] Exemplarily, the preprocessing at least includes data cleaning and feature extraction; first, the obtained behavior data and item data can be subjected to data cleaning; among them, data cleaning can include removing duplicate data, abnormal data, and irrelevant data, etc., to ensure the accuracy and reliability of the data; then, feature engineering is performed on the cleaned data to extract useful features to obtain a data set; this data set can include features such as the user's age, gender, occupation, and interests, as well as features such as the category and price of the item. Here, the purpose of feature engineering is to convert the data into a vector form for subsequent use.

[0050] Exemplarily, in terms of feature extraction, two methods of text mining and user behavior analysis can be adopted. Text mining mainly processes text information such as item data and user data, and converts the text information into numerical features through the TF-IDF method for subsequent model training and prediction; user behavior analysis is to process the behavior data generated by the user's search, browsing, purchase, collection and other behavior operations, and extract the user's interest features to recommend items that meet the user's needs. Through data cleaning and feature extraction, a preprocessed data set can be obtained, providing a basis for subsequent modeling and recommendation.

[0051] In some embodiments, training an initial list recommendation model based on a preprocessed data set to obtain a list recommendation model may include: dividing the preprocessed data set to obtain a training set and a test set; training the initial list recommendation model using the training set to obtain a trained initial list recommendation model; and testing the trained initial prediction model using the test set to obtain a list recommendation model.

[0052] In one embodiment, the preprocessed data set may be divided into a training set and a test set according to a set ratio; here, the value of the set ratio can be determined according to the actual situation, and the embodiments of the present application do not make specific limitations thereon. For example, the set ratio may be 9:1 or 8:2, etc.

[0053] In another embodiment, the preprocessed data set may be divided into a training set and a test set by using the K-fold cross-validation method; here, the value of K is not limited. For example, the 10-fold cross-validation method may be adopted, where 9 parts are used to learn and train the constructed initial list recommendation model, and 1 part is used to test the constructed initial list recommendation model, and the experiment is repeated 10 times.

[0054] In the embodiments of the present application, after obtaining the training set and the test set, the initial list recommendation model is first trained using the training set. Exemplarily, the iterative training may be performed using the gradient descent method. After obtaining the trained initial list recommendation model, the trained initial prediction model is then tested using the test set to obtain a trained list recommendation model.

[0055] Exemplarily, testing the trained initial prediction model using the test set may obtain a test result. According to this test result, the prediction accuracy of the initial list recommendation model can be further determined; comparing the prediction accuracy with a set value, when the prediction accuracy is greater than or equal to the set value, at this time, a trained list recommendation model can be obtained; when the prediction accuracy is less than the set value, the model parameters of the trained initial prediction model are updated until the prediction accuracy of the model reaches the set value; it can be understood that after obtaining the trained initial prediction model, testing and validating the initial list recommendation model through the test set can reduce the risk of overfitting of the list recommendation model.

[0056] Here, the value of the set value can be set according to the actual situation, and the embodiments of the present application do not make specific limitations thereon. For example, it may be 90% or 95%, etc.

[0057] In some embodiments, referring to Figure 2, the initial list recommendation model may include an initial list generation network and an initial list evaluation network; training the initial list recommendation model using a training set to obtain a trained initial list recommendation model may include: processing the training set by the initial list generation network to obtain multiple recommendation lists; processing the multiple recommendation lists by the initial list evaluation network to obtain a target recommendation list that meets the set requirements; adjusting the parameters of the initial list generation network and the initial list evaluation network according to the target recommendation list and label data to obtain a trained initial list recommendation model.

[0058] Exemplarily, the training set may be input into the initial list generation network, and the initial list generation network generates an item distribution strategy based on the user characteristics in the training set and generates multiple recommendation lists according to the item distribution strategy; wherein, each recommendation list includes the sorting results of at least two items. As Figure 2 shown, the initial list generation network generates four recommendation lists, and each recommendation list includes the sorting results of four different items, and these four items are represented by different numbers (corresponding Figure 2 to 9, 8, 1, and 3 in).

[0059] Furthermore, input the multiple recommendation lists generated by the initial list generation network into the initial list evaluation network, and the initial list evaluation network can evaluate each recommendation list according to key factors such as the similarity between the items in each recommendation list and the user interest, and obtain a target recommendation list that meets the set requirements (corresponding Figure 2 to 8913 in); here, the set requirements can be correspondingly set according to key factors such as the similarity between the items and the user interest.

[0060] Exemplarily, the training set may include label data of each item, and the label data is used to indicate the actual sorting position of each item; after obtaining the target recommendation list, the parameters of the initial list generation network and the initial list evaluation network can be adjusted according to the target recommendation list and the label data to obtain a trained initial list recommendation model.

[0061] It should be noted that the initial list recommendation model is equivalent to a combined recommendation model, and this model adopts the listwise sorting method. Compared with the pointwise sorting method, the listwise sorting method can consider multiple items and their previous association relationships at the same time, find the optimal order between the items, and make the model have stronger generalization ability.

[0062] Step 102: Use the list recommendation model as the teacher model and the initial auction model as the student model to construct a loss function; train the initial auction model according to the loss function to obtain a target auction model.

[0063] Exemplarily, according to the above, it can be known that the output of the list recommendation model is a sorted result list of multiple items, which solves the problem of recommendation ranking of multiple items, while the output of the initial auction model is a single item, that is, this model does not have the combination ability; to solve the problem of incompatibility between the list recommendation model and the initial auction model, after obtaining the list recommendation model according to the above steps, the list recommendation model can be used as the teacher model, and the pre-constructed initial auction model can be used as the student model to construct the loss function; that is, the initial auction model is trained in the way of teacher-student distillation, so that the list recommendation model transfers its combination ability to the initial auction model. That is to say, the knowledge of the list recommendation model is transferred to the initial auction model through the way of knowledge distillation.

[0064] It can be understood that through the way of knowledge distillation, the performance of the initial auction model can be made close to the performance of the list recommendation model. In this way, the complexity of the model can be simplified to a certain extent, and the model processing efficiency can be improved.

[0065] Exemplarily, to improve the training effect of the initial auction model, after obtaining the list recommendation model, the list recommendation model can be optimized and adjusted first. After obtaining the target list recommendation model after optimization and adjustment, the target list recommendation model can be used as the teacher model, and the initial auction model can be used as the student model to construct the loss function. The following is an exemplary description of the optimization and adjustment process of the list recommendation model.

[0066] In some embodiments, the above method may further include: after obtaining the target recommendation list that meets the set requirements, the target recommendation list is displayed, and during the display process, the feedback data of each user on different items in the target recommendation list is obtained; the list recommendation model is optimized and adjusted according to the feedback data to obtain the target list recommendation model.

[0067] Exemplarily, after obtaining the above target item list, the target item list can be displayed in a set area; here, the set area refers to the area for displaying the target item list. For example, it can be the home page of an e-commerce platform or the item details page, etc.

[0068] It can be understood that in order to continuously optimize the item recommendation effect of the list recommendation model, the feedback data of different users on the target item list can be collected, such as the feedback data of users' clicks, purchases, collections, etc. on the items; after obtaining the feedback data of each user on different items in the target item list, the list recommendation model can be optimized and adjusted according to the feedback data of each user to obtain the target list recommendation model.

[0069] In some embodiments, optimizing and adjusting the list recommendation model according to the feedback data to obtain the target list recommendation model may include: constructing a first reward function of the list recommendation model according to the feedback data; and optimizing and adjusting the list recommendation model according to the first reward function to obtain the target list recommendation model.

[0070] Exemplarily, after obtaining the feedback data of each user on different items in the target item list, a first reward function of the list recommendation model may be constructed according to the feedback data, as shown in formula (1):

[0071]

[0072] Here, P[i,:] represents the feedback data corresponding to the i-th user, N represents the number of types of operation behaviors, l l represents the coefficient of each operation behavior, and f l 1 represents the feedback data of the i-th user's first behavior operation on the l-th item, and f l N represents the feedback data of the i-th user's N-th behavior operation on the l-th item.

[0073] Exemplarily, after obtaining the first reward function, the reward value of the list recommendation model may be calculated according to the reward function. Furthermore, the list recommendation model may be supervised and optimized according to the reward value; by continuously iterating the above process, the target list recommendation model may be obtained; it can be seen that through the above training process, the quality of item recommendation can be gradually improved to ensure the recommendation effect of the model.

[0074] Exemplarily, after obtaining the target list recommendation model, a loss function between the target list recommendation model and the initial auction model may be further constructed, as shown in formula (2):

[0075] L Teacher-Student =L(I model , I comd ) (2)

[0076] where, I model represents the output of the initial auction model, and I comd represents the output of the target list recommendation model.

[0077] Exemplarily, after obtaining the loss function L Teacher-Student a usual model parameter update algorithm, such as the BackPropagation (BP) algorithm, may be used to adjust the model parameters of the initial auction model to obtain the trained target auction model.

[0078] Here, the type of the loss function is not specifically limited. For example, it can be a cross - entropy loss function or other types of loss functions, etc.

[0079] In the embodiments of the present application, the network structure of the initial auction model is not limited. For example, the initial auction model can be a neural network model including multiple input neurons and output neurons, as Figure 3 shown.

[0080] It can be understood that to reduce the risk of overfitting of the model and improve the generalization ability of the model, after obtaining the loss function, a regularization term can be added to the loss function as the updated loss function, and the initial auction model can be trained based on the updated loss function to obtain the trained target auction model. Since the regularization term can give weight penalties, making the weights of some features tend to zero or even equal to zero, thus reducing the impact of the weight error generated during the training process on the model accuracy and improving the model accuracy.

[0081] In some embodiments, training the initial auction model according to the loss function to obtain the target auction model may include: training the initial auction model according to the loss function to obtain the trained initial auction model; obtaining the behavior data and item data of multiple users in the second historical period; constructing the second reward function of the initial auction model according to the behavior data and item data of multiple users in the second historical period; and optimizing and adjusting the trained initial auction model according to the second reward function to obtain the target auction model.

[0082] Here, the second historical period is the period before the first historical period; the duration of the second period can be set according to the actual situation, and the embodiments of the present application do not limit this; the duration of the second historical period can be the same as or different from that of the first historical period. For example, assuming that the current day is the T - th day, the period from the (T - 180)-th day to the (T - 90)-th day can be determined as the second historical period, and the period from the (T - 89)-th day to the (T - 30)-th day can be determined as the first historical period.

[0083] Exemplarily, after obtaining the behavior data and item data of multiple users in the second historical period, pre - processing can be performed first. The process of pre - processing has been described above and will not be elaborated here; then, the second reward function of the initial auction model can be constructed according to the pre - processed data. The second reward function is shown in formula (3):

[0084] L long-term =r 1 +γr 2 +γ 2 r 3 +... (3)

[0085] Here, r1 ,r 2 ,r 3 respectively represent the reward values for the user's behavioral operations on the item in different time periods, and γ represents the attenuation coefficient, whose value range is between 0 and 1.

[0086] It can be seen that in the embodiment of the present application, when training the initial auction model using the behavioral data in the first historical time period, the behavioral data in the second historical time period will also be used for optimization and adjustment. That is, the training process of the model uses the user's long-term behavioral data. In this way, the training effect of the model can be improved.

[0087] Exemplarily, in practical applications, the training of the initial auction model often faces the problems of sparse rewards and unstable training. To solve this problem, the cross-entropy loss function can be used to stabilize the training process.

[0088] The embodiment of the present application proposes a model training method, which includes: obtaining the behavioral data and item data of multiple users in the first historical time period, where the behavioral data includes the data generated by the user's behavioral operations on the item; preprocessing the behavioral data and item data to obtain a preprocessed data set; training the initial list recommendation model according to the preprocessed data set to obtain a list recommendation model; using the list recommendation model as the teacher model and the initial auction model as the student model to construct a loss function; training the initial auction model according to the loss function to obtain the target auction model. It can be seen that by preprocessing the obtained historical behavioral data and item data of the user, the accuracy and reliability of the data can be ensured for subsequent model training; after obtaining the preprocessed data set, the initial list recommendation model is first trained. Since this model can consider the user's behavioral operations on the item during the training process, that is, this model can help identify the user's interest preferences for the item, the list recommendation model is used as the teacher model to train the initial auction model with a relatively small scale and convenient deployment, which can ensure that the output of the initial auction model also considers the user's interest preferences for the item. In this way, when using the auction model to auction the display positions corresponding to each item, such as advertising positions, in the subsequent process, the effectiveness of the auction results can be ensured.

[0089] In order to better reflect the purpose of the present application, further examples are given based on the above embodiments of the present application.

[0090] Figure 4 is a schematic flowchart of an auction method according to an embodiment of the present application. As Figure 4 shown, the method includes the following steps:

[0091] Step 200: Obtain the behavioral data of multiple users in the current time period and the candidate item set.

[0092] Here, the set of candidate items includes at least two candidate items; the current time period, also known as the prediction time period, represents a time period of a set duration including the current time, and the current time period is after the above-mentioned first historical time period; for example, assuming the current day is the Tth day, the time period from the (T - 89)th day to the (T - 30)th day can be determined as the first historical time period, and the time period from the (T - 29)th day to the Tth day can be determined as the current time period.

[0093] Step 201: Input the behavior data of multiple users in the current time period and the set of candidate items into the target auction model to obtain a predicted item list.

[0094] In the embodiments of the present application, after obtaining the behavior data of multiple users in the current time period and the set of candidate items, preprocessing can be performed first. The implementation process of the preprocessing has been described above and will not be elaborated here; then the preprocessed data is input into the target auction model to obtain a predicted item list; here, the predicted item list includes the sorting results of at least two candidate items. The target auction model is obtained according to the above-mentioned steps 100 to 102 and will not be elaborated here.

[0095] Step 202: Auction the display positions corresponding to at least two candidate items according to the sorting results.

[0096] Here, in the e-commerce scenario, the display position corresponding to a candidate item can be the advertisement position corresponding to the candidate item; it can be seen that after obtaining the target auction model in the embodiments of the present application, the sorting results of each item can be output according to the target auction model; since the determination of the sorting results takes into account the user feedback data and long-term behavior data, therefore, the sorting of each item can well reflect the user's interest preferences. Auctioning the advertisement positions of items based on the sorting results can improve the auction effect.

[0097] Figure 5 It is a schematic structural diagram of a model training framework in the embodiments of the present application. As Figure 5 shown, the input of the initial list recommendation model is a data set, and the output is a recommendation list (corresponding Figure 58913 in it, each digit corresponding to an item), after obtaining the trained list recommendation model, using the list recommendation model as the teacher model and the pre-constructed initial auction model as the student model to construct a loss function for training the initial auction model; that is, training the initial auction model in a teacher-student distillation manner so that the list recommendation model transfers its combination ability to the initial auction model. During the training process of the list recommendation model, feedback data of each user on different items in the recommendation list is collected, the reward value of the list recommendation model is determined according to the feedback data, and the list recommendation model is optimized and adjusted according to the reward value; during the training process of the initial auction model, the reward value of the initial auction model is determined according to the long-term behavior data of the user, and the initial auction model is optimized and adjusted according to the reward value. The specific training process has been described in the above embodiments and will not be elaborated here.

[0098] Figure 6 is a schematic structural diagram of a model training device according to an embodiment of the present application, as Figure 6 shown, the device includes: a first acquisition module 300, a first training module 301, and a second training module 302, where:

[0099] The first acquisition module 300 is configured to acquire the behavior data and item data of multiple users in the first historical time period, and the behavior data includes the data generated by the user's behavior operations on the items;

[0100] The first training module 301 is configured to preprocess the behavior data and item data to obtain a preprocessed data set; train an initial list recommendation model according to the preprocessed data set to obtain a list recommendation model;

[0101] The second training module 302 is configured to use the list recommendation model as the teacher model and the initial auction model as the student model to construct a loss function; train the initial auction model according to the loss function to obtain a target auction model.

[0102] In some embodiments, the first training module 301 is further configured to:

[0103] Divide the preprocessed data set to obtain a training set and a test set;

[0104] Train the initial list recommendation model using the training set to obtain a trained initial list recommendation model;

[0105] Test the trained initial prediction model using the test set to obtain the list recommendation model.

[0106] In some embodiments, the initial list recommendation model includes an initial list generation network and an initial list evaluation network; the first training module 301 is further configured to:

[0107] Process the training set according to the initial list generation network to obtain a plurality of recommendation lists; each recommendation list includes the sorting results of at least two items, and the training set includes the label data of each item, and the label data is used to indicate the actual sorting position of each item;

[0108] Process the plurality of recommendation lists according to the initial list evaluation network to obtain a target recommendation list that meets the set requirements;

[0109] Adjust the parameters of the initial list generation network and the initial list evaluation network according to the target recommendation list and the label data to obtain the trained initial list recommendation model.

[0110] In some embodiments, the first training module 301 is further configured to:

[0111] After obtaining the target recommendation list that meets the set requirements, display the target recommendation list, and during the display process, obtain the feedback data of each user on different items in the target recommendation list;

[0112] Optimize and adjust the list recommendation model according to the feedback data to obtain the target list recommendation model;

[0113] The second training module 302 is further configured to:

[0114] Use the target list recommendation model as the teacher model and the initial auction model as the student model to construct a loss function.

[0115] In some embodiments, the first training module 301 is further configured to:

[0116] Construct a first reward function of the list recommendation model according to the feedback data;

[0117] Optimize and adjust the list recommendation model according to the first reward function to obtain the target list recommendation model.

[0118] In some embodiments, the second training module 302 is further configured to:

[0119] Train the initial auction model according to the loss function to obtain the trained initial auction model;

[0120] Obtain the behavior data and item data of the multiple users in the second historical period; the second historical period is the period before the first historical period;

[0121] Construct a second reward function for the initial auction model according to the behavior data and item data of the multiple users in the second historical period;

[0122] Optimize and adjust the trained initial auction model according to the second reward function to obtain the target auction model.

[0123] Figure 7 It is a schematic structural diagram of a composition of an auction device according to an embodiment of the present application, as Figure 7 shown. The device includes: a second acquisition module 400, a prediction module 401, and a position auction module 402, where:

[0124] The second acquisition module 400 is configured to acquire the behavior data of multiple users in the current period and a set of candidate items; the set of candidate items includes at least two candidate items;

[0125] The prediction module 401 is configured to input the behavior data of the multiple users in the current period and the set of candidate items into the target auction model to obtain a predicted item list; the predicted item list includes the sorting results of the at least two candidate items;

[0126] The position auction module 402 is configured to auction the display positions corresponding to the at least two candidate items according to the sorting results; wherein, the target auction model is obtained according to the model training method provided by the foregoing one or more technical solutions.

[0127] In practical applications, the foregoing first acquisition module 300, first training module 301, second training module 302, second acquisition module 400, prediction module 401, and position auction module 402 can all be implemented by a processor located in an electronic device, and the processor can be at least one of ASIC, DSP, DSPD, PLD, FPGA, CPU, controller, microcontroller, and microprocessor.

[0128] In addition, each functional module in this embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional module.

[0129] When an integrated unit is implemented in the form of a software functional module and is not sold or used as an independent article, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, in essence, or the part that contributes to the related technology, or all or part of this technical solution, can be embodied in the form of a software article. This computer software article is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the method of this embodiment. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0130] Specifically, the computer program instructions corresponding to a model training method or an auction method in this embodiment can be stored on storage media such as optical discs, hard disks, and USB flash drives. When the computer program instructions corresponding to a model training method or an auction method in the storage medium are read or executed by an electronic device, any one of the model training methods or auction methods of the foregoing embodiments is implemented.

[0131] Based on the same technical concept as the foregoing embodiments, refer to Figure 8 , which shows the electronic device 500 provided by this application, and may include: a memory 501 and a processor 502; wherein,

[0132] The memory 501 is used to store computer programs and data;

[0133] The processor 502 is used to execute the computer program stored in the memory to implement any one of the model training methods or auction methods of the foregoing embodiments.

[0134] In practical applications, the aforementioned memory 501 can be a volatile memory, such as RAM; or a non-volatile memory, such as ROM, flash memory, hard disk drive (HDD), or solid-state drive (SSD); or a combination of the aforementioned types of memories, and provide instructions and data to the processor 502.

[0135] The above-mentioned processor 502 may be at least one of an ASIC, a DSP, a DSPD, a PLD, an FPGA, a CPU, a controller, a microcontroller, and a microprocessor. It can be understood that the electronic device for implementing the functions of the above-mentioned processor may also be other, and the embodiments of the present application do not make specific limitations.

[0136] In some embodiments, the functions or modules included in the device provided by the embodiments of the present application may be used to execute the methods described in the above method embodiments. The specific implementation may refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0137] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. The same or similar parts can be referred to each other. For the sake of brevity, they will not be repeated here.

[0138] The methods disclosed in the method embodiments provided by the present application can be arbitrarily combined without conflict to obtain new method embodiments.

[0139] The features disclosed in the article embodiments provided by the present application can be arbitrarily combined without conflict to obtain new article embodiments.

[0140] The features disclosed in the method or device embodiments provided by the present application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0141] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program articles. Therefore, the present application can adopt the form of hardware embodiments, software embodiments, or embodiments combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program article implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program code.

[0142] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program articles according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable model training devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable model training devices generate means for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0143] These computer program instructions can also be loaded onto a computer or other programmable model training device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, thereby the instructions executed on the computer or other programmable device provide for implementing in the process Figure 1 one process or multiple processes and / or blocks Figure 1 steps for the functions specified in one block or multiple blocks.

[0144] The above are only the preferred embodiments of the present application and are not intended to limit the protection scope of the present application.

Claims

1. A model training method, characterized in that, the method includes: Obtain the behavior data and item data of multiple users in the first historical time period, where the behavior data includes the data generated by the users' behavior operations on the items; Preprocess the behavior data and item data to obtain a preprocessed data set; train an initial list recommendation model according to the preprocessed data set to obtain a list recommendation model; Use the list recommendation model as the teacher model and the initial auction model as the student model to construct a loss function; train the initial auction model according to the loss function to obtain a target auction model.

2. The method according to claim 1, characterized in that, the training the initial list recommendation model according to the preprocessed data set to obtain a list recommendation model includes: Divide the preprocessed data set to obtain a training set and a test set; Use the training set to train the initial list recommendation model to obtain a trained initial list recommendation model; Use the test set to test the trained initial prediction model to obtain the list recommendation model.

3. The method according to claim 2, characterized in that, the initial list recommendation model includes an initial list generation network and an initial list evaluation network; the training the initial list recommendation model according to the training set to obtain a trained initial list recommendation model includes: Process the training set according to the initial list generation network to obtain multiple recommendation lists; each recommendation list includes the sorting results of at least two items, and the training set includes the label data of each item, and the label data is used to indicate the actual sorting position of each item; Process the multiple recommendation lists according to the initial list evaluation network to obtain a target recommendation list that meets the set requirements; Adjust the parameters of the initial list generation network and the initial list evaluation network according to the target recommendation list and the label data to obtain the trained initial list recommendation model.

4. The method according to claim 3, characterized in that, the method further includes: After obtaining the target recommendation list that meets the set requirements, display the target recommendation list, and during the display process, obtain the feedback data of each user on different items in the target recommendation list; Optimize and adjust the list recommendation model according to the feedback data to obtain a target list recommendation model; Correspondingly, the using the list recommendation model as the teacher model and the initial auction model as the student model to construct a loss function includes: Use the target list recommendation model as the teacher model and the initial auction model as the student model to construct a loss function.

5. The method according to claim 4, characterized in that, the optimizing and adjusting the list recommendation model according to the feedback data to obtain a target list recommendation model includes: Construct a first reward function of the list recommendation model according to the feedback data; Optimize and adjust the list recommendation model according to the first reward function to obtain the target list recommendation model.

6. The method according to claim 1, wherein, the training of the initial auction model according to the loss function to obtain the target auction model includes: training the initial auction model according to the loss function to obtain the trained initial auction model; acquire the behavior data and item data of the multiple users in the second historical period; the second historical period is the period before the first historical period; construct a second reward function of the initial auction model according to the behavior data and item data of the multiple users in the second historical period; optimize and adjust the trained initial auction model according to the second reward function to obtain the target auction model.

7. An auction method, wherein, the method includes: acquire the behavior data of multiple users in the current period and a set of candidate items; the set of candidate items includes at least two candidate items; input the behavior data of the multiple users in the current period and the set of candidate items into the target auction model to obtain a predicted item list; the predicted item list includes the sorting results of the at least two candidate items; auction the display positions corresponding to the at least two candidate items according to the sorting results; wherein, the target auction model is obtained by the method according to any one of claims 1 to 7.

8. A model training device, wherein, the device includes: a first acquisition module, configured to acquire the behavior data and item data of multiple users in the first historical period, where the behavior data includes the data generated by the user's behavior operations on the item; a first training module, configured to preprocess the behavior data and item data to obtain a preprocessed data set; train an initial list recommendation model according to the preprocessed data set to obtain a list recommendation model; a second training module, configured to use the list recommendation model as a teacher model and an initial auction model as a student model to construct a loss function; train the initial auction model according to the loss function to obtain a target auction model.

9. An electronic device, wherein, the device includes a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the method according to any one of claims 1 to 7.

10. A computer storage medium, on which a computer program is stored, wherein, when the computer program is executed by a processor, it implements the method according to any one of claims 1 to 7.