A recommendation system training method, a recommendation method, a device, an electronic device, and a storage medium

By determining the initial account status and behavioral data in the recommender system, and using a data simulator for sparse-to-dense transformation and training, the problem of long-term inaccuracy of metrics in existing recommender systems is solved, and a recommender system training that meets the metric requirements is achieved.

CN115146152BActive Publication Date: 2026-01-02BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210640273.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-07
Publication Date
2026-01-02
Estimated Expiration
2042-06-07

AI Technical Summary

Technical Problem

Existing recommendation methods are inaccurate in long-term metrics such as retention rate, causing recommendation systems trained based on long-term metrics to fail to meet requirements.

Method used

By determining the initial account status data and behavioral data, and simulating indicator information within a preset time period using a data simulator, a sparse-to-dense transformation and training are performed to obtain the target recommendation system that meets the preset indicator conditions.

Benefits of technology

This method decomposes sparse long-term metrics into dense long-term metrics, ensuring that the recommendation system meets the requirements during long-term metric training and improving the overall value of the recommendation system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115146152B_ABST
    Figure CN115146152B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a recommendation system training method, a recommendation method, a device, an electronic device and a storage medium, comprising: determining an initial recommendation object corresponding to a recommendation system based on initial account state data and initial account behavior data; determining preset index information of the recommendation system on a preset time period, transition account state data and transition account behavior data according to the initial account state data, the initial account behavior data and the initial recommendation object; performing sparse-to-dense conversion on the preset index information on the preset time period to obtain preset index information of the recommendation system at each time step in the preset time period; and training the recommendation system based on the preset index information at each time step, the transition account state data and the transition account behavior data to obtain a target recommendation system. The present application can decompose sparse long-term indicators into dense long-term indicators that can be improved, and thus the recommendation system trained based on the long-term indicators meets the index requirements.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of Internet, and particularly relates to a recommendation system training method and device, a recommendation method and device, an electronic device and a storage medium. BACKGROUND

[0002] With the rapid development of the current mobile Internet, in the recommendation system, recall, sorting and strategy have become standard paradigms. Personalized recommendation system services need to filter and score a large number of items in the candidate set, recommend high-quality objects to users, thereby improving user satisfaction and retention rate, and thus realizing the comprehensive value and sustainable development of the recommendation system.

[0003] The existing recommendation method usually calculates the corresponding score by means of the instant feedback prediction value of various users, and sorts the objects based on the score. This essentially improves the short-term indicators of users, such as click rate and viewing rate, and indirectly affects the long-term indicators such as retention rate, which makes the long-term indicators incorrect, and thus the recommendation system trained based on the long-term indicators is not required. SUMMARY

[0004] The present disclosure provides a recommendation system training method, a recommendation method, a device, an electronic device and a storage medium. The technical solution of the present disclosure is as follows:

[0005] According to a first aspect of an embodiment of the present disclosure, a recommendation system training method is provided, comprising:

[0006] determining initial account state data and initial account behavior data of the recommendation system;

[0007] determining an initial recommendation object corresponding to the recommendation system based on the initial account state data and the initial account behavior data;

[0008] determining preset indicator information, transition account state data and transition account behavior data of the recommendation system in a preset time period according to the initial account state data, the initial account behavior data and the initial recommendation object; the transition account state data represents improvement data of the initial account state data based on the initial recommendation object; and the transition account behavior data represents improvement data of the initial account behavior data based on the initial recommendation object;

[0009] performing sparse-to-dense conversion on the preset indicator information in the preset time period to obtain preset indicator information of each time step in the preset time period of the recommendation system;

[0010] training the recommendation system based on the preset indicator information of each time step, the transition account state data and the transition account behavior data to obtain a target recommendation system; the preset indicator information corresponding to the target recommendation system satisfies a preset indicator condition.

[0011] In some possible embodiments, the preset index information of the recommendation system over the preset time period, the transition account state data and the transition account behavior data are determined according to the initial account state data, the initial account behavior data and the initial recommendation object, comprising:

[0012] inputting the initial account state data, the initial account behavior data and the initial recommendation object into a data simulator to obtain first account state data, first account behavior data and a first time step of the recommendation system;

[0013] determining a first recommendation object corresponding to the recommendation system based on the first account state data and the first account behavior data; a loop step: inputting the first account state data, the first account behavior data and the first recommendation object into the data simulator to obtain second account state data, second account behavior data and a second time step of the recommendation system; until the preset index information of the recommendation system over the preset time period, the transition account state data and the transition account behavior data are obtained;

[0014] wherein the preset time period is composed of a plurality of time steps in the loop, and the plurality of time steps include the first time step and the second time step.

[0015] In some possible embodiments, the preset index information over the preset time period is converted from sparse to dense to obtain the preset index information of the recommendation system at each time step in the preset time period, comprising:

[0016] determining recommendation system state data corresponding to each time step based on account state data and account behavior data corresponding to each time step in the preset time period;

[0017] determining recommendation system behavior data corresponding to each time step based on a recommendation object corresponding to each time step in the preset time period;

[0018] inputting the preset time period, the preset index information over the preset time period, the recommendation system state data corresponding to each time step and the recommendation system behavior data corresponding to each time step into a trained index information decomposer to obtain the preset index information of the recommendation system at each time step.

[0019] In some possible embodiments, the method further comprises:

[0020] constructing an original information decomposer;

[0021] determining reference index information corresponding to each time step based on the preset time period and the preset index information over the preset time period;

[0022] The recommendation system state data and the behavior data corresponding to each time step are input into the original information decomposer to obtain prediction index information corresponding to each time step.

[0023] The original information analyzer is trained based on the reference index information corresponding to each time step and the prediction index information corresponding to each time step.

[0024] In a case where the iteration termination condition is met, the index information decomposer is obtained.

[0025] In some possible embodiments, in a case where the iteration termination condition is met, the index information decomposer is obtained, and the method comprises the following steps of:

[0026] In a case where a difference between the prediction index information corresponding to any time step in each time step and the reference index information is less than or equal to a first preset difference value, and a difference between the cumulative index information and preset index information in a preset time period is less than or equal to a second preset difference value, the training of the original information decomposer is terminated.

[0027] The trained original information decomposer is determined as the index information decomposer.

[0028] In some possible embodiments, the method further comprises the following steps of:

[0029] Obtaining a sample data set and initial data of an original generator;

[0030] Training the original generator and the discriminator based on the sample data set and the initial data to obtain a target generator; the target generator comprises a data simulator.

[0031] In some possible embodiments, the step of obtaining the sample data set and the initial data of the original generator comprises the following steps of:

[0032] Obtaining historical offline data of the recommendation system;

[0033] Dividing the historical offline data into a plurality of sample data based on access rounds; the number of sample data is the same as the number of access rounds; each sample data in the plurality of sample data comprises sample account state data, sample account behavior data and sample system behavior data.

[0034] Sampling the historical offline data of the recommendation system to obtain the initial data of the original generator.

[0035] In some possible embodiments, the step of training the original generator and the discriminator based on the sample data set and the initial data to obtain the target generator comprises the following steps of:

[0036] Inputting the sample data corresponding to the first access round and the initial data into the discriminator to obtain a first discrimination result of the sample data corresponding to the first access round and a second discrimination result of the initial data.

[0037] determining a target loss of the discriminator based on the label information of the sample data corresponding to the first access round, the first label information of the initial data corresponding to the discriminator, the first discrimination result and the second discrimination result;

[0038] training the discriminator based on the target loss of the discriminator;

[0039] determining a target loss of the original generator based on the second discrimination result and the first label information of the initial data corresponding to the original generator;

[0040] training the original generator based on the target loss of the original generator;

[0041] generating first generated data based on the original generator and the initial data;

[0042] inputting the sample data corresponding to the second access round and the first generated data into the discriminator to obtain a first discrimination result of the sample data corresponding to the second access round and a second discrimination result of the first generated data; the loop step includes: determining a target loss of the discriminator based on the label information of the sample data corresponding to the second access round, the first label information of the first generated data corresponding to the discriminator, the first discrimination result and the second discrimination result;

[0043] obtaining the target generator when the iteration termination condition is met.

[0044] In some possible embodiments, the original generator includes an original simulator and an original recommender; the first generated data is generated based on the original generator and the initial data, including:

[0045] inputting the sample account state data and the sample account behavior data in the initial data into the original recommender to obtain first system behavior data;

[0046] inputting the first system behavior data, the sample account state data and the sample account behavior data in the initial data into the original simulator to obtain first account state data and first account behavior data;

[0047] determining the first generated data based on the first account state data, the first account behavior data and the first system behavior data.

[0048] In some possible embodiments, the preset index information of the recommendation system on the preset time period, the transition account state data and the transition account behavior data are determined according to the initial account state data, the initial account behavior data and the initial recommendation object, including:

[0049] The second data determination module is configured to determine preset index information, transition account state data and transition account behavior data of the recommendation system in the preset time period according to the initial account state data, the initial account behavior data and the initial recommendation object; the transition account state data represents improvement data of the initial account state data based on the initial recommendation object; and the transition account behavior data represents improvement data of the initial account behavior data based on the initial recommendation object.

[0050] According to a second aspect of the embodiments of the present disclosure, a recommendation method is provided, comprising:

[0051] obtaining account state data and account execution data of a target account;

[0052] inputting the account state data and the account execution data of the target account into the target recommendation system trained according to any one of the recommendation system training methods in claims 1 to 10, to obtain a target recommendation object.

[0053] According to a third aspect of the embodiments of the present disclosure, a recommendation system training apparatus is provided, comprising:

[0054] The first data determination module is configured to determine initial account state data and initial account behavior data of the recommendation system.

[0055] The object determination module is configured to determine an initial recommendation object corresponding to the recommendation system based on the initial account state data and the initial account behavior data.

[0056] The second data determination module is configured to determine preset index information, transition account state data and transition account behavior data of the recommendation system in a preset time period according to the initial account state data, the initial account behavior data and the initial recommendation object; the transition account state data represents improvement data of the initial account state data based on the initial recommendation object; and the transition account behavior data represents improvement data of the initial account behavior data based on the initial recommendation object.

[0057] The information conversion module is configured to convert the preset index information in the preset time period from sparse to dense, to obtain preset index information of the recommendation system in each time step in the preset time period.

[0058] The training module is configured to train the recommendation system based on the preset index information, the transition account state data and the transition account behavior data of each time step, to obtain a target recommendation system; and the preset index information corresponding to the target recommendation system satisfies a preset index condition.

[0059] In some possible embodiments, the second data determination module is configured to:

[0060] input the initial account state data, the initial account behavior data and the initial recommendation object into a data simulator, to obtain first account state data, first account behavior data and a first time step of the recommendation system;

[0061] determining a first recommendation object corresponding to the recommendation system based on the first account state data and the first account behavior data; and repeating the steps of inputting the first account state data, the first account behavior data and the first recommendation object into the data simulator to obtain second account state data, second account behavior data and a second time step of the recommendation system until the preset index information, the transition account state data and the transition account behavior data of the recommendation system in the preset time period are obtained;

[0062] The preset time period is composed of a plurality of time steps in the loop, and the plurality of time steps include the first time step and the second time step.

[0063] In some possible embodiments, the information conversion module is configured to perform:

[0064] determining recommendation system state data corresponding to each time step based on the account state data and the account behavior data corresponding to each time step in the preset time period;

[0065] determining recommendation system behavior data corresponding to each time step based on the recommendation object corresponding to each time step in the preset time period;

[0066] inputting the preset time period, the preset index information in the preset time period, the recommendation system state data corresponding to each time step and the recommendation system behavior data corresponding to each time step into the trained index information decomposer to obtain the preset index information of the recommendation system in each time step.

[0067] In some possible embodiments, the apparatus further includes a decomposer training module configured to perform:

[0068] constructing an original information decomposer;

[0069] determining reference index information corresponding to each time step based on the preset time period and the preset index information in the preset time period;

[0070] inputting the recommendation system state data and behavior data corresponding to each time step into the original information decomposer to obtain predicted index information corresponding to each time step;

[0071] training the original information analyzer based on the reference index information corresponding to each time step and the predicted index information corresponding to each time step;

[0072] obtaining the index information decomposer when the iteration termination condition is met.

[0073] In some possible embodiments, the decomposer training module is configured to perform:

[0074] in a case that a difference between the prediction index information corresponding to any of the time steps and the reference index information is less than or equal to a first preset difference, and a difference between the cumulative index information and the preset index information in the preset time period is less than or equal to a second preset difference, terminating the training of the original information decomposer;

[0075] determining the trained original information decomposer as the index information decomposer.

[0076] In some possible embodiments, the apparatus further includes a target generator determination module configured to perform:

[0077] obtaining a sample data set and initial data of the original generator;

[0078] training the original generator and the discriminator based on the sample data set and the initial data to obtain the target generator; the target generator includes a data simulator.

[0079] In some possible embodiments, the target generator determination module is configured to perform:

[0080] obtaining historical offline data of the recommendation system;

[0081] dividing the historical offline data into a plurality of sample data based on access rounds; a number of the sample data is the same as a number of the access rounds; each of the plurality of sample data includes sample account state data, sample account behavior data, and sample system behavior data;

[0082] sampling the historical offline data of the recommendation system to obtain the initial data of the original generator.

[0083] In some possible embodiments, the target generator determination module is configured to perform:

[0084] inputting the sample data corresponding to the first access round and the initial data into the discriminator to obtain a first discrimination result of the sample data corresponding to the first access round and a second discrimination result of the initial data;

[0085] determining a target loss of the discriminator based on the labeled information of the sample data corresponding to the first access round, the first labeled information corresponding to the initial data to the discriminator, the first discrimination result, and the second discrimination result;

[0086] training the discriminator based on the target loss of the discriminator;

[0087] determining a target loss of the original generator based on the second discrimination result and the first labeled information corresponding to the initial data to the original generator;

[0088] training the original generator based on the target loss of the original generator;

[0089] generating first generated data based on the original generator and the initial data;

[0090] inputting the sample data corresponding to the second access round and the first generated data into the discriminator to obtain a first discrimination result of the sample data corresponding to the second access round and a second discrimination result of the first generated data; and

[0091] in a case where the iteration termination condition is met, obtaining the target generator.

[0092] In some possible embodiments, the original generator comprises an original simulator and an original recommender; and the target generator determination module is configured to perform:

[0093] inputting the sample account state data and the sample account behavior data in the initial data into the original recommender to obtain first system behavior data;

[0094] inputting the first system behavior data, the sample account state data and the sample account behavior data in the initial data into the original simulator to obtain first account state data and first account behavior data;

[0095] determining the first generated data based on the first account state data, the first account behavior data and the first system behavior data.

[0096] In some possible embodiments, the second data determination module is configured to perform:

[0097] determining, according to the initial account state data, the initial account behavior data and the initial recommendation object, preset index information corresponding to a retention rate and / or an account preference degree of the recommendation system in a preset time period, transition account state data and transition account behavior data.

[0098] According to a fourth aspect of the embodiments of the present disclosure, a recommendation device is provided, comprising:

[0099] the data acquisition module is configured to perform acquiring account state data and account execution data of a target account;

[0100] the object recommendation module is configured to perform inputting the account state data and the account execution data of the target account into the target recommendation system trained by the recommendation system training device to obtain a target recommendation object.

[0101] According to a fifth aspect of the embodiments of the present disclosure, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the method of any one of the above first aspect or second aspect.

[0102] According to a sixth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can execute the method of any one of the first aspect or the second aspect of the embodiments of the present disclosure.

[0103] According to a seventh aspect of the embodiments of the present disclosure, a computer program product is provided, the computer program product comprises a computer program stored in a readable storage medium, at least one processor of a computer device reads and executes the computer program from the readable storage medium, so that the computer device executes the method of any one of the first aspect or the second aspect of the embodiments of the present disclosure.

[0104] The embodiments of the present disclosure at least have the following beneficial effects:

[0105] The initial account state data and the initial account behavior data of the recommendation system are determined, the initial recommendation object corresponding to the recommendation system is determined based on the initial account state data and the initial account behavior data, the preset index information of the recommendation system in the preset time period, the transition account state data and the transition account behavior data are determined according to the initial account state data, the initial account behavior data and the initial recommendation object, the transition account state data represents the improvement data of the initial account state data based on the initial recommendation object; the transition account behavior data represents the improvement data of the initial account behavior data based on the initial recommendation object, the preset index information in the preset time period is converted from sparse to dense to obtain the preset index information of the recommendation system in each time step in the preset time period, the target recommendation system is obtained by training the recommendation system based on the preset index information of each time step, the transition account state data and the transition account behavior data, and the preset index information corresponding to the target recommendation system satisfies the preset index condition. The long-term index of the sparse can be decomposed into the long-term index of the dense which can be improved, and then the recommendation system trained based on the long-term index meets the index requirement.

[0106] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0107] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the specification, serve to explain the principles of the present disclosure, and do not constitute an improper limitation on the present disclosure.

[0108] Figure 1 is a schematic diagram of an application environment according to an exemplary embodiment;

[0109] Figure 2This is a flowchart illustrating a recommendation system training method according to an exemplary embodiment;

[0110] Figure 3 This is a flowchart illustrating an exemplary embodiment of obtaining a sample dataset and initial data for the original generator;

[0111] Figure 4 This is a flowchart illustrating the training of a generator and a discriminator according to an exemplary embodiment;

[0112] Figure 5 This is a framework diagram illustrating a generator and a discriminator according to an exemplary embodiment;

[0113] Figure 6 This is a flowchart illustrating a data simulator application according to an exemplary embodiment;

[0114] Figure 7 This is a flowchart illustrating an exemplary embodiment of obtaining preset indicator information over a preset time period;

[0115] Figure 8 This is a flowchart illustrating a method for obtaining preset indicator information for each time step, according to an exemplary embodiment.

[0116] Figure 9 This is a flowchart illustrating the training of an index decomposer according to an exemplary embodiment;

[0117] Figure 10 This is a schematic diagram illustrating a training structure for a recommendation system according to an exemplary embodiment;

[0118] Figure 11 This is a flowchart illustrating a recommendation method according to an exemplary embodiment;

[0119] Figure 12 This is a block diagram illustrating a recommendation system training apparatus according to an exemplary embodiment;

[0120] Figure 13 This is a block diagram of a recommendation system training apparatus according to an exemplary embodiment;

[0121] Figure 14 This is a block diagram illustrating an electronic device for training or recommending a recommendation system, according to an exemplary embodiment. Detailed Implementation

[0122] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0123] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-described drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation described in the following exemplary embodiments does not represent all implementations consistent with the present disclosure. Rather, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0124] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, analyzed data, etc.) involved in the present disclosure are all information and data authorized by the user or authorized by all parties.

[0125] Please refer to Figure 1 , Figure 1 is a schematic diagram of an application environment of a recommendation system training method according to an exemplary embodiment, as shown in Figure 1 , the application environment can include a recommendation system training client 01 and a server 02.

[0126] In the embodiments of the present application, the client 01 can obtain a sample data set through interaction with the server 02.

[0127] Optionally, the client 01 can include but is not limited to devices such as smartphones, desktop computers, tablets, laptops, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, smart wearable devices, etc. It can also be software running on the above devices, such as applications, applets, etc. Optionally, the operating system running on the device can include but is not limited to Android system, IOS system, Linux, Windows, Unix, etc.

[0128] Optionally, the server 02 determines initial account state data and initial account behavior data of the recommendation system, determines an initial recommendation object corresponding to the recommendation system based on the initial account state data and the initial account behavior data, determines preset index information of the recommendation system on a preset time period, transition account state data and transition account behavior data according to the initial account state data, the initial account behavior data and the initial recommendation object, the transition account state data representing improvement data of the initial account state data based on the initial recommendation object; the transition account behavior data representing improvement data of the initial account behavior data based on the initial recommendation object, performing sparse-to-dense conversion on the preset index information on the preset time period to obtain preset index information of the recommendation system at each time step in the preset time period, training the recommendation system based on the preset index information at each time step, the transition account state data and the transition account behavior data to obtain a target recommendation system, and the preset index information corresponding to the target recommendation system satisfying a preset index condition. After the server 02 obtains the target recommendation system, the account state data and the account execution data of the target account are obtained; the account state data and the account execution data of the target account are input into the obtained target recommendation system to obtain a target recommendation object.

[0129] The server 02 can include a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms. The operating system running on the server can include but is not limited to Android system, IOS system, linux, windows, Unix, etc.

[0130] In addition, it should be noted that Figure 1 The above merely illustrates an application environment of the recommendation system training method provided by the present disclosure. In actual application, other application environments can also be included.

[0131] Figure 2 is a flowchart of a recommendation system training method according to an exemplary embodiment, as Figure 2 As shown, the recommendation system training method can be applied to a server or a client, and includes the following steps:

[0132] In step S201, initial account state data and initial account behavior data of the recommendation system are determined.

[0133] In the embodiments of the present application, the server can determine initial account state data and initial account behavior data of the recommendation system. Optionally, the initial account state data includes a label list of a plurality of objects (videos) clicked by a user among all objects recommended by the recommendation system in a corresponding access round; and the initial account behavior data includes a viewing feedback representation of a client corresponding to the account information on all objects recommended by the recommendation system in the corresponding access round, and an interval time from the current access round to a next access round.

[0134] In the embodiments of the present application, the recommendation system can be an object recommendation system, and the object can include a video, music, information, and the like.

[0135] In an optional embodiment, the initial account state data and the initial account behavior data of the recommendation system can be initial account state data and initial account behavior data of one account, and can also be initial account state data and initial account behavior data of a plurality of accounts.

[0136] In step S203, initial recommended objects corresponding to the recommendation system are determined based on the initial account state data and the initial account behavior data.

[0137] In the embodiments of the present application, the server can determine initial recommended objects corresponding to the recommendation system based on the initial account state data and the initial account behavior data.

[0138] Optionally, the recommendation system at this time is not trained, and therefore, the server can input the initial account state data and the initial account behavior data into the recommendation system without training to obtain initial recommended objects output by the recommendation system.

[0139] In step S205, preset index information of the recommendation system on a preset time period, transition account state data, and transition account behavior data are determined according to the initial account state data, the initial account behavior data, and the initial recommended objects; the transition account state data represents improvement data of the initial recommended objects on the initial account state data; and the transition account behavior data represents improvement data of the initial recommended objects on the initial account behavior data.

[0140] Optionally, the transition account state data includes a label list of a plurality of objects (videos) clicked by a user among all objects recommended by the recommendation system in a corresponding access round, which is based on the initial recommended objects on the initial account state data; and the transition account behavior data includes a viewing feedback representation of a client corresponding to the account information on all objects recommended by the recommendation system in the corresponding access round, which is based on the initial recommended objects on the initial account behavior data, and an interval time from the current access round to a next access round.

[0141] In an alternative embodiment, the server can put the whole process on the platform carrying the recommendation system, i.e. the server acquires the initial account state data and the initial account behavior data of the recommendation system, then inputs the initial account state data and the initial account behavior data into the recommendation system without training, obtains the initial recommended objects output by the recommendation system, and determines the preset index information of the recommendation system in the preset time period based on the feedback (such as click viewing) of the user to the initial recommended objects, and acquires the new transition account state data and the transition account behavior data.

[0142] In another alternative embodiment, the server can use the data simulator to simulate the feedback of the user, and accelerate the whole training process.

[0143] The following describes an embodiment of determining the data simulator. In an alternative embodiment, the data simulator is part of the target generator trained based on the original generator.

[0144] In the embodiment of the application, the server can acquire the sample data set and the initial data of the original generator. Then, the server can train the original generator and the discriminator based on the sample data set and the initial data to obtain the target generator; wherein the target generator includes the data simulator.

[0145] Optionally, the original generator can be a newly constructed generator without any training. It can be a previously constructed generator that has been trained to a certain extent but not completed.

[0146] In the embodiment of the application, the sample data set can include a plurality of sample data. Optionally, the plurality of sample data can be sample data corresponding to one account information. Optionally, the plurality of sample data can be sample data corresponding to a plurality of account information, i.e. each account information in the plurality of account information corresponds to a plurality of sample data, and the plurality of sample data corresponding to each account information constitutes the sample data set.

[0147] The following describes an embodiment of acquiring the sample data set and the initial data of the original generator. Figure 3 is a flowchart of acquiring the sample data set and the initial data of the original generator according to an exemplary embodiment, as shown in Figure 3 , which includes:

[0148] In step S301, historical offline data of the recommendation system is acquired.

[0149] In the embodiment of the application, the recommendation system can be an object recommendation system, and the object can include videos, music, information, etc. The following describes the recommendation system as a video recommendation system.

[0150] In step S303, the historical offline data is divided into a plurality of sample data based on the access rounds; the number of the sample data is the same as the number of the access rounds; each of the plurality of sample data comprises sample account state data, sample account behavior data and sample system behavior data.

[0151] In the embodiment of the application, the server can obtain historical offline data corresponding to one account information or a plurality of account information from the recommendation system.

[0152] In an optional embodiment, taking the historical offline data as data corresponding to one account information as an example, in the embodiment of the application, the server can obtain historical offline data of the account information in a period of time from the recommendation system. Optionally, the period of time can be any length of time, such as a week, a month, a quarter, etc.

[0153] In the embodiment of the application, the historical offline data can be a historical record retained by the recommendation system after recommending a video to the account information when the client corresponding to the account information accesses the recommendation system. Optionally, the server can divide the historical offline data into a plurality of sample data based on the access rounds.

[0154] Optionally, one access round can include a process in which the application corresponding to the recommendation system is started by the client corresponding to the account information to the process in which the application is closed, and the process can include a process in which the application is hung in the background of the client.

[0155] In this way, the server can divide the historical offline data of the account information in the period of time into a plurality of sample data with the same number as the number of rounds based on the number of rounds of the access rounds. For example, assuming that the access rounds of the recommendation system accessed by the client corresponding to the account information in the period of time are 25 times, the server can divide the historical offline data in the period of time into 25 sample data.

[0156] In another optional embodiment, taking the historical offline data as data corresponding to a plurality of account information as an example, in the embodiment of the application, the server can obtain historical offline data of each of the plurality of account information in a period of time from the recommendation system. Optionally, the period of time can be any length of time, such as a week, a month, a quarter, etc.

[0157] In the embodiment of the application, the historical offline data can be a historical record retained by the recommendation system after recommending a video to the account information when the client corresponding to the account information accesses the recommendation system. Optionally, the server can divide the historical offline data into a plurality of sample data based on the access rounds.

[0158] Thus, the server can divide the historical offline data of each account information in the period of time into a plurality of sample data according to the number of access rounds. For example, assuming that the first account information in the plurality of account information corresponds to a client that accesses the recommendation system for 25 times in the period of time, the server can divide the historical offline data of the first account information in the period of time into 25 sample data; the second account information in the plurality of account information corresponds to a client that accesses the recommendation system for 20 times in the period of time, the server can divide the historical offline data of the second account information in the period of time into 20 sample data; the third account information in the plurality of account information corresponds to a client that accesses the recommendation system for 18 times in the period of time, the server can divide the historical offline data of the third account information in the period of time into 18 sample data.

[0159] In the embodiments of the present application, each sample data can include sample account state data, sample account behavior data and sample system behavior data in the corresponding access round.

[0160] Optionally, the sample account state data includes a label list of a plurality of objects (videos) clicked by the user in all objects recommended by the recommendation system in the corresponding access round. Each label c in the label list k is used to represent the label of each object clicked by the user. For example, the label of a certain object clicked by the user is sports, and the label c k of the object can be represented by “0”. Of course, the c k represented by “0” is an optional representation, and the actual representation is not limited. Wherein, u in the above formula refers to the label list related to the account, and t refers to the number of interactions between the account and the recommendation system about the object in the corresponding access round, which can be related to the time t.

[0161] Since each label c k is a discrete numerical value, in order to facilitate the subsequent embodiments, the server can convert c k into a high-dimensional continuous vector Thus, the sample account state data in each sample data can be represented as:

[0162]

[0163] Optionally, the sample account behavior data includes a viewing feedback representation of all objects recommended by the recommendation system by the client corresponding to the account information in the corresponding access round, and an interval time from the access round to the next access round.

[0164] Suppose that the recommendation system recommends N objects in this access round, the viewing feedback representation of the N objects can be expressed as , where the viewing feedback representation of each object can be expressed as 0 / 1, for example, "0" represents viewing, and "1" represents not viewing, k = 1, 2,..., N.

[0165] In this way, the sample account behavior data in each sample data can be expressed as:

[0166]

[0167] Optionally, the sample system behavior data includes the representation of all objects recommended by the recommendation system in the corresponding access round. Assuming that the object is a video, the identification of the video can include the label representation of the video (such as travel, game, or fitness, etc.), the content representation (elements contained in each frame of the video, including house, food, and plant, etc.), and the statistical characteristics (number of views, number of likes, number of forwards, etc.). In the embodiment of the application, the sample system behavior data can be expressed as .

[0168] In order to facilitate subsequent training and highlight the connection between the recommendation system and the account information, the server can construct the sample system state data based on the sample account state data and the sample account behavior data, that is, the combination of the sample account state data and the sample account behavior data at time t, and the expression is as follows:

[0169]

[0170] In step S305, the historical offline data of the recommendation system is sampled to obtain the initial data of the original generator.

[0171] In an optional embodiment, the server can sample the historical offline data, such as taking the sample data at time t as the initial data of the original generator, or combining part of the sample data at different times to obtain the initial data of the original generator.

[0172] In this way, the server can obtain the sample data set and the initial data of the original generator, and provide training effective data for subsequent training of the original generator and the discriminator.

[0173] In the embodiment of the application, the server can train the original generator and the discriminator based on the sample data set corresponding to one account information, or train the original generator and the discriminator based on the sample data set corresponding to multiple account information, to obtain the trained target generator and discriminator, and then obtain the data simulator in the target generator.

[0174] The following is described by taking a sample data set corresponding to one account information as an example, and the application of a sample data set corresponding to multiple account information can refer to the application of a sample data set corresponding to one account information, which will not be described here.

[0175] In order to enable the data simulator in the trained target generator to follow the time sequence when simulating data, the original generator and the discriminator can be trained by using the sample data set and the initial data. The multiple samples in the sample data set can be sorted according to the time sequence, and the sorted sample data set can be expressed as:

[0176] D = (τ1, τ2,... τ n ) …… Equation (4)

[0177] Each sample data in the sample data set can be expressed as:

[0178]

[0179] Figure 4 is a flowchart of training a generator and a discriminator according to an example embodiment, as shown in Figure 4 , including:

[0180] In step S401, the sample data corresponding to the first access round and the initial data are input into the discriminator to obtain the first discrimination result of the sample data corresponding to the first access round and the second discrimination result of the initial data.

[0181] In the embodiment of the application, the original generator can be constructed based on a deep neural network.

[0182] Figure 5 is a framework diagram of a generator and a discriminator according to an example embodiment, as shown in Figure 5 , including an offline data module, an original generator and a discriminator, wherein the original generator includes an original recommender and an original simulator. The following describes the training process in combination with Figure 5 .

[0183] In the embodiment of the application, the server can input the sample data corresponding to the first access round output by the offline data module into the discriminator, and at the same time, the original generator inputs the initial data into the discriminator to obtain the first discrimination result of the sample data corresponding to the first access round and the second discrimination result of the initial data.

[0184] Optionally, the discriminator usually represents the probability of data being true or false in 0 to 1. In fact, since the initial data is sampled or synthesized, and the original generator has not been trained at this time, the second discrimination result of the initial data tends to be 0 (the closer to 0, the more false), and the first discrimination result of the sample data corresponding to the first access round tends to be 1 (the closer to 1, the more true). For example, assuming that the first discrimination result is 0.85 and the second discrimination result is 0.25.

[0185] In step S402, the target loss of the discriminator is determined based on the label information of the sample data corresponding to the first access round, the first label information of the initial data corresponding to the discriminator, the first discrimination result and the second discrimination result.

[0186] In an optional embodiment, the label information of the sample data corresponding to the first access round and the first label information of the initial data corresponding to the discriminator can be obtained first. The first label information of the initial data corresponding to the discriminator and the label information of the sample data corresponding to the first access round are preset according to actual conditions. Since the initial data is false with respect to the discriminator, the first label information of the initial data corresponding to the discriminator is 0. Since the sample data corresponding to the first access round is true with respect to the discriminator, the label information of the sample data corresponding to the first access round is 1.

[0187] Optionally, the first discrimination loss can be determined based on the first discrimination result and the label information of the sample data corresponding to the first access round, the second discrimination loss can be determined based on the second discrimination result and the first label information of the initial data corresponding to the discriminator, and the target loss of the discriminator can be determined according to the target loss function of the discriminator, the first discrimination loss and the second discrimination loss.

[0188] In step S403, the discriminator is trained based on the target loss of the discriminator.

[0189] In this way, the server can train the discriminator based on the target loss of the discriminator, and complete the first round of training of the discriminator.

[0190] Subsequently, the discriminator can return a feedback to the offline data module and the original generator, to inform the offline data module and the original generator that the discriminator has completed a round of training, so that the offline data module sends the sample data corresponding to the second access round to the discriminator, and the original generator then completes its own training.

[0191] In step S404, the target loss of the original generator is determined based on the second discrimination result and the first label information of the initial data corresponding to the original generator.

[0192] In the embodiment of the present application, since the initial data is true with respect to the original generator, the first label information corresponding to the initial data of the original generator is 1, and the server can substitute the second discrimination result, the first label information corresponding to the initial data of the original generator, into the target loss function of the original generator to obtain the target loss of the original generator.

[0193] In step S405, the original generator is trained based on the target loss of the original generator.

[0194] In this way, the server can train the original generator based on the target loss of the original generator, and complete the first round of training of the original generator.

[0195] In step S406, the first generated data is generated based on the original generator and the initial data.

[0196] After the original generator completes the first round of training, the first generated data can be generated based on the original generator and the initial data.

[0197] In the embodiment of the present application, after completing a round of training, the server can input the sample account state data and the sample account behavior data in the initial data into the original recommender, and the original recommender can learn the recommendation behavior of the recommendation system to recommend N objects to obtain the first system behavior data. The first system behavior data is the label representation, the content representation and the statistical representation of each object in the newly recommended N objects. That is, the sample account state data and the sample account behavior data are used to generate the first system behavior data at t+1 time, and the expression is as follows:

[0198]

[0199] Subsequently, the server combines the sample account state data and the sample account behavior data in the initial data into sample system state data, and inputs the first system behavior data and the sample system state data into the original simulator to obtain the first account state data and the first account behavior data, and the expression is as follows:

[0200]

[0201] Optionally, the original generator determines the first generated data based on the first account state data, the first account behavior data and the first system behavior data. In this way, the original recommender and the original simulator iteratively interact in time sequence to generate user recommendation system interaction data, i.e., generated data corresponding to sample data.

[0202] In step S407, the second access round corresponding sample data and the first generated data are input into the discriminator to obtain the first discrimination result of the second access round corresponding sample data and the second discrimination result of the first generated data; the step is cycled: the target loss of the discriminator is determined based on the annotation information of the second access round corresponding sample data, the first annotation information corresponding to the discriminator of the first generated data, the first discrimination result and the second discrimination result.

[0203] In the embodiment of the application, after obtaining the first generated data, the server can input the second access round corresponding sample data output by the offline data module into the discriminator, and input the first generated data into the discriminator to obtain the first discrimination result of the second access round corresponding sample data and the second discrimination result of the first generated data. Then, referring to the specific process described above, the target loss of the discriminator is determined based on the annotation information of the second access round corresponding sample data, the first annotation information corresponding to the discriminator of the first generated data, the first discrimination result of the second access round corresponding sample data and the second discrimination result of the first generated data, the discriminator is trained based on the target loss of the discriminator, and thus the second round of discriminator training is completed.

[0204] Subsequently, the server can determine the target loss of the original generator based on the second discrimination result of the first generated data and the first annotation information corresponding to the original generator of the first generated data, train the original generator based on the target loss of the original generator, and thus complete the training of the original generator in the second round.

[0205] In step S408, the target generator is obtained when the iteration termination condition is met.

[0206] Optionally, the iteration can be terminated to obtain the target generator after the server completes a preset number of iterations or the discriminator converges.

[0207] After referring to the completion of the second round of training and subsequent multiple times of training of the discriminator and the original generator, the trained original generator, that is, the target generator, and the trained discriminator can be obtained when the iteration termination condition is met.

[0208] Since the trained original generator is obtained, the original recommender and the original simulator contained therein are also trained, and thus the trained object recommender corresponding to the original recommender and the trained data simulator corresponding to the original simulator are obtained.

[0209] Optionally, the target loss function of the discriminator can be represented as:

[0210]

[0211] wherein, for the sample data, for the generated data, for the discrimination result.

[0212] In an alternative embodiment, the iterative interaction of the generator and the discriminator is such that the training of the discriminator and the generator is completed in each iteration, and the generator can be trained first, and then the discriminator, or the discriminator can be trained first, and then the generator, as shown above.

[0213] Optionally, in order to verify whether the training process is sequential, the server can use the verification data set to verify the discriminator and the target generator after completing several rounds of training in the training process, and then verify the ability of the data simulator.

[0214] From the above, it can be seen that inputting the system behavior data and the system state data of the last round into the data simulator can obtain the account state data and the account behavior data of the next round, wherein the account behavior data can include the viewing feedback representation of each object, and the interval time between the current access round and the next access round Therefore, the trained data simulator can be used to simulate the process of the account accessing the recommendation system to be analyzed.

[0215] As known from the above, the initial account state data and the initial account behavior data of the recommendation system can be the initial account state data and the initial account behavior data of an account, or the initial account state data and the initial account behavior data of multiple accounts. Optionally, if the personal indicators of a to-be-analyzed account, including the monthly revisit frequency, are obtained through the data simulator, the initial account state data and the initial account behavior data of the recommendation system can be the initial account state data and the initial account behavior data of an account. Optionally, if the next-day retention rate or the monthly retention rate of the entire recommendation system is obtained through the data simulator, the initial account state data and the initial account behavior data of the recommendation system can be the initial account state data and the initial account behavior data of multiple accounts.

[0216] In the embodiment of the application, it is assumed that the next-day retention rate and / or the account preference degree of the recommendation system are to be analyzed, and the analysis is combined with the analysis of the account behavior data of the recommendation system. Figure 6 is a flowchart of a data simulator application according to an example embodiment. The server can obtain the initial account state data and the initial account behavior data of the recommendation system from the sample data set.

[0217] As Figure 6As shown, in some possible embodiments, the server can determine an initial recommendation object corresponding to the recommendation system based on the initial account state data and the initial account behavior data. Specifically, the server can determine a plurality of initial recommendation objects corresponding to each account based on the initial account state data and the initial account behavior data of each account.

[0218] Figure 7 is a flowchart of obtaining preset index information in a preset time period, such as a month corresponding to a month retention rate, according to an exemplary embodiment, which comprises:

[0219] In step S701, the initial account state data, the initial account behavior data and the initial recommendation object are input into the data simulator to obtain the first account state data, the first account behavior data and the first time step of the recommendation system.

[0220] As shown, the server inputs the initial account state data, the initial account behavior data and the initial recommendation object into the data simulator to obtain the first account state data, the first account behavior data and the first time step of the recommendation system. Figure 6

[0221] The first time step can be the interval time t1 between the first visit round and the next visit round in the Figure 6

[0222] In this way, the server simulates the visit process of the recommendation system of the recommendation system by using the data simulator.

[0223] In step S703, the first recommendation object corresponding to the recommendation system is determined based on the first account state data and the first account behavior data; the step is looped: the first account state data, the first account behavior data and the first recommendation object are input into the data simulator to obtain the second account state data, the second account behavior data and the second time step of the recommendation system; until the preset index information of the recommendation system in the preset time period, the transition account state data and the transition account behavior data are obtained; wherein the preset time period is composed of a plurality of time steps in the loop, and the plurality of time steps include the first time step and the second time step.

[0224] Since the new account state information and the new account behavior information, i.e. the first account state data and the first account behavior data, are obtained in step S701, the server can determine one or more first recommendation objects corresponding to the recommendation system based on the first account state data and the first account behavior data.

[0225] Subsequently, the server can input the first account state data, the first account behavior data and the first recommendation object into the data simulator to obtain the second account state data, the second account behavior data and the second time step of the recommendation system.​​​

[0226] wherein the second time step can be Figure 6 the revisit time t2 in the formula (2), i.e. the interval time between the second visit round and the next visit round.

[0227] Thus, the server simulates the second visit process of the recommendation system by using the data simulator. According to the above process, the server can obtain the time step (revisit time), account state information and account behavior information corresponding to each cycle by using the data simulator, until the total time length of each time step reaches the preset time period, such as one month corresponding to the monthly retention rate, and the cycle simulation process is ended. At this time, the server also obtains the transition account state data and the transition account behavior data corresponding to the time point of the preset time period.

[0228] In the embodiments of the present application, in terms of retention rate, in the process simulated by the data simulator, there can be two parts of accounts in the accounts of the recommendation system (such as 10000 accounts), including accounts (such as 6500 accounts) whose revisit time is 0, which are simulated for a period of time and have not reached the preset time period, and accounts (such as 3500 accounts) whose revisit time appears when the simulation reaches the preset time period. At this time, the server can determine that the monthly retention rate is 35%.

[0229] Suppose the sum of the revisit time t1, the revisit time t2, …, the revisit time t i is the preset time period, which can be the continuous time steps from the last time when a non-zero value such as the monthly retention rate appears to the next time when a non-zero value such as the monthly retention rate appears, and the specific formula is as follows.

[0230]

[0231] wherein r k = 0, k≠ T i and only wherein s refers to the recommendation system state data, a refers to the recommendation system behavior data, and the subscripts of s, a and r in the formula (9) refer to the time step.

[0232] Taking the monthly retention rate as an example, the preset time period is one month, and each time step r k in the preset time period is a non-zero value that is expected to appear, i.e. whether the monthly retention rate is obtained. Generally, only at the end of the month can the monthly retention rate be obtained, i.e. only in the formula is non-zero, such as 35%, and other r k are zero, and there is no retention rate.

[0233] Therefore, the data simulator can quickly simulate the access process of the recommendation system, and compared with the real data of the recommendation system, the time for completion can be greatly shortened, and the efficiency of subsequent data analysis can be improved.

[0234] In step S207, the preset index information on the preset time period is converted from sparse to dense to obtain the preset index information of the recommendation system at each time step in the preset time period.

[0235] In some possible embodiments, for the long-term index information such as the monthly retention rate in the above, the long-term index information can be obtained only at the preset time period. This makes it difficult to obtain the preset index information such as the monthly retention rate at a certain time step in the preset time period, and further makes it impossible to use the preset index information at a certain time step as data to improve and update the recommendation system. Based on this, the preset index information on the preset time period can be converted from sparse to dense to obtain the preset index information of the recommendation system at each time step in the preset time period.

[0236] Figure 8 is a flowchart for obtaining the preset index information at each time step according to an example embodiment, including:

[0237] In step S801, the recommendation system state data corresponding to each time step is determined based on the account state data and the account behavior data corresponding to each time step in the preset time period.

[0238] In step S803, the recommendation system behavior data corresponding to each time step is determined based on the recommendation object corresponding to each time step in the preset time period.

[0239] In step S805, the preset time period, the preset index information on the preset time period, the recommendation system state data corresponding to each time step, and the recommendation system behavior data corresponding to each time step are input into the trained index information decomposer to obtain the preset index information of the recommendation system at each time step.

[0240] The application also provides a training method of an index information decomposer, Figure 9 is a flowchart for training an index decomposer according to an example embodiment, including

[0241] In step S901, an original information decomposer is constructed.

[0242] In the embodiment of the application, the original information decomposer can be constructed based on a deep neural network.

[0243] In step S902, the reference index information corresponding to each time step is determined based on the preset time period and the preset index information on the preset time period.

[0244] In the embodiments of the present application, the server can use a uniform strategy to uniformly distribute the sparse numerical values to each time step of the sparse numerical values to obtain smooth non-zero numerical values, which is specifically shown as follows:

[0245]

[0246] For example, the monthly retention rate of 35% in the above is distributed to each time step of the sparse numerical values to obtain smooth non-zero numerical values.

[0247] In step S903, the recommendation system state data and the behavior data corresponding to each time step are input into the original information decomposer to obtain the prediction index information corresponding to each time step.

[0248] The original information decomposer R θ : S x A -> R, where S and A represent the recommendation system state data and the recommendation system behavior data in reinforcement learning, respectively. The server can input the recommendation system state data and the behavior data corresponding to each time step into the original information combination to assign a prediction index information R θ (s t , a t ).

[0249] In step S904, the original information analyzer is trained based on the reference index information corresponding to each time step and the prediction index information corresponding to each time step.

[0250] Subsequently, the server can train the original information analyzer based on the reference index information corresponding to each time step and the prediction index information corresponding to each time step.

[0251] In step S905, the index information decomposer is obtained when the iteration termination condition is met.

[0252] Therefore, the server determines that the difference between the prediction index information corresponding to any time step in each time step and the reference index information is less than or equal to a first preset difference, and the difference between the cumulative index information and the preset index information in the preset time period is less than or equal to a second preset difference, and terminates the training of the original information decomposer. The trained original information decomposer is determined as the index information decomposer. In this way, the sum of the preset index information of each time step after decomposition is equal to the value of the preset index information in the preset time period as much as possible, and the preset index information is more easily improved in the subsequent improvement process under the condition of ensuring the numerical value.

[0253] ​The server determines that the difference between the prediction index information and the reference index information corresponding to any of the time steps is less than or equal to the first preset difference, because the training target is to make the prediction index information corresponding to each time step as close as possible to the smooth non-zero value of each time step, which is achieved by the following formula:

[0254]

[0255] The server determines that the difference between the cumulative index information and the preset index information in the preset time period is less than or equal to the second preset difference, because it ensures that the cumulative value in the preset time period before and after the transformation is unchanged. Optionally, the cumulative value in the preset time period after the transformation can be ensured to be consistent by scaling the value in each time step, that is:

[0256] The scaling change formula is:

[0257]

[0258] In this way, the trained index information decomposer can obtain the monthly retention rate corresponding to each time step:

[0259]

[0260] wherein, is the monthly retention rate corresponding to each time step.

[0261] In this way, the embodiments of the present application can use the index information decomposer to decompose the sparse value into a dense value that is easy to improve, effectively improving the long-term index that is difficult to improve.

[0262] In addition, the present application proposes a user model learning method based on adversarial learning. Compared with ordinary supervised learning models or non-explicit user model modeling methods, the learned user feedback and behavior expression has stronger generalization, and can represent the user's preferences at different times and different behaviors.

[0263] In step S209, the recommendation system is trained based on the preset index information, the transition account state data and the transition account behavior data of each time step, and a target recommendation system is obtained; the preset index information corresponding to the target recommendation system satisfies the preset index condition.

[0264] In an optional embodiment, Figure 10 is a schematic diagram of a recommendation system training structure according to an exemplary embodiment, as shown in Figure 10As shown, the server can train the recommendation system by using the preset index information, the transition account state data and the transition account behavior data of each time step through the previously trained data simulator and the recommendation system, until the preset index information corresponding to the recommendation system meets the preset index condition, the training is stopped, and the target recommendation system is obtained.

[0265] From the above, the first round of preset index information, transition account state data and transition account behavior data in the preset time period are obtained through the user simulation behavior of the data simulator in the preset time period. Then, based on the fact that the preset index information in the preset time period is difficult to improve, the server can use the index decomposer to decompose the preset index information in the preset time period into the preset index information of each time step in the preset time period which is easy to improve. Subsequently, the preset index information, the transition account state data and the transition account behavior data of each time step are input into the recommendation system, and the recommendation system is trained to obtain the transition recommendation object of this round. In this way, the training of the first round of the recommendation system is completed.

[0266] Then, the server can input the transition account state data, the transition account behavior data and the transition recommendation object corresponding to the first round into the data simulator, and obtain the preset index information, the transition account state data and the transition account behavior data in the preset time period corresponding to the second round according to the flowchart applied by the data simulator. Figure 6 The preset index information, the transition account state data and the transition account behavior data in the preset time period corresponding to the second round are obtained according to the flowchart applied by the data simulator. Subsequently, the server can use the index decomposer to decompose the preset index information in the preset time period corresponding to the second round into the preset index information of each time step in the preset time period which is easy to improve. The preset index information, the transition account state data and the transition account behavior data of each time step are input into the recommendation system, and the recommendation system is trained to obtain the transition recommendation object of this round. In this way, the training of the second round of the recommendation system is completed.

[0267] Referring to the training process of the first round and the second round, the recommendation system can be continuously trained until the preset index information corresponding to the target recommendation system meets the preset index condition, such as the monthly retention rate reaching 50%, the training of the recommendation system is stopped, and the target recommendation system is obtained.

[0268] In another optional embodiment, the server can train the recommendation system by using the preset index information, the transition account state data and the transition account behavior data of each time step through the actual platform feedback system and the recommendation system, until the preset index information corresponding to the recommendation system meets the preset index condition, the training is stopped, and the target recommendation system is obtained. Optionally, in this embodiment, the server can use the actual platform feedback system to replace the data simulator in the previous embodiment to realize the training of the target recommendation system.

[0269] Thus, compared with the first embodiment, since the actual platform feedback system is slower than the data simulator in terms of feedback time, the server needs more time to obtain the trained target recommendation object. However, the actual platform feedback system is more accurate than the data simulator in terms of the simulated transition account state data, the transition account behavior data, and the preset indicator information in the preset time period, and thus a target recommendation system that meets the needs can be obtained.

[0270] The initial account state data, the first account state data, the second account state data, and the transition account state data (first round, second round,...) include a label list of a plurality of objects (videos) clicked by the user among all objects recommended by the recommendation system in the corresponding access round. The initial account behavior data, the first account behavior data, the second account behavior data, and the transition account behavior data (first round, second round,...) include a viewing feedback representation of the client corresponding to the account information on all objects recommended by the recommendation system, and an interval time from the current access round to the next access round. That is, the account state data represents the same data, but is given different names in different large cycle rounds for the convenience of writing. Similarly, the account behavior data represents the same data, but is given different names in different large cycle rounds for the convenience of writing.

[0271] Based on this, the trained target recommendation system can be used to recommend objects for a certain account.

[0272] Figure 11 is a flowchart of a recommendation method according to an exemplary embodiment, as shown in Figure 11 As shown, the recommendation system training method can be applied to a server or a client, and includes the following steps:

[0273] In step S1101, the account state data and the account execution data of the target account are obtained.

[0274] In step S1103, the account state data and the account execution data of the target account are input into the trained target recommendation system to obtain the target recommendation object.

[0275] The server can obtain the account state data and the account execution data of the target account, input the account state data and the account execution data of the target account into the trained target recommendation system, and obtain the target recommendation object. As mentioned above, the preset indicator information of the trained target recommendation system meets the preset indicator condition, and thus the target recommendation object recommended by the server for the target account meets the needs of the target account.

[0276] Figure 12is a block diagram of a recommendation system training apparatus according to an exemplary embodiment. Referring to Figure 12 The apparatus comprises:

[0277] A first data determination module 1201 is configured to determine initial account state data and initial account behavior data of a recommendation system.

[0278] An object determination module 1202 is configured to determine initial recommendation objects corresponding to the recommendation system based on the initial account state data and the initial account behavior data.

[0279] A second data determination module 1203 is configured to determine preset index information, transition account state data and transition account behavior data of the recommendation system over a preset time period according to the initial account state data, the initial account behavior data and the initial recommendation objects; the transition account state data represents improvement data of the initial account state data based on the initial recommendation objects; and the transition account behavior data represents improvement data of the initial account behavior data based on the initial recommendation objects.

[0280] An information conversion module 1204 is configured to convert the preset index information over the preset time period from sparse to dense to obtain preset index information of each time step of the recommendation system in the preset time period.

[0281] A training module 1205 is configured to train the recommendation system based on the preset index information, the transition account state data and the transition account behavior data of each time step to obtain a target recommendation system; and the preset index information corresponding to the target recommendation system satisfies a preset index condition.

[0282] In some possible embodiments, the second data determination module is configured to:

[0283] input the initial account state data, the initial account behavior data and the initial recommendation objects into a data simulator to obtain first account state data, first account behavior data and a first time step of the recommendation system;

[0284] determine first recommendation objects corresponding to the recommendation system based on the first account state data and the first account behavior data; and

[0285] wherein the preset time period is composed of a plurality of time steps in the loop, and the plurality of time steps include the first time step and the second time step.

[0286] In some possible embodiments, the information conversion module is configured to perform:

[0287] determining, based on the account state data and the account behavior data corresponding to each time step in the preset time period, the recommendation system state data corresponding to each time step;

[0288] determining, based on the recommendation object corresponding to each time step in the preset time period, the recommendation system behavior data corresponding to each time step;

[0289] inputting the preset time period, the preset index information on the preset time period, the recommendation system state data corresponding to each time step, and the recommendation system behavior data corresponding to each time step into the trained index information decomposer to obtain the preset index information of the recommendation system at each time step.

[0290] In some possible embodiments, the apparatus further includes a decomposer training module configured to perform:

[0291] constructing an original information decomposer;

[0292] determining, based on the preset time period and the preset index information on the preset time period, the reference index information corresponding to each time step;

[0293] inputting the recommendation system state data and the behavior data corresponding to each time step into the original information decomposer to obtain the predicted index information corresponding to each time step;

[0294] training the original information analyzer based on the reference index information corresponding to each time step and the predicted index information corresponding to each time step;

[0295] in a case where the iteration termination condition is met, obtaining the index information decomposer.

[0296] In some possible embodiments, the decomposer training module is configured to perform:

[0297] terminating the training of the original information decomposer in a case where a difference between the predicted index information and the reference index information corresponding to any time step in each time step is less than or equal to a first preset difference value, and a difference between the accumulated index information and the preset index information on the preset time period is less than or equal to a second preset difference value;

[0298] determining the trained original information decomposer as the index information decomposer.

[0299] In some possible embodiments, the apparatus further includes a target generator determination module configured to perform:

[0300] obtaining a sample data set and initial data of an original generator;

[0301] The original generator and the discriminator are trained based on the sample data set and the initial data to obtain a target generator; the target generator includes a data simulator.

[0302] In some possible embodiments, the target generator determination module is configured to perform:

[0303] Obtain historical offline data of a recommendation system;

[0304] Divide the historical offline data into a plurality of sample data based on access rounds; the number of the sample data is the same as the number of the access rounds; each sample data in the plurality of sample data includes sample account state data, sample account behavior data and sample system behavior data;

[0305] Sample the historical offline data of the recommendation system to obtain initial data of the original generator.

[0306] In some possible embodiments, the target generator determination module is configured to perform:

[0307] Input the sample data corresponding to the first access round and the initial data into the discriminator to obtain a first discrimination result of the sample data corresponding to the first access round and a second discrimination result of the initial data;

[0308] Determine a target loss of the discriminator based on the labeled information of the sample data corresponding to the first access round, the first labeled information of the initial data corresponding to the discriminator, the first discrimination result and the second discrimination result;

[0309] Train the discriminator based on the target loss of the discriminator;

[0310] Determine a target loss of the original generator based on the second discrimination result and the first labeled information of the initial data corresponding to the original generator;

[0311] Train the original generator based on the target loss of the original generator;

[0312] Generate first generated data based on the original generator and the initial data;

[0313] Input the sample data corresponding to the second access round and the first generated data into the discriminator to obtain a first discrimination result of the sample data corresponding to the second access round and a second discrimination result of the first generated data; the loop step is to determine a target loss of the discriminator based on the labeled information of the sample data corresponding to the second access round, the first labeled information of the first generated data corresponding to the discriminator, the first discrimination result and the second discrimination result;

[0314] In a case where an iteration termination condition is met, a target generator is obtained.

[0315] In some possible embodiments, the original generator comprises an original simulator and an original recommender; the target generator determination module is configured to perform:

[0316] inputting the sample account state data and the sample account behavior data in the initial data into the original recommender to obtain first system behavior data;

[0317] inputting the first system behavior data, the sample account state data and the sample account behavior data in the initial data into the original simulator to obtain first account state data and first account behavior data;

[0318] determining first generation data based on the first account state data, the first account behavior data and the first system behavior data.

[0319] In some possible embodiments, the second data determination module is configured to perform:

[0320] determining, according to the initial account state data, the initial account behavior data and the initial recommendation object, preset index information corresponding to a retention rate and / or an account preference degree of the recommendation system in a preset time period, transition account state data and transition account behavior data.

[0321] As to the apparatus in the above embodiments, the specific manners in which various modules perform operations have been described in detail in the embodiments of the method, and will not be described here in detail.

[0322] Figure 13 is a block diagram of a recommendation system training apparatus according to an example embodiment. Referring to Figure 13 The apparatus comprises:

[0323] The data acquisition module 1301 is configured to perform acquisition of account state data and account execution data of a target account.

[0324] The object recommendation module 1302 is configured to perform input of the account state data and the account execution data of the target account into a target recommendation system trained by the above-mentioned recommendation system training apparatus to obtain a target recommendation object.

[0325] As to the apparatus in the above embodiments, the specific manners in which various modules perform operations have been described in detail in the embodiments of the method, and will not be described here in detail.

[0326] Figure 14 is a block diagram of an electronic device 2000 for recommendation system training or recommendation according to an example embodiment. For example, the apparatus 2000 can be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0327] Referring to Figure 14 The device 2000 can include one or more of the following components: a processing component 2002, a memory 2004, a power supply component 2006, a multimedia component 2008, an audio component 2010, an input / output (I / O) interface 2012, a sensor component 2014, and a communication component 2016.

[0328] The processing component 2002 usually controls overall operations of the device 2000, such as operations associated with displaying, making phone calls, data communications, camera operations and recording operations. The processing component 2002 can include one or more processors 2020 to execute instructions to complete all or part of steps of the above methods. In addition, the processing component 2002 can include one or more modules to facilitate interaction between the processing component 2002 and other components. For example, the processing component 2002 can include a multimedia module to facilitate the interaction between the multimedia component 2008 and the processing component 2002.

[0329] The memory 2004 is configured to store various types of data to support operations of the device 2000. Examples of these data include instructions for any application or method operating on the device 2000, contact data, phonebook data, messages, pictures, videos, and the like. The memory 2004 can be implemented by any type of volatile or non-volatile storage devices or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0330] The power supply component 2006 provides power for various components of the device 2000. The power supply component 2006 can include a power supply management system, one or more power supplies, and other components associated with generating, managing and distributing power for the device 2000.

[0331] The multimedia component 2008 includes a screen providing an output interface between the device 2000 and a user. In some embodiments, the screen can include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors for sensing a touch, a slide and a gesture on the touch panel. The touch sensor can not only sense a boundary of a touching or sliding action, but also detect duration and pressure related to the touching or sliding action. In some embodiments, the multimedia component 2008 includes a front camera and / or a rear camera. When the device 2000 is in an operation mode, such as a camera mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front and rear camera can be a fixed optical lens system or have a focal length and optical zooming capability.

[0332] The audio component 2010 is configured to output and / or input audio signals. For example, the audio component 2010 includes a microphone (MIC) configured to receive an external audio signal when the device 2000 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 2004 or transmitted via the communication component 2016. In some embodiments, the audio component 2010 further includes a speaker for outputting audio signals.

[0333] The I / O interface 2012 provides an interface between the processing component 2002 and peripheral interface modules, such as a keypad, a click wheel, buttons, and so on. The buttons can include, but are not limited to, a home button, a volume button, a start button, and a lock button.

[0334] The sensor component 2014 includes one or more sensors configured to provide various state assessments for the device 2000. For example, the sensor component 2014 can detect an open / closed state of the device 2000, relative positioning of components, such as a display and a keypad of the device 2000, a change in position of the device 2000 or a component of the device 2000, presence or absence of user contact with the device 2000, an orientation or acceleration / deceleration of the device 2000, and a temperature change of the device 2000. The sensor component 2014 can include a proximity sensor configured to detect presence of a nearby object without any physical touch. The sensor component 2014 can further include a light sensor, such as a CMOS or CCD image sensor, for use in an imaging application. In some embodiments, the sensor component 2014 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0335] The communication component 2016 is configured to facilitate wired or wireless communication between the device 2000 and other devices. The device 2000 can access a wireless network based on a communication standard, such as WiFi, a cellular network (e.g., 2G, 3G, 4G, or 5G), or a combination thereof. In an exemplary embodiment, the communication component 2016 receives a broadcast signal or broadcast related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 2016 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) techniques, infrared data association (IrDA) techniques, ultra-wideband (UWB) techniques, Bluetooth (BT) techniques and other techniques.

[0336] In an exemplary embodiment, the device 2000 can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, micro-controllers, microprocessors or other electronic elements, for performing the above-described methods.

[0337] In an exemplary embodiment, a storage medium including instructions, such as the memory 2004 including instructions, is also provided, which can be executed by the processor 2020 of the device 2000 to complete the above-described methods. Optionally, the storage medium can be a non-transitory computer-readable storage medium, such as ROM, random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk and an optical data storage device, etc.

[0338] It should be noted that the above-mentioned sequence of the embodiments of the application is only for description, and does not represent the advantages and disadvantages of the embodiments. The above describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from the order in the embodiments and still achieve the desired result. In addition, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are possible or can be advantageous.

[0339] Each of the embodiments in the present specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiments.

[0340] Those skilled in the art can understand that all or part of the steps of the above-mentioned embodiments can be completed by hardware, or can be instructed to relevant hardware by program. The program can be stored in a computer readable storage medium, and the storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.

[0341] The above description is merely preferred embodiments of the present application, and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for training a recommendation system, characterized in that, include: Determine the initial account status data and initial account behavior data for the recommendation system; The initial recommendation object corresponding to the recommendation system is determined based on the initial account status data and the initial account behavior data; Based on the initial account status data, the initial account behavior data, and the initial recommendation object, the recommendation system determines the preset indicator information, transition account status data, and transition account behavior data for a preset time period. The transitional account status data is represented based on the improved data of the initial recommended object for the initial account status data; The transitional account behavior data representation is based on the improved data of the initial recommendation object for the initial account behavior data; The preset indicator information over the preset time period is transformed from sparse to dense to obtain the preset indicator information of the recommendation system at each time step within the preset time period. The recommendation system is trained based on the preset indicator information for each time step, the transition account status data, and the transition account behavior data to obtain a target recommendation system; the preset indicator information corresponding to the target recommendation system satisfies the preset indicator conditions. The step of transforming the preset indicator information over the preset time period from sparse to dense to obtain the preset indicator information of the recommendation system at each time step within the preset time period includes: The recommendation system status data corresponding to each time step is determined based on the account status data and account behavior data corresponding to each time step in the preset time period; The recommendation system behavior data corresponding to each time step is determined based on the recommended object corresponding to each time step in the preset time period; The preset time period, the preset indicator information within the preset time period, the recommendation system state data and recommendation system behavior data corresponding to each time step are input into the trained indicator information decomposer to obtain the preset indicator information of the recommendation system at each time step.

2. The recommendation system training method according to claim 1, characterized in that, The step of determining the preset indicator information, transitional account status data, and transitional account behavior data of the recommendation system within a preset time period based on the initial account status data, the initial account behavior data, and the initial recommendation object includes: The initial account status data, the initial account behavior data, and the initial recommendation object are input into a data simulator to obtain the first account status data, the first account behavior data, and the first time step of the recommendation system. Based on the first account status data and the first account behavior data, the first recommended object corresponding to the recommendation system is determined; the loop steps are: inputting the first account status data, the first account behavior data and the first recommended object into the data simulator to obtain the second account status data, the second account behavior data and the second time step of the recommendation system; until the preset indicator information, the transition account status data and the transition account behavior data of the recommendation system in the preset time period are obtained; The preset time period consists of multiple time steps in a loop, and the multiple time steps include the first time step and the second time step.

3. The recommendation system training method according to claim 1, characterized in that, The method further includes: Construct a raw information decomposer; Based on the preset time period and the preset indicator information on the preset time period, the reference indicator information corresponding to each time step is determined; The recommendation system state data and behavior data corresponding to each time step are input into the original information decomposer to obtain the prediction index information corresponding to each time step; The original information decomposer is trained based on the reference index information corresponding to each time step and the prediction index information corresponding to each time step. The index information decomposer is obtained when the iteration termination condition is met.

4. The recommendation system training method according to claim 3, characterized in that, The process of obtaining the index information decomposer under the condition of satisfying the iteration termination includes: If the difference between the predicted indicator information and the reference indicator information corresponding to any time step in each time step is less than or equal to a first preset difference, and the difference between the cumulative indicator information and the preset indicator information in the preset time period is less than or equal to a second preset difference, the training of the original information decomposer is terminated. The trained original information decomposer is determined as the index information decomposer.

5. The recommendation system training method according to claim 2, characterized in that, The method further includes: Obtain the sample dataset and the initial data for the original generator; The original generator and discriminator are trained based on the sample dataset and the initial data to obtain the target generator; the target generator includes a data simulator.

6. The recommendation system training method according to claim 5, characterized in that, The process of obtaining the sample dataset and the initial data of the original generator includes: Obtain the historical offline data of the recommendation system; The historical offline data is divided into multiple sample data based on the number of access rounds; the number of sample data is the same as the number of access rounds; each sample data includes sample account status data, sample account behavior data, and sample system behavior data. The initial data of the original generator is obtained by sampling the historical offline data of the recommendation system.

7. The recommendation system training method according to claim 6, characterized in that, The process of training the original generator and discriminator based on the sample dataset and the initial data to obtain the target generator includes: The sample data corresponding to the first access round and the initial data are input into the discriminator to obtain a first discrimination result of the sample data corresponding to the first access round and a second discrimination result of the initial data; The target loss of the discriminator is determined based on the annotation information of the sample data corresponding to the first access round, the first annotation information of the initial data corresponding to the discriminator, the first discrimination result, and the second discrimination result; The discriminator is trained based on the target loss of the discriminator; The target loss of the original generator is determined based on the second discrimination result and the first annotation information corresponding to the original generator in the initial data; The original generator is trained based on the target loss of the original generator; First generated data is generated based on the original generator and the initial data; The sample data corresponding to the second access round and the first generated data are input into the discriminator to obtain the first discrimination result of the sample data corresponding to the second access round and the second discrimination result of the first generated data; the loop step is to determine the target loss of the discriminator based on the annotation information of the sample data corresponding to the second access round, the first annotation information of the first generated data corresponding to the discriminator, the first discrimination result and the second discrimination result; The target generator is obtained when the iteration termination condition is met.

8. The recommendation system training method according to claim 7, characterized in that, The original generator includes an original simulator and an original recommender; the generation of first generated data based on the original generator and the initial data includes: Input the sample account status data and sample account behavior data from the initial data into the original recommender to obtain the first system behavior data; The first system behavior data, the sample account status data and the sample account behavior data in the initial data are input into the original simulator to obtain the first account status data and the first account behavior data. The first generated data is determined based on the first account status data, the first account behavior data, and the first system behavior data.

9. The recommendation system training method according to any one of claims 1-8, characterized in that, The step of determining the preset indicator information, transitional account status data, and transitional account behavior data of the recommendation system within a preset time period based on the initial account status data, the initial account behavior data, and the initial recommendation object includes: Based on the initial account status data, the initial account behavior data, and the initial recommendation object, the system determines the preset indicator information corresponding to the retention rate and / or account preference over a preset time period, as well as the transition account status data and the transition account behavior data.

10. A recommendation method, characterized in that, include: Obtain the target account's account status data and account execution data; The account status data and account execution data of the target account are input into the target recommendation system trained according to any one of the recommendation system training methods according to claims 1 to 9 to obtain the target recommendation object.

11. A recommendation system training device, characterized in that, include: The first data determination module is configured to determine the initial account status data and initial account behavior data of the recommendation system. The object determination module is configured to determine the initial recommendation object corresponding to the recommendation system based on the initial account status data and the initial account behavior data. The second data determination module is configured to determine, based on the initial account status data, the initial account behavior data, and the initial recommendation object, the preset indicator information, transition account status data, and transition account behavior data of the recommendation system within a preset time period; The transitional account status data is represented based on the improved data of the initial recommended object for the initial account status data; The transitional account behavior data representation is based on the improved data of the initial recommendation object for the initial account behavior data; The information conversion module is configured to perform a sparse-to-dense conversion on the preset indicator information over the preset time period to obtain the preset indicator information of the recommendation system at each time step within the preset time period. The training module is configured to train the recommendation system based on preset indicator information, transition account status data, and transition account behavior data for each time step to obtain a target recommendation system; the preset indicator information corresponding to the target recommendation system satisfies preset indicator conditions. The information conversion module is configured to execute: The recommendation system status data corresponding to each time step is determined based on the account status data and account behavior data corresponding to each time step in the preset time period; The recommendation system behavior data corresponding to each time step is determined based on the recommended object corresponding to each time step in the preset time period; The preset time period, the preset indicator information within the preset time period, the recommendation system state data and recommendation system behavior data corresponding to each time step are input into the trained indicator information decomposer to obtain the preset indicator information of the recommendation system at each time step.

12. A recommendation device, characterized in that, include: The data acquisition module is configured to acquire account status data and account execution data of the target account; The object recommendation module is configured to input the account status data and account execution data of the target account into the target recommendation system trained by the recommendation system training device according to claim 11, and obtain the target recommendation object.

13. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the recommendation system training method as described in any one of claims 1 to 9 or the recommendation method as described in claim 10.

14. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is able to perform the recommendation system training method as described in any one of claims 1 to 9 or the recommendation method as described in claim 10.

15. A computer program product, characterized in that, The computer program product includes a computer program stored in a readable storage medium, wherein at least one processor of a computer device reads from the readable storage medium and executes the computer program, causing the computer device to perform the recommendation system training method as claimed in any one of claims 1 to 9 or the recommendation method as claimed in claim 10.

Citation Information

Patent Citations

  • Click rate prediction method, recommendation method, click rate prediction model, click rate prediction device and equipment

    CN111046294A

  • Recommendation model training method and device, recommendation method and device, server and medium

    CN114462584A