Recommendation network model training method, device, electronic device and storage medium

By training recommended network models in virtual e-commerce platforms, and using the combination of virtual and real interactive behaviors, the problem of how to effectively verify the algorithm model and strategy of e-commerce platforms is solved, achieving improvements in accuracy and cost-effectiveness.

CN114841773BActive Publication Date: 2025-05-23JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210471495.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-28
Publication Date
2025-05-23
Estimated Expiration
2042-04-28

AI Technical Summary

Technical Problem

In e-commerce platforms, how to effectively verify the updated algorithm models and strategies, avoid the direct impact on users' browsing and purchasing behavior, and reduce trial and error costs.

Method used

A training method for recommending network models is proposed. By obtaining the virtual user characteristics, search keywords and real interaction behavior of sample users, the product sorting results and user behavior are determined based on the initial recommendation network model in the virtual e-commerce platform, the virtual interaction behavior is formed, and combined with the real interaction behavior, the initial recommendation network model is trained until it converges to the target recommendation network model.

Benefits of technology

It realizes efficient training of recommended network models in virtual e-commerce platforms to ensure the accuracy of verification results, reduces the impact on real e-commerce platforms, and reduces trial and error costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114841773B_ABST
    Figure CN114841773B_ABST
Patent Text Reader

Abstract

The present application proposes a training method for a recommendation network model, which includes: obtaining virtual user characteristics, search keywords and real interaction behaviors of sample users; determining the second commodity ranking result corresponding to the virtual user characteristics and search keywords according to the initial recommendation network model in the virtual e-commerce platform; determining the second user behavior of the sample user for the second commodity ranking result; forming a virtual interaction behavior according to the virtual user characteristics, the second commodity ranking result and the second user behavior; training the initial recommendation network model according to the virtual interaction behavior and the real interaction behavior to obtain a target recommendation network model. Thus, combined with the real interaction behavior of the sample user in the real e-commerce platform, the recommendation network model in the virtual e-commerce platform is trained, thereby ensuring the accuracy of the verification results when the virtual e-commerce platform is subsequently verified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a training method, device, electronic device and storage medium for a recommendation network model. Background Art

[0002] In order to better match users' personalized needs and improve their shopping experience and efficiency, the algorithm models and strategies within the e-commerce platform need to be updated and verified as necessary.

[0003] However, if the updated algorithm model and strategy are directly moved into the real e-commerce platform for verification, it will have a direct impact on the user's browsing and purchasing behavior, and the trial and error cost is very high. In the related art, the real e-commerce platform is usually simulated to obtain a virtual e-commerce platform, and the updated algorithm model and strategy are verified through the virtual e-commerce platform. In this case, how to train the recommendation network model in the virtual e-commerce platform is very important for verifying the accuracy of the algorithm model and strategy through the virtual e-commerce platform. Summary of the invention

[0004] The present application proposes a training method for a recommendation network model, an object testing method, an object verification method, a device, an electronic device and a storage medium.

[0005] On the one hand, an embodiment of the present application proposes a training method for a recommendation network model, the method comprising: obtaining virtual user characteristics, search keywords and real interaction behaviors of sample users; determining a second product ranking result corresponding to the virtual user characteristics and the search keywords according to an initial recommendation network model in a virtual e-commerce platform; determining a second user behavior of the sample user with respect to the second product ranking result; forming a virtual interaction behavior according to the virtual user characteristics, the second product ranking result and the second user behavior; and training the initial recommendation network model according to the virtual interaction behavior and the real interaction behavior to obtain a target recommendation network model.

[0006] In one embodiment of the present application, the initial recommendation network model is trained according to the virtual interaction behavior and the real interaction behavior to obtain a target recommendation network model, including:

[0007] Inputting the virtual interaction behavior into a discriminator network model to obtain a first probability value of the authenticity of the virtual interaction behavior;

[0008] Inputting the real interaction behavior into the discriminator network model to obtain a second probability value of the authenticity of the real interaction behavior;

[0009] The initial recommendation network model is trained according to the first probability value and the second probability value to obtain a target recommendation network model.

[0010] In one embodiment of the present application, the initial recommendation network model is trained according to the first probability value and the second probability value to obtain a target recommendation network model, including:

[0011] When the absolute value of the difference between the first probability value and the second probability value is greater than a preset threshold, inputting the virtual interaction behavior into a value network to obtain a value corresponding to the virtual interaction behavior;

[0012] According to the first probability value and the value, updating the model parameters of the initial recommendation network model until the absolute value of the difference between the probability value of the authenticity corresponding to the real interaction behavior and the probability value of the authenticity corresponding to the updated virtual interaction behavior is less than or equal to the preset threshold;

[0013] The recommended network model obtained when the absolute value of the difference is less than or equal to the preset threshold is used as the target recommended network model.

[0014] In one embodiment of the present application, updating the model parameters of the initial recommendation network model according to the first probability value and the value includes:

[0015] determining a corresponding advantage estimate value based on the first probability value and the value;

[0016] Obtaining an objective function corresponding to the initial recommendation network model;

[0017] Determining a corresponding objective function value according to the advantage estimate and the objective function;

[0018] According to the objective function value, the model parameters of the initial recommendation network model are updated.

[0019] In one embodiment of the present application, a proximal strategy optimization algorithm is used for the initial recommendation network model, and the obtaining of an objective function corresponding to the initial recommendation network model includes:

[0020] Obtaining an objective function of the proximal strategy optimization algorithm;

[0021] According to the objective function of the proximal strategy optimization algorithm, an objective function corresponding to the initial recommendation network model is determined.

[0022] The training method of the recommendation network model of the embodiment of the present application obtains the virtual user characteristics, search keywords and real interaction behaviors of the sample user; determines the second commodity ranking result corresponding to the virtual user characteristics and search keywords according to the initial recommendation network model in the virtual e-commerce platform; determines the second user behavior of the sample user for the second commodity ranking result; forms a virtual interaction behavior according to the virtual user characteristics, the second commodity ranking result and the second user behavior; trains the initial recommendation network model according to the virtual interaction behavior and the real interaction behavior to obtain the target recommendation network model. Thus, combined with the real interaction behavior of the sample user in the real e-commerce platform, the training of the recommendation network model in the virtual e-commerce platform is realized, thereby ensuring the accuracy of the verification results when the virtual e-commerce platform is subsequently verified.

[0023] Another aspect of the present application provides a method for testing an object, the method comprising:

[0024] Deploy a test object corresponding to the current object on the virtual e-commerce platform, wherein the test object is obtained by optimizing the current object, and the target recommendation network model in the virtual e-commerce platform is trained according to a training method for a recommendation network model;

[0025] An AB test is performed on the current object and the test object on the virtual e-commerce platform.

[0026] The object testing method of the embodiment of the present application can use a virtual e-commerce platform to test the test object to achieve the same effect as testing on a real e-commerce platform.

[0027] Another aspect of the present application provides a method for verifying an object, the method comprising:

[0028] Deploy multiple objects to be verified corresponding to the current object on the virtual e-commerce platform, wherein the objects to be verified are obtained by optimizing the current object, and the target recommendation network model in the virtual e-commerce platform is obtained by training according to a training method for a recommendation network model;

[0029] After the virtual e-commerce platform runs the plurality of objects to be verified for a specified time, determining a service quality result of each of the objects to be verified;

[0030] According to the service quality result, a target object with the best service quality result is selected from the multiple objects to be verified.

[0031] The object verification method of the embodiment of the present application can use a virtual e-commerce platform to verify the object to be verified, so as to achieve the same effect as verification on a real e-commerce platform.

[0032] In another aspect, an embodiment of the present application provides a training device for a virtual e-commerce platform, the device comprising:

[0033] An acquisition module is used to acquire virtual user characteristics, search keywords and real interaction behaviors of sample users, wherein the real interaction behaviors are formed according to the real user characteristics of the sample users, the first commodity ranking results and the first user behaviors, the first commodity ranking results are obtained by searching on a real e-commerce platform based on the real user characteristics and the search keywords, and the first user behaviors are user behaviors generated by the sample users on the first commodity ranking results;

[0034] A first determination module is used to determine a second commodity ranking result corresponding to the virtual user feature and the search keyword according to an initial recommendation network model in a virtual e-commerce platform, wherein the virtual e-commerce platform is obtained by simulating the real e-commerce platform;

[0035] A second determination module, configured to determine a second user behavior of the sample user with respect to the second commodity ranking result;

[0036] A forming module, configured to form a virtual interactive behavior according to the virtual user characteristics, the second commodity ranking result and the second user behavior;

[0037] A training module is used to train the initial recommendation network model according to the virtual interaction behavior and the real interaction behavior to obtain a target recommendation network model.

[0038] In one embodiment of the present application, the training module includes:

[0039] A first input unit, used to input the virtual interaction behavior into a discriminator network model to obtain a first probability value of the authenticity of the virtual interaction behavior;

[0040] A second input unit, used for inputting the real interaction behavior into the discriminator network model to obtain a second probability value of the authenticity of the real interaction behavior;

[0041] A training unit is used to train the initial recommendation network model according to the first probability value and the second probability value to obtain a target recommendation network model.

[0042] In one embodiment of the present application, the training unit is specifically used for:

[0043] When the absolute value of the difference between the first probability value and the second probability value is greater than a preset threshold, inputting the virtual interaction behavior into a value network to obtain a value corresponding to the virtual interaction behavior;

[0044] According to the first probability value and the value, updating the model parameters of the initial recommendation network model until the absolute value of the difference between the probability value of the authenticity corresponding to the real interaction behavior and the probability value of the authenticity corresponding to the updated virtual interaction behavior is less than or equal to the preset threshold;

[0045] The recommended network model obtained when the absolute value of the difference is less than or equal to the preset threshold is used as the target recommended network model.

[0046] In one embodiment of the present application, the training unit updates the model parameters of the initial recommendation network model according to the first probability value and the value, including:

[0047] determining a corresponding advantage estimate value based on the first probability value and the value;

[0048] Obtaining an objective function corresponding to the initial recommendation network model;

[0049] Determining a corresponding objective function value according to the advantage estimate and the objective function;

[0050] According to the objective function value, model parameters of the recommendation network model and the user behavior prediction model are updated respectively.

[0051] In one embodiment of the present application, a proximal strategy optimization algorithm is used for the initial recommendation network model, and the training unit obtains an objective function corresponding to the initial recommendation network model, including:

[0052] Obtaining an objective function of the proximal strategy optimization algorithm;

[0053] According to the objective function of the proximal strategy optimization algorithm, an objective function corresponding to the initial recommendation network model is determined.

[0054] The training device of the recommendation network model of the embodiment of the present application obtains the virtual user characteristics, search keywords and real interaction behaviors of the sample user; determines the second commodity ranking result corresponding to the virtual user characteristics and search keywords according to the initial recommendation network model in the virtual e-commerce platform; determines the second user behavior of the sample user for the second commodity ranking result; forms a virtual interaction behavior according to the virtual user characteristics, the second commodity ranking result and the second user behavior; trains the initial recommendation network model according to the virtual interaction behavior and the real interaction behavior to obtain the target recommendation network model. Thus, combined with the real interaction behavior of the sample user in the real e-commerce platform, the training of the recommendation network model in the virtual e-commerce platform is realized, thereby ensuring the accuracy of the verification results when the virtual e-commerce platform is subsequently verified.

[0055] In another aspect of the present application, an embodiment provides a testing device for an object, the device comprising:

[0056] A first deployment module is used to deploy a test object corresponding to the current object on a virtual e-commerce platform, wherein the test object is obtained by optimizing the current object, and the virtual e-commerce platform is trained according to a training method for a recommendation network model;

[0057] The testing module is used to perform AB testing on the current object and the test object on the virtual e-commerce platform.

[0058] The testing device of the object of the embodiment of the present application can use a virtual e-commerce platform to test the test object to achieve the same effect as testing on a real e-commerce platform.

[0059] In another aspect of the present application, an object verification device is provided, the device comprising:

[0060] In the second deployment module, the user deploys multiple objects to be verified corresponding to the current object on the virtual e-commerce platform, wherein the objects to be verified are obtained by optimizing the current object, and the virtual e-commerce platform is trained according to a training method for a recommendation network model;

[0061] A third determination module, after the virtual e-commerce platform runs the plurality of objects to be verified for a specified time, determines a service quality result of each of the objects to be verified;

[0062] A selection module is used to select a target object with the best service quality result from a plurality of objects to be verified according to the service quality result.

[0063] The object verification device of the embodiment of the present application can use a virtual e-commerce platform to verify the object to be verified, so as to achieve the same effect as verification on a real e-commerce platform.

[0064] On the other hand, an embodiment of the present application proposes an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements any of the above-mentioned recommended network model training methods, object testing methods, or object verification methods of the embodiments of the present application.

[0065] On the other hand, an embodiment of the present application proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-mentioned recommended network model training methods, object testing methods, or object verification methods of the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present application.

[0067] Figure 1 is a flowchart of a training method for a recommendation network model according to an embodiment of the present application;

[0068] Figure 2 is a flowchart of a training method for a recommendation network model according to another embodiment of the present application;

[0069] Figure 3 is a flowchart of a training method for a recommendation network model according to another embodiment of the present application;

[0070] Figure 4 is a flowchart of a training method for a recommendation network model according to another embodiment of the present application;

[0071] Figure 5 is a flow chart of a method for testing an object according to an embodiment of the present application;

[0072] Figure 6 is a flowchart of a method for verifying an object according to an embodiment of the present application;

[0073] Figure 7 is a structural diagram of a training device for a recommendation network model according to an embodiment of the present application;

[0074] Figure 8 is a structural schematic diagram of a training device for a recommendation network model according to another embodiment of the present application;

[0075] Fig. 9 is a schematic diagram of the structure of a testing device for an object according to an embodiment of the present application;

[0076] Fig.10 is a schematic diagram of the structure of an object verification device according to an embodiment of the present application;

[0077] Fig.11 It is a schematic diagram of the structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0078] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be construed as limiting the present invention.

[0079] The following describes the training method of the recommendation network model, the object testing method, the object verification method, the device, the electronic device and the storage medium of the embodiments of the present application with reference to the accompanying drawings.

[0080] Figure 1 It is a flowchart of a training method for a recommendation network model according to an embodiment of the present application.

[0081] Among them, it should be noted that the training method of the recommended network model provided in this embodiment is applied to the training device of the recommended network model. The training device of the recommended network model can be implemented by software and / or hardware. The training device of the recommended network model can be an electronic device or can be configured in an electronic device. The electronic device in this embodiment can be a PC (Personal Computer), a mobile device, a tablet computer, a terminal device or a server and other devices, which are not specifically limited here.

[0082] like Figure 1 As shown, the training method of the recommendation network model may include:

[0083] Step 101, obtaining virtual user features, search keywords and real interaction behaviors of sample users.

[0084] The real interactive behavior is formed based on the real user characteristics of the sample user, the first product ranking result and the first user behavior.

[0085] Among them, the first product ranking result is obtained by searching on a real e-commerce platform based on real user characteristics and search keywords.

[0086] The first user behavior is the user behavior of the sample user on the first product ranking result in the real e-commerce platform, wherein the first user behavior may include but is not limited to click, browse, purchase, favorite, comment, etc. It is understandable that the first user behavior may also include the time the sample user browses certain products, whether certain products are removed from the first product ranking result, etc.

[0087] The above-mentioned real user characteristics are determined by analyzing the historical behavior data of the sample users on the real e-commerce platform. For example, based on the historical behavior data, the size, style and color of the coats purchased by the sample users multiple times are analyzed to analyze the relevant characteristic data of the sample users.

[0088] In one embodiment of the present application, a possible implementation method for obtaining the first product ranking result is: searching according to the search keyword to obtain multiple products, and determining the recommendation values ​​of the multiple products according to the real user characteristics, and arranging the multiple products in descending order of the recommendation values ​​to obtain the first product ranking result.

[0089] In one embodiment of the present application, another possible implementation method for obtaining the first product ranking result is: searching for multiple products based on real user characteristics and search keywords, and randomly sorting the multiple products to obtain the first product ranking result.

[0090] It should be noted that the virtual user features of the sample users are randomly generated based on the generated network.

[0091] In one embodiment of the present application, a possible implementation method for obtaining the search keywords of the sample user is: extracting keywords based on the words entered by the sample user when searching for goods on a real e-commerce platform. For example, if the word entered by the sample user when searching for goods is "children's toys", the search keywords may be "children" and "toys".

[0092] In one embodiment of the present application, another possible implementation method of obtaining the search keywords of the sample user is: obtaining them according to the user characteristics of the sample user corresponding to the preset search keywords. For example, when the user characteristic of the sample user is "age 60 years old", the preset search keywords "calcium tablets", "blood pressure monitor" and the like corresponding to "age 60 years old" can be obtained.

[0093] It can be understood that, in this embodiment, the search keywords corresponding to the virtual user features and the real user features of the sample users are the same.

[0094] Step 102: Determine the second product ranking result corresponding to the virtual user characteristics and the search keyword according to the initial recommendation network model in the virtual e-commerce platform.

[0095] In one embodiment of the present application, an initial recommendation network model is included in the virtual e-commerce platform, wherein the initial recommendation network model can recommend commodities to sample users.

[0096] Specifically, the initial recommendation network model can recommend products to sample users based on virtual user features and search keywords, and the order of the recommended products is the second product ranking result.

[0097] Among them, the virtual e-commerce platform is obtained by simulating the real e-commerce platform.

[0098] Next, we can introduce the initial recommendation network model in the virtual e-commerce platform with a formula as follows:

[0099] action engine =f 1 (u_feature)

[0100] Among them, f 1 (u_feature) represents the initial recommendation network model in the virtual e-commerce platform, which is input by the virtual user features and search keywords of sample users. engine As output, the second product ranking result.

[0101] For example, the initial recommendation network model f 1 (u_feature) Input virtual user feature is "age is 60 years old", search keyword is "calcium tablets", then the second product sorting result action engine Calcium tablets suitable for people over 60 are available in different brands and prices.

[0102] Step 103: determining a second user behavior of the sample user with respect to the second commodity ranking result.

[0103] The second user behavior of the sample user with respect to the second commodity ranking result may include clicking, browsing, collecting, purchasing, commenting, etc.

[0104] Among them, you can use action engine Indicates the sorting result of the second product, action user Indicates the second user behavior.

[0105] In some exemplary embodiments, the virtual user features and the second commodity ranking results may be input into a pre-trained user behavior prediction model to obtain the second user behavior of the sample user with respect to the second commodity ranking results.

[0106] The user behavior prediction model in this embodiment is pre-trained based on the real user characteristics of the sample user, the first product ranking result, and the first user behavior. In some exemplary implementations, the exemplary process of training the user behavior prediction model may be: the real user characteristics of the sample user and the first product ranking result may be used as inputs of the user behavior prediction model, and the first user behavior may be used as outputs of the user behavior prediction model to train the user behavior prediction model.

[0107] Continuing with the above example, the sample user gets the second product ranking result of calcium tablets. engineAfter that, users can browse calcium tablets of different brands and prices, click to learn more, comment on them, or buy calcium tablets. These behaviors constitute the second user action. user .

[0108] Step 104 , forming a virtual interactive behavior according to the virtual user characteristics, the second commodity ranking result, and the second user behavior.

[0109] In one embodiment of the present application, the virtual interaction behavior is composed of the virtual user characteristics, the second commodity ranking result, and the second user behavior, which can also be expressed by a formula as shown below:

[0110] fake tra =(u feature ,action engine ,action user )

[0111] Among them, fake tra represents a virtual interaction behavior, u_feature represents a virtual user feature, action engine Indicates the sorting result of the second product, action user Indicates the second user behavior.

[0112] Step 105 , training the initial recommendation network model according to the virtual interaction behavior and the real interaction behavior to obtain a target recommendation network model.

[0113] In one embodiment of the present application, the initial recommendation network model may be trained according to the virtual interaction behavior and the real interaction behavior until the initial recommendation network model is trained and converges to the target recommendation network model.

[0114] The training method of the recommendation network model of the embodiment of the present application obtains the virtual user characteristics, search keywords and real interaction behaviors of the sample user; determines the second commodity ranking result corresponding to the virtual user characteristics and search keywords according to the initial recommendation network model in the virtual e-commerce platform; determines the second user behavior of the sample user for the second commodity ranking result; forms a virtual interaction behavior according to the virtual user characteristics, the second commodity ranking result and the second user behavior; trains the initial recommendation network model according to the virtual interaction behavior and the real interaction behavior to obtain the target recommendation network model. Thus, combined with the real interaction behavior of the sample user in the real e-commerce platform, the training of the recommendation network model in the virtual e-commerce platform is realized, thereby ensuring the accuracy of the verification results when the virtual e-commerce platform is subsequently verified.

[0115] In one embodiment of the present application, step 105 trains the initial recommendation network model according to the virtual interaction behavior and the real interaction behavior to obtain a possible implementation of the target recommendation network model, such as Figure 2 As shown, including:

[0116] Step 201: input the virtual interaction behavior into the discriminator network model to obtain a first probability value of the authenticity of the virtual interaction behavior.

[0117] In order to more accurately determine whether the training of the recommendation network model is completed, the virtual interaction behavior can be input into the discriminator network model to obtain a first probability value of the authenticity of the virtual interaction behavior.

[0118] Specifically, the function of the discriminator network model of virtual interaction behavior is as follows:

[0119] p fake =D(real tra )

[0120] Among them, p fake represents the probability value of the discriminator network model judging the virtual interaction behavior as true. For example, p fake It can be 0.3, which means that the probability value of the discriminator network model judging the virtual interaction behavior as true is 0.3.

[0121] Step 202: input the real interaction behavior into the discriminator network model to obtain a second probability value of the authenticity of the real interaction behavior.

[0122] In order to more accurately determine whether the training of the recommendation network model is completed, the real interaction behavior can be input into the discriminator network model to obtain a second probability value of the authenticity of the real interaction behavior, and use it as a reference.

[0123] Specifically, the function of the discriminator network model of the real interaction behavior is as follows:

[0124] p real =D(real tra )

[0125] Among them, p real represents the probability value of the discriminator network model judging the real interaction behavior as true. For example, p real It can be 0.7, which means that the probability value of the discriminator network model distinguishing the real interaction behavior as true is 0.7.

[0126] Step 203: training the initial recommendation network model according to the first probability value and the second probability value to obtain a target recommendation network model.

[0127] In one embodiment of the present application, whether the initial recommendation network model converges to the target recommendation network model may be determined based on the absolute value relationship of the difference between the first probability value and the second probability value obtained above.

[0128] In one embodiment of the present application, based on any of the above embodiments, in step 203, the initial recommendation network model is trained according to the first probability value and the second probability value to obtain a possible implementation of the target recommendation network model, such as Figure 3 As shown, including:

[0129] Step 301, when the absolute value of the difference between the first probability value and the second probability value is greater than a preset threshold, the virtual interaction behavior is input into the value network to obtain the value corresponding to the virtual interaction behavior.

[0130] In one embodiment of the present application, when the absolute value of the difference between the first probability value and the second probability value is greater than a preset threshold, it indicates that the initial recommendation network model at this time has not converged to the target recommendation network model.

[0131] In order to determine whether the virtual interaction behavior at this time has available value, the virtual interaction behavior can be input into the value network to obtain the value corresponding to the virtual interaction behavior.

[0132] The value network formula is as follows:

[0133] value=V(fake tra )

[0134] Among them, V(fake tra ) represents the value network function, and value is the calculated value.

[0135] It should be noted that the preset threshold is a critical value of the absolute value of the difference between the first probability value and the second probability value set in advance. It can be understood that in actual applications, the value of the preset threshold can be set according to actual application requirements, and this embodiment does not specifically limit this.

[0136] Step 302, based on the first probability value and the value, the model parameters of the initial recommendation network model are updated until the absolute value of the difference between the probability value of the authenticity corresponding to the real interaction behavior and the probability value of the authenticity corresponding to the updated virtual interaction behavior is less than or equal to a preset threshold.

[0137] In one embodiment of the present application, when the absolute value of the difference between the first probability value and the second probability value is less than or equal to a preset threshold, it indicates that the initial recommendation network model at this time has converged to the target recommendation network model, that is, the initial recommendation network model at this time is infinitely close to or equal to the target recommendation network model.

[0138] When the above purpose has not been achieved, it is necessary to update the model parameters of the initial recommendation network model according to the first probability value and value to obtain the updated virtual interaction behavior, and then train the updated recommendation network model according to the updated virtual interaction behavior and the real interaction behavior until the absolute value of the difference between the probability value of the authenticity corresponding to the real interaction behavior and the probability value of the authenticity corresponding to the updated virtual interaction behavior is less than or equal to the preset threshold.

[0139] Among them, a possible implementation method of updating the model parameters of the initial recommendation network model according to the first probability value and the value is: using a proximal policy optimization algorithm (PPO) to complete the model parameter update of the algorithm model.

[0140] Among them, the updated virtual interaction behavior is determined as follows:

[0141] Based on the updated recommendation network model, the virtual user characteristics and search keywords are processed to obtain a third product ranking result; a third user behavior generated by the sample user for the third product ranking result is determined; and an updated virtual interactive behavior is determined based on the virtual user characteristics, the third product ranking result and the third user behavior.

[0142] Therefore, it can be understood that the third product ranking result and the third user behavior are determined in the same manner as the second product ranking result and the second user behavior. However, since the model parameters of the initial recommendation network model are changed, the third product ranking result obtained is different from the second product ranking result, and the third user behavior is different from the second user behavior. Therefore, the updated virtual interaction behavior can be determined according to the virtual user characteristics, the third product ranking result and the third user behavior.

[0143] The updated virtual interaction behavior can be input into the discriminator network model again to obtain the third probability value of the authenticity of the virtual interaction behavior, and then the relationship between the absolute value of the difference between the third probability value of the authenticity of the virtual interaction behavior and the second probability value of the authenticity of the real interaction behavior and the preset threshold is judged again: if the absolute value of the difference between the third probability value of the authenticity of the virtual interaction behavior and the second probability value of the authenticity of the real interaction behavior at this time is less than or equal to the preset threshold, it indicates that the recommended network model at this time can replace the target recommended network model; if the absolute value of the difference between the third probability value of the authenticity of the virtual interaction behavior and the second probability value of the authenticity of the real interaction behavior at this time is greater than the preset threshold, it indicates that the recommended network model at this time has not yet replaced the target recommended network model, then the updated virtual interaction behavior will continue to be input into the value network to obtain the updated value corresponding to the updated virtual interaction behavior, and the model parameters of the recommended network model at this time will be updated again according to the third probability value and the updated value.

[0144] This method is repeated until the absolute value of the difference between the probability value of authenticity corresponding to the real interaction behavior and the probability value of authenticity corresponding to the updated virtual interaction behavior is less than or equal to the preset threshold.

[0145] Step 303: The recommended network model obtained when the absolute value of the difference is less than or equal to the preset threshold is used as the target recommended network model.

[0146] In one embodiment of the present application, based on any of the above embodiments, a possible implementation method of updating the model parameters of the initial recommendation network model according to the first probability value and the value in step 302 is as follows: Figure 4 As shown, including:

[0147] Step 401, determining a corresponding advantage estimate value according to a first probability value and a value.

[0148] In one embodiment of the present application, the advantage estimate value can be obtained using a generalized advantage estimate function, such as the following formula:

[0149] adv=GAE(p fake ,value,γ,λ)

[0150]

[0151] Among them, GAE (p fake,value,γ,λ) is the obtained advantage estimation value, γ and λ are constants, and both are set to 0.95 during the training process. Among them, γ∈[0,1] represents the discount factor, which is used to reflect the impact of time delay on the estimated value of the generalization advantage estimation function. λ∈[0,1] represents the hyperparameter. Reasonable adjustment of the value of λ can effectively balance the variance and bias of the generalization advantage estimation function.

[0152] Step 402: Obtain an objective function corresponding to the initial recommendation network model.

[0153] In one embodiment of the present application, the objective function corresponding to the initial recommendation network model can be determined based on the determined advantage estimation value, with the advantage estimation value as the sample user feature and the search keyword.

[0154] Step 403: Determine the corresponding objective function value according to the advantage estimate and the objective function.

[0155] In one embodiment of the present application, a new objective function value may be obtained based on the objective function of the initial recommendation network model, wherein the new objective function value corresponds to the third commodity ranking result.

[0156] Step 404: Update the model parameters of the initial recommendation network model according to the objective function value.

[0157] In one embodiment of the present application, a proximal strategy optimization algorithm may also be used for the initial recommendation network model. Based on any of the above embodiments, a possible implementation method of obtaining the objective function corresponding to the initial recommendation network model in step 402 includes:

[0158] Obtain an objective function of the proximal policy optimization algorithm; and determine an objective function corresponding to the initial recommendation network model based on the objective function of the proximal policy optimization algorithm.

[0159] Among them, the proximal policy optimization algorithm (PPO) can complete the model parameter update of the algorithm model.

[0160] The training method of the recommendation network model of the embodiment of the present application is composed of virtual user characteristics, second commodity ranking results and second user behavior to form a virtual interactive behavior. The initial recommendation network model is trained through virtual interactive behavior and real interactive behavior until it converges to the target recommendation network model. In the process of training the initial recommendation network model, it is also accompanied by the value calculation of the virtual interactive behavior and the real-time update of the initial recommendation network model using the proximal strategy optimization algorithm. Therefore, combined with the real interactive behavior of sample users in the real e-commerce platform, the recommendation network model in the virtual e-commerce platform is trained, thereby ensuring the accuracy of the verification results when the virtual e-commerce platform is subsequently verified.

[0161] The present application also provides a method for testing an object. Figure 5 As shown, the method includes:

[0162] Step 501: deploy a test object corresponding to the current object on the virtual e-commerce platform.

[0163] The test object is obtained by optimizing the current object, and the target recommendation network model in the virtual e-commerce platform is trained according to the training method of the above recommendation network model.

[0164] It is understandable that the test object may be one or more algorithm models and logic strategies optimized in the training method of the recommendation network model.

[0165] Step 502: Perform AB testing on the current object and the test object on the virtual e-commerce platform.

[0166] Specifically, the current object and the test object can be tested multiple times on the virtual e-commerce platform, that is, multiple tests of one or more algorithm models and logic strategies can be completed. Through multiple tests of one or more algorithm models and logic strategies, it can be tested whether the tested algorithm or logic is applicable to the virtual e-commerce platform.

[0167] The testing method of the object of the embodiment of the present application can achieve the same results as those tested on a real e-commerce platform by testing the algorithm model and logical strategy on a virtual e-commerce platform, thereby greatly reducing the risks and costs of testing on a real e-commerce platform without affecting the normal operation of the real e-commerce platform.

[0168] The present application also proposes a method for verifying an object. Figure 6 As shown, the method includes:

[0169] Step 601: deploy multiple objects to be verified corresponding to the current object on the virtual e-commerce platform.

[0170] The object to be verified is obtained by optimizing the current object, and the target recommendation network model in the virtual e-commerce platform is trained according to the training method of the recommendation network model.

[0171] It is understandable that the object to be verified may be an algorithm model and a logic strategy optimized in the training method of the recommendation network model.

[0172] Step 602 , after the virtual e-commerce platform runs multiple objects to be verified for a specified time, determines the service quality results of each object to be verified.

[0173] Specifically, a plurality of objects to be verified may be input into the virtual electronic commodity platform, and the objects may be run in the virtual electronic commodity platform for a specified time, and the service quality result of each object to be verified may be obtained.

[0174] The above-mentioned specified time is a fixed period of time set in advance.

[0175] Step 603: Select a target object with the best service quality result from multiple objects to be verified according to the service quality result.

[0176] For example, the service quality results of multiple objects to be verified may be arranged in descending order, and the object to be verified corresponding to the service quality result ranked first is taken as the target object.

[0177] The object verification method of the embodiment of the present application can achieve the same result as verification on a real e-commerce platform by verifying the algorithm model and logical strategy on a virtual e-commerce platform, thereby greatly reducing the risk and cost of verification on a real e-commerce platform without affecting the normal operation of the real e-commerce platform.

[0178] Figure 7 700 is a schematic diagram of a training device for a recommendation network model according to an embodiment of the present application. The device 700 includes: an acquisition module 710, a first determination module 720, a second determination module 730, a formation module 740 and a training module 750. Among them:

[0179] Acquisition module 710 is used to obtain virtual user characteristics, search keywords and real interactive behaviors of sample users, wherein the real interactive behaviors are formed based on the real user characteristics of the sample users, the first product ranking results and the first user behaviors, the first product ranking results are obtained by searching on a real e-commerce platform based on the real user characteristics and search keywords, and the first user behaviors are the user behaviors generated by the sample users on the first product ranking results.

[0180] The first determination module 720 is used to determine the second product ranking result corresponding to the virtual user characteristics and the search keyword according to the initial recommendation network model in the virtual e-commerce platform, wherein the virtual e-commerce platform is obtained by simulating a real e-commerce platform.

[0181] The second determination module 730 is used to determine a second user behavior of the sample user with respect to the second commodity ranking result.

[0182] The forming module 740 is used to form a virtual interactive behavior according to the virtual user characteristics, the second commodity ranking result and the second user behavior.

[0183] The training module 750 is used to train the initial recommendation network model according to the virtual interaction behavior and the real interaction behavior to obtain the target recommendation network model.

[0184] The training device of the recommendation network model of the embodiment of the present application obtains the virtual user characteristics, search keywords and real interaction behaviors of the sample user; determines the second commodity ranking result corresponding to the virtual user characteristics and search keywords according to the initial recommendation network model in the virtual e-commerce platform; determines the second user behavior of the sample user for the second commodity ranking result; forms a virtual interaction behavior according to the virtual user characteristics, the second commodity ranking result and the second user behavior; trains the initial recommendation network model according to the virtual interaction behavior and the real interaction behavior to obtain the target recommendation network model. Thus, combined with the real interaction behavior of the sample user in the real e-commerce platform, the training of the recommendation network model in the virtual e-commerce platform is realized, thereby ensuring the accuracy of the verification results when the virtual e-commerce platform is subsequently verified.

[0185] Figure 8 FIG. 1 is a schematic diagram of a training device for a recommendation network model according to another embodiment of the present application. Figure 8 As shown, the device 800 further includes: a first input unit 851, a second input unit 852 and a training unit 853. Among them:

[0186] The first input unit 851 is used to input the virtual interaction behavior into the discriminator network model to obtain a first probability value of the authenticity of the virtual interaction behavior.

[0187] The second input unit 852 is used to input the real interaction behavior into the discriminator network model to obtain a second probability value of the authenticity of the real interaction behavior.

[0188] The training unit 853 is used to train the initial recommendation network model according to the first probability value and the second probability value to obtain a target recommendation network model.

[0189] In one embodiment of the present application, the training unit 853 is further configured to:

[0190] When the absolute value of the difference between the first probability value and the second probability value is greater than a preset threshold, the virtual interaction behavior is input into the value network to obtain the value corresponding to the virtual interaction behavior;

[0191] According to the first probability value and the value, the model parameters of the initial recommendation network model are updated until the absolute value of the difference between the probability value of the authenticity corresponding to the real interaction behavior and the probability value of the authenticity corresponding to the updated virtual interaction behavior is less than or equal to a preset threshold;

[0192] The recommended network model obtained when the absolute value of the difference is less than or equal to the preset threshold is used as the target recommended network model.

[0193] In one embodiment of the present application, the training unit 853 updates the model parameters of the initial recommendation network model according to the first probability value and the value, including:

[0194] determining a corresponding advantage estimate based on the first probability value and the value;

[0195] Obtaining the objective function corresponding to the initial recommendation network model;

[0196] According to the advantage estimate and the objective function, the corresponding objective function value is determined;

[0197] According to the objective function value, the model parameters of the initial recommendation network model are updated respectively.

[0198] In one embodiment of the present application, a proximal strategy optimization algorithm is used for the initial recommendation network model, and the training unit 853 obtains an objective function corresponding to the initial recommendation network model, including:

[0199] Obtain the objective function of the proximal policy optimization algorithm;

[0200] According to the objective function of the proximal strategy optimization algorithm, the objective function corresponding to the initial recommendation network model is determined.

[0201] The training device of the recommendation network model of the embodiment of the present application is composed of virtual user characteristics, second commodity ranking results and second user behavior to form a virtual interactive behavior. The initial recommendation network model is trained through virtual interactive behavior and real interactive behavior until it converges to the target recommendation network model. In the process of training the initial recommendation network model, it is also accompanied by the value calculation of the virtual interactive behavior and the real-time update of the initial recommendation network model using the proximal strategy optimization algorithm. Therefore, combined with the real interactive behavior of sample users in the real e-commerce platform, the recommendation network model in the virtual e-commerce platform is trained, thereby ensuring the accuracy of the verification results when the virtual e-commerce platform is subsequently verified.

[0202] Fig. 9 1 is a schematic diagram of a test device for an object according to an embodiment of the present application. Fig. 9 As shown, the device 900 includes: a first deployment module 910 and a test module 920. Among them:

[0203] A first deployment module 910 is used to deploy a test object corresponding to the current object on the virtual e-commerce platform, wherein the test object is obtained by optimizing the current object, and the target recommendation network model in the virtual e-commerce platform is trained according to a training method for a recommendation network model;

[0204] The testing module 920 is used to perform AB testing on the current object and the test object on the virtual e-commerce platform.

[0205] The testing device of the object of the embodiment of the present application can achieve the same results as those tested on a real e-commerce platform by testing the algorithm model and logic strategy on a virtual e-commerce platform, thereby greatly reducing the risks and costs of testing on a real e-commerce platform without affecting the normal operation of the real e-commerce platform.

[0206] Fig.10 FIG. 1 is a schematic diagram of a structure of an object verification device according to an embodiment of the present application. Fig.10 As shown, the device 1000 includes: a second deployment module 1010, a third determination module 1020 and a selection module 1030. Among them:

[0207] The second deployment module 1010 is used to deploy multiple objects to be verified corresponding to the current object on the virtual e-commerce platform, wherein the objects to be verified are obtained by optimizing the current object, and the target recommendation network model in the virtual e-commerce platform is trained according to a training method for a recommendation network model;

[0208] The third determination module 1020 is used to determine the service quality result of each object to be verified after the virtual e-commerce platform runs the multiple objects to be verified for a specified time;

[0209] The selection module 1030 is used to select a target object with the best service quality result from multiple objects to be verified according to the service quality result.

[0210] The verification device of the object of the embodiment of the present application can achieve the same result as verification on a real e-commerce platform by verifying the algorithm model and logical strategy on a virtual e-commerce platform, thereby greatly reducing the risk and cost of verification on a real e-commerce platform without affecting the normal operation of the real e-commerce platform.

[0211] According to an embodiment of the present application, the present application also provides an electronic device and a readable storage medium.

[0212] Fig.11 It is a structural block diagram of an electronic device according to an embodiment of the present application.

[0213] like Fig.11 As shown, the electronic device 1100 includes: a memory 1110, a processor 1120, and computer instructions stored in the memory 1110 and executable on the processor 1120.

[0214] When the processor 1120 executes the instruction, it implements any recommended network model training method, object testing method or object verification method provided in the above embodiments.

[0215] Furthermore, the electronic device 1100 further includes:

[0216] The communication interface 1130 is used for communication between the memory 1110 and the processor 1120 .

[0217] The memory 1110 is used to store computer instructions that can be executed on the processor 1120 .

[0218] The memory 1110 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0219] The processor 1120 is used to implement any of the recommended network model training methods, object testing methods, or object verification methods of the above embodiments when executing the program.

[0220] If the memory 1110, the processor 1120 and the communication interface 1130 are implemented independently, the communication interface 1130, the memory 1110 and the processor 1120 can be connected to each other through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.11 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0221] Optionally, in a specific implementation, if the memory 1110, the processor 1120 and the communication interface 1130 are integrated on a chip, the memory 1110, the processor 1120 and the communication interface 1130 can communicate with each other through an internal interface.

[0222] The processor 1120 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0223] On the other hand, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any recommended network model training method, object testing method, or object verification method of the embodiments of the present application.

[0224] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0225] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.

Claims

1. A training method for a recommendation network model, It is characterized in that The method comprises: Acquire virtual user characteristics, search keywords, and real interaction behaviors of the sample user, wherein the real interaction behaviors are formed according to the real user characteristics of the sample user, the first commodity ranking result, and the first user behavior, the first commodity ranking result is obtained by searching on a real e-commerce platform based on the real user characteristics and the search keywords, and the first user behavior is the user behavior generated by the sample user on the first commodity ranking result; Determining a second product ranking result corresponding to the virtual user feature and the search keyword according to an initial recommendation network model in a virtual e-commerce platform, wherein the virtual e-commerce platform is obtained by simulating the real e-commerce platform; Determining a second user behavior of the sample user with respect to the second commodity ranking result; forming a virtual interactive behavior according to the virtual user characteristics, the second commodity ranking result, and the second user behavior; Inputting the virtual interaction behavior into a discriminator network model to obtain a first probability value of the authenticity of the virtual interaction behavior; Inputting the real interaction behavior into the discriminator network model to obtain a second probability value of the authenticity of the real interaction behavior; When the absolute value of the difference between the first probability value and the second probability value is greater than a preset threshold, inputting the virtual interaction behavior into a value network to obtain a value corresponding to the virtual interaction behavior; According to the first probability value and the value, updating the model parameters of the initial recommendation network model until the absolute value of the difference between the probability value of the authenticity corresponding to the real interaction behavior and the probability value of the authenticity corresponding to the updated virtual interaction behavior is less than or equal to the preset threshold; The recommended network model obtained when the absolute value of the difference is less than or equal to the preset threshold is used as the target recommended network model.

2. The method according to claim 1, It is characterized in that The updating of the model parameters of the initial recommendation network model according to the first probability value and the value includes: Substituting the first probability value and the value into a generalized advantage estimation function to obtain a corresponding advantage estimation value; Obtaining an objective function corresponding to the initial recommendation network model; Determining a corresponding objective function value according to the advantage estimate and the objective function; According to the objective function value, the model parameters of the initial recommendation network model are updated.

3. The method according to claim 2, It is characterized in that The initial recommendation network model is optimized by a proximal strategy algorithm, and the obtaining of an objective function corresponding to the initial recommendation network model includes: Obtaining an objective function of the proximal strategy optimization algorithm; According to the objective function of the proximal strategy optimization algorithm, an objective function corresponding to the initial recommendation network model is determined.

4. A test method for an object, It is characterized in that include: Deploy a test object corresponding to the current object on a virtual e-commerce platform, wherein the test object is obtained by optimizing the current object, and the target recommendation network model in the virtual e-commerce platform is trained by the method according to any one of claims 1 to 3; An AB test is performed on the current object and the test object on the virtual e-commerce platform.

5. A method for verifying an object, It is characterized in that include: Deploy multiple objects to be verified corresponding to the current object on the virtual e-commerce platform, wherein the objects to be verified are obtained by optimizing the current object, and the target recommendation network model in the virtual e-commerce platform is trained by the method according to any one of claims 1 to 3; After the virtual e-commerce platform runs the plurality of objects to be verified for a specified time, determining a service quality result of each of the objects to be verified; According to the service quality result, a target object with the best service quality result is selected from the multiple objects to be verified.

6. A training device for a recommendation network model, It is characterized in that The device comprises: An acquisition module is used to acquire virtual user characteristics, search keywords and real interaction behaviors of sample users, wherein the real interaction behaviors are formed according to the real user characteristics of the sample users, the first commodity ranking results and the first user behaviors, the first commodity ranking results are obtained by searching on a real e-commerce platform based on the real user characteristics and the search keywords, and the first user behaviors are user behaviors generated by the sample users on the first commodity ranking results; A first determination module is used to determine a second commodity ranking result corresponding to the virtual user feature and the search keyword according to an initial recommendation network model in a virtual e-commerce platform, wherein the virtual e-commerce platform is obtained by simulating the real e-commerce platform; A second determination module, configured to determine a second user behavior of the sample user with respect to the second commodity ranking result; A forming module, configured to form a virtual interactive behavior according to the virtual user characteristics, the second commodity ranking result and the second user behavior; A training module, used for training the initial recommendation network model according to the virtual interaction behavior and the real interaction behavior to obtain a target recommendation network model; The training module comprises: A first input unit, used to input the virtual interaction behavior into a discriminator network model to obtain a first probability value of the authenticity of the virtual interaction behavior; A second input unit, used for inputting the real interaction behavior into the discriminator network model to obtain a second probability value of the authenticity of the real interaction behavior; A training unit, configured to train the initial recommendation network model according to the first probability value and the second probability value to obtain a target recommendation network model; The training unit is specifically used for: When the absolute value of the difference between the first probability value and the second probability value is greater than a preset threshold, inputting the virtual interaction behavior into a value network to obtain a value corresponding to the virtual interaction behavior; Update the model parameters of the initial recommendation network model according to the first probability value and the value until the absolute value of the difference between the probability value of authenticity corresponding to the real interaction behavior and the probability value of authenticity corresponding to the updated virtual interaction behavior is less than or equal to the preset threshold; Use the recommendation network model obtained when the absolute value of the difference is less than or equal to the preset threshold as the target recommendation network model.

7. A testing device for an object, characterized in that, it includes: A first deployment module, configured to deploy a test object corresponding to the current object on a virtual e-commerce platform, where the test object is obtained by optimizing the current object, and the target recommendation network model in the virtual e-commerce platform is trained according to the method described in any one of claims 1-3; A testing module, configured to perform an AB test on the current object and the test object on the virtual e-commerce platform.

8. A verification device for an object, characterized in that, it includes: A second deployment module, configured to deploy a plurality of objects to be verified corresponding to the current object on a virtual e-commerce platform, where the objects to be verified are obtained by optimizing the current object, and the target recommendation network model in the virtual e-commerce platform is trained according to the method described in any one of claims 1-3; A third determination module, configured to determine the service quality results of each of the objects to be verified after the virtual e-commerce platform runs the plurality of objects to be verified for a specified time; A selection module, configured to select a target object with the best service quality result from the plurality of objects to be verified according to the service quality results.

9. An electronic device, characterized in that, it includes: A memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that when the processor executes the program, it implements a training method of a recommendation network model described in any one of claims 1-3 or a testing method of an object described in claim 4 or a verification method of an object described in claim 5.

10. A computer-readable storage medium, on which a computer program is stored, characterized in that, when the program is executed by a processor, it implements a training method of a recommendation network model described in any one of claims 1-3 or a testing method of an object described in claim 4 or a verification method of an object described in claim 5.

Citation Information

Patent Citations

  • Virtual sample generation method, terminal device and storage medium

    CN111275205A

  • Neural network training method, data processing method and related equipment

    CN113159315A