Recommendation and model training methods, devices, equipment, media and program products

By reusing the debiasing network structure in upstream and downstream recommendation models, the estimation bias problem caused by upstream model iteration is solved, and the accuracy of the recommendation system is improved.

CN114491280BActive Publication Date: 2025-09-09BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210179883.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-25
Publication Date
2025-09-09
Estimated Expiration
2042-02-25

AI Technical Summary

Technical Problem

In the recommendation system, the update iteration of the upstream recommendation model causes changes in the distribution of estimated click-through rate or high-order semantic features, introducing estimation bias and affecting the accuracy of recommendations.

Method used

The debiasing network structure is reused in upstream and downstream recommendation models. When the upstream model is updated and iterated, the debiasing network structure reused in the downstream model is also iterated to avoid changes in the distribution of estimated click-through rate or high-order semantic features.

Benefits of technology

This improves the accuracy of resource recommendations and avoids online estimation deviations introduced by upstream model updates and iterations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114491280B_ABST
    Figure CN114491280B_ABST
Patent Text Reader

Abstract

The present disclosure provides a recommendation method and apparatus, a model training method and apparatus, an electronic device, a non-transient computer-readable storage medium storing computer instructions, and a computer program product, and relates to the field of computer technology, and in particular, to the field of data processing technology. The present disclosure reuses a debiasing network structure in an upstream recommendation model and a downstream recommendation model. When the upstream recommendation model is updated and iterated, the debiasing network structure reused in the downstream recommendation model is simultaneously driven to iterate together, thereby avoiding the problem of online estimation bias introduced by the upstream recommendation model update and iteration, which causes changes in the distribution of estimated click-through rates or high-order semantic features, and effectively improves the accuracy of resource recommendations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, in particular to the field of data processing technology, and discloses a recommendation method and device, a model training method and device, an electronic device, a non-transitory computer-readable storage medium storing computer instructions, and a computer program product. Background Art

[0002] In current recommendation systems, different recommendation models are used at different stages of the system funnel to help the recommendation system select resources that best match user preferences. Current recommendation systems often apply the upstream recommendation model's estimated click-through rate (CTR) or high-level semantic features for each resource to downstream recommendation models. This approach can lead to significant changes in the distribution of these estimated CTRs or high-level semantic features due to updates and iterations of the upstream recommendation model, introducing bias and affecting estimation accuracy. Summary of the Invention

[0003] The present disclosure at least provides a recommendation method and device, a model training method and device, an electronic device, a program product, and a storage medium.

[0004] According to one aspect of the present disclosure, a recommendation method is provided, comprising:

[0005] Input the preset scene features into the debiasing network structure to obtain the first debiasing features;

[0006] Based on the first debiasing feature, initial recommended resources are selected from multiple resources to be recommended;

[0007] Input the scene features corresponding to the initial recommended resources into the debiasing network structure to obtain the second debiasing features;

[0008] Based on the second debiasing feature, target recommended resources are screened from multiple initial recommended resources.

[0009] According to another aspect of the present disclosure, a model training method is provided, comprising:

[0010] Get upstream and downstream models;

[0011] Determine a target network structure corresponding to the preset function based on the network structure with the preset function in the upstream model and the network structure with the preset function in the downstream model;

[0012] The upstream main network structure and target network structure except the network structure of the preset function in the upstream model and the downstream main network structure except the network structure of the preset function in the downstream model are jointly trained to obtain the trained upstream main network structure, target network structure and downstream main network structure.

[0013] According to another aspect of the present disclosure, there is provided a recommendation device, comprising:

[0014] A first depolarization module, configured to input a preset scene feature into a depolarization network structure to obtain a first depolarization feature;

[0015] An initial screening module, configured to screen out initial recommended resources from a plurality of resources to be recommended based on the first debiasing feature;

[0016] The second debiasing module is used to input the scene features corresponding to the initial recommended resources into the debiasing network structure to obtain the second debiasing features;

[0017] The target screening module is used to screen target recommended resources from multiple initial recommended resources based on the second debiasing feature.

[0018] According to another aspect of the present disclosure, there is provided a model training device, comprising:

[0019] Model acquisition module, used to obtain upstream models and downstream models;

[0020] A network integration module is used to determine a target network structure corresponding to the preset function based on the network structure with the preset function in the upstream model and the network structure with the preset function in the downstream model;

[0021] The training module is used to jointly train the upstream main network structure and the target network structure in the upstream model except the network structure with preset functions, and the downstream main network structure in the downstream model except the network structure with preset functions, to obtain the trained upstream main network structure, target network structure and downstream main network structure.

[0022] According to another aspect of the present disclosure, there is provided an electronic device, comprising:

[0023] at least one processor; and

[0024] a memory communicatively connected to the at least one processor; wherein,

[0025] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method in any embodiment of the present disclosure.

[0026] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided. The computer instructions are used to cause a computer to execute the method in any embodiment of the present disclosure.

[0027] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program / instruction, which implements the method in any embodiment of the present disclosure when the computer program / instruction is executed by a processor.

[0028] According to the technology disclosed in the present invention, the debiasing network structure is reused in the upstream recommendation model and the downstream recommendation model. When the upstream recommendation model is updated and iterated, the debiasing network structure reused in the downstream recommendation model will be iterated at the same time, thereby avoiding the problem of online estimation deviation introduced by the update and iteration of the upstream recommendation model, which leads to changes in the distribution of estimated click-through rate or high-order semantic features, and effectively improves the accuracy of resource recommendations.

[0029] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The accompanying drawings are used to better understand the present invention and do not constitute a limitation of the present invention.

[0031] Figure 1 is a flow chart of a recommended method according to the present disclosure;

[0032] Figure 2 Schematic diagram of the upstream recommendation model and the downstream recommendation model after multiplexing the debiasing network structure according to the present disclosure;

[0033] Figure 3 is a schematic diagram of the initial structure of the upstream recommendation model and the downstream recommendation model according to the present disclosure;

[0034] Figure 4 is a flow chart of the model training method according to the present disclosure;

[0035] Figure 5 It is a structural diagram of a recommended device according to the present disclosure;

[0036] Figure 6 is a schematic structural diagram of a model training device according to the present disclosure;

[0037] Figure 7 Schematic diagram of the structure of an electronic device according to the present disclosure. DETAILED DESCRIPTION

[0038] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0039] In current recommendation systems, different recommendation models are used at different stages of the system funnel to assist the recommendation system in selecting resources that best match user preferences. A key area of ​​model iteration for recommendation systems is how to complement the strengths of different recommendation models and maximize the benefits of the entire recommendation system. In current recommendation systems, the upstream recommendation model's estimated click-through rate (CTR) or high-level semantic features for each resource are often applied to the downstream recommendation model. This approach can lead to significant changes in the distribution of the estimated CTR or high-level semantic features due to updates to the upstream recommendation model, introducing bias in the estimates and affecting their accuracy.

[0040] In response to the above technical deficiencies, the present disclosure provides at least a recommendation method and apparatus, a model training method and apparatus, an electronic device, a non-transient computer-readable storage medium storing computer instructions, and a computer program product. The present disclosure reuses a debiasing network structure in the upstream recommendation model and the downstream recommendation model. When the upstream recommendation model is updated and iterated, the debiasing network structure reused in the downstream recommendation model is also iterated, thereby avoiding the problem of online estimation bias introduced by the upstream recommendation model update and iteration, which causes changes in the estimated click-through rate or the distribution of high-order semantic features. This effectively improves the accuracy of resource recommendations.

[0041] The recommendation method and model training method disclosed in the present invention are described below through specific embodiments.

[0042] Figure 1 The flowchart of the recommended method of the embodiment of the present disclosure is shown. The execution subject of the embodiment can be a device with computing capabilities. Figure 1 As shown, the recommended method of the embodiment of the present disclosure may include the following steps:

[0043] S110: Input the preset scene feature into the depolarization network structure to obtain a first depolarization feature.

[0044] The scene features in this step may be location features and / or context features. The above-mentioned location features may include location features of each preset location and features of target recommended resources corresponding to each preset location. The recommendation system in the present disclosure is used to recommend resources for each preset location, and the location features of each preset location have been determined when the preset location is set. Exemplarily, the location features may include an identifier of the preset location, location information, etc. The features of the target recommended resources corresponding to the preset location may include features such as an identifier of the target recommended resource and resource attributes. A preset location for which no target recommended resource is determined does not have the above-mentioned location features.

[0045] The above debiasing network structure is used to debias scene features, such as position debiasing. Figure 2As shown, the debiasing network structure can specifically be a network used to learn the above-mentioned scene features, for example, it can be the network structure corresponding to the shallow tower in the upstream recommendation model.

[0046] Since the upstream recommendation model has not yet determined the target recommended resources for any preset location when processing information and cannot obtain the actual scene features, the scene features in this step are preset.

[0047] S120 : Filtering out initial recommended resources from a plurality of resources to be recommended based on the first debiasing feature.

[0048] The aforementioned initially recommended resources are initial recommendation results, which can be screened based on the first estimated click-through rates of each resource to be recommended output by the upstream recommendation model. For example, the upstream recommendation model can determine the first estimated click-through rate of each resource to be recommended based on the first debiasing feature and resource information of each resource to be recommended.

[0049] like Figure 2 As shown, the main tower in the upstream recommendation model can be specifically used to determine a first user preference feature based on the resource information of each resource to be recommended input into the main tower and the user information corresponding to each preset location. The upstream recommendation model then determines a first estimated click-through rate for each resource to be recommended based on the first debiasing feature and the first user preference feature. The above main tower can also be referred to as the upstream main network structure described below.

[0050] The upstream recommendation model can be used to more accurately determine the first user preference feature. Then, combining the first user preference feature and the first debiasing feature can more accurately determine the first estimated click-through rate of each resource to be recommended. Using this more accurate first estimated click-through rate, the initial recommended resources that the user is more likely to pay attention to can be screened out from multiple resources to be recommended.

[0051] S130: Input the scene features corresponding to the initial recommended resources into the debiasing network structure to obtain a second debiasing feature.

[0052] The scene features corresponding to the aforementioned initial recommended resources include scene features corresponding to the preset locations of the determined target recommended resources, which are screened from the initial recommended resources. The preset locations of undetermined target recommended resources do not have location features.

[0053] The scene features in this step may also be location features and / or context features. The above-mentioned location features may include the location features of each preset location of the target recommended resource and the features of the target recommended resource corresponding to each preset location. Exemplarily, the location features may include the identifier of the preset location, location information, etc. The features of the target recommended resource corresponding to the preset location may include the identifier of the target recommended resource, resource attributes, and other features. The preset location where the target recommended resource has not been determined does not have the above-mentioned location features.

[0054] Since the downstream recommendation model has determined the target recommended resources for each preset location in a certain order when processing information, the scene features in this step are real scene features.

[0055] The above scene features can more accurately represent the scene information of the corresponding preset location, and based on the scene features, it is helpful to match more suitable target recommendation resources for the preset location.

[0056] S140 : Filtering target recommended resources from the multiple initial recommended resources based on the second debiasing feature.

[0057] The target recommended resource is the final recommendation result determined for the current preset location, and can be screened based on the second estimated click-through rate of each initial recommended resource output by the downstream recommendation model. For example, the downstream recommendation model can determine the second estimated click-through rate of each initial recommended resource based on the second debiasing feature and resource information of each initial recommended resource.

[0058] like Figure 2 As shown, the main tower in the downstream recommendation model can be specifically used to determine the second user preference feature based on the resource information of each initially recommended resource input to the main tower and the user information corresponding to each preset position. The downstream recommendation model then determines the second estimated click-through rate of each initially recommended resource based on the second debiasing feature and the second user preference feature. The above main tower can also be referred to as the downstream main network structure described below.

[0059] The downstream recommendation model can be used to more accurately determine the second user preference feature. Then, combining the second user preference feature and the second debiasing feature can more accurately determine the second estimated click-through rate of each initial recommended resource. Using this more accurate second estimated click-through rate, the target recommended resources that the user is more likely to pay attention to can be screened out from multiple initial recommended resources.

[0060] The above-mentioned depolarization network structure is a network structure with a depolarization function. Specifically, the following steps can be used to determine the above-mentioned network structure with a depolarization function:

[0061] First, obtain the upstream recommendation model and the downstream recommendation model; the upstream recommendation model includes the upstream main network structure and the scene feature learning network structure, such as Figure 3 The main tower and shallow tower in

[15] ; the downstream recommendation model includes a general feature learning network structure and a scene feature influence network structure, such as Figure 3 Then, the network structure with debiasing function in the upstream recommendation model and the network structure with debiasing function in the downstream recommendation model are determined, and the two determined network structures are integrated to obtain the above-mentioned debiasing network structure.

[0062] After the debiasing network structure is determined, the network structures in the downstream recommendation model, excluding the network structures with debiasing functions, form the downstream main network structure. For example, since only a portion of the network structures in the context tower in the downstream recommendation model have debiasing functions, only this portion of the network structures is combined when forming the debiasing network structure. The remaining network structures in the context tower are merged with the common tower to obtain the aforementioned downstream main network structure.

[0063] Context Tower can effectively characterize the impact of scene features such as location and context on a resource's estimated click-through rate. The downstream recommendation model has high performance requirements when determining target recommended resources for each preset location. Therefore, the downstream recommendation model has a relatively simple network structure and is lightweight overall. For example, the downstream recommendation model can be a contextDNN model.

[0064] After the debiasing network structure is determined, the network structures in the upstream recommendation model other than the network structure with the debiasing function form the upstream main network structure. For example, since the shallow tower in the upstream recommendation model has the debiasing function as a whole, the main tower in the upstream recommendation model is the upstream main network structure.

[0065] The upstream main network structure in the upstream recommendation model is responsible for learning features in the user-item dimension of user preferences. This model has a complex structure and effectively represents user preferences. A shallow tower is used to learn scene representations such as location features and context features. The aforementioned debiasing network structure is used to semantically debias scene features such as location and context. The upstream recommendation model can be a refined ranking model.

[0066] After determining the debiasing network structure, the structures of the upstream recommendation model and the downstream recommendation model are as follows: Figure 2As shown, the upstream and downstream recommendation models reuse the debiasing network structure. During model training, the upstream main network structure, the debiasing network structure, and the downstream main network structure are jointly trained to obtain the trained upstream main network structure, the debiasing network structure, and the downstream main network structure. The debiasing network structure is reused in both recommendation models and is updated and iterated along with the upstream recommendation model, overcoming the aforementioned shortcomings of the existing technology and improving the accuracy of resource recommendations.

[0067] In some embodiments, the model training can be performed using the following steps:

[0068] First, the first loss information is determined using standard scene features and multiple sample resources to be recommended, and the initial recommended sample resources are screened out from the multiple sample resources to be recommended; then, the second loss information is determined using the multiple initial recommended sample resources; finally, based on the first loss information and the second loss information, the upstream main network structure, the debiasing network structure, and the downstream main network structure are jointly trained until the preset cutoff conditions are met, thereby obtaining the trained upstream main network structure, the debiasing network structure, and the downstream main network structure.

[0069] The above-mentioned standard scene features can be determined based on the target recommendation sample resources of each preset location and used as labels. The target recommendation sample resources of each preset location can be stored in the form of a list.

[0070] During actual training, the resource information corresponding to multiple sample resources to be recommended and the user information corresponding to each preset position can be input into the upstream main network structure, which processes the input information and outputs a first predicted preference feature; the standard scene feature is input into the debiasing network structure, which processes the input information and outputs a predicted debiasing feature; then, the first loss information is determined based on the predicted debiasing feature and the first predicted preference feature. Exemplarily, the upstream recommendation network processes the predicted debiasing feature and the first predicted preference feature to obtain a first estimated sample click-through rate corresponding to each sample resource to be recommended; the first loss information is determined based on the first estimated sample click-through rate and the target recommended sample resource at each preset position.

[0071] According to the above steps, the first loss information can be determined more accurately by combining the first estimated sample click rate and the target recommended sample resources at each preset position. The first loss information can be used to more effectively adjust the parameters of the upstream recommendation model including the debiased network structure.

[0072] During actual training, resource information corresponding to multiple initial recommended sample resources and user information corresponding to each preset position can be input into the downstream main network structure. The downstream main network structure processes the input information and outputs a second predicted preference feature. Thereafter, the second loss information is determined based on the predicted debiasing feature and the second predicted preference feature. Exemplarily, the downstream recommendation network processes the predicted debiasing feature and the second predicted preference feature to obtain a second estimated sample click-through rate corresponding to each sample resource to be recommended. The second loss information is determined based on the second estimated sample click-through rate and the target recommended sample resource at each preset position.

[0073] According to the above steps, the second loss information can be determined more accurately by combining the second estimated sample click rate and the target recommended sample resources at each preset position. The second loss information can be used to more effectively adjust the parameters of the downstream recommendation model including the debiased network structure.

[0074] The above embodiment combines the first loss information and the second loss information to jointly train the upstream recommendation model and the downstream recommendation model, which can ensure that the parameters of the debiasing network structure used by the upstream recommendation model and the debiasing network structure used by the downstream recommendation model are adjusted synchronously, thereby improving the accuracy of resource recommendations.

[0075] like Figure 4 As shown, the present disclosure also provides a model training method, the execution subject of the method is a device with computing capabilities, and the model training method may specifically include the following steps:

[0076] S410: Obtain an upstream model and a downstream model.

[0077] The upstream model may include an upstream recommendation model; the downstream model may include a downstream recommendation model.

[0078] S420: Determine a target network structure corresponding to the preset function based on the network structure with the preset function in the upstream model and the network structure with the preset function in the downstream model.

[0079] The preset function may include a debiasing function, and the target network structure may include a debiasing network structure.

[0080] The method for determining the target network structure here is the same as the method for determining the debiased network structure in the embodiment, and the same contents will not be repeated here.

[0081] S430. Jointly train the upstream main network structure and the target network structure except the network structure with preset functions in the upstream model, and the downstream main network structure except the network structure with preset functions in the downstream model to obtain trained upstream main network structure, target network structure and downstream main network structure.

[0082] The method of model training here is the same as the method of training the upstream main network structure, the debiasing network structure, and the downstream main network structure in the above embodiment, and the same content will not be repeated.

[0083] By reusing the target network structure in both upstream and downstream models, updates to the upstream model will also drive iterations of the reused target network structure in the downstream model, thus avoiding online estimation bias introduced by upstream model updates. Applying this training method to resource recommendation model training also avoids online estimation bias introduced by upstream model updates, improving the accuracy of resource recommendations.

[0084] In some embodiments, the above-mentioned joint training of the upstream main network structure other than the network structure of the preset function in the upstream model, the target network structure, and the downstream main network structure other than the network structure of the preset function in the downstream model to obtain the trained upstream main network structure, the target network structure, and the downstream main network structure can be specifically implemented using the following steps:

[0085] First, the first loss information is determined by using standard scene features and multiple sample resources to be recommended, and the initial recommended sample resources are screened out from the multiple sample resources to be recommended; then, the second loss information is determined by using the multiple initial recommended sample resources; finally, based on the first loss information and the second loss information, the upstream main network structure and the target network structure except the network structure of the preset function in the upstream model, and the downstream main network structure except the network structure of the preset function in the downstream model are jointly trained to obtain the trained upstream main network structure, target network structure and downstream main network structure.

[0086] The steps for model training in this embodiment are the same as those for model training in the embodiment of the above-mentioned recommended method, and the same contents will not be repeated here.

[0087] The above embodiment combines the first loss information and the second loss information to jointly train the upstream model and the downstream model, which can ensure that the parameters of the target network structure used by the upstream model and the target network structure used by the downstream model are adjusted synchronously, avoiding online estimation deviations caused by upstream model update iterations.

[0088] The above-mentioned determination of the first loss information by utilizing the standard scenario features and the plurality of sample resources to be recommended may specifically include:

[0089] First, the standard scenario features are input into the target network structure to obtain the predicted debiasing features; then, the resource information corresponding to each sample resource to be recommended is input into the upstream main network structure to obtain the first predicted preference features; finally, the first loss information is determined based on the predicted debiasing features and the first predicted preference features.

[0090] The steps of determining the first loss information in this embodiment are the same as the steps of determining the first loss information in the embodiment of the above-mentioned recommended method, and the same contents will not be repeated here.

[0091] According to the above steps, the first loss information can be determined more accurately, and the first loss information can be used to more effectively adjust the parameters of the upstream model including the target network structure.

[0092] The above-mentioned use of multiple initial recommended sample resources to determine the second loss information may specifically include:

[0093] First, the resource information corresponding to each initial recommended sample resource is input into the downstream main network structure to obtain the second prediction preference feature; then, the second loss information is determined based on the prediction debiasing feature and the second prediction preference feature.

[0094] The steps of determining the second loss information in this embodiment are the same as the steps of determining the second loss information in the embodiment of the above-mentioned recommended method, and the same contents will not be repeated here.

[0095] According to the above steps, the second loss information can be determined more accurately, and the second loss information can be used to more effectively adjust the parameters of the downstream model including the target network structure.

[0096] In the above embodiment, by jointly training the upstream model and the downstream model, the advantages of the models at different stages can be complementary, the capabilities of each model in feature representation and feature modeling can be fully reused, and the respective estimation effects can be improved. At the same time, the estimation deviation problem caused by the introduction of the features of the upstream model can be solved, thereby realizing the link optimization of the funnel model and maximizing the system benefits.

[0097] Based on the same inventive concept, an embodiment of the present disclosure also provides a recommendation device corresponding to a recommendation method. Since the principle of solving the problem by the device in the embodiment of the present disclosure is similar to that of the above-mentioned recommendation method in the embodiment of the present disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.

[0098] like Figure 5 FIG. 1 is a schematic diagram of the structure of the recommendation device provided by an embodiment of the present disclosure, including:

[0099] The first depolarization module 510 is configured to input a preset scene feature into a depolarization network structure to obtain a first depolarization feature.

[0100] The initial screening module 520 is configured to screen out initial recommended resources from a plurality of resources to be recommended based on the first debiasing feature.

[0101] The second debiasing module 530 is configured to input the scene features corresponding to the initial recommended resources into the debiasing network structure to obtain second debiased features.

[0102] The target screening module 540 is configured to screen target recommended resources from the multiple initial recommended resources based on the second debiasing feature.

[0103] In some embodiments, when the initial screening module 530 screens out the initial recommended resources from the plurality of resources to be recommended based on the first debiasing feature, it is configured to:

[0104] Inputting resource information corresponding to each resource to be recommended into the upstream main network structure of the upstream recommendation model to obtain a first user preference feature;

[0105] Determining a first estimated click-through rate of each to-be-recommended resource based on the first debiasing feature and the first user preference feature;

[0106] An initial recommended resource is selected from the plurality of resources to be recommended according to the first estimated click-through rate.

[0107] In some embodiments, when the target screening module 540 screens out a target recommended resource from a plurality of initially recommended resources based on the second debiasing feature, it is configured to:

[0108] Inputting resource information corresponding to each initially recommended resource into the downstream main network structure of the downstream recommendation model to obtain a second user preference feature;

[0109] determining a second estimated click-through rate for each initially recommended resource based on the second debiasing feature and the second user preference feature;

[0110] According to the second estimated click-through rate, a target recommended resource is screened out from the multiple initial recommended resources.

[0111] In some embodiments, the scene features include location features and / or context features; wherein the location features include location features of each preset location and features of target recommended resources corresponding to each preset location.

[0112] In some embodiments, a model training module 550 is further included for:

[0113] Obtain upstream recommendation models and downstream recommendation models;

[0114] Determine the debiasing network structure based on the network structure with debiasing function in the upstream recommendation model and the network structure with debiasing function in the downstream recommendation model;

[0115] The upstream main network structure and the debiasing network structure except the network structure with debiasing function in the upstream recommendation model, and the downstream main network structure except the network structure with debiasing function in the downstream recommendation model are jointly trained to obtain the trained upstream main network structure, debiasing network structure and downstream main network structure.

[0116] In some embodiments, the model training module 550 is specifically configured to:

[0117] Determining first loss information by using standard scenario features and a plurality of sample resources to be recommended, and selecting an initial recommended sample resource from the plurality of sample resources to be recommended;

[0118] Determining second loss information using a plurality of initial recommended sample resources;

[0119] According to the first loss information and the second loss information, the upstream main network structure and the debiasing network structure except the network structure with debiasing function in the upstream recommendation model, and the downstream main network structure except the network structure with debiasing function in the downstream recommendation model are jointly trained to obtain the trained upstream main network structure, debiasing network structure and downstream main network structure.

[0120] In some embodiments, when the model training module 550 determines the first loss information using the standard scene features and the plurality of sample resources to be recommended, it is configured to:

[0121] Input the standard scene features into the debiasing network structure to obtain the predicted debiasing features;

[0122] Inputting resource information corresponding to each sample resource to be recommended into the upstream main network structure to obtain a first predicted preference feature;

[0123] First loss information is determined according to the predicted debiasing feature and the first predicted preference feature.

[0124] In some embodiments, when the model training module 550 uses multiple initial recommendation sample resources to determine the second loss information, it is configured to:

[0125] Input the resource information corresponding to each initial recommended sample resource into the downstream main network structure to obtain the second prediction preference feature;

[0126] Second loss information is determined based on the predicted debiasing feature and the second predicted preference feature.

[0127] Based on the same inventive concept, a model training device corresponding to a model training method is also provided in an embodiment of the present disclosure. Since the principle of solving the problem by the device in the embodiment of the present disclosure is similar to the above-mentioned model training method in the embodiment of the present disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.

[0128] like Figure 6 FIG. 1 is a schematic diagram of the structure of the model training device provided in an embodiment of the present disclosure, including:

[0129] The model acquisition module 610 is used to acquire the upstream model and the downstream model.

[0130] The network integration module 620 is configured to determine a target network structure corresponding to the preset function based on the network structure with the preset function in the upstream model and the network structure with the preset function in the downstream model.

[0131] The training module 630 is used to jointly train the upstream main network structure and the target network structure in the upstream model except the network structure with preset functions, and the downstream main network structure in the downstream model except the network structure with preset functions, to obtain the trained upstream main network structure, target network structure and downstream main network structure.

[0132] In some embodiments, the preset function includes a depolarization function; and / or,

[0133] The upstream model includes an upstream recommendation model; and / or,

[0134] Downstream models include downstream recommendation models.

[0135] In some embodiments, the training module 630 is specifically configured to:

[0136] Determining first loss information by using standard scenario features and a plurality of sample resources to be recommended, and selecting an initial recommended sample resource from the plurality of sample resources to be recommended;

[0137] Determining second loss information using a plurality of initial recommended sample resources;

[0138] According to the first loss information and the second loss information, the upstream main network structure and the target network structure except the network structure of the preset function in the upstream model, and the downstream main network structure except the network structure of the preset function in the downstream model are jointly trained to obtain the trained upstream main network structure, target network structure and downstream main network structure.

[0139] In some embodiments, when determining the first loss information using the standard scene features and a plurality of sample resources to be recommended, the training module 630 is configured to:

[0140] Input the standard scene features into the target network structure to obtain the predicted debiased features;

[0141] Inputting resource information corresponding to each sample resource to be recommended into the upstream main network structure to obtain a first predicted preference feature;

[0142] First loss information is determined according to the predicted debiasing feature and the first predicted preference feature.

[0143] In some embodiments, when the training module 630 uses a plurality of initial recommended sample resources to determine the second loss information, it is configured to:

[0144] Input the resource information corresponding to each initial recommended sample resource into the downstream main network structure to obtain the second prediction preference feature;

[0145] Second loss information is determined based on the predicted debiasing feature and the second predicted preference feature.

[0146] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0147] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0148] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0149] like Figure 7 As shown, the device 700 includes a computing unit 710, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 720 or a computer program loaded from a storage unit 780 into a random access memory (RAM) 730. Various programs and data required for the operation of the device 700 can also be stored in the RAM 730. The computing unit 710, the ROM 720, and the RAM 730 are connected to each other via a bus 740. An input / output (I / O) interface 750 is also connected to the bus 740.

[0150] Various components in device 700 are connected to I / O interface 750, including an input unit 760, such as a keyboard, mouse, etc.; an output unit 770, such as various types of displays, speakers, etc.; a storage unit 780, such as a magnetic disk, optical disk, etc.; and a communication unit 790, such as a network card, modem, wireless communication transceiver, etc. Communication unit 790 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0151] The computing unit 710 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 710 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 710 performs the various methods and processes described above, such as the recommendation method or model training method. For example, in some embodiments, the recommendation method or model training method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 780. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 700 via the ROM 720 and / or the communication unit 790. When the computer program is loaded into the RAM 730 and executed by the computing unit 710, one or more steps of the recommendation method or model training method described above can be performed. Alternatively, in other embodiments, the computing unit 710 can be configured to perform the recommendation method or model training method by any other appropriate means (e.g., by means of firmware).

[0152] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0153] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0154] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0155] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0156] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0157] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0158] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0159] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A recommendation method, comprising: Input the preset scene features into the debiasing network structure to obtain the first debiasing features; Based on the first debiasing feature, initial recommended resources are selected from multiple resources to be recommended; Inputting the scene features corresponding to the initial recommended resources into a debiasing network structure to obtain a second debiasing feature; Based on the second debiasing feature, a target recommended resource is screened out from the multiple initial recommended resources.

2. The method according to claim 1, wherein The step of selecting an initial recommended resource from a plurality of resources to be recommended based on the first debiasing feature includes: Inputting resource information corresponding to each resource to be recommended into an upstream main network structure of an upstream recommendation model to obtain a first user preference feature; determining a first estimated click-through rate of each to-be-recommended resource based on the first debiasing feature and the first user preference feature; Initially recommended resources are screened out from the multiple resources to be recommended according to the first estimated click-through rate.

3. The method according to claim 1 or 2, wherein: The step of selecting a target recommended resource from the plurality of initial recommended resources based on the second debiasing feature includes: Inputting resource information corresponding to each of the initially recommended resources into a downstream main network structure of a downstream recommendation model to obtain a second user preference feature; determining a second estimated click-through rate of each initially recommended resource based on the second debiasing feature and the second user preference feature; Filter out target recommended resources from the multiple initial recommended resources according to the second estimated click-through rate.

4. The method according to claim 1 or 2, wherein: The scene features include location features and / or context features; wherein the location features include location features of each preset location and features of target recommended resources corresponding to each preset location.

5. The method according to claim 1 or 2, further comprising: Obtain upstream recommendation models and downstream recommendation models; Determining the debiasing network structure based on the network structure with debiasing function in the upstream recommendation model and the network structure with debiasing function in the downstream recommendation model; The upstream main network structure excluding the network structure with debiasing function in the upstream recommendation model, the debiasing network structure, and the downstream main network structure excluding the network structure with debiasing function in the downstream recommendation model are jointly trained to obtain the trained upstream main network structure, debiasing network structure and downstream main network structure.

6. The method according to claim 5, wherein: The upstream main network structure other than the network structure with a debiasing function in the upstream recommendation model, the debiasing network structure, and the downstream main network structure other than the network structure with a debiasing function in the downstream recommendation model are jointly trained to obtain the trained upstream main network structure, the debiasing network structure, and the downstream main network structure, including: Determining first loss information by using standard scenario features and a plurality of sample resources to be recommended, and screening out initial recommended sample resources from the plurality of sample resources to be recommended; Determining second loss information using a plurality of initial recommended sample resources; According to the first loss information and the second loss information, the upstream main network structure except the network structure with debiasing function in the upstream recommendation model, the debiasing network structure, and the downstream main network structure except the network structure with debiasing function in the downstream recommendation model are jointly trained to obtain the trained upstream main network structure, debiasing network structure and downstream main network structure.

7. The method according to claim 6, wherein: The determining of the first loss information by using the standard scene feature and the plurality of sample resources to be recommended includes: Input the standard scene features into the debiasing network structure to obtain the predicted debiasing features; Inputting resource information corresponding to each sample resource to be recommended into the upstream main network structure to obtain a first predicted preference feature; The first loss information is determined according to the predicted debiasing feature and the first predicted preference feature.

8. The method according to claim 7, wherein: The determining of the second loss information by using the plurality of initial recommended sample resources includes: Input the resource information corresponding to each initial recommended sample resource into the downstream main network structure to obtain the second prediction preference feature; Second loss information is determined based on the predicted debiasing feature and the second predicted preference feature.

9. A model training method comprising: Get upstream and downstream models; Determining a target network structure corresponding to the preset function based on a network structure with the preset function in an upstream model and a network structure with the preset function in a downstream model; The upstream main network structure other than the network structure of the preset function in the upstream model, the target network structure, and the downstream main network structure other than the network structure of the preset function in the downstream model are jointly trained to obtain the trained upstream main network structure, target network structure, and downstream main network structure; wherein, The preset function includes a depolarization function; and / or, The upstream model includes an upstream recommendation model; and / or, The downstream model includes a downstream recommendation model.

10. The method according to claim 9, wherein: The jointly training the upstream main network structure other than the network structure of the preset function in the upstream model, the target network structure, and the downstream main network structure other than the network structure of the preset function in the downstream model to obtain the trained upstream main network structure, target network structure, and downstream main network structure includes: Determining first loss information by using standard scenario features and a plurality of sample resources to be recommended, and screening out initial recommended sample resources from the plurality of sample resources to be recommended; Determining second loss information using a plurality of initial recommended sample resources; According to the first loss information and the second loss information, the upstream main network structure other than the network structure of the preset function in the upstream model, the target network structure, and the downstream main network structure other than the network structure of the preset function in the downstream model are jointly trained to obtain the trained upstream main network structure, target network structure and downstream main network structure.

11. The method according to claim 10, wherein: The determining of the first loss information by using the standard scene feature and the plurality of sample resources to be recommended includes: Input the standard scene features into the target network structure to obtain the predicted debiased features; Inputting resource information corresponding to each sample resource to be recommended into the upstream main network structure to obtain a first predicted preference feature; The first loss information is determined according to the predicted debiasing feature and the first predicted preference feature.

12. The method according to claim 11, wherein The determining of the second loss information by using the plurality of initial recommended sample resources includes: Input the resource information corresponding to each initial recommended sample resource into the downstream main network structure to obtain the second prediction preference feature; Second loss information is determined based on the predicted debiasing feature and the second predicted preference feature.

13. A recommendation device, comprising: A first depolarization module, configured to input a preset scene feature into a depolarization network structure to obtain a first depolarization feature; An initial screening module, configured to screen out initial recommended resources from a plurality of resources to be recommended based on the first debiasing feature; A second debiasing module is configured to input the scene features corresponding to the initial recommended resources into a debiasing network structure to obtain a second debiasing feature; The target screening module is configured to screen target recommended resources from the plurality of initial recommended resources based on the second debiasing feature.

14. The device according to claim 13, wherein When the initial screening module screens out the initial recommended resources from the plurality of resources to be recommended based on the first debiasing feature, it is configured to: Inputting resource information corresponding to each resource to be recommended into an upstream main network structure of an upstream recommendation model to obtain a first user preference feature; determining a first estimated click-through rate of each to-be-recommended resource based on the first debiasing feature and the first user preference feature; Initially recommended resources are screened out from the multiple resources to be recommended according to the first estimated click-through rate.

15. The device according to claim 13 or 14, wherein When the target screening module screens out the target recommended resource from the multiple initial recommended resources based on the second debiasing feature, it is configured to: Inputting resource information corresponding to each of the initially recommended resources into a downstream main network structure of a downstream recommendation model to obtain a second user preference feature; determining a second estimated click-through rate of each initially recommended resource based on the second debiasing feature and the second user preference feature; Filter out target recommended resources from the multiple initial recommended resources according to the second estimated click-through rate.

16. The device according to claim 13 or 14, wherein The scene features include location features and / or context features; wherein the location features include location features of each preset location and features of target recommended resources corresponding to each preset location.

17. The apparatus according to claim 13 or 14, further comprising a model training module, configured to: Obtain upstream recommendation models and downstream recommendation models; Determining the debiasing network structure based on the network structure with debiasing function in the upstream recommendation model and the network structure with debiasing function in the downstream recommendation model; The upstream main network structure excluding the network structure with debiasing function in the upstream recommendation model, the debiasing network structure, and the downstream main network structure excluding the network structure with debiasing function in the downstream recommendation model are jointly trained to obtain the trained upstream main network structure, debiasing network structure and downstream main network structure.

18. The device according to claim 17, wherein The model training module is specifically used for: Determining first loss information by using standard scenario features and a plurality of sample resources to be recommended, and screening out initial recommended sample resources from the plurality of sample resources to be recommended; Determining second loss information using a plurality of initial recommended sample resources; According to the first loss information and the second loss information, the upstream main network structure except the network structure with debiasing function in the upstream recommendation model, the debiasing network structure, and the downstream main network structure except the network structure with debiasing function in the downstream recommendation model are jointly trained to obtain the trained upstream main network structure, debiasing network structure and downstream main network structure.

19. The device according to claim 18, wherein When the model training module determines the first loss information using the standard scene features and the plurality of sample resources to be recommended, it is used to: Input the standard scene features into the debiasing network structure to obtain the predicted debiasing features; Inputting resource information corresponding to each sample resource to be recommended into the upstream main network structure to obtain a first predicted preference feature; The first loss information is determined according to the predicted debiasing feature and the first predicted preference feature.

20. The device according to claim 19, wherein When the model training module uses a plurality of initial recommended sample resources to determine the second loss information, it is used to: Input the resource information corresponding to each initial recommended sample resource into the downstream main network structure to obtain the second prediction preference feature; Second loss information is determined based on the predicted debiasing feature and the second predicted preference feature.

21. A model training device comprising: Model acquisition module, used to obtain upstream models and downstream models; A network integration module, configured to determine a target network structure corresponding to the preset function based on a network structure with the preset function in an upstream model and a network structure with the preset function in a downstream model; A training module is used to jointly train the upstream main network structure other than the network structure of the preset function in the upstream model, the target network structure, and the downstream main network structure other than the network structure of the preset function in the downstream model to obtain the trained upstream main network structure, target network structure, and downstream main network structure; The preset function includes a depolarization function; and / or, The upstream model includes an upstream recommendation model; and / or, The downstream model includes a downstream recommendation model.

22. The device according to claim 21, wherein The training module is specifically used for: Determining first loss information by using standard scenario features and a plurality of sample resources to be recommended, and screening out initial recommended sample resources from the plurality of sample resources to be recommended; Determining second loss information using a plurality of initial recommended sample resources; According to the first loss information and the second loss information, the upstream main network structure other than the network structure of the preset function in the upstream model, the target network structure, and the downstream main network structure other than the network structure of the preset function in the downstream model are jointly trained to obtain the trained upstream main network structure, target network structure and downstream main network structure.

23. The device according to claim 22, wherein When the training module determines the first loss information using the standard scene features and the plurality of sample resources to be recommended, it is configured to: Input the standard scene features into the target network structure to obtain the predicted debiased features; Inputting resource information corresponding to each sample resource to be recommended into the upstream main network structure to obtain a first predicted preference feature; The first loss information is determined according to the predicted debiasing feature and the first predicted preference feature.

24. The device according to claim 23, wherein When the training module uses a plurality of initial recommended sample resources to determine the second loss information, it is configured to: Input the resource information corresponding to each initial recommended sample resource into the downstream main network structure to obtain the second prediction preference feature; Second loss information is determined based on the predicted debiasing feature and the second predicted preference feature.

25. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 12.

26. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 12.

27. A computer program product comprising a computer program / instructions, wherein: When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 12 is implemented.

Citation Information

Patent Citations

  • Click rate estimation model training method and device, and click rate estimation method and device

    CN113420227A

  • Recommendation model training method, recommendation method and recommendation system

    CN113987358A