Model Training, Resource Recommendation Method, Device, Electronic Device and Storage Medium

Through the ES model training method, combined with the preset preference representation model and the preset regression network, the user resource preferences are accurately portrayed, and the problem of unrefined resource proportion control in the existing recommendation system is solved, more personalized resource recommendations are achieved, and user experience is improved.

CN114564644BActive Publication Date: 2025-07-18BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210187902.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-28
Publication Date
2025-07-18
Estimated Expiration
2042-02-28

AI Technical Summary

Technical Problem

In the personalized resource recommendation system, especially in hybrid resource recommendation systems, it is difficult to accurately characterize users' preferences for different resources, resulting in insufficient refinement of resource proportion control.

Method used

The ES model training method is adopted, combining the preset preference representation model, the preset regression network and the preset twin network. By obtaining the attribute characteristics of the training object, supervised and unsupervised training is carried out, the user's resource preference representation is identified, and resource recommendations are made based on the click proportion and fusion parameters.

Benefits of technology

It improves the personalization level of resource recommendations, accurately controls the resource proportion, improves the user's browsing experience, and reduces manual intervention and costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114564644B_ABST
    Figure CN114564644B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, apparatus, electronic device, and storage medium for model training and resource recommendation, which relate to the field of Internet technologies, and particularly to the fields of artificial intelligence and deep learning. The specific implementation solution is as follows: obtaining the training attribute features of a training object; inputting the training attribute features into a pre-trained preset preference representation model to obtain a training resource preference representation corresponding to the training object output by a preset hidden layer of the preset preference representation model; inputting the training resource preference representation into a preset evolution strategy model to obtain training fusion parameters for various resource types; determining a training recommendation score for each candidate resource of each resource type according to the training fusion parameters of each resource type; recommending each candidate resource to the training object based on the training recommendation scores of each candidate resource, and collecting feedback behaviors of the training object on the recommended candidate resources; and training the preset evolution strategy model by using the feedback behaviors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of Internet technologies, and in particular to the fields of artificial intelligence and deep learning. Background Art

[0002] With the development of Internet technologies, recommendation systems have achieved rapid development. By leveraging machine learning technologies, recommendation systems can, based on the exploration of user behaviors, gain insights into users' interest preferences, and then automatically generate personalized content recommendations for users based on their interest preferences, improving users' browsing experiences. Summary of the Invention

[0003] The present disclosure provides a method, apparatus, electronic device, and storage medium for model training and resource recommendation.

[0004] According to one aspect of the present disclosure, a method for training an ES (Evolutionary Strategy) model is provided, including:

[0005] Obtaining training attribute features of a training object;

[0006] Inputting the training attribute features into a pre-trained preset preference representation model to obtain a training resource preference representation corresponding to the training object output by a preset hidden layer of the preset preference representation model;

[0007] Inputting the training resource preference representation into a preset evolutionary strategy model to obtain training fusion parameters for various resource types;

[0008] Determining a training recommendation score for each candidate resource of each resource type according to the training fusion parameters of each resource type;

[0009] Based on the training recommendation scores of each candidate resource, recommending each candidate resource to the training object and collecting feedback behaviors of the training object regarding the recommended candidate resources;

[0010] Training the preset evolutionary strategy model using the feedback behaviors.

[0011] According to another aspect of the present disclosure, a method for training a resource allocation model is provided, including:

[0012] Obtaining training attribute features of a training object and a first label of the training object, where the first label indicates the true click ratio of the training object for each resource type;

[0013] Inputting the training attribute features into a pre-trained preset preference representation model to obtain a training resource preference representation corresponding to the training object output by a preset hidden layer of the preset preference representation model;

[0014] Input the training resource preference representation into a preset regression network to obtain the predicted click proportion of each resource type output by the preset regression network;

[0015] Train the preset regression network according to the actual click proportion and the predicted click proportion of each resource type. The trained preset regression network is a resource allocation model.

[0016] According to another aspect of the present disclosure, a method for training a preference representation model is provided, including:

[0017] Obtain a training pair and a second label of the training pair. The training pair includes training attribute features of two training objects. The second label indicates the actual parameter relationship between the two training objects in the training pair. The type of the actual parameter is consistent with the type of the output parameters of the two sub-networks of a preset siamese network;

[0018] Input the training attribute features of the two training objects into the two sub-networks of the preset siamese network to obtain the predicted parameters corresponding to the training pair output by the two sub-networks;

[0019] Determine the predicted parameter relationship of the training pair based on the predicted parameters corresponding to the training pair;

[0020] Train the two sub-networks of the preset siamese network based on the second label and the predicted parameter relationship of the training pair. The trained preset siamese network is a preference representation model. The output of the preset hidden layer of the preference representation model is the resource preference representation corresponding to the training object.

[0021] According to another aspect of the present disclosure, a resource recommendation method is provided, including:

[0022] Obtain the target attribute features of the target object;

[0023] Input the target attribute features into the above-mentioned preference representation model to obtain the target resource preference representation output by the preset hidden layer of the preference representation model;

[0024] Input the target resource preference representation into the above-mentioned resource allocation model and the above-mentioned evolutionary strategy model respectively to obtain the target click proportion and the target fusion parameter of each resource type;

[0025] Obtain the candidate resources of each resource type according to the target click proportion of each resource type. The proportion of the candidate resources of each resource type among all candidate resources is consistent with the target click proportion of this resource type;

[0026] Determine the recommendation score of each candidate resource of each resource type according to the target fusion parameter of each resource type;

[0027] Recommend each candidate resource to the target object based on the recommendation score of each candidate resource.

[0028] According to another aspect of the present disclosure, there is provided a training device for an evolutionary strategy model, including:

[0029] An acquisition unit configured to acquire training attribute features of a training object;

[0030] A first input unit configured to input the training attribute features into a pre-trained preset preference representation model to obtain a training resource preference representation corresponding to the training object of the preset hidden layer output of the preset preference representation model;

[0031] A second input unit configured to input the training resource preference representation into a preset evolutionary strategy model to obtain training fusion parameters for various resource types;

[0032] A determination unit configured to determine training recommendation scores for each candidate resource of each resource type according to the training fusion parameters of each resource type;

[0033] A recommendation unit configured to recommend each candidate resource to the training object based on the training recommendation scores of each candidate resource, and collect feedback behaviors of the training object for the recommended candidate resources;

[0034] A training unit configured to train the preset evolutionary strategy model by using the feedback behaviors.

[0035] According to another aspect of the present disclosure, there is provided a training device for a resource allocation model, including:

[0036] An acquisition unit configured to acquire training attribute features of a training object and a first label of the training object, where the first label indicates the true click ratio of the training object for each resource type;

[0037] A first input unit configured to input the training attribute features into a pre-trained preset preference representation model to obtain a training resource preference representation corresponding to the training object of the preset hidden layer output of the preset preference representation model;

[0038] A second input unit configured to input the training resource preference representation into a preset regression network to obtain a predicted click ratio for each resource type output by the preset regression network;

[0039] A training unit configured to train the preset regression network according to the true click ratio and the predicted click ratio of each resource type, and the trained preset regression network is a resource allocation model.

[0040] According to another aspect of the present disclosure, there is provided a training device for a preference representation model, including:

[0041] An acquisition unit, configured to acquire a training pair and a second label of the training pair, where the training pair includes training attribute features of two training objects, and the second label indicates a true parameter relationship between the two training objects in the training pair, and the type of the true parameter is consistent with the type of output parameters of two sub-networks of a preset siamese network;

[0042] An input unit, configured to input the training attribute features of the two training objects into the two sub-networks of the preset siamese network, and obtain prediction parameters corresponding to the training pair output by the two sub-networks;

[0043] A determination unit, configured to determine a prediction parameter relationship of the training pair based on the prediction parameters corresponding to the training pair;

[0044] A training unit, configured to train the two sub-networks of the preset siamese network based on the second label and the prediction parameter relationship of the training pair, and the trained preset siamese network is a preference representation model, and an output of a preset hidden layer of the preference representation model is a resource preference representation corresponding to the training object.

[0045] According to another aspect of the present disclosure, there is provided a resource recommendation device, including:

[0046] A first acquisition unit, configured to acquire target attribute features of a target object;

[0047] A first input unit, configured to input the target attribute features into the above-mentioned preference representation model, and obtain a target resource preference representation output by a preset hidden layer of the preference representation model;

[0048] A second input unit, configured to input the target resource preference representation into the above-mentioned resource allocation model and the above-mentioned evolutionary strategy model respectively, and obtain a target click ratio and a target fusion parameter of each resource type;

[0049] A second acquisition unit, configured to acquire candidate resources of each resource type according to the target click ratio of each resource type, and the proportion of candidate resources of each resource type in all candidate resources is consistent with the target click ratio of this resource type;

[0050] A determination unit, configured to determine a recommendation score of each candidate resource of each resource type according to the target fusion parameter of each resource type;

[0051] A recommendation unit, configured to recommend each candidate resource to the target object based on the recommendation scores of each candidate resource.

[0052] According to another aspect of the present disclosure, there is provided an electronic device, including:

[0053] At least one processor; and

[0054] A memory communicatively connected to the at least one processor; wherein,

[0055] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute any of the above-mentioned methods.

[0056] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute any of the above-mentioned methods.

[0057] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, which when executed by a processor implements any of the above-mentioned methods.

[0058] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings

[0059] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0060] Figure 1 is a schematic flowchart of a method for training a preference representation model provided by an embodiment of the present disclosure;

[0061] Figure 2 is a schematic structural diagram of a preset siamese network provided by an embodiment of the present disclosure;

[0062] Figure 3 is based on Figure 2 a schematic diagram of a clustering result of resource preference representations extracted from the obtained preference representation model;

[0063] Figure 4 is a schematic flowchart of a method for training a resource allocation model provided by an embodiment of the present disclosure;

[0064] Figure 5 is a schematic diagram of a network structure of a recommendation system provided by an embodiment of the present disclosure;

[0065] Figure 6 is a first schematic flowchart of a method for training an ES model provided by an embodiment of the present disclosure;

[0066] Figure 7 is a second schematic flowchart of a method for training an ES model provided by an embodiment of the present disclosure;

[0067] Figure 8 It is the third schematic flowchart of the training method of the ES model provided by the embodiments of the present disclosure;

[0068] Figure 9 It is the fourth schematic flowchart of the training method of the ES model provided by the embodiments of the present disclosure;

[0069] Figure 10 It is the fifth schematic flowchart of the training method of the ES model provided by the embodiments of the present disclosure;

[0070] Figure 11 It is a schematic diagram of a scenario of parameter evolution provided by the embodiments of the present disclosure;

[0071] Figure 12 It is a schematic flowchart of a resource recommendation method provided by the embodiments of the present disclosure;

[0072] Figure 13 It is a schematic structural diagram of a training device of the ES model provided by the embodiments of the present disclosure;

[0073] Figure 14 It is a schematic structural diagram of a training device of a resource allocation model provided by the embodiments of the present disclosure;

[0074] Figure 15 It is a schematic structural diagram of a training device of a preference representation model provided by the embodiments of the present disclosure;

[0075] Figure 16 It is a schematic structural diagram of a resource recommendation device provided by the embodiments of the present disclosure;

[0076] Figure 17 It is the first schematic block diagram of an electronic device for implementing the model training method or the resource recommendation method of the embodiments of the present disclosure;

[0077] Figure 18 It is the second schematic block diagram of an electronic device for implementing the model training method or the resource recommendation method of the embodiments of the present disclosure. Detailed implementation manners

[0078] The following makes an explanation of exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to help understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0079] For ease of understanding, the following explains the terms that appear in the embodiments of the present disclosure.

[0080] The list page refers to the page that displays a list of resources recommended by the recommendation system to an object.

[0081] The landing page refers to the page that displays resources, such as the page displayed after an object clicks on a resource shown on the above-mentioned list page.

[0082] Refresh refers to the update of the list page.

[0083] With the development of Internet technology, recommendation systems have achieved rapid development. By leveraging machine learning techniques and mining user behavior, recommendation systems can gain insights into users' interest preferences and automatically generate personalized content recommendations for users.

[0084] However, current recommendation systems mostly focus on personalized recommendations from the resource dimension. The learning of users' preferences for different resources is relatively indirect and not well characterized. For some recommendation systems, especially for hybrid resource recommendation systems such as feed stream recommendation systems, accurately characterizing users' preferences for resources is particularly important for more personalized control and adjustment of the proportions of different resources.

[0085] To achieve more personalized control and adjustment of the proportions of different resources and improve the personalization of resource recommendations, the embodiments of the present disclosure provide a method for training a preference representation model. This method can be applied to electronic devices such as servers and mobile terminals, which are not limited herein. For ease of description, the following will be described with the recommendation system as the execution subject, which is not limiting. This preference representation model can be used to identify the preference representation of an object for various resource types, that is, resource preference representation. Among them, resource types can include but are not limited to picture and text types, short video types, small video types, and long video types, etc.

[0086] As Figure 1 shown, the method for training this preference representation model includes the following steps:

[0087] Step S11, obtaining a training pair and a second label of the training pair. The training pair includes the training attribute features of two training objects, and the second label indicates the true parameter relationship between the two training objects in the training pair. The type of the true parameter is the same as the type of the output parameters of the two sub-networks of the preset siamese network.

[0088] In the embodiments of the present disclosure, an object refers to the target that requests to obtain recommended resources. For example, when user 1 requests to obtain resources from the recommendation system, this user 1 is the object. A training object is an object used when training the model. The attribute features of the training object can be called training attribute features, and the attribute features can include at least one of identification features, boundary information, behavior features, and scenario features.

[0089] Among them, the identification feature can be represented by the UID (User Identity) feature; the boundary information can be represented by Side info, and the Side info can include information such as age, gender, geographical location, and registration time.

[0090] The behavior feature is the behavior feature within multiple time period durations. Among them, the length and quantity of the time period duration can be set according to actual needs. For example, the behavior feature can be divided into long-term behavior features, medium-term behavior features, and short-term behavior features. Among them, the cycle duration of the long-term behavior feature is greater than that of the medium-term behavior feature, and the cycle duration of the medium-term behavior feature is greater than that of the short-term behavior feature. For example, the cycle duration of the long-term behavior feature is 7 days, the cycle duration of the medium-term behavior feature is 1 day, and the cycle duration of the short-term behavior feature is one session. For another example, in addition to the above long-term behavior features, medium-term behavior features, and short-term behavior features, the behavior feature can also include the behavior feature of one refresh. The behavior feature can include but is not limited to the number of clicks and browsing duration, etc. Among them, the browsing duration can include the landing page browsing duration, list page browsing duration, etc.

[0091] In the embodiments of the present disclosure, the recommendation system divides the behavior feature in detail, that is, divides the behavior feature into the behavior feature within multiple time period durations, which can more accurately capture the change of the short-term resource preference of the object and improve the accuracy of resource preference representation.

[0092] The preset siamese network is a two-tower structure. The network structures of the two towers (i.e., the two sub-networks) of the preset siamese network are the same. Each sub-network includes an input layer, a hidden layer, and an output layer. The input of the preset siamese network is the attribute feature of the object, and the output is the specified parameter, such as the click proportion of the resource type and the duration proportion of the resource type, etc.

[0093] The preset siamese network can adopt the Pairwise idea in LTR (Learning To Rank), and learn the user resource preference representation by constructing the pair of the duration proportion and click proportion of the resource within the object and between objects. In this case, the preset siamese network can also be called a Pairwise network.

[0094] Optionally, the preset siamese network can adopt the ESMM (Entire Space Multi-task Model) structure, such as Figure 2 shown. Figure 2Among them, PCTCR represents the click - after click ratio, PCR represents the predicted click ratio of the output of the preset siamese network; PCTDR represents the click - after duration ratio, PDR represents the predicted duration ratio of the output of the preset siamese network, the circled multiplication sign represents multiplication processing, and the circled minus sign represents subtraction and multiplication processing. Figure 2 Taking the training of the preset siamese network only with the output click ratio and duration ratio as an example, it does not play a limiting role. The preset siamese network of the ESMM structure takes the attribute features of two groups of objects as inputs respectively, such as Figure 2 the object features, refresh features, cross - features, etc. shown in it, and completes the training based on this, which can effectively reduce the impact of sample imbalance caused by a large number of 0 - value samples in the click ratio, and further improve the accuracy of the preset siamese network.

[0095] Taking the resource type as the picture - text type, the specified parameter output by the preset siamese network is the click ratio of the picture - text type, and taking the training pair 1 including object 1 and object 2 as an example, the click ratio of object 1 for the picture - text type is a, and the click ratio of object 2 for the picture - text type is b. In practice, if a > b, the second label of training pair 1 indicates a > b; if a < b in practice, the second label of training pair 1 indicates a < b.

[0096] To save storage resources and computing resources, the second label can be represented by binary 0 and 1. For example, when a > b as above, the first label of training pair 1 is 1, and when a < b, the first label of training pair 1 is 0.

[0097] In the embodiments of the present disclosure, in the case of multiple resource types, the second label can be represented by a binary string, where each bit of the string represents the real - parameter relationship of the corresponding resource type. For example, the specified parameter output by the preset siamese network is the click ratio of the picture - text type, the resource types include the picture - text type and the short - video type, the first bit of the second label corresponds to the picture - text type, the second bit of the second label corresponds to the short - video type, binary 0 indicates that the value of object 1 is greater than the value of object 2, and 1 indicates that the value of object 1 is less than or equal to the value of object 2. Then when the second label of training pair 1 is 10, it means that the click ratio of object 1 for the picture - text type is greater than the click ratio of object 2 for the picture - text type, and the click ratio of object 1 for the short - video type is less than or equal to the click ratio of object 2 for the short - video type.

[0098] In the embodiments of the present disclosure, the number of training pairs obtained by the recommendation system can be set according to actual needs. For example, when the performance of the recommendation system is low, the number of training pairs obtained is small; when the recommendation system has a high precision requirement for extracting resource preferences represented by the preset siamese network, the number of training pairs obtained is large.

[0099] In one embodiment of the present disclosure, the training attribute features include identification features. In this case, the recommendation system can obtain the identification features of the training object, expand the dimensions of the identification features of the training object, and obtain the training attribute features of the training object.

[0100] In the embodiments of the present disclosure, the identification feature is a low-dimensional feature. The recommendation system expands the dimensions of the identification features of the training object, and based on the expanded identification features and other features such as behavior features and scenario features, obtains the training attribute features of the training object. Then, step S11 can be executed.

[0101] The dimension expansion here can also be called Slot dimension expansion. The dimension of the identification feature and the dimension of the expanded identification feature can be set according to implementation requirements. Optionally, the identification feature can be an 8-dimensional embedding feature, and the recommendation system can expand the identification feature into a 32-dimensional feature, that is, the expanded identification feature is a 32-dimensional feature.

[0102] In the technical solution provided by the embodiments of the present disclosure, the recommendation system expands the dimensions of the identification features, improves the personalized expression of the object identification features, improves the distinguishability of the objects, and more precisely learns the personalized resource preference representation of the objects.

[0103] In another embodiment of the present disclosure, the training attribute features include identification features and boundary information. Based on this, the recommendation system can obtain the expanded features of multiple other objects whose boundary information matches the boundary information of the training object. The expanded feature is a feature obtained by expanding the identification features of other objects, such as the expanded identification feature obtained by the above-mentioned Slot dimension expansion.

[0104] The recommendation system matches the boundary information of the training object with the boundary information of other objects. If the two match, it obtains the expanded features of the other object, that is, obtains the expanded features of multiple other objects whose boundary information matches the boundary information of the training object. The matching mentioned here can be that the boundary information is the same, or the similarity of the boundary information is greater than a preset similarity threshold, and this is not limited.

[0105] The recommendation system fuses the expanded features of multiple other objects, combines other features such as behavior features and scenario features, and obtains the training attribute features of the training object. In the embodiments of the present disclosure, the recommendation system can use weighted pooling (SumPooling) to implement the fusion of the expanded features. Then, step S11 can be executed.

[0106] For example, the recommendation system can perform weighted averaging on the expanded features of multiple other objects, and the weighted averaged expanded features are the training attribute features of the training object.

[0107] For another example, the recommendation system may randomly select a specified number of extended features from the extended features of multiple other objects, and perform weighted averaging on the selected extended features to obtain the training attribute features of the training object.

[0108] In the technical solution provided by the embodiments of the present disclosure, the recommendation system extends the object identification features, improves the personalized expression of the object identification features, improves the distinguishability of the objects, and more precisely learns the personalized resource preference representation of the objects.

[0109] In addition, in the technical solution provided by the embodiments of the present disclosure, through the boundary information, the cold start of the object identification features is optimized; the vector expressions of the object identification features of multiple similar other objects are fused, and the object identification features of the target object are extended to high-dimensional features, such as the above-mentioned 32-dimensional features, improving the accuracy of the resource preference representation learning of the cold start object and improving the generalization.

[0110] Step S12: Input the training attribute features of two training objects into two subnets of a preset siamese network to obtain the prediction parameters corresponding to the training pair output by the two subnets.

[0111] The structure of the preset siamese network can be referred to Figure 2 As shown, the structure of the preset siamese network can also be in other forms, which is not limited herein.

[0112] The recommendation system inputs the training attribute features of one training object included in the training pair into one subnet of the preset siamese network. After the subnet processes the input training attribute features, the output layer of the subnet outputs the prediction parameters of the training object; the training attribute features of the other training object included in the training pair are input into the other subnet of the preset siamese network. After the subnet processes the input training attribute features, the output layer of the subnet outputs the prediction parameters of the training object. Here, the prediction parameters output by the two subnets are the prediction parameters corresponding to the training pair.

[0113] Step S13: Determine the prediction parameter relationship of the training pair based on the prediction parameters corresponding to the training pair.

[0114] The recommendation system can calculate the difference between the prediction parameters of the two training objects included in the training pair to obtain the prediction parameter relationship of the training pair.

[0115] Step S14: Train the two subnets of the preset siamese network based on the second label of the training pair and the prediction parameter relationship. The trained preset siamese network is a preference representation model, and the output of the preset hidden layer of the preference representation model is the resource preference representation corresponding to the training object.

[0116] The sub-network may include one or more hidden layers, and the preset hidden layer may be any hidden layer in the sub-network. During the process of the sub-networks of the preset siamese network processing the training attribute features, the recommendation system may extract the output of the preset hidden layer in the sub-network, and this output is the resource preference representation corresponding to the training object.

[0117] In the embodiments of the present disclosure, the process of training the two sub-networks of the preset siamese network in step S14 may include: the recommendation system determines the loss value of the parameter relationship prediction based on the second label and the prediction parameter relationship of the training pair; if the loss value is less than the set loss threshold, it is determined that the preset siamese network converges, and the training of the two sub-networks of the preset siamese network is ended, that is, the training of the preset siamese network is ended; otherwise, it is determined that the preset siamese network does not converge, the network parameters of the two sub-networks are adjusted, and step S12 is returned for execution until the preset siamese network converges.

[0118] The inventor found that Figure 1 The preference representation model trained by the method shown can extract the resource preference representation of the object, which can make different objects have better distinguishability. The recommendation system can use tsne to perform dimensionality reduction clustering on the resource preference representations of the preset hidden layer of the preset siamese network, and the clustering result is as Figure 3 shown, Figure 3 where each point represents the resource preference representation of an object. From Figure 3 it can be seen that the resource preference representations of text and picture heavy objects, short video heavy objects, and small video heavy objects are respectively concentrated in different regions. In the embodiments of the present disclosure, the so-called "heavy" refers to that the click ratio is greater than the preset click ratio threshold.

[0119] In the technical solution provided by the embodiments of the present disclosure, the recommendation system directly uses the preference representation model to process the attribute features, and extracts the output of the preset hidden layer in the sub-network of the preset siamese network as the resource preference representation, which reduces the influence of artificial subjective factors, avoids manual intervention, reduces the labor cost, and improves the accuracy of the resource preference representation.

[0120] In addition, in the technical solution provided by the embodiments of the present disclosure, the recommendation system uses the training pair to perform supervised training on the preset siamese network, so that the preset siamese network can fully learn the data change law of the resource preference representation of the object. The trained preset siamese network can accurately identify the resource preference representation of the object, thereby improving the personalization of resource recommendation and the browsing experience of the object.

[0121] To achieve more personalized control and adjust the proportions of different resources, and improve the personalization of resource recommendations, embodiments of the present disclosure provide a method for training a resource allocation model. This method can be applied to electronic devices such as servers and mobile terminals, and is not limited thereto. For ease of description, the following will be described with a recommendation system as the execution subject, which does not serve as a limitation. This resource allocation model can be used to determine the allocation ratios of various resource types. Among them, resource types can include but are not limited to picture and text types, short video types, small video types, long video types, etc.

[0122] As Figure 4 shown, the training method of this preference representation model includes the following steps:

[0123] Step S41, obtain the training attribute features of the training object and the first label of the training object, where the first label indicates the true click ratio of the training object for each resource type.

[0124] In embodiments of the present disclosure, the click ratio of a resource type represents the ratio of the number of times an object clicks on this resource type to the number of times the object clicks on all resource types.

[0125] The training attribute features can include but are not limited to at least one of identification features, boundary information, behavior features, and scene features.

[0126] Optionally, the recommendation system can obtain the identification features of the training object; perform dimension expansion on the identification features to obtain the training attribute features of the training object. The recommendation system can also obtain the expanded features of multiple other objects whose boundary information matches the boundary information of the training object, where the expanded features are the features obtained by performing dimension expansion on the identification features of other objects; fuse the expanded features of multiple other objects to obtain the training attribute features of the training object. For specific reference, see the relevant description in the above step 11 part, and details will not be repeated here.

[0127] The number of training attribute features can be set according to actual needs. For example, when the performance of the recommendation system is low, the number of sample attribute features obtained is small; when the recommendation system has a high precision requirement for determining the resource allocation ratio by the preset regression network, the number of sample attribute features obtained is large.

[0128] Step S42, input the training attribute features into a pre-trained preset preference representation model to obtain the training resource preference representation corresponding to the training object output by the preset hidden layer of the preset preference representation model. The training process of the preset preference representation model can be found in the relevant description in the above steps S11 - S14 part, and details will not be repeated here.

[0129] Step S43, input the training resource preference representation into the preset regression network to obtain the predicted click ratio of each resource type output by the preset regression network.

[0130] In the embodiments of the present disclosure, the preset regression network may be a Pointwise network based on a multi-task framework. The preset regression network is used for resource quota allocation in three stages: resource aggregation before fine ranking and sequence candidate. The input of the preset regression network is the resource preference representation, and the output is the click proportion of each resource type at the refresh level. As Figure 5 shown in the Pointwise network, its input is the resource preference representation output by the preset hidden layer of the Pairwise network, and the output is the click proportion of picture and text types, the click proportion of short video types, the click proportion of small video types, etc.

[0131] After the recommendation system obtains the training resource preference representation, it inputs the training resource preference representation into the preset regression network. After the preset regression network processes the training resource preference representation, it predicts the click proportion of each resource type and outputs the predicted click proportion of each resource type, that is, the predicted click proportion of each resource type.

[0132] Step S44: Train the preset regression network according to the actual click proportion and the predicted click proportion of each resource type. The trained preset regression network is the resource allocation model.

[0133] In the embodiments of the present disclosure, the process of training the preset regression network in step S44 may include: the recommendation system determines the loss value of click proportion prediction based on the first label of each resource type and the predicted click proportion; if the loss value is less than the set loss threshold, it is determined that the preset regression network converges, and the training of the preset regression network ends; otherwise, it is determined that the preset regression network does not converge, the network parameters of the preset regression network are adjusted, and step S43 is returned for execution until the preset regression network converges.

[0134] In the technical solution provided by the embodiments of the present disclosure, the recommendation system performs supervised training on the preset regression network using the training resource preference representation, so that the preset regression network can fully learn the data change law of the click proportion of the object. The trained preset regression network (i.e., the resource allocation model) can accurately identify the click proportion of each resource type corresponding to different resource preference representations, thereby improving the personalization of resource recommendation and the browsing experience of the object.

[0135] After the resource allocation model is trained, the recommendation system inputs the resource preference representation of the target object into the trained resource allocation model. After the resource allocation model processes the resource preference representation, it outputs the target click proportion of each resource type. The target click proportion of each resource type here can be used for quota allocation among resources in each layer of the funnel.

[0136] The recommendation system processes the resource preference representation using a resource allocation model to obtain the target click proportion of each resource type, and accordingly conducts quota allocation of resources, achieving the supply of different resource quantities with scene personalization. For example, for objects with a relatively high click proportion of the short video type, they are video-intensive objects, and there will be more candidate video resources; for objects with a relatively high click proportion of the graphic and text type, they are graphic and text-intensive objects, and there will be more candidate graphic and text resources.

[0137] In addition, the recommendation system processes the resource preference representation using a resource allocation model to obtain the target click proportion of each resource type, reducing the influence of human subjective factors, avoiding manual intervention, reducing labor costs, and improving the accuracy of resource preference representation.

[0138] To achieve more personalized control and adjustment of the proportions of different resources and improve the personalization of resource recommendations, the embodiments of the present disclosure provide a training method for an ES model. This method can be applied to electronic devices such as servers and mobile terminals, and is not limited thereto. For ease of description, the following will be described with the recommendation system as the execution subject, which is not limiting. This ES model can be used to determine the fusion parameters of various resource types. Among them, the resource types can include but are not limited to graphic and text types, short video types, short video types, and long video types, etc.

[0139] As Figure 6 shown, the training method of this preference representation model includes the following steps:

[0140] Step S61, obtain the training attribute features of the training object.

[0141] In the embodiments of the present disclosure, the training attribute features can include but are not limited to at least one of identification features, boundary information, behavior features, and scene features.

[0142] Optionally, step S61 can be: obtain the identification features of the training object; expand the dimension of the identification features to obtain the training attribute features of the training object. For specific reference, see the relevant description in part 11 of the above steps, and details will not be repeated here.

[0143] Optionally, step S61 can also be: obtain the expanded features of multiple other objects whose boundary information matches the boundary information of the training object, where the expanded features are the features obtained by expanding the identification features of other objects; fuse the expanded features of multiple other objects to obtain the training attribute features of the training object. For specific reference, see the relevant description in part 11 of the above steps, and details will not be repeated here.

[0144] Step S62: Input the training attribute features into a pre-trained preset preference representation model to obtain the training resource preference representation corresponding to the training object output by the preset hidden layer of the preset preference representation model. For the training process of the preset preference representation model, refer to the relevant descriptions in the above Steps S11 - S14, which will not be elaborated here.

[0145] Step S63: Input the training resource preference representation into a preset ES model to obtain the training fusion parameters for various resource types.

[0146] In the embodiments of the present disclosure, the preset ES model can also be referred to as an evolutionary fusion model, and this preset ES model performs unsupervised evolutionary learning. The output of the preset ES model is the fusion parameter of the resource type, which is used for multi-objective fusion in the resource aggregation and sequence candidate stages. For the ranking stages of resources such as resource aggregation and sequence candidate, different resource types use different fusion parameters. For example Figure 5 As shown in the ES model, its input is the resource preference representation output by the preset hidden layer of the Pairwise network, and the output is the fusion parameters of the picture and text type, the fusion parameters of the short video type, the fusion parameters of the small video type, etc.

[0147] The recommendation system inputs the training resource preference representation into the preset ES model. After the preset ES model processes the training resource preference representation, it outputs the fusion parameters of various resource types, that is, the training fusion parameters.

[0148] Step S64: Determine the training recommendation scores of each candidate resource for each resource type according to the training fusion parameters of each resource type.

[0149] In the embodiments of the present disclosure, the candidate resources can be resources of various resource types obtained from the resource pool according to the click ratio of each resource type of the resource allocation model shown above Figure 4 or resources obtained by other means, which are not limited herein. The number of candidate resources for each resource type can be one or more.

[0150] The fusion parameter of a resource type is the weighted parameter based on which the scoring parameters of this resource type are fused. The fusion parameter of a resource type can be one or more, and the specific number can be determined according to the number of scoring parameters required for this resource type. For example, if a resource type requires 2 scoring parameters, then the fusion parameters of this resource type are 2, corresponding to each scoring parameter respectively. The scoring parameters can include but are not limited to the estimated click ratio and the estimated expansion duration, etc.

[0151] The fusion parameters are used in the ranking stage of resources such as resource aggregation and sequence candidates. For each resource type, according to the training fusion parameters of that resource type, the recommendation system scores each candidate resource of that resource type to determine the recommendation score of each candidate resource of that resource type.

[0152] In one example, the recommendation system can use the following formula (1) to determine the recommendation score of each candidate resource of that resource type.

[0153] score = ∏x y (1)

[0154] In formula (1), score is the recommendation score of the candidate resource, x is the scoring parameter of the candidate resource, and y is the corresponding fusion parameter of the candidate resource.

[0155] For example, the scoring parameters of the resource type include the estimated click-through rate and the estimated expansion duration. In this case, the recommendation system can determine the recommendation score of the candidate resources of that resource type.

[0156] score = crt t1 *dur t2

[0157] Among them, score is the recommendation score of the candidate resource, crt is the estimated click-through rate of the candidate resource, dur is the estimated expansion duration of the candidate resource, t1 is the fusion parameter corresponding to the estimated click-through rate, and t2 is the fusion parameter corresponding to the estimated expansion duration.

[0158] In another example, the recommendation system can use the following formula (2) to determine the recommendation score of each candidate resource of that resource type.

[0159] score = ∑x y (2)

[0160] In formula (2), score is the recommendation score of the candidate resource, x is the scoring parameter of the candidate resource, and y is the corresponding fusion parameter of the candidate resource.

[0161] For example, the scoring parameters of the resource type include the estimated click-through rate and the estimated expansion duration. In this case, the recommendation system can determine the recommendation score of the candidate resources of that resource type.

[0162] score = crt t1 +dur t2

[0163] Among them, score is the recommended score of the candidate resource, crt is the estimated click proportion of the candidate resource, dur is the estimated expansion duration of the candidate resource, t1 is the fusion parameter corresponding to the estimated click proportion, and t2 is the fusion parameter corresponding to the estimated expansion duration.

[0164] In the embodiments of the present disclosure, the recommendation system may also use other methods to determine the recommended score of each candidate resource of each resource type, which is not limited herein.

[0165] Step S65: Based on the training recommended scores of each candidate resource, recommend each candidate resource to the training object, and collect the feedback behaviors of the training object on the recommended candidate resources.

[0166] After obtaining the recommended scores of each candidate resource, the recommendation system may sort the candidate resources in descending order of the recommended scores, recommend the top preset number of candidate resources to the training object, or recommend the candidate resources with recommended scores higher than the preset score to the target object, which is not limited herein.

[0167] After recommending each candidate resource to the training object, the training object will give their feedback based on the recommended list refreshed once, such as swiping, clicking, list page browsing, landing page viewing, commenting, liking, sharing, etc. The recommendation system can collect the feedback of the training object, that is, collect the feedback behaviors of the training object on the recommended candidate resources.

[0168] Step S66: Use the feedback behaviors to train the preset ES model.

[0169] In the embodiments of the present disclosure, the feedback behaviors may include one or more of the landing page browsing duration, list page browsing duration, average page browsing duration of each resource type, and click times. The training attribute features of the training object can also be re-extracted from the feedback behaviors, so as to return to execute step S62 and re-iteratively train the preset ES model.

[0170] The process of training the preset ES model in step S66 may include: the recommendation system uses the feedback behaviors to determine the reward value of the preset ES model, and optimizes each network parameter in the preset ES model based on the reward value. This optimization process can be continuously executed, that is, when the preset ES model is applied to recommend resources to real objects, each network parameter in the preset ES model can also be optimized, so that the preset ES model is more suitable for recommending resources to real objects and improves the accuracy of resource recommendation.

[0171] To reduce the burden on the recommendation system, the recommendation system uses the feedback behaviors to train the preset ES model and optimize each network parameter in the preset ES model until the current reward value is less than the preset reward threshold.

[0172] In the technical solution provided by the embodiments of the present disclosure, the recommendation system performs unsupervised training on a preset ES model by using the training attribute features of the training object, so that the preset ES model can fully learn the data change rules of the training attribute features of the object. The trained preset ES model can accurately identify the fusion parameters of the object, thereby improving the personalization of resource recommendation and enhancing the browsing experience of the object.

[0173] After training the ES model, the recommendation system inputs the resource preference representation of the target object into the trained preset ES model. After the preset ES model processes the resource preference representation, it outputs the target fusion parameters for each resource type.

[0174] The recommendation system uses the preset ES model to process the resource preference representation, obtains the target fusion parameters for each resource type, reduces the influence of human subjective factors, avoids manual intervention, reduces labor costs, and improves the accuracy of the resource preference representation.

[0175] In an embodiment of the present disclosure, the preset ES model may adopt the network structure of MMOE (Multi-gate Mixture of Experts). Specifically, the preset ES model may include a secondary network corresponding to each resource type, and the secondary network corresponding to each resource type may include a tertiary network corresponding to different preference degrees for this resource type.

[0176] In one example, objects can be divided into groups with different preference degrees for resource types according to the proportion of the landing page duration of each resource type in the specified cycle duration. Correspondingly, in the preset ES model, the secondary network corresponding to this resource type includes tertiary networks with different preference degrees. The specified cycle duration can be 7 days or 1 day, etc.

[0177] For example, the resource types include picture and text type, short video type, and small video type. The preference degrees can be divided into low activity, mild, moderate, and severe. In this case, the preset ES model may include 3 secondary networks, namely secondary network 1 corresponding to the picture and text type, secondary network 2 corresponding to the short video type, and secondary network 3 corresponding to the small video type; secondary network 1 includes a tertiary network corresponding to the low activity preference degree, a tertiary network corresponding to the mild preference degree, a tertiary network corresponding to the moderate preference degree, and a tertiary network corresponding to the severe preference degree, as Figure 5 shown.

[0178] In this case, the embodiments of the present disclosure provide a training method for the ES model, as Figure 7As shown, the method may include steps S71 - S77, where steps S71 - S72, S75 - S77 are the same as steps S61 - S62, S64 - S66 above. Steps S73 - S74 are an implementable way of step S63.

[0179] Step S73: Input the training resource preference representation into each three - level network corresponding to each resource type to obtain the output parameters of each three - level network.

[0180] For each resource type, the recommendation system inputs the training resource preference representation into each three - level network corresponding to this resource type to obtain the output parameters of each three - level network corresponding to this resource type.

[0181] Step S74: For each resource type, fuse the output parameters of each three - level network of this resource type to obtain the fused parameters of this resource type.

[0182] For each resource type, the recommendation system performs weighted fusion on the output parameters of each three - level network of this resource type to obtain the training fused parameters of this resource type. Taking Figure 5 the ES model shown as an example, for the graphic - text type, after the three - level networks of low - activity, mild, moderate, and severe degrees of the graphic - text type respectively output parameters, the fusion unit of the graphic - text type, that is, Figure 5 the graphic - text gate in it fuses these 4 output parameters to obtain the training fused parameters of the graphic - text type; similarly, the target fused parameters of the short - video type and the training fused parameters of the small - video type are obtained, which will not be elaborated here.

[0183] Among them, the weights of the output parameters of each three - level network of each resource type can be controlled by a gating unit to improve the accuracy and precision of fusion.

[0184] In the technical solution provided by the embodiments of the present disclosure, the recommendation system uses the three - level networks corresponding to different preference degrees of each resource type to process the training resource preference representation, and obtains the training fused parameters of each resource type, avoiding the problem that due to the unknown different preference degrees of the training object for various resource types, the training resource preference representation of this training object is input into a three - level network that does not match the preference degree, resulting in a low precision of the training fused parameters, and improving the precision of the fused parameters.

[0185] After the training of the ES model is completed according to the Figure 7 method shown and the ES model is put on the line, the recommendation system can still adopt the Figure 7 method shown above to optimize each network parameter in the ES model.

[0186] Based on the preset ES model adopting the MMOE network structure, the embodiments of the present disclosure provide a training method for the ES model. As Figure 8 shown, the method may include steps S81 - S87. Among them, steps S81 - S82, S85 - S87 are the same as steps S61 - S62, S64 - S66 above. Steps S83 - S84 are an implementable manner of step S63.

[0187] Step S83, for each resource type, input the training resource preference representation into the target three - level network corresponding to this resource type to obtain the output parameters of the target three - level network. The preference degree corresponding to the target three - level network is consistent with the preference degree of the training object for this resource type.

[0188] After obtaining the training resource preference representation, for each resource type, the recommendation system can determine the three - level network corresponding to the preference degree of the training object for this resource type as the target three - level network corresponding to this resource type; input the training resource preference representation into the target three - level network corresponding to this resource type.

[0189] Step S84, use the output of the target three - level network as the training fusion parameter for this resource type.

[0190] The target three - level network processes the training resource preference representation and outputs the training fusion parameter for this resource type. The recommendation system uses the output of the target three - level network as the training fusion parameter for this resource type.

[0191] In the technical solution provided by the embodiments of the present disclosure, when the preference degree of the training object for each resource type is known, the recommendation system uses the three - level network that matches the preference degree of this resource type to process the training resource preference representation, obtains the training fusion parameter for this resource type, avoids the influence of the fusion parameter output by the three - level network with a matching preference degree, and improves the accuracy of the fusion parameter.

[0192] In addition, when the preference degree of the training object for each resource type is known, the recommendation system can use the training resource preference representation of this training object to perform unsupervised training on the corresponding target three - level network to further improve the accuracy of the fusion parameter output by the target three - level network.

[0193] After the training of the ES model is completed according to the Figure 8 shown method and the ES model is launched, the recommendation system can still adopt the Figure 7 shown method to optimize each network parameter in the ES model.

[0194] In an embodiment of the present disclosure, the embodiments of the present disclosure also provide a training method for the ES model. As Figure 9As shown, the method may include steps S91 - S98, where steps S91 - S95 are the same as steps S61 - S65 above. Steps S96 - S98 are an implementable way of step S66.

[0195] Step S96: Extract reward parameters from the feedback behavior.

[0196] In the embodiments of the present disclosure, the reward parameters may include one or more of the landing page browsing duration, the list page browsing duration, the average page browsing duration of each resource type, and the number of clicks.

[0197] Step S97: Determine the reward value of the preset ES model based on the reward parameters.

[0198] The recommendation system calculates the reward value (reward) of the target object, that is, the reward of the preset ES model, based on information such as the landing page browsing duration, the list page browsing duration, the average page browsing duration of each resource type, and the number of clicks included in the target attribute features. The recommendation system can collect the feedback of the target object at the hourly or daily level and determine the reward, which is not limited herein.

[0199] In one example, taking one refresh as the calculation period of the reward, the recommendation system can calculate the reward of the preset ES model using the following formula (3).

[0200] reward=(t1 + t2)+t3*n (3)

[0201] In formula (3), t1 represents the landing page browsing duration, t2 represents the list page browsing duration, t3 represents the average page browsing duration of each resource type, and n represents the number of clicks.

[0202] In the embodiments of the present disclosure, the recommendation system can also use other methods to determine the reward, which is not limited herein.

[0203] Step S98: Update the network parameters of the preset ES model using the reward value.

[0204] The recommendation system uses the reward value to re - update the network parameters of the preset ES model through the ES algorithm to complete the evolution of the parameters. Here, the evolution of the parameters is equivalent to the training of the ES model. When the ES network includes multiple three - level networks, the evolution of the parameters here is the evolution of the parameters of each three - level network, that is, unsupervised training of the three - level network.

[0205] In the technical solution provided by the embodiments of the present disclosure, based on the feedback behavior of the training object, using the ES algorithm for parameter optimization of evolutionary learning, there is no need to design a complex policy network, which greatly improves the efficiency and effect of recommendation multi - strategy optimization and significantly reduces the human and resource costs.

[0206] In one embodiment of the present disclosure, the embodiments of the present disclosure further provide a training method for an ES model. As Figure 10 shown, the method may include steps S101-S109, where steps S101-S107 are the same as the above steps S91-S97. Step S109 is an implementable manner of step S98.

[0207] Step S108, obtain the perturbation parameters of each network parameter in the preset ES model.

[0208] In the embodiments of the present disclosure, the perturbation parameters (such as noise) of each network parameter in the preset ES model can be preset in the recommendation system, and the recommendation system obtains the perturbation parameters of each network parameter in the preset ES model.

[0209] The recommendation system can also randomly generate perturbation parameters according to a specified algorithm. For example, the recommendation system obtains scenario information such as the CUID (Called User Identity) of the training object, date, and hour, performs a hash calculation on the scenario information of the training object to obtain a random number seed (such as seed); adopts a Gaussian distribution, and uses the random number seed to generate the perturbation parameters of each network parameter in the preset ES model.

[0210] After obtaining the perturbation parameters, the recommendation system can save the random number seed for offline restoration of this set of perturbation parameters.

[0211] In the embodiments of the present disclosure, the recommendation system can also obtain perturbation parameters in other ways, which are not limited herein.

[0212] Step S109, update each network parameter of the preset evolutionary strategy model by using the reward value and the perturbation parameters of each network parameter.

[0213] For the perturbation parameters of each network parameter, the recommendation system uses the reward value and the perturbation parameters of the network parameter to re-update the network parameter of the preset ES model through the ES algorithm to complete the evolution of the parameters.

[0214] In the embodiments of the present disclosure, for each original network parameter in the preset ES model, since no feedback behavior of the training object is collected, that is, the reward value is empty, the recommendation system can ignore the reward value, use the perturbation parameter of the network parameter to perturb the network parameter, obtain the new network parameter corresponding to the network parameter, and update each network parameter of the preset ES model to the corresponding new network parameter.

[0215] After updating the network parameters of the preset ES model, the recommendation system can input the training resource preference representation into the updated preset ES model to complete parameter evolution and push it online. The updated preset ES model performs forward calculation to obtain the fusion parameters of various resource types.

[0216] In different stages of the entire recommendation process, the recommendation system respectively uses the newly calculated fusion parameters after adding perturbations and applies them to each stage and each strategy, thereby obtaining candidate resources, and generating a refreshed recommendation list recommended to the object from the candidate resources.

[0217] In the embodiments of the present disclosure, the process of parameter evolution can be referred to Figure 11 as shown Figure 11 In, the policy network can be understood as the recommendation system, and h represents hours. The policy network can collect the feedback of the object at the hourly level, such as scene features, the behavioral features of the object (which can also be called the immersion state), etc., and determine the reward. The policy network uses the reward to complete online learning, updates the network parameters of the preset ES model, obtains a newly refreshed recommendation list, and recommends it to the target object to realize the application and exploration of resource recommendation. The object will give its own feedback based on the newly refreshed recommendation list. Then the policy network can collect the feedback of the object at the hourly level and determine the reward. In this way, the cycle continues, constantly performing parameter evolution and optimization.

[0218] In the technical solution provided by the embodiments of the present disclosure, based on the feedback behavior of the training object, the ES algorithm is used to perform parameter optimization for evolutionary learning, without the need to design a complex policy network, greatly improving the efficiency and effect of multi-strategy optimization of recommendations, and greatly reducing the human and resource costs.

[0219] In the technical solution provided by the embodiments of the present disclosure, when the recommendation system recommends resources to the training object, it uses the perturbation parameters and reward values to perturb the network parameters of the preset ES model, so that the network parameters conform to the scene, thereby improving the accuracy of the fusion parameters.

[0220] In the embodiments of the present disclosure, the recommendation system can simultaneously adopt a preset twin network, a preset regression network, and a preset ES model, as shown above Figure 5 At this time, the recommendation system can be divided into two parts: evolutionary fusion and resource preference representation. The resource preference representation is composed of a preset twin network and a preset regression network, and this part performs supervised training and learning. The evolutionary fusion is composed of a preset ES model, and this part performs unsupervised training and learning. The preset regression network and the preset ES model share the resource preference representation output by the hidden layer in the preset twin network.

[0221] Figure 5The recommended system shown combines supervision and evolution, optimizes the strategy combination based on the deep evolutionary learning, and can be applied to the online system based on the Feed flow, which can greatly improve the efficiency and effect of multi-strategy optimization of recommendations and greatly reduce the human and resource costs.

[0222] Based on the preference representation model, resource allocation model, and ES model obtained through the above training, the embodiments of the present disclosure provide a resource recommendation method, which can be applied to electronic devices such as servers and mobile terminals, and is not limited thereto. For ease of description, the following uses the recommended system as the execution subject for illustration, which does not serve as a limitation. As Figure 12 shown, the resource recommendation method includes the following steps:

[0223] Step S121, obtain the target attribute features of the target object.

[0224] Step S122, input the target attribute features into the preference representation model to obtain the target resource preference representation of the preset hidden layer output of the preference representation model. The training process of the preference representation model can be referred to the relevant description in the above Figure 1 part.

[0225] Step S123, input the target resource preference representation into the resource allocation model and the ES model respectively to obtain the target click ratio of each resource type and the target fusion parameter. The training process of the resource allocation model can be referred to the relevant description in the above Figure 4 part, and the training process of the ES model can be referred to the relevant description in the above Figure 6 - 10 part.

[0226] Step S124, obtain the candidate resources of each resource type according to the target click ratio of each resource type, and the ratio of the candidate resources of each resource type in all candidate resources is consistent with the target click ratio of this resource type.

[0227] After determining the target click ratio of each resource type, for each resource type, the recommended system obtains resources that match the target click ratio of this resource type as candidate resources. The ratio of the candidate resources of each resource type in the candidate resources of all resource types obtained is consistent with the target click ratio of this resource type.

[0228] For example, the resource types include picture and text type, short video type, and small video type. The click ratios of the picture and text type, short video type, and small video type are 0.1, 0.2, and 0.7 respectively, that is, the click ratio of the picture and text type: the click ratio of the short video type: the click ratio of the small video type is 1:2:7. Based on the above click ratios, the recommended system obtains the candidate resources of each resource type, where the candidate resources of the picture and text type: the candidate resources of the short video type: the candidate resources of the small video type are 1:2:7.

[0229] Step S125: Determine the recommended score for each candidate resource of each resource type according to the target fusion parameter of each resource type. Step S125 is similar to the above-mentioned step S64, and for specific details, please refer to the relevant description in the above-mentioned step S64 part.

[0230] Step S126: Recommend each candidate resource to the target object based on the recommended score of each candidate resource. Step S126 is similar to the above-mentioned step S65, and for specific details, please refer to the relevant description in the above-mentioned step S65 part.

[0231] In the technical solution provided by the embodiments of the present disclosure, the resource preference representation of the target object is extracted, and the click ratio of each resource type is determined by using the resource preference representation to achieve the quota allocation between resources. In addition, the fusion parameter of each resource type is determined by using the resource preference representation, and each candidate resource of the resource type is scored and recommended by using the fusion parameter of each resource type. It can be seen that in the technical solution provided by the embodiments of the present disclosure, when performing resource recommendation, the preference of the object for different resource types is accurately characterized, the proportion of different resources is more personalized controlled and adjusted based on the resource preference of the object, which is more suitable for the comprehensive recommendation of mixed resources such as Feed flow recommendation, improving the personalization of resource recommendation, and further improving the browsing experience of the object.

[0232] Corresponding to the above-mentioned training method of the ES model, the embodiments of the present disclosure also provide a training device for the ES model, as Figure 13 shown, including:

[0233] An acquisition unit 131, configured to acquire the training attribute features of the training object;

[0234] A first input unit 132, configured to input the training attribute features into a pre-trained preset preference representation model to obtain the training resource preference representation corresponding to the training object output by the preset hidden layer of the preset preference representation model;

[0235] A second input unit 133, configured to input the training resource preference representation into the preset ES model to obtain the training fusion parameters of various resource types;

[0236] A determination unit 134, configured to determine the training recommended score of each candidate resource of each resource type according to the training fusion parameter of each resource type;

[0237] A recommendation unit 135, configured to recommend each candidate resource to the training object based on the training recommended score of each candidate resource, and collect the feedback behavior of the training object for the recommended candidate resources;

[0238] A training unit 136, configured to train the preset ES model by using the feedback behavior.

[0239] Among them, the preset ES model includes a secondary network corresponding to each resource type, and the secondary network corresponding to each resource type includes a tertiary network corresponding to different preference degrees for the resource type;

[0240] The second input unit 133 can specifically be used for:

[0241] Input the training resource preference representation into each tertiary network corresponding to each resource type to obtain the output parameters of each tertiary network;

[0242] For each resource type, fuse the output parameters of each tertiary network of the resource type to obtain the training fusion parameters of the resource type.

[0243] Among them, the preset ES model includes a secondary network corresponding to each resource type, and the secondary network corresponding to each resource type includes a tertiary network corresponding to different preference degrees for the resource type;

[0244] The second input unit 133 can specifically be used for:

[0245] For each resource type, input the training resource preference representation into the target tertiary network corresponding to the resource type to obtain the output parameters of the target tertiary network, and the preference degree corresponding to the target tertiary network is consistent with the preference degree of the training object for the resource type;

[0246] Take the output of the target tertiary network as the training fusion parameter of the resource type.

[0247] Among them, the training unit 136 can specifically be used for:

[0248] Extract the reward parameter from the feedback behavior;

[0249] Based on the reward parameter, determine the reward value of the preset ES model;

[0250] Use the reward value to update the network parameters of the preset ES model.

[0251] Among them, the reward parameter includes one or more of the landing page browsing duration, the list page browsing duration, the average page browsing duration of each resource type, and the number of clicks.

[0252] Among them, the acquisition unit 131 can also be used to acquire the perturbation parameter of each network parameter in the preset ES model;

[0253] The training unit 136 can specifically be used for:

[0254] Use the reward value and the perturbation parameter of each network parameter to update each network parameter of the preset evolutionary strategy model.

[0255] Among them, the obtaining unit 131 can specifically be used for:

[0256] Performing a hash calculation on the scene information of the training object to obtain a random number seed;

[0257] Adopting a Gaussian distribution and using the random number seed to generate perturbation parameters for each network parameter in the preset ES model.

[0258] Among them, the obtaining unit 131 can specifically be used for:

[0259] Obtaining the identification feature of the training object;

[0260] Expanding the dimension of the identification feature to obtain the training attribute feature of the training object.

[0261] Among them, the obtaining unit 131 can specifically be used for:

[0262] Obtaining the expanded features of multiple other objects whose boundary information matches the boundary information of the training object, where the expanded features are the features obtained by expanding the identification features of the other objects;

[0263] Fusing the expanded features of multiple other objects to obtain the training attribute feature of the training object.

[0264] Among them, the training attribute feature includes at least one of an object identification feature, boundary information, behavior feature, and scene feature;

[0265] The behavior feature includes behavior features within multiple time period durations.

[0266] In the technical solution provided by the embodiments of the present disclosure, the recommendation system performs unsupervised training on the preset ES model by using the training attribute feature of the training object, so that the preset ES model can fully learn the data change law of the training attribute feature of the object. The trained preset ES model can accurately identify the fusion parameters of the object, thereby improving the personalization of resource recommendation and improving the browsing experience of the object.

[0267] After training the ES model, the recommendation system inputs the resource preference representation of the target object into the trained preset ES model. After the preset ES model processes the resource preference representation, it outputs the target fusion parameters of each resource type.

[0268] The recommendation system processes the resource preference representation by using the preset ES model to obtain the target fusion parameters of each resource type, reduces the influence of artificial subjective factors, avoids manual intervention, reduces the labor cost, and improves the accuracy of the resource preference representation.

[0269] Corresponding to the above training method of the resource allocation model, an embodiment of the present disclosure further provides a training device for a resource allocation model, as Figure 14 shown, including:

[0270] A first acquisition unit 141, configured to acquire training attribute features of a training object and a first label of the training object, where the first label indicates the true click ratio of the training object for each resource type;

[0271] A first input unit 142, configured to input the training attribute features into a pre-trained preset preference representation model, and obtain a training resource preference representation corresponding to the training object output by a preset hidden layer of the preset preference representation model;

[0272] A second input unit 143, configured to input the training resource preference representation into a preset regression network, and obtain a predicted click ratio of each resource type output by the preset regression network;

[0273] A training unit 144, configured to train the preset regression network according to the true click ratio and the predicted click ratio of each resource type, and the trained preset regression network is the resource allocation model.

[0274] Wherein, the training device for the above resource allocation model may further include:

[0275] A second acquisition unit, configured to acquire identification features of the training object;

[0276] A dimension expansion unit, configured to expand the identification features to obtain training attribute features of the training object.

[0277] Wherein, the training device for the above resource allocation model may further include:

[0278] A third acquisition unit, configured to acquire dimension-expanded features of multiple other objects whose boundary information matches the boundary information of the training object, where the dimension-expanded features are features obtained by expanding the identification features of the other objects;

[0279] A fusion unit, configured to fuse the dimension-expanded features of multiple other objects to obtain training attribute features of the training object.

[0280] Wherein, the training attribute features include at least one of identification features, boundary information, behavior features, and scene features;

[0281] The behavior features include behavior features within multiple time period durations.

[0282] In the technical solution provided by the embodiments of the present disclosure, the recommendation system uses the training resource preference representation to perform supervised training on a preset regression network, so that the preset regression network can fully learn the data change law of the click ratio of the object. The trained preset regression network (i.e., the resource allocation model) can accurately identify the click ratio of each resource type corresponding to different resource preference representations, thereby improving the personalization of resource recommendation and the browsing experience of the object.

[0283] After the resource allocation model is trained, the recommendation system inputs the resource preference representation of the target object into the trained resource allocation model. After the resource allocation model processes the resource preference representation, it outputs the target click ratio of each resource type. The target click ratio of each resource type here can be used for quota allocation among resources in each layer of the funnel.

[0284] The recommendation system uses the resource allocation model to process the resource preference representation, obtains the target click ratio of each resource type, and performs quota allocation of resources accordingly, realizing the supply of different resource quantities with scene personalization. For example, an object with a relatively high click ratio of the short video type is a video heavy user, and there will be more candidate video resources; an object with a relatively high click ratio of the graphic type is a graphic heavy user, and there will be more candidate graphic resources for the graphic heavy user.

[0285] In addition, the recommendation system uses the resource allocation model to process the resource preference representation, obtains the target click ratio of each resource type, reduces the influence of human subjective factors, avoids manual intervention, reduces labor costs, and improves the accuracy of the resource preference representation.

[0286] Corresponding to the above training method of the preference representation model, the embodiments of the present disclosure also provide a training device for the preference representation model, as Figure 15 shown, including:

[0287] A first acquisition unit 151, configured to acquire a training pair and a second label of the training pair. The training pair includes the training attribute features of two training objects, and the second label indicates the true parameter relationship between the two training objects in the training pair. The type of the true parameter is the same as the type of the output parameters of the two subnets of the preset siamese network;

[0288] An input unit 152, configured to input the training attribute features of the two training objects into the two subnets of the preset siamese network to obtain the predicted parameters corresponding to the training pair output by the two subnets;

[0289] A determination unit 153, configured to determine the predicted parameter relationship of the training pair based on the predicted parameters corresponding to the training pair;

[0290] A training unit 154, configured to train two sub-networks of a preset siamese network based on the relationship between the second label and the prediction parameters of the training pair. The trained preset siamese network is a preference representation model, and the output of the preset hidden layer of the preference representation model is the resource preference representation corresponding to the training object.

[0291] Wherein, the training device of the above-mentioned preference representation model may further include:

[0292] A second acquisition unit, configured to acquire the identification feature of the training object;

[0293] A dimension expansion unit, configured to expand the dimension of the identification feature to obtain the training attribute feature of the training object.

[0294] Wherein, the training device of the above-mentioned preference representation model may further include:

[0295] A third acquisition unit, configured to acquire the expanded features of multiple other objects whose boundary information matches the boundary information of the training object, where the expanded features are the features obtained by expanding the identification features of the other objects;

[0296] A fusion unit, configured to fuse the expanded features of multiple other objects to obtain the training attribute feature of the training object.

[0297] Wherein, the training attribute feature includes at least one of an identification feature, boundary information, behavior feature, and scene feature;

[0298] The behavior feature includes the behavior features within multiple time period durations.

[0299] In the technical solution provided by the embodiments of the present disclosure, the recommendation system directly processes the attribute features using the preference representation model, and extracts the output of the preset hidden layer in the sub-network of the preset siamese network as the resource preference representation, reducing the influence of artificial subjective factors, avoiding manual intervention, reducing labor costs, and improving the accuracy of the resource preference representation.

[0300] In addition, in the technical solution provided by the embodiments of the present disclosure, the recommendation system performs supervised training on the preset siamese network using the training pair, so that the preset siamese network can fully learn the data change law of the resource preference representation of the object. The trained preset siamese network can accurately identify the resource preference representation of the object, thereby improving the personalization of resource recommendation and the browsing experience of the object.

[0301] Corresponding to the above resource recommendation method, the embodiments of the present disclosure further provide a resource recommendation device, as Figure 16 shown, including:

[0302] A first acquisition unit 161, configured to acquire the target attribute feature of the target object;

[0303] A first input unit 162, configured to input target attribute features into the above-mentioned preference representation model, so as to obtain a preset hidden layer output target resource preference representation of the preference representation model;

[0304] A second input unit 163, configured to input the target resource preference representation into the above-mentioned resource allocation model and the above-mentioned evolutionary strategy model respectively, so as to obtain a target click ratio and a target fusion parameter of each resource type;

[0305] A second acquisition unit 164, configured to acquire candidate resources of each resource type according to the target click ratio of each resource type, and the proportion of candidate resources of each resource type in all candidate resources is consistent with the target click ratio of this resource type;

[0306] A determination unit 165, configured to determine a recommendation score for each candidate resource of each resource type according to the target fusion parameter of each resource type;

[0307] A recommendation unit 166, configured to recommend each candidate resource to a target object based on the recommendation score of each candidate resource.

[0308] In the technical solution provided by the embodiments of the present disclosure, the resource preference representation of the target object is extracted, and the click ratio of each resource type is determined by using the resource preference representation to achieve quota allocation among resources. In addition, the fusion parameter of each resource type is determined by using the resource preference representation, and each candidate resource of this resource type is scored and recommended by using the fusion parameter of each resource type. It can be seen that in the technical solution provided by the embodiments of the present disclosure, when resource recommendation is performed, the preference of the object for different resource types is accurately characterized, and based on the resource preference of the object, the proportion of different resources is more personalized controlled and adjusted, which is more suitable for the comprehensive recommendation of hybrid resources such as Feed flow recommendation, improving the personalization of resource recommendation, and thus improving the browsing experience of the object.

[0309] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved are all in compliance with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0310] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0311] Figure 17FIG. 0 shows a schematic block diagram of an electronic device 1700 for the resource recommendation method according to an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0312] As Figure 17 shown, the device 1700 includes a computing unit 1701 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1702 or a computer program loaded from a storage unit 1708 into a random access memory (RAM) 1703. In the RAM 1703, various programs and data required for the operation of the device 1700 can also be stored. The computing unit 1701, the ROM 1702, and the RAM 1703 are connected to each other via a bus 1704. An input / output (I / O) interface 1705 is also connected to the bus 1704.

[0313] A plurality of components in the device 1700 are connected to the I / O interface 1705, including: an input unit 1706, such as a keyboard, a mouse, etc.; an output unit 1707, such as various types of displays, speakers, etc.; a storage unit 1708, such as a magnetic disk, an optical disk, etc.; and a communication unit 1709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1709 allows the device 1700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0314] The computing unit 1701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1701 executes the various methods and processes described above, such as model training or resource recommendation methods. For example, in some embodiments, the model training or resource recommendation method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 1708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1700 via the ROM 1702 and / or the communication unit 1709. When the computer program is loaded into the RAM 1703 and executed by the computing unit 1701, one or more steps of the model training or resource recommendation method described above can be executed. Alternatively, in other embodiments, the computing unit 1701 can be configured to execute the model training or resource recommendation method in any other suitable way (e.g., by means of firmware).

[0315] Figure 18 FIG. shows a schematic block diagram of an electronic device 1800 for the model training or resource recommendation method according to an embodiment of the present disclosure. The electronic device includes:

[0316] At least one processor 1801; and

[0317] A memory 1802 communicatively connected to the at least one processor 1801; wherein,

[0318] The memory 1802 stores instructions executable by the at least one processor 1801, and the instructions are executed by the at least one processor 1801 to enable the at least one processor 1801 to execute any one of the above-described model training or resource recommendation methods.

[0319] An embodiment of the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute any one of the above-described model training or resource recommendation methods.

[0320] An embodiment of the present disclosure also provides a computer program product, including a computer program, where the computer program, when executed by a processor, implements any one of the above-described model training or resource recommendation methods.

[0321] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0322] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0323] In the context of this disclosure, a machine-readable medium may be a tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0324] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input received from the user can be in any form (including acoustic input, voice input, or tactile input).

[0325] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0326] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server incorporating blockchain.

[0327] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.

[0328] The above specific implementation manners do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A training method for an evolutionary strategy model, comprising: Obtaining training attribute features of a training object, where the training object is an object used when training a model, the object indicates a target for requesting recommended resources, and the training attribute features include at least one of identification features, boundary information, behavior features, and scenario features; Inputting the training attribute features into a pre-trained preset preference representation model to obtain a training resource preference representation corresponding to the training object output by a preset hidden layer of the preset preference representation model. The training resource preference representation is a preference representation of the object for various resource types. The preset preference representation model is obtained by training two sub-networks of a preset siamese network using training pairs and second labels of the training pairs. The training pairs include training attribute features of two training objects, and the second label indicates the true parameter relationship between the two training objects in the training pair. The type of the true parameter is consistent with the type of the output parameters of the two sub-networks of the preset siamese network; Inputting the training resource preference representation into a preset evolutionary strategy model to obtain training fusion parameters for various resource types; Determining a training recommendation score for each candidate resource of each resource type according to the training fusion parameters of each resource type; Based on the training recommendation scores of each candidate resource, recommending each candidate resource to the training object and collecting feedback behaviors of the training object for the recommended candidate resources; Using the feedback behaviors to train the preset evolutionary strategy model. The preset evolutionary strategy model outputs fusion parameters for various resource types. The fusion parameter of a resource type is a weighted parameter based on which the scoring parameters of the resource type are fused.

2. The method according to claim 1, wherein, The preset evolutionary strategy model includes a secondary network corresponding to each resource type, and the secondary network corresponding to each resource type includes a tertiary network corresponding to different preference degrees for the resource type; The step of inputting the training resource preference representation into a preset evolutionary strategy model to obtain training fusion parameters for various resource types includes: Inputting the training resource preference representation into each tertiary network corresponding to each resource type to obtain output parameters of each tertiary network; For each resource type, fusing the output parameters of each tertiary network of the resource type to obtain the training fusion parameter of the resource type.

3. The method according to claim 1, wherein, The preset evolutionary strategy model includes a secondary network corresponding to each resource type, and the secondary network corresponding to each resource type includes a tertiary network corresponding to different preference degrees for the resource type; The step of inputting the training resource preference representation into a preset evolutionary strategy model to obtain training fusion parameters for various resource types includes: For each resource type, inputting the training resource preference representation into the target tertiary network corresponding to the resource type to obtain output parameters of the target tertiary network. The preference degree corresponding to the target tertiary network is consistent with the preference degree of the training object for the resource type; Taking the output of the target tertiary network as the training fusion parameter of the resource type.

4. The method according to claim 1, wherein, The step of using the feedback behaviors to train the preset evolutionary strategy model includes: Extract a reward parameter from the feedback behavior; Determine a reward value of the preset evolutionary strategy model based on the reward parameter; Update network parameters of the preset evolutionary strategy model by using the reward value.

5. The method according to claim 4, wherein The reward parameter includes one or more of a landing page browsing duration, a list page browsing duration, an average page browsing duration of each resource type, and a number of clicks.

6. The method according to claim 4, further comprising: Obtain a perturbation parameter of each network parameter in the preset evolutionary strategy model; The step of updating network parameters of the preset evolutionary strategy model by using the reward value includes: Update each network parameter of the preset evolutionary strategy model by using the reward value and the perturbation parameter of each network parameter.

7. The method according to claim 6, wherein, The step of obtaining a perturbation parameter of each network parameter in the preset evolutionary strategy model includes: Perform a hash calculation on the scenario information of the training object to obtain a random number seed; Generate a perturbation parameter of each network parameter in the preset evolutionary strategy model by using the random number seed and adopting a Gaussian distribution.

8. The method according to claim 1, wherein The step of obtaining training attribute features of a training object includes: Obtain an identification feature of the training object; Perform dimension expansion on the identification feature to obtain the training attribute features of the training object.

9. The method according to claim 1, wherein The step of obtaining training attribute features of a training object includes: Obtain dimension expansion features of a plurality of other objects whose boundary information matches the boundary information of the training object, where the dimension expansion feature is a feature obtained by performing dimension expansion on the identification feature of the other object; Fuse the dimension expansion features of the plurality of other objects to obtain the training attribute features of the training object.

10. The method according to any one of claims 1-9, wherein, The training attribute features include at least one of an object identification feature, boundary information, behavior features, and scenario features; The behavior features include behavior features within a plurality of time period durations.

11. A method for training a resource allocation model, comprising: Obtain training attribute features of a training object and a first label of the training object, where the first label indicates a true click ratio of the training object for each resource type, the training object is an object used when training a model, the object indicates a target for requesting recommended resources, and the training attribute features include at least one of an identification feature, boundary information, behavior features, and scenario features; Input the training attribute features into a pre-trained preset preference representation model to obtain a training resource preference representation corresponding to the training object output by a preset hidden layer of the preset preference representation model, where the training resource preference representation is a preference representation of the object for various resource types, and the preset preference representation model is trained by using training pairs and a second label of the training pairs to train two sub-networks of a preset siamese network, the training pairs include training attribute features of two training objects, and the second label indicates a true parameter relationship between the two training objects in the training pair, and the type of the true parameter is consistent with the type of output parameters of the two sub-networks of the preset siamese network; Input the training resource preference representation into a preset regression network to obtain a predicted click ratio of each resource type output by the preset regression network; Train the preset regression network according to the actual click ratio and predicted click ratio of each resource type. The trained preset regression network is a resource allocation model, and the resource allocation model identifies the click ratio of each resource type corresponding to different resource preference representations.

12. The method according to claim 11, further comprising: Obtain the identification feature of the training object; Dimensionality expand the identification feature to obtain the training attribute feature of the training object.

13. The method according to claim 11, further comprising: Obtain the dimensionality-expanded features of multiple other objects whose boundary information matches the boundary information of the training object, where the dimensionality-expanded features are the features obtained by dimensionality expanding the identification features of the other objects; Fuse the dimensionality-expanded features of the multiple other objects to obtain the training attribute feature of the training object.

14. The method according to any one of claims 11-13, wherein, The training attribute feature includes at least one of an identification feature, boundary information, behavior feature, and scenario feature; The behavior feature includes behavior features within multiple time period durations.

15. A method for training a preference representation model, comprising: Obtain a training pair and a second label of the training pair, where the training pair includes the training attribute features of two training objects, and the second label indicates the true parameter relationship between the two training objects in the training pair. The type of the true parameter is the same as the type of the output parameters of the two sub-networks of the preset siamese network. The training object is the object used when training the model, and the object indicates the target of requesting to obtain recommended resources. The training attribute feature includes at least one of an identification feature, boundary information, behavior feature, and scenario feature; Input the training attribute features of the two training objects into the two sub-networks of the preset siamese network to obtain the predicted parameters corresponding to the training pair output by the two sub-networks; Based on the predicted parameters corresponding to the training pair, determine the predicted parameter relationship of the training pair; Based on the second label and the predicted parameter relationship of the training pair, train the two sub-networks of the preset siamese network. The trained preset siamese network is a preference representation model, and the output of the preset hidden layer of the preference representation model is the resource preference representation corresponding to the training object. The resource preference representation is the preference representation of the object for various resource types.

16. The method according to claim 15, further comprising: Obtain the identification feature of the training object; Dimensionality expand the identification feature to obtain the training attribute feature of the training object.

17. The method according to claim 15, further comprising: Obtain the dimensionality-expanded features of multiple other objects whose boundary information matches the boundary information of the training object, where the dimensionality-expanded features are the features obtained by dimensionality expanding the identification features of the other objects; Fuse the dimensionality-expanded features of the multiple other objects to obtain the training attribute feature of the training object.

18. The method according to any one of claims 15 - 17, wherein, The training attribute feature includes at least one of an identification feature, boundary information, behavior feature, and scenario feature; The behavior feature includes behavior features within multiple time period durations.

19. A resource recommendation method, comprising: Obtain the target attribute feature of the target object; Input the target attribute features into the preference representation model obtained by the method according to any one of claims 15-18, and obtain the target resource preference representation of the preset hidden layer output of the preference representation model; Input the target resource preference representation into the resource allocation model obtained by the method according to any one of claims 11-14 and the evolutionary strategy model obtained by the method according to any one of claims 1-10 respectively, and obtain the target click ratio and target fusion parameter of each resource type; According to the target click ratio of each resource type, obtain the candidate resources of each resource type, and the proportion of the candidate resources of each resource type in all candidate resources is consistent with the target click ratio of this resource type; According to the target fusion parameter of each resource type, determine the recommendation score of each candidate resource of each resource type; Based on the recommendation scores of each candidate resource, recommend each candidate resource to the target object.

20. A training device for an evolutionary strategy model, comprising: An acquisition unit, configured to acquire the training attribute features of a training object, where the training object is an object used when training a model, the object indicates the target of requesting to obtain recommended resources, and the training attribute features include at least one of identification features, boundary information, behavior features, and scenario features; A first input unit, configured to input the training attribute features into a pre-trained preset preference representation model, and obtain the training resource preference representation corresponding to the training object of the preset hidden layer output of the preset preference representation model. The training resource preference representation is the preference representation of the object for various resource types. The preset preference representation model is obtained by training two sub-networks of a preset siamese network by using training pairs and second labels of the training pairs. The training pairs include the training attribute features of two training objects, and the second label indicates the true parameter relationship between the two training objects in the training pair. The type of the true parameter is consistent with the type of the output parameters of the two sub-networks of the preset siamese network; A second input unit, configured to input the training resource preference representation into a preset evolutionary strategy model, and obtain the training fusion parameters of various resource types; A determination unit, configured to determine the training recommendation score of each candidate resource of each resource type according to the training fusion parameter of each resource type; A recommendation unit, configured to recommend each candidate resource to the training object based on the training recommendation scores of each candidate resource, and collect the feedback behaviors of the training object for the recommended candidate resources; A training unit, configured to use the feedback behaviors to train the preset evolutionary strategy model. The preset evolutionary strategy model outputs the fusion parameters of various resource types, and the fusion parameter of a resource type is a weighted parameter based on which the scoring parameters of this resource type are fused.

21. A training device for a resource allocation model, comprising: A first acquisition unit for acquiring the training attribute features of a training object and a first label of the training object, where the first label indicates the true click ratio of the training object for each resource type, the training object is an object used for training a model, the object indicates the target of requesting to obtain recommended resources, and the training attribute features include at least one of identification features, boundary information, behavior features, and scenario features; A first input unit for inputting the training attribute features into a pre-trained preset preference representation model to obtain a training resource preference representation corresponding to the training object output by a preset hidden layer of the preset preference representation model. The training resource preference representation is a preference representation of the object for various resource types. The preset preference representation model is obtained by training two sub-networks of a preset siamese network using training pairs and second labels of the training pairs. The training pairs include the training attribute features of two training objects, and the second label indicates the true parameter relationship between the two training objects in the training pair. The type of the true parameter is the same as the type of the output parameters of the two sub-networks of the preset siamese network; A second input unit for inputting the training resource preference representation into a preset regression network to obtain a predicted click ratio of each resource type output by the preset regression network; A training unit for training the preset regression network according to the true click ratio and the predicted click ratio of each resource type. The trained preset regression network is a resource allocation model, and the resource allocation model identifies the click ratio of each resource type corresponding to different resource preference representations.

22. A training device for a preference representation model, comprising: A first acquisition unit for acquiring training pairs and second labels of the training pairs. The training pairs include the training attribute features of two training objects, and the second label indicates the true parameter relationship between the two training objects in the training pair. The type of the true parameter is the same as the type of the output parameters of the two sub-networks of the preset siamese network. The training object is an object used for training a model, the object indicates the target of requesting to obtain recommended resources, and the training attribute features include at least one of identification features, boundary information, behavior features, and scenario features; An input unit for inputting the training attribute features of the two training objects into the two sub-networks of the preset siamese network to obtain predicted parameters corresponding to the training pair output by the two sub-networks; A determination unit for determining the predicted parameter relationship of the training pair based on the predicted parameters corresponding to the training pair; A training unit for training the two sub-networks of the preset siamese network based on the second label and the predicted parameter relationship of the training pair. The trained preset siamese network is a preference representation model, and the output of the preset hidden layer of the preference representation model is the resource preference representation corresponding to the training object. The resource preference representation is a preference representation of the object for various resource types.

23. A resource recommendation device, comprising: A first acquisition unit for acquiring the target attribute features of a target object; A first input unit, configured to input the target attribute features into a preference representation model obtained by the device according to claim 22, so as to obtain a preset hidden layer output target resource preference representation of the preference representation model; A second input unit, configured to input the target resource preference representation into a resource allocation model obtained by the device according to claim 21 and an evolutionary strategy model obtained by the device according to claim 20 respectively, so as to obtain a target click ratio and a target fusion parameter of each resource type; A second acquisition unit, configured to acquire candidate resources of each resource type according to the target click ratio of each resource type, and the proportion of the candidate resources of each resource type in all candidate resources is consistent with the target click ratio of this resource type; A determination unit, configured to determine a recommendation score of each candidate resource of each resource type according to the target fusion parameter of each resource type; A recommendation unit, configured to recommend each candidate resource to the target object based on the recommendation scores of each candidate resource.

24. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1-10, or is enabled to execute the method according to any one of claims 11-14, or is enabled to execute the method according to any one of claims 15-18, or is enabled to execute the method according to claim 19.

25. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-10, or are enabled to execute the method according to any one of claims 11-14, or are enabled to execute the method according to any one of claims 15-18, or are enabled to execute the method according to claim 19.

26. A computer program product, comprising a computer program, where the computer program, when executed by a processor, implements the method according to any one of claims 1-10, or is enabled to execute the method according to any one of claims 11-14, or is enabled to execute the method according to any one of claims 15-18, or is enabled to execute the method according to claim 19.

Citation Information

Patent Citations

  • Co-evolution-based personalized recommendation method and device

    CN109190040A

  • Multi-objective fused educational resource personalized recommendation system and method

    CN110795619A