Training method, device, equipment and medium for resource recommendation and resource allocation model
Through deep learning technology and the generator + discriminator training paradigm, combined with object features and resource interaction records, personalized resource allocation ratios are generated, which solves the problems of insufficient resource quality distinction and inconsistent prediction results in existing resource recommendation systems, and improves the accuracy of resource recommendations and user experience.
Patent Information
- Application Number
- CN202411304268.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-18
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2044-09-18
AI Technical Summary
Existing resource recommendation systems have difficulty effectively distinguishing low-quality resources when faced with complex and changeable user needs and resource characteristics, resulting in inefficient resource recommendation. The model prediction results are inconsistent with actual application needs, and the complex factors that affect user preferences cannot be fully captured. The feature extraction capabilities are insufficient, which affects the recommendation effect and user experience.
Using deep learning technology, combined with the object characteristics, resource characteristics and resource interaction records of the target object, the resource allocation ratios of various resource types are generated. Through the training paradigm of the generator and discriminator, the prediction accuracy and adaptability of the resource allocation model are improved, high-quality candidate resources are screened out, and a personalized resource recommendation list is generated.
It improves the rationality and accuracy of resource allocation ratio prediction, meets users' personalized resource preference needs, improves the accuracy of resource recommendations and user experience, and improves the overall performance of the resource recommendation system.
Smart Images

Figure CN119336985B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of AI (Artificial Intelligence), specifically to technical fields such as NLP (Natural Language Processing), deep learning, and big data, and especially to training methods, devices, equipment, and media for resource recommendation and resource allocation models. Background Art
[0002] With the rapid development of information technology and the explosive growth of internet data, people are gradually moving from an era of information scarcity to an era of information overload. Against this backdrop, resource recommendation systems have emerged. They aim to intelligently analyze user behavior and interests to recommend personalized resources tailored to their preferences, effectively alleviating the problem of information overload and improving the user experience in resource recommendation scenarios. Summary of the Invention
[0003] The present disclosure provides a method, apparatus, device, and medium for training a resource recommendation and resource allocation model.
[0004] According to a first aspect of the present disclosure, a resource recommendation method is provided, comprising:
[0005] Obtaining an object feature and a resource feature sequence of a target object; wherein the resource feature sequence includes resource features of candidate resources of multiple resource types in a candidate resource set;
[0006] Acquire a record feature sequence; wherein the record feature sequence includes record features of at least one resource interaction record of the target object within a set time period;
[0007] generating resource allocation proportions of the multiple resource types according to the object characteristics, the resource characteristic sequence, and the record characteristic sequence;
[0008] A resource recommendation list is determined from the candidate resource set according to the resource allocation proportions of the multiple resource types, so as to recommend the resource recommendation list to the target object.
[0009] According to a second aspect of the present disclosure, a method for training a resource allocation model is provided, comprising:
[0010] Obtaining a first sample; wherein the first sample includes a first object feature of a first sample object, a first resource feature sequence, and a first record feature sequence; the first resource feature sequence includes resource features of candidate resources of multiple resource types in a candidate resource set; and the first record feature sequence includes record features of at least one resource interaction record of the first sample object within a first sample period;
[0011] Using the generator of the resource allocation model to generate resource allocation proportions of the multiple resource types according to the first sample;
[0012] Determining a resource recommendation list from the candidate resource set according to resource allocation proportions of the multiple resource types;
[0013] using the discriminator of the resource allocation model to predict, based on the resource recommendation list, a first satisfaction level of the first sample subject with respect to the resource allocation mode of the resource recommendation list;
[0014] The resource allocation model is trained based on the first satisfaction level.
[0015] According to a third aspect of the present disclosure, a resource recommendation device is provided, comprising:
[0016] A first acquisition module is configured to acquire object features and a resource feature sequence of a target object; wherein the resource feature sequence includes resource features of candidate resources of multiple resource types in a candidate resource set;
[0017] A second acquisition module is configured to acquire a record feature sequence, wherein the record feature sequence includes the record features of at least one resource interaction record of the target object within a set time period;
[0018] A generating module, configured to generate resource allocation ratios of the multiple resource types according to the object characteristics, the resource characteristic sequence, and the record characteristic sequence;
[0019] A determination module is configured to determine a resource recommendation list from the candidate resource set according to the resource allocation proportions of the multiple resource types, so as to recommend the resource recommendation list to the target object.
[0020] According to a fourth aspect of the present disclosure, a training device for a resource allocation model is provided, comprising:
[0021] A first acquisition module is configured to acquire a first sample; wherein the first sample includes a first object feature of a first sample object, a first resource feature sequence, and a first record feature sequence; the first resource feature sequence includes resource features of candidate resources of multiple resource types in a candidate resource set; and the first record feature sequence includes record features of at least one resource interaction record of the first sample object within a first sample period;
[0022] a generating module, configured to generate resource allocation proportions of the plurality of resource types according to the first sample using a generator of the resource allocation model;
[0023] a determination module, configured to determine a resource recommendation list from the candidate resource set according to resource allocation proportions of the multiple resource types;
[0024] a first prediction module, configured to use the discriminator of the resource allocation model to predict, based on the resource recommendation list, a first satisfaction level of the first sample subject with respect to the resource allocation mode of the resource recommendation list;
[0025] A first training module is used to train the resource allocation model based on the first satisfaction level.
[0026] According to a fifth aspect of the present disclosure, there is provided an electronic device, including:
[0027] at least one processor; and
[0028] a memory communicatively connected to the at least one processor; wherein,
[0029] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the resource recommendation method proposed in the first aspect of the present disclosure, or execute the resource allocation model training method proposed in the second aspect of the present disclosure.
[0030] According to the sixth aspect of the present disclosure, a non-transitory computer-readable storage medium of computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the resource recommendation method proposed in the first aspect of the present disclosure, or to execute the resource allocation model training method proposed in the second aspect of the present disclosure.
[0031] According to the seventh aspect of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the resource recommendation method proposed in the first aspect of the present disclosure, or, when executed, implements the resource allocation model training method proposed in the second aspect of the present disclosure.
[0032] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0034] Figure 1 This is a flowchart of the resource recommendation method provided in the first embodiment of the present disclosure;
[0035] Figure 2This is a flowchart of the resource recommendation method provided in the second embodiment of the present disclosure;
[0036] Figure 3 This is a flowchart of the resource recommendation method provided in the third embodiment of the present disclosure;
[0037] Figure 4 This is a flowchart of the resource recommendation method provided in the fourth embodiment of the present disclosure;
[0038] Figure 5 A flowchart of a method for training a resource allocation model provided in the fifth embodiment of the present disclosure;
[0039] Figure 6 A flowchart of a method for training a resource allocation model provided in Embodiment 6 of the present disclosure;
[0040] Figure 7 This is a flow chart of a method for training a resource allocation model provided in the seventh embodiment of the present disclosure;
[0041] Figure 8 A flowchart of a method for training a resource allocation model provided in the eighth embodiment of the present disclosure;
[0042] Figure 9 A schematic diagram illustrating the implementation principle of the resource preference estimation algorithm provided in an embodiment of the present disclosure;
[0043] Figure 10 This is a schematic diagram of the structure of the resource recommendation device provided in the ninth embodiment of the present disclosure;
[0044] Figure 11 A schematic diagram of the structure of a training device for a resource allocation model provided in the tenth embodiment of the present disclosure;
[0045] Figure 12 A schematic block diagram of an example electronic device that can be used to implement embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0046] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0047] Currently, resource recommendation systems recommend resources to users based on simple user behavior statistics or rule-based matching strategies. In other words, resource preference estimation methods often rely on simple user behavior statistics or rule-based matching strategies. These methods are inadequate for addressing the complex and ever-changing user needs and resource characteristics. For example, user resource preferences are not static but are influenced by a variety of factors, including personal characteristics, historical behavior, and real-time context. Furthermore, the diversity and dynamic nature of resources also require resource recommendation systems to capture and respond to these changes in real time.
[0048] Existing resource preference modeling methods have begun to attempt to use user characteristics and resource consumption characteristics as feature inputs for neural network models, leveraging the powerful feature extraction and pattern recognition capabilities of deep learning to accurately estimate user preferences. While this resource preference modeling method can estimate user resource preferences to a certain extent, it still has many shortcomings, seriously restricting the recommendation effectiveness and user experience of resource recommendation systems.
[0049] First, existing resource preference modeling methods primarily focus on estimating user resource preferences, but ignore the quality differences of the resources within the candidate resource set. In practical applications, candidate resource sets may contain a large number of low-quality resources, and existing methods often fail to effectively distinguish low-quality resources. As a result, these low-quality resources are also allocated the same resource quota, wasting valuable recommendation resources and reducing recommendation efficiency.
[0050] Secondly, existing resource preference modeling methods misalign their target estimates with actual application needs. Models typically predict the proportion of various resources in the presentation phase. However, in the actual operation of resource recommendation systems, the more critical issue is how to rationally allocate resources across the funnel, such as coarse and fine ranking, to optimize user experience and conversion efficiency. This misalignment makes the model's predictions difficult to directly apply to real-world business scenarios.
[0051] Furthermore, existing resource preference modeling methods have relatively limited feature dimensions and cannot fully capture the complex factors influencing user preferences. In the digital age, user behavior and preferences are influenced by a variety of factors, including personal characteristics, historical behavior, and real-time context. Existing methods often only consider a subset of these factors, limiting the model's predictive accuracy.
[0052] Finally, multi-layer, fully connected neural networks, one of the core technologies in existing resource preference modeling methods, lack sufficient feature extraction capabilities. While neural networks (NNs) offer significant advantages in feature extraction and pattern recognition, their extraction capabilities still need to be improved when faced with complex and ever-changing user and resource data. This limits the model's ability to extract potential information from the data, which in turn impacts the overall performance of the resource recommendation system.
[0053] Therefore, to address at least one of the above-mentioned problems, the present disclosure proposes a method, apparatus, device, and medium for training a resource recommendation and resource allocation model.
[0054] The following describes the resource recommendation and resource allocation model training method, apparatus, device, and medium of the embodiments of the present disclosure with reference to the accompanying drawings.
[0055] Figure 1 This is a flowchart of the resource recommendation method provided in the first embodiment of the present disclosure.
[0056] The embodiment of the present disclosure is described by taking the resource recommendation method configured in a resource recommendation device as an example. The resource recommendation device can be applied to any electronic device so that the electronic device can perform a resource recommendation function.
[0057] Among them, the electronic device can be any device with computing capabilities, such as a personal computer, mobile terminal, server, etc. The mobile terminal can be, for example, a mobile phone, tablet computer, personal digital assistant, wearable device, etc., which are hardware devices with various operating systems, touch screens and / or display screens.
[0058] Exemplarily, the resource recommendation method can be applied to a service end (or server, cloud end).
[0059] like Figure 1 As shown, the resource recommendation method may include the following steps S101 to S104:
[0060] Step S101 : obtaining object features and a resource feature sequence of a target object; wherein the resource feature sequence includes resource features of candidate resources of various resource types in a candidate resource set.
[0061] Among them, resource types include but are not limited to: long videos, short videos, dynamic (wherein, dynamic resources include images and texts, and the text contains fewer characters and a larger number of images), graphics and text (wherein, graphics and text resources include images and texts, and the text contains a larger number of characters and a smaller number of images), etc.
[0062] The target object may be any user object, and the object characteristics of the target object are used to represent the resource preference of the target object.
[0063] The resource feature sequence is generated by extracting features from candidate resources of multiple resource types in the candidate resource set to obtain resource features of the candidate resources of multiple resource types and based on the resource features of the candidate resources of multiple resource types.
[0064] Step S102: obtaining a record feature sequence; wherein the record feature sequence includes the record features of at least one resource interaction record of the target object within a set time period.
[0065] The set period is a pre-set time period, the end time of the set period may be the current time, and the time difference between the end time and the start time of the set period is pre-set, that is, the set period may be a recent time period.
[0066] Among them, resource interaction records (or resource consumption records) are used to record the resources that the target object has historically interacted with, as well as the target object's interaction information on the resource (including but not limited to: browsing time, comment information, sharing information, like information, etc.).
[0067] In the embodiment of the present disclosure, the resource interaction records of the target object within a set time period can be obtained, and features of each resource interaction record can be extracted to obtain the record features of each resource interaction record, and a record feature sequence can be generated based on the record features of each resource interaction record.
[0068] Step S103: Generate resource allocation ratios of multiple resource types based on the object features, resource feature sequences, and record feature sequences.
[0069] In the embodiment of the present disclosure, deep learning technology can be used to predict the resource allocation proportions of multiple resource types preferred by the target object based on the object characteristics, resource feature sequence and record feature sequence of the target object.
[0070] Step S104 : determining a resource recommendation list from the candidate resource set according to the resource allocation ratios of the multiple resource types, and recommending the resource recommendation list to the target object.
[0071] In the disclosed embodiment, a resource recommendation list can be determined from a candidate resource set based on the resource allocation ratios of various resource types. For example, if the resource allocation ratio of long videos is 20%, the resource allocation ratio of short videos is 30%, the resource allocation ratio of dynamic resources is 40%, and the resource allocation ratio of graphics and text is 10%, then the resource recommendation list will have a ratio of long videos, a ratio of short videos, a ratio of dynamic resources, and a ratio of graphics and text resources.
[0072] In an embodiment of the present disclosure, after determining the resource recommendation list, the resource recommendation list may be recommended to a target object, so that the target object browses or consumes recommended resources of various resource types in the resource recommendation list.
[0073] The resource recommendation method of the embodiment of the present disclosure combines the object characteristics of the target object, the resource characteristics of candidate resources of multiple resource types, and the record characteristics of the resource interaction records that characterize the user's resource preferences or immediate preferences within a set time period to predict the resource allocation ratios of multiple resource types preferred by the target object. This can improve the rationality and reliability of the resource allocation ratio prediction, and then recommend a resource recommendation list to the target object based on the reliable resource allocation ratio. This can ensure that the resource allocation method of the recommended resource recommendation list can meet the personalized resource preference needs of the target object, and improve the accuracy of resource recommendations in the resource recommendation list, thereby improving the user's usage experience in the resource recommendation scenario.
[0074] It should be noted that in the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information are all carried out with the user's consent, and are in compliance with relevant laws and regulations and do not violate public order and good morals.
[0075] In order to clearly illustrate how any embodiment of the present disclosure generates resource allocation ratios of multiple resource types based on object features, resource feature sequences, and record feature sequences, the present disclosure also proposes a resource recommendation method.
[0076] Figure 2 This is a flowchart of the resource recommendation method provided in the second embodiment of the present disclosure.
[0077] like Figure 2 As shown, the resource recommendation method may include the following steps S201 to S207:
[0078] Step S201 : obtaining object features and a resource feature sequence of a target object; wherein the resource feature sequence includes resource features of candidate resources of various resource types in a candidate resource set.
[0079] For explanation of step S201, please refer to the relevant description in any embodiment of the present disclosure, and will not be repeated here.
[0080] In any embodiment of the present disclosure, the candidate resource set may be obtained by screening from massive resources based on object features of the target object.
[0081] As an example, first, resource features of multiple initial resources can be obtained, wherein the resource features are obtained by performing feature extraction on the resource information of the initial resources. Afterwards, a trained resource recommendation model can be used to predict the resource scores of the multiple initial resources based on the object features of the target object and the resource features of the multiple initial resources, wherein the resource score of each initial resource is used to indicate the quality of the initial resource and the degree of preference of the target object for the initial resource. Thus, in the present disclosure, a candidate resource set can be determined from the multiple initial resources based on the resource scores of the multiple initial resources.
[0082] As an example, the total number of candidate resources in the marked candidate resource set is M. In the present disclosure, multiple initial resources can be sorted from high to low according to the values of the resource scores to obtain a sorting sequence, and a candidate resource set can be generated based on the M initial resources ranked first in the sorting sequence.
[0083] As another example, the total number of candidate resources in the candidate resource set is marked as M. In the present disclosure, the screening ratio of each resource type can be configured, and the screening number of each resource type can be determined based on M and the screening ratio, so that the candidate resource set can be determined from multiple initial resources based on the screening number of each resource type and the resource score of the initial resources under each resource type.
[0084] For example, for any resource type, the number of screenings for that resource type is marked as M. i , where i is a positive integer and the value of i is not greater than the number of resource types (or species), then in this disclosure, the initial resources under the resource type can be sorted from high to low according to the resource score value to obtain a sorted sequence, and the M ranked at the top of the sorted sequence is selected from the sorted sequence. i initial resources, so in the present disclosure, a candidate resource set can be generated according to the initial resources selected under each resource type.
[0085] In summary, based on deep learning technology, predicting the resource scores of massive initial resources (which represent the quality of the initial resources and the target object's preference for the initial resources) can improve the prediction efficiency and the accuracy of the prediction results. Then, based on the resource scores of massive initial resources, resources can be screened to obtain a candidate resource set, which can ensure that the screened candidate resource set can take into account both resource quality and the target object's resource preference.
[0086] Step S202 : obtaining a record feature sequence; wherein the record feature sequence includes record features of at least one resource interaction record of the target object within a set time period.
[0087] For explanation of step S202, please refer to the relevant description in any embodiment of the present disclosure, and will not be repeated here.
[0088] Step S203 , obtaining a resource score of each candidate resource in the candidate resource set; wherein the resource score is used to indicate the quality of the candidate resource and the target object's preference for the candidate resource.
[0089] In an embodiment of the present disclosure, a resource score of each candidate resource in the candidate resource set may be obtained; wherein the resource score of each candidate resource is used to indicate the quality of the candidate resource and the target object's preference for the candidate resource.
[0090] As an example, a trained resource recommendation model can be used to predict the resource score of each candidate resource. For example, for any candidate resource, the resource recommendation model can be used to predict the resource score of the candidate resource based on the object characteristics of the target object and the resource characteristics of the candidate resource.
[0091] Step S204: determining a type score of the same resource type according to the average of the resource scores of the candidate resources belonging to the same resource type.
[0092] Exemplarily, for any resource type, the average of the resource scores of multiple candidate resources belonging to the resource type can be used as the score of the resource type (referred to as the type score in this disclosure); wherein the type score is used to indicate the target object's preference for the resource type.
[0093] Step S205 : Generate resource allocation proportions of the various resource types based on the object features, the resource feature sequence, the record feature sequence, and the type scores of the various resource types.
[0094] In the embodiment of the present disclosure, deep learning technology can be used to predict the resource allocation proportions of multiple resource types preferred by the target object based on the object characteristics, resource feature sequence, record feature sequence and type scores of multiple resource types of the target object.
[0095] Step S206 : determining a resource recommendation list from the candidate resource set according to the resource allocation proportions of the multiple resource types.
[0096] For explanation of step S206, please refer to the relevant description in any embodiment of the present disclosure, which will not be repeated here.
[0097] In any one of the embodiments of the present disclosure, the method for determining the resource recommendation list is, for example: first, the number of resources recommended to the target object is obtained, and then the recommended number of each resource type can be calculated based on the number of resources and the resource allocation ratio of multiple resource types. Therefore, in this application, the resource recommendation list can be determined from the candidate resources of multiple types of resources based on the recommended numbers of multiple resource types and the resource scores of the candidate resources of multiple resource types.
[0098] As an example, the number of marked resources is S, and the resource allocation proportion of any resource type is q%, then the recommended number of this resource type can be S*q%. In this disclosure, the candidate resources can be concentrated on the candidate resources of this resource type, and sorted from high to low according to the value of the resource score to obtain a sorting sequence, so that a resource recommendation list can be generated based on the S*q% candidate resources ranked at the top in the sorting sequence.
[0099] In summary, the resource recommendation list recommended to the target object is determined by comprehensively considering the resource allocation ratio of each resource type and the resource score of each candidate resource under each resource type. This can not only improve the resource quality in the resource recommendation list, but also ensure that the resource allocation method of the recommended resource recommendation list can meet the user's personalized resource preference needs and improve the user's resource consumption experience.
[0100] Step S207: recommending a resource recommendation list to the target object.
[0101] In an embodiment of the present disclosure, the resource recommendation list may be recommended to the target object. For example, the resource recommendation list may be pushed to an account bound to the target object.
[0102] The resource recommendation method of the embodiment of the present disclosure predicts the resource allocation proportions of multiple resource types preferred by the target object based on the object characteristics of the target object, the resource characteristics and resource scores of candidate resources of multiple resource types (characterizing the quality of the candidate resources and the target object's preference for the candidate resources), and the record characteristics of the resource interaction records that characterize the target object's resource preferences or immediate preferences within a set time period. This can improve the rationality and accuracy of the resource allocation proportion prediction, thereby further meeting the target object's personalized resource allocation needs and resource preference needs, and improving the accuracy of resource recommendations in the resource recommendation list, thereby improving the user experience in the resource recommendation scenario.
[0103] In order to clearly illustrate how the resource allocation ratios of multiple resource types are generated based on the object characteristics of the target object, the resource feature sequence, the record feature sequence and the type scores of multiple resource types in the above embodiment, the present disclosure also proposes a resource recommendation method.
[0104] Figure 3 This is a flowchart of the resource recommendation method provided in Example 3 of the present disclosure.
[0105] like Figure 3 As shown, the resource recommendation method may include the following steps S301 to S308:
[0106] Step S301: Acquire the object characteristics, resource characteristic sequence, and record characteristic sequence of the target object.
[0107] The resource feature sequence includes resource features of candidate resources of various resource types in the candidate resource set.
[0108] The record feature sequence includes the record features of at least one resource interaction record of the target object within a set time period.
[0109] Step S302: Obtain the resource score of each candidate resource in the candidate resource set.
[0110] The resource score of any candidate resource is used to indicate the quality of the candidate resource and the target object's preference for the candidate resource.
[0111] Step S303: Determine the type score of the same resource type according to the average of the resource scores of the candidate resources belonging to the same resource type.
[0112] For explanations of steps S301 to S303 , reference may be made to the relevant descriptions in any embodiment of the present disclosure, and no further details will be given here.
[0113] Step S304: Encode the object feature and resource feature sequence using the first encoding network in the generator of the resource allocation model to obtain object resource features.
[0114] Among them, the resource allocation model can be a generative artificial intelligence model, which can include a generator and a discriminator (or discriminator).
[0115] There is no limitation on the network structure of the first coding network. For example, the first coding network may be a Transformer-based coding network.
[0116] In the embodiment of the present disclosure, the first encoding network in the generator in the resource allocation model can be used to encode the object feature and resource feature sequence of the target object to obtain the object resource feature that integrates the object feature and the resource feature.
[0117] Exemplarily, the first encoding network may use an attention mechanism to encode the object features and resource feature sequences of the target object to obtain object resource features.
[0118] Step S305: Use the second encoding network in the generator to encode the record feature sequence to obtain resource common features.
[0119] The network structure of the second coding network is not limited. For example, the second coding network can also be a Transformer-based coding network. It should be noted that the network structure of the second coding network can be the same as or different from the network structure of the first coding network, and this is not limited in the present embodiment.
[0120] Among them, the resource common characteristics refer to the common characteristics between multiple record characteristics in the record feature sequence. The resource common characteristics are used to characterize the immediate preferences of the target object within a set time period, or to characterize the development trend and potential needs of the target object's preferences.
[0121] In an embodiment of the present disclosure, the second encoding network in the generator may be used to encode the record feature sequence to obtain common features among multiple record features, which are referred to as resource common features in the present disclosure.
[0122] Step S306 , concatenating the object resource features, the resource common features, and the type scores of the multiple resource types to obtain concatenated features.
[0123] In the embodiment of the present disclosure, object resource features, resource common features, and type scores of multiple resource types may be sequentially spliced together to obtain spliced features.
[0124] Step S307: Use the prediction network in the generator to process the splicing features to obtain resource allocation ratios of multiple resource types.
[0125] There is no limitation on the network structure of the prediction network. For example, the prediction network may be a DNN (Deep Neural Network).
[0126] In the embodiment of the present disclosure, the prediction network in the generator can be used to process the splicing features to obtain the resource allocation ratios of multiple resource types.
[0127] Exemplarily, the prediction network may include multiple DNN networks (wherein the number of DNN networks is consistent with the number of resource types), and the spliced features may be input into multiple DNN networks for processing. The output of each DNN network passes through a softmax (normalized exponential function or exponential normalization function) layer to obtain the quota ratio that should be allocated to a resource type, which is recorded as the resource allocation ratio in this disclosure.
[0128] Step S308 : determining a resource recommendation list from the candidate resource set according to the resource allocation proportions of the multiple resource types, and recommending the resource recommendation list to the target object.
[0129] For explanation of step S308, please refer to the relevant description in any embodiment of the present disclosure, and will not be repeated here.
[0130] The resource recommendation method of the embodiment of the present disclosure adopts deep learning technology, and predicts the resource allocation proportions of multiple resource types preferred by the target object based on the object characteristics of the target object, the resource characteristics and resource scores of candidate resources of multiple resource types (characterizing the quality of the candidate resources and the target object's preference for the candidate resources), and the record characteristics of the resource interaction records that characterize the target object's resource preferences or immediate preferences within a set time period, thereby improving the efficiency, rationality and accuracy of the resource allocation proportion prediction.
[0131] In order to clearly illustrate any embodiment of the present disclosure, the present disclosure also proposes a resource recommendation method.
[0132] Figure 4 This is a flowchart of the resource recommendation method provided in the fourth embodiment of the present disclosure.
[0133] like Figure 4 As shown, the resource recommendation method may include the following steps S401 to S407:
[0134] Step S401: Acquire the object characteristics and resource characteristics sequence of the target object.
[0135] The resource feature sequence includes resource features of candidate resources of various resource types in the candidate resource set.
[0136] For explanation of step S401, please refer to the relevant description in any embodiment of the present disclosure, and will not be repeated here.
[0137] In any embodiment of the present disclosure, the object features of the target object can be obtained by following steps A to C:
[0138] Step A: Obtain historical behavior information and preference setting information associated with the target object.
[0139] Among them, historical behavior information is used to indicate the records of the target object's interaction with resources or products in the historical period, search records, etc., such as the resources the target object has browsed, the links clicked, the products purchased, the videos watched, the keywords searched, etc.
[0140] Among them, preference setting information refers to the preference options actively set or selected by the target object, such as language preference, interface style, notification settings, content filtering conditions, etc.
[0141] Step B: Obtain object information and context information of the target object.
[0142] The target information includes but is not limited to the target's age, gender, occupation, education level and other basic information.
[0143] The context information includes, but is not limited to, the current situation information (such as time and location) of the target object, device information, network status, geographic location, etc.
[0144] Step C: Extract features from historical behavior information, preference setting information, object information, and context information to obtain object features of the target object.
[0145] Exemplarily, feature extraction technology may be used to extract features from historical behavior information, preference setting information, object information, and context information of the target object to obtain object features of the target object.
[0146] In summary, by building rich object features, we can fully understand the needs, preferences, and contexts of the target objects, thereby providing more personalized and accurate resource recommendation services, which can not only improve the satisfaction of the target objects, but also enhance the competitiveness and market share of the enterprise.
[0147] Step S402: Obtain at least one resource interaction record of the target object within a set time period.
[0148] It should be noted that the explanations on setting time periods and resource interaction records in the aforementioned embodiments are also applicable to this embodiment and are not described in detail here.
[0149] Step S403 : determining the target object's interaction satisfaction with the resource indicated by each resource interaction record based on the interaction information of multiple dimensions in each resource interaction record.
[0150] Among them, interactive information includes but is not limited to: browsing information (or playback information), comment information, sharing information, like information, etc.
[0151] In an embodiment of the present disclosure, for any resource interaction record, interaction indicators of multiple dimensions can be calculated based on the interaction information of multiple dimensions in the resource interaction record, and the target object's satisfaction with the resource indicated by the resource interaction record (referred to as interaction satisfaction in this disclosure) can be calculated based on the interaction indicators of multiple dimensions.
[0152] As an example, the weighted sum, average, and sum of interaction indicators in multiple dimensions can be used as the target object's interaction satisfaction with the resource indicated by the resource interaction record.
[0153] Step S404: determining a target interaction record from each resource interaction record according to the interaction satisfaction level of each resource interaction record.
[0154] In the embodiment of the present disclosure, resource interaction records with an interaction satisfaction level greater than a satisfaction level threshold may be regarded as target interaction records satisfying the target object.
[0155] Step S405: extract features from each target interaction record to obtain a record feature sequence.
[0156] In the embodiment of the present disclosure, feature extraction can be performed on each target interaction record based on a feature extraction algorithm to obtain the record features of each target interaction record, and a record feature sequence can be generated based on the record features of each target interaction record; wherein the record feature sequence includes the record features of each target interaction record.
[0157] Step S406: Generate resource allocation ratios of multiple resource types based on the object features, resource feature sequences, and record feature sequences.
[0158] Step S407 : determining a resource recommendation list from the candidate resource set according to the resource allocation proportions of the multiple resource types, and recommending the resource recommendation list to the target object.
[0159] The explanation of steps S406 to S407 can be found in the relevant description of any embodiment of the present disclosure and will not be repeated here.
[0160] The resource recommendation method of the embodiment of the present disclosure calculates the target object's interaction satisfaction with the resources indicated by each resource interaction record based on the target object's actual interaction information with the resources, which can improve the accuracy and rationality of the interaction satisfaction calculation, and thus determine the target interaction record that satisfies the target object from each resource interaction record based on the accurate interaction satisfaction. It can accurately screen out the target interaction record that satisfies the target object from massive resource interaction records, and then generate a record feature sequence based on the record features of the target interaction record that satisfies the target object. This can enable the record feature sequence to accurately represent the target object's personalized resource preference needs, thereby improving the accuracy and efficiency of subsequent resource recommendations.
[0161] The above are various embodiments corresponding to the application method of the resource allocation model (ie, the resource recommendation method). The present disclosure also proposes a training method for the resource allocation model.
[0162] Figure 5 This is a flowchart of the training method of the resource allocation model provided in the fifth embodiment of the present disclosure.
[0163] like Figure 5 As shown, the resource allocation model training method may include the following steps S501 to S505:
[0164] Step S501, obtaining a first sample; wherein the first sample includes a first object feature, a first resource feature sequence, and a first record feature sequence of a first sample object; the first resource feature sequence includes resource features of candidate resources of multiple resource types in a candidate resource set; and the first record feature sequence includes record features of at least one resource interaction record of the first sample object within a first sample period.
[0165] Among them, resource types include but are not limited to: long videos, short videos, dynamics, pictures and texts, etc.
[0166] The first sample object may be any user object. The method for obtaining the first object feature of the first sample object is similar to the method for obtaining the object feature of the target object, and will not be elaborated herein.
[0167] The first resource feature sequence is generated by extracting features from candidate resources of multiple resource types in the candidate resource set to obtain resource features of the candidate resources of multiple resource types, and based on the resource features of the candidate resources of multiple resource types.
[0168] Among them, the first sample period is a pre-set time period, and each resource interaction record (or resource consumption record) of the first sample object within the first sample period is used to record the resources that the first sample object has historically interacted with, as well as the first sample object's interaction information with the resource (including but not limited to: browsing time, comment information, sharing information, like information, etc.).
[0169] It should be noted that the implementation principle of step S501 is similar to that of steps S101 to S102, and will not be described in detail here.
[0170] Step S502 : Using a generator of a resource allocation model to generate resource allocation ratios of multiple resource types based on the first sample.
[0171] Among them, the resource allocation model can be a generative artificial intelligence model, which can include a generator and a discriminator (or discriminator).
[0172] In an embodiment of the present disclosure, a generator of a resource allocation model may be used to generate resource allocation ratios of multiple resource types based on the first sample.
[0173] Step S503 : determining a resource recommendation list from the candidate resource set according to resource allocation proportions of the multiple resource types.
[0174] It should be noted that the implementation principle of step S503 is similar to that of step S206 and will not be described in detail here.
[0175] Step S504 : using the discriminator of the resource allocation model to predict, based on the resource recommendation list, a first satisfaction level of the first sample subject with respect to the resource allocation mode in the resource recommendation list.
[0176] In an embodiment of the present disclosure, the discriminator of the resource allocation model can be used to predict the first sample object's satisfaction (or satisfaction probability, referred to as first satisfaction in the present disclosure) with the resource allocation method of the resource recommendation list based on the resource recommendation list.
[0177] Step S505: training the resource allocation model based on the first satisfaction level.
[0178] In the embodiment of the present disclosure, the resource allocation model may be trained based on the first satisfaction level.
[0179] In any embodiment of the present disclosure, the discriminator in the resource allocation model may be a pre-trained discriminator. In this case, the generator of the resource allocation model may be trained based on the first satisfaction level. For example, the first satisfaction level output by the discriminator may be used as a reward value, and an optimization algorithm such as reinforcement learning or gradient descent may be used to adjust the model parameters in the generator. Through continuous iterative training, the generator may gradually learn how to generate a resource allocation method that maximizes user satisfaction (or satisfaction probability), thereby improving the prediction accuracy of the resource allocation model and satisfying the personalized resource preferences of different users.
[0180] The training method of the resource allocation model of the embodiment of the present disclosure introduces a generator + discriminator training paradigm. This paradigm decouples the resource generation and evaluation processes, allowing the resource allocation model to more flexibly adapt to complex and changing user preferences. This not only improves the prediction accuracy of the resource allocation model, but also enhances the adaptability and robustness of the resource allocation model in practical applications.
[0181] In order to clearly illustrate how a resource allocation model generator is used in any embodiment of the present disclosure to generate resource allocation ratios of multiple resource types based on a first sample, the present disclosure also proposes a resource allocation model training method.
[0182] Figure 6 This is a flowchart of the training method of the resource allocation model provided in Example 6 of the present disclosure.
[0183] like Figure 6 As shown, the resource allocation model training method may include the following steps S601 to S606:
[0184] Step S601: obtain a first sample and obtain a resource score of each candidate resource in a candidate resource set.
[0185] Among them, the first sample includes the first object feature, the first resource feature sequence and the first record feature sequence of the first sample object; the first resource feature sequence contains the resource features of candidate resources of multiple resource types in the candidate resource set; the first record feature sequence contains the record features of at least one resource interaction record of the first sample object within the first sample period.
[0186] The resource score of any candidate resource is used to indicate the quality of the candidate resource and the preference of the first sample subject for the candidate resource.
[0187] It should be noted that the implementation principle of step S601 is similar to the implementation principle of steps S201 to S203, and will not be described in detail here.
[0188] Step S602: Determine a first score of the same resource type according to the average of the resource scores of the candidate resources belonging to the same resource type.
[0189] It should be noted that the implementation principle of step S602 is similar to that of step S204 and will not be described in detail here.
[0190] Step S603 : Using a generator to generate resource allocation proportions of multiple resource types according to the first object feature, the first resource feature sequence, the first record feature sequence, and the first scores of the multiple resource types.
[0191] It should be noted that the implementation principle of step S603 is similar to that of step S205 and will not be described in detail here.
[0192] In any embodiment of the present disclosure, the generator may include a first encoding network, a second encoding network, and a first prediction network. In the present disclosure, the first encoding network in the generator may be used to encode the first object feature and the first resource feature sequence to obtain a first encoding feature; the second encoding network in the generator may be used to encode the first record feature sequence to obtain a second encoding feature; the first encoding feature, the second encoding feature, and the first scores of multiple resource types may be spliced to obtain a first spliced feature; the first prediction network in the generator may be used to process the first spliced feature to obtain the resource allocation ratio of multiple resource types. The implementation principle is similar to that of steps S304 to S307 and will not be elaborated here.
[0193] Therefore, by adopting deep learning technology, the resource allocation proportions of multiple resource types preferred by the first sample object are predicted based on the first object characteristics of the first sample object, the resource characteristics and resource scores of candidate resources of multiple resource types (characterizing the quality of the candidate resources and the degree of preference of the first sample object for the candidate resources), and the record characteristics of the resource interaction records that characterize the resource preferences or immediate preferences of the first sample object within a set time period, thereby improving the efficiency, rationality and accuracy of the resource allocation proportion prediction.
[0194] Step S604 : determining a resource recommendation list from the candidate resource set according to resource allocation proportions of the multiple resource types.
[0195] Step S605 : using the discriminator of the resource allocation model to predict the first satisfaction of the first sample subject with respect to the resource allocation mode in the resource recommendation list according to the resource recommendation list.
[0196] Step S606: training the resource allocation model based on the first satisfaction level.
[0197] For explanations of steps S604 to S606 , reference may be made to the relevant descriptions in any embodiment of the present disclosure, and no further details will be given here.
[0198] The training method of the resource allocation model of the embodiment of the present disclosure predicts the resource allocation proportions of multiple resource types preferred by the first sample object based on the first object characteristics of the first sample object, the resource characteristics and resource scores of candidate resources of multiple resource types (characterizing the quality of the candidate resources and the degree of preference of the first sample object for the candidate resources), and the record characteristics of the resource interaction records that characterize the resource preferences or immediate preferences of the first sample object within a set time period. This can improve the rationality and accuracy of the resource allocation proportion prediction, thereby improving the prediction accuracy of the resource allocation model, meeting the personalized resource allocation needs and resource preference needs of different users, and thereby improving the user experience in the resource recommendation scenario.
[0199] In order to clearly illustrate how the discriminator of the resource allocation model in any embodiment of the present disclosure uses the resource recommendation list to predict the first satisfaction of the first sample object with the resource allocation method of the resource recommendation list, the present disclosure also proposes a training method for the resource allocation model.
[0200] Figure 7 This is a flowchart of the training method of the resource allocation model provided in Example 7 of the present disclosure.
[0201] like Figure 7 As shown, the resource allocation model training method may include the following steps S701 to S706:
[0202] Step S701: obtain a first sample, and use a generator of a resource allocation model to generate resource allocation ratios of multiple resource types based on the first sample.
[0203] Among them, the first sample includes the first object feature, the first resource feature sequence and the first record feature sequence of the first sample object; the first resource feature sequence contains the resource features of candidate resources of multiple resource types in the candidate resource set; the first record feature sequence contains the record features of at least one resource interaction record of the first sample object within the first sample period.
[0204] Step S702 : determining a resource recommendation list from the candidate resource set according to resource allocation proportions of multiple resource types.
[0205] The explanation of steps S701 to S702 can be found in the relevant description of any embodiment of the present disclosure and will not be repeated here.
[0206] Step S703 : obtaining a resource score of each recommended resource in the resource recommendation list, and determining a second score of the same resource type according to an average of the resource scores of the recommended resources belonging to the same resource type.
[0207] The resource score of any recommended resource is used to indicate the quality of the recommended resource and the preference of the first sample subject for the recommended resource.
[0208] It should be noted that the implementation principle of step S703 is similar to that of steps S601 to S602, and will not be described in detail here.
[0209] Step S704: Obtain a second resource feature sequence; wherein the second resource feature sequence includes resource features of each recommended resource in the resource recommendation list.
[0210] It should be noted that the method for obtaining the second resource feature sequence is similar to the method for obtaining the first resource feature sequence, and will not be described in detail here.
[0211] Step S705 : input the second scores of the multiple resource types, the second resource feature sequence, the first record feature sequence, and the first object feature into a discriminator to obtain a first satisfaction score output by the discriminator.
[0212] The first satisfaction level is used to indicate the probability that the first sample object is satisfied with the resource allocation method in the resource recommendation list.
[0213] In the embodiment of the present disclosure, the second scores, second resource feature sequences, first record feature sequences and first object features of multiple resource types may be input into the discriminator for processing to obtain a first satisfaction level output by the discriminator.
[0214] In any embodiment of the present disclosure, the discriminator may include a third encoding network, a fourth encoding network, and a second prediction network. In the present disclosure, the discriminator may use the following steps a to d to predict a first satisfaction level of the first sample subject with respect to the resource allocation method of the resource recommendation list:
[0215] Step a: Use the third encoding network of the discriminator to encode the first object feature and the second resource feature sequence to obtain a third encoding feature. The implementation principle is similar to that of step S304 and will not be repeated here.
[0216] There is no restriction on the network structure of the third coding network. For example, the third coding network may be a Transformer-based coding network.
[0217] Step b: The discriminator's fourth encoding network is used to encode the first record feature sequence to obtain a fourth encoding feature. The fourth encoding feature refers to a common feature among multiple record features in the first record feature sequence. This fourth encoding feature is used to characterize the immediate preferences of the first sample subject during the first sample period, or to characterize the development trend and potential demand of the first sample subject's preferences. The implementation principle is similar to that of step S305 and is not further described here.
[0218] The network structure of the fourth coding network is not limited. For example, the fourth coding network can also be a Transformer-based coding network. It should be noted that the network structure of the fourth coding network can be the same as or different from the network structure of the third coding network, and this is not limited in the embodiments of the present disclosure.
[0219] Step c: Concatenate the third coding feature, the fourth coding feature, and the second scores of the multiple resource types to obtain a second concatenated feature. The implementation principle is similar to that of step S306 and will not be described in detail here.
[0220] Step d: Processing the second concatenated features using the second prediction network of the discriminator to obtain a first satisfaction score.
[0221] Exemplarily, the second prediction network can be a DNN network, and the second splicing feature can be input into the DNN network for processing. The output of the DNN network passes through the softmax layer to obtain the probability that the first sample object is satisfied with the resource allocation method of the resource recommendation list, which is recorded as the first satisfaction in this disclosure.
[0222] In summary, by using deep learning technology, and simultaneously predicting the first satisfaction of the first sample object with the resource allocation methods of various recommended resources based on the first object characteristics of the first sample object, the resource characteristics and resource scores of recommended resources of multiple resource types (characterizing the quality of the recommended resources and the degree of preference of the first sample object for the recommended resources), and the record characteristics of the resource interaction records that characterize the resource preferences or immediate preferences of the first sample object during the first sample period, the efficiency, rationality and accuracy of the first satisfaction prediction can be improved.
[0223] Step S706: training the resource allocation model based on the first satisfaction level.
[0224] For explanation of step S706, please refer to the relevant description in any embodiment of the present disclosure, and will not be repeated here.
[0225] The training method of the resource allocation model of the embodiment of the present disclosure predicts the first satisfaction of the first sample object with the resource allocation method of the resource recommendation list based on the first object characteristics of the first sample object, the resource characteristics and resource scores of various recommended resources in the resource recommendation list (characterizing the quality of the recommended resources and the degree of preference of the first sample object for the recommended resources), and the record characteristics of the resource interaction records that characterize the resource preferences or immediate preferences of the first sample object in the first sample period, thereby improving the rationality and accuracy of the first satisfaction prediction.
[0226] In any embodiment of the present disclosure, the discriminator in the resource allocation model may be a trained discriminator. In order to clearly illustrate how the discriminator is trained in any embodiment of the present disclosure, the present disclosure also proposes a training method for the resource allocation model.
[0227] Figure 8 This is a flowchart of the training method of the resource allocation model provided in the eighth embodiment of the present disclosure.
[0228] like Figure 8 As shown, the discriminator in the resource allocation model can be trained by following steps S801 to S803:
[0229] Step S801, obtaining a second sample; wherein the second sample includes a second object feature, a third resource feature sequence, and a second record feature sequence of a second sample object; the third resource feature sequence includes resource features of interactive resources of multiple resource types in an interactive resource set; and the second record feature sequence includes record features of at least one resource interaction record of the second sample object within a second sample period.
[0230] The second sample object and the first sample object may be the same user object, or may be different user objects, which is not limited in the embodiment of the present disclosure.
[0231] It should be noted that the implementation principle of step S801 is similar to that of step S501 and will not be described in detail here.
[0232] Step S802: using a discriminator to predict a second satisfaction level of a second sample object with respect to a resource allocation mode of an interactive resource set based on the second sample.
[0233] In an embodiment of the present disclosure, the second sample may be input into the discriminator so that the discriminator can predict the satisfaction (or satisfaction probability, referred to as second satisfaction in the present disclosure) of the second sample object with respect to the resource allocation mode of the interactive resource set.
[0234] In any embodiment of the present disclosure, the discriminator may use the following steps 1 to 3 to predict the second satisfaction of the second sample object with respect to the resource allocation mode of the interactive resource set:
[0235] Step 1: Obtain the resource score of each interactive resource in the interactive resource set; wherein the resource score of any interactive resource is used to indicate the quality of the interactive resource and the preference of the second sample object for the interactive resource. The implementation principle is similar to step S203 and will not be repeated here.
[0236] Step 2: Determine the third score of the same resource type based on the average of the resource scores of the interactive resources of the same resource type. The implementation principle is similar to that of step S204 and will not be described in detail here.
[0237] Step 3: Input the third scores, third resource feature sequences, second record feature sequences, and second object features of the multiple resource types into the discriminator to obtain the second satisfaction level output by the discriminator. The implementation principle is similar to step S705 and will not be described in detail here.
[0238] Therefore, based on the second object characteristics of the second sample object, the resource characteristics and resource scores of various interactive resources in the interactive resource set (characterizing the quality of the interactive resources and the degree of preference of the second sample object for the interactive resources), and the record characteristics of the resource interaction records that characterize the resource preferences or immediate preferences of the second sample object in the second sample period, the second satisfaction of the second sample object with the resource allocation method of the interactive resource set is predicted, which can improve the rationality and accuracy of the second satisfaction prediction.
[0239] As an example, the discriminator may include a third encoding network, a fourth encoding network, and a second prediction network. In the present disclosure, the discriminator may use the following steps 3.1 to 3.4 to predict the second satisfaction of the second sample object with respect to the resource allocation mode of the interactive resource set:
[0240] Step a: Use the third encoding network of the discriminator to encode the second object feature and the third resource feature sequence to obtain the fifth encoded feature. The implementation principle is similar to that of step S304 and will not be repeated here.
[0241] There is no restriction on the network structure of the third coding network. For example, the third coding network may be a Transformer-based coding network.
[0242] Step b: The discriminator's fourth encoding network is used to encode the second record feature sequence to obtain a sixth encoded feature. The sixth encoded feature refers to a common feature among multiple record features in the second record feature sequence. This sixth encoded feature is used to characterize the immediate preferences of the second sample subject during the second sample period, or to characterize the development trend and potential demand of the second sample subject's preferences. The implementation principle is similar to that of step S305 and is not further described here.
[0243] The network structure of the fourth coding network is not limited. For example, the fourth coding network can also be a Transformer-based coding network. It should be noted that the network structure of the fourth coding network can be the same as or different from the network structure of the third coding network, and this is not limited in the embodiments of the present disclosure.
[0244] Step c: Concatenate the fifth coding feature, the sixth coding feature, and the third scores of the multiple resource types to obtain a third concatenated feature. The implementation principle is similar to that of step S306 and will not be described in detail here.
[0245] Step d: Processing the third concatenated feature using the second prediction network of the discriminator to obtain a second satisfaction score.
[0246] Exemplarily, the second prediction network can be a DNN network, and the third splicing feature can be input into the DNN network for processing. The output of the DNN network passes through the softmax layer to obtain the satisfaction probability of the second sample object with the resource allocation method of the interactive resource set, which is recorded as the second satisfaction in this disclosure.
[0247] In summary, the encoding technology in deep learning technology is used to encode the second object characteristics of the second sample object, the resource characteristics of interactive resources of multiple resource types, and the record characteristics of the resource interaction records that characterize the resource preferences or immediate preferences of the second sample object in the second sample period, and the encoded characteristics are fused with the third scores of multiple resource types determined based on the resource scores of multiple interactive resources (characterizing the quality of the interactive resources and the degree of preference of the second sample object for the interactive resources) to obtain a third splicing feature. The third splicing feature is processed to obtain the second satisfaction of the second sample object with the resource allocation methods of various interactive resources, which can improve the efficiency, effectiveness, rationality and accuracy of the second satisfaction prediction.
[0248] Step S803 : training the discriminator according to the difference between the second satisfaction level and the true satisfaction level annotated by the second sample.
[0249] The real satisfaction level is determined based on the actual interaction situation of the second sample object with the interactive resources in the interactive resource set, and is used to indicate the actual satisfaction probability of the second sample object with the resource allocation mode of the interactive resource set.
[0250] In an embodiment of the present disclosure, the value of the loss function (referred to as the loss value in the present disclosure) can be determined based on the difference between the second satisfaction level and the true satisfaction level annotated by the second sample, so that the discriminator can be trained based on the loss value to minimize the loss value.
[0251] Among them, the loss value is positively correlated with the above difference, that is, the greater the difference, the greater the loss value, and conversely, the smaller the difference, the smaller the loss value.
[0252] It should be noted that the above only uses the training termination condition of the discriminator as an example of minimizing the loss value. In actual application, other training termination conditions can also be set. For example, the training termination condition can also be: the training time reaches the set time, the number of training times reaches the set number of times, etc. The present disclosure does not limit this.
[0253] The training method of the resource allocation model of the embodiment of the present disclosure pre-trains the discriminator based on a supervised training method, which can improve the prediction accuracy of the discriminator, thereby improving the prediction accuracy of the resource allocation model and enhancing the robustness of the resource allocation model in practical applications.
[0254] In any of the embodiments of the present disclosure, the present disclosure proposes a resource preference modeling technology, which constructs a more accurate, efficient, and practical application-oriented resource preference estimation model (referred to as a resource allocation model in the present disclosure) by comprehensively considering the resource quality of the candidate resource set, adjusting the estimation target of the model, enriching the feature dimensions, and using a model structure with stronger feature extraction capabilities. This will help improve the resource recommendation effect and user experience of the resource recommendation system and create greater value for the enterprise. The present disclosure aims to solve the problem of how to accurately capture and model the user's resource preferences in a highly interactive and immersive environment, so as to provide a more accurate basis for the resource allocation of the resource recommendation system. This technology is of great significance for improving user experience, optimizing resource utilization efficiency, and enhancing the degree of personalization of the resource recommendation system.
[0255] As an example, in order to solve the deficiencies in the prior art, the present disclosure proposes an improved resource preference estimation algorithm, the implementation principle of which can be as follows: Figure 9 As shown, it mainly includes the following two parts:
[0256] Part 1: Training the Evaluator: The Evaluator can be trained based on actual user interaction data on resources. The Evaluator is used to evaluate the quality of the generated resource recommendation list (or resource recommendation sequence).
[0257] At the beginning of the algorithm, an Evaluator needs to be trained first. This Evaluator is used to evaluate the user's satisfaction with the resource allocation method of a given set of candidate resources. The training of the Evaluator is based on the following key features:
[0258] 1. Representation of the candidate resource set, denoted as a resource feature sequence in this disclosure: The N1 candidate resources in the candidate resource set can be encoded using the Transformer encoder structure, and key information such as the content, type, and quality of the candidate resources can be extracted as resource features. Based on the resource features of all candidate resources, a resource feature sequence is generated.
[0259] 3. User representation, referred to as object features in this disclosure: Based on the user's historical behavior information, personal information, preference setting information, etc., a comprehensive representation of the user (i.e., object features) is constructed.
[0260] 4. A sequence representation containing the user's N2 satisfactory resource interaction records, referred to as a record feature sequence in this disclosure: Specifically focusing on the user's recent satisfactory resource interaction (or consumption) records, the common features of these resource interaction records are extracted through sequence modeling technology to reflect the user's immediate preferences.
[0261] The goal of the Evaluater is to predict whether the resource allocation method for a given set of candidate resources meets the user's playback or consumption satisfaction. To this end, a large amount of labeled sample data can be used to train the Evaluater so that it can accurately evaluate the user's satisfaction with different resource allocation methods.
[0262] The second part is the training of the Generator. The Generator can generate the resource allocation ratio preferred by the user. Then, the Evaluater scores the resource recommendation sequence determined based on this resource allocation ratio and uses it as the reward. Finally, the Generator adjusts the model parameters according to the reward to generate a more accurate resource allocation ratio preferred by the user.
[0263] The Generator is responsible for generating a list of resource recommendations that the user may prefer and optimizing the resource allocation ratio. Its training process includes the following steps:
[0264] 1. Encoding object features and resource feature sequences: Through the first encoding network, the object features and the resource feature sequence of resource features containing N candidate resources are encoded to obtain encoded features that fuse context information, user personalization information, and resource information. These features include but are not limited to the user's historical behavior patterns, current context information (such as time, location), the type and quality of candidate resources, etc.
[0265] 2. Encoding record feature sequences: Model the record features of the user's recently satisfied resource interactions through the second encoding network to obtain resource common features (reflecting the development trend and potential needs of the user's preferences).
[0266] 3. Estimation of resource allocation ratio: Concatenate the output of the first encoding network, the output of the second encoding network, and the scores of various resource types ( Figure 9 denoted as refined ranking Q in Figure 9 taking the prediction network including multiple DNNs as an example), and then process it through the prediction network ( taking the prediction network including multiple DNNs as an example). The output of each DNN passes through a softmax layer to obtain the quota ratio of a resource type to be allocated.
[0267] 4. Evaluater evaluation and feedback: Based on the estimated resource allocation ratios of various resource types, determine the resource recommendation list item list (including N3 recommended resources, where N3 < N1), and use the pre-trained Evaluater to evaluate the user's satisfaction (reward) with the resource allocation method of the resource recommendation list. This step realizes the real-time evaluation and feedback of the resource allocation method.
[0268] 5. Model parameter update: Based on the reward value output by the Evaluator, the Generator's model parameters are adjusted using optimization algorithms such as reinforcement learning or gradient descent. Through continuous iterative training, the Generator gradually learns how to generate resource allocation methods that maximize user satisfaction.
[0269] It should be noted that Figure 9 X in cls Refers to the CLS (Classification) label. In the field of deep learning, especially in the fields of NLP and computer vision, CLS is often used as an identifier or label for classification tasks. cls It uses the attention mechanism to cls The H cls User information and resource information have been integrated; X user refers to the object characteristics, H user It uses the attention mechanism to user The H user Also integrates resource information; X nid1 To X nidN1 Refers to the N1 resource features in the resource feature sequence (or the first resource feature sequence).
[0270] It should be noted that C cls The meaning of X cls Similar, I will not elaborate here; C nid1 to C nidN2 Refers to the N2 record features in the record feature sequence; X nid1 To X nidN3 It refers to the N3 resource features in the resource feature sequence after quota allocation (referred to as the second resource feature sequence in this disclosure).
[0271] In summary, by deeply analyzing the user's historical behavior information and contextual information, an accurate user preference model (referred to as a resource allocation model in this disclosure) is constructed and applied to multiple key stages of the resource recommendation system. In the funnel link from coarse sorting to fine sorting and the funnel link between fine sorting and re-sorting, the resource allocation ratio of each resource type is allocated based on user preferences to achieve efficient resource allocation and personalized recommendations.
[0272] Furthermore, by introducing the Generator+Evaluater training paradigm, this paradigm decouples the resource generation and evaluation processes, allowing the model to more flexibly adapt to complex and changing user preferences. The Generator is responsible for generating resource allocation ratios based on the user's historical behavior information and current contextual information, and for generating a resource recommendation list based on the resource allocation ratios, while the Evaluater focuses on evaluating the potential attractiveness and quality of the recommended resources in the resource recommendation list, thereby screening out the resources that best meet the user's preferences. This division of labor and cooperation model not only improves the model's prediction accuracy, but also enhances its adaptability and robustness in practical applications. Furthermore, the Transformer's Encoder can be used to capture changes and long-term trends in users' historically satisfactory resource interaction records, and by building a rich object feature system, the model can more accurately understand users' resource preference needs.
[0273] With the above Figures 1 to 4 Corresponding to the resource recommendation method provided in the embodiment, the present disclosure also provides a resource recommendation device. Figures 1 to 4 The resource recommendation method provided in the embodiment corresponds to the resource recommendation method, so the implementation of the resource recommendation method is also applicable to the resource recommendation device provided in the embodiment of the present disclosure, and will not be described in detail in the embodiment of the present disclosure.
[0274] Figure 10 This is a structural diagram of the resource recommendation device provided in Example 9 of the present disclosure.
[0275] like Figure 10 As shown, the resource recommendation device 1000 may include: a first acquisition module 1010 , a second acquisition module 1020 , a generation module 1030 and a determination module 1040 .
[0276] The first acquisition module 1010 is configured to acquire the object features and resource feature sequence of the target object; wherein the resource feature sequence includes resource features of candidate resources of multiple resource types in the candidate resource set;
[0277] The second acquisition module 1020 is used to acquire a record feature sequence; wherein the record feature sequence includes the record features of at least one resource interaction record of the target object within a set time period;
[0278] A generating module 1030 is used to generate resource allocation ratios of multiple resource types based on object features, resource feature sequences, and record feature sequences;
[0279] The determination module 1040 is configured to determine a resource recommendation list from the candidate resource set according to resource allocation ratios of multiple resource types, so as to recommend the resource recommendation list to the target object.
[0280] In a possible implementation of the embodiment of the present disclosure, the generation module 1030 is used to: obtain the resource score of each candidate resource in the candidate resource set; wherein the resource score is used to indicate the quality of the candidate resource and the target object's preference for the candidate resource; determine the type score of the same resource type based on the average of the resource scores of the candidate resources belonging to the same resource type; generate the resource allocation ratio of multiple resource types based on object characteristics, resource feature sequence, record feature sequence and type scores of multiple resource types.
[0281] In a possible implementation of the embodiment of the present disclosure, the generation module 1030 is used to: use the first encoding network in the generator of the resource allocation model to encode the object feature and resource feature sequence to obtain the object resource feature; use the second encoding network in the generator to encode the record feature sequence to obtain the resource common feature; splice the object resource feature, the resource common feature and the type scores of multiple resource types to obtain the spliced feature; use the prediction network in the generator to process the spliced feature to obtain the resource allocation ratio of multiple resource types.
[0282] In a possible implementation of the embodiment of the present disclosure, the determination module 1040 is used to: obtain the number of resources recommended to the target object; determine the recommended number of multiple resource types based on the number of resources and the resource allocation ratio of multiple resource types; determine a resource recommendation list from the candidate resources of multiple types of resources based on the recommended number of multiple resource types and the resource scores of the candidate resources of multiple resource types.
[0283] In a possible implementation of the embodiment of the present disclosure, the second acquisition module 1020 is used to: obtain at least one resource interaction record of the target object within a set time period; determine the target object's interaction satisfaction with the resources indicated by each resource interaction record based on the interaction information of multiple dimensions in each resource interaction record; determine the target interaction record from each resource interaction record based on the interaction satisfaction of each resource interaction record; and perform feature extraction on each target interaction record to obtain a record feature sequence.
[0284] In a possible implementation of the embodiment of the present disclosure, the candidate resource set is determined using the following modules:
[0285] A third acquisition module is used to obtain resource characteristics of multiple initial resources;
[0286] A prediction module, configured to predict resource scores of the plurality of initial resources based on the object characteristics and the resource characteristics of the plurality of initial resources using a resource allocation model;
[0287] The screening module is used to determine a candidate resource set from the multiple initial resources according to the resource scores of the multiple initial resources.
[0288] In a possible implementation of the embodiment of the present disclosure, the first acquisition module 1010 is used to: obtain historical behavior information and preference setting information associated with the target object; obtain object information and context information of the target object; perform feature extraction on the historical behavior information, preference setting information, object information and context information to obtain object features of the target object.
[0289] The resource recommendation device of the embodiment of the present disclosure combines the object characteristics of the target object, the resource characteristics of candidate resources of multiple resource types, and the record characteristics of the resource interaction records that characterize the user's resource preferences or immediate preferences within a set time period to predict the resource allocation ratios of multiple resource types preferred by the target object. This can improve the rationality and reliability of the resource allocation ratio prediction, and then recommend a resource recommendation list to the target object based on the reliable resource allocation ratio. This can ensure that the resource allocation method of the recommended resource recommendation list can meet the personalized resource preference needs of the target object, and improve the accuracy of resource recommendations in the resource recommendation list, thereby improving the user's usage experience in the resource recommendation scenario.
[0290] With the above Figures 5 to 8 Corresponding to the training method of the resource allocation model provided in the embodiment, the present disclosure also provides a training device for the resource allocation model. Figures 5 to 8 The training method of the resource allocation model provided in the embodiment corresponds to the embodiment, so the implementation method of the resource allocation model training method is also applicable to the training device of the resource allocation model provided in the embodiment of the present disclosure, and will not be described in detail in the embodiment of the present disclosure.
[0291] Figure 11 This is a structural diagram of the training device for the resource allocation model provided in the tenth embodiment of the present disclosure.
[0292] like Figure 11 As shown, the resource allocation model training device 1100 may include: a first acquisition module 1110 , a generation module 1120 , a determination module 1130 , a first prediction module 1140 and a first training module 1150 .
[0293] The first acquisition module 1110 is configured to acquire a first sample; wherein the first sample includes a first object feature of a first sample object, a first resource feature sequence, and a first record feature sequence; the first resource feature sequence includes resource features of candidate resources of multiple resource types in a candidate resource set; and the first record feature sequence includes record features of at least one resource interaction record of the first sample object within a first sample period.
[0294] A generating module 1120 is configured to generate resource allocation ratios of multiple resource types based on the first sample using a generator of a resource allocation model;
[0295] A determination module 1130 is configured to determine a resource recommendation list from the candidate resource set based on resource allocation ratios of multiple resource types;
[0296] A first prediction module 1140 is configured to predict a first satisfaction level of a first sample subject with respect to a resource allocation method in the resource recommendation list using a discriminator of the resource allocation model according to the resource recommendation list;
[0297] The first training module 1150 is configured to train the resource allocation model based on the first satisfaction level.
[0298] In a possible implementation of the embodiment of the present disclosure, the generation module 1120 is used to: obtain the resource score of each candidate resource in the candidate resource set; wherein the resource score is used to indicate the quality of the candidate resource and the preference of the first sample object for the candidate resource; determine the first score of the same resource type based on the average of the resource scores of the candidate resources belonging to the same resource type; and use a generator to generate the resource allocation ratio of multiple resource types based on the first object feature, the first resource feature sequence, the first record feature sequence and the first scores of multiple resource types.
[0299] In a possible implementation of the embodiment of the present disclosure, the generation module 1120 is used to: use the first encoding network in the generator to encode the first object feature and the first resource feature sequence to obtain a first encoding feature; use the second encoding network in the generator to encode the first record feature sequence to obtain a second encoding feature; splice the first encoding feature, the second encoding feature and the first scores of multiple resource types to obtain a first spliced feature; use the first prediction network in the generator to process the first spliced feature to obtain the resource allocation ratio of multiple resource types.
[0300] In a possible implementation of the embodiment of the present disclosure, the first prediction module 1140 is used to: obtain a resource score for each recommended resource in the resource recommendation list, and determine a second score for the same resource type based on the average of the resource scores of the recommended resources belonging to the same resource type; obtain a second resource feature sequence; wherein the second resource feature sequence includes resource features of each recommended resource in the resource recommendation list; input the second scores of multiple resource types, the second resource feature sequence, the first record feature sequence, and the first object feature into a discriminator to obtain a first satisfaction level output by the discriminator.
[0301] In a possible implementation of the embodiment of the present disclosure, the first prediction module 1140 is used to: use the third encoding network of the discriminator to encode the first object feature and the second resource feature sequence to obtain a third encoded feature; use the fourth encoding network of the discriminator to encode the first record feature sequence to obtain a fourth encoded feature; splice the third encoded feature, the fourth encoded feature and the second scores of multiple resource types to obtain a second spliced feature; and use the second prediction network of the discriminator to process the second spliced feature to obtain a first satisfaction level.
[0302] In a possible implementation of the embodiment of the present disclosure, the discriminator is a trained discriminator, and the discriminator is trained using the following modules:
[0303] A second acquisition module is configured to acquire a second sample; wherein the second sample includes a second object feature of a second sample object, a third resource feature sequence, and a second record feature sequence; the third resource feature sequence includes resource features of interactive resources of multiple resource types in the interactive resource set; and the second record feature sequence includes record features of at least one resource interaction record of the second sample object within the second sample period;
[0304] A second prediction module is configured to use the discriminator to predict, based on the second sample, a second satisfaction level of the second sample object with respect to the resource allocation mode of the interactive resource set;
[0305] The second training module is used to train the discriminator according to the difference between the second satisfaction level and the true satisfaction level marked by the second sample.
[0306] In a possible implementation of the embodiment of the present disclosure, the second prediction module is used to: obtain the resource score of each interactive resource in the interactive resource set; determine the third score of the same resource type based on the average of the resource scores of the interactive resources belonging to the same resource type; input the third scores, third resource feature sequences, second record feature sequences and second object features of multiple resource types into the discriminator to obtain a second satisfaction level output by the discriminator.
[0307] In a possible implementation of the embodiment of the present disclosure, the second prediction module is used to: use the third encoding network of the discriminator to encode the second object feature and the third resource feature sequence to obtain a fifth encoding feature; use the fourth encoding network of the discriminator to encode the second record feature sequence to obtain a sixth encoding feature; splice the fifth encoding feature, the sixth encoding feature and the third scores of multiple resource types to obtain a third spliced feature; and use the second prediction network of the discriminator to process the third spliced feature to obtain a second satisfaction level.
[0308] In a possible implementation of the embodiment of the present disclosure, the first training module 1150 is configured to train a generator of a resource allocation model using a first satisfaction level.
[0309] The training device for the resource allocation model of the embodiment of the present disclosure introduces a generator + discriminator training paradigm. This paradigm decouples the resource generation and evaluation processes, enabling the resource allocation model to more flexibly adapt to complex and changing user preferences. This not only improves the prediction accuracy of the resource allocation model, but also enhances the adaptability and robustness of the resource allocation model in practical applications.
[0310] In order to implement the above embodiments, the present disclosure also provides an electronic device, which may include at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the resource recommendation method or resource allocation model training method proposed in any of the above embodiments of the present disclosure.
[0311] In order to implement the above embodiments, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the resource recommendation method or resource allocation model training method proposed in any of the above embodiments of the present disclosure.
[0312] In order to implement the above embodiments, the present disclosure further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the resource recommendation method or resource allocation model training method proposed in any of the above embodiments of the present disclosure.
[0313] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0314] Figure 12 A schematic block diagram of an example electronic device that can be used to implement an embodiment of the present disclosure is shown. The electronic device may include the server and client in the above-mentioned embodiments. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0315] like Figure 12As shown, the electronic device 1200 includes a computing unit 1201, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 1202 or a computer program loaded from a storage unit 1207 into a RAM (Random Access Memory) 1203. Various programs and data required for the operation of the device 1200 can also be stored in the RAM 1203. The computing unit 1201, the ROM 1202, and the RAM 1203 are connected to each other via a bus 1204. An I / O (Input / Output) interface 1205 is also connected to the bus 1204.
[0316] Various components in device 1200 are connected to I / O interface 1205, including an input unit 1206, such as a keyboard and mouse; an output unit 1207, such as various types of displays and speakers; a storage unit 1208, such as a magnetic disk and optical disk; and a communication unit 1209, such as a network card, a modem, a wireless communication transceiver, etc. Communication unit 1209 allows device 1200 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0317] The computing unit 1201 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 1201 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units that run machine learning model algorithms, a DSP (Digital Signal Processor), and any appropriate processor, controller, microcontroller, etc. The computing unit 1201 performs the various methods and processes described above, such as the resource recommendation method or resource allocation model training method described above. For example, in some embodiments, the resource recommendation method or resource allocation model training method described above can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1208. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 1200 via the ROM 1202 and / or the communication unit 1209. When the computer program is loaded into RAM 1203 and executed by computing unit 1201, one or more steps of the resource recommendation method or resource allocation model training method described above may be performed. Alternatively, in other embodiments, computing unit 1201 may be configured to perform the resource recommendation method or resource allocation model training method described above in any other appropriate manner (e.g., via firmware).
[0318] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System on Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0319] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0320] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0321] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0322] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.
[0323] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. The client-server relationship is established by computer programs running on the respective computers and establishing a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service ecosystem that addresses the management difficulties and poor scalability of traditional physical hosts and VPS services. The server may also be a server in a distributed system or a server integrated with blockchain.
[0324] It's important to note that artificial intelligence (AI) is the study of how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). This encompasses both hardware and software technologies. AI hardware technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily encompass computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graphs.
[0325] According to the technical solution of the embodiment of the present disclosure, the resource allocation ratios of the various resource types preferred by the target object are predicted by combining the object characteristics of the target object, the resource characteristics of the candidate resources of various resource types, and the record characteristics of the resource interaction records that characterize the user's resource preferences or immediate preferences within a set time period. This can improve the rationality and reliability of the resource allocation ratio prediction, and then recommend a resource recommendation list to the target object based on the reliable resource allocation ratio. This can enable the resource allocation method of the recommended resource recommendation list to meet the personalized resource preference needs of the target object, and improve the accuracy of the resource recommendation in the resource recommendation list, thereby improving the user's usage experience in the resource recommendation scenario.
[0326] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0327] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A method for training a resource allocation model, comprising: Obtaining a first sample; wherein the first sample includes a first object feature of a first sample object, a first resource feature sequence, and a first record feature sequence; the first resource feature sequence includes resource features of first candidate resources of multiple resource types in a first candidate resource set; and the first record feature sequence includes record features of at least one resource interaction record of the first sample object within a first sample period; Using the generator of the resource allocation model to generate first resource allocation ratios of the multiple resource types according to the first sample; Determining a first resource recommendation list from the first candidate resource set according to the first resource allocation proportions of the multiple resource types; determining a second score for the same resource type based on an average of the resource scores of the recommended resources of the same resource type in the first resource recommendation list; wherein the resource score of any recommended resource is used to indicate the quality of the recommended resource and the preference of the first sample subject for the recommended resource; Inputting the second scores of the multiple resource types, the second resource feature sequence, the first record feature sequence, and the first object feature into a discriminator of the resource allocation model, and obtaining a first satisfaction level of the first sample object with respect to the resource allocation method of the first resource recommendation list output by the discriminator; wherein the second resource feature sequence includes the resource features of each of the recommended resources in the first resource recommendation list; The resource allocation model is trained based on the first satisfaction level.
2. The method according to claim 1, wherein The generator using the resource allocation model generates first resource allocation ratios of the multiple resource types according to the first sample, including: Obtaining a resource score of each first candidate resource in the first candidate resource set; wherein the resource score of the first candidate resource is used to indicate the quality of the first candidate resource and the preference of the first sample subject for the first candidate resource; Determining a first score of the same resource type according to an average of resource scores of first candidate resources belonging to the same resource type; The generator is used to generate first resource allocation ratios of the multiple resource types according to the first object feature, the first resource feature sequence, the first record feature sequence, and the first scores of the multiple resource types.
3. The method according to claim 2, wherein: The generator is used to generate first resource allocation proportions of the multiple resource types according to the first object feature, the first resource feature sequence, the first record feature sequence, and the first scores of the multiple resource types, including: Encoding the first object feature and the first resource feature sequence using a first encoding network in the generator to obtain a first encoding feature; Encoding the first record feature sequence using a second encoding network in the generator to obtain a second encoding feature; concatenating the first coding feature, the second coding feature, and the first scores of the multiple resource types to obtain a first concatenated feature; The first splicing feature is processed using a first prediction network in the generator to obtain first resource allocation ratios of the multiple resource types.
4. The method according to claim 1, wherein The step of inputting the second scores of the plurality of resource types, the second resource feature sequence, the first record feature sequence, and the first object feature into a discriminator of the resource allocation model, and obtaining a first satisfaction level of the first sample object with respect to the resource allocation mode of the first resource recommendation list output by the discriminator, includes: encoding the first object feature and the second resource feature sequence using a third encoding network of the discriminator to obtain a third encoded feature; encoding the first record feature sequence using a fourth encoding network of the discriminator to obtain a fourth encoding feature; concatenating the third coding feature, the fourth coding feature, and the second scores of the multiple resource types to obtain a second concatenated feature; The second splicing feature is processed using a second prediction network of the discriminator to obtain the first satisfaction level.
5. The method according to any one of claims 1 to 4, wherein The discriminator is a trained discriminator, and the discriminator is trained by the following steps: Obtaining a second sample; wherein the second sample includes a second object feature of a second sample object, a third resource feature sequence, and a second record feature sequence; the third resource feature sequence includes resource features of interactive resources of multiple resource types in the interactive resource set; and the second record feature sequence includes record features of at least one resource interaction record of the second sample object within a second sample period; using the discriminator to predict, based on the second sample, a second satisfaction level of the second sample object with respect to the resource allocation mode of the interactive resource set; The discriminator is trained according to the difference between the second satisfaction level and the true satisfaction level labeled by the second sample.
6. The method according to claim 5, wherein: The step of using the discriminator to predict, based on the second sample, a second satisfaction level of the second sample object with respect to the resource allocation mode of the interactive resource set includes: Obtaining a resource score of each interactive resource in the interactive resource set; determining a third score of the same resource type according to an average of resource scores of interactive resources belonging to the same resource type; The third scores of the multiple resource types, the third resource feature sequence, the second record feature sequence, and the second object feature are input into the discriminator to obtain a second satisfaction level output by the discriminator.
7. The method according to claim 6, wherein: The step of inputting the third scores of the plurality of resource types, the third resource feature sequence, the second record feature sequence, and the second object feature into the discriminator to obtain a second satisfaction level output by the discriminator includes: encoding the second object feature and the third resource feature sequence using a third encoding network of the discriminator to obtain a fifth encoded feature; encoding the second record feature sequence using a fourth encoding network of the discriminator to obtain a sixth encoding feature; concatenating the fifth coding feature, the sixth coding feature, and the third scores of the multiple resource types to obtain a third concatenated feature; The third splicing feature is processed using a second prediction network of the discriminator to obtain the second satisfaction level.
8. The method according to claim 5, wherein The training of the resource allocation model according to the first satisfaction level includes: The first satisfaction level is used to train a generator of the resource allocation model.
9. A resource recommendation method, comprising: Obtaining an object feature and a resource feature sequence of a target object; wherein the resource feature sequence includes resource features of second candidate resources of multiple resource types in a second candidate resource set; Acquire a record feature sequence; wherein the record feature sequence includes record features of at least one resource interaction record of the target object within a set time period; A resource allocation model is used to generate a second resource allocation ratio for the multiple resource types according to the object feature, the resource feature sequence, and the record feature sequence; wherein the resource allocation model is trained using the training method according to any one of claims 1 to 8; According to the second resource allocation proportions of the multiple resource types, a second resource recommendation list is determined from the second candidate resource set, so as to recommend the second resource recommendation list to the target object.
10. The method according to claim 9, wherein: The adopting of the resource allocation model to generate the second resource allocation ratios of the multiple resource types according to the object feature, the resource feature sequence, and the record feature sequence includes: Obtaining a resource score of each second candidate resource in the second candidate resource set; wherein the resource score of the second candidate resource is used to indicate the quality of the second candidate resource and the target object's preference for the second candidate resource; Determining a type score of the same resource type according to an average of resource scores of second candidate resources belonging to the same resource type; The resource allocation model is adopted to generate second resource allocation proportions of the multiple resource types according to the object characteristics, the resource characteristic sequence, the record characteristic sequence, and the type scores of the multiple resource types.
11. The method according to claim 10, wherein: The adopting the resource allocation model to generate second resource allocation proportions for the multiple resource types according to the object feature, the resource feature sequence, the record feature sequence, and the type scores of the multiple resource types includes: Encoding the object feature and the resource feature sequence using a first encoding network in a generator of the resource allocation model to obtain object resource features; Using the second encoding network in the generator to encode the record feature sequence to obtain resource common features; Splicing the object resource feature, the resource common feature, and the type scores of the multiple resource types to obtain a spliced feature; The prediction network in the generator is used to process the splicing features to obtain second resource allocation ratios of the multiple resource types.
12. The method according to claim 10, wherein: The determining a second resource recommendation list from the second candidate resource set according to the second resource allocation proportions of the multiple resource types includes: Obtaining the number of resources recommended to the target object; Determining recommended quantities of the multiple resource types based on the resource quantity and the second resource allocation proportions of the multiple resource types; The second resource recommendation list is determined from the second candidate resources of the multiple resource types according to the recommendation quantities of the multiple resource types and the resource scores of the second candidate resources of the multiple resource types.
13. The method according to claim 9, wherein: The obtaining of the record feature sequence includes: Obtain at least one resource interaction record of the target object within the set time period; Determining the target object's interaction satisfaction with the resource indicated by each resource interaction record based on the interaction information of multiple dimensions in each resource interaction record; determining a target interaction record from each of the resource interaction records according to the interaction satisfaction level of each of the resource interaction records; Feature extraction is performed on each target interaction record to obtain the record feature sequence.
14. The method according to any one of claims 9 to 13, wherein: The second candidate resource set is determined by the following steps: Obtain resource characteristics of multiple initial resources; Using a resource recommendation model to predict resource scores of the multiple initial resources based on the object characteristics and the resource characteristics of the multiple initial resources; The second candidate resource set is determined from the multiple initial resources according to the resource scores of the multiple initial resources.
15. The method according to any one of claims 9 to 14, wherein: The object features of the target object are obtained by following the steps below: Obtaining historical behavior information and preference setting information associated with the target object; Acquire object information and context information of the target object; Feature extraction is performed on the historical behavior information, the preference setting information, the object information, and the context information to obtain object features of the target object.
16. A training device for a resource allocation model, comprising: A first acquisition module is configured to acquire a first sample; wherein the first sample includes a first object feature of a first sample object, a first resource feature sequence, and a first record feature sequence; the first resource feature sequence includes resource features of first candidate resources of multiple resource types in a first candidate resource set; and the first record feature sequence includes record features of at least one resource interaction record of the first sample object within a first sample period; A first generating module, configured to generate, by using the generator of the resource allocation model, first resource allocation ratios of the plurality of resource types according to the first sample; A first determining module, configured to determine a first resource recommendation list from the first candidate resource set according to first resource allocation proportions of the multiple resource types; A first prediction module is configured to determine a second score of the same resource type based on the average of the resource scores of the recommended resources of the same resource type in the first resource recommendation list, and input the second scores of the multiple resource types, the second resource feature sequence, the first record feature sequence, and the first object feature into a discriminator of the resource allocation model to obtain a first satisfaction score of the first sample subject with the resource allocation method of the first resource recommendation list output by the discriminator; wherein the resource score of any recommended resource is used to indicate the quality of the any recommended resource and the preference of the first sample subject for the any recommended resource; and the second resource feature sequence includes the resource features of each of the recommended resources in the first resource recommendation list; A first training module is used to train the resource allocation model based on the first satisfaction level.
17. The device according to claim 16, wherein The first generating module is configured to: Obtaining a resource score of each first candidate resource in the first candidate resource set; wherein the resource score of the first candidate resource is used to indicate the quality of the first candidate resource and the preference of the first sample subject for the first candidate resource; Determining a first score of the same resource type according to an average of resource scores of first candidate resources belonging to the same resource type; The generator is used to generate first resource allocation ratios of the multiple resource types according to the first object feature, the first resource feature sequence, the first record feature sequence, and the first scores of the multiple resource types.
18. The device according to claim 16, wherein The first generating module is configured to: Encoding the first object feature and the first resource feature sequence using a first encoding network in the generator to obtain a first encoding feature; Encoding the first record feature sequence using a second encoding network in the generator to obtain a second encoding feature; concatenating the first coding feature, the second coding feature, and the first scores of the multiple resource types to obtain a first concatenated feature; The first splicing feature is processed using a first prediction network in the generator to obtain first resource allocation ratios of the multiple resource types.
19. The device according to claim 16, wherein The first prediction module is used to: encoding the first object feature and the second resource feature sequence using a third encoding network of the discriminator to obtain a third encoded feature; encoding the first record feature sequence using a fourth encoding network of the discriminator to obtain a fourth encoding feature; concatenating the third coding feature, the fourth coding feature, and the second scores of the multiple resource types to obtain a second concatenated feature; The second splicing feature is processed using a second prediction network of the discriminator to obtain the first satisfaction level.
20. The device according to any one of claims 16 to 19, wherein The discriminator is a trained discriminator, which is trained using the following modules: A second acquisition module is configured to acquire a second sample; wherein the second sample includes a second object feature of a second sample object, a third resource feature sequence, and a second record feature sequence; the third resource feature sequence includes resource features of interactive resources of multiple resource types in the interactive resource set; and the second record feature sequence includes record features of at least one resource interaction record of the second sample object within a second sample period. a second prediction module, configured to use the discriminator to predict, based on the second sample, a second satisfaction level of the second sample object with respect to the resource allocation mode of the interactive resource set; The second training module is used to train the discriminator according to the difference between the second satisfaction level and the true satisfaction level marked by the second sample.
21. The device according to claim 20, wherein The second prediction module is used to: Obtaining a resource score of each interactive resource in the interactive resource set; determining a third score of the same resource type according to an average of resource scores of interactive resources belonging to the same resource type; The third scores of the multiple resource types, the third resource feature sequence, the second record feature sequence, and the second object feature are input into the discriminator to obtain a second satisfaction level output by the discriminator.
22. The device according to claim 21, wherein The second prediction module is used to: encoding the second object feature and the third resource feature sequence using a third encoding network of the discriminator to obtain a fifth encoded feature; encoding the second record feature sequence using a fourth encoding network of the discriminator to obtain a sixth encoding feature; concatenating the fifth coding feature, the sixth coding feature, and the third scores of the multiple resource types to obtain a third concatenated feature; The third splicing feature is processed using a second prediction network of the discriminator to obtain the second satisfaction level.
23. The apparatus according to claim 20, wherein The first training module is used to: The first satisfaction level is used to train a generator of the resource allocation model.
24. A resource recommendation device, comprising: A third acquisition module is configured to acquire an object feature and a resource feature sequence of a target object; wherein the resource feature sequence includes resource features of second candidate resources of multiple resource types in the second candidate resource set; A fourth acquisition module is configured to acquire a record feature sequence, wherein the record feature sequence includes record features of at least one resource interaction record of the target object within a set time period; a second generating module, configured to generate, based on the object features, the resource feature sequence, and the record feature sequence, a second resource allocation ratio for the plurality of resource types using a resource allocation model; wherein the resource allocation model is obtained by training using the training device according to any one of claims 16 to 23; The second determining module is configured to determine a second resource recommendation list from the second candidate resource set according to the second resource allocation proportions of the multiple resource types, so as to recommend the second resource recommendation list to the target object.
25. The apparatus according to claim 24, wherein The second generating module is used to: Obtaining a resource score of each second candidate resource in the second candidate resource set; wherein the resource score of the second candidate resource is used to indicate the quality of the second candidate resource and the target object's preference for the second candidate resource; Determining a type score of the same resource type according to an average of resource scores of second candidate resources belonging to the same resource type; The resource allocation model is adopted to generate second resource allocation proportions of the multiple resource types according to the object characteristics, the resource characteristic sequence, the record characteristic sequence, and the type scores of the multiple resource types.
26. The device according to claim 25, wherein The second generating module is used to: Encoding the object feature and the resource feature sequence using a first encoding network in a generator of the resource allocation model to obtain object resource features; Using the second encoding network in the generator to encode the record feature sequence to obtain resource common features; Splicing the object resource feature, the resource common feature, and the type scores of the multiple resource types to obtain a spliced feature; The prediction network in the generator is used to process the splicing features to obtain second resource allocation ratios of the multiple resource types.
27. The apparatus according to claim 25, wherein The second determining module is configured to: Obtaining the number of resources recommended to the target object; Determining recommended quantities of the multiple resource types based on the resource quantity and the second resource allocation proportions of the multiple resource types; The second resource recommendation list is determined from the second candidate resources of the multiple resource types according to the recommendation quantities of the multiple resource types and the resource scores of the second candidate resources of the multiple resource types.
28. The apparatus according to claim 24, wherein The fourth acquisition module is configured to: Obtain at least one resource interaction record of the target object within the set time period; Determining the target object's interaction satisfaction with the resource indicated by each resource interaction record based on the interaction information of multiple dimensions in each resource interaction record; determining a target interaction record from each of the resource interaction records according to the interaction satisfaction level of each of the resource interaction records; Feature extraction is performed on each target interaction record to obtain the record feature sequence.
29. The device according to any one of claims 24 to 28, wherein The second candidate resource set is determined by using the following modules: A fifth acquisition module, configured to acquire resource characteristics of a plurality of initial resources; a third prediction module, configured to predict resource scores of the plurality of initial resources based on the object features and the resource features of the plurality of initial resources using a resource recommendation model; A screening module is configured to determine the second candidate resource set from the multiple initial resources according to the resource scores of the multiple initial resources.
30. The device according to any one of claims 24 to 28, wherein The third acquisition module is used to: Obtaining historical behavior information and preference setting information associated with the target object; Acquire object information and context information of the target object; Feature extraction is performed on the historical behavior information, the preference setting information, the object information, and the context information to obtain object features of the target object.
31. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the resource recommendation method described in any one of claims 1-8, or execute the resource allocation model training method described in any one of claims 9-15.
32. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable the computer to execute the resource recommendation method according to any one of claims 1-8, or to execute the resource allocation model training method according to any one of claims 9-15.
33. A computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the resource recommendation method according to any one of claims 1 to 8, or, when executed, implements the resource allocation model training method according to any one of claims 9 to 15.
Citation Information
Patent Citations
Reading resource comment pushing method and system
CN107798012A
Model training method and device, resource recommendation method and device, electronic equipment and storage medium
CN114564644A