Model training method, resource recommendation method, sample generation method and device
Patent Information
- Application Number
- CN202310590388.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-23
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-05-23
AI Technical Summary
[0002]视频网站、图书网站等平台可以根据用户的历史行为,向用户推荐视频、文本等资源,然而目前资源推荐的效果较差,影响了用户体验
Smart Images

Figure CN116662652B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and more particularly to the fields of information flow and intelligent recommendation. More specifically, this disclosure provides a method for training a resource recommendation model, a resource recommendation method, a method for generating training samples, an apparatus, an electronic device, a storage medium, and a computer program product. Background Technology
[0002] Video websites, book websites, and other platforms can recommend resources such as videos and texts to users based on their historical behavior. However, the current resource recommendation effect is poor, which affects the user experience. Summary of the Invention
[0003] This disclosure provides a method for training a resource recommendation model, a resource recommendation method, a method for generating training samples, an apparatus, an electronic device, a storage medium, and a computer program product.
[0004] According to one aspect of this disclosure, a training method for a resource recommendation model is provided. The resource recommendation model includes a first sub-model and a second sub-model. The method includes: acquiring training samples; the training samples include object features of an object, first resource features of a first resource, second resource features of a second resource, a first label, a second label, and a third label, wherein the first label represents the difference between the object's preference for the first resource and the object's preference for the second resource, the second label represents the object's preference for the first resource, and the third label represents the object's preference for the second resource; processing the object features, first resource features, and second resource features using the first sub-model to obtain a first evaluation value; inputting the object features and first resource features into the second sub-model to obtain a second evaluation value; inputting the object features and second resource features into the second sub-model to obtain a third evaluation value; and training the first sub-model and the second sub-model based on a first difference between the first evaluation value and the first label, a second difference between the second evaluation value and the second label, and a third difference between the third evaluation value and the third label.
[0005] According to another aspect of this disclosure, a resource recommendation method is provided, comprising: determining a target object and a plurality of candidate resources to be recommended; for each candidate resource, processing the target object features of the target object and the candidate resource features of the candidate resource using a resource recommendation model to obtain a recommendation evaluation value for the candidate resource; determining the target resource from the plurality of candidate resources based on the plurality of recommendation evaluation values of the plurality of candidate resources; and recommending the target resource to the target object; wherein the resource recommendation model is trained using the above-described resource recommendation model training method.
[0006] According to another aspect of this disclosure, a method for generating training samples is provided, comprising: dividing multiple resources into multiple resource sets based on the behavior of an object toward multiple resources; the multiple resources being resources that have been shown to the object; and generating training samples based on at least one of the multiple resource sets; wherein the training samples include object features of the object, first resource features of a first resource, second resource features of a second resource, a first label, a second label, and a third label, the first label representing the difference between the object's preference for the first resource and the object's preference for the second resource, the second label representing the object's preference for the first resource, and the third label representing the object's preference for the second resource.
[0007] According to another aspect of this disclosure, a training apparatus for a resource recommendation model is provided. The resource recommendation model includes a first sub-model and a second sub-model. The apparatus includes: a sample acquisition module for acquiring training samples; the training samples include object features of an object, first resource features of a first resource, second resource features of a second resource, a first label, a second label, and a third label, wherein the first label represents the difference between the object's preference for the first resource and the object's preference for the second resource, the second label represents the object's preference for the first resource, and the third label represents the object's preference for the second resource; a first evaluation value determination module for processing the object features, the first resource features, and the second resource features using the first sub-model to obtain a first evaluation value; a second evaluation value determination module for inputting the object features and the first resource features into the second sub-model to obtain a second evaluation value; a third evaluation value determination module for inputting the object features and the second resource features into the second sub-model to obtain a third evaluation value; and a training module for training the first sub-model and the second sub-model based on a first difference between the first evaluation value and the first label, a second difference between the second evaluation value and the second label, and a third difference between the third evaluation value and the third label.
[0008] According to another aspect of this disclosure, a resource recommendation apparatus is provided, comprising: an information determination module for determining a target object and a plurality of candidate resources to be recommended; a recommendation evaluation value determination module for processing the target object features of the target object and the candidate resource features of the candidate resources using a resource recommendation model for each candidate resource to obtain a recommendation evaluation value for the candidate resource; a target resource determination module for determining a target resource from the plurality of candidate resources based on the plurality of recommendation evaluation values of the plurality of candidate resources; and a recommendation module for recommending the target resource to the target object; wherein the resource recommendation model is obtained using the above-described training.
[0009] According to another aspect of this disclosure, an apparatus for generating training samples is provided, comprising: a partitioning module for partitioning multiple resources into multiple resource sets based on the behavior of an object toward multiple resources; the multiple resources being resources that have been presented to the object; and a generation module for generating training samples based on at least one of the multiple resource sets; wherein the training samples include object features of the object, first resource features of a first resource, second resource features of a second resource, a first label, a second label, and a third label, the first label representing the difference between the object's preference for the first resource and the object's preference for the second resource, the second label representing the object's preference for the first resource, and the third label representing the object's preference for the second resource.
[0010] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the methods provided in this disclosure.
[0011] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause a computer to perform the methods provided in this disclosure.
[0012] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the methods provided in this disclosure.
[0013] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0014] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0015] Figure 1 This is a schematic diagram illustrating application scenarios of the method for generating training samples, the training method for the resource recommendation model, the resource recommendation method, and the apparatus according to embodiments of this disclosure.
[0016] Figure 2 This is a schematic flowchart of a method for generating training samples according to an embodiment of the present disclosure;
[0017] Figure 3 This is a schematic diagram illustrating the principle of partitioning a resource set according to an embodiment of this disclosure;
[0018] Figure 4This is a schematic flowchart of a training method for a resource recommendation model according to an embodiment of the present disclosure;
[0019] Figure 5 This is a schematic diagram illustrating the training method of a resource recommendation model according to an embodiment of the present disclosure;
[0020] Figure 6 This is a schematic flowchart of a resource recommendation method according to an embodiment of the present disclosure;
[0021] Figure 7 This is a schematic structural block diagram of an apparatus for generating training samples according to an embodiment of the present disclosure;
[0022] Figure 8 This is a schematic structural block diagram of a training device for a resource recommendation model according to an embodiment of the present disclosure;
[0023] Figure 9 This is a schematic structural block diagram of a resource recommendation device according to embodiments of the present disclosure; and
[0024] Figure 10 This is a structural block diagram of an electronic device used to implement the method for generating training samples, the method for training a resource recommendation model, and the method for recommending resources in accordance with the embodiments of this disclosure. Detailed Implementation
[0025] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0026] In some recommendation scenarios, a multi-factor integrated information flow recommendation method is adopted. This method determines the probability of the target user’s behavior towards candidate resources in multiple dimensions such as clicks, reading time, and interaction. Then, based on multiple probabilities in multiple dimensions, it determines the target user’s preference for candidate resources. Finally, it ranks multiple candidate resources based on preference and recommends the top-ranked candidate resources to the target user.
[0027] For example, we can determine the first probability that a target user will click on a candidate resource, the second probability that the target user will read the candidate resource for a long time, and the third probability that the target user will interact with the candidate resource. Then, based on the first, second, and third probabilities, we can determine the target user's preference for the candidate resource and make recommendations based on the preference.
[0028] However, the resources recommended to users using the above recommendation methods are the result of a multi-factor balance. Information flow recommendation methods that rely on a combination of factors are prone to the problem of a single factor having a significant impact, leading to the determined preference level failing to reflect the user's overall satisfaction with the candidate resources. For example, a candidate resource might be determined to have a high preference level, but in reality, it might have a high click probability but a short reading time, or a high interaction probability but a low click probability.
[0029] Therefore, the information flow recommendation method with the combined effect of the above-mentioned multi-factors has poor recommendation effect, users cannot obtain satisfactory resources, and the user experience is reduced.
[0030] This embodiment aims to provide a method for generating training samples, a method for training a resource recommendation model, and a method for recommending resources. This method comprehensively characterizes multi-dimensional user behavior information, models the user's overall satisfaction with resources, recommends resources with high satisfaction to users, alleviates the problem of poor recommendation effect caused by an excessively large single factor under the combined effect of multiple factors, and improves the overall user experience.
[0031] The method provided in this embodiment can be applied to information flow recommendation, and can also be more widely applied to various recommendation systems.
[0032] The technical solutions provided in this disclosure will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0033] Figure 1 This is a schematic diagram illustrating an application scenario of the method for generating training samples, the training method for a resource recommendation model, the resource recommendation method, and the apparatus according to embodiments of this disclosure.
[0034] like Figure 1 As shown, the application scenario 100 of this embodiment may include an electronic device 110, which may be any electronic device with processing capabilities, including but not limited to smartphones, tablets, laptops, desktop computers, and servers.
[0035] According to embodiments of this disclosure, such as Figure 1 As shown, the application scenario 100 may also include a server 140. The electronic device 110 can communicate with the server 140 via a network, which may include a wireless or wired communication link.
[0036] According to embodiments of this disclosure, such as Figure 1 As shown, the application scenario 100 may also include a database 160, which can maintain a massive number of training samples. These training samples may have labels, such as a first label, a second label, and a third label. Training samples can be generated using methods for generating training samples, and these training samples can be stored in the database 160.
[0037] For example, server 140 can be used to train a resource recommendation model. Server 140 can access database 160 and extract a portion of training samples from database 160 to train the resource recommendation model. When training resource recommendation model 150, the total loss of the resource recommendation model can be determined by using a loss function based on the first evaluation value, second evaluation value, and third evaluation value output by the model, as well as the first label, second label, and third label. The training of the model is completed by minimizing the total loss of the model.
[0038] For example, server 140 can be used to train a resource recommendation model and, in response to a model acquisition request sent by electronic device 110, send the trained resource recommendation model 150 to electronic device 110 to facilitate resource recommendation by electronic device 110. In one embodiment, the server can also determine the recommendation evaluation value of candidate resources based on the trained resource recommendation model.
[0039] For example, the electronic device 110 can determine the recommended evaluation value of the candidate resource based on the target object characteristics and candidate resource characteristics of the target object 120, and then determine the target resource 130 based on the recommended evaluation value and recommend it to the target object 120.
[0040] It should be noted that the methods for generating training samples, training the resource recommendation model, and recommending resources provided in this disclosure can be executed by the electronic device 110 or the server 140.
[0041] It should be understood that Figure 1 The number and types of electronic devices, servers, and databases shown are merely illustrative. Depending on implementation needs, any number and type of electronic devices, servers, and databases can be included.
[0042] The following combination Figures 2-3 The method for generating training samples will be explained.
[0043] Figure 2 This is a schematic flowchart of a method for generating training samples according to an embodiment of the present disclosure.
[0044] like Figure 2 As shown, the method 200 for generating training samples may include operations S210 to S220.
[0045] In operation S210, based on the actions of the object toward multiple resources, the multiple resources are divided into multiple resource sets; the multiple resources are the resources that have been presented to the object.
[0046] For example, the object can be a user.
[0047] For example, resources can include videos, text, images, music, etc. Resources are those that have already been displayed to an object, such as resources shown to a user using a display screen or other display device.
[0048] For example, user actions related to resources can include clicking, browsing, and interacting. Browsing can be categorized by duration into long-term and short-term browsing, while interactions can include liking, commenting, saving, and sharing. Accordingly, multiple resource sets can include a set of clicks, a first set of browsing with longer browsing time, a second set of browsing with shorter browsing time, and an interaction set. The interaction set can include a set of likes, a set of shares, etc.
[0049] In operation S220, training samples are generated based on at least one of a plurality of resource sets.
[0050] For example, training samples include: object features of the object, first resource features of the first resource, second resource features of the second resource, first label, second label, and third label. For example, the first and second resources can be selected from the same resource set, or from different resource sets.
[0051] For example, the first label represents the difference between the object's preference level P1 for the first resource and the object's preference level P2 for the second resource. The degree of preference can include satisfaction and dissatisfaction, and the difference between the two degree of preference P1 and P2 can be: which of the first and second resources the object is more satisfied with. For example, if the degree of preference P1 is higher than the degree of preference P2, it means that the object is more satisfied with the first resource.
[0052] For example, the second label represents the object's preference for the first resource, and the third label represents the object's preference for the second resource. For instance, the object's preference for a resource could be whether the user is satisfied with the resource; that is, the second label represents whether the object is satisfied with the first resource, and the third label represents whether the object is satisfied with the second resource.
[0053] According to the technical solution provided in this disclosure, the solution first divides the resource set, and then constructs resource pairs based on the resource set. A resource pair refers to a first resource and a second resource. Therefore, manual sample labeling is unnecessary, reducing labeling costs. Furthermore, the training samples can be used to train a resource recommendation model. The resource recommendation model trained using these training samples can accurately assess the overall preference of a target object for candidate resources.
[0054] Figure 3 This is a schematic diagram illustrating the principle of partitioning resource sets according to an embodiment of this disclosure.
[0055] The following combination Figure 3The method described above, which divides multiple resources into multiple resource sets based on the behavior generated by an object for multiple resources, will be explained.
[0056] For example, it can be determined whether the user clicked on a resource. If not, the resource is added to the display set 301. If the user clicked, it can be determined whether the user interacted with the resource. If an interaction occurred, the resource is added to the interaction set 304. If no interaction occurred, it can be determined whether the browsing time is greater than or equal to a predetermined time, such as 5 seconds. If it is greater than or equal to the predetermined time, the resource is added to the first browsing set 302; otherwise, the resource is added to the second browsing set 303.
[0057] It should be noted that this embodiment does not limit the order of the above judgments. Generally speaking, in response to the detection that an object has not clicked on a resource, the resource can be added to the display set. In response to the detection that an object has clicked on a resource and the object has interacted with the resource, the resource can be added to the interaction set. In response to the detection that an object has clicked on a resource, the object has not interacted with the resource, and the object's browsing time for the resource is greater than or equal to a predetermined time, the resource can be added to the first browsing set. In response to the detection that an object has clicked on a resource, the object has not interacted with the resource, and the object's browsing time for the resource is less than a predetermined time, the resource can be added to the second browsing set.
[0058] This embodiment of the disclosure divides resource sets according to the object's click, interaction, browsing and other behaviors. The above behaviors can accurately reflect the object's overall preference for resources. For example, the preference for the interaction set, the first browsing set, the display set and the second browsing set decreases in that order. Then, the resource recommendation model is trained using training samples generated from these resource sets, which can enable the resource recommendation model to accurately evaluate the target object's overall preference for candidate resources.
[0059] The method for determining the first resource, the second resource, and the first tag will be described below with reference to the embodiments.
[0060] In one example, a first resource and a second resource can be selected from different sets of resources.
[0061] It should be noted that resource sets can correspond to levels of preference, and the resource set to which a resource belongs can reflect the degree of user preference for that resource. For example, the degree of preference for the interaction set A, the first browsing set B, the display set C, and the second browsing set D decreases in that order. That is, the user's preference for resources that generate interactive behavior, resources that are viewed for a long time, resources that have been displayed, and resources that are viewed for a short time decreases in that order.
[0062] For example, one resource can be selected from any two resource sets among multiple resource sets, designated as the first resource and the second resource, and then the first tag can be determined based on the resource set to which the first resource belongs and the resource set to which the second resource belongs.
[0063] As can be seen, the first and second resources in the training samples constitute a resource pair. When determining resource pairs, a resource from the interaction set A and a resource from the first browsing set B can form a resource pair; a resource from the interaction set A and a resource from the display set C can form a resource pair; a resource from the interaction set A and a resource from the second browsing set D can form a resource pair; a resource from the first browsing set B and a resource from the display set C can form a resource pair; a resource from the first browsing set B and a resource from the second browsing set D can form a resource pair; and a resource from the display set C and the first resource from the second browsing set D can form a resource pair.
[0064] The value of the first label can be 1 or 0. 1 indicates that the user prefers the first resource more than the second resource, and 0 indicates that the user prefers the first resource less than the second resource.
[0065] This embodiment determines resource pairs from different resource sets and determines the value of the first label based on the resource set to which the resource belongs. Therefore, there is no need for manual sample labeling, training samples can be generated relatively conveniently, and the first label accurately represents the deviation between the object's preference for the first resource and the user's preference for the second resource.
[0066] In one example, a first resource and a second resource can be selected from the same set of resources.
[0067] For example, a first resource and a second resource can be determined from the browsing collection, and then a first tag can be determined based on the browsing duration corresponding to the first resource and the browsing duration corresponding to the second resource.
[0068] As can be seen, browsing time can reflect a user's level of preference for a resource from another dimension; that is, the longer the browsing time, the higher the user's preference. The value of the first tag can be 1 or 0. 1 indicates that the user's preference for the first resource is higher than their preference for the second resource, for example, the browsing time for the first resource is longer than the browsing time for the second resource. 0 indicates that the user's preference for the first resource is lower than their preference for the second resource, for example, the browsing time for the first resource is shorter than the browsing time for the second resource.
[0069] For example, a first resource and a second resource can be identified from the interaction set. Then, based on the interaction categories corresponding to the first resource and the second resource, a first tag can be determined. For instance, if the interaction category for the first resource is "share" and the interaction category for the second resource is "comment," the value of the first tag can be determined to be 1.
[0070] This embodiment determines resource pairs from the same resource set and determines the value of the first tag based on the browsing duration or interaction category of the resource. Therefore, there is no need for manual sample labeling, training samples can be generated relatively easily, and the first tag accurately represents the deviation between the object's preference for the first resource and the user's preference for the second resource.
[0071] The method for determining the first resource, the second resource, and the first tag has been described above. The method for determining the second tag and the third tag will be described below with reference to embodiments.
[0072] For example, regarding the first resource, if the object interacts with the first resource, it can be determined that the object is satisfied with the first resource. If the object does not interact with the first resource, and the object's completion rate for the first resource is greater than or equal to the completion rate threshold, it can be determined that the object is satisfied with the first resource. If the object does not interact with the first resource, and the object's completion rate for the first resource is less than the completion rate threshold, it can be determined that the object is dissatisfied with the first resource. If the object is satisfied with the first resource, the value of the second label can be 1. If the object is dissatisfied with the first resource, the value of the second label can be 0.
[0073] For the second resource, we can determine whether the object is satisfied with it, and then determine the third tag. The specific method for determining the third tag can be referred to the second tag, and will not be repeated here.
[0074] For example, when the resource is a text-based resource, the completion rate is determined based on the browsing time and the amount of text in the resource, such as using the ratio between browsing time and the amount of text as the completion rate.
[0075] For example, when the resource is a video resource, the completion rate is determined based on the viewing time and the video length of the resource, such as using the ratio between the viewing time and the video length as the completion rate.
[0076] In embodiments of this disclosure, when objects engage in mutual behavior or have a high completion rate, it is determined that the objects are satisfied with the resources, thus enabling accurate determination of the values of the second and third tags.
[0077] The following combination Figures 4-5 The training method for the resource recommendation model is explained.
[0078] Figure 4This is a schematic flowchart of a training method for a resource recommendation model according to an embodiment of the present disclosure.
[0079] like Figure 4 As shown, the training method 400 of the resource recommendation model may include operations S410 to S440.
[0080] Resource recommendation models can be LTR (Learning to Rank) models. A resource recommendation model can include a first sub-model and a second sub-model. The first sub-model can include convolutional neural networks, and the second sub-model can also include convolutional neural networks, etc.
[0081] Operate S410 to obtain training samples.
[0082] In operation S420, the object features, the first resource features, and the second resource features are processed using the first sub-model to obtain the first evaluation value.
[0083] In operation S430, the object characteristics and the first resource characteristics are input into the second sub-model to obtain the second evaluation value.
[0084] In operation S440, the object characteristics and the second resource characteristics are input into the second sub-model to obtain the third evaluation value.
[0085] In operation S450, the first sub-model and the second sub-model are trained based on the first difference between the first evaluation value and the first label, the second difference between the second evaluation value and the second label, and the third difference between the third evaluation value and the third label.
[0086] For example, methods for generating training samples can be used to generate the training samples needed to train a resource recommendation model.
[0087] For example, training samples may include object features of an object, first resource features of a first resource, second resource features of a second resource, a first label, a second label, and a third label. The first label represents the difference between the object's preference for the first resource and the object's preference for the second resource. The second label represents the object's preference for the first resource, and the third label represents the object's preference for the second resource.
[0088] For example, the degree of preference can include satisfaction and dissatisfaction; that is, the degree of preference reflects whether the object is satisfied with the resource. A high degree of preference indicates that the object is satisfied with the resource, while a low degree of preference indicates that the object is dissatisfied with the resource.
[0089] According to the embodiments provided in this disclosure, the first label reflects which of the first and second resources the object is more satisfied with; the second label reflects whether the object is satisfied with the first resource; and the third label reflects whether the object is satisfied with the second resource. It can be seen that the above three labels can accurately reflect the object's overall preference for the first and second resources. Furthermore, based on the first evaluation value, a resource recommendation model can be trained from a pair-wise perspective; based on the second and third evaluation values, a resource recommendation model can be trained from a point-wise perspective. Therefore, the trained resource recommendation model can accurately assess the object's overall preference for resources, thereby improving the recommendation effect.
[0090] According to another embodiment of this disclosure, the method for processing object features, first resource features, and second resource features using a first sub-model to obtain a first evaluation value may include the following operations: inputting the object features and the first resource features into the first sub-model to obtain a first sub-evaluation value; inputting the object features and the second resource features into the first sub-model to obtain a second sub-evaluation value; and then determining the first evaluation value based on the first and second sub-evaluation values.
[0091] For example, the first evaluation value can characterize the object's preference for the first resource as estimated by the first sub-model, and the second evaluation value can characterize the object's preference for the second resource as estimated by the first sub-model.
[0092] For example, the difference between the first sub-evaluation value and the second sub-evaluation value can be calculated as the first evaluation value. The first evaluation value can characterize the estimated difference between the object's preference for the first resource and the object's preference for the second resource, and this estimated difference is output by the first sub-model.
[0093] In this embodiment, a first sub-evaluation value and a second sub-evaluation value are determined separately, and then a first evaluation value is determined based on the two sub-evaluation values. The first evaluation value can reflect the estimated difference between the object's preference for the first resource and the second resource, thus improving the training effect of the resource recommendation model.
[0094] According to another embodiment of this disclosure, the method for training a first sub-model and a second sub-model based on a first difference between a first evaluation value and a first label, a second difference between a second evaluation value and a second label, and a third difference between a third evaluation value and a third label may include the following operations: determining a first loss based on the first difference between the first evaluation value and the first label; determining a second loss based on the second difference between the second evaluation value and the second label; determining a third loss based on the third difference between the third evaluation value and the third label; determining a total loss based on the first loss, the second loss, and the third loss; and adjusting the parameters of the first sub-model and the second sub-model based on the total loss.
[0095] For example, the losses mentioned above can be cross-entropy loss, mean squared error loss, etc. This embodiment does not limit the loss function.
[0096] For example, the total loss can be a weighted sum of the first, second, and third losses, where the weights of the first, second, and third losses can be equal. If the total loss is less than or equal to a loss threshold, the resource recommendation model has converged; otherwise, it has not converged and needs further training using training samples. For instance, the network gradient can be calculated based on the total loss, and the parameters of the resource recommendation model can be adjusted using gradient descent until the model converges.
[0097] In this embodiment, a first loss, a second loss, and a third loss are determined respectively, and then a total loss is determined based on these three losses to adjust the parameters of the resource recommendation model, thereby ensuring the training effect of the resource recommendation model.
[0098] Figure 5 This is a schematic diagram illustrating the training method of a resource recommendation model according to an embodiment of the present disclosure.
[0099] like Figure 5 As shown, the resource recommendation model 520 in this embodiment may include two first sub-models 521 and 522 and two second sub-models 523 and 524. The parameters of the two first sub-models 521 and 522 may be the same, and the parameters of the two second sub-models 523 and 524 may be the same. The training process of the resource recommendation model 520 is described below.
[0100] The first input information 511 (including object feature u and first resource feature i) can be input into the first sub-model 521, and the first sub-model 521 outputs a first sub-evaluation value 531. The second input information 512 (including object feature u and second resource feature j) can be input into the second sub-model 522, and the second sub-model 522 outputs a second sub-evaluation value 532. Based on the first sub-evaluation value 531 and the second sub-evaluation value 532, a first evaluation value 541 is determined. Based on the difference between the first evaluation value 541 and the first label, a first loss 551 is determined.
[0101] The first input information 511 can be input into the first second sub-model 523, and the second first sub-model 522 outputs a second evaluation value 542. The second loss 552 is determined based on the difference between the second evaluation value 542 and the second label.
[0102] The second input information 512 can be input into the second sub-model 524, and the second sub-model 524 outputs a third evaluation value 543. The third loss 553 is determined based on the difference between the third evaluation value 543 and the third label.
[0103] The total loss 560 is determined based on the first loss 551, the second loss 552, and the third loss 553. Then, the parameters of the two first sub-models 521 and 522 and the two second sub-models 523 and 524 are adjusted based on the total loss 560.
[0104] It should be noted that in the above embodiments, both the first sub-model and the second sub-model adopt a dual-tower structure, that is, there are two first sub-models and two sub-models with the same parameters. In other embodiments, the number of first sub-models can also be one, in which case the first input information 511 and the second input information 512 can be input into the first sub-model sequentially. Similarly, the number of second sub-models can also be one, in which case the first input information 511 and the second input information 512 can be input into the second sub-model sequentially.
[0105] In other embodiments, the resource recommendation model 520 described above may omit the second sub-model, and correspondingly, the labels of the training samples may omit the second and third labels.
[0106] The following combination Figure 6 The training method for the resource recommendation model is explained.
[0107] Figure 6 This is a schematic flowchart of a resource recommendation method according to an embodiment of the present disclosure.
[0108] like Figure 6 As shown, the resource recommendation method 600 may include operations S610 to S640.
[0109] In operation S610, the target object and multiple candidate resources to be recommended are identified.
[0110] For example, a pre-defined recall algorithm can be used to recall multiple candidate resources from the database. This embodiment does not limit the recall algorithm.
[0111] In operation S620, for each of the multiple candidate resources, the resource recommendation model is used to process the target object features of the target object and the candidate resource features of the candidate resources to obtain the recommendation evaluation value for the candidate resource.
[0112] For example, the resource recommendation model is trained using the above training method, and the resource recommendation model may include a first sub-model and a second sub-model.
[0113] For example, at least one of the first and second sub-models can be used to determine the recommendation rating, which represents the overall preference of the target object for the resource.
[0114] In operation S630, the target resource is determined from multiple candidate resources based on multiple recommended evaluation values of multiple candidate resources.
[0115] For example, multiple candidate resources are ranked according to their recommendation evaluation values, and then a predetermined number of candidate resources in the top order are determined as target resources.
[0116] When operating S640, target resources are recommended to the target object.
[0117] This embodiment of the disclosure utilizes the above-described resource recommendation model to process candidate resources, thereby accurately assessing the overall preference of the target object for candidate resources, ensuring that the target object has a high overall preference for the target resources, and improving the recommendation effect.
[0118] The following describes a method for determining the recommended evaluation value of candidate resources, with reference to specific embodiments.
[0119] In one example, the recommendation evaluation value can be determined using only the first sub-model. For instance, the target object features and candidate resource features can be input into the first sub-model of the resource recommendation model. The first sub-model outputs a first recommendation sub-evaluation value, which can then be used as the recommendation evaluation value. This embodiment determines the recommendation evaluation value based solely on the first sub-model, a simple and convenient method. Furthermore, the trained first sub-model exhibits high processing performance, thus ensuring that the recommendation evaluation value accurately reflects the user's preference for candidate resources.
[0120] In another example, the recommendation evaluation value can be determined using only the second sub-model. For instance, the target object features and candidate resource features can be input into the second sub-model of the resource recommendation model. The second sub-model outputs a second recommendation sub-evaluation value, which can be used as the recommendation evaluation value. This embodiment determines the recommendation evaluation value based solely on the second sub-model, a simple and convenient method. Furthermore, the trained second sub-model exhibits high processing performance, thus ensuring that the recommendation evaluation value accurately reflects the user's preference for candidate resources.
[0121] In another example, a first sub-model and a second sub-model can be used to determine the recommendation evaluation value. For example, the recommendation evaluation value can be determined based on the first recommendation sub-evaluation value and the second recommendation sub-evaluation value. Alternatively, the weighted sum of the first and second recommendation sub-evaluation values can be used as the recommendation evaluation value, with the weights of the first and second recommendation sub-evaluation values being equal. This embodiment determines the recommendation evaluation value based on the first and second sub-models, ensuring that the recommendation evaluation value accurately reflects the user's preference for candidate resources, thereby guaranteeing the recommendation effect.
[0122] Figure 7 This is a schematic structural block diagram of an apparatus for generating training samples according to an embodiment of the present disclosure.
[0123] like Figure 7 As shown, the apparatus 700 for generating training samples may include a partitioning module 710 and a generation module 720.
[0124] The partitioning module 710 is used to partition multiple resources into multiple resource sets based on the behavior generated by the object towards multiple resources; the multiple resources are the resources that have been presented to the object;
[0125] The generation module 720 is used to generate training samples based on at least one of a plurality of resource sets; wherein the training samples include object features of an object, first resource features of a first resource, second resource features of a second resource, a first label, a second label, and a third label, wherein the first label represents the difference between the object's preference for the first resource and the object's preference for the second resource, the second label represents the object's preference for the first resource, and the third label represents the object's preference for the second resource.
[0126] According to another embodiment of this disclosure, the partitioning module includes: a first adding submodule, a second adding submodule, a third adding submodule, and a fourth adding submodule. The first adding submodule is used to add the resource to a display set in response to the detection that an object has not clicked on the resource; the second adding submodule is used to add the resource to an interaction set in response to the detection that an object has clicked on the resource and the object has interacted with the resource; the third adding submodule is used to add the resource to a first browsing set in response to the detection that an object has clicked on the resource, the object has not interacted with the resource, and the object's browsing time for the resource is greater than or equal to a predetermined duration; the fourth adding submodule is used to add the resource to a second browsing set in response to the detection that an object has clicked on the resource, the object has not interacted with the resource, and the object's browsing time for the resource is less than a predetermined duration.
[0127] According to another embodiment of this disclosure, the generation module includes: a first determining submodule configured to, for each of the first resource and the second resource: determine that the object is satisfied with the resource in response to detecting that the object has generated an interactive behavior with the resource; determine that the object is satisfied with the resource in response to detecting that the object has not generated an interactive behavior with the resource and that the object's completion rate with respect to the resource is greater than or equal to a completion rate threshold; determine that the object is dissatisfied with the resource in response to detecting that the object has not generated an interactive behavior with the resource and that the object's completion rate with respect to the resource is less than a completion rate threshold; wherein, when the resource is a text-type resource, the completion rate is determined based on the browsing time and the amount of text in the resource; when the resource is a video-type resource, the completion rate is determined based on the browsing time and the video duration of the resource.
[0128] According to another embodiment of this disclosure, the generation module includes a second determining submodule and a third determining submodule. The second determining submodule is used to determine one resource from any two resource sets among a plurality of resource sets, as the first resource and the second resource; the third determining submodule is used to determine a first tag based on the resource set to which the first resource belongs and the resource set to which the second resource belongs.
[0129] According to another embodiment of this disclosure, the multiple resource sets include a browsing set, and the resources in the browsing set correspond to browsing durations; the generation module includes a fourth determining submodule and a fifth determining submodule. The fourth determining submodule is used to determine a first resource and a second resource from the browsing set; the fifth determining submodule is used to determine a first tag based on the browsing duration corresponding to the first resource and the browsing duration corresponding to the second resource.
[0130] Figure 8 This is a schematic structural block diagram of a training apparatus for a resource recommendation model according to an embodiment of the present disclosure.
[0131] like Figure 8 As shown, the resource recommendation model includes a first sub-model and a second sub-model. The training device 800 of the resource recommendation model may include a sample acquisition module 810, a first evaluation value determination module 820, a second evaluation value determination module 830, a third evaluation value determination module 840, and a training module 850.
[0132] The sample acquisition module 810 is used to acquire training samples; the training samples include object features of the object, first resource features of the first resource, second resource features of the second resource, first label, second label and third label. The first label represents the difference between the object's preference for the first resource and the object's preference for the second resource. The second label represents the object's preference for the first resource and the third label represents the object's preference for the second resource.
[0133] The first evaluation value determination module 820 is used to process object features, first resource features and second resource features using the first sub-model to obtain the first evaluation value.
[0134] The second evaluation value determination module 830 is used to input the object characteristics and the first resource characteristics into the second sub-model to obtain the second evaluation value.
[0135] The third evaluation value determination module 840 is used to input the object characteristics and the second resource characteristics into the second sub-model to obtain the third evaluation value.
[0136] Training module 850 is used to train a first sub-model and a second sub-model based on a first difference between a first evaluation value and a first label, a second difference between a second evaluation value and a second label, and a third difference between a third evaluation value and a third label.
[0137] According to another embodiment of this disclosure, the first evaluation value determination module includes: a first sub-evaluation value determination sub-module, a second sub-evaluation value determination sub-module, and a first evaluation value determination sub-module. The first sub-evaluation value determination sub-module is used to input object features and a first resource feature into a first sub-model to obtain a first sub-evaluation value; the second sub-evaluation value determination sub-module is used to input object features and a second resource feature into the first sub-model to obtain a second sub-evaluation value; and the first evaluation value determination sub-module is used to determine a first evaluation value based on the first sub-evaluation value and the second sub-evaluation value.
[0138] According to another embodiment of this disclosure, the training module includes: a first loss determination submodule, a second loss determination submodule, a third loss determination submodule, a total loss determination submodule, and a parameter adjustment submodule. The first loss determination submodule is used to determine a first loss based on a first difference between a first evaluation value and a first label; the second loss determination submodule is used to determine a second loss based on a second difference between a second evaluation value and a second label; the third loss determination submodule is used to determine a third loss based on a third difference between a third evaluation value and a third label; the total loss determination submodule is used to determine a total loss based on the first loss, the second loss, and the third loss; and the parameter adjustment submodule is used to adjust the parameters of the first sub-model and the parameters of the second sub-model based on the total loss.
[0139] Figure 9 This is a schematic structural block diagram of a resource recommendation device according to an embodiment of the present disclosure.
[0140] like Figure 9 As shown, the resource recommendation device 900 may include an information determination module 910, a recommendation evaluation value determination module 920, a target resource determination module 930, and a recommendation module 940.
[0141] The information determination module 910 is used to determine the target object and multiple candidate resources to be recommended.
[0142] The recommendation evaluation value determination module 920 is used to process the target object features of the target object and the candidate resource features of the candidate resources for each of the multiple candidate resources using the resource recommendation model, and obtain the recommendation evaluation value for the candidate resource.
[0143] The target resource determination module 930 is used to determine the target resource from multiple candidate resources based on multiple recommended evaluation values of multiple candidate resources.
[0144] The recommendation module 940 is used to recommend target resources to the target object; wherein, the resource recommendation model is trained by the training device of the above-mentioned resource recommendation model.
[0145] According to another embodiment of this disclosure, the recommendation evaluation value determination module includes: a first input submodule, a second input submodule, and a recommendation evaluation value determination submodule. The first input submodule is used to input the target object features and candidate resource features of candidate resources into a first submodel in the resource recommendation model to obtain a first recommendation sub-evaluation value; the second input submodule is used to input the target object features and candidate resource features of candidate resources into a second submodel in the resource recommendation model to obtain a second recommendation sub-evaluation value; the recommendation evaluation value determination submodule is used to determine a recommendation evaluation value based on the first and second recommendation sub-evaluation values.
[0146] According to another embodiment of this disclosure, the recommendation evaluation value determination module includes: a third input submodule, used to input the target object features and the candidate resource features of the candidate resources into the first sub-model in the resource recommendation model to obtain a first recommendation sub-evaluation value, and use the first recommendation sub-evaluation value as the recommendation evaluation value.
[0147] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0148] In the technical solution disclosed herein, the user's authorization or consent is obtained before acquiring or collecting the user's personal information.
[0149] According to embodiments of this disclosure, this disclosure also provides an electronic device, including at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform at least one of the above-described methods for generating training samples, training a resource recommendation model, and recommending resources.
[0150] According to embodiments of this disclosure, this disclosure also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute at least one of the above-described methods for generating training samples, training a resource recommendation model, and recommending resources.
[0151] According to embodiments of this disclosure, this disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements at least one of the above-described methods for generating training samples, training a resource recommendation model, and recommending resources.
[0152] Figure 10 This is a structural block diagram of an electronic device used to implement the methods for generating training samples, training resource recommendation models, and resource recommendation methods according to embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0153] like Figure 10 As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1002 or a computer program loaded from storage unit 1008 into random access memory (RAM) 1003. The RAM 1003 may also store various programs and data required for the operation of device 1000. The computing unit 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. Input / output (I / O) interface 1005 is also connected to bus 1004.
[0154] Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, etc.; output unit 1007, such as various types of monitors, speakers, etc.; storage unit 1008, such as disk, optical disk, etc.; and communication unit 1009, such as network card, modem, wireless transceiver, etc. Communication unit 1009 allows device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0155] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as at least one of the methods for generating training samples, training resource recommendation models, and resource recommendation methods described above. For example, in some embodiments, at least one of the methods for generating training samples, training resource recommendation models, and resource recommendation methods described above can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by computing unit 1001, it can perform one or more steps of at least one of the methods described above for generating training samples, training resource recommendation models, and resource recommendation methods. Alternatively, in other embodiments, computing unit 1001 can be configured by any other suitable means (e.g., by means of firmware) to execute at least one of the methods described above for generating training samples, training resource recommendation models, and resource recommendation methods.
[0156] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0157] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0158] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0159] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0160] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0161] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.
[0162] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0163] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for training a resource recommendation model, the resource recommendation model comprising a first sub-model and a second sub-model, the method comprising: Obtain training samples; The training samples include object features of an object, first resource features of a first resource, second resource features of a second resource, a first label, a second label, and a third label. The first label represents the difference between the object's preference for the first resource and the object's preference for the second resource. The second label represents the object's preference for the first resource. The third label represents the object's preference for the second resource. The first resource and the second resource come from different resource sets in a plurality of resource sets. The plurality of resource sets are obtained by dividing the plurality of resources according to the behavior generated by the object for the plurality of resources. Each of the plurality of resource sets corresponds to a degree of preference. The first tag is determined according to the resource set to which the first resource belongs and the resource set to which the second resource belongs. The resource includes at least one of video, text, image and music. The object features, the first resource features, and the second resource features are processed using the first sub-model to obtain a first evaluation value; The object features and the first resource features are input into the second sub-model to obtain the second evaluation value; The object features and the second resource features are input into the second sub-model to obtain the third evaluation value; as well as The first sub-model and the second sub-model are trained based on the first difference between the first evaluation value and the first label, the second difference between the second evaluation value and the second label, and the third difference between the third evaluation value and the third label.
2. The method according to claim 1, wherein, The process of using the first sub-model to process the object features, the first resource features, and the second resource features to obtain the first evaluation value includes: The object features and the first resource features are input into the first sub-model to obtain the first sub-evaluation value; The object features and the second resource features are input into the first sub-model to obtain the second sub-evaluation value; and The first evaluation value is determined based on the first sub-evaluation value and the second sub-evaluation value.
3. The method according to claim 1, wherein, The step of training the first sub-model and the second sub-model based on the first difference between the first evaluation value and the first label, the second difference between the second evaluation value and the second label, and the third difference between the third evaluation value and the third label includes: A first loss is determined based on a first difference between the first evaluation value and the first label; The second loss is determined based on the second difference between the second evaluation value and the second label; The third loss is determined based on the third difference between the third evaluation value and the third label; Based on the first loss, the second loss, and the third loss, determine the total loss; and Based on the total loss, adjust the parameters of the first sub-model and the second sub-model.
4. A resource recommendation method, comprising: Identify the target object and multiple candidate resources to be recommended; For each of the multiple candidate resources, the target object features of the target object and the candidate resource features of the candidate resource are processed using a resource recommendation model to obtain a recommendation evaluation value for the candidate resource. Based on the multiple recommended evaluation values of the multiple candidate resources, the target resource is determined from the multiple candidate resources; as well as Recommend the target resource to the target object; The resource recommendation model is trained using the method described in any one of claims 1 to 3.
5. The method according to claim 4, wherein, The resource recommendation model is used to process the target object features of the target object and the candidate resource features of the candidate resources to obtain recommendation evaluation values for the candidate resources, including: The target object features and the candidate resource features are input into the first sub-model of the resource recommendation model to obtain the first recommendation sub-evaluation value; The target object features and the candidate resource features are input into the second sub-model of the resource recommendation model to obtain the second recommendation sub-evaluation value; and The recommendation evaluation value is determined based on the first recommendation sub-evaluation value and the second recommendation sub-evaluation value.
6. The method according to claim 4, wherein, The resource recommendation model is used to process the target object features of the target object and the candidate resource features of the candidate resources to obtain recommendation evaluation values for the candidate resources, including: The target object features and the candidate resource features are input into the first sub-model of the resource recommendation model to obtain the first recommendation sub-evaluation value, and the first recommendation sub-evaluation value is used as the recommendation evaluation value.
7. A method for generating training samples, comprising: Based on the behavior of an object towards multiple resources, the multiple resources are divided into multiple resource sets; The aforementioned resources are those that have already been presented to the object; as well as A training sample is generated based on at least one of the plurality of resource sets, the training sample being used in the method of any one of claims 1 to 3; The training samples include object features of an object, first resource features of a first resource, second resource features of a second resource, a first label, a second label, and a third label. The first label represents the difference between the object's preference for the first resource and the object's preference for the second resource. The second label represents the object's preference for the first resource. The third label represents the object's preference for the second resource. The first resource and the second resource come from different resource sets among multiple resource sets, each of which has a corresponding degree of preference. The first tag is determined based on the resource set to which the first resource belongs and the resource set to which the second resource belongs.
8. The method according to claim 7, wherein, The step of dividing the multiple resources into multiple resource sets based on the behavior generated by the object in relation to multiple resources includes: In response to the detection that the object has not clicked the resource, the resource is added to the display collection; In response to detecting that the object clicks on the resource and the object performs an interactive behavior on the resource, the resource is added to the interaction set; In response to detecting that the object clicked on the resource, and the object did not perform any interactive behavior on the resource, and the object's browsing time for the resource was greater than or equal to a predetermined duration, the resource was added to a first browsing set; and In response to detecting that the object clicked on the resource, and the object did not perform any interactive behavior on the resource, and the browsing time of the object on the resource was less than the predetermined time, the resource was added to the second browsing set.
9. The method according to claim 7, wherein, The step of generating training samples based on at least one of the plurality of resource sets includes: For each of the first resource and the second resource: In response to detecting that the object has generated an interactive behavior toward the resource, it is determined that the object is satisfied with the resource; In response to detecting that the object has not generated any interactive behavior towards the resource, and that the object's completion rate for the resource is greater than or equal to a completion rate threshold, it is determined that the object is satisfied with the resource; and In response to detecting that the object has not generated any interactive behavior for the resource, and the object's completion rate for the resource is less than a completion rate threshold, it is determined that the object is not satisfied with the resource; Wherein, if the resource is a text-based resource, the completion rate is determined based on the browsing time and the amount of text in the resource; if the resource is a video-based resource, the completion rate is determined based on the browsing time and the video duration of the resource.
10. The method according to claim 7, wherein, The step of generating training samples based on at least one of the plurality of resource sets includes: From any two resource sets among the plurality of resource sets, one resource is determined as the first resource and the second resource, respectively; and The first tag is determined based on the resource set to which the first resource belongs and the resource set to which the second resource belongs.
11. The method according to claim 7, wherein, The plurality of resource sets include a browsing set, and the resources in the browsing set correspond to browsing durations; The step of generating training samples based on at least one of the plurality of resource sets includes: The first resource and the second resource are determined from the browsing set; as well as The first tag is determined based on the browsing duration corresponding to the first resource and the browsing duration corresponding to the second resource.
12. A training apparatus for a resource recommendation model, the resource recommendation model comprising a first sub-model and a second sub-model, the apparatus comprising: The sample acquisition module is used to acquire training samples; The training samples include object features of an object, first resource features of a first resource, second resource features of a second resource, a first label, a second label, and a third label. The first label represents the difference between the object's preference for the first resource and the object's preference for the second resource. The second label represents the object's preference for the first resource. The third label represents the object's preference for the second resource. The first resource and the second resource come from different resource sets in a plurality of resource sets. The plurality of resource sets are obtained by dividing the plurality of resources according to the behavior generated by the object for the plurality of resources. Each of the plurality of resource sets corresponds to a degree of preference. The first tag is determined according to the resource set to which the first resource belongs and the resource set to which the second resource belongs. The resource includes at least one of video, text, image and music. The first evaluation value determination module is used to process the object features, the first resource features and the second resource features using the first sub-model to obtain a first evaluation value. The second evaluation value determination module is used to input the object features and the first resource features into the second sub-model to obtain the second evaluation value; The third evaluation value determination module is used to input the object characteristics and the second resource characteristics into the second sub-model to obtain the third evaluation value; and The training module is used to train the first sub-model and the second sub-model based on the first difference between the first evaluation value and the first label, the second difference between the second evaluation value and the second label, and the third difference between the third evaluation value and the third label.
13. The apparatus according to claim 12, wherein, The first evaluation value determination module includes: The first sub-evaluation value determination submodule is used to input the object features and the first resource features into the first sub-model to obtain the first sub-evaluation value; The second sub-evaluation value determination submodule is used to input the object features and the second resource features into the first sub-model to obtain the second sub-evaluation value; and The first evaluation value determination submodule is used to determine the first evaluation value based on the first sub-evaluation value and the second sub-evaluation value.
14. The apparatus according to claim 12, wherein, The training module includes: The first loss determination submodule is used to determine a first loss based on a first difference between the first evaluation value and the first label; The second loss determination submodule is used to determine the second loss based on the second difference between the second evaluation value and the second label; The third loss determination submodule is used to determine the third loss based on the third difference between the third evaluation value and the third label; The total loss determination submodule is used to determine the total loss based on the first loss, the second loss, and the third loss; and The parameter adjustment submodule is used to adjust the parameters of the first sub-model and the parameters of the second sub-model according to the total loss.
15. A resource recommendation device, comprising: The information determination module is used to determine the target object and multiple candidate resources to be recommended; The recommendation evaluation value determination module is used to process the target object features of the target object and the candidate resource features of the candidate resource using a resource recommendation model for each candidate resource among multiple candidate resources, and obtain a recommendation evaluation value for the candidate resource. The target resource determination module is used to determine the target resource from the plurality of candidate resources based on the plurality of recommended evaluation values of the plurality of candidate resources; as well as The recommendation module is used to recommend the target resource to the target object; The resource recommendation model is trained using the apparatus described in any one of claims 12 to 14.
16. The apparatus according to claim 15, wherein, The module for determining the recommended evaluation value includes: The first input submodule is used to input the target object features and the candidate resource features of the candidate resources into the first sub-model of the resource recommendation model to obtain the first recommendation sub-evaluation value; The second input submodule is used to input the target object features and the candidate resource features into the second sub-model of the resource recommendation model to obtain a second recommendation sub-evaluation value; and The recommendation evaluation value determination submodule is used to determine the recommendation evaluation value based on the first recommendation sub-evaluation value and the second recommendation sub-evaluation value.
17. The apparatus according to claim 15, wherein, The module for determining the recommended evaluation value includes: The third input submodule is used to input the target object features and the candidate resource features of the candidate resources into the first sub-model of the resource recommendation model to obtain the first recommendation sub-evaluation value, and use the first recommendation sub-evaluation value as the recommendation evaluation value.
18. An apparatus for generating training samples, comprising: The partitioning module is used to divide the multiple resources into multiple resource sets based on the behavior generated by the object in response to multiple resources; The aforementioned resources are those that have already been presented to the object; as well as A generation module is configured to generate training samples based on at least one of the plurality of resource sets, the training samples being used in the apparatus of any one of claims 12 to 14; The training samples include object features of an object, first resource features of a first resource, second resource features of a second resource, a first label, a second label, and a third label. The first label represents the difference between the object's preference for the first resource and the object's preference for the second resource. The second label represents the object's preference for the first resource. The third label represents the object's preference for the second resource. The first resource and the second resource come from different resource sets among multiple resource sets, each of which has a corresponding degree of preference. The first tag is determined based on the resource set to which the first resource belongs and the resource set to which the second resource belongs.
19. The apparatus according to claim 18, wherein, The partitioning module includes: The first addition submodule is used to add the resource to the display collection in response to the detection that the object has not clicked the resource; The second adding submodule is used to add the resource to the interaction set in response to detecting that the object clicks on the resource and the object generates an interactive behavior for the resource; The third adding submodule is used to add the resource to the first browsing set in response to detecting that the object clicked on the resource, and the object did not generate any interactive behavior towards the resource, and the object's browsing time for the resource is greater than or equal to a predetermined time; and The fourth addition submodule is used to add the resource to the second browsing set in response to the detection that the object clicks on the resource, the object does not generate any interactive behavior for the resource, and the browsing time of the object for the resource is less than the predetermined time.
20. The apparatus according to claim 18, wherein, The generation module includes: A first determining submodule is configured to determine each of the first resource and the second resource: In response to detecting that the object has generated an interactive behavior toward the resource, it is determined that the object is satisfied with the resource; In response to detecting that the object has not generated any interactive behavior towards the resource, and that the object's completion rate for the resource is greater than or equal to a completion rate threshold, it is determined that the object is satisfied with the resource; and In response to detecting that the object has not generated any interactive behavior for the resource, and the object's completion rate for the resource is less than a completion rate threshold, it is determined that the object is not satisfied with the resource; Wherein, if the resource is a text-based resource, the completion rate is determined based on the browsing time and the amount of text in the resource; if the resource is a video-based resource, the completion rate is determined based on the browsing time and the video duration of the resource.
21. The apparatus according to claim 18, wherein, The generation module includes: The second determining submodule is configured to determine one resource from any two resource sets among the plurality of resource sets, as the first resource and the second resource, respectively; and The third determining submodule is used to determine the first tag based on the resource set to which the first resource belongs and the resource set to which the second resource belongs.
22. The apparatus according to claim 18, wherein, The plurality of resource sets include a browsing set, and the resources in the browsing set correspond to browsing durations; The generation module includes: The fourth determining submodule is used to determine the first resource and the second resource from the browsing set; as well as The fifth determining submodule is used to determine the first tag based on the browsing duration corresponding to the first resource and the browsing duration corresponding to the second resource.
23. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 11.
24. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 11.
25. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Model training method and device, intention recognition method and device and electronic equipment
CN114330364A