Method, device and equipment for training multimedia resource recommendation model and storage medium
By integrating prediction models and penalty adjustment methods, a multimedia resource recommendation model is trained, which solves the uncertainty problem of recommendation models with limited resources and improves the diversity and accuracy of recommendations.
Patent Information
- Application Number
- CN202310788076.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-29
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2043-06-29
AI Technical Summary
Existing multimedia resource recommendation models suffer from uncertainty when recommending a small number of multimedia resources, leading to the Matthew effect and affecting the diversity and novelty of recommendations.
By integrating prediction models to predict the interaction behavior between target objects and multimedia resources, and adjusting the feedback values by combining the first and second penalty degrees, a multimedia resource recommendation model is trained to avoid overly conservative recommendation strategies.
It mitigates the Matthew effect, improves the diversity and novelty of multimedia resource recommendations, and enhances the accuracy of recommendations.
Smart Images

Figure CN116595264B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a training method, apparatus, device, and storage medium for a multimedia resource recommendation model. Background Technology
[0002] With the development of internet technology, most applications can recommend various types of multimedia resources to users through multimedia resource recommendation systems, such as videos, audio, news, and items, to meet users' different interests. To improve long-term user satisfaction with multimedia resource recommendation systems, offline reinforcement learning can be introduced. The core idea of offline reinforcement learning is to train the multimedia resource recommendation model using historical data (historical interaction records between users and multimedia resources). However, due to the diverse types of multimedia resources in historical data and the uneven distribution of different types, the prediction results of the multimedia resource recommendation model for a relatively small number of certain types of multimedia resources are uncertain. This uncertainty refers to the uncertainty of whether the user is interested in the multimedia resources recommended by the model. Therefore, how to solve the problem of uncertainty in the prediction results of multimedia resource recommendation models for a relatively small number of multimedia resources in historical data is a key research focus in the recommendation field.
[0003] In related technologies, conservative recommendation strategies are typically introduced in offline reinforcement learning to reduce the reliance of multimedia resource recommendation models on a limited number of multimedia resources in historical data. Specifically, when recommending multimedia resources to users, the model tries to avoid recommending multimedia resources of the same type as those with high uncertainty in historical data. Instead, it only recommends multimedia resources of the same type as those with high certainty in historical data. This avoids recommending multimedia resources that the user may not be interested in, thus improving the accuracy of the multimedia resource recommendation model.
[0004] However, the conservative approach of the above scheme will create a severe Matthew effect in the recommendation field. This will lead to a higher and higher recommendation frequency for a small number of popular or mainstream multimedia resources, while the recommendation frequency for most less popular multimedia resources will decrease, thus sacrificing the diversity and novelty of recommended multimedia resources and reducing user satisfaction with the recommended multimedia resources. Summary of the Invention
[0005] This disclosure provides a training method, apparatus, device, and storage medium for a multimedia resource recommendation model, which can avoid the multimedia resource recommendation model learning an overly conservative recommendation strategy and alleviate the Matthew effect when recommending multimedia resources. The technical solution of this disclosure is as follows:
[0006] According to one aspect of the embodiments of this disclosure, a method for training a multimedia resource recommendation model is provided, comprising:
[0007] An integrated prediction model is obtained, which is trained based on historical data. The historical data includes object features of multiple sample objects, multimedia resource features of multiple sample multimedia resources, and historical interaction behavior between the multiple sample objects and the multiple sample multimedia resources. The integrated prediction model is used to predict the interaction behavior between objects and multimedia resources.
[0008] Based on the integrated prediction model, the object features of the target object and the multimedia resource features of the target multimedia resource are predicted to obtain a first feedback value of the target object to the target multimedia resource and a first penalty degree of the first feedback value. The target object is any object among the plurality of sample objects, and the target multimedia resource is the sample multimedia resource recommended to the target object by the multimedia resource recommendation model among the plurality of sample multimedia resources. The first feedback value is used to indicate the predicted interaction behavior between the target object and the target multimedia resource, and the first penalty degree is used to indicate the dispersion of the first feedback value. The multimedia resource recommendation model is used to predict the multimedia resources recommended to the object.
[0009] The first feedback value is adjusted based on the first penalty degree and the second penalty degree of the first feedback value, wherein the second penalty degree is used to indicate the degree of randomness of the occurrence of the target multimedia resource among the plurality of sample multimedia resources;
[0010] The multimedia resource recommendation model is trained based on the adjusted first feedback value, the object features of the target object, and the multimedia resource features of the target multimedia resource.
[0011] According to another aspect of the embodiments of this disclosure, a training apparatus for a multimedia resource recommendation model is provided, comprising:
[0012] The acquisition unit is configured to acquire an integrated prediction model, which is trained based on historical data. The historical data includes object features of multiple sample objects, multimedia resource features of multiple sample multimedia resources, and historical interaction behavior between the multiple sample objects and the multiple sample multimedia resources. The integrated prediction model is used to predict the interaction behavior between the objects and the multimedia resources.
[0013] The first prediction unit is configured to predict the object features of the target object and the multimedia resource features of the target multimedia resource based on the integrated prediction model, and obtain a first feedback value of the target object to the target multimedia resource and a first penalty degree of the first feedback value. The target object is any object among the plurality of sample objects, and the target multimedia resource is a sample multimedia resource recommended to the target object by the multimedia resource recommendation model among the plurality of sample multimedia resources. The first feedback value is used to indicate the predicted interaction behavior between the target object and the target multimedia resource, and the first penalty degree is used to indicate the dispersion of the first feedback value. The multimedia resource recommendation model is used to predict the multimedia resources recommended to the object.
[0014] The adjustment unit is configured to adjust the first feedback value based on the first penalty degree and a second penalty degree of the first feedback value, wherein the second penalty degree is used to indicate the degree of randomness of the occurrence of the target multimedia resource among the plurality of sample multimedia resources;
[0015] The training unit is configured to train the multimedia resource recommendation model based on the adjusted first feedback value, the object features of the target object, and the multimedia resource features of the target multimedia resource.
[0016] In some embodiments, the integrated prediction model includes multiple Gaussian probability models, which are used to predict the distribution of feedback values of objects to multimedia resources.
[0017] The prediction unit includes:
[0018] The prediction subunit is configured to predict, based on any Gaussian probability model, the object features of the target object and the multimedia resource features of the target multimedia resource to obtain the mean and variance of the second feedback value of the target object to the target multimedia resource, wherein the second feedback value is used to indicate the initial interaction behavior of the target object to the target multimedia resource.
[0019] The first determining subunit is configured to determine the first feedback value as the average of multiple averages of the second feedback value;
[0020] The second determining subunit is configured to determine the maximum value of a plurality of variances of the second feedback value as the first penalty degree.
[0021] In some embodiments, the apparatus further includes:
[0022] The second prediction unit is configured to predict the object features of the plurality of sample objects and the multimedia resource features of the plurality of sample multimedia resources based on the Gaussian probability model for any Gaussian probability model, and obtain the mean and variance of the third feedback values of the plurality of sample objects to the plurality of sample multimedia resources. The third feedback values are used to indicate the predicted interaction behavior between the sample objects and the sample multimedia resources.
[0023] The first determining unit is configured to determine the training loss of the Gaussian probability model based on the historical interaction behavior of the plurality of sample objects and the plurality of sample multimedia resources, the mean of the third feedback value and the variance of the third feedback value, wherein the training loss is used to indicate the difference between the historical interaction behavior and the predicted interaction behavior.
[0024] The update unit is configured to update the model parameters of the Gaussian probability model based on the training loss.
[0025] In some embodiments, the apparatus further includes:
[0026] The second determining unit is configured to determine sample multimedia resources associated with the object identifier from the historical data based on the object identifier of the target object;
[0027] The third determining unit is configured to determine the relative entropy of the multimedia resource recommendation model based on the sample multimedia resources and the target multimedia resources, wherein the relative entropy is used to indicate the difference between the target multimedia resources and the sample multimedia resources;
[0028] The fourth determining unit is configured to determine the second penalty degree based on the relative entropy, wherein the relative entropy is positively correlated with the second penalty degree.
[0029] In some embodiments, the adjustment unit is configured to determine the weight of the second penalty degree based on the relative entropy, the weight being negatively correlated with the relative entropy; and to perform a weighted summation of the first feedback value, the first penalty degree, and the second penalty degree based on the weight of the second penalty degree to obtain the adjusted first feedback value.
[0030] In some embodiments, the training unit includes:
[0031] The third determining subunit is configured to determine the preference information of the target object based on the adjusted first feedback value, the object characteristics of the target object, and the multimedia resource characteristics of the target multimedia resource. The preference information is used to indicate whether the target object is interested in the target multimedia resource.
[0032] The adjustment subunit is configured to adjust the model parameters of the multimedia resource recommendation model based on the preference information and the adjusted first feedback value.
[0033] In some embodiments, the apparatus further includes:
[0034] The fifth determining unit is configured to determine, based on the object identifier of the target object, sample multimedia resources associated with the object identifier and the historical interaction behavior between the target object and the sample multimedia resources from the historical data;
[0035] The sixth determining unit is configured to determine the historical preference information of the target object based on the object characteristics of the target object, the multimedia resource characteristics of the sample multimedia resource, and the historical interaction behavior. The historical preference information is used to indicate whether the target object is interested in the sample multimedia resource.
[0036] The processing unit is configured to process the historical preference information based on the multimedia resource recommendation model to obtain sample multimedia resources recommended by the multimedia resource recommendation model to the target object.
[0037] According to another aspect of the embodiments of this disclosure, an electronic device is provided, the electronic device comprising:
[0038] One or more processors;
[0039] Memory used to store the executable program code of the processor;
[0040] The processor is configured to execute the program code to implement the training method of the multimedia resource recommendation model described above.
[0041] According to another aspect of the present disclosure, a computer-readable storage medium is provided, which, when the program code in the computer-readable storage medium is executed by the processor of an electronic device, enables the electronic device to perform the training method of the multimedia resource recommendation model described above.
[0042] According to another aspect of the present disclosure, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the training method of the multimedia resource recommendation model described above.
[0043] This disclosure provides a training method for a multimedia resource recommendation model. Based on an ensemble prediction model, it predicts the object features of a target object and the multimedia resource features of a target multimedia resource. This yields a first feedback value indicating the interaction behavior between the target object and the target multimedia resource, and a first penalty value indicating the degree of dispersion of the first feedback value. Since the target multimedia resource is a sample multimedia resource recommended by the multimedia resource recommendation model for the target object, adjusting the first feedback value based on the first penalty value and a second penalty value indicating the randomness of the target multimedia resource's appearance among multiple sample multimedia resources can penalize the first feedback value if the recommended target multimedia resource is too conservative. This avoids the trained multimedia resource recommendation model learning an overly conservative recommendation strategy, alleviates the Matthew effect in multimedia resource recommendation models, and improves the diversity of multimedia resources recommended by the multimedia resource recommendation model.
[0044] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0045] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0046] Figure 1 This is a schematic diagram illustrating the implementation environment of a training method for a multimedia resource recommendation model according to an exemplary embodiment.
[0047] Figure 2 This is a flowchart illustrating a training method for a multimedia resource recommendation model according to an exemplary embodiment.
[0048] Figure 3 This is a flowchart illustrating another method for training a multimedia resource recommendation model according to an exemplary embodiment.
[0049] Figure 4 This is a schematic diagram illustrating the workflow of an entropy penalizer according to an exemplary embodiment.
[0050] Figure 5 This is a framework diagram illustrating a training method for a multimedia resource recommendation model according to an exemplary embodiment.
[0051] Figure 6 This is a schematic diagram illustrating the test results of a multimedia resource recommendation model according to an exemplary embodiment.
[0052] Figure 7This is a schematic diagram illustrating experimental results of a training method for a multimedia resource recommendation model based on an exemplary embodiment, targeting the Matthew effect.
[0053] Figure 8 This is a schematic diagram illustrating the experimental results of a training method for a multimedia resource recommendation model under different exit conditions, according to an exemplary embodiment.
[0054] Figure 9 This is a block diagram illustrating a training apparatus for a multimedia resource recommendation model according to an exemplary embodiment.
[0055] Figure 10 This is a block diagram illustrating a training apparatus for another multimedia resource recommendation model according to an exemplary embodiment.
[0056] Figure 11 This is a block diagram illustrating a terminal according to an exemplary embodiment.
[0057] Figure 12 This is a block diagram illustrating a server according to an exemplary embodiment. Detailed Implementation
[0058] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0059] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0060] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this disclosure are authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the historical data involved in this disclosure was obtained with full authorization.
[0061] Figure 1 This is a schematic diagram illustrating the implementation environment of a training method for a multimedia resource recommendation model according to an exemplary embodiment. See also... Figure 1 The implementation environment specifically includes: terminal 101 and server 102.
[0062] Terminal 101 can be at least one of the following devices: smartphone, smartwatch, desktop computer, laptop, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), and laptop computer. An application for displaying multimedia resources can be installed and run on terminal 101. Users can log in to this application through terminal 101 to access the services provided by the application. This application is associated with server 102, which provides background services. Terminal 101 can connect to server 102 via a wireless network or a wired network.
[0063] Terminal 101 can refer to one of a plurality of terminals; this embodiment uses terminal 101 as an example only. Those skilled in the art will understand that the number of terminals can be more or less. For example, there may be only a few terminals, or dozens or hundreds, or even more. This disclosure does not limit the number of terminals or the type of device.
[0064] Server 102 can be at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. Server 102 can be connected to terminal 101 and other terminals via a wireless network or a wired network. Optionally, the number of servers can be more or less, and this disclosure does not limit this. Of course, server 102 may also include other functional servers to provide more comprehensive and diversified services.
[0065] Figure 2 This is a flowchart illustrating a training method for a multimedia resource recommendation model according to an exemplary embodiment, such as... Figure 2 As shown, the method is performed by an electronic device and includes the following steps:
[0066] In step S201, an integrated prediction model is obtained. The integrated prediction model is trained based on historical data, which includes object features of multiple sample objects, multimedia resource features of multiple sample multimedia resources, and historical interaction behavior between multiple sample objects and multiple sample multimedia resources. The integrated prediction model is used to predict the interaction behavior between objects and multimedia resources.
[0067] In this embodiment, the ensemble prediction model is trained based on historical data. The historical data is labeled training data, including object features of multiple sample objects, multimedia resource features of multiple sample multimedia resources, and historical interaction behaviors between multiple sample objects and multiple sample multimedia resources. The electronic device can use the historical interaction behaviors between multiple sample objects and multiple sample multimedia resources as label information, and perform supervised training on the ensemble prediction model based on the historical data. The sample objects are accounts logged into on multimedia resource clients installed on the terminal. Through the sample objects, various multimedia resources, such as videos, audio, and items, can be published or viewed. The sample multimedia resources can be multimedia resources viewed by the sample objects or multimedia resources not viewed by the sample objects. The historical interaction behaviors between the sample objects and sample multimedia resources include clicks, viewing duration, and likes by the sample objects. The object features of the sample objects include object identifiers and object activity levels; the multimedia resource features of the sample multimedia resources include multimedia resource identifiers, multimedia resource types, multimedia resource durations, and multimedia resource popularity. Electronic devices can obtain historical data from a local database, from other electronic devices, or from multimedia resources published by the terminal. This disclosure does not limit the scope of such historical data.
[0068] In step S202, based on the ensemble prediction model, the object features of the target object and the multimedia resource features of the target multimedia resource are predicted to obtain the first feedback value of the target object to the target multimedia resource and the first penalty degree of the first feedback value. The target object is any object among multiple sample objects, and the target multimedia resource is the sample multimedia resource recommended to the target object by the multimedia resource recommendation model among multiple sample multimedia resources. The first feedback value is used to indicate the predicted interaction behavior between the target object and the target multimedia resource, and the first penalty degree is used to indicate the degree of dispersion of the first feedback value. The multimedia resource recommendation model is used to predict the multimedia resources recommended to the object.
[0069] In this embodiment of the disclosure, the multimedia resource recommendation model to be trained is used to predict multimedia resources recommended to an object, and the electronic device is able to obtain the target multimedia resources recommended by the multimedia resource recommendation model to the target object. The target object is any object among multiple sample objects in historical data, and the target multimedia resource is a sample multimedia resource among multiple sample multimedia resources in historical data. The ensemble prediction model is used to predict the interaction behavior between the object and the multimedia resource. Based on the ensemble prediction model, the electronic device can predict the object features of the target object and the multimedia resource features of the target multimedia resource, obtaining a first feedback value of the target object to the target multimedia resource and a first penalty degree for the first feedback value. The first feedback value indicates the predicted interaction behavior between the target object and the target multimedia resource, such as the target object's clicks, browsing duration, and likes. The first penalty degree indicates the dispersion of the first feedback value, that is, the degree of uncertainty regarding the first feedback value predicted by the ensemble prediction model. Therefore, the first penalty degree can indirectly reflect the credibility of the predicted interaction behavior between the target object and the target multimedia resource predicted by the integrated prediction model. The first penalty degree is positively correlated with the degree of dispersion. The larger the first penalty degree, the greater the dispersion of the first feedback value, and the lower the credibility of the predicted interaction behavior between the target object and the target multimedia resource predicted by the integrated prediction model.
[0070] In step S203, the first feedback value is adjusted based on the first penalty degree and the second penalty degree of the first feedback value. The second penalty degree is used to indicate the degree of randomness in the appearance of the target multimedia resource among multiple sample multimedia resources.
[0071] In this embodiment, the electronic device determines a second penalty degree for the first feedback value based on the target multimedia resources recommended by the multimedia resource recommendation model and multiple sample multimedia resources in historical data. The second penalty degree indicates the degree of randomness in which the target multimedia resources recommended by the multimedia resource recommendation model appear among the multiple sample multimedia resources. Specifically, the second penalty degree is positively correlated with the degree of randomness; the larger the second penalty degree, the greater the degree of randomness in the appearance of the target multimedia resources among the multiple sample multimedia resources. This indicates that the multimedia resource recommendation model recommends a wider variety of target multimedia resources for the target object, not just limited to recommending sample multimedia resources from historical data that have a high frequency of interaction with the target object. By adjusting the first feedback value based on the first and second penalty degrees, the electronic device can avoid learning an overly conservative recommendation strategy during the subsequent training of the multimedia resource recommendation model.
[0072] In step S204, a multimedia resource recommendation model is trained based on the adjusted first feedback value, the object features of the target object, and the multimedia resource features of the target multimedia resource.
[0073] In this embodiment, the electronic device trains a multimedia resource recommendation model based on an adjusted first feedback value, object features of the target object, and multimedia resource features of the target multimedia resource. During training, the electronic device can adjust the recommendation strategy of the multimedia resource recommendation model based on the predicted interaction behavior between the target object and the target multimedia resource indicated by the adjusted first feedback value, the object features of the target object, and the multimedia resource features of the target multimedia resource. This avoids the trained multimedia resource recommendation model learning an overly conservative recommendation strategy and increases the diversity of multimedia resources recommended by the model.
[0074] This disclosure provides a training method for a multimedia resource recommendation model. Based on an ensemble prediction model, it predicts the object features of a target object and the multimedia resource features of a target multimedia resource. This yields a first feedback value indicating the interaction behavior between the target object and the target multimedia resource, and a first penalty value indicating the degree of dispersion of the first feedback value. Since the target multimedia resource is a sample multimedia resource recommended by the multimedia resource recommendation model for the target object, adjusting the first feedback value based on the first penalty value and a second penalty value indicating the randomness of the target multimedia resource's appearance among multiple sample multimedia resources can penalize the first feedback value if the recommended target multimedia resource is too conservative. This avoids the trained multimedia resource recommendation model learning an overly conservative recommendation strategy, alleviates the Matthew effect in multimedia resource recommendation models, and improves the diversity of multimedia resources recommended by the multimedia resource recommendation model.
[0075] In some embodiments, the integrated prediction model includes multiple Gaussian probability models, which are used to predict the distribution of the object's feedback values to multimedia resources.
[0076] Based on an ensemble prediction model, the object features of the target object and the multimedia resource features of the target multimedia resource are predicted to obtain the first feedback value of the target object to the target multimedia resource and the first penalty degree of the first feedback value, including:
[0077] For any Gaussian probability model, based on the Gaussian probability model, the object characteristics of the target object and the multimedia resource characteristics of the target multimedia resource are predicted to obtain the mean and variance of the second feedback value of the target object to the target multimedia resource. The second feedback value is used to indicate the initial interaction behavior of the target object to the target multimedia resource.
[0078] The average of multiple means of the second feedback value is determined as the first feedback value;
[0079] The maximum value of the multiple variances of the second feedback value is determined as the first penalty degree.
[0080] In this embodiment of the disclosure, the electronic device can obtain a first feedback value, which indicates the predicted interaction behavior between the target object and the target multimedia resource, by averaging multiple means of the second feedback values output by multiple Gaussian probability models. By determining the maximum value among the multiple variances of the second feedback values output by the multiple Gaussian probability models as the first penalty degree of the first feedback value, the reliability of the predicted interaction behavior between the target object and the target multimedia resource predicted by the integrated prediction model can be indirectly determined by determining the dispersion of the first feedback value.
[0081] In some embodiments, the method further includes:
[0082] For any Gaussian probability model, based on the Gaussian probability model, the object features of multiple sample objects and the multimedia resource features of multiple sample multimedia resources are predicted to obtain the mean and variance of the third feedback values of multiple sample objects to multiple sample multimedia resources. The third feedback values are used to indicate the predicted interaction behavior between sample objects and sample multimedia resources.
[0083] Based on the historical interaction behavior of multiple sample objects and multiple sample multimedia resources, the mean of the third feedback value, and the variance of the third feedback value, the training loss of the Gaussian probability model is determined. The training loss is used to indicate the difference between the historical interaction behavior and the predicted interaction behavior.
[0084] The model parameters of the Gaussian probability model are updated based on the training loss.
[0085] In this embodiment, the electronic device, based on the mean and variance of the third feedback value predicted by the Gaussian probability model and the historical interaction behavior of the sample object with the sample multimedia resource, can determine the training loss of the Gaussian probability model used to indicate the difference between the historical and predicted interaction behavior of the sample object and the sample multimedia resource. The training loss is negatively correlated with the accuracy of the Gaussian probability model. The electronic device can update the model parameters of the Gaussian probability model based on the training loss to reduce the training loss and thus improve the accuracy of the Gaussian probability model.
[0086] In some embodiments, the method further includes:
[0087] Based on the object identifier of the target object, sample multimedia resources associated with the object identifier are determined from historical data;
[0088] Based on sample multimedia resources and target multimedia resources, the relative entropy of the multimedia resource recommendation model is determined. The relative entropy is used to indicate the difference between the target multimedia resources and sample multimedia resources.
[0089] The second penalty degree is determined based on the relative entropy, and the relative entropy is positively correlated with the second penalty degree.
[0090] In this embodiment, the electronic device, based on relative entropy, can determine the difference between the target multimedia resource and multiple sample multimedia resources historically recommended to the target object in historical data. By analyzing the difference between the target multimedia resource and the sample multimedia resources, the electronic device can determine the degree of randomness in the appearance of the target multimedia resource among the multiple sample multimedia resources, that is, the second penalty degree of the target object's first feedback value to the media resource. The larger the second penalty degree, the greater the degree of randomness in the appearance of the target multimedia resource among the multiple sample multimedia resources. This indicates that the multimedia resource recommendation model recommends a greater diversity of target multimedia resources to the target object, and the multimedia resource recommendation model is not limited to recommending sample multimedia resources with high interaction frequency with the target object in historical data.
[0091] In some embodiments, adjusting the first feedback value based on a first penalty degree and a second penalty degree of the first feedback value includes:
[0092] The weight of the second penalty degree is determined based on the relative entropy, and the weight is negatively correlated with the relative entropy;
[0093] Based on the weight of the second penalty degree, the first feedback value, the first penalty degree, and the second penalty degree are weighted and summed to obtain the adjusted first feedback value.
[0094] In this embodiment of the disclosure, the electronic device performs a weighted summation of the first feedback value, the first penalty degree, and the second penalty degree based on the weight of the second penalty degree to obtain the adjusted first feedback value. This can avoid learning an overly conservative recommendation strategy during the subsequent training of the multimedia resource recommendation model, alleviate the Matthew effect when recommending multimedia resources, and improve the diversity of multimedia resources recommended by the multimedia resource recommendation model.
[0095] In some embodiments, a multimedia resource recommendation model is trained based on the adjusted first feedback value, the object features of the target object, and the multimedia resource features of the target multimedia resource, including:
[0096] Based on the adjusted first feedback value, the object characteristics of the target object, and the multimedia resource characteristics of the target multimedia resource, the preference information of the target object is determined. The preference information is used to indicate whether the target object is interested in the target multimedia resource.
[0097] Based on preference information and the adjusted first feedback value, the model parameters of the multimedia resource recommendation model are adjusted.
[0098] In this embodiment, the electronic device can adjust the model parameters of the multimedia resource recommendation model based on the target object's preference information and the target object's first feedback value to the target multimedia resource recommended at the current moment. Based on the adjusted multimedia resource recommendation model, the electronic device can determine the target multimedia resource to recommend to the target object at the next moment from multiple sample multimedia resources in historical data. During training, the multimedia resource recommendation model can gradually learn the target object's preference information and adjust its model parameters based on the adjusted first feedback value of the target object to the target multimedia resource. This improves the accuracy of the multimedia resource recommendation model in recommending multimedia resources to the target object while avoiding learning overly conservative recommendation strategies, thus mitigating the Matthew effect when recommending multimedia resources.
[0099] In some embodiments, the method further includes:
[0100] Based on the object identifier of the target object, the sample multimedia resources associated with the object identifier and the historical interaction behavior between the target object and the sample multimedia resources are determined from historical data;
[0101] Based on the object characteristics of the target object, the multimedia resource characteristics of the sample multimedia resources, and historical interaction behavior, the historical preference information of the target object is determined. The historical preference information is used to indicate whether the target object is interested in the sample multimedia resources.
[0102] Based on the multimedia resource recommendation model, historical preference information is processed to obtain sample multimedia resources recommended by the multimedia resource recommendation model to the target object.
[0103] In this embodiment, the electronic device can determine the historical preference information of the target object based on the object characteristics of the target object, the multimedia resource characteristics of the sample multimedia resources historically recommended to the target object, and the historical interaction behavior between the target object and the sample multimedia resources. This historical preference information indicates whether the target object is interested in the sample multimedia resources. The electronic device can input the historical preference information into a multimedia resource recommendation model to be trained, and the multimedia resource recommendation model processes the historical preference information of the target object to obtain the target multimedia resources recommended to the target object.
[0104] The above Figure 2 The diagram illustrates the basic process of this disclosure. The following section will further elaborate on the solution provided in this disclosure based on a specific implementation method. Figure 3 This is a flowchart illustrating another method for training a multimedia resource recommendation model according to an exemplary embodiment. The method is performed by an electronic device; see [link to relevant documentation]. Figure 3 The method includes:
[0105] In step S301, an ensemble prediction model is trained based on historical data. The historical data includes object features of multiple sample objects, multimedia resource features of multiple sample multimedia resources, and historical interaction behaviors between multiple sample objects and multiple sample multimedia resources. The ensemble prediction model is used to predict the interaction behaviors between objects and multimedia resources.
[0106] In this embodiment, the historical data is labeled training data, including object features of multiple sample objects, multimedia resource features of multiple sample multimedia resources, and historical interaction behaviors between multiple sample objects and multiple sample multimedia resources. The electronic device can use the historical interaction behaviors between multiple sample objects and multiple sample multimedia resources as label information, and perform supervised training on the ensemble prediction model based on the historical data. The sample objects are accounts logged into on multimedia resource clients installed on the terminal. Through the sample objects, various multimedia resources, such as videos, audio, and items, can be published or viewed. The sample multimedia resources can be multimedia resources viewed by the sample objects or multimedia resources not viewed by the sample objects. The historical interaction behaviors between the sample objects and sample multimedia resources include the sample objects' clicks, viewing duration, and likes. The object features of the sample objects include object identifiers and object activity levels; the multimedia resource features of the sample multimedia resources include multimedia resource identifiers, multimedia resource types, multimedia resource durations, and multimedia resource popularity. The electronic device can obtain historical data from a local database, from other electronic devices, or based on multimedia resources published by the terminal; this embodiment does not limit the scope of the historical data.
[0107] In some embodiments, the integrated prediction model includes multiple Gaussian probabilistic models (GPMs). Accordingly, for any GPM, based on the GPM, object features of multiple sample objects and multimedia resource features of multiple sample multimedia resources are predicted to obtain the mean and variance of the third feedback values of multiple sample objects to multiple sample multimedia resources. The third feedback values are used to indicate the predicted interaction behavior between the sample objects and the sample multimedia resources. Based on the historical interaction behavior of multiple sample objects and multiple sample multimedia resources, the mean of the third feedback values, and the variance of the third feedback values, the training loss of the GPM is determined. The training loss is used to indicate the difference between historical interaction behavior and predicted interaction behavior. Based on the training loss, the model parameters of the GPM are updated. The GPM is used to predict the distribution of feedback values of objects to multimedia resources. Based on the GPM, the electronic device can predict the object features of sample objects and the resource features of sample multimedia resources in historical data, obtaining the mean and variance of the third feedback values of sample objects to sample multimedia resources. Since the third feedback value is used to indicate the predicted interaction behavior of the sample object with the sample multimedia resource, the electronic device can determine the training loss of the Gaussian probability model, which indicates the difference between the historical and predicted interaction behavior of the sample object and the sample multimedia resource, based on the mean and variance of the third feedback value and the historical interaction behavior of the sample object with the sample multimedia resource. The training loss is negatively correlated with the accuracy of the Gaussian probability model. The electronic device can update the model parameters of the Gaussian probability model based on the training loss to reduce the training loss and thus improve the accuracy of the Gaussian probability model.
[0108] In some embodiments, the third feedback value is represented as a sequence to indicate the sample object's clicks, browsing duration, and likes on the sample multimedia resource. The first bit of the sequence indicates whether the sample object clicked on the sample multimedia resource, the second bit indicates whether the sample object liked the sample multimedia resource, and the other bits in the sequence represent the browsing duration of the sample multimedia resource. Here, 1 indicates a click or a like, and 0 indicates no click or no like. For example, with a third feedback value of [1,1,1,1,0,1], it can be seen that the sample object clicked on the sample multimedia resource and liked it, and the browsing duration after the click was 2 seconds. 3 +2 2 +2 0 = 13 seconds.
[0109] In some embodiments, the electronic device determines the training loss of the Gaussian probability model based on the historical interaction behavior of multiple sample objects and multiple sample multimedia resources, the mean of the third feedback value, and the variance of the third feedback value using the following formula 1.
[0110] Formula 1:
[0111]
[0112] The ensemble prediction model includes K Gaussian probability models. For the Kth Gaussian probability model θ k The training loss is N, where N is the data pair x consisting of sample objects and sample multimedia resources in the historical data. i Quantity, y i For data pair x i Historical interaction behavior between sample objects and sample multimedia resources For data pair x i The mean of the third feedback values of the sample objects to the sample multimedia resources. The variance of the third feedback value.
[0113] In step S302, the integrated prediction model includes multiple Gaussian probability models. For any Gaussian probability model, based on the Gaussian probability model, the object features of the target object and the multimedia resource features of the target multimedia resource are predicted to obtain the mean and variance of the second feedback value of the target object to the target multimedia resource. The target object is any object among multiple sample objects, and the target multimedia resource is the sample multimedia resource recommended to the target object by the multimedia resource recommendation model among multiple sample multimedia resources. The multimedia resource recommendation model is used to predict the multimedia resources recommended to the object, and the second feedback value is used to indicate the initial interaction behavior of the target object to the target multimedia resource.
[0114] In this embodiment of the disclosure, the multimedia resource recommendation model to be trained is used to predict the multimedia resources recommended to an object, and the electronic device is able to obtain the target multimedia resources recommended by the multimedia resource recommendation model to the target object. The ensemble prediction model includes multiple Gaussian probability models, which are used to predict the distribution of the object's feedback values to the multimedia resources. During the process of the electronic device predicting the object features of the target object and the resource features of the target multimedia resources based on the ensemble prediction model, the input to each Gaussian probability model is the object features of the target object and the multimedia resource features of the target multimedia resource, and the output is the mean and variance of the second feedback value of the target object to the target multimedia resource. Here, the target object is any object among multiple sample objects in the historical data, and the target multimedia resource is a sample multimedia resource among multiple sample multimedia resources in the historical data. The second feedback value is used to indicate the initial interaction behavior of the target object with the target multimedia resource. The parameter values of the model parameters of each Gaussian probability model are different, and the mean and variance of the obtained second feedback values are also different.
[0115] In some embodiments, the multimedia resource recommendation model determines target multimedia resources to recommend to the target object based on the target object's historical preference information. Accordingly, the electronic device, based on the target object's object identifier, determines sample multimedia resources associated with the object identifier and the target object's historical interaction behavior with the sample multimedia resources from historical data; based on the target object's object features, the multimedia resource features of the sample multimedia resources, and the historical interaction behavior, it determines the target object's historical preference information, which indicates whether the target object is interested in the sample multimedia resources; based on the multimedia resource recommendation model, it processes the historical preference information to obtain the sample multimedia resources recommended to the target object by the multimedia resource recommendation model. Specifically, the sample multimedia resources associated with the target object's object identifier in the historical data are the historically recommended sample multimedia resources to the target object. The electronic device can determine the target object's historical preference information based on the target object's object features, the multimedia resource features of the historically recommended sample multimedia resources to the target object, and the target object's historical interaction behavior with the sample multimedia resources. This historical preference information indicates whether the target object is interested in the sample multimedia resources. The electronic device can obtain the target multimedia resources to recommend to the target object by inputting the historical preference information into the multimedia resource recommendation model to be trained, and by having the multimedia resource recommendation model process the target object's historical preference information.
[0116] In step S303, the average of multiple mean values of the second feedback value is determined as the first feedback value, which is used to indicate the predicted interaction behavior between the target object and the target multimedia resource.
[0117] In this embodiment of the disclosure, the electronic device, based on multiple Gaussian probability models, can predict multiple means of the second feedback value. The electronic device averages these multiple means to obtain a first feedback value of the target object on the target multimedia resource, output by the integrated prediction model. The first feedback value indicates the predicted interaction behavior between the target object and the target multimedia resource, such as the target object's click activity, browsing duration, or "like" behavior.
[0118] In some embodiments, the electronic device determines the first feedback value by averaging the multiple mean values of the second feedback value using the following formula 2.
[0119] Formula 2:
[0120]
[0121] in, As the first feedback value, a t For the multimedia resource recommendation model, the target multimedia resource recommended for the target object at the current time t is s. t To provide the target object's preference information at the current time t, the ensemble prediction model includes K Gaussian probability models. For the Kth Gaussian probability model θ k Data pairs x are composed of the object characteristics of the target object and the multimedia resource characteristics of the target multimedia resources. i The mean of the predicted second feedback values.
[0122] In step S304, the maximum value of the multiple variances of the second feedback value is determined as the first penalty degree, which is used to indicate the degree of dispersion of the first feedback value.
[0123] In this embodiment, the electronic device, based on multiple Gaussian probability models, can predict multiple variances of the second feedback value. These variances reflect the dispersion of the second feedback value and can be considered a representation of uncertainty. Since uncertainty indicates the degree of uncertainty regarding data, the variances can indirectly reflect the reliability of the second feedback value. The electronic device determines the maximum value among the multiple variances as the first penalty degree of the first feedback value. By determining the dispersion of the first feedback value, it can indirectly determine the reliability of the predicted interaction behavior between the target object and the target multimedia resource predicted by the integrated prediction model. Specifically, the first penalty degree is positively correlated with the dispersion; the larger the first penalty degree, the greater the dispersion of the first feedback value, and the lower the reliability of the predicted interaction behavior between the target object and the target multimedia resource predicted by the integrated prediction model.
[0124] In some embodiments, the electronic device determines the maximum value of a plurality of variances of the second feedback value as the first penalty degree using the following Formula 3.
[0125] Formula 3:
[0126]
[0127] Among them, P U As the first penalty degree, the ensemble prediction model consists of K Gaussian probability models. For the Kth Gaussian probability model θ k Data pairs x are composed of the object characteristics of the target object and the multimedia resource characteristics of the target multimedia resources. i The variance of the predicted second feedback value.
[0128] In step S305, based on the object identifier of the target object, sample multimedia resources associated with the object identifier are determined from historical data.
[0129] In this embodiment of the disclosure, the target object is any one of a plurality of sample objects. Therefore, based on the object identifier of the target object, the electronic device can determine the sample object with the same object identifier as the target object from historical data, and then determine the sample multimedia resource associated with the sample object. The sample multimedia resource can be a single multimedia resource or a collection of multimedia resources, including multiple multimedia resources; this application embodiment does not impose any limitation on this.
[0130] In step S306, based on the sample multimedia resources and the target multimedia resources, the relative entropy of the multimedia resource recommendation model is determined. The relative entropy is used to indicate the difference between the target multimedia resources and the sample multimedia resources.
[0131] In this embodiment, the target multimedia resource is a multimedia resource selected by the multimedia resource recommendation model from multiple sample multimedia resources in historical data and recommended to the target object. This target multimedia resource reflects the current recommendation strategy of the multimedia resource recommendation model for the target object. Sample multimedia resources in the historical data associated with the object identifier of the target object reflect historical recommendation strategies for the target object. Based on the sample multimedia resources and the target multimedia resource, the electronic device can calculate the entropy of the multimedia resource recommendation model's current recommendation strategy during the learning process of historical recommendation strategies, i.e., determine the relative entropy of the multimedia resource recommendation model. The relative entropy indicates the difference between the target multimedia resource and the sample multimedia resources, i.e., the difference between the current recommendation strategy and the historical recommendation strategy of the multimedia resource recommendation model.
[0132] In step S307, a second penalty degree is determined based on the relative entropy of the first feedback value. The relative entropy is positively correlated with the second penalty degree, which is used to indicate the degree of randomness in the appearance of the target multimedia resource among multiple sample multimedia resources.
[0133] In this embodiment, the electronic device, based on relative entropy, can determine the difference between the target multimedia resource and multiple sample multimedia resources historically recommended to the target object in historical data. Since the target multimedia resource is a sample multimedia resource selected by the multimedia resource recommendation model from historical data, the electronic device can determine the degree of randomness of the target multimedia resource's appearance among multiple sample multimedia resources by analyzing the difference between the target multimedia resource and the sample multimedia resources. This is also known as the second penalty degree of the target object's first feedback value to the media resource. The relative entropy is positively correlated with the second penalty degree; the greater the relative entropy, the greater the difference between the current recommendation strategy and the historical recommendation strategy of the multimedia resource recommendation model. Therefore, the randomness of the target multimedia resource recommended by the multimedia resource recommendation model is higher. The greater the degree of randomness of the target multimedia resource's appearance among multiple sample multimedia resources, the greater the second penalty degree. This indicates that the multimedia resource recommendation model recommends a higher diversity of target multimedia resources for the target object, and the multimedia resource recommendation model is not limited to recommending sample multimedia resources with high interaction frequency with the target object in historical data.
[0134] In some embodiments, the electronic device uses an entropy penalty mechanism to determine the relative entropy of the multimedia resource recommendation model. For example, Figure 4 This is a schematic diagram illustrating the workflow of an entropy penalty device according to an exemplary embodiment. See also Figure 4 Suppose we extract a historical interaction trajectory between a target object and sample multimedia resources from historical data as [1,0,9,4,2,1,7,3,8]. The multimedia resource prediction model predicts that the target multimedia resource interacting with the target object at the current time t is sample multimedia resource number 8 in the historical interaction trajectory. Then, the first-order entropy of the multimedia resource recommendation model is the entropy value of the multimedia resource corresponding to the "?" position after sample multimedia resource number 8 in the historical interaction trajectory of the target object and sample multimedia resources. That is, given that the target multimedia resource interacting with the target object at the current time is sample multimedia resource number 8, the randomness of the target multimedia resource appearing among multiple sample multimedia resources predicted by the multimedia resource recommendation model at the next time step can be statistically obtained from historical data. Similarly, the second-order and third-order entropies can be calculated using the same method, i.e., the entropy values of the multimedia resources corresponding to the "?" positions in the subsequences [3,8,?] and [7,3,8,?]. In any subsequence, the positions of the sample multimedia resources before the "?" can be freely interchanged; that is, [3,7,8,?], [7,8,3,?], and [7,3,8,?] are all indistinguishable. The electronic device sums the multi-order entropy to obtain the relative entropy of the multimedia resource recommendation model.
[0135] In step S308, the weight of the second penalty degree is determined based on the relative entropy, and the weight is negatively correlated with the relative entropy.
[0136] In this embodiment, the electronic device determines the weight of the second penalty degree based on relative entropy. When the relative entropy is large, the second penalty degree is large, resulting in a higher diversity of target multimedia resources recommended by the multimedia resource recommendation model. This indicates that the current recommendation strategy learned by the multimedia resource recommendation model from historical recommendation strategies is not a conservative one. Therefore, the electronic device can reduce the weight of the second penalty degree to reduce the penalty intensity on the first feedback value. This prevents the multimedia resource recommendation model trained based on the first feedback value from learning an overly conservative recommendation strategy, especially when the target multimedia resources recommended by the multimedia resource recommendation model are too conservative.
[0137] In step S309, the first feedback value, the first penalty degree, and the second penalty degree are weighted and summed based on the weight of the second penalty degree to obtain the adjusted first feedback value.
[0138] In this embodiment, the electronic device, based on a first penalty degree, can determine the dispersion of the first feedback value, i.e., the credibility of the target object's predicted interaction behavior with the target multimedia resources recommended by the multimedia resource recommendation model. Based on a second penalty degree, the electronic device can determine the randomness of the occurrence of the target multimedia resources among multiple sample multimedia resources, i.e., the diversity of the target multimedia resources recommended by the multimedia resource recommendation model for the target object. Based on the weight of the second penalty degree, the electronic device performs a weighted summation of the first feedback value, the first penalty degree, and the second penalty degree to obtain an adjusted first feedback value. This avoids learning an overly conservative recommendation strategy during subsequent training of the multimedia resource recommendation model, alleviates the Matthew effect in recommending multimedia resources, and improves the diversity of multimedia resources recommended by the multimedia resource recommendation model.
[0139] In some embodiments, the electronic device uses the following formula (Formula 4) to perform a weighted summation of the first feedback value, the first penalty degree, and the second penalty degree to obtain the adjusted first feedback value.
[0140] Formula 4:
[0141]
[0142] in, The first feedback value after adjustment. As the first feedback value, P U As the first penalty degree, P E Let λ1 and λ2 be the weights of the first and second penalty degrees, respectively, and let a be the second penalty degree. tFor the multimedia resource recommendation model, the target multimedia resource recommended for the target object at the current time t is s. t This refers to the preference information of the target object at the current time t.
[0143] In step S310, based on the adjusted first feedback value, the object characteristics of the target object, and the multimedia resource characteristics of the target multimedia resource, the preference information of the target object is determined. The preference information is used to indicate whether the target object is interested in the target multimedia resource.
[0144] In this embodiment of the disclosure, the first feedback value is used to indicate the predicted interaction behavior of the target object with the target multimedia resource recommended at the current moment. Based on the predicted interaction behavior of the target object with the target multimedia resource, the object characteristics of the target object, and the multimedia resource characteristics of the target multimedia resource, the electronic device can determine the target object's preference information for the next moment. This preference information is used to indicate whether the target object is interested in the target multimedia resource recommended by the multimedia resource recommendation model at the next moment.
[0145] In some embodiments, the electronic device determines the preference information of the target object based on the adjusted first feedback value, the object characteristics of the target object, and the multimedia resource characteristics of the target multimedia resource using the following Formula 5.
[0146] Formula 5:
[0147]
[0148] in, a represents the representation vector of the target object's preference information at the next time step t+1. n This represents the target multimedia resources recommended by the multimedia resource recommendation model for the target object. This represents a at time t-N+1. n The representation vector, This is the first feedback value after adjustment at time t-N+1, where N is a window size representing the number of a values to be calculated. n , The quantity.
[0149] In step S311, the model parameters of the multimedia resource recommendation model are adjusted based on preference information and the adjusted first feedback value.
[0150] In this embodiment, during the training of the multimedia resource recommendation model, the electronic device inputs the adjusted first feedback value and the target object's preference information for the next time step into the multimedia resource recommendation model. The multimedia resource recommendation model processes the adjusted first feedback value and preference information to determine the target multimedia resource to be recommended to the target object for the next time step. Specifically, the electronic device can adjust the model parameters of the multimedia resource recommendation model based on the target object's preference information for the next time step and the target object's first feedback value for the target multimedia resource recommended at the current time step. Based on the adjusted multimedia resource recommendation model, the electronic device can determine the target multimedia resource to be recommended to the target object for the next time step from multiple sample multimedia resources in historical data. During the training process, the multimedia resource recommendation model can gradually learn the target object's preference information and adjust the model parameters based on the adjusted first feedback value of the target object for the target multimedia resource. This improves the accuracy of the multimedia resource recommendation model in recommending multimedia resources to the target object while avoiding learning overly conservative recommendation strategies, thus mitigating the Matthew effect when recommending multimedia resources.
[0151] For example, Figure 5 This is a framework diagram illustrating a training method for a multimedia resource recommendation model according to an exemplary embodiment. For example... Figure 5 As shown in the diagram, the framework includes four key modules: an ensemble prediction model, an entropy penalizer, a state tracker, and a multimedia resource recommendation model. The ensemble prediction model models the environment for the multimedia resource recommendation process. Here, the environment is a real-world object, so the ensemble prediction model also acts as an object simulator. The ensemble prediction model consists of K Gaussian probability models, trained on historical data. The outputs of the K Gaussian probability models have the same meaning. The input to the k-th Gaussian probability model is a data pair x consisting of a sample object and a sample multimedia resource, and y representing the historical interaction behavior y between the sample object and the sample multimedia resource. The output is two statistics representing the feedback value of the sample object to the sample multimedia resource: the mean and the mean. and variance After the ensemble prediction model is trained, its parameters need to be fixed. Then, based on the first feedback value of the target object to the target multimedia resource provided by the ensemble prediction model, a multimedia resource recommendation model is trained. During the training process, the multimedia resource recommendation model determines the target multimedia resource 'a' to recommend to the target object based on the target object's historical preference information 's'. The ensemble prediction model can predict the object features of the target object and the multimedia resource features of the target multimedia resource 'a', obtaining the first feedback value. And the first penalty degree P of the first feedback valueU The first feedback value is the average of multiple means output by the Gaussian probability model, and the first penalty is the maximum of multiple variances output by the Gaussian probability model. The entropy penalizer is used to determine the entropy of the multimedia resource recommendation model's current recommendation strategy during the learning process of historical recommendation strategies, i.e., the relative entropy of the multimedia resource recommendation model. Then, based on this relative entropy, the second penalty P of the first feedback value is determined. E The state tracker is used to recommend target multimedia resources 'a' based on the target object and the adjusted first feedback value. The system determines the target object's preference information for the next time step. This state tracker can be implemented using any temporal neural network, such as a recurrent neural network (RNN), a convolutional neural network (CNN), or an attention-based Transformer network.
[0152] This disclosure provides a training method for a multimedia resource recommendation model. Based on an ensemble prediction model, it predicts the object features of a target object and the multimedia resource features of a target multimedia resource. This yields a first feedback value indicating the interaction behavior between the target object and the target multimedia resource, and a first penalty value indicating the degree of dispersion of the first feedback value. Since the target multimedia resource is a sample multimedia resource recommended by the multimedia resource recommendation model for the target object, adjusting the first feedback value based on the first penalty value and a second penalty value indicating the randomness of the target multimedia resource's appearance among multiple sample multimedia resources can penalize the first feedback value if the recommended target multimedia resource is too conservative. This avoids the trained multimedia resource recommendation model learning an overly conservative recommendation strategy, alleviates the Matthew effect in multimedia resource recommendation models, and improves the diversity of multimedia resources recommended by the multimedia resource recommendation model.
[0153] All of the above-mentioned optional technical solutions can be combined in any way to form optional embodiments of this disclosure, and will not be described in detail here.
[0154] Based on the above embodiments, this disclosure also evaluates the effectiveness of the multimedia resource recommendation model on the datasets KuaiRec and KuaiRand. Figure 6 This is a schematic diagram illustrating the test results of a multimedia resource recommendation model according to an exemplary embodiment, such as... Figure 6 As shown, Figure 6The vertical axes (A1) and (B1) represent the cumulative reward of the target object, which indicates the target object's cumulative satisfaction with multiple multimedia resources recommended by the multimedia resource recommendation model. Figure 6 (A1) and Figure 6 As shown in the curve distribution in (B1), compared with the current mainstream training methods, the training method (DORL) provided in this embodiment of the disclosure achieves the best results in terms of cumulative gain. It obtained the highest gain curve after 200 rounds of training. Figure 6 The vertical axes of (A2) and (B2) represent the length of the interaction trajectory between the target object and the multimedia resources, that is, the number of multimedia resources interacted with by the target object. It can be seen that the training method provided in this embodiment has the longest interaction trajectory, which means that the target object has a high degree of satisfaction with the multimedia resources recommended by the multimedia resource recommendation model during the interaction with the sample multimedia resources, and will not easily quit. Figure 6 The vertical axes (A3) and (B3) represent the target object's satisfaction level in a single round of interaction, i.e., the target object's satisfaction with a single multimedia resource. Compared to other mainstream methods, the training method (DORL) provided in this embodiment also achieves good single-round satisfaction. As can be seen from the above, the multimedia resource recommendation model trained using the training method provided in this embodiment is more effective. Figure 7 This is a schematic diagram illustrating experimental results of a training method for a multimedia resource recommendation model based on an exemplary embodiment, addressing the Matthew effect. (Example:) Figure 7 As shown, Figure 7 In diagrams (1) and (2), the broken lines represent the length of the interaction trajectory between the target object and the multimedia resources, while the bar charts represent the proportion of the most mainstream features among the multimedia resources recommended to the target object. As the weight λ2 of the second penalty degree (represented by the horizontal axis) increases, the penalty for the target object's feedback value to the multimedia resources gradually intensifies. Figure 7 It can be seen that the proportion of mainstream features in the multimedia resources recommended by the multimedia resource recommendation model to the target object is gradually decreasing. While the diversity of multimedia resources is gradually increasing, a longer interaction trajectory can be obtained. Therefore, it can be seen that the training method provided in this embodiment has achieved a good effect in alleviating the Matthew effect. Figure 8 This is a schematic diagram illustrating experimental results of a training method for a multimedia resource recommendation model under different exit conditions, according to an exemplary embodiment. Figure 8 As shown, Figure 8In (1) and (2), the horizontal axis N represents the number of multimedia resources used to calculate the exit condition in the simulation environment. Taking N=3 as an example, the similarity between the currently recommended multimedia resource and the three previously recommended multimedia resources will be calculated. If the similarity between the current multimedia resource and the three previously recommended multimedia resources is less than a certain threshold, it can be judged as a duplicate recommendation. Under certain duplicate recommendation conditions, the target audience will feel bored and exit browsing multimedia resources. As N increases, the user's sensitivity to duplicate recommendations will gradually increase, making it easier to exit interaction with multimedia resources. Figure 8 As shown, as N increases, the cumulative satisfaction of the target object obtained by the training method (DORL) provided in this embodiment of the present disclosure is not significantly affected, and all values exceed those of other comparative methods. Therefore, it can be concluded that the training method (DORL) provided in this embodiment of the present disclosure exhibits good robustness under different target object sensitivities.
[0155] In some embodiments, the electronic device determines the cumulative satisfaction level of the target object with multiple multimedia resources recommended by the multimedia resource recommendation model using the following Formula Six.
[0156] Formula Six:
[0157]
[0158] Wherein, J(π) θ π represents the cumulative satisfaction level of the target audience. θ The recommendation strategy for the multimedia resource recommendation model, π θ =π θ (a t |s t ) is used to indicate the preference information s of the target object at the current time t. t Recommend target multimedia resources to the target audience afterwards. t The probability τ is the probability of the multimedia resource recommendation model based on strategy π. θ The interaction trajectory left after the target multimedia resource recommended to the target object interacts with the target object, where H is the length of the interaction trajectory and γ is a discount factor in the range (0,1). This is the adjusted first feedback value.
[0159] Figure 9 This is a block diagram illustrating a training apparatus for a multimedia resource recommendation model according to an exemplary embodiment. (Refer to...) Figure 9 The device includes: an acquisition unit 901, a first prediction unit 902, an adjustment unit 903, and a training unit 904.
[0160] The acquisition unit 901 is configured to acquire an integrated prediction model. The integrated prediction model is trained based on historical data, which includes object features of multiple sample objects, multimedia resource features of multiple sample multimedia resources, and historical interaction behavior between multiple sample objects and multiple sample multimedia resources. The integrated prediction model is used to predict the interaction behavior between objects and multimedia resources.
[0161] The first prediction unit 902 is configured to predict the object features of the target object and the multimedia resource features of the target multimedia resource based on an integrated prediction model, and obtain a first feedback value of the target object to the target multimedia resource and a first penalty degree of the first feedback value. The target object is any object among multiple sample objects, and the target multimedia resource is the sample multimedia resource recommended to the target object by the multimedia resource recommendation model among multiple sample multimedia resources. The first feedback value is used to indicate the predicted interaction behavior between the target object and the target multimedia resource, the first penalty degree is used to indicate the degree of dispersion of the first feedback value, and the multimedia resource recommendation model is used to predict the multimedia resources recommended to the object.
[0162] The adjustment unit 903 is configured to adjust the first feedback value based on a first penalty degree and a second penalty degree of the first feedback value. The second penalty degree is used to indicate the degree of randomness of the occurrence of the target multimedia resource among multiple sample multimedia resources.
[0163] Training unit 904 is configured to train a multimedia resource recommendation model based on the adjusted first feedback value, the object features of the target object, and the multimedia resource features of the target multimedia resource.
[0164] In some embodiments, the integrated prediction model includes multiple Gaussian probability models, which are used to predict the distribution of the object's feedback values to multimedia resources.
[0165] Figure 10 This is a block diagram illustrating a training apparatus for another multimedia resource recommendation model according to an exemplary embodiment. See also Figure 10 Prediction unit 902 includes:
[0166] Prediction subunit 1001 is configured to predict the object features of the target object and the multimedia resource features of the target multimedia resource based on any Gaussian probability model, and obtain the mean and variance of the second feedback value of the target object to the target multimedia resource. The second feedback value is used to indicate the initial interaction behavior of the target object to the target multimedia resource.
[0167] The first determining subunit 1002 is configured to determine the average of multiple averages of the second feedback value as the first feedback value;
[0168] The second determining subunit 1003 is configured to determine the maximum value of multiple variances of the second feedback value as the first penalty degree.
[0169] In some embodiments, see continue to see Figure 10 The device also includes:
[0170] The second prediction unit 905 is configured to predict the object features of multiple sample objects and the multimedia resource features of multiple sample multimedia resources based on any Gaussian probability model, and obtain the mean and variance of the third feedback values of multiple sample objects to multiple sample multimedia resources. The third feedback values are used to indicate the predicted interaction behavior between the sample objects and the sample multimedia resources.
[0171] The first determining unit 906 is configured to determine the training loss of the Gaussian probability model based on the historical interaction behavior of multiple sample objects and multiple sample multimedia resources, the mean of the third feedback value, and the variance of the third feedback value. The training loss is used to indicate the difference between the historical interaction behavior and the predicted interaction behavior.
[0172] Update unit 907 is configured to update the model parameters of the Gaussian probability model based on the training loss.
[0173] In some embodiments, see continue to see Figure 10 The device also includes:
[0174] The second determining unit 908 is configured to determine sample multimedia resources associated with the object identifier from historical data based on the object identifier of the target object.
[0175] The third determining unit 909 is configured to determine the relative entropy of the multimedia resource recommendation model based on the sample multimedia resources and the target multimedia resources. The relative entropy is used to indicate the difference between the target multimedia resources and the sample multimedia resources.
[0176] The fourth determining unit 910 is configured to determine the second penalty degree based on the relative entropy, wherein the relative entropy is positively correlated with the second penalty degree.
[0177] In some embodiments, the adjustment unit 903 is configured to determine the weight of the second penalty degree based on the relative entropy, the weight being negatively correlated with the relative entropy; and to perform a weighted summation of the first feedback value, the first penalty degree, and the second penalty degree based on the weight of the second penalty degree to obtain the adjusted first feedback value.
[0178] In some embodiments, see continue to see Figure 10 Training unit 904 includes:
[0179] The third determining subunit 1004 is configured to determine the target object's preference information based on the adjusted first feedback value, the object characteristics of the target object, and the multimedia resource characteristics of the target multimedia resource. The preference information is used to indicate whether the target object is interested in the target multimedia resource.
[0180] Subunit 1005 is configured to adjust the model parameters of the multimedia resource recommendation model based on preference information and the adjusted first feedback value.
[0181] In some embodiments, see continue to see Figure 10 The device also includes:
[0182] The fifth determining unit 911 is configured to determine, based on the object identifier of the target object, the sample multimedia resources associated with the object identifier and the historical interaction behavior between the target object and the sample multimedia resources from historical data;
[0183] The sixth determining unit 912 is configured to determine the historical preference information of the target object based on the object characteristics of the target object, the multimedia resource characteristics of the sample multimedia resources, and the historical interaction behavior. The historical preference information is used to indicate whether the target object is interested in the sample multimedia resources.
[0184] The processing unit 913 is configured to process historical preference information based on the multimedia resource recommendation model to obtain sample multimedia resources recommended by the multimedia resource recommendation model to the target object.
[0185] This disclosure provides a training apparatus for a multimedia resource recommendation model. Based on an ensemble prediction model, it predicts the object features of a target object and the multimedia resource features of a target multimedia resource. This allows for the generation of a first feedback value indicating the interaction between the target object and the target multimedia resource, and a first penalty value indicating the degree of dispersion of the first feedback value. Since the target multimedia resource is a sample multimedia resource recommended by the multimedia resource recommendation model for the target object, adjusting the first feedback value based on the first penalty value and a second penalty value indicating the randomness of the target multimedia resource's appearance among multiple sample multimedia resources can penalize the first feedback value if the recommended target multimedia resource is too conservative. This avoids the trained multimedia resource recommendation model learning an overly conservative recommendation strategy, alleviates the Matthew effect in multimedia resource recommendation models, and improves the diversity of multimedia resources recommended by the multimedia resource recommendation model.
[0186] It should be noted that the training device for the multimedia resource recommendation model provided in the above embodiments is only illustrated by the division of the above functional units when running the application. In practical applications, the above functions can be assigned to different functional units as needed, that is, the internal structure of the electronic device can be divided into different functional units to complete all or part of the functions described above. In addition, the training device for the multimedia resource recommendation model provided in the above embodiments and the training method embodiments for the multimedia resource recommendation model belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0187] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0188] When an electronic device is provided as a terminal, Figure 11 This is a block diagram illustrating a terminal 1100 according to an exemplary embodiment. The terminal... Figure 11 A structural block diagram of a terminal 1100 provided in an exemplary embodiment of this disclosure is shown. The terminal 1100 may be a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. The terminal 1100 may also be referred to as a user device, portable terminal, laptop terminal, desktop terminal, or other names.
[0189] Typically, terminal 1100 includes a processor 1101 and a memory 1102.
[0190] Processor 1101 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1101 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1101 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1101 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 1101 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0191] The memory 1102 may include one or more computer-readable storage media, which may be non-transitory. The memory 1102 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1102 are used to store at least one program code, which is executed by the processor 1101 to implement the training method and multimedia resource recommendation method of the multimedia resource recommendation model provided in the method embodiments of this disclosure.
[0192] In some embodiments, the terminal 1100 may also optionally include a peripheral device interface 1103 and at least one peripheral device. The processor 1101, memory 1102, and peripheral device interface 1103 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 1103 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 1104, a display screen 1105, a camera assembly 1106, an audio circuit 1107, a positioning assembly 1108, and a power supply 1109.
[0193] Peripheral device interface 1103 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 1101 and memory 1102. In some embodiments, processor 1101, memory 1102 and peripheral device interface 1103 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 1101, memory 1102 and peripheral device interface 1103 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0194] The radio frequency (RF) circuit 1104 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1104 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1104 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 1104 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 1104 can communicate with other terminals via at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: metropolitan area networks (MANs), various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks (WLANs), and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 1104 may also include circuitry related to NFC (Near Field Communication), which is not limited in this disclosure.
[0195] Display screen 1105 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 1105 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 1101 for processing. In this case, display screen 1105 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1105, which serves as the front panel of terminal 1100; in other embodiments, there may be at least two display screens, respectively disposed on different surfaces of terminal 1100 or in a folded design; in still other embodiments, display screen 1105 may be a flexible display screen, disposed on a curved or folded surface of terminal 1100. Furthermore, display screen 1105 may also be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. The display screen 1105 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0196] The camera assembly 1106 is used to acquire images or videos. Optionally, the camera assembly 1106 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 1106 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.
[0197] The audio circuit 1107 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 1101 for processing, or input to the radio frequency circuit 1104 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each positioned at a different location on the terminal 1100. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 1101 or the radio frequency circuit 1104 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 1107 may also include a headphone jack.
[0198] Power supply 1108 is used to power the various components in terminal 1100. Power supply 1108 can be AC power, DC power, a disposable battery, or a rechargeable battery. When power supply 1108 includes a rechargeable battery, the rechargeable battery can support wired charging or wireless charging. The rechargeable battery can also be used to support fast charging technology.
[0199] In some embodiments, the terminal 1100 further includes one or more sensors 1109. The one or more sensors 1109 include, but are not limited to: an accelerometer 1110, a gyroscope 1111, a pressure sensor 1112, an optical sensor 99, and a proximity sensor 1114.
[0200] Accelerometer 1110 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established with terminal 1100. For example, accelerometer 1110 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 1101 can control display screen 1105 to display the user interface in either a landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 1110. Accelerometer 1110 can also be used for games or for acquiring user motion data.
[0201] The gyroscope sensor 1111 can detect the orientation and rotation angle of the terminal 1100. The gyroscope sensor 1111 can work in conjunction with the accelerometer sensor 1110 to collect the user's 3D movements on the terminal 1100. Based on the data collected by the gyroscope sensor 1111, the processor 1101 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.
[0202] The pressure sensor 1112 can be disposed on the side bezel of the terminal 1100 and / or on the lower layer of the display screen 1105. When the pressure sensor 1112 is disposed on the side bezel of the terminal 1100, it can detect the user's grip signal on the terminal 1100, and the processor 1101 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 1112. When the pressure sensor 1112 is disposed on the lower layer of the display screen 1105, the processor 1101 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 1105. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0203] An optical sensor 1113 is used to collect ambient light intensity. In one embodiment, the processor 1101 can control the display brightness of the display screen 1105 based on the ambient light intensity collected by the optical sensor 1113. Optionally, when the ambient light intensity is high, the display brightness of the display screen 1105 is increased; when the ambient light intensity is low, the display brightness of the display screen 1105 is decreased. In another embodiment, the processor 1101 can also dynamically adjust the shooting parameters of the camera assembly 1106 based on the ambient light intensity collected by the optical sensor 1113.
[0204] The proximity sensor 1114, also known as the distance sensor, is installed on the front panel of the terminal 1100. The proximity sensor 1114 is used to detect the distance between the user and the front of the terminal 1100. In one embodiment, when the proximity sensor 1114 detects that the distance between the user and the front of the terminal 1100 is gradually decreasing, the processor 1101 controls the display screen 1105 to switch from a screen-on state to a screen-off state; when the proximity sensor 1114 detects that the distance between the user and the front of the terminal 1100 is gradually increasing, the processor 1101 controls the display screen 1105 to switch from a screen-off state to a screen-on state.
[0205] Those skilled in the art will understand that Figure 11 The structure shown does not constitute a limitation on terminal 1100 and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0206] When electronic devices are provided as servers, Figure 12This is a block diagram illustrating a server 1200 according to an exemplary embodiment. The server 1200 can vary significantly due to differences in configuration or performance. It may include one or more Central Processing Units (CPUs) 1201 and one or more memories 1202. The memories 1202 store at least one line of program code, which is loaded and executed by the processor 1201 to implement the training method and multimedia resource recommendation method of the multimedia resource recommendation model provided in the various method embodiments described above. Of course, the server may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server 1200 may also include other components for implementing device functions, which will not be elaborated here.
[0207] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as memory 1102 or memory 1202 including instructions. These instructions can be executed by processor 1101 of terminal 1120 or processor 1201 of server 1200 to complete the training method for the multimedia resource recommendation model described above. Optionally, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device, etc.
[0208] A computer program product includes a computer program / instructions that, when executed by a processor, implement the training method for the aforementioned multimedia resource recommendation model.
[0209] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0210] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A training method for a multimedia resource recommendation model, characterized in that, The method includes: An integrated prediction model is obtained, which is trained based on historical data. The historical data includes object features of multiple sample objects, multimedia resource features of multiple sample multimedia resources, and historical interaction behavior between the multiple sample objects and the multiple sample multimedia resources. The integrated prediction model is used to predict the interaction behavior between objects and multimedia resources. Based on the integrated prediction model, the object features of the target object and the multimedia resource features of the target multimedia resource are predicted to obtain a first feedback value of the target object to the target multimedia resource and a first penalty degree of the first feedback value. The target object is any object among the plurality of sample objects, and the target multimedia resource is the sample multimedia resource recommended to the target object by the multimedia resource recommendation model among the plurality of sample multimedia resources. The first feedback value is used to indicate the predicted interaction behavior between the target object and the target multimedia resource, and the first penalty degree is used to indicate the dispersion of the first feedback value. The multimedia resource recommendation model is used to predict the multimedia resources recommended to the object. The first feedback value is adjusted based on the first penalty degree and the second penalty degree of the first feedback value, wherein the second penalty degree is used to indicate the degree of randomness of the occurrence of the target multimedia resource among the plurality of sample multimedia resources; The multimedia resource recommendation model is trained based on the adjusted first feedback value, the object features of the target object, and the multimedia resource features of the target multimedia resource.
2. The training method for the multimedia resource recommendation model according to claim 1, characterized in that, The integrated prediction model includes multiple Gaussian probability models, which are used to predict the distribution of the object's feedback values to multimedia resources. The step of predicting the object features of the target object and the multimedia resource features of the target multimedia resource based on the integrated prediction model to obtain the first feedback value of the target object to the target multimedia resource and the first penalty degree of the first feedback value includes: For any Gaussian probability model, based on the Gaussian probability model, the object features of the target object and the multimedia resource features of the target multimedia resource are predicted to obtain the mean and variance of the second feedback value of the target object to the target multimedia resource. The second feedback value is used to indicate the initial interaction behavior of the target object to the target multimedia resource. The average of multiple mean values of the second feedback value is determined as the first feedback value; The maximum value of the multiple variances of the second feedback value is determined as the first penalty degree.
3. The training method for the multimedia resource recommendation model according to claim 2, characterized in that, The method further includes: For any Gaussian probability model, based on the Gaussian probability model, the object features of the multiple sample objects and the multimedia resource features of the multiple sample multimedia resources are predicted to obtain the mean and variance of the third feedback values of the multiple sample objects to the multiple sample multimedia resources. The third feedback values are used to indicate the predicted interaction behavior between the sample objects and the sample multimedia resources. Based on the historical interaction behavior of the multiple sample objects and the multiple sample multimedia resources, the mean of the third feedback value and the variance of the third feedback value, the training loss of the Gaussian probability model is determined, and the training loss is used to indicate the difference between the historical interaction behavior and the predicted interaction behavior. The model parameters of the Gaussian probability model are updated based on the training loss.
4. The training method for the multimedia resource recommendation model according to claim 1, characterized in that, The method further includes: Based on the object identifier of the target object, sample multimedia resources associated with the object identifier are determined from the historical data; Based on the sample multimedia resources and the target multimedia resources, the relative entropy of the multimedia resource recommendation model is determined, and the relative entropy is used to indicate the difference between the target multimedia resources and the sample multimedia resources. The second penalty degree is determined based on the relative entropy, wherein the relative entropy is positively correlated with the second penalty degree.
5. The training method for the multimedia resource recommendation model according to claim 4, characterized in that, The adjustment of the first feedback value based on the second penalty degree and the first feedback value includes: Based on the relative entropy, the weight of the second penalty degree is determined, and the weight is negatively correlated with the relative entropy; Based on the weight of the second penalty degree, the first feedback value, the first penalty degree, and the second penalty degree are weighted and summed to obtain the adjusted first feedback value.
6. The training method for the multimedia resource recommendation model according to claim 1, characterized in that, The step of training the multimedia resource recommendation model based on the adjusted first feedback value, the object features of the target object, and the multimedia resource features of the target multimedia resource includes: Based on the adjusted first feedback value, the object characteristics of the target object, and the multimedia resource characteristics of the target multimedia resource, the preference information of the target object is determined, and the preference information is used to indicate whether the target object is interested in the target multimedia resource. Based on the preference information and the adjusted first feedback value, the model parameters of the multimedia resource recommendation model are adjusted.
7. The training method for the multimedia resource recommendation model according to claim 1, characterized in that, The method further includes: Based on the object identifier of the target object, the sample multimedia resources associated with the object identifier and the historical interaction behavior between the target object and the sample multimedia resources are determined from the historical data; Based on the object characteristics of the target object, the multimedia resource characteristics of the sample multimedia resources, and the historical interaction behavior, the historical preference information of the target object is determined, and the historical preference information is used to indicate whether the target object is interested in the sample multimedia resources. Based on the multimedia resource recommendation model, the historical preference information is processed to obtain sample multimedia resources recommended by the multimedia resource recommendation model to the target object.
8. A training device for a multimedia resource recommendation model, characterized in that, The device includes: The acquisition unit is configured to acquire an integrated prediction model, which is trained based on historical data. The historical data includes object features of multiple sample objects, multimedia resource features of multiple sample multimedia resources, and historical interaction behavior between the multiple sample objects and the multiple sample multimedia resources. The integrated prediction model is used to predict the interaction behavior between the objects and the multimedia resources. The prediction unit is configured to predict the object features of the target object and the multimedia resource features of the target multimedia resource based on the integrated prediction model, and obtain a first feedback value of the target object to the target multimedia resource and a first penalty degree of the first feedback value. The target object is any object among the plurality of sample objects, and the target multimedia resource is a sample multimedia resource recommended to the target object by the multimedia resource recommendation model among the plurality of sample multimedia resources. The first feedback value is used to indicate the predicted interaction behavior between the target object and the target multimedia resource, and the first penalty degree is used to indicate the dispersion of the first feedback value. The multimedia resource recommendation model is used to predict the multimedia resources recommended to the object. The adjustment unit is configured to adjust the first feedback value based on the first penalty degree and a second penalty degree of the first feedback value, wherein the second penalty degree is used to indicate the degree of randomness of the occurrence of the target multimedia resource among the plurality of sample multimedia resources; The training unit is configured to train the multimedia resource recommendation model based on the adjusted first feedback value, the object features of the target object, and the multimedia resource features of the target multimedia resource.
9. An electronic device, characterized in that, The electronic device includes: One or more processors; Memory used to store the executable program code of the processor; The processor is configured to execute the program code to implement the training method of the multimedia resource recommendation model as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is able to perform the training method of the multimedia resource recommendation model as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Information generation model training method, information generation method, device and equipment
CN114547266A
Resource recommendation method, and multi-target fusion model training method and apparatus
WO2023040494A1