Media resource recommendation method and apparatus, and electronic device

By acquiring the triplet information of media resources and using a text generation model to generate resource descriptions, combined with user history and image information, the problem of poor recommendation performance for trending resources is solved, and more accurate media resource recommendations are achieved.

CN116955658BActive Publication Date: 2025-11-28CHINA MOBILE COMM LTD RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311011297.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-11
Publication Date
2025-11-28
Estimated Expiration
2043-08-11

AI Technical Summary

Technical Problem

Among existing media resource recommendation methods, the recommendation of trending resources is less effective and fails to accurately reflect user interests.

Method used

By acquiring the triplet information of media resources, a text generation model is used to generate resource descriptions. Combined with user history records and image information of media resources, recommended media resources are determined.

Benefits of technology

It improves the accuracy of media resource recommendations and user experience, making recommended resources more closely matched with user interests.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116955658B_ABST
    Figure CN116955658B_ABST
Patent Text Reader

Abstract

The application provides a media resource recommendation method and device and electronic equipment. The method comprises the following steps: obtaining triple information of each media resource in a first media resource set, wherein the triple information of the media resource comprises first attribute information of the media resource, second attribute information associated with the media resource, and a relationship between the first attribute information and the second attribute information; inputting the triple information of each media resource in the first media resource set into a trained text generation model to generate resource description content of each media resource in the first media resource set; obtaining resource description content and image information of media resources in a second media resource set, wherein the target media resource comprises the first media resource set and the second media resource set; and determining recommended media resources of a user to be recommended from the target media resource set based on historical behavior records of the user to be recommended and the resource description content and the image information of the media resources in the target media resource set, so as to improve the media resource recommendation effect.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of information processing, and particularly relates to a media resource recommendation method and device and electronic equipment. BACKGROUND

[0002] With the development of Internet technology, various applications emerge in an endless stream. For example, the emergence of media resource (media asset) applications provides convenience for users to watch media resources. For another example, televisions are becoming more and more intelligent, and can receive media resources provided by various media resource applications for television users to watch.

[0003] However, at present, for television users, the general recommendation method is to recommend hot media resources by using a hot ranking list. However, the recommended hot resources may not be of interest to the user, and the media resource recommendation effect is poor. SUMMARY

[0004] The embodiments of the present application provide a media resource recommendation method, device and electronic equipment to solve the problem of poor media resource recommendation effect.

[0005] To solve the above technical problems, the present application is implemented as follows:

[0006] In a first aspect, the embodiments of the present application provide a media resource recommendation method, and the method comprises the following steps:

[0007] Obtaining triple information of each media resource in a first media resource set, wherein the triple information of the media resource comprises first attribute information of the media resource, second attribute information associated with the media resource, and a relationship between the first attribute information and the second attribute information;

[0008] Inputting the triple information of each media resource in the first media resource set into a trained text generation model to generate resource description content of each media resource in the first media resource set;

[0009] Obtaining resource description content of a second media resource set and image information of a target media resource set, wherein the target media resource set comprises the first media resource set and the second media resource set;

[0010] Determining, based on historical behavior records of a user to be recommended and the resource description content and the image information of the target media resource set, a recommended media resource of the user to be recommended from the target media resource set.

[0011] In a second aspect, the embodiments of the present application provide a media resource recommendation device, and the device comprises:

[0012] The first information acquisition module is configured to acquire the triple information of each media resource in the first media resource set, wherein the triple information of the media resource comprises first attribute information of the media resource, second attribute information associated with the media resource, and a relationship between the first attribute information and the second attribute information.

[0013] The description content generation module is configured to input the triple information of each media resource in the first media resource set into a trained text generation model, and generate resource description content of each media resource in the first media resource set.

[0014] The second information acquisition module is configured to acquire resource description content of media resources in a second media resource set and image information of media resources in a target media resource set, wherein the target media resource set comprises the first media resource set and the second media resource set.

[0015] The recommended resource determination module is configured to determine, based on historical behavior records of a user to be recommended and the resource description content and the image information of the media resources in the target media resource set, recommended media resources of the user to be recommended from the target media resource set.

[0016] In a third aspect, an embodiment of the present application provides an electronic device, comprising a transceiver and a processor,

[0017] The processor is configured to:

[0018] acquire triple information of each media resource in a first media resource set, wherein the triple information of the media resource comprises first attribute information of the media resource, second attribute information associated with the media resource, and a relationship between the first attribute information and the second attribute information;

[0019] input the triple information of each media resource in the first media resource set into a trained text generation model, and generate resource description content of each media resource in the first media resource set;

[0020] acquire resource description content of media resources in a second media resource set and image information of media resources in a target media resource set, wherein the target media resource set comprises the first media resource set and the second media resource set;

[0021] determine, based on historical behavior records of a user to be recommended and the resource description content and the image information of the media resources in the target media resource set, recommended media resources of the user to be recommended from the target media resource set.

[0022] In a fourth aspect, an electronic device is provided, which includes a processor, a memory, and a program stored in the memory and capable of running on the processor. When the program is executed by the processor, the steps of the media resource recommendation method in the first aspect are implemented.

[0023] In a fifth aspect, a computer readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the media resource recommendation method in the first aspect are implemented.

[0024] In the recommendation method of the present embodiment, instead of using the hot spot recommendation method to recommend hot spot resources to the user, the triple information of each media resource in the first media resource set is used to generate the resource description content of each media resource by the text generation model. On the basis of the historical behavior record of the user to be recommended, the resource description content of the second media resource set and the image information of the target media resource set are fused to determine the recommended media resource of the user to be recommended. That is, in the recommendation method of the present embodiment, the association between the recommended resource and the user can be improved through the historical behavior record of the user. Moreover, on the basis of the historical behavior record of the user, the resource description content and the image information of the media resource are combined to make the recommended media resource more suitable for the user, thereby improving the media resource recommendation effect. BRIEF DESCRIPTION OF DRAWINGS

[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the description of the embodiments of the present application will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0026] Figure 1 is one of the flowcharts of the media resource recommendation method provided by the embodiments of the present application;

[0027] Figure 2 is the second flowchart of the media resource recommendation method provided by the embodiments of the present application;

[0028] Figure 3 is a structure diagram of an ItemSage model;

[0029] Figure 4 is a structure diagram of the media resource recommendation device provided by the embodiments of the present application;

[0030] Figure 5 is a structure diagram of the electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION

[0031] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.

[0032] Referring to Figure 1 , Figure 1 is a flowchart of a media resource recommendation method provided by an embodiment of the present application, which can be applied to an electronic device for execution, where the electronic device can be a terminal (which can be a mobile device or a non-mobile device) or a server device. As shown in Figure 1 , the media resource recommendation method provided by the embodiment includes the following steps:

[0033] Step 101: Obtain the triple information of each media resource in the first set of media resources, where the triple information of the media resource includes the first attribute information of the media resource, the second attribute information associated with the media resource, and the relationship between the first attribute information and the second attribute information.

[0034] It should be noted that the media resource in the embodiment of the present application can be an image, a video, etc. The first attribute information can include but is not limited to a resource name, etc. The second attribute information can include but is not limited to a resource-associated user name (for example, an actor name of a movie video) or a resource belonging to a resource category (which can also be understood as a resource tag or a resource type to which the resource belongs, for example, for a video resource, the resource category can include but is not limited to a comedy category, a thriller category, a romance category, an animation category, and an action category, etc.).

[0035] Step 102: Input the triple information of each media resource in the media resource into a trained text generation model to generate the resource description content of each media resource in the first set of media resources.

[0036] It should be noted that the text generation model used in the embodiment of the present application is not specifically limited, for example, in one example, the text generation model can be but is not limited to a language model, for example, a generative pre-trained transformer (Transformer) model 3 (GTP3) can be used. In another example, the first set of media resources includes media resources without / missing media resource descriptions, in this embodiment, the text generation model can be used to generate the resource description content corresponding to the media resource based on the triple information of the media resource in the first set of media resources.

[0037] Step 103: obtaining resource description content of media resources in the second media resource set and image information of media resources in the target media resource set, wherein the target media resource set includes the first media resource set and the second media resource set.

[0038] It can be understood that the media resource provider (content provider) provides the resource description content of the media resources in the second media resource set in advance, that is, the resource description content of the media resources in the second media resource set is the description content provided by the media resource provider. In this embodiment, the resource description content of the media resources in the second media resource set can be directly obtained. In addition, in addition to the resource description content of the first media resource set and the second media resource set, image information of the first media resource set and the second media resource set also needs to be obtained for resource recommendation. In an example, the image information of the media resource can include but is not limited to a poster of the media resource, an image frame in the media resource, etc.

[0039] Step 104: determining the recommended media resource of the user to be recommended from the target media resource set based on the historical behavior record of the user to be recommended and the resource description content and the image information of the media resources in the target media resource set.

[0040] It should be noted that the historical behavior record can be understood as a media asset historical behavior record, including but not limited to at least one of a media asset browsing record, a media asset collection record, a media asset watching record, a media asset like record, and a media asset comment record.

[0041] In this embodiment, the resource description content of each media resource in the first media resource set can be generated by using the triple information of each media resource in the first media resource set through a text generation model. Based on the historical behavior record of the user to be recommended, the resource description content of the target media resource set and the image information of the target media resource set can be fused to determine the recommended media resource of the user to be recommended, so as to improve the media resource recommendation effect.

[0042] In one embodiment, the recommended media resource of the user to be recommended is determined from the target media resource set based on the historical behavior record of the user to be recommended and the resource description content and the image information of the media resources in the target media resource set, including:

[0043] obtaining a first semantic feature vector of the resource description content of each media resource in the target media resource set;

[0044] obtaining a second semantic feature vector of the image information of each media resource in the target media resource set;

[0045] obtain a fusion semantic feature vector of the media resources in the target media resource set based on the first semantic feature vector of the resource description content of the media resources in the target media resource set and the second semantic feature vector of the image information of the media resources in the target media resource set;

[0046] determine a historical behavior resource set of the user to be recommended from the historical behavior record of the user to be recommended, and obtain a third semantic feature vector of each behavior resource in the historical behavior resource set, the behavior resource in the historical behavior resource set being a media resource having a behavior record in the historical behavior record of the user to be recommended;

[0047] determine the recommended media resource of the user to be recommended based on the fusion semantic feature vector of the media resources in the target media resource set and the third semantic feature vector of the behavior resource in the historical behavior resource set.

[0048] That is, the resource description content of the media resource can be subjected to semantic analysis to extract the first semantic feature vector of the resource description content, and the resource description content of each media resource in the target media resource set has a corresponding first semantic feature vector, and the image information of the media resource can also be subjected to semantic analysis to extract the second semantic feature vector of the image information of the media resource, and the image information of each media resource in the target media resource set has a corresponding second semantic feature vector. Then, the first semantic feature vector of the resource description content of the media resources in the target media resource set and the second semantic feature vector of the image information of the media resources in the target media resource set are fused to obtain the fusion semantic feature vector of the target media resource set, that is, the fusion semantic feature vector of each media resource in the target media resource set. It should be noted that in this embodiment, the fusion semantic feature vector of the media resource can be obtained by adding or splicing the first semantic feature vector of the resource description content of the media resource and the second semantic feature vector of the image information of the media resource. In addition, the fusion semantic feature vector of the media resource can also be obtained by inputting the first semantic feature vector of the resource description content of the media resource and the second semantic feature vector of the image information of the media resource into a trained model for feature fusion.

[0049] The historical behavior resource set of the user to be recommended can be understood as a resource on which the user to be recommended has performed historical behavior, that is, the behavior resource in the historical behavior resource set is a media resource having a behavior record in the historical behavior record of the user to be recommended, for example, at least one of a resource on which the user to be recommended has performed browsing behavior, a resource on which the user to be recommended has performed collection behavior, a resource on which the user to be recommended has performed watching behavior, a resource on which the user to be recommended has performed like behavior, and a resource on which the user to be recommended has performed comment behavior.

[0050] After the set of historical behavior resources of the user to be recommended is determined based on the historical behavior records of the user to be recommended, a third semantic feature vector of each behavior resource in the set of historical behavior resources can be extracted. For example, if the behavior resource has media description content, semantic analysis can be performed on the media description content of the behavior resource to determine the third semantic feature vector of the behavior resource. Alternatively, semantic analysis can be performed on the image information of the behavior resource to determine the third semantic feature vector of the behavior resource.

[0051] After the third semantic feature vector of each behavior resource in the set of historical behavior resources is obtained, the set of target media resources can be used to determine the recommended media resources of the user to be recommended by using the fusion semantic feature vector of the set of target media resources and the third semantic feature vector of the set of historical behavior resources.

[0052] That is, in this embodiment, the first semantic feature vector of the resource description content and the second semantic feature vector of the image information of each media resource in the set of target media resources are determined through semantic analysis. The first semantic feature vector and the second semantic feature vector of each media resource in the set of target media resources can be fused to obtain the fusion semantic feature vector of the resource description content of each media resource in the set of target media resources. Then, the set of target media resources can be used to determine the recommended media resources of the user to be recommended by using the fusion semantic feature vector of the set of target media resources and the third semantic feature vector of the set of historical behavior resources, so as to improve the effect of recommending media resources for the user to be recommended.

[0053] In one embodiment, the set of target media resources includes M media resources, and the set of historical behavior resources includes K behavior resources, M is an integer greater than 1, and K is a positive integer.

[0054] Based on the fusion semantic feature vector of the media resources in the set of target media resources and the third semantic feature vector of the behavior resources in the set of historical behavior resources, the recommended media resources of the user to be recommended are determined, including:

[0055] For each behavior resource in the set of historical behavior resources, the similarity between the third semantic feature vector of the behavior resource and the fusion semantic feature vector of each media resource in the set of target media resources is calculated respectively to obtain M similarities of the behavior resource.

[0056] The media resources corresponding to the first N similarities in the set of similarities are determined as the recommended media resources. The set of similarities includes K×M similarities, the first N similarities are greater than the remaining similarities, the remaining similarities are the similarities other than the first N similarities in the set of similarities, and N is a positive integer.

[0057] It should be noted that there are various ways to calculate the similarity, and the embodiments of the present application do not limit the way of calculating the similarity. For example, the cosine distance between the third semantic feature vector of the behavior resource and the fusion feature vector of the media resource can be calculated, and the cosine distance is taken as the similarity between the third semantic feature vector of the behavior resource and the fusion semantic feature vector of the media resource. For each historical resource, the similarity between the third semantic feature vector thereof and the fusion semantic feature vector of each media resource in each target media resource set is calculated. Thus, for each behavior resource, M corresponding similarities can be obtained, and there are K behavior resources, so that KxM similarities can be obtained. The media resources corresponding to the first N similarities in the KxM similarities can be taken as the recommended media resources. In this way, the adaptation degree between the recommended media resources and the historical behaviors of the user to be recommended can be improved, and the recommendation effect can be improved.

[0058] In one embodiment, the fusion semantic feature vector of the media resource in the target media resource set is obtained based on the first semantic feature vector of the resource description content of the media resource in the target media resource set and the second semantic feature vector of the image information of the media resource in the target media resource set, and the fusion semantic feature vector of the media resource in the target media resource set is obtained based on the first semantic feature vector of the resource description content of the media resource in the target media resource set and the second semantic feature vector of the image information of the media resource in the target media resource set.

[0059] The first semantic feature vector of the resource description content of the media resource in the target media resource set and the second semantic feature vector of the image information of the media resource in the target media resource set are input into the trained feature fusion model, and the feature fusion is performed through the trained feature fusion model to obtain the fusion semantic feature vector of the media resource in the target media resource set.

[0060] That is, in the present embodiment, for each media resource in the target media resource set, the first semantic feature vector of the resource description content of the media resource and the second semantic feature vector of the image information of the media resource can be fused through the trained feature fusion model to obtain the fusion semantic feature vector of the media resource. Thus, the fusion semantic feature vector of each media resource in the target media resource set can be obtained. In one example, the trained feature fusion model can adopt a trained ItemSage (product embedding model) model.

[0061] In one embodiment, the trained feature fusion model is trained in the following manner:

[0062] The first training sample set and the channel category to which the media resource sample in the first training sample set belongs are obtained, and the first training sample set includes the semantic feature vector of the resource description content of a plurality of first sample media resources and the semantic feature vector of the image information of the plurality of first sample media resources.

[0063] The initial feature fusion model is trained based on the first training sample set and a channel category to which the first training sample set belongs, to obtain a trained feature fusion model.

[0064] It should be noted that the channel category can include but is not limited to a children's channel, a movie channel, a TV series channel, and a variety show channel, etc. The channel category to which the media resource sample in the first training sample set belongs can be understood as including the channel category to which the media resource sample in the plurality of first sample media resources belongs. In the training process of the feature fusion model, the initial feature fusion model is trained by using the first training sample set and the channel category to which the first training sample set belongs. The first training sample set not only includes the semantic feature vector of the resource description content of the plurality of first sample media resources, but also includes the semantic feature vector of the image information of the plurality of first sample media resources, so as to improve the model training effect and improve the performance of the trained model.

[0065] In one embodiment, the loss function used in the training process of the initial feature fusion model is related to the first function, the channel category label of the media resource sample in the first training sample set, and the predicted channel category label. The predicted channel category label is the channel category label obtained by the initial feature fusion model in the category prediction of the media resource sample in the first training sample set.

[0066] The product of the first function and the product of the first step function and the second step function is inversely related, and the first function is also inversely related to the product of the third step function and the fourth step function.

[0067] The first step function is a step function about a first parameter, and the first parameter is the difference between the channel category label of the first media resource sample and the preset probability threshold. The second step function is a step function about a second parameter, and the second parameter is the difference between the predicted channel category label of the first media resource sample and the preset probability threshold. The third step function is a step function about a third parameter, and the third parameter is the difference between 1 and a first sum value, and the first sum value is the sum of the channel category label of the first media resource sample and the preset probability threshold. The fourth step function is a step function about a fourth parameter, and the fourth parameter is the difference between 1 and a second sum value, and the second sum value is the sum of the predicted channel category label of the first media resource sample and the preset probability threshold. The first media resource sample is a resource in the plurality of first sample media resources.

[0068] In one embodiment, the step function about the target parameter has a first value when the target parameter is greater than zero, a second value when the target parameter is zero, and a third value when the target parameter is less than zero, wherein the target parameter includes any one of the first parameter, the second parameter, the third parameter, and the fourth parameter.

[0069] In one embodiment, the trained text generation model is trained by the following way:

[0070] Obtaining the triple information of the plurality of second media resource samples and the resource description content of the plurality of second media resource samples;

[0071] Fusing the triple information of the plurality of second media resource samples to generate a media resource knowledge graph;

[0072] Training the text generation model by using the media resource knowledge graph and the resource description content of the plurality of second media resource samples to obtain the trained text generation model.

[0073] It should be noted that the triple information of the second media resource sample includes the first attribute information of the second media resource sample, the second attribute information associated with the second media resource sample, and the relationship between the first attribute information of the second media resource sample and the second attribute information of the second media resource sample. By fusing the triple information of the plurality of second media resource samples to generate a media resource knowledge graph, and training the text generation model by using the media resource knowledge graph and the resource description content of the plurality of second media resource samples to obtain the trained text generation model, the accuracy of generating the resource description content can be improved by using the trained text generation model obtained by the above training process to generate the resource description content of each media resource in the first media resource set based on the triple information of each media resource in the first media resource set.

[0074] In one embodiment, after inputting the triple information of each media resource in the first media resource set into the trained text generation model to generate the resource description content of each media resource in the first media resource set, the method further includes:

[0075] Verifying the resource description content of each media resource in the first media resource set to obtain a first probability of the resource description content of each media resource in the first media resource set, wherein the first probability of the resource description content of the media resource is used to represent the integrity of the resource description content of the media resource;

[0076] In a case where the first probability of the resource description content of the media resource is greater than a preset threshold, the resource description content of the media resource generated by the trained text generation model is reserved.

[0077] In a case where the first probability of the resource description content of the media resource is less than or equal to the preset threshold, information is filled in the preset resource description template according to the attribute information of the media resource, and the content obtained after the information is filled in the preset resource description template is determined as the resource description content of the media resource.

[0078] It should be noted that there are various verification methods, which are not specifically limited in the embodiment. For example, the resource description content of each media resource in the first media resource set can be verified by a pointer network. In the embodiment, the resource description content of each media resource in the first media resource set generated by the trained text generation model can be verified. In a case where the first probability of the resource description content of the media resource is greater than a preset threshold, the resource description content of the media resource generated by the trained text generation model is reserved, and in a case where the first probability of the resource description content of the media resource is less than or equal to the preset threshold, information can be filled in the preset resource description template according to the attribute information of the media resource, and the content obtained after the information is filled in the preset resource description template is determined as the resource description content of the media resource, so as to improve the accuracy of the resource description content of the media resource.

[0079] The process of the above method will be specifically described in the following specific embodiments.

[0080] The method of the embodiment of the application can solve the problem that the content-based recommendation for a cold start user does not fully utilize the text and poster semantic information of media assets in the current home television large screen recommendation scene. The application proposes a recommendation algorithm that fuses media asset text and poster information, improves the accuracy of personalized recommendation in this field, and at the same time, in order to solve the problem of missing media asset description, the application proposes a text generation model based on a knowledge graph to fill in the media asset description, so as to prevent the model bias problem caused by the missing media asset description. Combined with the filled media asset description and image (for example, poster) semantic information, the pre-training model fusion feature fusion model (ItemSage framework) is used to obtain the fusion semantic feature vector (embedding) of the media asset, so as to perform media asset recommendation, improve the recommendation accuracy, and improve the user experience.

[0081] As shown in Figure 2 The overall process of the embodiment of the application is as follows:

[0082] Step 201: According to the attribute information of the media asset such as the name, director, actor, label and channel, a triple knowledge graph is constructed, and a text generation model (for example, GPT3) is used to fill in the resource description content.

[0083] Step 202: Obtain a first semantic feature vector of the resource description content using a pre-trained model;

[0084] Step 203: Obtain a second semantic feature vector of the media poster using a pre-trained model;

[0085] Step 204: According to the feature fusion model, construct a classification task, and obtain a fusion semantic feature vector of the media;

[0086] Step 205: According to the fusion semantic feature vector of the media generated by the feature fusion model, use a similarity calculation formula (for example, but not limited to, cosine similarity, etc.) to calculate the similarity between the third semantic feature vector of the behavior resource in the historical behavior resource set and the fusion semantic feature vector of each media resource in the target media resource set, and select the top N media resources (TopN) with the highest similarity for recommendation.

[0087] The specific implementation is as follows:

[0088] Step 31: According to the name, director, actor, label, channel, and other attribute information of the media, construct a triple knowledge graph, and fill the resource description content based on a GPT3 text generation model;

[0089] Step 3101: According to the name, director, actor, label, channel, and other basic information of the media, construct triple structures such as: media name-director-director name, media name-actor-actor name, and media name-label-label name, construct a media knowledge graph through knowledge fusion, in the form of head entity-relation-tail entity (for example: movie A-director-actor a)

[0090] Step 3102: Input the triple information of the media knowledge graph into the text generation model (for example, GPT3), fine-tune the text generation model, and obtain a trained text generation model;

[0091] Step 3103: For media whose description is incomplete, input its related head node, tail node, and relationship into the trained text generation model to generate the resource description content;

[0092] Step 3104: Use the pointer network to check whether the generated resource description content covers the key information of the media;

[0093] Step 32: Obtain a first semantic feature vector of the media description using a pre-trained model;

[0094] Step 3201: Input the resource description content, and obtain its first semantic feature vector using the R2D2 multi-modal pre-trained model;

[0095] Step 33: Obtain the second semantic feature vector of the media poster using the pre-trained model;

[0096] Step 3301: Input the media poster and obtain its second semantic feature vector using the R2D2 multi-modal pre-trained model;

[0097] Step 34: According to the improved ItemSage framework, build a classification task and obtain the fusion semantic feature vector of the media;

[0098] As shown in Figure 3 ItemSage is a recommendation framework for obtaining product embedding information, which includes embedding layer, linear mapping layer (Linear), encoding layer (transformer Encoder), and output layer from bottom to top. The embedding layer includes product image information embedding layer (PinSAGE) and product feature information embedding layer (HashEmbedder). The encoding layer uses transformer network structure for feature vector extraction. According to the classification task, the final output is the semantic feature vector of the product. After the encoding layer extracts the feature vector, it is normalized by layer normalization (LayerNorm) and linear mapping (Linear). The obtained feature vector is then passed through the activation function (GELU) and linear mapping (Linear) to output the fusion semantic feature vector (Itemsage embedding);

[0099] Step 3401: Improve the loss function of the transformer in the Itemsage framework and build a classification task;

[0100] A channel classification task is built with movie, TV series, sports, and variety shows as labels. The input is the first semantic feature vector of the resource description content of the media and the second semantic feature vector of the poster information of the media, and the output is the channel of the media. In the related Itemsage framework, the loss function of the transformer network structure is the cross-entropy loss function, and the formula is:

[0101] Loss=-∑ y y true logy pred ;

[0102] Where y true represents the actual label (channel category) of the sample y, and y pred represents the label of the sample y predicted by the model.

[0103] In the task of media classification, on the one hand, the distinction is not great for the part of media belonging to the cartoon channel or the children's channel; on the other hand, it is easy to make classification errors for the children's film belonging to the children's channel or the film channel, or belonging to both the children's channel and the film channel. For such "ambiguous" samples, a probability threshold (for example, 0.6, in principle only greater than 0.5 can be required) can be set in advance, if the probability of the positive sample output by the model is greater than 0.6, the model is not updated, if the probability of the negative sample output by the model is less than 0.4, the model is also not updated, if the probability value of the sample output by the model is between 0.4 and 0.6, the model is updated, which can ensure higher classification accuracy, and the output of the fusion semantic feature vector of the media is more accurate. The unit step function θ(x) is introduced:

[0104]

[0105] The new loss function used in the embodiment of the application is:

[0106] Loss new =-∑ y∈Y α(y true ,y pred )y true logy pred ;

[0107] The first function is a function about the channel category label to which the media resource sample in the first training sample set belongs and the predicted channel category label. In the above formula, Y is the first training sample set, α(y true , y pred ) is the first function with parameters y true and y pred , y true represents the actual label (channel category label) of the sample y in the first training sample set, and y pred represents the label of the sample y in the first training sample set predicted by the model, wherein α(y true , y pred )=1-θ(y true -m)θ(y pred -m)-θ(1-m-y true )θ(1-m-y pred ), m represents a preset probability threshold, and the preset probability threshold is a value greater than 0 and less than or equal to 1. It can be understood that α(y true , y pred ) can be defined as a correction term, and the loss function of the embodiment of the application is to increase the correction term, that is, the first function, on the basis of the original loss function. If the sample y is a positive sample, y true =1, then:

[0108] α(1, ypred ) = 1 - 0(y pred -m) ;

[0109] If y pred > m, then a(l, y pred ) = 0, the cross-entropy of the loss function reaches a minimum, if y pred < m, then a(l, y pred ) = 1, the cross-entropy remains unchanged. That is, the output probability of the positive sample is greater than m, and the output probability of the negative sample is less than 1-m, the model is not updated. The output probability of the positive sample is less than m, and the output probability of the negative sample is greater than 1-m, the model is updated, which can solve the classification of the sample, so as to improve the accuracy of classification and embedding.

[0110] Step 3402: obtaining the fusion semantic feature vector of the media resource:

[0111] According to the second semantic feature vector of the poster information of the media resource obtained by the R2D2 pre-training model and the first semantic feature vector of the resource description content, the fituning is performed on the ItemSage model, the entire ItemSage framework is trained according to the channel label of the media resource, and finally the fusion semantic feature vector fused with the description information and the poster information of the media resource is obtained;

[0112] Step 35: according to the fusion semantic feature vector generated by the ItemSage framework, the similarity is calculated using the similarity calculation formula, and the top N media resources (TopN) with larger similarity are selected for recommendation:

[0113] Step 3501: according to the fusion semantic feature vector of the media resource fused with the description information and the poster information of the media resource generated by the ItemSage framework, the faiss (an open source library for clustering and similarity search, which provides high-efficiency similarity search and clustering for dense vectors, can support search of billions of vectors, and is an approximate nearest neighbor search library) can be used to construct an index, and the normalized cosine similarity is used to calculate the similarity between the media resources, and the most similar N media resources of each media resource are found out for recommendation, N is a positive integer.

[0114] In view of the problem that the content-based recommendation of the cold start user in the existing family TV large screen recommendation does not fully utilize the text and poster semantic information of the media resource. The application proposes a recommendation algorithm fusing media text and poster information, which can well obtain the fusion semantic feature vector of the media resource based on the improved ItemSage fusion pre-training model, and the vector can fully express the deeper semantic information of the media resource, thereby improving the accuracy of the family TV personalized recommendation.

[0115] In the existing large-screen TV recommendation, there is a problem of missing or large difference of media content information (director, actor, score, description) from multiple content providers (CP). The present application proposes a text generation model based on a knowledge graph to fill in the media description to prevent model bias problems caused by missing media description.

[0116] The present application can be applied to the recommendation scene of large-screen TV. At present, large-screen TV covers almost all families, and the recommendation effect of large-screen TV directly affects the income. The present application will be better than the existing recommendation scheme in recommendation effect and user experience, and will bring good effect to market promotion and application expansion.

[0117] Reference Figure 4 , Figure 4 is a structural schematic diagram of an electronic device of a media resource recommendation apparatus 400 provided by an embodiment of the present application, as shown in Figure 4 The media resource recommendation apparatus 400 comprises:

[0118] A first information acquisition module 401 is configured to acquire triple information of each media resource in a first media resource set, wherein the triple information of the media resource comprises first attribute information of the media resource, second attribute information associated with the media resource, and a relationship between the first attribute information and the second attribute information.

[0119] A description content generation module 402 is configured to input the triple information of each media resource in the first media resource set into a trained text generation model to generate resource description content of each media resource in the first media resource set.

[0120] A second information acquisition module 403 is configured to acquire resource description content of media resources in a second media resource set and image information of media resources in a target media resource set, wherein the target media resource set comprises the first media resource set and the second media resource set.

[0121] A recommended resource determination module 404 is configured to determine recommended media resources of a user to be recommended from the target media resource set based on historical behavior records of the user to be recommended and the resource description content and the image information of the media resources in the target media resource set.

[0122] The media resource recommendation apparatus 400 provided by the embodiment can realize each process of each embodiment of the above-mentioned media resource recommendation method, and the technical features are one-to-one correspondence and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0123] The embodiment of the present application further provides an electronic device, comprising a processor, a memory and a program stored in the memory and executable on the processor, when the program is executed by the processor, each process of the media resource recommendation method embodiment is realized, and the same technical effects can be achieved, to avoid repetition, which will not be repeated here.

[0124] Specifically, referring to Figure 5 The embodiment of the present application further provides an electronic device, comprising a bus 501, a transceiver 502, an antenna 503, a bus interface 504, a processor 505 and a memory 506.

[0125] The processor 505 is configured to:

[0126] Obtain the triple information of each media resource in the first media resource set, wherein the triple information of the media resource comprises first attribute information of the media resource, second attribute information associated with the media resource and a relationship between the first attribute information and the second attribute information;

[0127] Input the triple information of each media resource in the first media resource set into the trained text generation model to generate resource description content of each media resource in the first media resource set;

[0128] Obtain the resource description content of the media resource in the second media resource set and the image information of the media resource in the target media resource set, wherein the target media resource set comprises the first media resource set and the second media resource set;

[0129] Based on the historical behavior record of the user to be recommended and the resource description content and the image information of the media resource in the target media resource set, determine the recommended media resource of the user to be recommended from the target media resource set.

[0130] The processor 505 in the above electronic device can realize each process of each embodiment of the media resource recommendation method, and the technical features are one-to-one correspondence, and the same technical effects can be achieved, to avoid repetition, which will not be repeated here.

[0131] In Figure 5In particular embodiments, bus architecture (represented by bus 501) can include any number of interconnecting buses and bridges, and the bus 501 can link various circuitry from the one or more processors represented by processor 505, and the memory represented by memory 506. The bus 501 can also link various other circuitry, such as peripheral devices, voltage regulators, and power management circuitry, all of which are well known in the art, and therefore, not further described herein. Bus interface 504 provides an interface between bus 501 and transceiver 502. Transceiver 502, which can be a single element or multiple elements such as multiple receivers and transmitters, provides a means for communicating with various other apparatus over a transmission medium. Antenna 503 transmits data processed by processor 505 over a wireless medium, and further, antenna 503 receives data and communicates the data to processor 505.

[0132] Processor 505 is responsible for managing the bus 501 and general processing, and can also provide various functions including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory 506 can be used to store data used by processor 505 during execution of program instructions.

[0133] Optionally, processor 505 can be a CPU, ASIC, FPGA, or CPLD.

[0134] The embodiments of the present application also provide a computer readable storage medium, and the computer readable storage medium stores a computer program. The computer program is executed by a processor to implement various processes of the media resource recommendation method embodiments, and achieve the same technical effects. To avoid repetition, details are not described herein. The computer readable storage medium can be, for example, a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk or an optical disk, etc.

[0135] It should be noted that, in this document, the terms "comprising", "including", or any other variant thereof are intended to cover a non-exclusive inclusion, so that processes, methods, articles, or apparatuses that comprise a list of elements not only include those elements, but also include other elements that are not expressly listed, or other elements inherent in such processes, methods, articles, or apparatuses. Without more limitations, an element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0136] Those skilled in the art can clearly understand the above-mentioned embodiment method can be realized by means of software and necessary general hardware platform, of course, also can be through hardware, but many cases the former is the better embodiment. Based on such understanding, the technical solutions of the present application essentially or say the part of the prior art contribution can be embodied in the form of software product, the computer software product is stored in a storage medium (such as ROM / RAM, magnetic disc, optical disc), including a number of instructions to make a terminal (may be a mobile phone, computer, server, air conditioner, or network equipment, etc.) executes the method of each embodiment of the present application.

[0137] The embodiments of the present application are described above in combination with the drawings, but the present application is not limited to the above-mentioned specific embodiments, the above-mentioned specific embodiments are only illustrative, but not limited, those skilled in the art can make many forms without departing from the purpose of the present application and the scope of the claims under the inspiration of the present application, all belong to the protection of the present application.

Claims

1. A media resource recommendation method, characterized in that, The method comprises: obtaining the triple information of each media resource in the first media resource set, wherein the triple information of the media resource comprises the first attribute information of the media resource, the second attribute information associated with the media resource, and the relationship between the first attribute information and the second attribute information; inputting the triple information of each media resource in the first media resource set into the trained text generation model to generate the resource description content of each media resource in the first media resource set; obtaining the resource description content of the media resources in the second media resource set and the image information of the media resources in the target media resource set, wherein the target media resource set comprises the first media resource set and the second media resource set; determining the recommended media resource of the to-be-recommended user from the target media resource set based on the historical behavior record of the to-be-recommended user and the resource description content and image information of the media resources in the target media resource set; after inputting the triple information of each media resource in the first media resource set into the trained text generation model to generate the resource description content of each media resource in the first media resource set, the method further comprises: verifying the resource description content of each media resource in the first media resource set to obtain the first probability of the resource description content of each media resource in the first media resource set, wherein the first probability of the resource description content of the media resource is used to represent the integrity of the resource description content of the media resource; in the case where the first probability of the resource description content of the media resource is greater than a preset threshold, retaining the resource description content of the media resource generated by the trained text generation model; in the case where the first probability of the resource description content of the media resource is less than or equal to the preset threshold, filling information in a preset resource description template according to the attribute information of the media resource, and determining the content obtained after filling information in the preset resource description template as the resource description content of the media resource.

2. The media resource recommendation method of claim 1, wherein, The method for determining the recommended media resource of the to-be-recommended user from the target media resource set based on the historical behavior record of the to-be-recommended user and the resource description content and image information of the media resources in the target media resource set comprises: obtaining the first semantic feature vector of the resource description content of each media resource in the target media resource set; obtaining the second semantic feature vector of the image information of each media resource in the target media resource set; obtaining the fusion semantic feature vector of the media resources in the target media resource set based on the first semantic feature vector of the resource description content of the media resources in the target media resource set and the second semantic feature vector of the image information of the media resources in the target media resource set; determining the historical behavior resource set of the to-be-recommended user from the historical behavior record of the to-be-recommended user, and obtaining the third semantic feature vector of each behavior resource in the historical behavior resource set, wherein the behavior resources in the historical behavior resource set are media resources having behavior records in the historical behavior record of the to-be-recommended user; Determine the recommended media resource of the user to be recommended based on the fusion semantic feature vector of the media resource in the target media resource set and the third semantic feature vector of the behavior resource in the historical behavior resource set.

3. The media resource recommendation method of claim 2, wherein, The target media resource set includes M media resources, and the historical behavior resource set includes K behavior resources, M is an integer greater than 1, and K is a positive integer. The determination of the recommended media resource of the user to be recommended based on the fusion semantic feature vector of the media resource in the target media resource set and the third semantic feature vector of the behavior resource in the historical behavior resource set includes: For each behavior resource in the historical behavior resource set, the similarity between the third semantic feature vector of the behavior resource and the fusion semantic feature vector of each media resource in the target media resource set is calculated respectively to obtain M similarities of the behavior resource. The media resources corresponding to the first N similarities in the similarity set are determined as the recommended media resources, the similarity set includes K*M similarities, the first N similarities are greater than the remaining similarities, the remaining similarities are the similarities in the similarity set except the first N similarities, and N is a positive integer.

4. The media resource recommendation method of claim 2, wherein, The fusion semantic feature vector of the media resource in the target media resource set is obtained based on the first semantic feature vector of the resource description content of the media resource in the target media resource set and the second semantic feature vector of the image information of the media resource in the target media resource set, and the fusion semantic feature vector of the media resource in the target media resource set is obtained. The first semantic feature vector of the resource description content of the media resource in the target media resource set and the second semantic feature vector of the image information of the media resource in the target media resource set are input into a trained feature fusion model, and the fusion semantic feature vector of the media resource in the target media resource set is obtained by performing feature fusion on the trained feature fusion model.

5. The media resource recommendation method of claim 4, wherein, The trained feature fusion model is obtained by the following method: Obtain a first training sample set and a channel category to which a media resource sample in the first training sample set belongs, the first training sample set including semantic feature vectors of resource description contents of a plurality of first sample media resources and semantic feature vectors of image information of the plurality of first sample media resources; Train an initial feature fusion model based on the first training sample set and the channel category to which the first training sample set belongs to obtain the trained feature fusion model.

6. The media resource recommendation method of claim 5, wherein, The loss function used in the training process of the initial feature fusion model is related to a first function, a channel category label of a media resource sample in the first training sample set, and a predicted channel category label, and the predicted channel category label is a channel category label obtained by the initial feature fusion model predicting the category of the media resource sample in the first training sample set; The product of the first function and the product of the first and second step functions is inversely related, and the product of the first function and the product of the third and fourth step functions is inversely related. The first step function is a step function about a first parameter, the first parameter is a difference between a channel category label to which a first media resource sample belongs and a preset probability threshold; the second step function is a step function about a second parameter, the second parameter is a difference between a predicted channel category label of the first media resource sample and the preset probability threshold; the third step function is a step function about a third parameter, the third parameter is a difference between 1 and a first sum, the first sum is a sum of the channel category label to which the first media resource sample belongs and the preset probability threshold; the fourth step function is a step function about a fourth parameter, the fourth parameter is a difference between 1 and a second sum, the second sum is a sum of the predicted channel category label of the first media resource sample and the preset probability threshold, and the first media resource sample is a resource in the plurality of first sample media resources.

7. The media resource recommendation method of claim 6, wherein, The step function about the target parameter, in the case that the target parameter is greater than zero, the value of the step function about the target parameter is a first numerical value, in the case that the target parameter is zero, the value of the step function about the target parameter is a second numerical value, and in the case that the target parameter is less than zero, the value of the step function about the target parameter is a third numerical value, wherein the target parameter includes any one of the first parameter, the second parameter, the third parameter and the fourth parameter.

8. The media resource recommendation method of claim 1, wherein, The trained text generation model is obtained by the following manner: obtaining triple information of a plurality of second media resource samples and resource description content of the plurality of second media resource samples; fusing the triple information of the plurality of second media resource samples to generate a media resource knowledge graph; using the media resource knowledge graph and the resource description content of the plurality of second media resource samples to train a text generation model to obtain the trained text generation model.

9. A media resource recommendation apparatus, characterized by comprising: The device comprises: a first information acquisition module configured to acquire triple information of each media resource in a first media resource set, wherein the triple information of the media resource comprises first attribute information of the media resource, second attribute information associated with the media resource, and a relationship between the first attribute information and the second attribute information; a description content generation module configured to input the triple information of each media resource in the first media resource set into a trained text generation model to generate resource description content of each media resource in the first media resource set; a second information acquisition module configured to acquire resource description content of media resources in a second media resource set and image information of media resources in a target media resource set, wherein the target media resource set comprises the first media resource set and the second media resource set; a recommended resource determination module configured to determine recommended media resources of a to-be-recommended user from the target media resource set based on historical behavior records of the to-be-recommended user and the resource description content and the image information of the media resources in the target media resource set; wherein the device is further configured to: verify the resource description content of each media resource in the first media resource set to obtain a first probability of the resource description content of each media resource in the first media resource set, the first probability of the resource description content of the media resource being used to represent integrity of the resource description content of the media resource; in a case where the first probability of the resource description content of the media resource is greater than a preset threshold, retaining the resource description content of the media resource generated by the trained text generation model; in a case where the first probability of the resource description content of the media resource is less than or equal to the preset threshold, performing information filling on a preset resource description template according to attribute information of the media resource, and determining content obtained after the information filling on the preset resource description template as the resource description content of the media resource.

10. An electronic device, comprising: comprising a transceiver and a processor, the processor is configured to: obtain triple information of each media resource in a first media resource set, wherein the triple information of the media resource comprises first attribute information of the media resource, second attribute information associated with the media resource, and a relationship between the first attribute information and the second attribute information; input the triple information of each media resource in the first media resource set into a trained text generation model to generate resource description content of each media resource in the first media resource set; obtain resource description content of media resources in a second media resource set and image information of media resources in a target media resource set, wherein the target media resource set comprises the first media resource set and the second media resource set; determine, based on historical behavior records of a user to be recommended and the resource description content and the image information of the media resources in the target media resource set, a recommended media resource of the user to be recommended from the target media resource set; the processor is further configured to: verify the resource description content of each media resource in the first media resource set to obtain a first probability of the resource description content of each media resource in the first media resource set, the first probability of the resource description content of the media resource being used to represent integrity of the resource description content of the media resource; in a case where the first probability of the resource description content of the media resource is greater than a preset threshold, retaining the resource description content of the media resource generated by the trained text generation model; in a case where the first probability of the resource description content of the media resource is less than or equal to the preset threshold, performing information filling on a preset resource description template according to attribute information of the media resource, and determining content obtained after the information filling on the preset resource description template as the resource description content of the media resource.

11. An electronic device, comprising: comprise: a processor, a memory, and a program stored on the memory and executable on the processor, the program being executed by the processor to implement the steps of the method of any one of claims 1 to 8.

12. A computer readable storage medium having stored thereon a computer program which, when executed by a processor, implements the steps of the method of any one of claims 1-8.

Citation Information

Patent Citations

  • Video recommendation method and device, electronic equipment and storage medium

    CN112818251A

  • Video labeling method and device thereof, equipment, medium and product

    CN115359402A