Multimedia resource pushing method and device

By learning user resource interaction information through intelligent agent networks and combining it with diversity analysis, the problems of high repetition and poor accuracy in multimedia resource push systems are solved, thereby improving user retention rate and the performance of the resource push platform.

CN115186173BActive Publication Date: 2026-01-02BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210619242.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-01
Publication Date
2026-01-02
Estimated Expiration
2042-06-01

AI Technical Summary

Technical Problem

Existing multimedia resource push systems suffer from problems such as high repetition in resource pushes, poor accuracy, and declining user retention rates.

Method used

By acquiring the target resource interaction information of the target object and the attribute information of the multimedia resources to be pushed, resource interaction learning is carried out using intelligent agent networks. Combined with diversity analysis, multimedia resources are identified and pushed, including interest intelligent agent networks, hotspot intelligent agent networks and exploration intelligent agent networks, and the contribution of users staying on the resource push platform is learned.

Benefits of technology

It improved the accuracy and diversity of multimedia resource delivery, increased user retention, reduced system resource waste, and enhanced the performance of the resource delivery platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115186173B_ABST
    Figure CN115186173B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a multimedia resource pushing method and device, the method comprising: obtaining target resource interaction information of a target object and first resource attribute information of a to-be-pushed multimedia resource; inputting the target resource interaction information and the first resource attribute information into at least one agent network for resource interaction learning, and determining a pushed multimedia resource corresponding to the at least one agent network; the at least one agent network is used for learning a multimedia resource corresponding to at least one interaction reason, and a contribution of the target object to a stay on a resource pushing platform; performing diversity analysis on the pushed multimedia resource to obtain diversity index data; and based on the diversity index data, pushing a target multimedia resource in the pushed multimedia resource to the target object. The embodiments of the present disclosure can improve the diversity of the pushed multimedia resource on the basis of improving the user retention in the platform.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of artificial intelligence, and particularly relates to a multimedia resource pushing method and device. BACKGROUND

[0002] With the rapid development of artificial intelligence and Internet technology, multimedia resource pushing systems such as short videos bring great convenience and pleasure to people's life, can help users quickly filter out irrelevant information, and push possible interesting content to them to better meet user needs.

[0003] In the related art, the multimedia resource pushing is often combined with the correlation between the user attribute information and the resource attribute information of the multimedia resource. However, the user preference is not constant, and only combining the correlation cannot effectively capture the interest preference of the user, and it is easy to repeatedly push the same or similar multimedia resources to the user, and there is a high repetition of the pushed multimedia resources, which further leads to poor resource pushing accuracy and effect in the resource pushing platform, and problems such as a decrease in user retention rate. SUMMARY

[0004] The present disclosure provides a multimedia resource pushing method and device to at least solve the problems of high repetition of the pushed multimedia resources, poor resource pushing accuracy and effect, and a decrease in user retention rate in the related art. The technical solutions of the present disclosure are as follows:

[0005] According to a first aspect of an embodiment of the present disclosure, a multimedia resource pushing method is provided, comprising:

[0006] obtaining target resource interaction information of a target object and first resource attribute information of a to-be-pushed multimedia resource, the target resource interaction information representing at least one interaction operation information of the target object performed on a historical multimedia resource;

[0007] inputting the target resource interaction information and the first resource attribute information into at least one agent network for resource interaction learning to determine a pushed multimedia resource corresponding to the at least one agent network; the at least one agent network is used to learn at least one candidate multimedia resource, and the contribution of the target object to the resource pushing platform, the at least one candidate multimedia resource is a multimedia resource corresponding to at least one interaction reason in the to-be-pushed multimedia resource, and any interaction reason represents the reason for the target object to perform an interaction operation on the corresponding multimedia resource;

[0008] performing diversity analysis on the pushed multimedia resource to obtain diversity index data;

[0009] Based on the diversity index data, push a target multimedia resource in the push multimedia resource to the target object.

[0010] In an optional embodiment, the diversity analysis on the push multimedia resource to obtain the diversity index data comprises:

[0011] Determine the category index data of the push multimedia resource and a preset balance parameter, the preset balance parameter is used to balance the diversity and correlation between multimedia resources in the push multimedia resource;

[0012] Based on the first resource attribute information of the push multimedia resource, generate resource association data corresponding to the push multimedia resource;

[0013] Based on the resource association data, the category index data and the preset balance parameter, the diversity analysis on the push multimedia resource to obtain the diversity index data.

[0014] In an optional embodiment, the determination of the category index data of the push multimedia resource comprises:

[0015] Input the first resource attribute information of the push multimedia resource corresponding to any of the intelligent agent networks into the resource category identification network corresponding to any of the intelligent agent networks for resource category identification processing to obtain the category index data;

[0016] Wherein, the resource category identification network corresponding to any of the intelligent agent networks is obtained by performing resource category identification training on a preset neural network based on the second resource attribute information of the candidate multimedia resource corresponding to any of the intelligent agent networks and the preset category data of the candidate multimedia resource.

[0017] In an optional embodiment, the at least one intelligent agent network comprises at least one of an interest intelligent agent network, a hot spot intelligent agent network and an exploration intelligent agent network;

[0018] The interest intelligent agent network is configured to learn candidate multimedia resources corresponding to interest interaction reasons and contribute to the target object's stay on the resource pushing platform. The hotspot intelligent agent network is configured to learn candidate multimedia resources corresponding to hotspot interaction reasons and contribute to the target object's stay on the resource pushing platform. The exploration intelligent agent network is configured to learn candidate multimedia resources corresponding to preset interaction reasons and contribute to the target object's stay on the resource pushing platform in the case of exploration guidance of the target exploration network on the unpushed multimedia resources. The target exploration network is configured to guide the exploration intelligent agent network to explore the unpushed multimedia resources in the process of learning candidate multimedia resources corresponding to the preset interaction reasons and contributing to the target object's stay on the resource pushing platform. The unpushed multimedia resources are unpushed candidate multimedia resources in the candidate multimedia resources corresponding to the preset interaction reasons.

[0019] In an optional embodiment, the target resource interaction information is a target interaction graph with the object attribute information of the target object and the resource attribute information of the historical multimedia resources as nodes and at least one interaction operation performed by the target object on the historical multimedia resources as edges.

[0020] Any of the intelligent agent networks comprises a graph convolution network, an interaction state encoding network, a stay learning network, and a filtering network corresponding to the interaction reason of any of the intelligent agent networks. The target resource interaction information and the first resource attribute information are input into at least one intelligent agent network for resource interaction learning to determine the pushed multimedia resources corresponding to the at least one intelligent agent network, comprising:

[0021] The graph convolution network is configured to perform feature representation learning on the target interaction graph to obtain first graph feature information.

[0022] The interaction state encoding network is configured to perform interaction state encoding processing on the first graph feature information to obtain first state encoding information.

[0023] The first resource attribute information is input into the filtering network for interaction preference filtering to obtain the at least one candidate multimedia resource.

[0024] The first state encoding information and the candidate multimedia resource are input into the stay learning network for stay contribution learning to obtain first stay index data corresponding to the at least one interaction reason. The first stay index data represents the contribution of the at least one candidate multimedia resource to the target object's stay on the resource pushing platform.

[0025] The pushed multimedia resource is determined from the candidate multimedia resource based on the first stay index data.

[0026] In an optional embodiment, the pushing, to the target object, of the target multimedia resource in the push multimedia resource based on the diversity index data comprises:

[0027] determining, from the push multimedia resource, a target multimedia resource according to the diversity index data;

[0028] pushing, to the target object, the target multimedia resource.

[0029] In an optional embodiment, the method further comprises:

[0030] obtaining sample resource interaction information of a sample object and third resource attribute information of a sample multimedia resource, the sample resource interaction information representing at least one interaction operation information performed by the sample object on historical multimedia resources of the sample object;

[0031] inputting the sample resource interaction information and the third resource attribute information into at least one to-be-trained intelligent agent network for resource interaction learning to obtain second stay index data corresponding to the at least one interaction reason, the second stay index data representing a contribution of a first multimedia resource corresponding to the at least one interaction reason in the sample multimedia resource to a stay of the sample object on the resource pushing platform;

[0032] inputting the second stay index data into a preset center network for global stay contribution learning to obtain global stay index data;

[0033] determining, from the first multimedia resource, a second multimedia resource based on the second stay index data;

[0034] performing resource interaction analysis on the sample object based on third resource attribute information of the second multimedia resource and object attribute information of the sample object to obtain an interaction multimedia resource of the sample object;

[0035] training the at least one to-be-trained intelligent agent network based on the interaction multimedia resource and the global stay index data to obtain the at least one intelligent agent network.

[0036] In an optional embodiment, the sample resource interaction information is a sample interaction graph with object attribute information of the sample object and resource attribute information of the historical multimedia resource of the sample object as nodes, and at least one interaction operation performed by the sample object on the historical multimedia resource of the sample object as edges; any of the to-be-trained intelligent agent networks comprises a to-be-trained graph convolution network, a to-be-trained encoding network, a to-be-trained stay learning network, and a to-be-trained filtering network corresponding to an interaction reason of any of the to-be-trained intelligent agent networks; and the inputting of the sample resource interaction information and the third resource attribute information into at least one to-be-trained intelligent agent network for resource interaction learning to obtain the second stay index data corresponding to the at least one interaction reason comprises:

[0037] performing feature representation learning on the sample interaction graph based on the to-be-trained graph convolution network to obtain second graph feature information;

[0038] performing interaction state encoding processing on the second graph feature information based on the to-be-trained encoding network to obtain second state encoding information;

[0039] inputting the third resource attribute information into the to-be-trained filtering network for interaction preference filtering to obtain the first multimedia resource;

[0040] inputting the second state encoding information and third resource attribute information of the first multimedia resource into the to-be-trained stay learning network for stay contribution learning to obtain the second stay index data.

[0041] In an optional embodiment, in the case where the at least one intelligent agent network comprises an exploration intelligent agent network, the method further comprises:

[0042] inputting the second state encoding information and third resource attribute information of the first multimedia resource into a target exploration network corresponding to the exploration intelligent agent network for stay contribution learning to obtain third stay index data;

[0043] determining fourth stay index data based on the third stay index data and state incentive information, the state incentive information being determined based on an interaction state of the first multimedia resource corresponding to the at least one interaction reason in a training process of the at least one to-be-trained intelligent agent network, the interaction state representing a number of times that the first multimedia resource is determined as the second multimedia resource;

[0044] generating first loss information based on the fourth stay index data and the second stay index data;

[0045] updating network parameters of the to-be-trained stay learning network corresponding to the exploration intelligent agent network based on the first loss information.

[0046] In an optional embodiment, the training of the at least one agent network to be trained based on the interactive multimedia resource and the global dwell indicator data comprises:

[0047] updating the sample resource interaction information based on the interactive multimedia resource;

[0048] generating reward information corresponding to any of the agent networks to be trained based on the interactive operation information corresponding to the interactive multimedia resource;

[0049] determining second loss information corresponding to any of the agent networks to be trained based on the reward information corresponding to any of the agent networks to be trained and second dwell indicator data corresponding to any of the agent networks to be trained;

[0050] updating network parameters of the at least one agent network to be trained based on the second loss information and the global dwell indicator data;

[0051] repeating the inputting of the sample resource interaction information and the third resource attribute information into the at least one agent network to be trained for resource interaction learning to obtain second dwell indicator data corresponding to the at least one interaction reason based on the updated sample resource interaction information and the updated at least one agent network to be trained until the step of updating the network parameters of the at least one agent network to be trained based on the second loss information and the global dwell indicator data, until the reward information meets a preset condition;

[0052] using the at least one agent network to be trained meeting the preset condition as the at least one agent network.

[0053] In an optional embodiment, after the pushing of the target multimedia resource in the push multimedia resource to the target object based on the diversity indicator data, the method further comprises:

[0054] obtaining continuous interaction data of the target object, the continuous interaction data representing an operation quantity of at least one interactive operation continuously performed by the target object on the target multimedia resource;

[0055] determining first measurement indicator data based on the continuous interaction data and preset continuous interaction data, the first measurement indicator data being used to measure a contribution effect of the target multimedia resource on the dwell of the target object on the resource pushing platform.

[0056] In an optional embodiment, the target multimedia resource comprises a plurality of multimedia resources; and the method further comprises:

[0057] obtain fourth resource attribute information corresponding to the plurality of multimedia resources;

[0058] perform resource category identification on the plurality of multimedia resources based on the fourth resource attribute information to obtain target resource categories of the plurality of multimedia resources;

[0059] determine second measurement index data according to a number of resource categories in the target resource categories; the second measurement index data is used to measure diversity of the plurality of multimedia resources.

[0060] According to a second aspect of the embodiment of the present disclosure, a multimedia resource pushing device is provided, comprising:

[0061] a first information obtaining module configured to perform obtaining target resource interaction information of a target object and first resource attribute information of to-be-pushed multimedia resources, the target resource interaction information representing at least one interaction operation information performed by the target object on historical multimedia resources;

[0062] a first resource interaction learning module configured to perform inputting the target resource interaction information and the first resource attribute information into at least one intelligent agent network for resource interaction learning to determine pushed multimedia resources corresponding to the at least one intelligent agent network; the at least one intelligent agent network is used to learn at least one candidate multimedia resource, and the at least one candidate multimedia resource is a multimedia resource corresponding to at least one interaction reason in the to-be-pushed multimedia resources, any interaction reason representing a reason for performing an interaction operation on a corresponding multimedia resource by the target object;

[0063] a diversity analysis module configured to perform diversity analysis on the pushed multimedia resources to obtain diversity index data;

[0064] a resource pushing module configured to perform pushing a target multimedia resource in the pushed multimedia resources to the target object based on the diversity index data.

[0065] In an optional embodiment, the diversity analysis module comprises:

[0066] a data determining unit configured to perform determining category index data of the pushed multimedia resources and a preset balance parameter, the preset balance parameter being used to balance diversity and correlation among multimedia resources in the pushed multimedia resources;

[0067] a resource association data generating unit configured to perform generating resource association data corresponding to the pushed multimedia resources based on the first resource attribute information of the pushed multimedia resources;

[0068] The diversity analysis unit is configured to perform diversity analysis on the push multimedia resources based on the resource association data, the category index data, and the preset balance parameter, to obtain the diversity index data.

[0069] In an optional embodiment, the data determination unit comprises:

[0070] The resource category identification processing unit is configured to perform resource category identification processing on the first resource attribute information of the push multimedia resources corresponding to any of the intelligent agent networks by inputting the first resource attribute information into a resource category identification network corresponding to any of the intelligent agent networks, to obtain the category index data.

[0071] The resource category identification network corresponding to any of the intelligent agent networks is obtained by performing resource category identification training on a preset neural network based on the second resource attribute information of the candidate multimedia resources corresponding to any of the intelligent agent networks and preset category data of the candidate multimedia resources.

[0072] In an optional embodiment, the at least one intelligent agent network comprises at least one of an interest intelligent agent network, a hotspot intelligent agent network, and an exploration intelligent agent network.

[0073] The interest intelligent agent network is configured to learn candidate multimedia resources corresponding to interest interaction causes and contribute to the stay of the target object on the resource push platform; the hotspot intelligent agent network is configured to learn candidate multimedia resources corresponding to hotspot interaction causes and contribute to the stay of the target object on the resource push platform; the exploration intelligent agent network is configured to learn candidate multimedia resources corresponding to preset interaction causes and contribute to the stay of the target object on the resource push platform in a case where a target exploration network guides exploration of unpushed multimedia resources; and the target exploration network is configured to guide the exploration intelligent agent network to explore the unpushed multimedia resources in the process of learning the candidate multimedia resources corresponding to the preset interaction causes and contributing to the stay of the target object on the resource push platform, the unpushed multimedia resources being unpushed candidate multimedia resources corresponding to the preset interaction causes.

[0074] In an optional embodiment, the target resource interaction information is a target interaction graph with object attribute information of the target object and resource attribute information of the historical multimedia resources as nodes and at least one interaction operation performed by the target object on the historical multimedia resources as edges.

[0075] Any of the agent networks comprises a graph convolution network, an interaction state encoding network, a dwell learning network, and a filtering network of an interaction reason corresponding to any of the agent networks; the first resource interaction learning module comprises:

[0076] a first feature representation learning unit configured to perform feature representation learning on the target interaction graph based on the graph convolution network to obtain first graph feature information;

[0077] a first interaction state encoding processing unit configured to perform interaction state encoding processing on the first graph feature information based on the interaction state encoding network to obtain first state encoding information;

[0078] an interaction preference filtering unit configured to perform interaction preference filtering on the first resource attribute information input into the filtering network to obtain the at least one candidate multimedia resource;

[0079] a dwell contribution learning unit configured to perform dwell contribution learning on the first state encoding information and the candidate multimedia resource input into the dwell learning network to obtain first dwell indicator data corresponding to the at least one interaction reason, the first dwell indicator data representing a contribution of the at least one candidate multimedia resource to dwell of the target object on the resource pushing platform;

[0080] a pushed multimedia resource determination unit configured to determine the pushed multimedia resource from the candidate multimedia resource based on the first dwell indicator data.

[0081] In an optional search box, the resource pushing module comprises:

[0082] a target multimedia resource determination unit configured to determine a target multimedia resource from the pushed multimedia resource according to the diversity indicator data;

[0083] a resource pushing unit configured to push the target multimedia resource to the target object.

[0084] In an optional embodiment, the apparatus further comprises:

[0085] a second information acquisition module configured to acquire sample resource interaction information of a sample object and third resource attribute information of a sample multimedia resource, the sample resource interaction information representing at least one interaction operation information performed by the sample object on the historical multimedia resource of the sample object;

[0086] a second resource interaction learning module configured to perform inputting the sample resource interaction information and the third resource attribute information into at least one to-be-trained intelligent agent network for resource interaction learning, to obtain second stay index data corresponding to the at least one interaction reason, the second stay index data representing a first multimedia resource corresponding to the at least one interaction reason in the sample multimedia resource, and a stay contribution of the sample object on the resource pushing platform;

[0087] a global stay contribution learning module configured to perform inputting the second stay index data into a preset center network for global stay contribution learning, to obtain global stay index data;

[0088] a second multimedia resource determination module configured to perform determining a second multimedia resource from the first multimedia resource based on the second stay index data;

[0089] a resource interaction analysis module configured to perform resource interaction analysis on the sample object based on third resource attribute information of the second multimedia resource and object attribute information of the sample object, to obtain an interaction multimedia resource of the sample object;

[0090] a network training module configured to perform training the at least one to-be-trained intelligent agent network based on the interaction multimedia resource and the global stay index data, to obtain the at least one intelligent agent network.

[0091] In an optional embodiment, the sample resource interaction information is a sample interaction graph with object attribute information of the sample object and resource attribute information of historical multimedia resources of the sample object as nodes, and at least one interaction operation performed by the sample object on the historical multimedia resources of the sample object as edges; any to-be-trained intelligent agent network includes a to-be-trained graph convolution network, a to-be-trained encoding network, a to-be-trained stay learning network, and a to-be-trained filtering network corresponding to an interaction reason of any to-be-trained intelligent agent network; and the second resource interaction learning module includes:

[0092] a second feature representation learning unit configured to perform feature representation learning on the sample interaction graph based on the to-be-trained graph convolution network, to obtain second graph feature information;

[0093] a second interaction state encoding processing unit configured to perform interaction state encoding processing on the second graph feature information based on the to-be-trained encoding network, to obtain second state encoding information;

[0094] a second interaction preference filtering unit configured to perform inputting the third resource attribute information into the to-be-trained filtering network for interaction preference filtering, to obtain the first multimedia resource;

[0095] The second dwell contribution learning unit is configured to perform dwell contribution learning by inputting the second state encoding information and third resource attribute information of the first multimedia resource into the to-be-trained dwell learning network to obtain the second dwell indicator data.

[0096] In an optional embodiment, in the case that the at least one agent network comprises an exploration agent network, the apparatus further comprises:

[0097] The dwell contribution learning module is configured to perform dwell contribution learning by inputting the second state encoding information and third resource attribute information of the first multimedia resource into a target exploration network corresponding to the exploration agent network to obtain third dwell indicator data.

[0098] The fourth dwell indicator data determination module is configured to perform determination of fourth dwell indicator data based on the third dwell indicator data and state incentive information, the state incentive information being determined based on an interaction state of the first multimedia resource corresponding to the at least one interaction reason in a training process of the at least one to-be-trained agent network, the interaction state representing a number of times that the first multimedia resource is determined as the second multimedia resource.

[0099] The first loss information generation module is configured to perform generation of first loss information based on the fourth dwell indicator data and the second dwell indicator data.

[0100] The network parameter module is configured to perform update of network parameters of the to-be-trained dwell learning network corresponding to the exploration agent network based on the first loss information.

[0101] In an optional embodiment, the network training module comprises:

[0102] The sample resource interaction information update unit is configured to perform update of the sample resource interaction information based on the interaction multimedia resource.

[0103] The reward information generation unit is configured to perform generation of reward information corresponding to any of the to-be-trained agent networks based on interaction operation information corresponding to the interaction multimedia resource.

[0104] The second loss information determination unit is configured to perform determination of second loss information corresponding to any of the to-be-trained agent networks based on the reward information corresponding to any of the to-be-trained agent networks and the second dwell indicator data corresponding to any of the to-be-trained agent networks.

[0105] The network parameter update unit is configured to perform update of network parameters of the at least one to-be-trained agent network based on the second loss information and the global dwell indicator data.

[0106] The training processing unit is configured to perform the steps of repeatedly inputting the sample resource interaction information and the third resource attribute information into the at least one to-be-trained intelligent agent network for resource interaction learning to obtain second stay index data corresponding to the at least one interaction reason based on the updated sample resource interaction information and the updated at least one to-be-trained intelligent agent network, until the network parameters of the at least one to-be-trained intelligent agent network are updated based on the second loss information and the global stay index data, until the reward information meets a preset condition.

[0107] The intelligent agent network determination unit is configured to execute the at least one to-be-trained intelligent agent network that meets the preset condition as the at least one intelligent agent network.

[0108] In an optional embodiment, the apparatus further comprises:

[0109] The continuous interaction data acquisition module is configured to execute, after the target multimedia resource in the pushed multimedia resource is pushed to the target object based on the diversity index data, acquire continuous interaction data of the target object, the continuous interaction data representing an operation number of at least one interaction operation continuously performed by the target object on the target multimedia resource.

[0110] The first measurement index data determination module is configured to determine first measurement index data based on the continuous interaction data and preset continuous interaction data, the first measurement index data being used to measure a contribution effect of the target multimedia resource on the stay of the target object on the resource pushing platform.

[0111] In an optional embodiment, the target multimedia resource comprises a plurality of multimedia resources; and the apparatus further comprises:

[0112] The fourth resource attribute information acquisition module is configured to execute acquisition of fourth resource attribute information corresponding to the plurality of multimedia resources.

[0113] The resource category identification module is configured to perform resource category identification on the plurality of multimedia resources based on the fourth resource attribute information to obtain a target resource category of the plurality of multimedia resources.

[0114] The second measurement index data determination module is configured to determine second measurement index data according to a number of resource categories in the target resource category; the second measurement index data being used to measure diversity of the plurality of multimedia resources.

[0115] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, comprising: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the method according to any one of the first aspect.

[0116] According to a fourth aspect of the embodiments of the present disclosure, a computer readable storage medium is provided, when instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the method according to any one of the first aspect of the embodiments of the present disclosure.

[0117] According to a fifth aspect of the embodiments of the present disclosure, a computer program product containing instructions is provided, when it is run on a computer, the computer is enabled to perform the method according to any one of the first aspect of the embodiments of the present disclosure.

[0118] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:

[0119] In combination with at least one agent network, candidate multimedia resources corresponding to at least one interaction reason are learned from target resource interaction information of a target object, and the contribution of the target object to resource retention on a resource pushing platform can be combined with the retention contribution to more accurately capture different resource interaction preferences of the user, improve the effectiveness of the determined pushed multimedia resources, and in the case of obtaining at least one pushed multimedia resource corresponding to the interaction reason, the diversity index data can be obtained by analyzing the diversity of the pushed multimedia resources; and in combination with the diversity index data, the target multimedia resource pushed to the target object is filtered from the pushed multimedia resources, which can greatly improve the diversity of the pushed multimedia resources, enrich the pushed multimedia resources, and thus better meet the user demand, improve the resource pushing accuracy and effectiveness, reduce the system resource waste caused by invalid multimedia resource pushing, and greatly improve the system performance of the resource pushing platform.

[0120] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0121] The accompanying drawings incorporated in the specification and forming a part of it, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the disclosure without imposing undue limitation on the disclosure.

[0122] Figure 1 is a schematic diagram of an application environment according to an exemplary embodiment;

[0123] Figure 2 is a flowchart of a multimedia resource pushing method according to an exemplary embodiment;

[0124] Figure 3 is a flowchart of inputting target resource interaction information and first resource attribute information into at least one agent network for resource interaction learning according to an example embodiment;

[0125] Figure 4 is a flowchart of pre-training at least one agent network according to an example embodiment;

[0126] Figure 5 is a flowchart of inputting sample resource interaction information and third resource attribute information into at least one to-be-trained agent network for resource interaction learning to obtain second stay index data corresponding to at least one interaction reason according to an example embodiment;

[0127] Figure 6 is a flowchart of training at least one to-be-trained agent network based on interaction multimedia resources and global stay index data to obtain at least one agent network according to an example embodiment;

[0128] Figure 7 is a structural schematic diagram in a training process of at least one agent network according to an example embodiment;

[0129] Figure 8 is a flowchart of performing diversity analysis on pushed multimedia resources to obtain diversity index data according to an example embodiment;

[0130] Figure 9 is a block diagram of a multimedia resource pushing device according to an example embodiment;

[0131] Figure 10 is a block diagram of an electronic device for multimedia resource pushing according to an example embodiment;

[0132] Figure 11 is a block diagram of an electronic device for multimedia resource pushing according to an example embodiment. DETAILED DESCRIPTION

[0133] In order for those skilled in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings.

[0134] It should be noted that the terms "first", "second", and the like in the description and claims of the present disclosure and the above-described drawings are used to distinguish similar objects, and do not necessarily have to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation described in the following exemplary embodiments does not represent all implementations consistent with the present disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0135] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, analyzed data, etc.) involved in the present disclosure are all information and data authorized by the user or authorized by all parties.

[0136] Please refer to Figure 1 , Figure 1 is a schematic diagram of an application environment according to an exemplary embodiment, which can include a terminal 100 and a server 200.

[0137] In an optional embodiment, the terminal 100 can be used to provide a multimedia resource push service for any user. Specifically, the terminal 100 can include but is not limited to electronic devices such as smart phones, desktop computers, tablet computers, notebook computers, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, smart wearable devices, etc., and can also be software such as an application program running on the above-mentioned electronic devices. Optionally, the operating system running on the electronic device can include but is not limited to Android system, IOS system, Linux, Windows, etc.

[0138] In an optional embodiment, the server 200 can provide background services for the terminal 100. Specifically, the server 200 can be a standalone physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0139] In addition, it should be noted that Figure 1 The application environment shown is only one application environment provided by the present disclosure, and in actual application, it can also include other application environments, for example, it can include more terminals.

[0140] In the embodiments of the present disclosure, the terminal 100 and the server 200 described above can be directly or indirectly connected through wired or wireless communication, which is not limited by the present disclosure.

[0141] Figure 2 Fig. 1 is a flowchart of a multimedia resource pushing method according to an exemplary embodiment, as shown in Fig. 1, the multimedia resource pushing method is used in a terminal electronic device, including the following steps. Figure 2

[0142] In step S201, target resource interaction information of a target object and first resource attribute information of a multimedia resource to be pushed are acquired.

[0143] In one specific embodiment, the target object can be a pushing object of a multimedia resource in a resource pushing platform; specifically, the target object can be a user account in the resource pushing platform. The target resource interaction information can represent at least one interaction operation information of a historical multimedia resource performed by the target object; specifically, the at least one interaction operation information can include information of at least one of an effective browsing operation (a browsing duration greater than a preset duration), a clicking operation, a sharing operation, a like operation, a comment operation, a review operation, and a hot clicking operation (a clicking operation on a hot multimedia resource). The historical multimedia resource can be a multimedia resource on which the target object has performed at least one interaction operation within a preset historical time period. Specifically, the preset historical time period can be set in combination with actual applications, for example, one week, one month, etc. before the current time. Optionally, the multimedia resource to be pushed can be a multimedia resource in the resource pushing platform. The first resource attribute information can be information for describing the multimedia resource to be pushed. Taking a video as an example, the first resource attribute information can include video identification, a release date, a video frame image, audio information, a playing duration, title information, and interaction information (for example, the number of like, sharing, and other interaction operations) that can describe the video.

[0144] ​In a specific embodiment, the target resource interaction information can be a target interaction graph with the object attribute information of the target object and the resource attribute information of the historical multimedia resources as nodes, and at least one interaction operation performed by the target object on the historical multimedia resources as edges. Specifically, the object attribute information of the target object can be attribute information capable of representing the interest preference of the target object, such as user gender, etc. The resource attribute information can be information for describing the multimedia resources. Specifically, taking the like operation as an example, if the target object performs a like operation on multimedia resource 1 in the historical multimedia resources, it can be determined that the length of the edge between the target object and multimedia resource 1 in the like interaction dimension is 1, and correspondingly, in the process of constructing the target interaction graph, an edge between the node where the object attribute information is located and the node where the resource attribute information of multimedia resource 1 is located in the like interaction dimension can be established. Conversely, if the target object performs a share operation but not a like operation on multimedia resource 2 in the historical multimedia resources, it can be determined that the length of the edge between the target object and multimedia resource 1 in the like interaction dimension is 0, and correspondingly, in the process of constructing the target interaction graph, an edge between the node where the object attribute information is located and the node where the resource attribute information of multimedia resource 1 is located in the like interaction dimension does not need to be established.

[0145] In step S203, the target resource interaction information and the first resource attribute information are input into at least one agent network for resource interaction learning to determine the pushed multimedia resources corresponding to the at least one agent network.

[0146] In a specific embodiment, the at least one agent network described above can be used to learn at least one candidate multimedia resource that contributes to the stay of the target object on the resource pushing platform, and the at least one candidate multimedia resource can be a multimedia resource corresponding to at least one interaction reason in the to-be-pushed multimedia resources. The interaction reason can represent the reason why the target object performs an interaction operation on the corresponding multimedia resource. Optionally, the interaction reason can include an interest interaction reason, a hot interaction reason, etc. Optionally, taking the click operation as an example, the multimedia resource corresponding to the interest interaction reason can be a multimedia resource on which the target object will perform a click operation due to the interest in the multimedia resource; and the multimedia resource corresponding to the hot interaction reason can be a multimedia resource on which the target object will perform a click operation due to the fact that the multimedia resource is a hot resource. Specifically, the multimedia resource corresponding to any interaction reason can be obtained by filtering the to-be-pushed multimedia resources based on the filtering network corresponding to the interaction reason. Optionally, the filtering network corresponding to the interest interaction reason can be used to filter out the multimedia resources that are of interest to the target object, and the filtering network corresponding to the hot interaction reason can be used to filter out the multimedia resources that are hot resources. Optionally, the at least one agent network can include at least one of an interest agent network, a hot agent network, and an exploration agent network.

[0147] In one specific embodiment, the above interest intelligent agent network can be used to learn candidate multimedia resources corresponding to interest interaction reasons that contribute to the stay of target objects on the resource pushing platform. Specifically, in combination with the interest intelligent agent network for resource interaction learning, user personalized preferences can be combined to recommend multimedia resources that contribute to user retention in the platform.

[0148] In one specific embodiment, the above hotspot intelligent agent network is used to learn candidate multimedia resources corresponding to hotspot interaction reasons that contribute to the stay of target objects on the resource pushing platform. In actual applications, in addition to regular interactions that meet user personalized preferences (interests), users can also make non-personalized preference interactions at a certain time period, such as interactions with hotspot multimedia resources. Specifically, for example, a user has always been interested in football multimedia resources and not interested in game multimedia resources, but one day, due to the popularity of multimedia resources of a game match, the user will often perform interaction operations on the multimedia resources of the game match out of a herd mentality and the need to satisfy curiosity. Correspondingly, in combination with the hotspot intelligent agent network for resource interaction learning, multimedia resource recommendations that contribute to user retention in the platform can be made from the perspective of hotspot perception.

[0149] In one specific embodiment, the above exploration intelligent agent network is used to learn candidate multimedia resources corresponding to preset interaction reasons that contribute to the stay of target objects on the resource pushing platform in the case of target exploration network guiding exploration of unpushed multimedia resources. Specifically, the target exploration network is used to guide the exploration intelligent agent network to explore unpushed multimedia resources in the process of learning candidate multimedia resources corresponding to preset interaction reasons that contribute to the stay of target objects on the resource pushing platform. The unpushed multimedia resources are unpushed candidate multimedia resources in the candidate multimedia resources corresponding to the preset interaction reasons. Specifically, the preset interaction reasons can include interest interaction reasons, hotspot interaction reasons, etc.

[0150] In actual applications, the data volume of candidate multimedia resources is often large, resulting in a relatively large space that the intelligent agent needs to explore in the process of resource interaction learning, which is prone to interaction exploration deviation. Correspondingly, in combination with the target exploration network, the exploration intelligent agent network can be prompted to explore unknown state space in a certain rule, better improving the exploration accuracy of user resource interaction preferences.

[0151] In the above embodiments, in combination with at least one of the interest agent network, the hotspot agent network, and the exploration agent network, candidate multimedia resources corresponding to at least one interaction reason are learned, which contributes to the stay of the target object on the resource pushing platform. Based on the resource recommendation from the user's personalized preference, hotspot perception, and other dimensions, the exploration of the unpushed multimedia resources can be combined with the target exploration network to better improve the exploration accuracy of the user's resource interaction preference, and thus the resource pushing effect and the user retention rate in the resource pushing platform can be improved.

[0152] In an optional embodiment, in the case that the target resource interaction information is a target interaction graph with the object attribute information of the target object and the resource attribute information of the historical multimedia resources as nodes, and at least one interaction operation performed by the target object on the historical multimedia resources as edges, any agent network can include a graph convolution network, an interaction state encoding network, a stay learning network, and a filtering network of the interaction reason corresponding to any agent network; wherein the graph convolution network can be used to learn the feature information of the interaction graph. The interaction state encoding network can be used to convert the graph feature information into state encoding information for the stay learning network to learn the stay contribution. Specifically, the state encoding information can represent the access state (i.e., the interaction state) of the corresponding multimedia resource. The filtering network of the interaction reason corresponding to any agent network can be used to filter out the candidate multimedia resources corresponding to the interaction reason. The stay learning network is used to learn the index data of the stay contribution of the target object on the resource pushing platform in combination with the state encoding information and the candidate multimedia resources corresponding to the corresponding interaction reason. Optionally, as shown in Figure 3 The target resource interaction information and the first resource attribute information are input into at least one agent network for resource interaction learning, and the pushed multimedia resources corresponding to at least one agent network can include:

[0153] In step S2031, the target interaction graph is characterized based on the graph convolution network to obtain first graph feature information;

[0154] In step S2033, the first graph feature information is processed by the interaction state encoding network to obtain first state encoding information;

[0155] In step S2035, the first resource attribute information is input into the filtering network for interaction preference filtering to obtain at least one candidate multimedia resource;

[0156] In step S2037, the first state encoding information and the candidate multimedia resources are input into the stay learning network for stay contribution learning to obtain first stay index data corresponding to at least one interaction reason.

[0157] In step S2039, based on the first stay index data, a push multimedia resource is determined from the candidate multimedia resources.

[0158] In one specific embodiment, the target interaction graph corresponding to the interactive operation has a common representation of an adjacency list, an adjacency set, and an adjacency matrix, and the like. Optionally, in the case where the at least one interactive operation includes multiple interactive operations, the target interaction graph can be a composite edge graph structure (a multi-interactive operation dimensional graph structure). Optionally, the multi-interactive operation dimensional graph structure can be represented by an adjacency tensor matrix. wherein is a p-dimensional (interactive operation dimension) correlation matrix between node i and node j; further, after representing the target interaction graph in the form of an adjacency tensor matrix, the graph neural network can be input to learn the feature representation to obtain first graph feature information. Specifically, the first graph feature information can be based on the feature representation of the target interaction graph learned by the graph neural network.

[0159] In one specific embodiment, in order to fully mine the potential relationship information in the interaction graph other than the interaction between the object and the multimedia resource, two sub-networks are set in the graph neural network for learning the features of the nodes and edges, respectively. Correspondingly, the graph neural network can include an edge feature extraction network and a node feature extraction network. Optionally, in the case where the graph neural network includes multiple layers of networks, each layer of network can be provided with an edge feature extraction network and a node feature extraction network. Optionally, the adjacency tensor matrix corresponding to the target interaction graph can be input to the edge feature extraction network and the node feature extraction network in the first layer of network to extract the edge features and the node features, and the output of the edge feature extraction network of the previous layer is sequentially taken as the input of the edge feature extraction network of the next layer, and the output of the node feature extraction network of the previous layer is sequentially taken as the input of the node feature extraction network of the next layer. The edge feature extraction network and the node feature extraction network of the last layer are taken as the first graph feature information.

[0160] In one specific embodiment, the filtering network of the interaction reason corresponding to any agent network can be used to filter out the candidate multimedia resources corresponding to the interaction reason. Optionally, taking the hotspot agent network as an example, in the interest agent network, the interaction information in the first resource attribute information can be combined to determine the interaction index data (the numerical value of the interaction index data is positively correlated with the multimedia resource heat) that can represent the multimedia resource heat; then, based on the interaction index data, the candidate multimedia resources corresponding to the hotspot interaction reason can be filtered out from the to-be-pushed multimedia resources. Optionally, based on the interaction index data, the candidate multimedia resources corresponding to the hotspot interaction reason can be filtered out from the to-be-pushed multimedia resources, which can include: based on the interaction index data, the multimedia resources in the to-be-pushed multimedia resources are sorted in descending order; then, from the to-be-pushed multimedia resources, the multimedia resources with a preset number of the interaction index data sorting are filtered out as the candidate multimedia resources corresponding to the hotspot interaction reason. Optionally, based on the interaction index data, the candidate multimedia resources corresponding to the hotspot interaction reason can be filtered out from the to-be-pushed multimedia resources, which can include: the multimedia resources (hot resources) corresponding to the interaction index data greater than or equal to a preset threshold in the to-be-pushed multimedia resources are taken as the candidate multimedia resources corresponding to the hotspot interaction reason.

[0161] In one specific embodiment, taking the interest agent network as an example, the object attribute information of the target object and the first resource attribute information can be input into the interest agent network, and in the interest agent network, the similarity between the object attribute information and the first resource attribute information is calculated; then, based on the similarity, the candidate multimedia resources corresponding to the interest interaction reason are filtered out from the to-be-pushed multimedia resources. Specifically, the specific step refinement of filtering out the candidate multimedia resources corresponding to the interest interaction reason from the to-be-pushed multimedia resources based on the similarity can be referred to the specific step refinement of filtering out the candidate multimedia resources corresponding to the hotspot interaction reason from the to-be-pushed multimedia resources based on the interaction index data, which will not be described here.

[0162] In one specific embodiment, the first stay index data described above can represent that at least one candidate multimedia resource contributes to the stay of the target object on the resource pushing platform.

[0163] In one specific embodiment, the first stay index data corresponding to each interaction reason can be combined to determine the pushed multimedia resource corresponding to the interaction reason from the candidate multimedia resources corresponding to the interaction reason. Specifically, the specific step refinement of determining the pushed multimedia resource from the candidate multimedia resources based on the first stay index data can be referred to the specific step refinement of filtering out the candidate multimedia resources corresponding to the hotspot interaction reason from the to-be-pushed multimedia resources based on the interaction index data, which will not be described here.

[0164] In the above embodiments, the resource interaction information of the target interaction graph as the target object is subjected to feature representation learning, which can combine at least one interaction operation performed on the historical multimedia resource by the target object to more effectively learn the interaction features of the user; and in the resource interaction learning process, the at least one interaction reason filtering network is further combined to filter the interaction preferences of the to-be-pushed multimedia resource, so that the candidate multimedia resource corresponding to different interaction reasons can be more targeted to contribute to the retention of the target object on the resource pushing platform, thereby improving the effectiveness of the determined pushed multimedia resource and the resource pushing accuracy, improving the user retention rate of the resource pushing platform, and further reducing the system resource waste caused by invalid multimedia resource pushing, and greatly improving the system performance of the resource pushing platform.

[0165] In an optional embodiment, the above method can further include the step of pre-training at least one agent network, specifically, as shown in Figure 4 Pre-training at least one agent network can include the following steps:

[0166] In step S401, sample resource interaction information of a sample object and third resource attribute information of a sample multimedia resource are obtained.

[0167] In step S403, the sample resource interaction information and the third resource attribute information are input into at least one to-be-trained agent network for resource interaction learning to obtain second retention index data corresponding to at least one interaction reason;

[0168] In step S405, the second retention index data is input into a preset center network for global retention contribution learning to obtain global retention index data;

[0169] In step S407, based on the second retention index data, a second multimedia resource is determined from the first multimedia resource;

[0170] In step S409, based on the third resource attribute information of the second multimedia resource and the object attribute information of the sample object, resource interaction analysis is performed on the sample object to obtain an interaction multimedia resource of the sample object.

[0171] In step S411, based on the interaction multimedia resource and the global retention index data, at least one to-be-trained agent network is trained to obtain at least one agent network.

[0172] In a specific embodiment, the sample resource interaction information can represent at least one interaction operation information of the sample object on the historical multimedia resource of the sample object. Specifically, the third resource attribute information can be information for describing the sample multimedia resource. Specifically, the sample multimedia resource can be a multimedia resource in the resource pushing platform. Optionally, the sample multimedia resource and the to-be-pushed multimedia resource can include the same multimedia resource in the resource pushing platform, or can include different multimedia resources in the resource pushing platform.

[0173] In an optional embodiment, the sample resource interaction information is a sample interaction graph with the object attribute information of the sample object and the resource attribute information of the historical multimedia resource of the sample object as nodes, and at least one interaction operation performed by the sample object on the historical multimedia resource of the sample object as edges. Any to-be-trained intelligent agent network includes: a to-be-trained graph convolution network, a to-be-trained encoding network, a to-be-trained stay learning network, and a to-be-trained filtering network corresponding to the interaction reason of any to-be-trained intelligent agent network. As shown in Figure 5 The input of the sample resource interaction information and the third resource attribute information into the at least one to-be-trained intelligent agent network for resource interaction learning to obtain the second stay index data corresponding to the at least one interaction reason can include the following steps:

[0174] In step S501, the second graph feature information is obtained by performing feature representation learning on the sample interaction graph based on the to-be-trained graph convolution network.

[0175] In step S503, the second state encoding information is obtained by performing interaction state encoding processing on the second graph feature information based on the to-be-trained encoding network.

[0176] In step S505, the first multimedia resource is obtained by inputting the third resource attribute information into the to-be-trained filtering network for interaction preference filtering.

[0177] In step S507, the second stay index data is obtained by inputting the second state encoding information and the third resource attribute information of the first multimedia resource into the to-be-trained stay learning network for stay contribution learning.

[0178] In a specific embodiment, the second stay index data represents the first multimedia resource corresponding to at least one interaction reason in the sample multimedia resource, and the stay contribution of the sample object in the resource pushing platform. The specific steps of steps S501 to S507 can be referred to the above steps S2031-S2037, which will not be repeated here.

[0179] In the above embodiments, in the intelligent agent network training process, the resource interaction information of the sample interaction graph is taken as a sample object for feature representation learning, which can combine at least one interaction operation performed by the sample object on the historical multimedia resource to more effectively learn the interaction features of the sample user; and in the resource interaction learning process, the sample multimedia resource is also filtered by at least one interaction reason filtering network for interaction preference filtering, which can more specifically perform the contribution of the candidate multimedia resource corresponding to different interaction reasons to the sample object on the resource pushing platform, which can greatly improve the effectiveness and accuracy of the intelligent agent network trained for resource interaction learning, and thus can improve the user retention rate of the resource pushing platform in the subsequent resource pushing process based on at least one intelligent agent network, reduce the system resource waste caused by invalid multimedia resource pushing, and greatly improve the system performance of the resource pushing platform.

[0180] In one specific embodiment, the preset center network can be a network for global stay contribution learning in combination with the second stay indicator data output by the at least one intelligent agent network to be trained. Specifically, the second stay indicator data output by the at least one intelligent agent network to be trained can be weighted to obtain global stay indicator data. Specifically, the weight information corresponding to the second stay indicator data output by any intelligent agent network to be trained can be generated in combination with the input (i.e., state encoding information) of the intelligent agent network to be trained. Specifically, a weight generation network can be provided in the preset center network, which can include at least one linear layer and one nonlinear activation layer connected in sequence. Correspondingly, the corresponding weight information can be extracted from the state encoding information in combination with the at least one linear layer and the one nonlinear activation layer. Optionally, the weight information corresponding to any intelligent agent network to be trained can represent the importance of the intelligent agent network to be trained in the at least one intelligent agent network to be trained.

[0181] In one specific embodiment, the second multimedia resource corresponding to each interaction reason can be determined from the first multimedia resource corresponding to the interaction reason in combination with the second stay indicator data corresponding to the interaction reason. Specifically, the specific steps of determining the second multimedia resource from the first multimedia resource based on the second stay indicator data can be referred to the specific steps of filtering the candidate multimedia resource corresponding to the hot interaction reason from the to-be-pushed multimedia resource based on the interaction indicator data, which will not be described here.

[0182] In a specific embodiment, the third resource attribute information described above can be information for describing the second multimedia resource. Optionally, resource interaction analysis can be performed in combination with a resource interaction analysis network; specifically, the resource interaction analysis network can be obtained by performing resource interaction analysis training on a to-be-trained resource interaction analysis network based on object attribute information of a preset sample object and resource attribute information of historical multimedia resources of the preset sample object; optionally, the third resource attribute information of the second multimedia resource and the object attribute information of the sample object can be input into the resource interaction analysis network for resource interaction analysis to obtain interaction probability information corresponding to the second multimedia resource; and based on the interaction probability information, an interaction multimedia resource is filtered out from the second multimedia resource (the second multimedia resource includes multiple multimedia resources). Specifically, the interaction probability information corresponding to the second multimedia resource can represent a probability of the sample object performing an interaction operation on the second multimedia resource.

[0183] In an optional embodiment, as shown in Figure 6 Training at least one to-be-trained agent network based on the interaction multimedia resource and the global stay index data to obtain at least one agent network can include the following steps:

[0184] In step S601, sample resource interaction information is updated based on the interaction multimedia resource;

[0185] In step S603, reward information corresponding to any to-be-trained agent network is generated based on the interaction operation information corresponding to the interaction multimedia resource;

[0186] In step S605, second loss information corresponding to any to-be-trained agent network is determined according to the reward information corresponding to any to-be-trained agent network and second stay index data corresponding to any to-be-trained agent network;

[0187] In step S607, network parameters of at least one to-be-trained agent network are updated based on the second loss information and the global stay index data;

[0188] In step S609, based on the updated sample resource interaction information and the updated at least one to-be-trained agent network, the sample resource interaction information and the third resource attribute information are repeatedly input into the at least one to-be-trained agent network for resource interaction learning to obtain second stay index data corresponding to at least one interaction reason, until the step of updating the network parameters of the at least one to-be-trained agent network based on the second loss information and the global stay index data, until the reward information satisfies a preset condition;

[0189] In step S611, at least one to-be-trained agent network that satisfies the preset condition is used as at least one agent network.

[0190] In a specific embodiment, the resource attribute information of the interactive multimedia resource can be taken as a node, the sample object interacting with the interactive multimedia resource can be taken as an edge, the node corresponding to the interactive multimedia resource can be added to the sample interaction graph, and then the update of the sample resource interaction information can be realized.

[0191] In a specific embodiment, the reward information corresponding to the interaction operation information can be generated in combination with a preset reward generation function. The reward information can represent a stay contribution reward. Optionally, the reward information corresponding to any to-be-trained agent network can be added to the second stay index data corresponding to the to-be-trained agent network in combination with a preset linear function to obtain fifth stay index data; then, the second loss information corresponding to the second stay index data and the fifth stay index data can be calculated in combination with a preset loss function. Specifically, the second loss information can represent the difference between the second stay index data and the fifth stay index data.

[0192] In a specific embodiment, updating the network parameters of the at least one to-be-trained agent network based on the second loss information and the global stay index data can include updating the network parameters of the at least one to-be-trained agent network in combination with the global stay index data on the basis of updating the network parameters of the at least one to-be-trained agent network in combination with the second loss information. Specifically, the network parameters can be updated in combination with a gradient descent method or the like.

[0193] In a specific embodiment, the maximum value of the global reward information corresponding to the global stay index data can be consistent with the maximum value of the reward information corresponding to the at least one agent network; optionally, the reward information satisfying a preset condition can be the maximum reward information greater than or equal to a preset threshold.

[0194] In the above embodiments, during the training process of the at least one to-be-trained agent network, the network is updated in combination with the global stay index data while the network is updated based on the second loss information corresponding to the to-be-trained agent network itself, the resource interaction learning training can be performed from the global dimension, and then the resource interaction learning capability of the trained at least one agent network can be greatly improved, the user retention rate of the resource pushing platform in the subsequent resource pushing process based on the at least one agent network can be improved, the system resource waste caused by invalid multimedia resource pushing can be reduced, and the system performance of the resource pushing platform can be greatly improved.

[0195] In a specific embodiment, as shown in FIG. 1, Figure 7 Figure 7 is a structural schematic diagram in the training process of the at least one agent network according to an example embodiment. Optionally, the at least one agent network can include an exploration agent network, an interest agent network, and a hotspot agent network. ​

[0196] In the above embodiment, in the process of training the at least one agent network, the at least one agent network to be trained can learn the second stay index data corresponding to the at least one interaction reason from the sample resource interaction information and the third resource attribute information, then based on the second stay index data, the second multimedia resource corresponding to the at least one interaction reason can be selected from the first multimedia resource; and in combination with the third resource attribute information of the second multimedia resource and the object attribute information of the sample object, the sample object can be analyzed for resource interaction to obtain the interaction multimedia resource of the sample object, and then the network training can be performed in combination with the interaction multimedia resource and the global stay index data corresponding to the at least one agent network, which can greatly improve the effectiveness and accuracy of the at least one agent network trained for resource interaction learning, thereby improving the user retention rate of the resource pushing platform in the subsequent resource pushing process based on the at least one agent network, reducing the system resource waste caused by invalid multimedia resource pushing, and greatly improving the system performance of the resource pushing platform.

[0197] In an optional embodiment, in the case where the at least one agent network includes an exploration agent network, since the exploration agent network needs to be guided for exploration of the unpushed multimedia resource in combination with the corresponding target exploration network in the training process of the exploration agent network. Correspondingly, the above method can further include:

[0198] inputting the second state encoding information and the third resource attribute information of the first multimedia resource into the target exploration network corresponding to the exploration agent network for stay contribution learning to obtain third stay index data;

[0199] determining fourth stay index data based on the third stay index data and the state incentive information;

[0200] generating first loss information based on the fourth stay index data and the second stay index data;

[0201] updating the network parameters of the trained stay learning network corresponding to the exploration agent network based on the first loss information.

[0202] In a specific embodiment, the specific steps of inputting the second state encoding information and the third resource attribute information of the first multimedia resource into the target exploration network corresponding to the exploration agent network for stay contribution learning to obtain third stay index data can be referred to the specific steps of inputting the second state encoding information and the first multimedia resource into the trained stay learning network for stay contribution learning to obtain second stay index data, which will not be described herein.

[0203] In a specific embodiment, the state incentive information can be determined based on an interaction state of the first multimedia resource corresponding to the at least one interaction reason in the training process of the at least one to-be-trained agent network, and the interaction state represents a number of times that the first multimedia resource is determined as the second multimedia resource. Optionally, the state incentive information is greater when the number of times that the first multimedia resource is determined as the second multimedia resource is smaller. Optionally, the number of times that the first multimedia resource is determined as the second multimedia resource can be quantified into the state incentive information according to a preset quantification rule.

[0204] In a specific embodiment, the first loss information can be generated in combination with a preset loss function, and the network parameters can be adjusted in combination with a gradient descent method. In a specific embodiment, the network structure of the target exploration network is the same as that of the to-be-trained stay learning network, but the network parameters of the target exploration network are fixed and unchanged. Specifically, because the network parameters of the to-be-trained stay learning network change constantly during the training process, the outputs of the target exploration network and the to-be-trained stay learning network will be different (the stay index data is different) for the same input (the second state encoding information and the third resource attribute information of the first multimedia resource). Although the network parameters of the target exploration network are unchanged, the stay index data can be adjusted in combination with the state incentive information (i.e., the state incentive information is added to the third stay index data to obtain the fourth stay index data). Accordingly, during the training process, the first loss information is generated based on the fourth stay index data and the second stay index data, and the process of adjusting the to-be-trained stay learning network is constantly reduced, i.e., the difference between the outputs of the to-be-trained stay learning network and the target exploration network is constantly reduced. The to-be-trained stay learning network also learns the interaction information of the user on the unpushed multimedia resource.

[0205] In the above embodiments, the target exploration network is used to guide the exploration of the unpushed multimedia resource for the to-be-trained stay learning network in the to-be-trained exploration agent network, so that the to-be-trained stay learning network learns the interaction information of the user on the unpushed multimedia resource during the training process. Therefore, the resource interaction learning ability of the trained exploration agent network can be improved, the user retention rate of the resource pushing platform during the subsequent resource pushing process based on the exploration agent network can be improved, the system resource waste caused by invalid multimedia resource pushing can be reduced, and the system performance of the resource pushing platform can be greatly improved.

[0206] In step S205, the diversity of the pushed multimedia resource is analyzed to obtain diversity index data.

[0207] In a specific embodiment, the diversity index data can represent the diversity of the pushed multimedia resource. Optionally, as Figure 8As shown, the diversity analysis on the push multimedia resource to obtain the diversity index data can include the following steps:

[0208] In step S2051, the category index data of the push multimedia resource and the preset balance parameter are determined.

[0209] In step S2053, the resource association data corresponding to the push multimedia resource is generated based on the first resource attribute information of the push multimedia resource.

[0210] In step S2055, the push multimedia resource is analyzed for diversity based on the resource association data, the category index data, and the preset balance parameter, and diversity index data is obtained.

[0211] In one specific embodiment, the preset balance parameter can be used to balance the diversity and correlation between multimedia resources in the push multimedia resource. Optionally, the preset balance parameter can be set in combination with the demand for diversity and correlation in actual application. Specifically, the category index data can represent the probability that the multimedia resource belongs to the corresponding resource category.

[0212] In one optional embodiment, the determination of the category index data of the push multimedia resource can include:

[0213] The first resource attribute information of the push multimedia resource corresponding to any agent network is input into the resource category identification network corresponding to any agent network for resource category identification processing to obtain the category index data.

[0214] In one specific embodiment, the resource category identification network corresponding to any agent network can be obtained by performing resource category identification training on a preset neural network based on the second resource attribute information of the candidate multimedia resource corresponding to any agent network and the preset category data of the candidate multimedia resource. Optionally, the preset category data can be resource category information.

[0215] In the above embodiment, the resource category identification network trained in combination with the second resource attribute information of the candidate multimedia resource corresponding to the agent network performs resource category identification processing on the push multimedia resource corresponding to the interaction reason, which can better improve the accuracy and effectiveness of resource category identification.

[0216] In an optional embodiment, the push multimedia resource can include a plurality of multimedia resources, and the generating of the resource association data corresponding to the push multimedia resource based on the first resource attribute information of the push multimedia resource can include: determining the association data between the first resource attribute information of each two multimedia resources in the push multimedia resource, to obtain the resource association data. Specifically, the resource association data can include the similarity between each two multimedia resources in the push multimedia resource; optionally, the resource association data can be cosine similarity, Euclidean distance, etc. between the first resource attribute information.

[0217] In an optional embodiment, the diversity analysis on the push multimedia resource based on the resource association data, the category index data and the preset balance parameter can be performed to obtain the diversity index data, which can be combined with the following formula:

[0218]

[0219] wherein, the diversity index data, Diag is a function for constructing a diagonal matrix, is a preset balance coefficient, , the category index data; the resource association data.

[0220] In an optional embodiment, the diversity analysis on the push multimedia resource based on the resource association data, the category index data and the preset balance parameter can be performed to obtain the diversity index data, which can be combined with the following formula:

[0221]

[0222] wherein, the diversity index data, Diag is a function for constructing a diagonal matrix, is a preset balance coefficient, , the category index data; the resource association data.

[0223] In the above embodiments, the diversity analysis can be performed based on the resource association data, the category index data and the preset balance parameter corresponding to the push multimedia resource, to ensure the correlation of the multimedia resources, balance the diversity and correlation between the target multimedia resources determined based on the diversity index data, and further improve the effectiveness of the multimedia resource pushing, reduce the system resource waste caused by invalid multimedia resource pushing, and greatly improve the system performance of the resource pushing platform.

[0224] In addition, it should be noted that the specific refinement structure and number of layers of each network in the embodiments of the present specification can be set in combination with actual applications.

[0225] In step S207, based on the diversity index data, a target multimedia resource in the push multimedia resource is pushed to the target object.

[0226] In an optional embodiment, the above-mentioned pushing, based on the diversity index data, a target multimedia resource in the push multimedia resource to the target object comprises:

[0227] According to the diversity index data, a target multimedia resource is determined from the push multimedia resource;

[0228] The target multimedia resource is pushed to the target object.

[0229] In a specific embodiment, the above-mentioned determining, according to the diversity index data, a target multimedia resource from the push multimedia resource can refer to the specific step refinement of the above-mentioned filtering, based on the interaction index data, a candidate multimedia resource corresponding to a hot spot interaction reason from the to-be-pushed multimedia resource, which will not be described here.

[0230] In the above-mentioned embodiments, the target multimedia resource recommended to the target object is determined from the push multimedia resource in combination with the diversity index data, which can effectively improve the diversity of the target multimedia resource pushed to the target object, and further improve the effectiveness of multimedia resource recommendation, reduce the push of invalid multimedia resources, reduce system resource waste, and improve system performance.

[0231] In an optional embodiment, after the above-mentioned method pushes, based on the diversity index data, a target multimedia resource in the push multimedia resource to the target object, the method further comprises:

[0232] Obtaining continuous interaction data of the target object;

[0233] Determining first measurement index data based on the continuous interaction data and preset continuous interaction data.

[0234] In a specific embodiment, the above-mentioned continuous interaction data can represent the number of at least one interaction operation continuously performed by the target object on the target multimedia resource; optionally, the preset continuous interaction data can be set in combination with the demand for continuous interaction operation of the object in actual applications; optionally, the ratio of the continuous interaction data and the preset continuous interaction data can be taken as the first measurement index data. Specifically, the above-mentioned first measurement index data can be used to measure the contribution effect of the target multimedia resource to the stay of the target object on the resource push platform.

[0235] In the above embodiments, after the target multimedia resource is pushed to the target object, the first measurement index data is generated by combining the preset continuous interaction data with the continuous interaction data of the target object, which can measure the resource pushing effect from the contribution effect of the target multimedia resource to the stay of the target object on the resource pushing platform, greatly improve the effectiveness of the resource pushing effect, and further better improve the at least one intelligent agent network.

[0236] In an optional embodiment, the target multimedia resource includes a plurality of multimedia resources; optionally, the method can further include:

[0237] Obtaining fourth resource attribute information corresponding to the plurality of multimedia resources;

[0238] Performing resource category identification on the plurality of multimedia resources based on the fourth resource attribute information to obtain target resource categories of the plurality of multimedia resources;

[0239] Determining the second measurement index data according to the number of resource categories in the target resource categories.

[0240] In a specific embodiment, the fourth resource attribute information corresponding to any multimedia resource can be information for describing the multimedia resource. Specifically, resource category identification can be performed in combination with a pre-hung resource category identification network, and optionally, the division of resource categories can be set in combination with actual applications, such as tourism, food, etc.

[0241] In a specific embodiment, the second measurement index data can be used to measure the diversity of the plurality of multimedia resources. Optionally, the number of resource categories in the target resource categories of the plurality of multimedia resources can be counted, and the number is taken as the second measurement index data.

[0242] In the above embodiments, the second measurement index data is generated in combination with the number of resource categories of the plurality of multimedia resources pushed to the target object, which can measure the resource pushing effect from the resource category dimension, greatly improve the effectiveness of the resource pushing effect, and further better improve the at least one intelligent agent network.

[0243] According to the technical solutions provided by the embodiments of the present specification, the target object's target resource interaction information is learned from at least one intelligent agent network, candidate multimedia resources corresponding to at least one interaction reason are learned, the contribution of the target object to the resource pushing platform is determined, the different resource interaction preferences of the user can be more accurately captured in combination with the contribution, the effectiveness of the pushed multimedia resources determined can be improved, and in the case that at least one interaction reason corresponding to the pushed multimedia resources is obtained, the diversity index data can be obtained by performing diversity analysis on the pushed multimedia resources; and in combination with the diversity index data, the target multimedia resources pushed to the target object are selected from the pushed multimedia resources, the diversity of the pushed multimedia resources can be greatly improved on the basis of improving the user retention in the platform, the pushed multimedia resources are enriched, and thus the user demand can be better met, the resource pushing accuracy and effectiveness can be improved, the system resource waste caused by invalid multimedia resource pushing can be reduced, and the system performance of the resource pushing platform can be greatly improved.

[0244] Figure 9 is a block diagram of a multimedia resource pushing device according to an exemplary embodiment. Referring to Figure 9 , the device comprises:

[0245] The first information acquisition module 910 is configured to perform acquisition of target resource interaction information of a target object and first resource attribute information of a to-be-pushed multimedia resource, and the target resource interaction information represents at least one interaction operation information of the target object performed on historical multimedia resources;

[0246] The first resource interaction learning module 920 is configured to perform input of the target resource interaction information and the first resource attribute information into at least one intelligent agent network for resource interaction learning, and determine the pushed multimedia resources corresponding to the at least one intelligent agent network; the at least one intelligent agent network is used for learning at least one candidate multimedia resource, the contribution of the target object to the resource pushing platform, and the at least one candidate multimedia resource is a multimedia resource corresponding to at least one interaction reason in the to-be-pushed multimedia resource, and any interaction reason represents the reason for performing an interaction operation on the corresponding multimedia resource by the target object;

[0247] The diversity analysis module 930 is configured to perform diversity analysis on the pushed multimedia resources to obtain diversity index data;

[0248] The resource pushing module 940 is configured to perform, based on the diversity index data, pushing of target multimedia resources in the pushed multimedia resources to the target object.

[0249] In an optional embodiment, the diversity analysis module 930 comprises:

[0250] The data determination unit is configured to determine the category index data of the pushed multimedia resources and a preset balance parameter, the preset balance parameter being used to balance the diversity and correlation among the multimedia resources in the pushed multimedia resources.

[0251] The resource association data generation unit is configured to generate the resource association data corresponding to the pushed multimedia resources based on the first resource attribute information of the pushed multimedia resources.

[0252] The diversity analysis unit is configured to perform diversity analysis on the pushed multimedia resources based on the resource association data, the category index data and the preset balance parameter, and obtain diversity index data.

[0253] In an optional embodiment, the data determination unit comprises:

[0254] The resource category identification processing unit is configured to input the first resource attribute information of the pushed multimedia resources corresponding to any intelligent agent network into a resource category identification network corresponding to any intelligent agent network for resource category identification processing, and obtain the category index data.

[0255] The resource category identification network corresponding to any intelligent agent network is obtained by performing resource category identification training on a preset neural network based on the second resource attribute information of the candidate multimedia resources corresponding to any intelligent agent network and the preset category data of the candidate multimedia resources.

[0256] In an optional embodiment, the at least one intelligent agent network comprises at least one of an interest intelligent agent network, a hotspot intelligent agent network and an exploration intelligent agent network.

[0257] The interest intelligent agent network is used to learn the candidate multimedia resources corresponding to the interest interaction reasons and contribute to the stay of the target object on the resource pushing platform, the hotspot intelligent agent network is used to learn the candidate multimedia resources corresponding to the hotspot interaction reasons and contribute to the stay of the target object on the resource pushing platform, and the exploration intelligent agent network is used to learn the candidate multimedia resources corresponding to the preset interaction reasons in the case that the target exploration network guides the exploration of the unpushed multimedia resources, and contribute to the stay of the target object on the resource pushing platform.

[0258] In an optional embodiment, the target resource interaction information is a target interaction graph with nodes of object attribute information of the target object and resource attribute information of the historical multimedia resource and edges of at least one interaction operation performed by the target object on the historical multimedia resource;

[0259] The any-agent network comprises a graph convolution network, an interaction state encoding network, a stay learning network, and a filtering network of an interaction reason corresponding to the any-agent network; the first resource interaction learning module 920 comprises:

[0260] A first feature representation learning unit configured to perform feature representation learning on the target interaction graph based on the graph convolution network to obtain first graph feature information;

[0261] A first interaction state encoding processing unit configured to perform interaction state encoding processing on the first graph feature information based on the interaction state encoding network to obtain first state encoding information;

[0262] An interaction preference filtering unit configured to perform interaction preference filtering on the first resource attribute information input into the filtering network to obtain at least one candidate multimedia resource;

[0263] A stay contribution learning unit configured to perform stay contribution learning on the first state encoding information and the candidate multimedia resource input into the stay learning network to obtain first stay index data corresponding to at least one interaction reason, the first stay index data representing a stay contribution of the at least one candidate multimedia resource to the target object on the resource pushing platform;

[0264] A pushed multimedia resource determination unit configured to perform pushed multimedia resource determination based on the first stay index data from the candidate multimedia resource.

[0265] In an optional search box, the resource pushing module 940 comprises:

[0266] A target multimedia resource determination unit configured to perform target multimedia resource determination from the pushed multimedia resource according to the diversity index data;

[0267] A resource pushing unit configured to perform resource pushing of the target multimedia resource to the target object.

[0268] In an optional embodiment, the above device further comprises:

[0269] A second information acquisition module configured to perform acquisition of sample resource interaction information of a sample object and third resource attribute information of a sample multimedia resource, the sample resource interaction information representing at least one interaction operation information performed by the sample object on the historical multimedia resource of the sample object;

[0270] The second resource interaction learning module is configured to input the sample resource interaction information and the third resource attribute information into at least one to-be-trained intelligent agent network to perform resource interaction learning, to obtain second stay index data corresponding to at least one interaction reason, and the second stay index data represents a contribution of the sample object to stay on the resource pushing platform for a first multimedia resource corresponding to the at least one interaction reason in the sample multimedia resource.

[0271] The global stay contribution learning module is configured to input the second stay index data into a preset center network to perform global stay contribution learning, to obtain global stay index data.

[0272] The second multimedia resource determination module is configured to determine a second multimedia resource from the first multimedia resource based on the second stay index data.

[0273] The resource interaction analysis module is configured to perform resource interaction analysis on the sample object based on the third resource attribute information of the second multimedia resource and the object attribute information of the sample object, to obtain an interaction multimedia resource of the sample object.

[0274] The network training module is configured to train at least one to-be-trained intelligent agent network based on the interaction multimedia resource and the global stay index data, to obtain at least one intelligent agent network.

[0275] In an optional embodiment, the sample resource interaction information is a sample interaction graph with the object attribute information of the sample object and the resource attribute information of the historical multimedia resource of the sample object as nodes, and at least one interaction operation performed by the sample object on the historical multimedia resource of the sample object as edges; any to-be-trained intelligent agent network includes a to-be-trained graph convolution network, a to-be-trained encoding network, a to-be-trained stay learning network, and a to-be-trained filtering network corresponding to the interaction reason of any to-be-trained intelligent agent network; the second resource interaction learning module includes:

[0276] The second feature representation learning unit is configured to perform feature representation learning on the sample interaction graph based on the to-be-trained graph convolution network, to obtain second graph feature information.

[0277] The second interaction state encoding processing unit is configured to perform interaction state encoding processing on the second graph feature information based on the to-be-trained encoding network, to obtain second state encoding information.

[0278] The second interaction preference filtering unit is configured to input the third resource attribute information into the to-be-trained filtering network to perform interaction preference filtering, to obtain the first multimedia resource.

[0279] The second dwell contribution learning unit is configured to perform dwell contribution learning by inputting the second state encoding information and the third resource attribute information of the first multimedia resource into the trained dwell learning network to obtain second dwell indicator data.

[0280] In an optional embodiment, in the case that the at least one agent network comprises an exploration agent network, the apparatus further comprises:

[0281] The dwell contribution learning module is configured to perform dwell contribution learning by inputting the second state encoding information and the third resource attribute information of the first multimedia resource into the target exploration network corresponding to the exploration agent network to obtain third dwell indicator data.

[0282] The fourth dwell indicator data determination module is configured to perform determination of the fourth dwell indicator data based on the third dwell indicator data and state incentive information, the state incentive information being determined based on an interaction state of the first multimedia resource corresponding to at least one interaction reason in a training process of the at least one trained agent network, the interaction state representing a number of times that the first multimedia resource is determined as the second multimedia resource.

[0283] The first loss information generation module is configured to perform generation of the first loss information based on the fourth dwell indicator data and the second dwell indicator data.

[0284] The network parameter module is configured to perform update of the network parameters of the trained dwell learning network corresponding to the exploration agent network based on the first loss information.

[0285] In an optional embodiment, the network training module comprises:

[0286] The sample resource interaction information update unit is configured to perform update of the sample resource interaction information based on the interaction multimedia resource.

[0287] The reward information generation unit is configured to perform generation of the reward information corresponding to any trained agent network based on the interaction operation information corresponding to the interaction multimedia resource.

[0288] The second loss information determination unit is configured to perform determination of the second loss information corresponding to any trained agent network based on the reward information corresponding to any trained agent network and the second dwell indicator data corresponding to any trained agent network.

[0289] The network parameter update unit is configured to perform update of the network parameters of the at least one trained agent network based on the second loss information and the global dwell indicator data.

[0290] The training processing unit is configured to perform the step of repeatedly inputting the sample resource interaction information and the third resource attribute information into the at least one to-be-trained intelligent agent network for resource interaction learning based on the updated sample resource interaction information and the updated at least one to-be-trained intelligent agent network, to obtain second stay index data corresponding to at least one interaction reason, until the network parameters of the at least one to-be-trained intelligent agent network are updated based on the second loss information and the global stay index data, until the reward information meets the preset condition.

[0291] The intelligent agent network determination unit is configured to execute the at least one to-be-trained intelligent agent network meeting the preset condition as the at least one intelligent agent network.

[0292] In an optional embodiment, the apparatus described above further includes:

[0293] The continuous interaction data acquisition module is configured to execute the step of acquiring continuous interaction data of the target object after the target multimedia resource in the pushed multimedia resource is pushed to the target object based on the diversity index data, the continuous interaction data representing an operation quantity of at least one interaction operation continuously performed by the target object on the target multimedia resource.

[0294] The first measurement index data determination module is configured to execute the step of determining the first measurement index data based on the continuous interaction data and the preset continuous interaction data, the first measurement index data being used to measure a contribution effect of the target multimedia resource on the stay of the target object on the resource pushing platform.

[0295] In an optional embodiment, the target multimedia resource includes a plurality of multimedia resources; and the apparatus described above further includes:

[0296] The fourth resource attribute information acquisition module is configured to execute the step of acquiring fourth resource attribute information corresponding to the plurality of multimedia resources.

[0297] The resource category identification module is configured to execute the step of performing resource category identification on the plurality of multimedia resources based on the fourth resource attribute information, to obtain a target resource category of the plurality of multimedia resources.

[0298] The second measurement index data determination module is configured to execute the step of determining the second measurement index data according to a quantity of resource categories in the target resource category; the second measurement index data being used to measure the diversity of the plurality of multimedia resources.

[0299] As to the apparatus in the above-described embodiments, the specific manners in which various modules perform operations have been described in detail in the embodiments of the method, and will not be described herein in detail.

[0300] Figure 10is a block diagram of an electronic device for multimedia resource pushing according to an example embodiment. The electronic device can be a terminal, and its internal structure can be as shown in Figure 10 The electronic device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. The processor of the electronic device is configured to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is configured to communicate with an external terminal through a network connection. The computer program is executed by the processor to implement a multimedia resource pushing method. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device can be a touch layer overlaid on the display screen, or a key, trackball, or touchpad provided on the housing of the electronic device, or an external keyboard, touchpad, or mouse.

[0301] Figure 11 is a block diagram of an electronic device for multimedia resource pushing according to an example embodiment. The electronic device can be a terminal, and its internal structure can be as shown in Figure 11 The electronic device includes a processor, a memory, and a network interface connected through a system bus. The processor of the electronic device is configured to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is configured to communicate with an external terminal through a network connection. The computer program is executed by the processor to implement a multimedia resource pushing method.

[0302] Those skilled in the art can understand that Figure 10 or Figure 11 The structure shown in the above examples is only a block diagram of part of the structure related to the present disclosure, and does not constitute a limitation on the electronic device to which the present disclosure is applied. The specific electronic device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0303] In an example embodiment, an electronic device is also provided, including a processor, a memory for storing instructions executable by the processor, and wherein the processor is configured to execute the instructions to implement a multimedia resource pushing method as in the embodiments of the present disclosure.

[0304] In an example embodiment, a computer readable storage medium is also provided, which, when instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to perform the multimedia resource pushing method in the embodiments of the present disclosure.

[0305] In an example embodiment, a computer program product containing instructions, which, when run on a computer, enables the computer to perform the multimedia resource pushing method in the embodiments of the present disclosure, is also provided.

[0306] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing related hardware through a computer program, and the computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments of each method. Any reference to memory, storage, databases, or other media used to store data in embodiments provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in many forms such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0307] Other embodiments of the present disclosure will be apparent to those skilled in the art with the consideration of the specification and practice of the disclosed application. The present application is intended to cover any variations, uses, or adaptations of the present disclosure following the general principles thereof and including modifications and equivalents that are obvious to those skilled in the art. The specification and examples are to be regarded as exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.

[0308] It should be understood that the present disclosure is not limited to the precise structures as herein described and illustrated in the drawings, and that various modifications and changes can be made without departing from its scope. The scope of the present disclosure is limited only by the claims that follow.

Claims

1. A multimedia resource pushing method, characterized in that, include: Obtain target resource interaction information of the target object and first resource attribute information of the multimedia resource to be pushed, wherein the target resource interaction information represents at least one interaction operation information performed by the target object on historical multimedia resources; The target resource interaction information is a target interaction graph with the object attribute information of the target object and the resource attribute information of the historical multimedia resources as nodes, and at least one interaction operation performed by the target object on the historical multimedia resources as edges. The target resource interaction information and the first resource attribute information are input into at least one intelligent agent network for resource interaction learning to determine the push multimedia resources corresponding to the at least one intelligent agent network; the at least one intelligent agent network is used to learn at least one candidate multimedia resources to contribute to the target object's stay on the resource push platform, and the at least one candidate multimedia resource is a multimedia resource corresponding to at least one interaction reason among the multimedia resources to be pushed, and any one of the interaction reasons represents the reason why the target object performs an interaction operation on the corresponding multimedia resource; Any of the aforementioned intelligent agent networks includes: a graph convolutional network, an interaction state encoding network, a dwell learning network, and a filtering network for interaction reasons corresponding to any of the aforementioned intelligent agent networks; the step of inputting the target resource interaction information and the first resource attribute information into at least one intelligent agent network for resource interaction learning, and determining the push multimedia resource corresponding to the at least one intelligent agent network, includes: performing feature representation learning on the target interaction graph based on the graph convolutional network to obtain first graph feature information; performing interaction state encoding processing on the first graph feature information based on the interaction state encoding network to obtain first state encoding information; inputting the first resource attribute information into the filtering network for interaction preference filtering to obtain the at least one candidate multimedia resource; inputting the first state encoding information and the at least one candidate multimedia resource into the dwell learning network for dwell contribution learning to obtain first dwell index data corresponding to the at least one interaction reason, wherein the first dwell index data characterizes the at least one candidate multimedia resource for the dwell contribution of the target object on the resource push platform; and determining the push multimedia resource from the candidate multimedia resources based on the first dwell index data. A diversity analysis was performed on the pushed multimedia resources to obtain diversity index data; Based on the aforementioned diversity index data, the target multimedia resources from the pushed multimedia resources are pushed to the target object.

2. The multimedia resource push method according to claim 1, characterized in that, The diversity analysis of the pushed multimedia resources to obtain diversity index data includes: The category index data and preset balance parameters of the pushed multimedia resources are determined. The preset balance parameters are used to balance the diversity and correlation among the multimedia resources in the pushed multimedia resources. Based on the first resource attribute information of the pushed multimedia resource, resource association data corresponding to the pushed multimedia resource is generated; Based on the resource association data, the category index data, and the preset balance parameters, a diversity analysis is performed on the pushed multimedia resources to obtain the diversity index data.

3. The multimedia resource push method according to claim 2, characterized in that, The category index data for determining the pushed multimedia resources includes: The first resource attribute information of the multimedia resources pushed by any of the aforementioned intelligent agent networks is input into the resource category identification network corresponding to any of the aforementioned intelligent agent networks for resource category identification processing to obtain the category index data. The resource category identification network corresponding to any of the aforementioned agent networks is obtained by training a preset neural network to identify resource categories based on the second resource attribute information of the candidate multimedia resources corresponding to any of the aforementioned agent networks and the preset category data of the candidate multimedia resources.

4. The multimedia resource push method according to claim 1, characterized in that, The at least one agent network includes at least one of interest agent network, hotspot agent network and exploratory agent network; The interest-based intelligent agent network is used to learn candidate multimedia resources corresponding to interest-based interaction reasons, contributing to the target object's stay on the resource push platform; the hotspot intelligent agent network is used to learn candidate multimedia resources corresponding to hotspot interaction reasons, contributing to the target object's stay on the resource push platform; the exploration intelligent agent network is used to learn candidate multimedia resources corresponding to preset interaction reasons when the target exploration network guides the exploration of unpush multimedia resources, contributing to the target object's stay on the resource push platform; the target exploration network is used to guide the exploration intelligent agent network to explore and learn unpush multimedia resources during the process of the exploration intelligent agent network learning the candidate multimedia resources corresponding to the preset interaction reasons and contributing to the target object's stay on the resource push platform, whereby the unpush multimedia resources are the unpush candidate multimedia resources among the candidate multimedia resources corresponding to the preset interaction reasons.

5. The multimedia resource push method according to any one of claims 1 to 4, characterized in that, The step of pushing the target multimedia resource from the multimedia resources to the target object based on the diversity index data includes: Based on the diversity index data, target multimedia resources are determined from the pushed multimedia resources; The target multimedia resource is pushed to the target object.

6. The multimedia resource push method according to any one of claims 1 to 4, characterized in that, The method further includes: Obtain sample resource interaction information of a sample object and third resource attribute information of sample multimedia resources. The sample resource interaction information represents at least one interactive operation information performed by the sample object on the historical multimedia resources of the sample object. The sample resource interaction information and the third resource attribute information are input into at least one intelligent agent network to be trained for resource interaction learning, and the second dwell index data corresponding to the at least one interaction reason is obtained. The second dwell index data represents the first multimedia resource corresponding to the at least one interaction reason in the sample multimedia resources, and the contribution of the sample object to the dwell time on the resource push platform. The second stay index data is input into a preset central network to learn the global stay contribution, and the global stay index data is obtained. Based on the second dwell time index data, the second multimedia resource is determined from the first multimedia resource; Based on the third resource attribute information of the second multimedia resource and the object attribute information of the sample object, resource interaction analysis is performed on the sample object to obtain the interactive multimedia resources of the sample object. Based on the interactive multimedia resources and the global dwell time index data, the at least one agent network to be trained is trained to obtain the at least one agent network.

7. The multimedia resource push method according to claim 6, characterized in that, The sample resource interaction information is a sample interaction graph with the object attribute information of the sample object and the resource attribute information of the sample object's historical multimedia resources as nodes, and at least one interaction operation performed by the sample object on the sample object's historical multimedia resources as edges. Any of the agent networks to be trained includes: a graph convolutional network to be trained, an encoding network to be trained, a dwell learning network to be trained, and a filtering network to be trained for the interaction reasons corresponding to any of the agent networks to be trained; the step of inputting the sample resource interaction information and the third resource attribute information into at least one agent network to be trained for resource interaction learning to obtain the second dwell index data corresponding to the at least one interaction reason includes: Based on the convolutional network of the graph to be trained, feature representation learning is performed on the sample interaction graph to obtain the feature information of the second graph. Based on the coding network to be trained, the feature information of the second graph is processed by interactive state coding to obtain the second state coding information; The third resource attribute information is input into the filter network to be trained for interaction preference filtering to obtain the first multimedia resource. The second state encoding information and the third resource attribute information of the first multimedia resource are input into the stay learning network to be trained for stay contribution learning, so as to obtain the second stay index data.

8. The multimedia resource push method according to claim 7, characterized in that, In the case where the at least one agent network includes an exploratory agent network, the method further includes: The second state encoding information and the third resource attribute information of the first multimedia resource are input into the target exploration network corresponding to the exploration agent network to learn the dwell contribution and obtain the third dwell index data. Based on the third dwell index data and state incentive information, the fourth dwell index data is determined. The state incentive information is determined based on the interaction state of the first multimedia resource corresponding to the at least one interaction reason during the training of the at least one agent network to be trained. The interaction state represents the number of times the first multimedia resource is identified as the second multimedia resource. Based on the fourth stay index data and the second stay index data, first loss information is generated; The network parameters of the training dwell learning network corresponding to the exploratory agent network are updated based on the first loss information.

9. The multimedia resource push method according to claim 6, characterized in that, The step of training the at least one agent network based on the interactive multimedia resources and the global dwell time index data to obtain the at least one agent network includes: Based on the interactive multimedia resources, update the sample resource interaction information; Based on the interactive operation information corresponding to the interactive multimedia resources, generate reward information for any of the intelligent agent networks to be trained. Based on the reward information corresponding to any of the agent networks to be trained and the second dwell index data corresponding to any of the agent networks to be trained, determine the second loss information corresponding to any of the agent networks to be trained; The network parameters of the at least one agent network to be trained are updated based on the second loss information and the global dwell index data. Based on the updated sample resource interaction information and the updated at least one agent network to be trained, repeat the steps of inputting the sample resource interaction information and the third resource attribute information into the at least one agent network to be trained for resource interaction learning, to obtain the second dwell index data corresponding to the at least one interaction reason, until the steps of updating the network parameters of the at least one agent network to be trained based on the second loss information and the global dwell index data are repeated, until the reward information meets the preset conditions. At least one agent network that meets the preset conditions is selected as the at least one agent network.

10. The multimedia resource push method according to any one of claims 1 to 4, characterized in that, After pushing the target multimedia resource from the pushed multimedia resources to the target object based on the diversity index data, the method further includes: Acquire continuous interaction data of the target object, wherein the continuous interaction data represents the number of operations performed by the target object on the target multimedia resource in a continuous manner; Based on the continuous interaction data and the preset continuous interaction data, a first measurement index is determined. The first measurement index is used to measure the contribution effect of the target multimedia resource on the target object's stay on the resource push platform.

11. The multimedia resource push method according to any one of claims 1 to 4, characterized in that, The target multimedia resource includes multiple multimedia resources; the method further includes: Obtain the fourth resource attribute information corresponding to the plurality of multimedia resources; Based on the fourth resource attribute information, resource category identification is performed on the multiple multimedia resources to obtain the target resource category of the multiple multimedia resources; A second metric is determined based on the number of resource categories in the target resource category; the second metric is used to measure the diversity of the multiple multimedia resources.

12. A multimedia resource push device, characterized in that, include: The first information acquisition module is configured to acquire target resource interaction information of the target object and first resource attribute information of the multimedia resource to be pushed, wherein the target resource interaction information represents at least one interaction operation information performed by the target object on historical multimedia resources; The target resource interaction information is a target interaction graph with the object attribute information of the target object and the resource attribute information of the historical multimedia resources as nodes, and at least one interaction operation performed by the target object on the historical multimedia resources as edges. The first resource interaction learning module is configured to input the target resource interaction information and the first resource attribute information into at least one intelligent agent network for resource interaction learning, and determine the push multimedia resources corresponding to the at least one intelligent agent network; the at least one intelligent agent network is used to learn at least one candidate multimedia resources to contribute to the target object's stay on the resource push platform, and the at least one candidate multimedia resource is a multimedia resource corresponding to at least one interaction reason among the multimedia resources to be pushed, and any one of the interaction reasons represents the reason why the target object performs an interaction operation on the corresponding multimedia resource; Any of the aforementioned intelligent agent networks includes: a graph convolutional network, an interaction state encoding network, a dwell learning network, and a filtering network for interaction reasons corresponding to any of the aforementioned intelligent agent networks; the first resource interaction learning module includes: a first feature representation learning unit, configured to perform feature representation learning on the target interaction graph based on the graph convolutional network to obtain first graph feature information; a first interaction state encoding processing unit, configured to perform interaction state encoding processing on the first graph feature information based on the interaction state encoding network to obtain first state encoding information; an interaction preference filtering unit, configured to input the first resource attribute information into the filtering network to perform interaction preference filtering to obtain the at least one candidate multimedia resource; a dwell contribution learning unit, configured to input the first state encoding information and the at least one candidate multimedia resource into the dwell learning network to perform dwell contribution learning to obtain first dwell index data corresponding to the at least one interaction reason, wherein the first dwell index data represents the dwell contribution of the at least one candidate multimedia resource to the target object on the resource push platform; and a push multimedia resource determination unit, configured to determine the push multimedia resource from the candidate multimedia resources based on the first dwell index data; The diversity analysis module is configured to perform diversity analysis on the pushed multimedia resources to obtain diversity index data. The resource push module is configured to push target multimedia resources from the multimedia resources to the target object based on the diversity index data.

13. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the multimedia resource push method as described in any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is able to perform the multimedia resource push method as described in any one of claims 1 to 11.

15. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the multimedia resource push method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Method and device for determining pushed information

    CN103838756A

  • Multimedia resource pushing method and device, electronic equipment and storage medium

    CN114048392A