A recommendation method, device and equipment for multimedia resources and a storage medium

By acquiring the relationship type feature information between the target object and its neighboring objects, and calculating the representation feature information of the target object, the problem of low matching degree in existing multimedia resource recommendation algorithms is solved, and more accurate multimedia resource recommendation is achieved.

CN115495598BActive Publication Date: 2026-02-10TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111680341.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-06-17
Filing Date
2021-12-30
Publication Date
2026-02-10
Estimated Expiration
2041-12-30

AI Technical Summary

Technical Problem

Existing multimedia resource recommendation algorithms have low matching accuracy and cannot accurately recommend multimedia resources that users are interested in.

Method used

By obtaining the representation vector of the target object and the set of representation vectors of neighboring objects, and based on the relational feature information of K first relation types, the representation feature information of the target object is calculated, and multimedia resources with the highest matching degree are recommended.

Benefits of technology

It improved the accuracy of multimedia resource recommendations, uncovered the interests of the target audience, and enhanced the precision of the recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115495598B_ABST
    Figure CN115495598B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a multimedia resource recommendation method and device, equipment and a storage medium. The method comprises: obtaining a representation vector of a target object and a set of representation vectors of adjacent objects; obtaining representation feature information of the target object according to the representation vector of the target object and the set of representation vectors of adjacent objects, wherein the representation feature information is determined according to relationship feature information corresponding to each first relationship type in K first relationship types between the target object and adjacent objects; obtaining a set of multimedia resources; and recommending a first multimedia resource in the set of multimedia resources to the target object according to the representation feature information of the target object. It can be seen that the interest points of the target object are mined through the relationship types between the target object and adjacent objects, thereby improving the recommendation accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and more specifically to a method, apparatus, device, and computer-readable storage medium for recommending multimedia resources. Background Technology

[0002] With the development of computer technology, a massive amount of multimedia resources have emerged on the internet. Currently, most multimedia resource recommendation algorithms are not accurate enough in performing operations such as multimedia recommendations, resulting in low matching accuracy. Summary of the Invention

[0003] This invention provides a method, apparatus, device, and storage medium for recommending multimedia resources, which can significantly improve the accuracy of recommendations.

[0004] On the one hand, embodiments of this application provide a method for recommending multimedia resources, including:

[0005] Obtain the representation vector of the target object and the set of representation vectors of neighboring objects. The set of representation vectors of neighboring objects includes the representation vectors of the first neighboring objects of the target object. The target object and the first neighboring objects are divided into K first relation types, where K is a positive integer.

[0006] Based on the representation vector of the target object and the set of representation vectors of neighboring objects, the representation feature information of the target object is obtained. The representation feature information is determined based on the relation feature information corresponding to each of the K first relation types of the target object.

[0007] Obtain a set of multimedia resources, and recommend a first multimedia resource from the set to the target object based on the representational feature information of the target object.

[0008] On the one hand, embodiments of this application also provide a multimedia resource recommendation device, including:

[0009] The acquisition unit is used to acquire the representation vector of the target object and the set of representation vectors of neighboring objects. The set of representation vectors of neighboring objects includes the representation vectors of the first neighboring objects of the target object. The target object and the first neighboring objects are divided into K first relation types, where K is a positive integer.

[0010] The processing unit is configured to obtain representation feature information of the target object based on the representation vector of the target object and the set of representation vectors of neighboring objects, wherein the representation feature information is determined based on the relation feature information corresponding to each of the K first relation types of the target object; and to acquire a set of multimedia resources and recommend a first multimedia resource in the set of multimedia resources to the target object based on the representation feature information of the target object.

[0011] In one embodiment, the processing unit is specifically used for:

[0012] Based on the representation vector of the target object and the set of representation vectors of the adjacent objects, the relation feature information corresponding to each of the K first relation types is obtained;

[0013] The target object is analyzed by performing feature analysis on the K relation feature information corresponding to the K first relation types respectively, and the representation feature information of the target object is obtained.

[0014] In one embodiment, the processing unit is specifically used for:

[0015] Obtain the feature parameter set of the h-th first relation type among the K first relation types, the feature parameter set including the weight matrix and the bias vector of the h-th first relation type;

[0016] Using the feature parameter set of the h-th first relation type, calculate the target object intermediate features under the h-th first relation type, and calculate the intermediate features of each first neighboring object of the target object under the h-th first relation type;

[0017] Based on the intermediate features of the target object and the intermediate features of each adjacent object, the relation feature information of the h-th first relation type is obtained.

[0018] In one embodiment, the relation feature information of the h-th first relation type is the relation feature information of the h-th first relation type at the T-th iteration, where T is a positive integer; the processing unit is specifically used for:

[0019] Obtain the target probability that each of the first neighboring objects of the target object is classified into the h-th first relation type at the t-th iteration, where t is a positive integer and t is less than T;

[0020] Based on the target probability and the intermediate features of each neighboring object, calculate the aggregate features of the first neighboring object of the target object;

[0021] The intermediate features of the target object and the aggregate features of the first neighboring objects of the target object are processed to obtain the relation feature information of the h-th first relation type at the (t+1)-th iteration.

[0022] In one embodiment, the processing unit is specifically used for:

[0023] Obtain a first feature information set of the first neighboring objects of the target object, and a second feature information set of the second neighboring objects of the target object. The first feature information set includes: relationship feature information corresponding to each of the multiple second relationship types of the first neighboring objects of the target object. The second feature information set includes: relationship feature information corresponding to each of the multiple third relationship types of the second neighboring objects of the target object.

[0024] The relationship feature information corresponding to the K first relationship types of the target object, the first feature information set, and the second feature information set are used as inputs to the relationship prediction model to obtain the prediction result output by the relationship prediction model. The relationship prediction model includes L layers of graph convolutional network layers, where L is a positive integer.

[0025] Overfitting is applied to the prediction results to obtain the representational feature information of the target object;

[0026] In the relationship prediction model, the input data of the g-th graph convolutional network layer includes the data obtained after overfitting the output data of the (g-1)-th graph convolutional network layer.

[0027] In one embodiment, the processing unit is specifically used for:

[0028] Obtain viewing information of the second multimedia resource, wherein the viewing information includes object identifiers of Q objects that have viewed the second multimedia resource, where Q is a positive integer;

[0029] Based on the object identifiers of the Q objects, obtain the representation vectors of the Q objects;

[0030] The representation vectors of the Q objects are fused, and the result of the fusion is averaged and pooled to obtain the resource feature information of the second multimedia resource.

[0031] The multimedia resource set includes the resource feature information of the second multimedia resource.

[0032] In one embodiment, the processing unit is specifically used for:

[0033] Based on the representation feature information of the target object and the resource feature information of each multimedia resource in the multimedia resource set, the matching degree between the target object and each multimedia resource in the multimedia resource set is calculated.

[0034] A first multimedia resource is recommended to the target object, wherein the first multimedia resource is the multimedia resource in the multimedia resource set that has the highest matching degree with the target object.

[0035] In one embodiment, the processing unit is specifically used for:

[0036] The resource feature information of the second multimedia resource in the multimedia resource set is concatenated with the representation feature information of the target object to obtain a concatenated feature set;

[0037] A multilayer perceptron is used to process each splicing feature in the splicing feature set to obtain the relationship vector between each of the K first relationship types and the second multimedia resource;

[0038] Calculate the weight corresponding to each first relation type based on the relation vector between each first relation type and the second multimedia resource;

[0039] Based on the relationship vector between each of the K first relationship types and the second multimedia resource, and the weight corresponding to each first relationship type, the matching degree between the target object and the second multimedia resource is obtained.

[0040] In one embodiment, the processing unit is further configured to:

[0041] Obtain a set of association information, which includes a set of object information and a set of relationship information;

[0042] N network nodes are generated based on the object information set. Each of the N network nodes corresponds to an object, and each network node carries object information of the object corresponding to that network node. N is a positive integer.

[0043] If the set of relational information indicates that the object corresponding to the first network node and the object of the second network node among the N network nodes have interactive behavior, then the connection information between the first network node and the second network node is generated according to the interactive behavior to obtain the relational information network graph.

[0044] Based on the associated information network graph, the representation vectors of the N objects corresponding to the N network nodes are obtained.

[0045] In one embodiment, the edge weight in the edge information of the first network node and the second network node includes: a weight determined based on the correlation between the first network node and the second network node, wherein the edge weight is proportional to the correlation, and the correlation is determined based on the interaction information of the first network node and the second network node within a target time period, wherein the interaction information includes at least one of the following: cumulative number of interactions, cumulative interaction duration, interaction frequency, and interaction content.

[0046] In one embodiment, the processing unit is specifically used for:

[0047] Starting from the target network node corresponding to the target object, a random walk is performed in the associated information network graph to obtain M trajectories, each with a step size of P; where M and P are both positive integers.

[0048] Based on the object information carried in the M trajectories, the representation vector of the target object is obtained;

[0049] The probability of traversing from the i-th network node to the j-th network node is proportional to the target edge weight, which is the edge weight in the edge information of the i-th network node and the j-th network node. i and j are both positive integers, i is not equal to j, and i and j are both less than or equal to N.

[0050] Accordingly, embodiments of the present invention also provide an intelligent device, including: a storage device and a processor; the storage device stores a computer program; the processor executes the computer program to implement the above-described method for recommending multimedia resources.

[0051] Accordingly, embodiments of the present invention also provide a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the above-described method for recommending multimedia resources is implemented.

[0052] Accordingly, this application provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the aforementioned recommended method for multimedia resources.

[0053] In this embodiment, the representation vector of the target object and the set of representation vectors of neighboring objects are obtained. Based on the representation vector of the target object and the set of representation vectors of neighboring objects, the representation feature information of the target object is obtained. The representation feature information is determined based on the relationship feature information corresponding to each of the K first relationship types between the target object and its neighboring objects. A multimedia resource set is obtained, and based on the representation feature information of the target object, the first multimedia resource in the multimedia resource set is recommended to the target object. It is evident that by exploring the relationship types between the target object and its neighboring objects, the interest points of the target object are mined, thereby significantly improving the recommendation accuracy. Attached Figure Description

[0054] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0055] Figure 1 A recommended scenario diagram of multimedia resources provided for embodiments of this application;

[0056] Figure 2 A flowchart illustrating a method for recommending multimedia resources provided in an embodiment of this application;

[0057] Figure 3 A flowchart illustrating another method for recommending multimedia resources provided in this application embodiment;

[0058] Figure 4a A schematic diagram of an object association information network diagram provided in an embodiment of this application;

[0059] Figure 4b A schematic diagram of a graph convolution model based on a relational information network graph provided in an embodiment of this application;

[0060] Figure 5 A schematic diagram of a multimedia resource recommendation device provided in an embodiment of this application;

[0061] Figure 6 This is a schematic diagram of the structure of a smart device provided in an embodiment of this application. Detailed Implementation

[0062] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0063] This application relates to Artificial Intelligence (AI) and Machine Learning (ML). AI utilizes digital computers or computers-controlled machines to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results. In other words, artificial intelligence is a comprehensive technology within computer science; it primarily aims to understand the essence of intelligence and produce new intelligent devices that can react in a manner similar to human intelligence, enabling these devices to possess multiple functions such as perception, reasoning, and decision-making. The intelligent device provided in this application can recommend multimedia resources to a target object based on the relationship type between the target object and its neighboring objects.

[0064] AI technology is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, large application processing technologies, operating / interaction systems, and mechatronics. AI software technologies mainly include computer vision, speech processing, natural language processing, and machine learning / deep learning. In the process of constructing the interconnected information network graph in this application's embodiments, one or more of the aforementioned AI software technologies are involved when converting the interactive behaviors between objects into connection information; for example, computer vision technology is involved when extracting features of (short) videos sent between objects; speech processing technology is involved when extracting features of speech sent between objects; and natural language processing technology is involved when extracting features of text information sent between objects.

[0065] Computer vision technology typically includes image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and other technologies, as well as common biometric recognition technologies such as face recognition and fingerprint recognition.

[0066] Key technologies in speech technology include Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and voiceprint recognition. Enabling computers to hear, see, speak, and feel is the future direction of human-computer interaction, with speech emerging as one of the most promising methods.

[0067] Natural Language Processing (NLP) is an important field within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language people use in daily life—and thus it has a close relationship with linguistic research. NLP techniques typically include text processing, semantic understanding, machine translation, question answering, and knowledge graphs.

[0068] Machine learning is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of AI and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning / deep learning typically includes techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and formulaic learning. The embodiments of this application involve relevant machine learning techniques in the process of training a relationship prediction model based on an association information network graph.

[0069] Furthermore, this application's embodiments may also involve artificial intelligence cloud services and blockchain. Artificial intelligence cloud services are generally also referred to as AIaaS (AI as a Service). This is currently a mainstream service model for artificial intelligence platforms. Specifically, AIaaS platforms break down several common AI services and provide them as independent or packaged services in the cloud. This service model is similar to opening an AI-themed marketplace: all developers can access and use one or more AI services provided by the platform through API interfaces. Some experienced developers can also use the AI ​​framework and AI infrastructure provided by the platform to deploy and maintain their own dedicated cloud AI services. This application's embodiments mainly involve recommending multimedia resources to objects through a multimedia recommendation platform (i.e., artificial intelligence cloud services).

[0070] Blockchain is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, it is a decentralized database, a chain of data blocks linked using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer. In this embodiment, a smart device can obtain the association relationships of objects from the blockchain network and then recommend multimedia resources to target objects based on reliable association relationships; it can also upload the analyzed object representation feature information and multimedia resource resource resource resource feature information to the blockchain for subsequent use (for example, within a certain time period, the resource feature information of a certain multimedia resource may need to be matched with the representation feature information of multiple objects; uploading the multimedia resource resource resource feature information to the blockchain allows other network nodes assisting in multimedia resource recommendation to use it directly).

[0071] Please see Figure 1 , Figure 1 This is a recommended scenario diagram for multimedia resources provided in an embodiment of this application. For example... Figure 1 As shown, the recommended scenario for multimedia resources includes terminal device 101 and server 102. Terminal device 101 is the device used by the target object, and may include, but is not limited to, smartphones (such as Android phones, iOS phones, etc.), tablet computers, portable personal computers, mobile internet devices (MIDs), etc. The terminal device is equipped with a display device, which may also be a monitor, display screen, touch screen, etc. The touch screen may also be a touch screen, touch panel, etc., and this embodiment of the application does not limit the scope of the application.

[0072] Server 102 refers to a backend device capable of recommending personalized multimedia resources based on the identifier of the target object sent by terminal device 101. After determining the first multimedia resource to recommend to the target object based on the identifier of the target object sent by terminal device 101, server 102 can return the first multimedia resource to terminal device 101. Page 103 is a schematic diagram of a page displayed by terminal device 101 based on the first multimedia resource sent by server 102, as provided in this application. Server 102 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. In addition, multiple servers can be grouped into a blockchain network, with each server being a node in the blockchain network. Terminal device 101 and server 102 can be directly or indirectly connected through wired or wireless communication, which is not limited in this application.

[0073] It should be noted that, Figure 1 The number of terminal devices and servers in the recommended scenarios of the multimedia resources shown is only an example. For example, there can be multiple terminal devices and servers. This application does not limit the number of terminal devices and servers.

[0074] Optionally, the multimedia resource recommendation scenario may also include only the terminal device 101 equipped with a multimedia resource recommendation device. After the target opens the multimedia platform, the terminal device 101 recommends multimedia resources to the target target through the multimedia resource recommendation device (such as displaying the recommended multimedia resources in the interface 103).

[0075] It is understood that in the specific embodiments of this application, data related to object information, multimedia information, and object interaction behavior are involved. When the above embodiments of this application are applied to specific products or technologies, permission or consent from the object is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0076] Figure 1 In the multimedia resource recommendation scenario shown, the multimedia resource recommendation process mainly includes the following steps:

[0077] (1) Server 102 obtains the representation vector of the target object and the set of representation vectors of neighboring objects. In one embodiment, server 102 obtains the representation vector of the target object and the set of representation vectors of neighboring objects based on the association information network graph. The association information network graph is constructed based on the interaction behavior of various objects on the social platform. The set of representation vectors of neighboring objects includes the representation vector of the first neighboring object of the target object. The first neighboring object refers to the object that has interaction behavior with the target object. Similarly, the second neighboring object refers to the object that has interaction behavior with the first neighboring object of the target object but does not have interaction behavior with the target object. Further, the target object and the first neighboring object are divided into K first relation types, where K is a positive integer.

[0078] (2) Server 102 obtains the representation feature information of the target object based on the representation vector of the target object and the set of representation vectors of neighboring objects. The representation feature information is determined based on the relationship feature information corresponding to each of the K first relationship types of the target object. In one embodiment, server 102 obtains the relationship feature information corresponding to each of the K first relationship types based on the representation vector of the target object and the set of representation vectors of neighboring objects (i.e., calculates the relationship feature information corresponding to the first relationship type through the representation vector of the target object and the representation vectors of neighboring objects). The target object is represented by the K relationship feature information corresponding to the K first relationship types respectively, and the representation feature information of the target object is obtained. That is to say, the representation feature information of the target object is jointly represented by the K relationship feature information corresponding to the K first relationship types respectively (i.e., the representation feature information of the target object is obtained based on the relationship type between the target object and the neighboring objects of the target object).

[0079] (3) Server 102 obtains a multimedia resource set and recommends the first multimedia resource in the multimedia resource set (one or more multimedia resources in the multimedia resource set whose matching degree with the target object's representation feature information is higher than the matching degree threshold, i.e., the multimedia resources that the target object is most likely to be interested in) to the target object based on the representation feature information of the target object; In one embodiment, the resource feature information of each multimedia resource in the multimedia resource set is obtained by the representation vector of the object that has clicked on the multimedia resource.

[0080] In this embodiment, the representation vector of the target object and the set of representation vectors of neighboring objects are obtained. Based on the representation vector of the target object and the set of representation vectors of neighboring objects, the representation feature information of the target object is obtained. The representation feature information is determined based on the relationship feature information corresponding to each of the K first relationship types between the target object and its neighboring objects. A multimedia resource set is obtained, and based on the representation feature information of the target object, the first multimedia resource in the multimedia resource set is recommended to the target object. It is evident that by exploring the relationship types between the target object and its neighboring objects, the interest points of the target object are mined, thereby significantly improving the recommendation accuracy.

[0081] Please see Figure 2 , Figure 2 This is a flowchart illustrating a method for recommending multimedia resources provided in an embodiment of this application. The method described in this application is applied to a smart device, which may be, for example, a terminal device used by some of the aforementioned objects, or a server with special functions. The method includes the following steps.

[0082] S201: Obtain the representation vector of the target object and the set of representation vectors of neighboring objects. The set of representation vectors of neighboring objects includes the representation vectors of the first neighboring objects of the target object. The first neighboring object refers to an object that interacts with the target object. Similarly, the second neighboring object refers to an object that interacts with the first neighboring object of the target object but does not interact with the target object.

[0083] In one implementation, the server obtains the representation vector of the target object and the set of representation vectors of neighboring objects based on an association information network graph constructed based on the interaction behavior of various objects on a social platform. The target object and its first neighboring object are divided into K first relationship types, where K is a positive integer. It should be noted that the division into K first relationship types can be set according to actual needs; for example, the K first relationship types can be divided based on association relationships, the cumulative duration of interaction behavior, or the time of the first interaction between the target object and its first neighboring object, etc.

[0084] S202: Based on the representation vector of the target object and the set of representation vectors of its neighboring objects, obtain the representation feature information of the target object. The representation feature information is determined based on the relation feature information corresponding to each of the K first relation types for the target object.

[0085] The representation vector of a target object is used to represent the features of the target object. In one embodiment, the representation vector of a target object can be obtained based on its own object feature information, or based on the representation vectors of its first neighboring objects, or based on the representation vectors of its first to Sth neighboring objects, where S is a positive integer. It can be understood that S is proportional to the amount of object feature information carried in the representation vector of the target object. Similarly, in the set of neighboring object representation vectors, the representation vectors of the first neighboring objects of each target object can be obtained based on their own object feature information, or based on the representation vectors of the first neighboring objects of that first neighboring object, or based on the representation vectors of the first neighboring objects of that first neighboring object to the Sth neighboring object.

[0086] The representational feature information of the target object can be either a feature vector or a feature matrix, carrying the object feature information of the target object. In one implementation, the server obtains the relational feature information corresponding to each of the K first relation types based on the target object's representation vector and the set of neighboring object representation vectors (e.g., calculating the relational feature information corresponding to the first relation type using the target object's representation vector and the neighboring object representation vectors). After obtaining the relational feature information corresponding to each of the K first relation types, the server represents the target object using the K relational feature information corresponding to each of the K first relation types, thus obtaining the target object's representational feature information. In other words, the target object's representational feature information is jointly represented by the K relational feature information corresponding to each of the K first relation types (i.e., the target object's representational feature information is obtained based on the relation types between the target object and its neighboring objects).

[0087] S203: Obtain the multimedia resource set and, based on the representational characteristics of the target object, recommend the first multimedia resource from the set to the target object. The multimedia resource set can be preset or obtained in real time by the multimedia resource platform based on multimedia resources in the database.

[0088] In one implementation, the resource feature information of each multimedia resource in the multimedia resource set is obtained by examining the representation vector of the object of that multimedia resource. Based on the representation feature information of the target object and the resource feature information of each multimedia resource in the multimedia resource set, the server recommends a first multimedia resource in the multimedia resource set to the target object. The first multimedia resource is one or more multimedia resources in the multimedia resource set whose representation feature information matches the target object with a matching degree higher than a matching degree threshold; that is, the multimedia resources that the target object is most likely to be interested in.

[0089] In this embodiment, the representation vector of the target object and the set of representation vectors of neighboring objects are obtained. Based on the representation vector of the target object and the set of representation vectors of neighboring objects, the representation feature information of the target object is obtained. The representation feature information is determined based on the relationship feature information corresponding to each of the K first relationship types between the target object and its neighboring objects. A multimedia resource set is obtained, and based on the representation feature information of the target object, the first multimedia resource in the multimedia resource set is recommended to the target object. It is evident that by exploring the relationship types between the target object and its neighboring objects, the interest points of the target object are mined, thereby significantly improving the recommendation accuracy.

[0090] Please see Figure 3 , Figure 3 This is a flowchart illustrating another method for recommending multimedia resources provided in an embodiment of this application. The method described in this application is applied to a smart device, which may be, for example, a terminal device used by some of the aforementioned objects, or a server with special functions. The method includes the following steps.

[0091] S301: Obtain the set of relationship information and generate a relationship information network diagram based on the set of relationship information. The set of relationship information includes a set of object information and a set of relationship information.

[0092] In one implementation, the smart device generates N network nodes based on an object information set. Each of the N network nodes corresponds to an object, and each network node carries object information related to the object it corresponds to, where N is a positive integer. The edges between the network nodes are determined based on the interaction behavior between the objects corresponding to each network node. If the relationship information set indicates that the object corresponding to the first network node and the object of the second network node have interaction behavior, then the edge information between the first network node and the second network node is generated based on the interaction behavior, resulting in an association information network graph.

[0093] Furthermore, the intelligent device can determine the weight of each edge in the network graph of the associated information based on the set of relational information. Specifically, the edge weights in the edge information of the first network node and the second network node include: weights determined based on the degree of association between the first network node and the second network node. The edge weights are proportional to the degree of association, which is determined based on the interaction information of the first network node and the second network node within a target time period. The interaction information includes at least one of the following: cumulative number of interactions, cumulative interaction duration, and interaction frequency.

[0094] In one embodiment, an object association information network graph G = (A, X) is first defined. The total number of objects is N, A is the object association matrix, and X represents object feature information. Typically, the object association matrix A needs to include various information such as the number of chats, chat frequency, and interaction frequency. Based on the historical behavior of objects, the smart device connects the network nodes corresponding to each object to form the object association information network graph. The edges between objects are determined by the degree of object association. If two objects have many interactions, the connection between them has a higher weight; if they have few interactions, the connection has a lower weight. If two objects have no interactions, there is no connection between them. The measurement of object interaction behavior can be determined by variables such as the number of object interactions, the cumulative duration of interactions, the ranking of interaction frequencies, and the number of interaction days within a target time period.

[0095] Figure 4a This is a schematic diagram of an object association information network diagram provided in an embodiment of this application. For example... Figure 4a As shown, if the objects corresponding to network nodes u1 and u2 interact more frequently, while the objects corresponding to network nodes u1 and u3 interact less frequently, then the connection weight between network nodes u1 and u2 is higher than the connection weight between network nodes u1 and u3. If the object corresponding to network node u1 has no interaction with any other objects besides those corresponding to network nodes u2 and u3, then u1 has no connection with any other network nodes. Furthermore, to better describe the degree of association between object i and object j, the number of interactions between the two objects can be denoted as c. ij The relationship between object i and object j can be expressed as: log(1+c ij In other words, in the object association matrix A, A ij =log(1+c ij ).

[0096] Furthermore, after obtaining the association information network graph, the intelligent device obtains the representation vectors of N objects corresponding to N network nodes (i.e., describing each object through vectors) based on the association information network graph. The purpose is to describe the interaction behavior of objects through vectors, making the vector representations of objects with close relationships similar; correspondingly, the vector representations of objects with distant relationships differ significantly. Specifically, based on the association information network graph, a random walk is performed in the association information network graph starting from the target network node corresponding to the target object, resulting in M ​​trajectories, each with a step size of P; where M and P are both positive integers; based on the object information carried in the M trajectories, the representation vector of the target object is obtained; where the probability of walking from the i-th network node to the j-th network node is proportional to the target edge weight, the target edge weight is the edge weight in the edge information between the i-th and j-th network nodes, where i and j are both positive integers, i not equal to j, and i and j are both less than or equal to N. That is to say, A ij The larger the value, the greater the probability of traversing from network node i to network node j.

[0097] In one embodiment, smart devices represent objects using methods such as vectorized embedding. Commonly used vectorized embedding methods include unsupervised object embedding methods such as Node2Vec embedding. Taking Node2Vec embedding as an example, based on an association information network graph, multiple random walks are performed starting from the target network node in the graph. Subsequently, all the walked paths are used as a corpus and input into the word2vec word vector embedding algorithm model. The word2vec word vector embedding algorithm model processes the corpus to obtain the representation vector of the target object corresponding to the target network node. Since the connection weights between network nodes corresponding to different objects in the graph are different, the influence of weights can be considered during the vectorized embedding process, using weighted random walks (i.e., the probability of walking from network node i to network node j is proportional to A). ij (Proportional). Similarly, following the above method, the intelligent device can obtain a matrix composed of the representation vectors of all objects corresponding to all nodes in the associated information network graph, denoted as X, where X = {x1, x2, ..., x...} N}. Where x i This represents the representation vector of the i-th object.

[0098] S302: Obtain the representation vector of the target object and the set of representation vectors of adjacent objects.

[0099] In one implementation, after obtaining a matrix X consisting of the representation vectors of all objects corresponding to all nodes in the associated information network graph, the intelligent device can obtain the representation vector of the target object and the set of representation vectors of adjacent objects from the matrix X.

[0100] S303: Based on the representation vector of the target object and the set of representation vectors of adjacent objects, obtain the relation feature information corresponding to each of the K first relation types.

[0101] Figure 4b This is a schematic diagram of a graph convolution model based on a relational information network graph, provided as an embodiment of this application. For example... Figure 4b As shown, V0 is the target network node corresponding to the target object, and V1-V8 are the network nodes corresponding to the first adjacent objects of the target object. In one implementation, V1-V8 have the same aggregation weight and the same mapping function. Simply put, the correlation between V1-V8 and V0 is not considered, and it is assumed that V1-V8 have the same influence on V0.

[0102] In another implementation, since V1-V8 are objects that V0 has established relationships with through different methods, their influence on the object varies (e.g., adjacent objects with higher interaction frequency have a higher influence on the target object). Therefore, it cannot be simply assumed that V1-V8 have the same influence on V0. To address this, different objects need to be categorized (e.g., objects are divided into multiple categories based on factors such as the reason for establishing the connection, the type of relationship, the cumulative duration of the relationship, and the frequency of interaction). In practice, it has been found that it is often difficult to directly obtain the reasons for the formation of relationships in real-world applications.

[0103] In one embodiment, the smart device divides the first adjacent objects of the target object into K first relation types according to a preset rule, and obtains the feature parameter set of the h-th first relation type among the K first relation types. The feature parameter set includes the weight matrix of the h-th first relation type and the bias vector of the h-th first relation type (e.g., the weight matrix W of the h-th first relation type is initialized according to a preset rule). h and the bias vector b of the h-th first relation type h For example, the weight matrix W for the h-th first relation type. h and the bias vector b of the h-th first relation type h Random initialization is performed, and during training, gradient descent is used to adjust the weight matrix W of the h-th first relation type. h and the bias vector b of the h-th first relation type h The update is performed, ultimately yielding the updated weight matrix W for the h-th first relation type. h and the bias vector b of the h-th first relation type h Similarly, intelligent devices can obtain the weight matrix and bias vector for each first relation type based on the above method.

[0104] Furthermore, using the feature parameter set of the h-th first relation type, the intermediate features (i.e., the implicit representation of the target object) of the target object under the h-th first relation type are calculated, and the intermediate features (i.e., the implicit representations of the first adjacent objects) of each first adjacent object of the target object under the h-th first relation type are also calculated. The implicit representation z of object i in the h-th first relation type is... i,h It can be represented as:

[0105]

[0106] Where ||x||2 represents the calculation of the magnitude of x, and the division by the magnitude is to eliminate the influence of vector length on classification; σ(x) is the activation function (such as sigmoid function, tanh function, ReLU function, etc.). W is the weight matrix of the h-th first relation type. h The transpose of x i Let b be the representation vector of object i (obtained from matrix X in step S301). h Let h be the bias vector of the h-th first relation type. Based on Formula 1 above, the intelligent device can calculate the implicit representation of the target object and its first neighboring objects.

[0107] In one embodiment, after calculating the implicit representation of the target object and the implicit representation of the first neighboring objects of the target object, the intelligent device obtains the relation feature information c of the h-th first relation type based on the implicit representation of the target object and the implicit representation of the first neighboring objects of the target object. h Specifically, it can be expressed as:

[0108]

[0109] Wherein, network node u corresponds to Figure 4b V0 (i.e., the network node corresponding to the target object), (v|(u,v)∈G) corresponds to Figure 4b In the context of V1-V8 (i.e., the network nodes corresponding to the first adjacent objects of the target object), p v,h p represents the probability that object v is assigned to the h-th first relation type. v,h ≥0, and z u,h z is the implicit representation of the target object in the h-th first relation type. v,h p is the implicit representation of the first neighboring object of the target object in the h-th first relation type. v,h The value is determined by the number of first relation types; in one specific implementation... For example, suppose the number of first relation types for object A is 5 (i.e., the associations of object A are divided into 5 categories), then Based on Formula 2 above, the intelligent device can calculate the relation feature information of K first relation types.

[0110] In another embodiment, the relation feature information of the h-th first relation type is the relation feature information of the h-th first relation type at the T-th iteration, where T is a positive integer. The intelligent device obtains the target probability that each of the target object's first neighboring objects is classified into the h-th first relation type at the t-th iteration. t is a positive integer, and t is less than T; then the intelligent device determines the target probability. The intermediate features between each neighboring object of the target object (i.e., the implicit representation of the first neighboring object (z)) v,h )) Calculate the aggregation features of the first neighboring objects of the target object. For the intermediate features of the target object (i.e., the implicit representation of the target object (z) u,h The aggregation features of the target object and its first neighboring object are processed to obtain the relation feature information of the h-th first relation type at the (t+1)-th iteration. Specifically:

[0111]

[0112] Wherein, the exponential function exp(x) represents the calculation of the exponent of x. For z v,h The transpose of the matrix, This represents the relation feature information of the h-th first relation type at the t-th iteration. Further, It can be represented as:

[0113]

[0114] Wherein, network node u corresponds to Figure 4b V0 (i.e., the network node corresponding to the target object), (v|(u,v)∈G) corresponds to Figure 4b V1-V8 (i.e., the network nodes corresponding to the first adjacent objects of the target object), z is used to represent the probability that object v is assigned to the h-th first relation type at the (t-1)-th iteration. u,h z is the implicit representation of the target object in the h-th first relation type. v,h Let be the implicit representation of the first neighboring object of the target object in the h-th first relation type. Here, is the initial value of the probability that object v is assigned to the h-th first relation type. It is determined by the number of first relation types; in one specific implementation... For example, suppose the number of first relation types for object A is 5 (i.e., the associations of object A are divided into 5 categories), then Based on the iterative calculations using Formulas 3 and 4 above, the intelligent device can calculate the relational feature information of K first relation types (the relational feature information of each first relation type at the Tth iteration). Practice has shown that optimizing the probability of object v being assigned to the h-th first relation type through T iterations can distinguish the impact of different first relation types on object v, thus significantly improving the accuracy of multimedia resource recommendations.

[0115] S304: Represent the target object by using the K relation feature information corresponding to the K first relation types respectively, and obtain the representation feature information of the target object.

[0116] In one implementation, the smart device determines the relation feature information of each first relation type at the Tth iteration as the relation feature information of that first relation type (e.g., This allows us to obtain the representational feature information y of the target object. u =[c1,c2,…,c K ].

[0117] In another implementation, the smart device acquires a first feature information set of the first neighboring objects of the target object and a second feature information set of the second neighboring objects of the target object. The first feature information set includes relation feature information corresponding to each of the multiple second relation types of the first neighboring objects of the target object (e.g., [c′1, c′2, ..., c′...). R The second feature information set includes the relation feature information corresponding to each of the multiple third relation types of the second adjacent objects of the target object (such as [c″1, c″2, ..., c″). S R and S are positive integers, and R, S, and K can be the same or different (when they are the same, each object and its first neighboring object are divided into K first relation types; when they are different, each object and its first neighboring object are divided into different numbers of first relation types). The specific implementation of the intelligent device acquiring the first feature information set of the first neighboring objects of the target object and the second feature information set of the second neighboring objects of the target object can be found in steps S301-S303, and will not be repeated here. The representation feature information of the first neighboring objects of the target object can be obtained through the first feature information set. Similarly, the representational feature information of the second neighboring objects of the target object can be obtained through the second feature information set. The representational feature information of the Lth neighboring objects of the target object can be obtained from the Lth feature information set.

[0118] The relationship prediction model takes the relationship feature information corresponding to the K first relationship types of the target object and the first feature information set to the Lth feature information set as input to obtain the prediction result output by the relationship prediction model. The relationship prediction model includes L graph convolutional network layers, where L is a positive integer. Overfitting is performed on the prediction result to obtain the representation feature information of the target object. Among them, the input data of the gth graph convolutional network layer in the relationship prediction model includes the data obtained after overfitting the output data of the (g-1)th graph convolutional network layer.

[0119] Specifically, the input to the l-th layer of the relationship prediction model is the processing result of the target object in the (l-1)-th layer. And the set of processing results of the first adjacent object of the target object at level l-1. Let l be a positive integer, and l is less than or equal to L. Processing the input data using the l-th layer of the relation prediction model can be represented as:

[0120]

[0121] Furthermore, overfitting is applied to the output data of the l-th layer of the relationship prediction model to obtain the processing result of the target object at the l-th layer.

[0122]

[0123] Among them, f (l) (x) represents processing x through the l-th layer of the relation prediction model, and dropout(x) represents overfitting x. The value is initialized to x u The processing result of the target object at level L. It can be represented as:

[0124]

[0125] It should be noted that U u With y u The difference is that y u It is obtained through the representation vector of the target object and the representation vector of the first neighboring object of the target object, U. u It is obtained by using the representation vector of the target object and the representation vectors of the target object's first neighboring objects to the Lth neighboring objects (relative to y). u It covers more feature information.

[0126] S305: Obtain a set of multimedia resources and recommend the first multimedia resource in the set to the target object based on the representation feature information of the target object.

[0127] In one implementation, the multimedia resource set includes resource feature information of a second multimedia resource. The smart device acquires viewing information of the second multimedia resource, which includes object identifiers of Q objects that have viewed the second multimedia resource, where Q is a positive integer. Based on the object identifiers of the Q objects, it acquires representation vectors of the Q objects (e.g., from matrix X in step S301). It performs a fusion process on the representation vectors of the Q objects (e.g., superimposing the representation vectors of the Q objects), and performs mean pooling on the result of the fusion process to obtain the resource feature information of the second multimedia resource. Specifically, it assumes that the list of objects that have viewed the second multimedia resource is (u1, u2, u3, u4, ..., uH), where H is the total number of all objects that have viewed the video. If the resource feature information of the second multimedia resource is represented as i... m ,but:

[0128]

[0129] in, For object u h The representation vector. Similarly, based on Formula 8 above, smart devices can obtain the resource feature information of all multimedia resources in the multimedia resource set. This resource feature information can be recorded in matrix I, I = {i1, i2, ..., i...} N}

[0130] Furthermore, the intelligent device calculates the matching degree between the target object and each multimedia resource in the multimedia resource set based on the representation feature information of the target object and the resource feature information of each multimedia resource in the multimedia resource set.

[0131] In one implementation, the intelligent device can predict the relationship between each first relation type and multimedia resource i using a multi-layer percetron (MLP), and then use an attention mechanism to comprehensively consider the preferences of different first relation types to finally obtain the prediction results of the target object u and multimedia resource i.

[0132] The intelligent device concatenates the resource feature information of the second multimedia resource in the multimedia resource set with the representation feature information of the target object to obtain a concatenated feature set (that is, the resource feature information of the second multimedia resource and [c u,1 ,c u,2 ,…,c u,KThe relation feature information corresponding to each first relation type in [ ] is concatenated to obtain K concatenated features. A multilayer perceptron is then used to process each concatenated feature in the concatenated feature set to obtain the relation vector between each of the K first relation types and the second multimedia resource. The relation vector between the kth first relation type of the target object and multimedia resource i can be represented as:

[0133] r u,i,k =MLP1(c u,k ||I i ) Formula 9

[0134] Where x||y represents concatenating vectors x and y, and MLP1(x) represents processing x using a first multilayer perceptron. Based on the above formula 9, the intelligent device can process each concatenated feature in the concatenated feature set through the first multilayer perceptron to obtain the relationship vector between each first relation type and each multimedia resource.

[0135] After obtaining the relationship vectors between each first relationship type and the second multimedia resource, the intelligent device calculates the weight corresponding to each first relationship type based on the relationship vector between each first relationship type and the second multimedia resource; the weight of the k-th first relationship type of the target object and multimedia resource i can be expressed as:

[0136]

[0137] Where, the exponential function exp(x) represents the calculation of the exponent of x, σ(x) is the activation function (such as the sigmoid function), and a T Let a be the attention vector. T r u,i,k This indicates that the relationship vector and the attention vector are multiplied by a dot product. Based on Equation 10 above, the intelligent device can calculate the weights between each first relationship type of the target object and each multimedia resource.

[0138] After obtaining the weights between each first relation type of the target object and the second multimedia resource, the intelligent device calculates the matching degree between the target object and the second multimedia resource based on the relation vector between each of the K first relation types and the second multimedia resource, and the weight corresponding to each first relation type. The matching degree between the target object and multimedia resource i can be expressed as:

[0139]

[0140] Where MLP2(x) represents the processing of x using a second multilayer perceptron. Based on the above formula 11, the intelligent device can obtain the matching degree between the target object and each multimedia resource in the multimedia resource set through the second multilayer perceptron.

[0141] Furthermore, the smart device sorts the multimedia resources in the multimedia resource set according to their matching degree from high to low, and identifies one or more multimedia resources that appear before the target position as the first multimedia resource. The first multimedia resource is then recommended to the target object. In one embodiment, the first multimedia resource is the multimedia resource in the multimedia resource set that has the highest matching degree with the target object.

[0142] In another implementation, before recommending multimedia resources, the intelligent device can optimize the parameters in Formulas 1-11 using training data (i.e., compare the labeled data with the predicted data calculated by Formulas 1-11, and adjust the parameters in Formulas 1-11 using a loss function to reduce the difference between the labeled data and the predicted data until the loss function converges). After training, the intelligent device obtains the representation vectors of all objects. When a multimedia resource acquisition request for the target object u is detected, steps S301-S305 are executed to compare the similarity between the target object u and each multimedia resource in the multimedia resource set, and then multimedia resources that meet the recommendation criteria are recommended to the target object.

[0143] The embodiments of this application are as follows: Figure 2 Based on the previous implementation, a network graph of association information is constructed using a set of association relationship information, thereby obtaining the representation vector of the target object and a set of representation vectors of neighboring objects. Through the implicit representations of the target object and its first neighboring objects, the relational feature information corresponding to each of the first relation types is obtained, thus yielding the representational feature information of the target object. By analyzing the representation vectors of objects that have viewed multimedia resources, the resource feature information of the multimedia resources is obtained. Based on the representational feature information of the target object and the resource feature information of the multimedia resources, multimedia resources are recommended to the target object. It is evident that by exploring the relationship types between the target object and its neighboring objects, the target object's points of interest are mined, thereby significantly improving recommendation accuracy.

[0144] The methods of the embodiments of this application have been described in detail above. In order to facilitate better implementation of the above solutions of the embodiments of this application, the apparatus of the embodiments of this application is provided below.

[0145] Please see Figure 5 , Figure 5 This is a schematic diagram of a multimedia resource recommendation device provided in an embodiment of this application. The multimedia resource recommendation device 500 includes an acquisition unit 501 and a processing unit 502. The device can be mounted on a smart device, which may include a terminal device or a server. Figure 5 The multimedia resource recommendation device shown can be used to perform the above. Figure 2and Figure 3 The described method embodiments include some or all of the functionalities. The detailed descriptions of each unit are as follows:

[0146] The acquisition unit 501 is used to acquire the representation vector of the target object and the set of representation vectors of neighboring objects. The set of representation vectors of neighboring objects includes the representation vectors of the first neighboring objects of the target object. The target object and the first neighboring objects are divided into K first relation types, where K is a positive integer.

[0147] Processing unit 502 is configured to obtain representation feature information of the target object based on the representation vector of the target object and the set of representation vectors of neighboring objects, wherein the representation feature information is determined based on the relation feature information corresponding to each of the K first relation types of the target object; and to obtain a set of multimedia resources and recommend a first multimedia resource in the set of multimedia resources to the target object based on the representation feature information of the target object.

[0148] In one embodiment, the processing unit 502 is specifically used for:

[0149] Based on the representation vector of the target object and the set of representation vectors of the adjacent objects, the relation feature information corresponding to each of the K first relation types is obtained;

[0150] The target object is analyzed by performing feature analysis on the K relation feature information corresponding to the K first relation types respectively, and the representation feature information of the target object is obtained.

[0151] In one embodiment, the processing unit 502 is specifically used for:

[0152] Obtain the feature parameter set of the h-th first relation type among the K first relation types, the feature parameter set including the weight matrix and the bias vector of the h-th first relation type;

[0153] Using the feature parameter set of the h-th first relation type, calculate the target object intermediate features under the h-th first relation type, and calculate the intermediate features of each first neighboring object of the target object under the h-th first relation type;

[0154] Based on the intermediate features of the target object and the intermediate features of each adjacent object, the relation feature information of the h-th first relation type is obtained.

[0155] In one embodiment, the relation feature information of the h-th first relation type is the relation feature information of the h-th first relation type at the T-th iteration, where T is a positive integer; the processing unit 502 is specifically used for:

[0156] Obtain the target probability that each of the first neighboring objects of the target object is classified into the h-th first relation type at the t-th iteration, where t is a positive integer and t is less than T;

[0157] Based on the target probability and the intermediate features of each neighboring object, calculate the aggregate features of the first neighboring object of the target object;

[0158] The intermediate features of the target object and the aggregate features of the first neighboring objects of the target object are processed to obtain the relation feature information of the h-th first relation type at the (t+1)-th iteration.

[0159] In one embodiment, the processing unit 502 is specifically used for:

[0160] Obtain a first feature information set of the first neighboring objects of the target object, and a second feature information set of the second neighboring objects of the target object. The first feature information set includes: relationship feature information corresponding to each of the multiple second relationship types of the first neighboring objects of the target object. The second feature information set includes: relationship feature information corresponding to each of the multiple third relationship types of the second neighboring objects of the target object.

[0161] The relationship feature information corresponding to the K first relationship types of the target object, the first feature information set, and the second feature information set are used as inputs to the relationship prediction model to obtain the prediction result output by the relationship prediction model. The relationship prediction model includes L layers of graph convolutional network layers, where L is a positive integer.

[0162] Overfitting is applied to the prediction results to obtain the representational feature information of the target object;

[0163] In the relationship prediction model, the input data of the g-th graph convolutional network layer includes the data obtained after overfitting the output data of the (g-1)-th graph convolutional network layer.

[0164] In one embodiment, the processing unit 502 is specifically used for:

[0165] Obtain viewing information of the second multimedia resource, wherein the viewing information includes object identifiers of Q objects that have viewed the second multimedia resource, where Q is a positive integer;

[0166] Based on the object identifiers of the Q objects, obtain the representation vectors of the Q objects;

[0167] The representation vectors of the Q objects are fused, and the result of the fusion is averaged and pooled to obtain the resource feature information of the second multimedia resource.

[0168] The multimedia resource set includes the resource feature information of the second multimedia resource.

[0169] In one embodiment, the processing unit 502 is specifically used for:

[0170] Based on the representation feature information of the target object and the resource feature information of each multimedia resource in the multimedia resource set, the matching degree between the target object and each multimedia resource in the multimedia resource set is calculated.

[0171] A first multimedia resource is recommended to the target object, wherein the first multimedia resource is the multimedia resource in the multimedia resource set that has the highest matching degree with the target object.

[0172] In one embodiment, the processing unit 502 is specifically used for:

[0173] The resource feature information of the second multimedia resource in the multimedia resource set is concatenated with the representation feature information of the target object to obtain a concatenated feature set;

[0174] A multilayer perceptron is used to process each splicing feature in the splicing feature set to obtain the relationship vector between each of the K first relationship types and the second multimedia resource;

[0175] Calculate the weight corresponding to each first relation type based on the relation vector between each first relation type and the second multimedia resource;

[0176] Based on the relationship vector between each of the K first relationship types and the second multimedia resource, and the weight corresponding to each first relationship type, the matching degree between the target object and the second multimedia resource is obtained.

[0177] In one embodiment, the processing unit 502 is further configured to:

[0178] Obtain a set of association information, which includes a set of object information and a set of relationship information;

[0179] N network nodes are generated based on the object information set. Each of the N network nodes corresponds to an object, and each network node carries object information of the object corresponding to that network node. N is a positive integer.

[0180] If the set of relational information indicates that the object corresponding to the first network node and the object of the second network node among the N network nodes have interactive behavior, then the connection information between the first network node and the second network node is generated according to the interactive behavior to obtain the relational information network graph.

[0181] Based on the associated information network graph, the representation vectors of the N objects corresponding to the N network nodes are obtained.

[0182] In one embodiment, the edge weight in the edge information of the first network node and the second network node includes: a weight determined based on the correlation between the first network node and the second network node, wherein the edge weight is proportional to the correlation, and the correlation is determined based on the interaction information of the first network node and the second network node within a target time period, wherein the interaction information includes at least one of the following: cumulative number of interactions, cumulative interaction duration, interaction frequency, and interaction content.

[0183] In one embodiment, the processing unit 502 is specifically used for:

[0184] Starting from the target network node corresponding to the target object, a random walk is performed in the associated information network graph to obtain M trajectories, each with a step size of P; where M and P are both positive integers.

[0185] Based on the object information carried in the M trajectories, the representation vector of the target object is obtained;

[0186] The probability of traversing from the i-th network node to the j-th network node is proportional to the target edge weight, which is the edge weight in the edge information of the i-th network node and the j-th network node. i and j are both positive integers, i is not equal to j, and i and j are both less than or equal to N.

[0187] According to one embodiment of this application, Figure 2 and Figure 3 The steps involved in the recommended method for multimedia resources shown can be derived from... Figure 5 The recommendation of multimedia resources is performed by each unit in the illustrated device. For example, Figure 2 Step S201 shown can be performed by Figure 5 The acquisition unit 501 shown is executed, and steps S202 and S203 can be performed by... Figure 5 The processing unit 502 shown is executed. Figure 3 Steps S301 and S302 shown can be derived from... Figure 5 The acquisition unit 501 shown is executed, and steps S303-S305 can be performed by... Figure 5 The processing unit 502 shown is executed. Figure 5 The various units in the multimedia resource recommendation device shown can be individually or entirely merged into one or more other units, or some of the units can be further divided into multiple functionally smaller units. This achieves the same operation without affecting the technical effects of the embodiments of this application. The above units are based on logical function division. In practical applications, the function of one unit can also be implemented by multiple units, or the function of multiple units can be implemented by one unit. In other embodiments of this application, the multimedia resource recommendation device may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented collaboratively by multiple units.

[0188] According to another embodiment of this application, the following can be executed by running on a general-purpose computing device, such as a computer, which includes processing elements and storage elements such as a central processing unit (CPU), random access memory (RAM), and read-only memory (ROM). Figure 2 and Figure 3 The computer program (including program code) for each step involved in the corresponding method shown, to construct such... Figure 5 The diagram illustrates a multimedia resource recommendation apparatus and a method for implementing multimedia resource recommendation embodiments of this application. The computer program may be recorded on, for example, a computer-readable recording medium, loaded onto the aforementioned computing device via the same medium, and executed therein.

[0189] Based on the same inventive concept, the principle and beneficial effects of the multimedia resource recommendation device provided in the embodiments of this application are similar to those of the multimedia resource recommendation device in the method embodiments of this application. Please refer to the principle and beneficial effects of the method implementation. For the sake of brevity, they will not be repeated here.

[0190] Please see Figure 6 , Figure 6This is a schematic diagram of the structure of a smart device provided in an embodiment of this application. The smart device 600 includes at least a processor 601, a communication interface 602, and a memory 603. The processor 601, communication interface 602, and memory 603 can be connected via a bus or other means. The processor 601 (or Central Processing Unit, CPU) is the computing and control core of the terminal. It can parse various instructions within the terminal and process various data. For example, the CPU can parse power-on / off commands sent by an object to the terminal and control the terminal to perform power-on / off operations; it can also transmit various interactive data between internal structures of the terminal, and so on. The communication interface 602 may optionally include a standard wired interface or a wireless interface (such as Wi-Fi, mobile communication interface, etc.), and can be used to send and receive data under the control of the processor 601; the communication interface 602 can also be used for data transmission and interaction within the terminal. The memory 603 is a memory device in the terminal used to store programs and data. It is understood that the memory 603 here can include the terminal's built-in memory, or it can include extended memory supported by the terminal. The memory 603 provides storage space for storing the terminal's operating system, which may include, but is not limited to, Android, iOS, Windows Phone, etc. This application does not limit this.

[0191] In this embodiment of the application, the processor 601 performs the following operations by running executable program code in the memory 603:

[0192] The target object's representation vector and a set of neighboring object representation vectors are obtained through the communication interface 602. The set of neighboring object representation vectors includes the representation vectors of the first neighboring object of the target object. The target object and the first neighboring object are divided into K first relation types, where K is a positive integer.

[0193] Based on the representation vector of the target object and the set of representation vectors of neighboring objects, the representation feature information of the target object is obtained. The representation feature information is determined based on the relation feature information corresponding to each of the K first relation types of the target object.

[0194] Obtain a set of multimedia resources, and recommend a first multimedia resource from the set to the target object based on the representational feature information of the target object.

[0195] As an optional embodiment, the processor 601 obtains the representation feature information of the target object based on the representation vector of the target object and the set of representation vectors of neighboring objects in the following specific embodiment:

[0196] Based on the representation vector of the target object and the set of representation vectors of the adjacent objects, the relation feature information corresponding to each of the K first relation types is obtained;

[0197] The target object is analyzed by performing feature analysis on the K relation feature information corresponding to the K first relation types respectively, and the representation feature information of the target object is obtained.

[0198] As an optional embodiment, the processor 601 obtains the relation feature information corresponding to each of the K first relation types based on the representation vector of the target object and the set of representation vectors of neighboring objects. A specific embodiment of this is as follows:

[0199] Obtain the feature parameter set of the h-th first relation type among the K first relation types, the feature parameter set including the weight matrix and the bias vector of the h-th first relation type;

[0200] Using the feature parameter set of the h-th first relation type, calculate the target object intermediate features under the h-th first relation type, and calculate the intermediate features of each first neighboring object of the target object under the h-th first relation type;

[0201] Based on the intermediate features of the target object and the intermediate features of each adjacent object, the relation feature information of the h-th first relation type is obtained.

[0202] As an optional embodiment, the relation feature information of the h-th first relation type is the relation feature information of the h-th first relation type at the T-th iteration, where T is a positive integer; the specific embodiment in which the processor 601 obtains the relation feature information of the h-th first relation type based on the intermediate features of the target object and the intermediate features of each adjacent object is as follows:

[0203] Obtain the target probability that each of the first neighboring objects of the target object is classified into the h-th first relation type at the t-th iteration, where t is a positive integer and t is less than T;

[0204] Based on the target probability and the intermediate features of each neighboring object, calculate the aggregate features of the first neighboring object of the target object;

[0205] The intermediate features of the target object and the aggregate features of the first neighboring objects of the target object are processed to obtain the relation feature information of the h-th first relation type at the (t+1)-th iteration.

[0206] As an optional embodiment, the processor 601 performs feature analysis on the target object using the K relation feature information corresponding to the K first relation types, and obtains the representation feature information of the target object in the following specific embodiment:

[0207] Obtain a first feature information set of the first neighboring objects of the target object, and a second feature information set of the second neighboring objects of the target object. The first feature information set includes: relationship feature information corresponding to each of the multiple second relationship types of the first neighboring objects of the target object. The second feature information set includes: relationship feature information corresponding to each of the multiple third relationship types of the second neighboring objects of the target object.

[0208] The relationship feature information corresponding to the K first relationship types of the target object, the first feature information set, and the second feature information set are used as inputs to the relationship prediction model to obtain the prediction result output by the relationship prediction model. The relationship prediction model includes L layers of graph convolutional network layers, where L is a positive integer.

[0209] Overfitting is applied to the prediction results to obtain the representational feature information of the target object;

[0210] In the relationship prediction model, the input data of the g-th graph convolutional network layer includes the data obtained after overfitting the output data of the (g-1)-th graph convolutional network layer.

[0211] As an optional embodiment, the processor 601 acquires the multimedia resource set in the following specific embodiment:

[0212] Obtain viewing information of the second multimedia resource, wherein the viewing information includes object identifiers of Q objects that have viewed the second multimedia resource, where Q is a positive integer;

[0213] Based on the object identifiers of the Q objects, obtain the representation vectors of the Q objects;

[0214] The representation vectors of the Q objects are fused, and the result of the fusion is averaged and pooled to obtain the resource feature information of the second multimedia resource.

[0215] The multimedia resource set includes the resource feature information of the second multimedia resource.

[0216] As an optional embodiment, the processor 601 recommends a first multimedia resource from the multimedia resource set to the target object based on the representation feature information of the target object. A specific embodiment of this is as follows:

[0217] Based on the representation feature information of the target object and the resource feature information of each multimedia resource in the multimedia resource set, the matching degree between the target object and each multimedia resource in the multimedia resource set is calculated.

[0218] A first multimedia resource is recommended to the target object, wherein the first multimedia resource is the multimedia resource in the multimedia resource set that has the highest matching degree with the target object.

[0219] As an optional embodiment, the processor 601 calculates the matching degree between the target object and each multimedia resource in the multimedia resource set based on the representation feature information of the target object and the resource feature information of each multimedia resource in the multimedia resource set. A specific embodiment of this is as follows:

[0220] The resource feature information of the second multimedia resource in the multimedia resource set is concatenated with the representation feature information of the target object to obtain a concatenated feature set;

[0221] A multilayer perceptron is used to process each splicing feature in the splicing feature set to obtain the relationship vector between each of the K first relationship types and the second multimedia resource;

[0222] Calculate the weight corresponding to each first relation type based on the relation vector between each first relation type and the second multimedia resource;

[0223] Based on the relationship vector between each of the K first relationship types and the second multimedia resource, and the weight corresponding to each first relationship type, the matching degree between the target object and the second multimedia resource is obtained.

[0224] As an optional embodiment, the processor 601, by running executable program code in the memory 603, also performs the following operations:

[0225] Obtain a set of association information, which includes a set of object information and a set of relationship information;

[0226] N network nodes are generated based on the object information set. Each of the N network nodes corresponds to an object, and each network node carries object information of the object corresponding to that network node. N is a positive integer.

[0227] If the set of relational information indicates that the object corresponding to the first network node and the object of the second network node among the N network nodes have interactive behavior, then the connection information between the first network node and the second network node is generated according to the interactive behavior to obtain the relational information network graph.

[0228] Based on the associated information network graph, the representation vectors of the N objects corresponding to the N network nodes are obtained.

[0229] As an optional embodiment, the edge weight in the edge information of the first network node and the second network node includes: a weight determined according to the correlation between the first network node and the second network node, wherein the edge weight is proportional to the correlation, and the correlation is determined according to the interaction information of the first network node and the second network node within a target time period, wherein the interaction information includes at least one of the following: cumulative interaction count, cumulative interaction duration, interaction frequency, and interaction content.

[0230] As an optional embodiment, the processor 601 obtains the representation vectors of the N objects corresponding to the N network nodes based on the associated information network graph in the following specific embodiment:

[0231] Starting from the target network node corresponding to the target object, a random walk is performed in the associated information network graph to obtain M trajectories, each with a step size of P; where M and P are both positive integers.

[0232] Based on the object information carried in the M trajectories, the representation vector of the target object is obtained;

[0233] The probability of traversing from the i-th network node to the j-th network node is proportional to the target edge weight, which is the edge weight in the edge information of the i-th network node and the j-th network node. i and j are both positive integers, i is not equal to j, and i and j are both less than or equal to N.

[0234] Based on the same inventive concept, the principle and beneficial effects of the intelligent device provided in the embodiments of this application in solving the problem are similar to the principle and beneficial effects of the recommended method for multimedia resources in the embodiments of this application in solving the problem. For the sake of brevity, the principle and beneficial effects of the method implementation can be referred to.

[0235] This application also provides a computer-readable storage medium storing one or more instructions adapted for loading by a processor and executing the method for recommending multimedia resources described in the above method embodiments.

[0236] This application also provides a computer program product containing instructions that, when run on a computer, cause the computer to execute the method for recommending multimedia resources described in the above method embodiments.

[0237] This application also provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the aforementioned recommended method for multimedia resources.

[0238] The steps in the method of this application embodiment can be adjusted, combined, or deleted according to actual needs.

[0239] The modules in the device of this application embodiment can be merged, divided, and deleted according to actual needs.

[0240] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc.

[0241] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Those skilled in the art will understand that implementing all or part of the processes of the above embodiments and making equivalent changes in accordance with the claims of this application are still within the scope of the invention.

Claims

1. A method for recommending multimedia resources, characterized in that, include: Obtain the representation vector of the target object and the set of representation vectors of neighboring objects. The set of representation vectors of neighboring objects includes the representation vectors of the first neighboring objects of the target object. The target object and the first neighboring objects are divided into K first relation types, where K is a positive integer. The first neighboring objects are objects that have interactive behavior with the target object on social platforms. Based on the representation vector of the target object and the set of representation vectors of neighboring objects, the representation feature information of the target object is obtained. The representation feature information is determined based on the relation feature information corresponding to each of the K first relation types of the target object. Obtain a set of multimedia resources, and recommend a first multimedia resource from the set to the target object based on the representational feature information of the target object.

2. The method as described in claim 1, characterized in that, The step of obtaining the representation feature information of the target object based on the representation vector of the target object and the set of representation vectors of neighboring objects includes: Based on the representation vector of the target object and the set of representation vectors of the adjacent objects, the relation feature information corresponding to each of the K first relation types is obtained; The target object is analyzed by performing feature analysis on the K relation feature information corresponding to the K first relation types respectively, and the representation feature information of the target object is obtained.

3. The method as described in claim 2, characterized in that, The step of obtaining the relation feature information corresponding to each of the K first relation types based on the representation vector of the target object and the set of representation vectors of neighboring objects includes: Obtain the feature parameter set of the h-th first relation type among the K first relation types, the feature parameter set including the weight matrix and the bias vector of the h-th first relation type; Using the feature parameter set of the h-th first relation type, calculate the target object intermediate features under the h-th first relation type, and calculate the intermediate features of each first neighboring object of the target object under the h-th first relation type; Based on the intermediate features of the target object and the intermediate features of each adjacent object, the relation feature information of the h-th first relation type is obtained.

4. The method as described in claim 3, characterized in that, The relation feature information of the h-th first relation type is the relation feature information of the h-th first relation type at the T-th iteration, where T is a positive integer; obtaining the relation feature information of the h-th first relation type based on the intermediate features of the target object and the intermediate features of each adjacent object includes: Obtain the target probability that each of the first neighboring objects of the target object is classified into the h-th first relation type at the t-th iteration, where t is a positive integer and t is less than T; Based on the target probability and the intermediate features of each neighboring object, calculate the aggregate features of the first neighboring object of the target object; The intermediate features of the target object and the aggregate features of the first neighboring objects of the target object are processed to obtain the relation feature information of the h-th first relation type at the (t+1)-th iteration.

5. The method as described in claim 2, characterized in that, The step of performing feature analysis on the target object using the K relation feature information corresponding to the K first relation types to obtain the representation feature information of the target object includes: Obtain a first feature information set of the first neighboring objects of the target object, and a second feature information set of the second neighboring objects of the target object. The first feature information set includes: relationship feature information corresponding to each of the multiple second relationship types of the first neighboring objects of the target object. The second feature information set includes: relationship feature information corresponding to each of the multiple third relationship types of the second neighboring objects of the target object. The relationship feature information corresponding to the K first relationship types of the target object, the first feature information set, and the second feature information set are used as inputs to the relationship prediction model to obtain the prediction result output by the relationship prediction model. The relationship prediction model includes L layers of graph convolutional network layers, where L is a positive integer. Overfitting is applied to the prediction results to obtain the representational feature information of the target object; In the relationship prediction model, the input data of the g-th graph convolutional network layer includes the data obtained after overfitting the output data of the (g-1)-th graph convolutional network layer.

6. The method as described in claim 1, characterized in that, The acquisition of the multimedia resource set includes: Obtain viewing information of the second multimedia resource, wherein the viewing information includes object identifiers of Q objects that have viewed the second multimedia resource, where Q is a positive integer; Based on the object identifiers of the Q objects, obtain the representation vectors of the Q objects; The representation vectors of the Q objects are fused, and the result of the fusion is averaged and pooled to obtain the resource feature information of the second multimedia resource. The multimedia resource set includes the resource feature information of the second multimedia resource.

7. The method as described in claim 6, characterized in that, The step of recommending a first multimedia resource from the multimedia resource set to the target object based on the representation feature information of the target object includes: Based on the representation feature information of the target object and the resource feature information of each multimedia resource in the multimedia resource set, the matching degree between the target object and each multimedia resource in the multimedia resource set is calculated. A first multimedia resource is recommended to the target object, wherein the first multimedia resource is the multimedia resource in the multimedia resource set that has the highest matching degree with the target object.

8. The method as described in claim 7, characterized in that, The step of calculating the matching degree between the target object and each multimedia resource in the multimedia resource set based on the representation feature information of the target object and the resource feature information of each multimedia resource in the multimedia resource set includes: The resource feature information of the second multimedia resource in the multimedia resource set is concatenated with the representation feature information of the target object to obtain a concatenated feature set; A multilayer perceptron is used to process each splicing feature in the splicing feature set to obtain the relationship vector between each of the K first relationship types and the second multimedia resource; Calculate the weight corresponding to each first relation type based on the relation vector between each first relation type and the second multimedia resource; Based on the relationship vector between each of the K first relationship types and the second multimedia resource, and the weight corresponding to each first relationship type, the matching degree between the target object and the second multimedia resource is obtained.

9. The method as described in claim 1, characterized in that, The method further includes: Obtain a set of association information, which includes a set of object information and a set of relationship information; N network nodes are generated based on the object information set. Each of the N network nodes corresponds to an object, and each network node carries object information of the object corresponding to that network node. N is a positive integer. If the set of relational information indicates that the object corresponding to the first network node and the object of the second network node among the N network nodes have interactive behavior, then the connection information between the first network node and the second network node is generated according to the interactive behavior to obtain the relational information network graph. Based on the associated information network graph, the representation vectors of the N objects corresponding to the N network nodes are obtained.

10. The method as described in claim 9, characterized in that, The edge weights in the connection information between the first network node and the second network node include: weights determined based on the correlation between the first network node and the second network node, wherein the edge weights are proportional to the correlation, and the correlation is determined based on the interaction information between the first network node and the second network node within a target time period, wherein the interaction information includes at least one of the following: cumulative number of interactions, cumulative interaction duration, interaction frequency, and interaction content.

11. The method as described in claim 10, characterized in that, The step of obtaining the representation vectors of the N objects corresponding to the N network nodes based on the associated information network graph includes: Starting from the target network node corresponding to the target object, a random walk is performed in the associated information network graph to obtain M trajectories, each with a step size of P; where M and P are both positive integers. Based on the object information carried in the M trajectories, the representation vector of the target object is obtained; The probability of traversing from the i-th network node to the j-th network node is proportional to the target edge weight, which is the edge weight in the edge information of the i-th network node and the j-th network node. i and j are both positive integers, i is not equal to j, and i and j are both less than or equal to N.

12. A device for recommending multimedia resources, characterized in that, include: The acquisition unit is used to acquire the representation vector of the target object and the set of representation vectors of neighboring objects. The set of representation vectors of neighboring objects includes the representation vectors of the first neighboring objects of the target object. The target object and the first neighboring objects are divided into K first relation types, where K is a positive integer. The first neighboring objects are objects that have interactive behavior with the target object on the social platform. The processing unit is configured to obtain representation feature information of the target object based on the representation vector of the target object and the set of representation vectors of neighboring objects, wherein the representation feature information is determined based on the relation feature information corresponding to each of the K first relation types of the target object; And for acquiring a set of multimedia resources, and recommending a first multimedia resource from the set of multimedia resources to the target object based on the representation feature information of the target object.

13. A smart device, characterized in that, include: Storage devices and processors; The storage device stores a computer program; A processor that executes a computer program to implement the method for recommending multimedia resources as described in any one of claims 1-11.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the method for recommending multimedia resources as described in any one of claims 1-11.

15. A computer program product comprising computer instructions, characterized in that, The computer instructions are stored in a computer-readable storage medium, and when the computer instructions are read and executed by the processor of a computer device, the computer device performs the method for recommending multimedia resources as described in any one of claims 1-11.

Citation Information

Patent Citations

  • Method and device for training interaction prediction model and method and device for predicting interaction object

    CN112085293A

  • Recommendation Method and Apparatus

    US20200272913A1

  • Method, System, and Apparatus for Identifying and Revealing Selected Objects from Video

    US20200327378A1

  • Method, apparatus, device and medium for generating captioning information of multimedia data

    WO2020190112A1