A data processing method, device, and computer-readable storage medium

By constructing heterogeneous graphs and randomly sampling to generate video feature vectors containing multivariate information, the problem in existing technologies that video feature vectors cannot accurately represent the relationship between videos and actual application scenarios is solved, thereby improving the accuracy of video recommendation and clustering.

CN113761272BActive Publication Date: 2025-10-21TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110420502.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-19
Publication Date
2025-10-21
Estimated Expiration
2041-04-19

AI Technical Summary

Technical Problem

Existing video feature vectors are only based on video content, which makes it difficult to accurately represent the relationship between videos and actual application scenarios, resulting in low application accuracy in scenarios such as video recommendation, video recall, or video clustering.

Method used

By constructing a heterogeneous graph, obtaining video identifiers and identification nodes associated with heterogeneous information, random sampling is performed to generate heterogeneous sampling sequences and homogeneous sampling sequences, and the initial word encoding model is used to adjust and generate video feature vectors containing multivariate information.

Benefits of technology

It improves the accuracy of video application in actual application scenarios, can accurately represent the correlation between videos and actual application scenarios, and improves the effect of video recommendation and clustering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113761272B_ABST
    Figure CN113761272B_ABST
Patent Text Reader

Abstract

Embodiments of the application disclose a data processing method, device and computer readable storage medium associated with artificial intelligence, wherein the method comprises: obtaining a video identifier of a video and associated heterogeneous information associated with the video; the data attribute types of the two are different; the video identifier and the heterogeneous information identifier of the associated heterogeneous information are both determined as an identifier node, and a heterogeneous graph containing the identifier node is generated; the identifier nodes of the heterogeneous graph are sampled to obtain a heterogeneous sampling sequence and an isomorphic sampling sequence; the heterogeneous sampling sequence contains at least two identifier nodes belonging to different data attribute types, and the isomorphic sampling sequence contains at least two identifier nodes belonging to the same data attribute type; and a video feature vector corresponding to the video identifier is generated according to the heterogeneous sampling sequence and the isomorphic sampling sequence. By using the application, the video feature vector can contain rich multi-element information, and thus the application accuracy of the video in an actual application scenario can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of Internet technology, and in particular to a data processing method, device, and computer-readable storage medium. Background Art

[0002] In scenarios such as video recommendation, video recall, or video clustering, the vectorized expression of video features is crucial. For example, video feature vectors can be clustered to discover new video topics, similarity calculations can be performed on video feature vectors to recommend related videos, or video feature vectors can be applied to video recommendation models.

[0003] Existing methods for constructing video feature vectors mostly rely on prior information about the video content itself for supervised model training. These methods then select intermediate-layer features from the model as the video's representation vector (also known as a feature vector). For example, when constructing a video classification model based on video classification information, the model first trains the video's classification information. During prediction, the high-dimensional output vector from the model's intermediate layer is used as the video feature vector. Existing methods that rely on the model's output of video feature vectors can contain textual, visual, and audio information, but only within the video content itself. These video feature vectors, which only cover the video's textual or visual information, are only well-suited for video classification scenarios. Applications of these video feature vectors to other practical scenarios, such as video recommendation, video recall, or video clustering, are limited to accurately representing the relationships between videos and the actual application scenario due to their single-dimensional nature (containing only the video or the relationships between videos). Consequently, their accuracy in real-world applications is reduced. Summary of the Invention

[0004] The embodiments of the present application provide a data processing method, device, and computer-readable storage medium, which can enable a video feature vector to contain rich multivariate information, thereby improving the application accuracy of the video in actual application scenarios.

[0005] On the one hand, an embodiment of the present application provides a data processing method, including:

[0006] Obtaining a video identifier of a video and associated heterogeneous information associated with the video; the data attribute type of the video is different from the data attribute type of the associated heterogeneous information;

[0007] Determine the video identifier and the heterogeneous information identifier of the associated heterogeneous information as identification nodes, and generate a heterogeneous graph including the identification nodes;

[0008] Sampling identification nodes of the heterogeneous graph to obtain a heterogeneous sampling sequence and a homogeneous sampling sequence; the heterogeneous sampling sequence includes at least two identification nodes belonging to different data attribute types, and the homogeneous sampling sequence includes at least two identification nodes belonging to the same data attribute type;

[0009] A video feature vector corresponding to the video identifier is generated according to the heterogeneous sampling sequence and the homogeneous sampling sequence.

[0010] In one aspect, an embodiment of the present application provides a data processing device, including:

[0011] A data acquisition module is used to acquire a video identifier of a video and associated heterogeneous information associated with the video; the data attribute type of the video is different from the data attribute type of the associated heterogeneous information;

[0012] A first generating module is configured to determine the video identifier and the heterogeneous information identifier of the associated heterogeneous information as identification nodes, and generate a heterogeneous graph including the identification nodes;

[0013] A sampling node module is used to sample identification nodes of a heterogeneous graph to obtain a heterogeneous sampling sequence and a homogeneous sampling sequence; the heterogeneous sampling sequence contains at least two identification nodes belonging to different data attribute types, and the homogeneous sampling sequence contains at least two identification nodes belonging to the same data attribute type;

[0014] The second generating module is used to generate a video feature vector corresponding to the video identifier according to the heterogeneous sampling sequence and the homogeneous sampling sequence.

[0015] The number of videos is at least two, and the number of associated heterogeneous information is at least two;

[0016] The first generation module includes:

[0017] A first determining unit is configured to determine associated edges between identification nodes and edge weights of associated edges based on an associated relationship between each two videos, an associated relationship between each two pieces of associated heterogeneous information, and an associated relationship between a video and associated heterogeneous information;

[0018] The first generating unit is used to generate a heterogeneous graph according to the identification nodes, the associated edges and the edge weights of the associated edges.

[0019] The identification node includes a video identification node belonging to a video identification and a heterogeneous identification node belonging to a heterogeneous information identification; the association edge includes a first association edge, a second association edge and a third association edge;

[0020] The first determining unit includes:

[0021] A first determining subunit, configured to determine a first associated edge between video identification nodes and an edge weight of the first associated edge based on an association relationship between each two videos;

[0022] A second determining subunit is configured to determine a second associated edge between heterogeneous identification nodes and an edge weight of the second associated edge according to an associated relationship between each two associated heterogeneous information;

[0023] The third determining subunit is configured to determine a third associated edge between the video identification node and the heterogeneous identification node, and an edge weight of the third associated edge, based on the association relationship between the video and the associated heterogeneous information.

[0024] The video includes a first video and a second video; the video identification node includes a first video identification node corresponding to the first video and a second video identification node corresponding to the second video;

[0025] The first determining subunit includes:

[0026] The sequence acquisition subunit is used to acquire valid video sequences associated with N video browsing users respectively; N is a positive integer; N valid video sequences include valid video sequences L x , x is a positive integer and x is less than or equal to the total number of N valid video sequences; valid video sequence L x The valid videos in the list are sorted according to the time sequence of the associated video browsing users; the ratio of the effective browsing time of the valid videos to the total video time of the valid videos is greater than the browsing ratio threshold;

[0027] The position determination subunit is used to determine the position of the first video and the second video if they are in the valid video sequence L x The positions in are adjacent positions, then the valid video sequence L is determined x The first video and the second video have an adjacent position relationship;

[0028] The sequence determination subunit is used to determine, among the N valid video sequences, valid video sequences with adjacent positional relationships as associated valid video sequences; the associated valid video sequences are used to indicate that a first associated edge exists between the first video identification node and the second video identification node;

[0029] The sequence counting subunit is used to count the number of associated sequences associated with the valid video sequence, and determine the number of associated sequences as the edge weight of the first associated edge.

[0030] The associated heterogeneous information includes first associated heterogeneous information and second associated heterogeneous information; the heterogeneous identification node includes a first heterogeneous identification node corresponding to the first associated heterogeneous information and a second heterogeneous identification node corresponding to the second associated heterogeneous information;

[0031] The second determining subunit includes:

[0032] A video determination subunit is configured to determine the same video as an associated video if the video associated with the first associated heterogeneous information and the video associated with the second associated heterogeneous information are identical. The associated video is used to indicate that a second associated edge exists between the first heterogeneous identification node and the second heterogeneous identification node.

[0033] The video counting sub-unit is used to count the number of videos of the associated video and determine the number of videos as the edge weight of the second associated edge.

[0034] The associated heterogeneous information includes a video browsing user group; the heterogeneous identification node includes a user identification node corresponding to the video browsing user group;

[0035] The third determining subunit includes:

[0036] The acquisition number subunit is used to obtain the effective viewing number of the video by the video viewing user group within the video viewing cycle; the effective viewing number refers to the number of times the video is effectively viewed by the video viewing users in the video viewing user group;

[0037] a determining association subunit, configured to determine that a third association edge exists between the video identification node and the user identification node if the valid browsing count is greater than a valid browsing count threshold;

[0038] The weight determination subunit is used to determine the video browsing users who have effectively browsed the video as the associated video browsing users of the video in the video browsing user group, and determine the number of users of the associated video browsing users as the edge weight of the third associated edge.

[0039] The associated heterogeneous information includes at least two video accounts; the associated relationship includes an account associated relationship; and the heterogeneous identification node includes account identification nodes corresponding to the at least two video accounts respectively.

[0040] The third determining subunit includes:

[0041] An account acquisition subunit is used to acquire, from at least two video accounts, an associated video account that has an account association relationship with the video; the account association relationship is used to indicate that the video publishing user publishes the video through the associated video account;

[0042] An account subunit is determined, which is used to determine the account identification node corresponding to the associated video account as the associated account identification node; a third association edge exists between the video identification node and the associated account identification node, and the edge weight of the third association edge is a constant parameter.

[0043] The associated heterogeneous information includes at least two video tags; the associated relationship includes a tag associated relationship; the heterogeneous identification node includes a tag identification node corresponding to at least two video tags respectively;

[0044] The third determining subunit includes:

[0045] The tag acquisition subunit is used to acquire, from at least two video tags, an associated video tag that has a tag association relationship with the video; the tag association relationship is used to indicate that the video is marked with an associated video tag;

[0046] The label determination subunit is used to determine the label identification node corresponding to the associated video label as the associated label identification node; there is a third associated edge between the video identification node and the associated label identification node, and the edge weight of the third associated edge is a constant parameter.

[0047] Among them, the sampling node module includes:

[0048] The first acquisition unit is configured to acquire a random sampling heterogeneous path and a random sampling homogeneous path; the random sampling heterogeneous path is used to indicate a type sampling order of data attribute types sampled when sampling identification nodes of different data attribute types; the random sampling homogeneous path is used to indicate a type of data attribute sampled when sampling identification nodes of the same data attribute type;

[0049] A first sampling unit is configured to randomly sample the identified nodes in the heterogeneous graph according to the type sampling order indicated by the randomly sampled heterogeneous path to obtain a heterogeneous sampling sequence;

[0050] The second sampling unit is used to determine the data attribute type indicated by the randomly sampled isomorphic path as the target data attribute type, and randomly sample the identification nodes belonging to the target data attribute type in the heterogeneous graph to obtain a isomorphic sampling sequence.

[0051] Wherein, the first sampling unit includes:

[0052] a fourth determining subunit, configured to determine, according to a type sampling order, the jth data attribute type to be sampled in the random sampling heterogeneous path as the data attribute type to be sampled; j is a positive integer less than or equal to S, and S is the total number of nodes of the identification nodes to be sampled based on the random sampling heterogeneous path;

[0053] The sampling target subunit is used to sample a target identification node from the heterogeneous graph according to the sampled node set and the attribute type of the data to be sampled; the data attribute type to which the target identification node belongs is the attribute type of the data to be sampled; the sampled node set includes the sampled identification node;

[0054] Add node subunit, used to add the target identification node to the sampled node set if j is less than S;

[0055] The sequence generation subunit is used to generate a heterogeneous sampling sequence according to the sampled node set and the target identification node if j is equal to S. The target identification node is the last identification node in the heterogeneous sampling sequence.

[0056] The sampled node set includes a previous neighbor identification node, and the previous neighbor identification node is the last identification node in the sampled node set;

[0057] Sampling target subunits, including:

[0058] The node acquisition subunit is used to acquire w target pre-selected identification nodes that have associated edges with the previous neighbor identification node from the heterogeneous graph according to the attribute type of the data to be sampled; wherein w is a positive integer; the data attribute type to which the w target pre-selected identification nodes belong is the attribute type of the data to be sampled;

[0059] The summing weight subunit is used to obtain the edge weights of the associated edges between the previous neighbor identification node and each target pre-selected identification node, and sum the obtained edge weights to obtain the total edge weight; the w target pre-selected identification nodes include the target pre-selected identification node Y m , where m is a positive integer and m is less than or equal to w;

[0060] Determine the probability subunit, which is used to compare the previous neighbor identification node with the target pre-selected identification node Y m The edge weight Z of the associated edge between m , and the ratio of the total edge weight, is determined as the target pre-selected identification node Y m The random sampling probability of

[0061] The sampling node subunit is used to randomly sample w target pre-selected identification nodes according to the random sampling probability corresponding to each target pre-selected identification node to obtain the target identification node; in the heterogeneous sampling sequence, the previous neighbor identification node is the previous identification node of the target identification node.

[0062] The homogeneous sampling sequence includes a first homogeneous sampling sequence and a second homogeneous sampling sequence;

[0063] The second sampling unit includes:

[0064] A first generating subunit is configured to randomly sample the video identification nodes in the heterogeneous graph if the target data attribute type is the data attribute type corresponding to the video identification node, to generate a first isomorphic sampling sequence comprising at least two video identification nodes; the first isomorphic sampling sequence is used to represent the topological relationship between the video identification nodes in the heterogeneous graph;

[0065] The second generation subunit is used to randomly sample the heterogeneous identification nodes in the heterogeneous graph if the target data attribute type is the data attribute type corresponding to the heterogeneous identification node, and generate a second homogeneous sampling sequence including at least two heterogeneous identification nodes; the second homogeneous sampling sequence is used to characterize the topological relationship between the heterogeneous identification nodes in the heterogeneous graph.

[0066] The second generation module includes:

[0067] The second determining unit is used to determine the heterogeneous sampling sequence and the homogeneous sampling sequence into at least two random sampling sequences; the at least two random sampling sequences include the random sampling sequence S a , a is a positive integer and a is less than or equal to the total number of sequences of at least two random sampling sequences;

[0068] The second acquisition unit is used to obtain the random sampling sequence S a The true encoding label of each identified node in the random sampling sequence S a Including identification node D b , and the node D b There are adjacent identification nodes with position association relationship, b is a positive integer and b is less than or equal to the random sampling sequence S a The total number of nodes that identify nodes in the graph; the vector dimension of each true encoding label is equal to the total number of nodes that identify nodes in the heterogeneous graph;

[0069] The second acquisition unit is also used to identify the node D b The true encoding label C b Input into the initial word encoding model to obtain the predicted coding labels of adjacent identification nodes;

[0070] The model adjustment unit is used to adjust the model parameters in the initial word encoding model according to the real encoding labels of the adjacent identification nodes and the predicted encoding labels of the adjacent identification nodes to obtain the target word encoding model;

[0071] The second generating unit is configured to input the video identifier into the target word encoding model to obtain a video feature vector corresponding to the video identifier.

[0072] On one hand, the present application provides a computer device, including: a processor, a memory, and a network interface;

[0073] The above-mentioned processor is connected to the above-mentioned memory and the above-mentioned network interface, wherein the above-mentioned network interface is used to provide data communication function, the above-mentioned memory is used to store computer programs, and the above-mentioned processor is used to call the above-mentioned computer program so that the computer device executes the method in the embodiment of the present application.

[0074] On one hand, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. The computer program is suitable for being loaded by a processor and executing the method in the embodiment of the present application.

[0075] On the one hand, an embodiment of the present application provides a computer program product or a computer program, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method in the embodiment of the present application.

[0076] In an embodiment of the present application, a video feature vector is generated based on a heterogeneous sampling sequence and a homogeneous sampling sequence. Since the heterogeneous sampling sequence contains at least two identification nodes belonging to different data attribute types, the video feature vector may include the association relationship between identification nodes belonging to different data attribute types. Similarly, since the homogeneous sampling sequence contains at least two identification nodes belonging to the same data attribute type, the video feature vector may include the association relationship between identification nodes determined by associated heterogeneous information; the data attribute type of the above-mentioned associated heterogeneous information is different from the data attribute type of the video. Therefore, it can be seen that the video feature vector in the present application can cover the features of the associated heterogeneous information associated with the video, as well as the features of the association relationship between the video and the associated heterogeneous information, that is, the video feature vector contains multivariate information features; if the video feature vector in the present application is applied to actual scenarios, such as video recommendation, video recall or video clustering, because it contains diversified information features, it will be able to accurately characterize the association relationship between the video and the actual application scenario, so in the actual application scenario, it can improve the application accuracy of the video. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0078] Figure 1a This is a schematic diagram of a system architecture provided by an embodiment of the present application;

[0079] Figure 1b This is a scenario diagram of a data processing method provided in an embodiment of the present application;

[0080] Figure 2 This is an overall framework diagram of a video feature vector learning based on heterogeneous information provided by an embodiment of the present application;

[0081] Figure 3 This is a flow chart of a data processing method provided in an embodiment of the present application;

[0082] Figure 4 This is a flow chart of a data processing method provided in an embodiment of the present application;

[0083] Figure 5 This is a schematic diagram of a data processing scenario provided by an embodiment of the present application;

[0084] Figure 6 This is a schematic diagram of a data processing scenario provided by an embodiment of the present application;

[0085] Figure 7 This is a flow chart of a data processing method provided in an embodiment of the present application;

[0086] Figure 8 This is a schematic diagram of a data processing scenario provided by an embodiment of the present application;

[0087] Figure 9 is a structural diagram of a data processing device provided in an embodiment of the present application;

[0088] Figure 10 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0089] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0090] To facilitate understanding, we first provide a brief explanation of some nouns:

[0091] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive field within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making. AI technology is an interdisciplinary discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0092] Natural language processing (NLP) is a key area of ​​research in computer science and artificial intelligence. It studies the theories and methods that enable effective communication between humans and computers using natural language. Natural language processing (NLP) integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language we use in everyday life—and is closely linked to the study of linguistics. Natural language processing technologies typically include text processing, semantic understanding, machine translation, robotic question answering, and knowledge graphs.

[0093] A heterogeneous graph is a topological graph comprising at least one type of node and at least one type of edge. In the present application, the at least one type of node may include a video identification node and a heterogeneous identification node, wherein the video identification node is determined based on the video identification, the heterogeneous identification node is determined based on the heterogeneous information identification, and the data attribute type to which the video identification belongs is different from the data attribute type to which the associated heterogeneous identification belongs. The at least one type of edge may include a first associated edge between video identification nodes, a second associated edge between heterogeneous identification nodes, and a third associated edge between video identification nodes and heterogeneous identification nodes.

[0094] The solutions provided in the embodiments of this application involve technologies such as natural language processing and deep learning of artificial intelligence, which are specifically illustrated by the following embodiments.

[0095] See Figure 1a , Figure 1a This is a schematic diagram of a system architecture provided by an embodiment of the present application. Figure 1a As shown, the system may include a server 10a and a user terminal cluster, and the user terminal cluster may include: user terminal 10b, user terminal 10c, ..., user terminal 10d. It can be understood that the above system may include one or more user terminals, and the number of user terminals will not be limited here.

[0096] Among them, there can be a communication connection between the user terminal clusters, for example, there is a communication connection between user terminal 10b and user terminal 10c, and there is a communication connection between user terminal 10b and user terminal 10d. At the same time, any user terminal in the user terminal cluster can have a communication connection with server 10a, for example, there is a communication connection between user terminal 10b and server 10a, and there is a communication connection between user terminal 10c and server 10a. Among them, the above-mentioned communication connection does not limit the connection method, and can be directly or indirectly connected through wired communication, directly or indirectly connected through wireless communication, or through other methods, and this application does not limit it here.

[0097] It should be understood that Figure 1a Each user terminal in the user terminal cluster shown in FIG. 1 may be installed with an application client. When the application client is running in each user terminal, it may be respectively connected to the above-mentioned Figure 1a The server 10a shown in FIG. 1 is connected to the server 10a, i.e., the aforementioned communication connection. The application client may be a social client, multimedia client (e.g., a video client), entertainment client (e.g., a game client), educational client, live streaming client, or other application client capable of loading and playing videos. The application client may be a standalone client or an embedded sub-client integrated into a client (e.g., a social client, educational client, or multimedia client), without limitation.

[0098] Server 10a provides services to user terminal clusters through communication connection functions. When a user terminal (which can be user terminal 10b, user terminal 10c or user terminal 10d) obtains video A and needs to process video A, such as querying video C similar to video A, or obtaining classification information of video A (such as sports, movies, etc.), the user terminal can send video A to server 10a through the above application client. Please refer to Figure 1b , Figure 1b is a scenario diagram of a data processing method provided in an embodiment of the present application, Figure 1b The example scenario is to recommend videos related to video 10f. The video browsing user browses video 10f through user terminal 10b. When he wants to view videos related to video 10f, Figure 1bAs shown, the video browsing user can click the search control 10g in the user terminal 10b, and the user terminal 10b sends the video 10f to the server 10a through the communication connection function. After the server 10a obtains the video 10f, it obtains the video identifier of the video 10f, such as Figure 1b The video identifier vid1 shown in the example is understandable to be the identifier of the video 10f in the video database 10h.

[0099] The video database 10h includes a large number of videos (including video 10f) and a pre-trained target word encoding model. The server 10a inputs the video identifier vid1 into the target word encoding model to obtain a video feature vector 10e for video 10f. It is worth noting that the video feature vector 10e is not generated based on the content of video 10f itself. The vector may include information features of associated heterogeneous information associated with video 10f, such as Figure 1b The user features shown in (which can be understood as the browsing behavior characteristics of the user browsing the video above), account features (features of the account that posted video 10f), and tag features (i.e., features of the video tags carried by video 10f) can allow the video feature vector 10e to accurately represent the relationship between video 10f and other videos.

[0100] The server 10a obtains the video feature vector of each video in the video database 10h, for example Figure 1b The video feature vectors 10i, ..., and 10j are exemplified. Similarly, the video feature vectors 10i, ..., and 10j are similar to the video feature vector 10e and can all contain diversified information features. The server 10a can then perform similarity calculations on the video feature vectors 10e and 10i, ..., and 10j, and use the video feature vector with the highest similarity to the video feature vector 10e as the target video feature vector to obtain the video identifier corresponding to the target video feature vector, for example Figure 1b For example, the video identifiers vid5, vid10, and vid100 are used to obtain the videos corresponding to the above video identifiers. Figure 1b It can be understood that the examples of video 10m, video 10n and video 10p are not only determined based on video 10f, but can also be determined based on the video tags carried by video 10f, the video publishing account that published video 10f, or the video browsing user who browsed video 10f. That is, the video feature vector generated by this application can recommend popular video topics associated with information such as video tags or video publishing accounts to video browsing users.

[0101] Subsequently, server 10a sends videos 10m, 10n, and 10p to the application client of user terminal 10b. After receiving videos 10m, 10n, and 10p sent by server 10a, the application client of user terminal 10b can display videos 10m, 10n, and 10p on its corresponding screen. Server 10a can associate video 10f, video identifier vid1, videos 10m, 10n, and 10p and store them in video database 10h. When receiving video 10f sent by the user terminal again, the server can directly return videos 10m, 10n, and 10p to the user terminal that sent video 10f. The video database 10h can be thought of as an electronic filing cabinet—a place where electronic files (herein, video 10f, video identifier vid1, video 10m, video 10n, and video 10p) are stored. Server 10a can perform operations such as adding, querying, updating, and deleting these files. A "database" is a collection of data that is stored in a specific manner, can be shared with multiple users, has minimal redundancy, and is independent of any application.

[0102] Alternatively, if a pre-trained target word encoding model is stored locally on the user terminal, the user terminal can locally input the video identifier into the target word encoding model to obtain a video feature vector for the video, and then perform downstream tasks based on the video feature vector. Since training the target word encoding model involves a large amount of offline computation, the target word encoding model locally on the user terminal can be trained and sent to the user terminal by the server 10a.

[0103] Further, please see Figure 2 , Figure 2 This is an overall framework diagram of a video feature vector learning based on heterogeneous information provided by the embodiment of the present application. Figure 2 As shown, the framework can include three parts:

[0104] 1) Composition. This involves constructing a heterogeneous graph. First, a large number of videos and a large amount of associated heterogeneous information are obtained. This associated heterogeneous information is information associated with the videos, but with different data attribute types. It is understood that the associated heterogeneous information can include heterogeneous information of one data attribute type or multiple data attribute types. This application does not limit the number of data attribute types for the associated heterogeneous information, and this number can be set based on the actual application scenario.

[0105] For ease of understanding, the entire document uses video viewing users, video publishing accounts, and video tags to illustrate associated heterogeneous information. Specifically, the associated heterogeneous information in this application includes video viewing users, video publishing accounts, and video tags. The data attribute type corresponding to video viewing users is user attributes, the data attribute type corresponding to video publishing accounts is account attributes, and the data attribute type corresponding to video tags is tag attributes. Obviously, the data attribute types corresponding to various types of associated heterogeneous information are also different.

[0106] This application obtains the video identifier Vid of the video, the user identifier Gid of the video browsing user, the account identifier Pid of the video publishing account, and the tag identifier Tid of the video tag; it should be noted that the user identifier Gid in the heterogeneous graph is constructed for at least one video browsing user with the same attributes. For example, all 20-year-old female users in Shenzhen have one identifier, so the user identifier Gid can be understood as the identifier of the video browsing user group.

[0107] The identification node is generated by using the video identification Vid, the user identification Gid, the account identification Pid and the tag identification Tid. The identification node may include a video identification node generated according to the video identification Vid (such as Figure 2 The video identification node 1, video identification node 2 and video identification node 3 in the example are generated according to the user identification Gid (eg, Figure 2 The user identification node 4 and the user identification node 5 in the example are generated according to the account identification Pid (such as Figure 2 The example account identification node 6 and account number identification node 7), the tag identification node generated according to the tag identification Tid (such as Figure 2 The illustrated label identifies node 8 and the label identifies node 9).

[0108] This application uses the association relationship between each identification node to determine the association edge, and constructs a network heterogeneous graph based on the identification nodes and the association edge. Among them, each association edge carries an edge weight, and the edge weight is determined based on the association relationship between the two identification nodes corresponding to the association edge. Among them, the specific process of determining the association edge and the edge weight of the association edge based on the association relationship between the identification nodes is shown below. Figure 4 The corresponding embodiments will not be described in detail here.

[0109] 2) Random sampling. This application randomly samples the identification nodes in the heterogeneous graph according to the preset random sampling heterogeneous path, for example Figure 2The randomly sampled heterogeneous path Gid-Vid-Tid-Vid-Gid shown in the example randomly samples identification nodes in the heterogeneous graph according to the randomly sampled heterogeneous path Gid-Vid-Tid-Vid-Gid to generate a heterogeneous sampling sequence, such as user identification node 4-video identification node 1-label identification node 8-video identification node 2-user identification node 5 (which can be abbreviated as gid4-vid1-tid8-vid2-gid5). This sequence can reflect the topological relationship between video browsing users, videos, and video tags in the heterogeneous graph. Figure 2 An example of a randomly sampled heterogeneous path Gid-Vid-Pid-Vid-Gid is also given. According to the randomly sampled heterogeneous path Gid-Vid-Pid-Vid-Gid, the heterogeneous sampling sequence generated by the randomly sampled identification node can reflect the topological relationship between video browsing users, videos, and video publishing accounts in the heterogeneous graph.

[0110] This application randomly samples the identification nodes of the target data attribute type in the heterogeneous graph according to the preset random sampling isomorphic path, wherein the target data attribute type is the data attribute type indicated by the random sampling isomorphic path, for example Figure 2 The randomly sampled isomorphic path Vid-Vid-Vid-Vid-Vid shown in the example, that is, the target data attribute type is a video attribute. According to the randomly sampled isomorphic path Vid-Vid-Vid-Vid-Vid, the video identification nodes in the heterogeneous graph are randomly sampled to generate an isomorphic sampling sequence, such as vid1-vid2-vid3-vid2-vid1. This sequence can reflect the topological relationship between videos in the heterogeneous graph. Figure 2 An example of a randomly sampled isomorphic path Tid-Tid-Tid-Tid-Tid is also given. According to the randomly sampled isomorphic path Tid-Tid-Tid-Tid-Tid, the isomorphic sampling sequence generated by the randomly sampled label identification nodes can reflect the topological relationship between labels in the heterogeneous graph.

[0111] This application takes two randomly sampled heterogeneous paths and two randomly sampled homogeneous paths as examples, such as Figure 2 The random sampling heterogeneous path shown is Gid-Vid-Tid-Vid-Gid, the random sampling heterogeneous path is Gid-Vid-Pid-Vid-Gid, the random sampling homogeneous path is Vid-Vid-Vid-Vid-Vid, and the random sampling homogeneous path is Tid-Tid-Tid-Tid-Tid. The server can obtain a large number of random sampling sequences (including homogeneous sampling sequences and heterogeneous sampling sequences) based on the above four sampling paths.

[0112] 3) Training the initial word encoding model. This application uses the heterogeneous sampling sequences and homogeneous sampling sequences generated by random sampling as the input of the initial word encoding model. The contextual relationship of the nodes in each sequence is used as the constraint of the initial word encoding model to learn the representation vectors of the nodes in the heterogeneous graph. Figure 2 The initial word encoding model exemplified may be a skipgram model (a word vector learning model that predicts context based on a central word), that is, the current identification node may be used to predict its context identification node.

[0113] Please see again Figure 2 During a training process, the input layer of the initial word encoding model obtains a one-hot encoding vector for the identification node. The dimension of this one-hot encoding vector is equal to the total number of nodes V in the heterogeneous graph. After a hidden layer mapping, the representation features of the identification node are reduced to h dimensions. Finally, after normalization (softmax) processing at the output layer, an output probability distribution of V dimensions is obtained. Assuming the window size is k, based on the current identification node, the k identification nodes before and after it are predicted, and finally k V-dimensional probability distributions are output, such as Figure 2 shown.

[0114] Through the above training, this application can obtain the target word encoding model, which includes a mapping matrix for the video , through the mapping matrix As well as the video identifier, this application can obtain the video feature vector of each video in training.

[0115] To sum up, the video feature vector generated by this application can contain the topological relationship between various heterogeneous information, so the vector can be applied to various downstream business scenarios, such as video clustering, recommended video recall, related video recommendation, etc. It can assist the operation side in exploring new video topics, and can also be connected to the video recommendation system to improve consumption indicators such as clicks, duration, and retention.

[0116] in, Figure 1a The server 10a, user terminal 10b, user terminal 10c, ..., user terminal 10d may include a mobile phone, a tablet computer, a laptop computer, a PDA, a smart speaker, a mobile internet device (MID), a POS (Point Of Sales), a wearable device (such as a smart watch, a smart bracelet, etc.), etc.

[0117] It is understandable that the data processing method provided in the embodiments of the present application can be executed by a computer device, including but not limited to a terminal or a server. The above-mentioned server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The above-mentioned terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited to this. The terminal and the server can be directly or indirectly connected via wired or wireless communication, and this application does not limit this.

[0118] Further, see Figure 3 , Figure 3 This is a flow chart of a data processing method provided by an embodiment of the present application. The data processing method can be Figure 1a The server or user terminal can also execute the method together with the server and the user terminal. In the embodiment of the present application, the method is executed by the server as an example. Figure 3 As shown, the data processing process may include the following steps.

[0119] Step S101: obtaining a video identifier of a video and associated heterogeneous information associated with the video; the data attribute type of the video is different from the data attribute type of the associated heterogeneous information.

[0120] Specifically, the data used in this application can come from information flow products, and the data includes videos and associated heterogeneous information associated with the videos, wherein the data may include multiple videos, and this application does not limit the number of videos; the associated heterogeneous information may include heterogeneous information of one data attribute type or multiple data attribute types. This application does not limit the specific data attribute type of the associated heterogeneous information, and it can be set according to the actual application scenario. For example, the associated heterogeneous information is the user who browses the video, the account that publishes the video, the tags carried by the video, the introduction of the video, and the comments on the video.

[0121] In step S102 , the video identifier and the heterogeneous information identifier of the associated heterogeneous information are both determined as identification nodes, and a heterogeneous graph including the identification nodes is generated.

[0122] Specifically, the server obtains the video identifier of the video and the heterogeneous information identifier of the associated heterogeneous information, wherein the video identifier can be the video name, video URL, etc. of the video, which is not limited here, as long as it is unique; similarly, the heterogeneous information identifier can be any information that can be used to identify the associated heterogeneous information, and this application does not limit the method, as long as it is unique. The server determines the video identifier as a video identification node and the heterogeneous information identifier as a heterogeneous identification node. Obviously, the data attribute type corresponding to the video identification node is different from the data attribute type corresponding to the heterogeneous identification node.

[0123] The number of videos is at least two, and the number of associated heterogeneous information is at least two. Based on the association relationship between each two videos, the association relationship between each two associated heterogeneous information, and the association relationship between a video and the associated heterogeneous information, the association edges between the identification nodes and the edge weights of the association edges are determined. Based on the identification nodes, the association edges, and the edge weights of the association edges, a heterogeneous graph is generated. The association edges include a first association edge, a second association edge, and a third association edge. Obviously, the heterogeneous graph includes identification nodes belonging to different data attribute types and multiple types of association edges.

[0124] The specific process of determining the association edge between the identification nodes and the edge weight of the association edge may include: the server determines the first association edge between the video identification nodes and the edge weight of the first association edge according to the association relationship between each two videos, for example Figure 2 There is a first association edge between the video identification node 2 and the video identification node 3; according to the association relationship between each two associated heterogeneous information, the second association edge between the heterogeneous identification nodes and the edge weight of the second association edge are determined. It should be understood that the data attribute types corresponding to the two heterogeneous identification nodes connected by the second association edge are the same. For example, the two heterogeneous identification nodes connected by the second association edge are both label identification nodes, such as Figure 2 There is a second association edge between the label identification node 8 and the label identification node 9 shown in ; according to the association relationship between the video and the associated heterogeneous information, the third association edge between the video identification node and the heterogeneous identification node and the edge weight of the third association edge are determined, such as Figure 2 There is a third association edge between the video identification node 1 and the account identification node 6 shown in FIG. The edge weight of the first association edge, the edge weight of the second association edge, and the edge weight of the third association edge are determined as follows: Figure 4 The description of the corresponding embodiment is not elaborated here. It is understandable that in actual application, the associated heterogeneous information and the association relationship between the video and the associated heterogeneous information can be set according to the application scenario, and then different types of second associated edges and third associated edges can be obtained.

[0125] Step S103 , performing identification node sampling on the heterogeneous graph to obtain a heterogeneous sampling sequence and a homogeneous sampling sequence; the heterogeneous sampling sequence includes at least two identification nodes belonging to different data attribute types, and the homogeneous sampling sequence includes at least two identification nodes belonging to the same data attribute type.

[0126] Specifically, a random sampling heterogeneous path and a random sampling homogeneous path are obtained; the random sampling heterogeneous path is used to indicate the type sampling order of the data attribute types sampled when sampling identification nodes of different data attribute types; the random sampling homogeneous path is used to indicate the data attribute types sampled when sampling identification nodes of the same data attribute type.

[0127] According to the type sampling order indicated by the randomly sampled heterogeneous path, the identification nodes in the heterogeneous graph are randomly sampled to obtain a heterogeneous sampling sequence; the data attribute type indicated by the randomly sampled homogeneous path is determined as the target data attribute type, and the identification nodes belonging to the target data attribute type in the heterogeneous graph are randomly sampled to obtain a homogeneous sampling sequence.

[0128] The server obtains the pre-set random sampling heterogeneous paths and random sampling isomorphic paths. It can be understood that the random sampling heterogeneous paths and random sampling isomorphic paths can be set according to the actual application scenario. The present application does not limit the random sampling heterogeneous paths and random sampling isomorphic paths. The random sampling heterogeneous paths include different data attribute types. When randomly sampling the identification nodes in the heterogeneous graph, the sampling order is based on the type indicated by the random sampling heterogeneous paths. The present application can sample the heterogeneous graph multiple times based on a small number of random sampling heterogeneous paths and random sampling isomorphic paths, and then obtain a large number of heterogeneous sampling sequences and isomorphic sampling sequences, so that a large amount of training data can be prepared for the unsupervised training in step S104.

[0129] Step S104 : generating a video feature vector corresponding to the video identifier according to the heterogeneous sampling sequence and the homogeneous sampling sequence.

[0130] Specifically, the heterogeneous sampling sequence and the homogeneous sampling sequence are determined to be at least two random sampling sequences; the at least two random sampling sequences include the random sampling sequence S a , a is a positive integer and a is less than or equal to the total number of sequences of at least two random sampling sequences; obtain the random sampling sequence S a The true encoding label of each identified node in the random sampling sequence S a Including identification node D b , and the node D b There are adjacent identification nodes with position association relationship, b is a positive integer and b is less than or equal to the random sampling sequence S aThe total number of nodes in the identification node; the vector dimension of each true encoding label is equal to the total number of nodes in the heterogeneous graph identification node; the identification node D b The true encoding label C b Input into the initial word encoding model to obtain the predicted coding labels of adjacent identification nodes; adjust the model parameters in the initial word encoding model according to the actual coding labels of the adjacent identification nodes and the predicted coding labels of the adjacent identification nodes to obtain the target word encoding model; input the video identifier into the target word encoding model to obtain the video feature vector corresponding to the video identifier.

[0131] The homogeneous sampling sequence and the heterogeneous sampling sequence obtained in step S103 are both determined as random sampling sequences, and the random sampling sequences are used to perform unsupervised training on the word vector learning model (i.e., the initial word encoding model). It is understandable that the embodiment of the present application does not limit the model type to which the initial word encoding model belongs, and it can be any word vector learning model, such as the skipgram model or the CBOW model (a word vector learning model that predicts the central word based on the context). The server can obtain the one-hot encoding vector of each identification node in each random sampling sequence outside the initial word encoding model, or can obtain the one-hot encoding vector of each identification node in each random sampling sequence through the initial word encoding model.

[0132] For ease of understanding and description, the present embodiment is described using the training of a skipgram model. Assuming the random sampling sequence is {gid1, vid4, tid6, vid3, gid2}, k=2, during a training process, if label identification node 6 is used as the input of the skipgram model (equivalent to using the one-hot encoding vector corresponding to label identification node 6 as the input of the skipgram model), then its four upper and lower adjacent identification nodes, namely user identification node 1, video identification node 4, video identification node 3, and user identification node 2, are used as supervisory signals. The one-hot encoding vectors corresponding to the above four identification nodes are used as the true encoding labels for model training to obtain the predicted encoding vectors corresponding to the above four identification nodes, i.e., the predicted encoding labels. Then, based on the above four true encoding labels and the four predicted encoding labels, the model parameters in the skipgram model are adjusted to obtain the target word encoding model.

[0133] After the server obtains the target word encoding model, it can input the video identifier in training into the target word encoding model to obtain a video feature vector corresponding to the video identifier.

[0134] In an embodiment of the present application, a video feature vector is generated based on a heterogeneous sampling sequence and a homogeneous sampling sequence. Since the heterogeneous sampling sequence contains at least two identification nodes belonging to different data attribute types, the video feature vector may include the association relationship between identification nodes belonging to different data attribute types. Similarly, since the homogeneous sampling sequence contains at least two identification nodes belonging to the same data attribute type, the video feature vector may include the association relationship between identification nodes determined by associated heterogeneous information; the data attribute type of the above-mentioned associated heterogeneous information is different from the data attribute type of the video. Therefore, it can be seen that the video feature vector in the present application can cover the features of the associated heterogeneous information associated with the video, as well as the features of the association relationship between the video and the associated heterogeneous information, that is, the video feature vector contains multivariate information features; if the video feature vector in the present application is applied to actual scenarios, such as video recommendation, video recall or video clustering, because it contains diversified information features, it will be able to accurately characterize the association relationship between the video and the actual application scenario, so in the actual application scenario, it can improve the application accuracy of the video. In addition, this application avoids the existing supervised training method by constructing a heterogeneous graph to learn the topological structure relationship between videos and associated heterogeneous information, making the application range of video feature vectors wider.

[0135] Further, see Figure 4 , Figure 4 This is a flow chart of a data processing method provided by an embodiment of the present application. The data processing method can be Figure 1a The server or user terminal can also execute the method together with the server and the user terminal. In the embodiment of the present application, the method is executed by the server as an example. Figure 4 As shown, the data processing process may include the following steps.

[0136] Step S201, obtaining a video and associated heterogeneous information associated with the video; the associated heterogeneous information includes the video browsing user, the video publishing account, and the video tag; the data attribute types corresponding to the video, the video browsing user, the video publishing account, and the video tag are all different.

[0137] Specifically, in the embodiments of this application, heterogeneous information is associated with the user who views the video, the account that posts the video (also referred to as the video posting account), and the tags carried by the video (also referred to as video tags). This application aggregates video-viewing users according to the age-gender-geographic city triple. That is, a single user identification node includes all video-viewing users with the same gender, age, and region. For example, "20-year-old Shenzhen female" is one user identification node, and "30-year-old Dongguan male" is another user identification node. This results in a large number of user identification nodes. Therefore, unless otherwise specified, the video-viewing user in this article is a user group that includes at least one video-viewing user.

[0138] Please also see Figure 5 , Figure 5 This is a schematic diagram of a data processing scenario provided by an embodiment of the present application. Figure 5 As shown, the server 40a obtains 3 videos, namely Figure 5 The video 401b, video 401c and video 401d are shown, wherein the video 401b is created by the account 404b (eg Figure 5 The video is published by the video publishing account 100000) and carries the video tag 403b (such as Figure 5 Sports in), user group 402b has viewed video 401b; video 401c is viewed by account 404c (such as Figure 5 The video is published by the video publishing account 200000) and carries the video tag 403b (such as Figure 5 Sports in ) and video tag 403c (such as Figure 5 ), user group 402b and user group 402c have browsed video 401c respectively; video 401d is published by account 404c and carries video tag 403c (such as Figure 5 ), user group 402c has viewed video 401d.

[0139] It is understandable that Figure 5 The example numbers (such as 3 videos, 2 user groups, 2 video publishing accounts and 2 video tags) are assumed for ease of understanding and description. In actual application, the server 40a needs to obtain a large number of videos and related heterogeneous information to extract various topological structural relationships in the heterogeneous graph.

[0140] Step S202: The video ID of the video, the user ID of the video browsing user, the account ID of the video publishing account, and the tag ID of the video tag are all determined as identification nodes; the number of videos is at least two, and the number of video tags is at least two.

[0141] Specifically, the identification node includes a video identification node belonging to the video identification, a user identification node belonging to the user identification, an account identification node belonging to the account identification, and a tag identification node belonging to the tag identification.

[0142] This embodiment of the application uses serial numbers to name video identifiers, user identifiers, account identifiers, and tag identifiers. Please refer to Figure 5 , the server 40a identifies the video 401b as the video identifier vid1, identifies the video 401c as the video identifier vid2, and identifies the video 401d as the video identifier vid3; the server 40a identifies the account 404b (ie Figure 5 The video publishing account 100000 in the video is identified as account ID pid6, and the account 404c (i.e. Figure 5 The video publishing account 200000 in the video is identified as account ID pid7; the server 40a identifies the user group 402b as user ID gid4 and the user group 402c as user ID gid5; the server 40a identifies the video tag 403b (ie Figure 5 Sports in) is identified as tag identifier tid8, and video tag 403c (ie Figure 5 The movie in is identified by tag identifier tid9.

[0143] The server 40a determines the video ID, user ID, account ID and tag ID as identification nodes. Figure 5 The video identifier vid1, video identifier vid2, video identifier vid3, account identifier pid6, account identifier pid7, user identifier gid4, user identifier gid5, tag identifier tid8 and tag identifier tid9 in are all determined as identification nodes. Among them, the server 40a generates video identification nodes according to the video identifier vid1, video identifier vid2 and video identifier vid3 respectively. Figure 5 The circular nodes are used to indicate the video identification nodes, such as Figure 5 The example video identification node 1, video identification node 2 and video identification node 3; the server 40a generates account identification nodes according to the account identification pid6 and account identification pid7 respectively. Figure 5 The diamond-shaped nodes indicate the account identification nodes, such as Figure 5 The example account identification node 6 and the account identification node 7; the server 40a generates a user identification node according to the user identification gid4 and the user identification gid5 respectively, Figure 5 The rectangular nodes are used to indicate user identification nodes, such as Figure 5 The example user identification node 4 and user identification node 5; the server 40a generates a tag identification node according to the tag identification tid8 and the tag identification tid9 respectively, Figure 5The nodes are marked with triangle nodes as shown in the following example: Figure 5 The example shows an account identification node 8 and an account number identification node 9.

[0144] Step S203, based on the association relationship between each two videos, the association relationship between each two video tags, and the association relationship between the video and the associated heterogeneous information, determine the association edges between the identification nodes and the edge weights of the association edges, and generate a heterogeneous graph based on the identification nodes, the association edges and the edge weights of the association edges.

[0145] Specifically, the associated edges include a first associated edge, a second associated edge, and a third associated edge.

[0146] The server 40a determines the first associated edge between the corresponding two video identification nodes and the edge weight of the first associated edge according to the association relationship between each two videos, for example Figure 5 The first associated edge between the video identification node 2 and the video identification node 3 in the example; according to the association relationship between each two video tags, the second associated edge between the corresponding two tag identification nodes and the edge weight of the second associated edge are determined by Figure 5 It can be seen that there is a second association edge between the label identification node 1 and the label identification node 2; the server 40a can determine the third association edge between the video node and the heterogeneous identification node, and the edge weight of the third association edge according to the association relationship between the video and the associated heterogeneous information, for example Figure 5 In the example, there is a third association edge between the video identification node 1 and the user identification node 4, a third association edge between the video identification node 1 and the account identification node 6, and a third association edge between the video identification node 1 and the tag identification node 8. It is understood that in actual applications, heterogeneous information can be associated according to the application scenario, thereby obtaining different types of third association edges.

[0147] Please see again Figure 5 , the server 40a generates a heterogeneous graph based on the video identification node, the user identification node, the account identification node, the tag identification node, the first associated edge, the second associated edge, and the third associated edge. Each associated edge carries an edge weight. The specific process of determining the edge weight of the associated edge is shown below. Figure 7 The corresponding embodiments will not be described in detail this time.

[0148] Step S204, obtaining a randomly sampled heterogeneous path and a randomly sampled homogeneous path; the randomly sampled heterogeneous path is used to indicate the type sampling order of the data attribute types sampled when sampling identification nodes of different data attribute types; the randomly sampled homogeneous path is used to indicate the data attribute types sampled when sampling identification nodes of the same data attribute type.

[0149] Specifically, the server obtains a pre-set random sampling heterogeneous path and a random sampling homogeneous path. It is understandable that the random sampling heterogeneous path and the random sampling homogeneous path can be set according to the actual application scenario. This application does not limit the random sampling heterogeneous path and the random sampling homogeneous path. The random sampling heterogeneous path includes at least two data attribute types. When randomly sampling the identification node in the heterogeneous graph, the sampling order is based on the type indicated by the random sampling heterogeneous path.

[0150] Step S205 , randomly sampling the identification nodes in the heterogeneous graph according to the type sampling order indicated by the random sampling heterogeneous path to obtain a heterogeneous sampling sequence; the heterogeneous sampling sequence includes at least two identification nodes belonging to different data attribute types.

[0151] Specifically, according to the type sampling order, the data attribute type required to be sampled for the jth item in the random sampling heterogeneous path is determined as the data attribute type to be sampled; j is a positive integer less than or equal to S, and S is the total number of nodes of the identification nodes required to be sampled based on the random sampling heterogeneous path; according to the sampled node set and the data attribute type to be sampled, the target identification node is sampled from the heterogeneous graph; the data attribute type to which the target identification node belongs is the data attribute type to be sampled; the sampled node set includes the sampled identification node; if j is less than S, the target identification node is added to the sampled node set; if j is equal to S, a heterogeneous sampling sequence is generated according to the sampled node set and the target identification node, and the target identification node is the last identification node in the heterogeneous sampling sequence.

[0152] The sampled node set includes a preceding neighbor identification node, which is the last identification node in the sampled node set; the specific process of sampling a target identification node from a heterogeneous graph may include: obtaining w target preselected identification nodes having associated edges with the preceding neighbor identification node from the heterogeneous graph according to the attribute type of the data to be sampled; wherein w is a positive integer; the data attribute type to which the w target preselected identification nodes belong is the attribute type of the data to be sampled; respectively obtaining the edge weights of the associated edges between the preceding neighbor identification node and each target preselected identification node, summing the obtained edge weights to obtain a total edge weight; the w target preselected identification nodes include the target preselected identification node Y m , where m is a positive integer and m is less than or equal to w; the previous neighbor identification node and the target pre-selected identification node Y m The edge weight Z of the associated edge between m , and the ratio of the total edge weight, is determined as the target pre-selected identification node Y mrandom sampling probability; randomly sample w target pre-selected identification nodes according to the random sampling probability corresponding to each target pre-selected identification node to obtain the target identification node; in the heterogeneous sampling sequence, the previous neighbor identification node is the previous identification node of the target identification node.

[0153] Please also see Figure 6 , Figure 6 This is a schematic diagram of a data processing scenario provided by an embodiment of the present application. Figure 6 As shown, the embodiment of the present application assumes that the random sampling heterogeneous path is Gid-Vid-Tid-Vid-Gid. It can be seen that the server will first randomly sample the identification node belonging to the user attribute (i.e., the user identification node), and then randomly sample the identification node belonging to the video attribute (i.e., the video identification node), and then randomly sample the identification node belonging to the tag attribute (i.e., the tag identification node), and then randomly sample the identification node belonging to the video attribute (i.e., the video identification node), and finally randomly sample the identification node belonging to the user attribute (i.e., the user identification node). The resulting heterogeneous sampling sequence includes 5 identification nodes and 3 data attribute types.

[0154] Please see again Figure 6 , Figure 6 Each associated edge in carries an edge weight. For example, there is a third associated edge between user identification node 4 and video identification node 1, and the edge weight of the third associated edge between user identification node 4 and video identification node 1 is E (gid4,vid1) There is a third associated edge between the user identification node 4 and the video identification node 2, and the edge weight of the third associated edge between the user identification node 4 and the video identification node 2 is E (gid4,vid2) =4; there is a third associated edge between the account identification node 6 and the video identification node 1, and the edge weight of the third associated edge between the account identification node 6 and the video identification node 1 is E (vid1, pid6) There is a third associated edge between the label identification node 8 and the video identification node 1, and the edge weight of the third associated edge between the label identification node 8 and the video identification node 1 is E (vid1, tid8) Equal to 1.

[0155] In a sampling of identification nodes, such as Figure 6As shown, the identification nodes in the heterogeneous graph are randomly sampled according to the random sampling heterogeneous path Gid-Vid-Tid-Vid-Gid; according to the type sampling order, the first data attribute type to be sampled in the random sampling heterogeneous path Gid-Vid-Tid-Vid-Gid is first determined as the data attribute type to be sampled, that is, the data attribute type to be sampled is the user attribute. Since the current randomly sampled identification node is the first identification node, the sampled node set is an empty set, that is, there is no previous neighbor identification node. At this time, the user identification nodes belonging to the user attribute can be randomly sampled on average, which can also be understood as traversing the user identification nodes one by one. For example, according to the random sampling heterogeneous path Gid-Vid-Tid-Vid-Gid, the identification nodes in the heterogeneous graph are randomly sampled 100 times, then Figure 6 The user identification node 4 in Figure 6 The user identification node 5 in is randomly sampled 50 times.

[0156] Assume that the target identification node sampled from the heterogeneous graph this time is user identification node 4, such as Figure 6 As shown, since j = 1, which is less than S (S = 5), the user identification node 4 is added to the sampled node set. At this time, according to the type sampling order indicated by the randomly sampled heterogeneous path Gid-Vid-Tid-Vid-Gid, the first identification node is successfully sampled, and then the data attribute type required to be sampled for the second random sampling heterogeneous path Gid-Vid-Tid-Vid-Gid is determined as the data attribute type to be sampled, that is, the data attribute type to be sampled is the video attribute. Obviously, the current randomly sampled identification node is the second identification node, and the sampled node set includes the user identification node 4, so the previous neighbor identification node is the user identification node 4. At this time, a random walk (deepwalk) algorithm can be used to randomly sample the video identification nodes belonging to the video attribute. It can be understood that this application does not limit the random walk algorithm, and any random walk algorithm can be used. This application uses the metapath2vec algorithm (a vertex embedding method for heterogeneous information networks). The specific sampling process is described as follows.

[0157] Please see again Figure 6 , it is known that the current previous neighbor identification node is user identification node 4, and the attribute type of the data to be sampled is video attribute. The server can obtain the target pre-selected identification node from the heterogeneous graph, that is, Figure 6 The video identification node 1 belonging to the video attribute and the video identification node 2 belonging to the video attribute. Further, the server obtains the edge weights of the associated edges between the above two target pre-selected identification nodes and the previous neighbor identification node (ie, user identification node 4), such as Figure 6As shown, the edge weight E of the third association edge between the user identification node 4 and the video identification node 1 is (gid4,vid1) Equal to 3, the edge weight E of the third association edge between user identification node 4 and video identification node 2 (gid4,vid2) =4, so the total edge weight is 7; the random sampling probability of video identification node 1 is p(vid1|gid4)=3 / 7, and the random sampling probability of video identification node 2 is p(vid2|gid4)=4 / 7. The server randomly samples the target pre-selected identification nodes based on the random sampling probability of each target pre-selected identification node. In this embodiment of the present application, if the identification nodes in the heterogeneous graph are randomly sampled 100 times according to the random sampling heterogeneous path Gid-Vid-Tid-Vid-Gid, the second identification node in the heterogeneous sampling sequence is video identification node 1 approximately 43 times and video identification node 2 approximately 57 times.

[0158] If the above process of randomly sampling the identification nodes in the heterogeneous graph based on the random sampling of heterogeneous paths is understood as a random walk, it can be expressed by formula (1), which is as follows:

[0159] (1)

[0160] Where P represents a randomly sampled heterogeneous path, such as Figure 6 The random sampling heterogeneous path shown in the example is Gid-Vid-Tid-Vid-Gid; i represents the number of random walk steps. For example, i=1 means that the first identification node sampled randomly walks to the second identification node sampled, for example, the user identification node 4 mentioned above randomly walks to the video identification node 2; express Random walk to The random walk probability corresponds to the random sampling probability mentioned above; v represents the identification node, , represents the identification node of the tth type; express Among the neighbor identification nodes of , the identification node belongs to type t+1, It is the target identification node and belongs to the t+1 type, corresponding to the attribute type of the data to be sampled; express and The edge weights between express The node type; express as well as There are related edges between them.

[0161] According to formula (1), when performing the random walk in step i, only the predefined node types in the randomly sampled heterogeneous paths, i.e., the data attribute types, are considered. The larger the weight of the connected edge is, the greater the random walk probability of the identified node is.

[0162] Please see again Figure 6 Assume that the target identification node sampled from the heterogeneous graph for the second time is video identification node 2. Since j = 2, which is less than S (S = 5), video identification node 2 is added to the sampled node set. The current sampled node set includes user identification node 4 and video identification node 2. It is understandable that the subsequent random sampling is consistent with the above-mentioned sampling identification node process, so it will not be described in detail here. Until j equals S, a heterogeneous sampling sequence is generated based on the sampled node set and the target identification node.

[0163] Step S206 , determining the data attribute type indicated by the randomly sampled isomorphic path as the target data attribute type, and randomly sampling the identification nodes belonging to the target data attribute type in the heterogeneous graph to obtain an isomorphic sampling sequence; the isomorphic sampling sequence includes at least two identification nodes belonging to the same data attribute type.

[0164] Specifically, the isomorphic sampling sequence includes a first isomorphic sampling sequence and a second isomorphic sampling sequence; the specific process of obtaining the isomorphic sampling sequence may include: if the target data attribute type is the data attribute type corresponding to the video identification node, then randomly sampling the video identification nodes in the heterogeneous graph to generate a first isomorphic sampling sequence containing at least two video identification nodes; the first isomorphic sampling sequence is used to characterize the topological relationship between the video identification nodes in the heterogeneous graph; if the target data attribute type is the data attribute type corresponding to the heterogeneous identification node, then randomly sampling the heterogeneous identification nodes in the heterogeneous graph to generate a second isomorphic sampling sequence containing at least two heterogeneous identification nodes; the second isomorphic sampling sequence is used to characterize the topological relationship between the heterogeneous identification nodes in the heterogeneous graph.

[0165] First, the target data attribute type indicated in the randomly sampled isomorphic path is determined. The embodiment of the present application cites two types: video attributes and label attributes. The server can then employ a random walk algorithm to randomly sample video identification nodes belonging to the video attribute, generating a first isomorphic sampling sequence comprising at least two video identification nodes. The first isomorphic sampling sequence is used to characterize the topological relationship between video identification nodes in the heterogeneous graph. The server can employ a random walk algorithm to randomly sample label identification nodes belonging to the label attribute, generating a second isomorphic sampling sequence comprising at least two label identification nodes. The second isomorphic sampling sequence is used to characterize the topological relationship between label identification nodes in the heterogeneous graph.

[0166] It can be understood that this application does not limit the random walk algorithm, and any random walk algorithm can be used, for example, it can be any random walk algorithm among the Node2Vec algorithm (an algorithm that uses vector modeling for nodes in a network graph), the second-order PageRank algorithm (a link analysis algorithm), the second-order SimRank algorithm (a collaborative filtering recommendation algorithm), and the second-order RWR algorithm (a restart random walk algorithm).

[0167] Step S207 : generating a video feature vector corresponding to the video identifier according to the heterogeneous sampling sequence and the homogeneous sampling sequence.

[0168] For details on the implementation of step S207, please refer to the Figure 3 The description of step S104 in the corresponding embodiment will not be repeated here.

[0169] In summary, the embodiment of the present application discloses a video vectorization method based on graph representation learning, which aims to utilize a variety of heterogeneous information, such as video browsing users, video publishing accounts, video tags and videos, etc., as well as the correlation between a variety of heterogeneous information, to construct a heterogeneous graph, and then perform unsupervised graph representation learning based on the heterogeneous graph. For example, the two random walk algorithms of metapath2vec and node2vec are integrated to perform deepwalk on the heterogeneous graph to generate a random sampling sequence, and skipgram is executed on the random sampling sequence results to learn the topological structure of the network heterogeneous graph to obtain a vectorized expression of the video. The video feature vector obtained by the above method contains the topological relationship between a variety of heterogeneous information, so the vector can be applied to a variety of downstream business scenarios, such as video clustering, recommended video recall, associated video recommendation, etc. It can assist the operation side in exploring new video topics, and can also be connected to the recommendation system to improve consumption indicators such as clicks, duration, and retention.

[0170] In an embodiment of the present application, a video feature vector is generated based on a heterogeneous sampling sequence and a homogeneous sampling sequence. Since the heterogeneous sampling sequence contains at least two identification nodes belonging to different data attribute types, the video feature vector may include the association relationship between identification nodes belonging to different data attribute types. Similarly, since the homogeneous sampling sequence contains at least two identification nodes belonging to the same data attribute type, the video feature vector may include the association relationship between identification nodes determined by associated heterogeneous information; the data attribute type of the above-mentioned associated heterogeneous information is different from the data attribute type of the video. Therefore, it can be seen that the video feature vector in the present application can cover the features of the associated heterogeneous information associated with the video, as well as the features of the association relationship between the video and the associated heterogeneous information, that is, the video feature vector contains multivariate information features; if the video feature vector in the present application is applied to actual scenarios, such as video recommendation, video recall or video clustering, because it contains diversified information features, it will be able to accurately characterize the association relationship between the video and the actual application scenario, so in the actual application scenario, it can improve the application accuracy of the video. In addition, this application avoids the existing supervised training method by constructing a heterogeneous graph to learn the topological structure relationship between videos and associated heterogeneous information, making the application range of video feature vectors wider.

[0171] Further, see Figure 7 , Figure 7 This is a flow chart of a data processing method provided by an embodiment of the present application. Figure 7 As shown, the data processing process may include the following steps S2031-S2033, and steps S2031-S2033 are Figure 4 A specific embodiment of step S203 in the corresponding embodiment.

[0172] Step S2031 : Determine a first associated edge between video identification nodes and an edge weight of the first associated edge based on the association relationship between each two videos.

[0173] Specifically, the video includes a first video and a second video; the video identification node includes a first video identification node corresponding to the first video and a second video identification node corresponding to the second video; obtain valid video sequences associated with N video browsing users respectively; N is a positive integer; N valid video sequences include valid video sequences L x , x is a positive integer and x is less than or equal to the total number of N valid video sequences; valid video sequence L x The valid videos in the sequence are sorted according to the time sequence of the associated video browsing users browsing the videos; the ratio of the effective browsing time of the valid videos to the total video time of the valid videos is greater than the browsing ratio threshold; if the first video and the second video are respectively in the valid video sequence L xThe positions in are adjacent positions, then the valid video sequence L is determined x The first video and the second video have an adjacent position relationship; among N valid video sequences, the valid video sequence with the adjacent position relationship is determined as an associated valid video sequence; the associated valid video sequence is used to characterize the existence of a first associated edge between the first video identification node and the second video identification node; the number of associated sequences of the associated valid video sequence is counted, and the number of associated sequences is determined as the edge weight of the first associated edge.

[0174] Please also see Figure 8 , Figure 8 This is a schematic diagram of a data processing scenario provided by an embodiment of the present application. Figure 8 As shown, the server obtains the valid video sequences associated with N video browsing users. Figure 8 In the example, let N be 2, that is, 2 valid video sequences, respectively Figure 8 Valid video sequence 1 and valid video sequence 2. Valid video sequence 1 is for video browsing user 70a, and valid video sequence 2 is for video browsing user 70c. Video browsing user 70a corresponds to video browsing user group 70d (including one or more video browsing users with attributes related to video browsing user 70a, for example, all are 20-year-old females from Shenzhen), and video browsing user 70c corresponds to video browsing user group 70e (including one or more video browsing users with attributes related to video browsing user 70c, for example, all are 30-year-old females from Nanjing).

[0175] Among them, the effective video sequence 1 includes 3 effective videos, such as Figure 8 As shown, they are video 701b, video 702b and video 703b respectively. Video browsing user 70a browses video 701b at 19:15 and watches the entire video. Video 701b is published by video publishing account 200000, and the video tags it carries include funny and movie. Video browsing user 70a browses video 702b at 19:00 and only watches the first half of the video content. Video 702b is published by video publishing account 200000, and the video tags it carries include funny and movie. Video browsing user 70a browses video 703b at 18:36, and the effective time for watching video 703b is 25 seconds. Video 703b is published by video publishing account 100000, and the video tags it carries include sports. Valid video sequence 2 also includes 3 valid videos, such as Figure 8 As shown, they are video 701b, video 704b and video 702b respectively. The basic situation of each video is the same as the basic situation of the video in the valid video sequence 1, so they are not described here one by one. You can refer to the description of the video in the valid video sequence 1 and Figure 8 .

[0176] It is understandable that the embodiment of the present application only takes two valid video sequences as an example. The total number of valid video sequences for actually constructing a heterogeneous graph can be any number and is not limited here. After the server obtains the above two valid video sequences, it can construct a heterogeneous graph based on the two valid video sequences. First, it determines a variety of information for constructing identification nodes. The embodiment of the present application uses videos, video browsing users, video publishing accounts, and video tags as examples to illustrate heterogeneous information. After determining a variety of heterogeneous information, the server performs identification processing on the various heterogeneous information, such as Figure 8 As shown, the server identifies video 701b as video identifier vid1, determines video identifier vid1 as video identifier node 1, identifies video 702b as video identifier vid2, determines video identifier vid2 as video identifier node 2, identifies video 703b as video identifier vid3, determines video identifier vid3 as video identifier node 3, identifies video 704b as video identifier vid4, and determines video identifier vid4 as video identifier node 4. The server identifies video browsing user group 70d as user identifier gid5, determines user identifier gid5 as user identifier node 5, identifies video browsing user group 70e as user identifier gid6, and determines user identifier gid6 as user identifier node 6. The processing process of video publishing accounts and video tags is the same as above, so they will not be repeated one by one. Please refer to Figure 8 After obtaining the identification nodes of the heterogeneous information, the server determines the associated edges between the identification nodes and the edge weights of the associated edges based on the association relationship between the heterogeneous information.

[0177] The embodiments of this application are Figure 8 The video 701b in the example is regarded as the first video and the video 702b is regarded as the second video. The relationship between other videos can refer to the description of video 701b and video 702b below. Therefore, the video identification node can include video identification node 1 corresponding to video 701b and video identification node 2 corresponding to video 702b. Figure 8 There are two valid video sequences in the video sequence. Obviously, the positions of video 701b and video 702b in valid video sequence 1 are adjacent, so the server can determine that valid video type 1 is video 701b and the associated valid video sequence of 701b; in valid video sequence 2, video 701b and video 702b are not adjacent, so it can be determined that valid video type 2 is not video 701b and the associated valid video sequence of 701b; the server counts the number of associated sequences of the associated valid video sequence, that is, 1, and determines 1 as the edge weight of the first associated edge between video identification node 1 and video identification node 2.

[0178] For the sake of beauty and clarity, the embodiments of this application are Figure 8 In the constructed heterogeneous graph, if the edge weight of the associated edge is 1, it is not depicted in the heterogeneous graph. Only edge weights greater than 1 are depicted.

[0179] Step S2032: Determine the second associated edge between the tag identification nodes and the edge weight of the second associated edge according to the association relationship between each two video tags.

[0180] Specifically, the associated heterogeneous information includes at least two video tags; the heterogeneous identification node includes a first tag identification node corresponding to the first video tag, and a second tag identification node corresponding to the second video tag; if the same video exists between the video associated with the first video tag and the video associated with the second video tag, the same video is determined as an associated video; the associated video is used to characterize the existence of a second associated edge between the first tag identification node and the second tag identification node; the number of videos of the associated video is counted, and the number of videos is determined as the edge weight of the second associated edge.

[0181] It is understandable that the associated heterogeneous information can be any information associated with the video, and the present application embodiment does not limit this. For the sake of ease of description and understanding, the present application embodiment takes the video tag as an example, and the heterogeneous identification node can include the tag identification node. Please refer to Figure 8 , Figure 8 Three video tags are given as examples: sports, movies, and funny. Figure 8 From the four videos in the figure, namely video 701b, video 702b, video 703b and video 704b, it can be seen that there are the same videos between the two video tags of funny and movie, namely video 701b and video 702b. Therefore, it can be determined that video 701b and video 702b are related videos of funny and movie. Furthermore, it can be determined that there is a second related edge between the label identification node 12 corresponding to funny and the label identification node 11 corresponding to movie, and the edge weight of the edge is the number of videos of the related video, namely 2.

[0182] Step S2033 : Determine a third associated edge between the video identification node and the heterogeneous identification node, and an edge weight of the third associated edge, based on the association relationship between the video and the associated heterogeneous information.

[0183] Specifically, the associated heterogeneous information includes a video browsing user group; the heterogeneous identification node includes a user identification node corresponding to the video browsing user group; within a video browsing cycle, the number of valid browsing times of the video by the video browsing user group is obtained; the valid browsing times refers to the number of times the video is effectively browsed by the video browsing users in the video browsing user group; if the valid browsing times is greater than the valid browsing times threshold, it is determined that a third associated edge exists between the video identification node and the user identification node; in the video browsing user group, the video browsing users who have effectively browsed the video are determined as the associated video browsing users of the video, and the number of users of the associated video browsing users is determined as the edge weight of the third associated edge.

[0184] In this step, assume that the video browsing user group is a group of 20-year-old female video browsers in Shenzhen, that is, the group may include one or more 20-year-old female video browsers located in Shenzhen; and the video browsing cycle is 1 week. Then, count the number of effective views of videos by 20-year-old female video browsers located in Shenzhen within a week. If video browser user a has 2 effective views of video q within a week, video browser user b has 1 effective view of video q within a week, and video browser user c has 1 effective view of video q within a week, and the other video browsers in the video browsing user group have not viewed video q within a week, or have ineffectively viewed video q, then the number of effective views of video q by this video browsing user group within a week is 4.

[0185] If the effective browsing count threshold is equal to or greater than 4 (for example, 100), it is determined that there is no third association edge between the user identification node corresponding to the video browsing user group and the video identification node corresponding to video q; if the effective browsing count threshold is less than 4 (for example, 2), it is determined that there is a third association edge between the user identification node corresponding to the video browsing user group and the video identification node corresponding to video q. At this time, video browsing user a, video browsing user b, and video browsing user c are determined as associated video browsing users of video q. Further, the edge weight of the above-mentioned third association edge can be determined to be 3.

[0186] It is understandable that the above figures are only assumed for the sake of ease of understanding and description and have no actual effect.

[0187] Optionally, the associated heterogeneous information includes at least two video accounts; the association relationship includes an account association relationship; the heterogeneous identification node includes account identification nodes corresponding to at least two video accounts respectively; among the at least two video accounts, an associated video account that has an account association relationship with the video is obtained; the account association relationship is used to characterize that the video publishing user publishes the video through the associated video account; the account identification node corresponding to the associated video account is determined as the associated account identification node; there is a third association edge between the video identification node and the associated account identification node, and the edge weight of the third association edge is a constant parameter.

[0188] The above video account can be a video publishing account, that is, the account that publishes the video. The server obtains the video publishing account of each video, please refer to Figure 8 , video publishing account 200000 has published video 701b and video 702b, then there is a third associated edge between the account identification node 8 corresponding to video publishing account 200000 and the video identification node 1 corresponding to video 701b, and the edge weight of this edge can be a constant parameter, such as 1; there is a third associated edge between the account identification node 8 corresponding to video publishing account 200000 and the video identification node 2 corresponding to video 702b. Video publishing account 100000 has published video 703b, then there is a third associated edge between the account identification node 7 corresponding to video publishing account 100000 and the video identification node 3 corresponding to video 703b, and the edge weight of this edge is 1; video publishing account 300000 has published video 704b, then there is a third associated edge between the account identification node 9 corresponding to video publishing account 300000 and the video identification node 4 corresponding to video 704b, and the edge weight of this edge is 1. The topological relationship between the above video identification nodes and account identification nodes is as follows. Figure 8 As shown in the heterogeneous diagram.

[0189] Optionally, the associated heterogeneous information includes at least two video tags; the association relationship includes a tag association relationship; the heterogeneous identification node includes tag identification nodes corresponding to at least two video tags respectively; among the at least two video tags, an associated video tag that has a tag association relationship with the video is obtained; the tag association relationship is used to characterize that the video is marked with an associated video tag; the tag identification node corresponding to the associated video tag is determined as the associated tag identification node; a third association edge exists between the video identification node and the associated tag identification node, and the edge weight of the third association edge is a constant parameter.

[0190] Please see again Figure 8The server obtains the video tags carried by each video. Video 701b carries two video tags, namely, funny and movie. Then, there is a third associated edge between the tag identification node 12 corresponding to funny and the video identification node 1 corresponding to video 701b. The edge weight of this edge can be a constant parameter, such as 1; there is a third associated edge between the tag identification node 11 corresponding to movie and the video identification node 1 corresponding to video 701b, and the edge weight of this edge is 1; video 702b carries two video tags, namely, funny and movie. Then, there is a third associated edge between the tag identification node 12 corresponding to funny and the video identification node 2 corresponding to video 702b. The edge weight of this edge is 1; there is a third associated edge between the label identification node 11 corresponding to the movie and the video identification node 2 corresponding to video 702b, and the edge weight of this edge is 1; video 703b carries a sports video label, then there is a third associated edge between the label identification node 10 corresponding to sports and the video identification node 3 corresponding to video 703b, and the edge weight of this edge is 1; video 704b carries a sports video label, then there is a third associated edge between the label identification node 10 corresponding to sports and the video identification node 4 corresponding to video 704b, and the edge weight of this edge is 1; the topological relationship between the above video identification nodes and label identification nodes is as follows Figure 8 As shown in the heterogeneous diagram.

[0191] In an embodiment of the present application, a video feature vector is generated based on a heterogeneous sampling sequence and a homogeneous sampling sequence. Since the heterogeneous sampling sequence contains at least two identification nodes belonging to different data attribute types, the video feature vector may include the association relationship between identification nodes belonging to different data attribute types. Similarly, since the homogeneous sampling sequence contains at least two identification nodes belonging to the same data attribute type, the video feature vector may include the association relationship between identification nodes determined by associated heterogeneous information; the data attribute type of the above-mentioned associated heterogeneous information is different from the data attribute type of the video. Therefore, it can be seen that the video feature vector in the present application can cover the features of the associated heterogeneous information associated with the video, as well as the features of the association relationship between the video and the associated heterogeneous information, that is, the video feature vector contains multivariate information features; if the video feature vector in the present application is applied to actual scenarios, such as video recommendation, video recall or video clustering, because it contains diversified information features, it will be able to accurately characterize the association relationship between the video and the actual application scenario, so in the actual application scenario, it can improve the application accuracy of the video. In addition, this application avoids the existing supervised training method by constructing a heterogeneous graph to learn the topological structure relationship between videos and associated heterogeneous information, making the application range of video feature vectors wider.

[0192] Further, see Figure 9 , Figure 9This is a structural diagram of a data processing device provided in an embodiment of the present application. The above-mentioned data processing device can be a computer program (including program code) running on a computer device, for example, the data processing device is an application software; the device can be used to execute the corresponding steps of the method provided in an embodiment of the present application. Figure 9 As shown, the data processing device 1 may include: a data acquisition module 11 , a first generation module 12 , a sampling node module 13 and a second generation module 14 .

[0193] The data acquisition module 11 is used to acquire the video identifier of the video and the associated heterogeneous information associated with the video; the data attribute type of the video is different from the data attribute type of the associated heterogeneous information;

[0194] A first generating module 12 is configured to determine the video identifier and the heterogeneous information identifier of the associated heterogeneous information as identification nodes, and generate a heterogeneous graph including the identification nodes;

[0195] The sampling node module 13 is used to perform identification node sampling on the heterogeneous graph to obtain a heterogeneous sampling sequence and a homogeneous sampling sequence; the heterogeneous sampling sequence includes at least two identification nodes belonging to different data attribute types, and the homogeneous sampling sequence includes at least two identification nodes belonging to the same data attribute type;

[0196] The second generating module 14 is configured to generate a video feature vector corresponding to the video identifier according to the heterogeneous sampling sequence and the homogeneous sampling sequence.

[0197] The specific functional implementation of the data acquisition module 11, the first generation module 12, the sampling node module 13 and the second generation module 14 can be found in the above Figure 3 Steps S101 to S104 in the corresponding embodiment are not described again here.

[0198] See also Figure 9 , the number of videos is at least two, and the number of associated heterogeneous information is at least two;

[0199] The first generating module 12 may include a first determining unit 121 and a first generating unit 122 .

[0200] A first determining unit 121 is configured to determine associated edges between identification nodes and edge weights of associated edges based on an associated relationship between each two videos, an associated relationship between each two pieces of associated heterogeneous information, and an associated relationship between a video and associated heterogeneous information;

[0201] The first generating unit 122 is configured to generate a heterogeneous graph according to the identified nodes, associated edges, and edge weights of the associated edges.

[0202] The specific functional implementation of the first determining unit 121 and the first generating unit 122 can be found in the above Figure 3 The step S102 in the corresponding embodiment will not be described again here.

[0203] See also Figure 9 , the identification node includes a video identification node belonging to the video identification and a heterogeneous identification node belonging to the heterogeneous information identification; the association edge includes a first association edge, a second association edge and a third association edge;

[0204] The first determining unit 121 may include a first determining subunit 1211 , a second determining subunit 1212 , and a third determining subunit 1213 .

[0205] A first determining subunit 1211 is configured to determine a first associated edge between video identification nodes and an edge weight of the first associated edge based on an association relationship between each two videos;

[0206] The second determining subunit 1212 is configured to determine a second associated edge between heterogeneous identification nodes and an edge weight of the second associated edge according to the association relationship between each two associated heterogeneous information;

[0207] The third determining subunit 1213 is configured to determine a third associated edge between the video identification node and the heterogeneous identification node, and an edge weight of the third associated edge, according to the association relationship between the video and the associated heterogeneous information.

[0208] The specific functional implementation of the first determining subunit 1211, the second determining subunit 1212 and the third determining subunit 1213 can be found in the above Figure 7 Steps S1021 to S1023 in the corresponding embodiment are not described again here.

[0209] See also Figure 9 , the video includes a first video and a second video; the video identification node includes a first video identification node corresponding to the first video, and a second video identification node corresponding to the second video;

[0210] The first determining subunit 1211 may include: an acquiring sequence subunit 12111 , a position determining subunit 12112 , a determining sequence subunit 12113 , and a counting sequence subunit 12114 .

[0211] The acquisition sequence subunit 12111 is used to acquire valid video sequences associated with N video browsing users respectively; N is a positive integer; the N valid video sequences include valid video sequences L x , x is a positive integer and x is less than or equal to the total number of N valid video sequences; valid video sequence L xThe valid videos in the list are sorted according to the time sequence of the associated video browsing users; the ratio of the effective browsing time of the valid videos to the total video time of the valid videos is greater than the browsing ratio threshold;

[0212] The position determination subunit 12112 is used to determine the position of the first video and the second video if they are in the valid video sequence L x The positions in are adjacent positions, then the valid video sequence L is determined x The first video and the second video have an adjacent position relationship;

[0213] The sequence determination subunit 12113 is configured to determine, among the N valid video sequences, valid video sequences having adjacent positional relationships as associated valid video sequences; the associated valid video sequences are used to indicate that a first associated edge exists between the first video identification node and the second video identification node;

[0214] The sequence counting subunit 12114 is configured to count the number of associated sequences associated with the valid video sequence, and determine the number of associated sequences as the edge weight of the first associated edge.

[0215] The specific functional implementation of the acquisition sequence subunit 12111, the determination position subunit 12112, the determination sequence subunit 12113 and the statistical sequence subunit 12114 can be found in the above Figure 7 Step S1021 in the corresponding embodiment will not be described again here.

[0216] See also Figure 9 , the associated heterogeneous information includes first associated heterogeneous information and second associated heterogeneous information; the heterogeneous identification node includes a first heterogeneous identification node corresponding to the first associated heterogeneous information, and a second heterogeneous identification node corresponding to the second associated heterogeneous information;

[0217] The second determining subunit 1212 may include: a video determining subunit 12121 and a video counting subunit 12122 .

[0218] The video determination subunit 12121 is configured to determine the same video as the associated video if the video associated with the first associated heterogeneous information and the video associated with the second associated heterogeneous information are identical. The associated video is used to indicate that a second associated edge exists between the first heterogeneous identification node and the second heterogeneous identification node.

[0219] The video counting sub-unit 12122 is used to count the number of videos of the associated video and determine the number of videos as the edge weight of the second associated edge.

[0220] The specific functional implementation of the video subunit 12121 and the statistical video subunit 12122 can be found in the above Figure 7 Step S1022 in the corresponding embodiment will not be described again here.

[0221] See also Figure 9 , the associated heterogeneous information includes a video browsing user group; the heterogeneous identification node includes a user identification node corresponding to the video browsing user group;

[0222] The third determining subunit 1213 may include: an acquiring times subunit 12131 , an association determining subunit 12132 , and a weight determining subunit 12133 .

[0223] The acquisition times subunit 12131 is used to obtain the effective browsing times of the video by the video browsing user group within the video browsing cycle; the effective browsing times refer to the number of times the video is effectively browsed by the video browsing users in the video browsing user group;

[0224] The association determination subunit 12132 is configured to determine that a third association edge exists between the video identification node and the user identification node if the valid browsing count is greater than the valid browsing count threshold;

[0225] The weight determination subunit 12133 is used to determine the video browsing users who have effectively browsed the video in the video browsing user group as the associated video browsing users of the video, and determine the number of users of the associated video browsing users as the edge weight of the third associated edge.

[0226] The specific functional implementation of the acquisition times subunit 12131, the determination association subunit 12132 and the determination weight subunit 12133 can be found in the above Figure 7 Step S1023 in the corresponding embodiment will not be described again here.

[0227] See also Figure 9 , the associated heterogeneous information includes at least two video accounts; the associated relationship includes an account associated relationship; the heterogeneous identification node includes an account identification node corresponding to at least two video accounts respectively;

[0228] The third determining subunit 1213 may include: an acquiring account subunit 12134 and an determining account subunit 12135 .

[0229] The account acquisition subunit 12134 is used to acquire, from at least two video accounts, an associated video account that has an account association relationship with the video; the account association relationship is used to indicate that the video publishing user publishes the video through the associated video account;

[0230] The account determination subunit 12135 is used to determine the account identification node corresponding to the associated video account as the associated account identification node; there is a third association edge between the video identification node and the associated account identification node, and the edge weight of the third association edge is a constant parameter.

[0231] The specific functional implementation of the acquisition account subunit 12134 and the determination account subunit 12135 can be found in the above Figure 7 Step S1023 in the corresponding embodiment will not be described again here.

[0232] See also Figure 9 , the associated heterogeneous information includes at least two video tags; the association relationship includes a tag association relationship; the heterogeneous identification node includes a tag identification node corresponding to at least two video tags respectively;

[0233] The third determining subunit 1213 may include: an acquiring tag subunit 12136 and a determining tag subunit 12137 .

[0234] The tag acquisition subunit 12136 is configured to acquire, from at least two video tags, an associated video tag that has a tag association relationship with the video; the tag association relationship is used to indicate that the video is tagged with the associated video tag;

[0235] The label determination subunit 12137 is used to determine the label identification node corresponding to the associated video label as the associated label identification node; there is a third associated edge between the video identification node and the associated label identification node, and the edge weight of the third associated edge is a constant parameter.

[0236] The specific functional implementation of the obtaining label subunit 12136 and the determining label subunit 12137 can be found in the above Figure 7 Step S1023 in the corresponding embodiment will not be described again here.

[0237] See also Figure 9 The sampling node module 13 may include: a first acquiring unit 131 , a first sampling unit 132 , and a second sampling unit 133 .

[0238] The first acquisition unit 131 is used to acquire a random sampling heterogeneous path and a random sampling homogeneous path; the random sampling heterogeneous path is used to indicate the type sampling order of the data attribute types sampled when sampling identification nodes of different data attribute types; the random sampling homogeneous path is used to indicate the data attribute types sampled when sampling identification nodes of the same data attribute type;

[0239] A first sampling unit 132 is configured to randomly sample the identified nodes in the heterogeneous graph according to the type sampling order indicated by the randomly sampled heterogeneous path to obtain a heterogeneous sampling sequence;

[0240] The second sampling unit 133 is configured to determine the data attribute type indicated by the randomly sampled isomorphic path as the target data attribute type, and randomly sample the identification nodes belonging to the target data attribute type in the heterogeneous graph to obtain a isomorphic sampling sequence.

[0241] The specific functional implementation of the first acquisition unit 131, the first sampling unit 132 and the second sampling unit 133 can be found in the above Figure 4 Steps S204 to S206 in the corresponding embodiment are not described again here.

[0242] See also Figure 9 The first sampling unit 132 may include: a fourth determining subunit 1321 , a sampling target subunit 1322 , a adding node subunit 1323 , and a generating sequence subunit 1324 .

[0243] The fourth determining subunit 1321 is configured to determine, according to a type sampling order, the j-th data attribute type to be sampled in the random sampling heterogeneous path as the data attribute type to be sampled; j is a positive integer less than or equal to S, and S is the total number of nodes of the identification nodes to be sampled based on the random sampling heterogeneous path;

[0244] The sampling target subunit 1322 is configured to sample a target identification node from the heterogeneous graph according to the sampled node set and the attribute type of the data to be sampled; the data attribute type to which the target identification node belongs is the attribute type of the data to be sampled; the sampled node set includes the sampled identification node;

[0245] The adding node subunit 1323 is used to add the target identification node to the sampled node set if j is less than S;

[0246] The sequence generation subunit 1324 is configured to generate a heterogeneous sampling sequence according to the sampled node set and the target identification node if j is equal to S, where the target identification node is the last identification node in the heterogeneous sampling sequence.

[0247] The specific functional implementation of the fourth determination subunit 1321, the sampling target subunit 1322, the adding node subunit 1323 and the generating sequence subunit 1324 can be referred to above. Figure 4 Step S205 in the corresponding embodiment will not be described again here.

[0248] See also Figure 9 , the sampled node set includes a previous neighbor identification node, and the previous neighbor identification node is the last identification node in the sampled node set;

[0249] The sampling target subunit 1322 may include: an acquisition node subunit 13221 , a sum weight subunit 13222 , a determination probability subunit 13223 , and a sampling node subunit 13224 .

[0250] The node acquisition subunit 13221 is used to acquire, from the heterogeneous graph, w target preselected identification nodes that have associated edges with the preceding neighbor identification node according to the attribute type of the data to be sampled; wherein w is a positive integer; and the data attribute type of the w target preselected identification nodes is the attribute type of the data to be sampled;

[0251] The summing weight subunit 13222 is used to obtain the edge weights of the associated edges between the previous neighbor identification node and each target pre-selected identification node, and sum the obtained edge weights to obtain the total edge weight; the w target pre-selected identification nodes include the target pre-selected identification node Y m , where m is a positive integer and m is less than or equal to w;

[0252] Determine probability subunit 13223, used to compare the previous neighbor identification node with the target pre-selected identification node Y m The edge weight Z of the associated edge between m , and the ratio of the total edge weight, is determined as the target pre-selected identification node Y m The random sampling probability of

[0253] The sampling node subunit 13224 is used to randomly sample w target pre-selected identification nodes according to the random sampling probability corresponding to each target pre-selected identification node to obtain a target identification node; in the heterogeneous sampling sequence, the previous identification node is the previous identification node of the target identification node.

[0254] The specific functional implementation of the acquisition node subunit 13221, the sum weight subunit 13222, the determination probability subunit 13223 and the sampling node subunit 13224 can be found in the above Figure 4 Step S205 in the corresponding embodiment will not be described again here.

[0255] See also Figure 9 , the homogeneous sampling sequence includes a first homogeneous sampling sequence and a second homogeneous sampling sequence;

[0256] The second sampling unit 133 may include a first generating subunit 1331 and a second generating subunit 1332 .

[0257] The first generating subunit 1331 is configured to randomly sample the video identification nodes in the heterogeneous graph to generate a first homogeneous sampling sequence including at least two video identification nodes if the target data attribute type is the data attribute type corresponding to the video identification node; the first homogeneous sampling sequence is used to represent the topological relationship between the video identification nodes in the heterogeneous graph;

[0258] The second generating subunit 1332 is used to randomly sample the heterogeneous identification nodes in the heterogeneous graph if the target data attribute type is the data attribute type corresponding to the heterogeneous identification node, and generate a second homogeneous sampling sequence including at least two heterogeneous identification nodes; the second homogeneous sampling sequence is used to characterize the topological relationship between the heterogeneous identification nodes in the heterogeneous graph.

[0259] The specific functional implementation of the first generation subunit 1331 and the second generation subunit 1332 can be found in the above Figure 4 Step S206 in the corresponding embodiment will not be described again here.

[0260] See also Figure 9 The second generating module 14 may include: a second determining unit 141 , a second acquiring unit 142 , a model adjusting unit 143 and a second generating unit 144 .

[0261] The second determining unit 141 is configured to determine the heterogeneous sampling sequence and the homogeneous sampling sequence into at least two random sampling sequences; the at least two random sampling sequences include the random sampling sequence S a , a is a positive integer and a is less than or equal to the total number of sequences of at least two random sampling sequences;

[0262] The second acquisition unit 142 is used to acquire the random sampling sequence S a The true encoding label of each identified node in the random sampling sequence S a Including identification node D b , and the node D b There are adjacent identification nodes with position association relationship, b is a positive integer and b is less than or equal to the random sampling sequence S a The total number of nodes that identify nodes in the graph; the vector dimension of each true encoding label is equal to the total number of nodes that identify nodes in the heterogeneous graph;

[0263] The second acquiring unit 142 is further configured to identify the node D b The true encoding label C b Input into the initial word encoding model to obtain the predicted coding labels of adjacent identification nodes;

[0264] The model adjustment unit 143 is used to adjust the model parameters in the initial word encoding model according to the real encoding labels of the adjacent identification nodes and the predicted encoding labels of the adjacent identification nodes to obtain the target word encoding model;

[0265] The second generating unit 144 is configured to input the video identifier into the target word encoding model to obtain a video feature vector corresponding to the video identifier.

[0266] The specific functional implementation of the second determining unit 141, the second acquiring unit 142, the adjusting model unit 143 and the second generating unit 144 can be referred to above. Figure 3 The corresponding step S104 in the embodiment will not be described again here.

[0267] In an embodiment of the present application, a video feature vector is generated based on a heterogeneous sampling sequence and a homogeneous sampling sequence. Since the heterogeneous sampling sequence contains at least two identification nodes belonging to different data attribute types, the video feature vector may include the association relationship between identification nodes belonging to different data attribute types. Similarly, since the homogeneous sampling sequence contains at least two identification nodes belonging to the same data attribute type, the video feature vector may include the association relationship between identification nodes determined by associated heterogeneous information; the data attribute type of the above-mentioned associated heterogeneous information is different from the data attribute type of the video. Therefore, it can be seen that the video feature vector in the present application can cover the features of the associated heterogeneous information associated with the video, as well as the features of the association relationship between the video and the associated heterogeneous information, that is, the video feature vector contains multivariate information features; if the video feature vector in the present application is applied to actual scenarios, such as video recommendation, video recall or video clustering, because it contains diversified information features, it will be able to accurately characterize the association relationship between the video and the actual application scenario, so in the actual application scenario, it can improve the application accuracy of the video. In addition, this application avoids the existing supervised training method by constructing a heterogeneous graph to learn the topological structure relationship between videos and associated heterogeneous information, making the application range of video feature vectors wider.

[0268] Further, see Figure 10 , Figure 10 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. Figure 10 As shown, the computer device 1000 can be the above Figure 3Corresponding to the server in the embodiment, the computer device 1000 may include: at least one processor 1001, such as a CPU, at least one network interface 1004, a user interface 1003, a memory 1005, and at least one communication bus 1002. The communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), a keyboard (Keyboard), and the network interface 1004 may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory, or a non-volatile memory (non-volatile memory), such as at least one disk storage. The memory 1005 may optionally also be at least one storage device located away from the aforementioned processor 1001. As Figure 10 As shown, the memory 1005 as a computer storage medium may include an operating system, a network communication module, a user interface module, and a device control application.

[0269] exist Figure 10 In the computer device 1000 shown, the network interface 1004 can provide network communication functions; the user interface 1003 is mainly used to provide an interface for user input; and the processor 1001 can be used to call the device control application stored in the memory 1005 to achieve:

[0270] Obtaining a video identifier of a video and associated heterogeneous information associated with the video; the data attribute type of the video is different from the data attribute type of the associated heterogeneous information;

[0271] Determine the video identifier and the heterogeneous information identifier of the associated heterogeneous information as identification nodes, and generate a heterogeneous graph including the identification nodes;

[0272] Sampling identification nodes of the heterogeneous graph to obtain a heterogeneous sampling sequence and a homogeneous sampling sequence; the heterogeneous sampling sequence includes at least two identification nodes belonging to different data attribute types, and the homogeneous sampling sequence includes at least two identification nodes belonging to the same data attribute type;

[0273] A video feature vector corresponding to the video identifier is generated according to the heterogeneous sampling sequence and the homogeneous sampling sequence.

[0274] It should be understood that the computer device 1000 described in the embodiment of the present application can execute the above Figure 3 、 Figure 4 as well as Figure 7 The description of the data processing method in the corresponding embodiment can also be performed as described above. Figure 9 The description of the data processing device 1 in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of adopting the same method will not be repeated here either.

[0275] The present invention also provides a computer-readable storage medium that stores a computer program. The computer program includes program instructions that are executed by a processor to implement Figure 3 、 Figure 4 as well as Figure 7 The data processing methods provided in each step are detailed in the above Figure 3 、 Figure 4 as well as Figure 7 The implementation methods provided by each step will not be described in detail here. In addition, the description of the beneficial effects of adopting the same method will not be described in detail here either.

[0276] The computer-readable storage medium can be the data processing apparatus provided in any of the aforementioned embodiments, or the internal storage unit of the computer device, such as the computer device's hard drive or memory. The computer-readable storage medium can also be an external storage device of the computer device, such as a plug-in hard drive, a smart media card (SMC), a secure digital (SD) card, a flash memory card, etc. Furthermore, the computer-readable storage medium can include both the internal storage unit of the computer device and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium can also be used to temporarily store data that has been output or is about to be output.

[0277] The present application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, so that the computer device can execute the aforementioned Figure 3 、 Figure 4 as well as Figure 7 The description of the data processing method in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of adopting the same method will not be repeated here either.

[0278] The terms "first", "second", etc. in the description, claims, and drawings of the embodiments of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product, or device comprising a series of steps or units is not limited to the listed steps or modules, but may optionally include steps or modules not listed, or may optionally include other step units inherent to these processes, methods, apparatuses, products, or devices.

[0279] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0280] The method and related apparatus provided in the embodiment of the present application are described with reference to the method flow chart and / or structural diagram provided in the embodiment of the present application, and specifically can be implemented by computer program instructions for each process and / or box of the method flow chart and / or structural diagram, and the combination of the process and / or box in the flow chart and / or block diagram. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the function specified in one process or multiple processes of the flow chart and / or one box or multiple boxes of the structural diagram. These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device, and the instruction device implements the function specified in one process or multiple processes of the flow chart and / or one box or multiple boxes of the structural diagram. These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the structural diagram.

[0281] The above disclosure is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.

Claims

1. A data processing method, characterized in that: include: Obtaining a video identifier of a video and associated heterogeneous information associated with the video; The data attribute type of the video is different from the data attribute type of the associated heterogeneous information; Determine the video identifier and the heterogeneous information identifier of the associated heterogeneous information as identification nodes, and generate a heterogeneous graph including the identification nodes; Obtaining a random sampling heterogeneous path and a random sampling homogeneous path; the random sampling heterogeneous path is used to indicate the type sampling order of the data attribute types sampled when sampling identification nodes of different data attribute types; the random sampling homogeneous path is used to indicate the data attribute types sampled when sampling identification nodes of the same data attribute type; Randomly sampling the identification nodes in the heterogeneous graph according to the type sampling order indicated by the random sampling heterogeneous path to obtain a heterogeneous sampling sequence; the heterogeneous sampling sequence includes at least two identification nodes belonging to different data attribute types; Determining the data attribute type indicated by the randomly sampled isomorphic path as a target data attribute type, and randomly sampling the identification nodes belonging to the target data attribute type in the heterogeneous graph to obtain an isomorphic sampling sequence; the isomorphic sampling sequence includes at least two identification nodes belonging to the same data attribute type; A video feature vector corresponding to the video identifier is generated according to the heterogeneous sampling sequence and the homogeneous sampling sequence.

2. The method according to claim 1, characterized in that The number of the videos is at least two, and the number of the associated heterogeneous information is at least two; The generating of a heterogeneous graph including the identified node includes: Determining the associated edges between the identified nodes and the edge weights of the associated edges according to the associated relationship between each two videos, the associated relationship between each two pieces of associated heterogeneous information, and the associated relationship between the videos and the associated heterogeneous information; The heterogeneous graph is generated according to the identified nodes, the associated edges, and the edge weights of the associated edges.

3. The method according to claim 2, characterized in that The identification node includes a video identification node belonging to the video identification and a heterogeneous identification node belonging to the heterogeneous information identification; the associated edge includes a first associated edge, a second associated edge and a third associated edge; The determining, based on the association relationship between each two videos, the association relationship between each two associated heterogeneous information, and the association relationship between the videos and the associated heterogeneous information, of the association edges between the identified nodes and the edge weights of the association edges includes: Determining the first associated edge between the video identification nodes and the edge weight of the first associated edge according to the association relationship between each two videos; Determining the second association edge between the heterogeneous identification nodes and the edge weight of the second association edge according to the association relationship between each two associated heterogeneous information; The third associated edge between the video identification node and the heterogeneous identification node, and an edge weight of the third associated edge are determined according to the association relationship between the video and the associated heterogeneous information.

4. The method according to claim 3, characterized in that The video includes a first video and a second video; the video identification node includes a first video identification node corresponding to the first video, and a second video identification node corresponding to the second video; The determining, based on the association relationship between each two videos, the first association edge between the video identification nodes and the edge weight of the first association edge includes: Get the valid video sequences associated with N video browsing users respectively; N is a positive integer; N valid video sequences include valid video sequence L x , x is a positive integer and x is less than or equal to the total number of the N valid video sequences; the valid video sequence L x The valid videos in the list are sorted according to the time sequence of the associated video browsing users browsing the videos; the ratio of the effective browsing time of the videos corresponding to the valid videos to the total video time of the valid videos is greater than the browsing ratio threshold; If the first video and the second video are respectively in the valid video sequence L x The positions in the video sequence L are adjacent positions, then the effective video sequence L is determined. x The first video and the second video have an adjacent position relationship; Among the N valid video sequences, a valid video sequence having the adjacent position relationship is determined as an associated valid video sequence; the associated valid video sequence is used to represent that the first associated edge exists between the first video identification node and the second video identification node; The number of associated sequences associated with the valid video sequence is counted, and the number of associated sequences is determined as the edge weight of the first associated edge.

5. The method according to claim 3, characterized in that The associated heterogeneous information includes first associated heterogeneous information and second associated heterogeneous information; the heterogeneous identification node includes a first heterogeneous identification node corresponding to the first associated heterogeneous information and a second heterogeneous identification node corresponding to the second associated heterogeneous information; The determining, based on the association relationship between each two associated heterogeneous information, the second association edge between the heterogeneous identification nodes and the edge weight of the second association edge includes: If the video associated with the first associated heterogeneous information and the video associated with the second associated heterogeneous information are identical, the identical videos are determined as associated videos; the associated videos are used to represent the existence of the second associated edge between the first heterogeneous identification node and the second heterogeneous identification node; The number of videos of the associated videos is counted, and the number of videos is determined as the edge weight of the second associated edge.

6. The method according to claim 3, characterized in that The associated heterogeneous information includes a video browsing user group; the heterogeneous identification node includes a user identification node corresponding to the video browsing user group; The determining, based on the association relationship between the video and the associated heterogeneous information, the third association edge between the video identification node and the heterogeneous identification node, and the edge weight of the third association edge, includes: During a video browsing period, obtaining the number of valid browsing times of the video by the video browsing user group; the valid browsing times refers to the number of times the video is effectively browsed by the video browsing users in the video browsing user group; If the valid browsing times are greater than the valid browsing times threshold, determining that the third association edge exists between the video identification node and the user identification node; In the video browsing user group, video browsing users who have effectively browsed the video are determined as associated video browsing users of the video, and the number of users of the associated video browsing users is determined as the edge weight of the third associated edge.

7. The method according to claim 3, characterized in that The associated heterogeneous information includes at least two video accounts; the associated relationship includes an account association relationship; the heterogeneous identification node includes the account identification nodes corresponding to the at least two video accounts respectively; The determining, based on the association relationship between the video and the associated heterogeneous information, the third association edge between the video identification node and the heterogeneous identification node, and the edge weight of the third association edge, includes: From the at least two video accounts, obtain an associated video account that has the account association relationship with the video; the account association relationship is used to indicate that the video publishing user publishes the video through the associated video account; The account identification node corresponding to the associated video account is determined as the associated account identification node; the third associated edge exists between the video identification node and the associated account identification node, and the edge weight of the third associated edge is a constant parameter.

8. The method according to claim 3, characterized in that The associated heterogeneous information includes at least two video tags; the associated relationship includes a tag association relationship; the heterogeneous identification node includes a tag identification node corresponding to the at least two video tags respectively; The determining, based on the association relationship between the video and the associated heterogeneous information, the third association edge between the video identification node and the heterogeneous identification node, and the edge weight of the third association edge, includes: From the at least two video tags, obtain an associated video tag that has the tag association relationship with the video; the tag association relationship is used to indicate that the video is marked with the associated video tag; The tag identification node corresponding to the associated video tag is determined as the associated tag identification node; the third associated edge exists between the video identification node and the associated tag identification node, and the edge weight of the third associated edge is a constant parameter.

9. The method according to claim 1, characterized in that The randomly sampling the identified nodes in the heterogeneous graph according to the type sampling order indicated by the randomly sampled heterogeneous path to obtain a heterogeneous sampling sequence includes: According to the type sampling order, the j-th data attribute type required to be sampled in the random sampling heterogeneous path is determined as the data attribute type to be sampled; j is a positive integer less than or equal to S, and S is the total number of nodes of the identification nodes required to be sampled based on the random sampling heterogeneous path; According to the sampled node set and the data attribute type to be sampled, a target identification node is sampled from the heterogeneous graph; the data attribute type to which the target identification node belongs is the data attribute type to be sampled; the sampled node set includes the sampled identification nodes; If j is less than S, the target identification node is added to the sampled node set; If j is equal to S, the heterogeneous sampling sequence is generated according to the sampled node set and the target identification node, and the target identification node is the last identification node in the heterogeneous sampling sequence.

10. The method according to claim 9, characterized in that The sampled node set includes a previous neighbor identification node, and the previous neighbor identification node is the last identification node in the sampled node set; The step of sampling a target identification node from the heterogeneous graph according to the sampled node set and the attribute type of the data to be sampled includes: According to the attribute type of the data to be sampled, w target pre-selected identification nodes having associated edges with the preceding identification node are obtained from the heterogeneous graph; wherein w is a positive integer; and the data attribute type of the w target pre-selected identification nodes is the attribute type of the data to be sampled; The edge weights of the associated edges between the preceding neighbor identification node and each target preselected identification node are obtained respectively, and the obtained edge weights are summed to obtain the total edge weight; the w target preselected identification nodes include the target preselected identification node Y m , where m is a positive integer and m is less than or equal to w; The preceding neighbor identification node and the target preselected identification node Y m The edge weight Z of the associated edge between m , and the ratio of the total edge weight, is determined as the target preselected identification node Y m The random sampling probability of The w target preselected identification nodes are randomly sampled according to the random sampling probability corresponding to each target preselected identification node to obtain the target identification node; in the heterogeneous sampling sequence, the previous identification node is the previous identification node of the target identification node.

11. The method according to claim 1, wherein The homogeneous sampling sequence includes a first homogeneous sampling sequence and a second homogeneous sampling sequence; The randomly sampling the identification nodes belonging to the target data attribute type in the heterogeneous graph to obtain a homogeneous sampling sequence includes: If the target data attribute type is a data attribute type corresponding to a video identification node, randomly sampling the video identification nodes in the heterogeneous graph to generate a first isomorphic sampling sequence including at least two video identification nodes; the first isomorphic sampling sequence is used to represent the topological relationship between the video identification nodes in the heterogeneous graph; If the target data attribute type is a data attribute type corresponding to a heterogeneous identification node, the heterogeneous identification nodes in the heterogeneous graph are randomly sampled to generate a second homogeneous sampling sequence including at least two heterogeneous identification nodes; the second homogeneous sampling sequence is used to characterize the topological relationship between the heterogeneous identification nodes in the heterogeneous graph.

12. The method according to claim 1, characterized in that Generating a video feature vector corresponding to the video identifier according to the heterogeneous sampling sequence and the homogeneous sampling sequence includes: The heterogeneous sampling sequence and the homogeneous sampling sequence are determined to be at least two random sampling sequences; the at least two random sampling sequences include a random sampling sequence S a , a is a positive integer and a is less than or equal to the total number of sequences of the at least two random sampling sequences; Get the random sampling sequence S a The true encoding label of each identified node in the random sampling sequence S a Including identification node D b , and the node D b There are adjacent identification nodes with position association relationship, b is a positive integer and b is less than or equal to the random sampling sequence S a The total number of nodes that identify the nodes in the heterogeneous graph; the vector dimension of each true encoding label is equal to the total number of nodes that identify the nodes in the heterogeneous graph; The identification node D b The true encoding label C b Input into the initial word encoding model to obtain the predicted encoding label of the adjacent identification node; Adjusting the model parameters in the initial word encoding model according to the real encoding labels of the adjacent identification nodes and the predicted encoding labels of the adjacent identification nodes to obtain a target word encoding model; The video identifier is input into the target word encoding model to obtain a video feature vector corresponding to the video identifier.

13. A computer device, characterized in that: include: processor, memory, and network interface; The processor is connected to the memory and the network interface, wherein the network interface is used to provide a data communication function, the memory is used to store a computer program, and the processor is used to call the computer program so that the computer device executes the method according to any one of claims 1 to 12.

14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which is suitable for being loaded and executed by a processor, so that a computer device having the processor executes the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Social recommendation method based on heterogeneous information and isomorphic information network fusion

    CN112182424A