Video recommendation method and device based on dynamic bipartite network representation
By capturing snapshots, filling virtual nodes, and capturing explicit and implicit features in a dynamic bipartite network, and transforming them into low-dimensional vector representations, the problems of inaccurate recommendation results and high computational resource consumption in dynamic bipartite network representation learning are solved, achieving efficient and accurate personalized video recommendation.
Patent Information
- Application Number
- CN202511066021.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-14
AI Technical Summary
Existing dynamic binary network representation learning methods have low accuracy in personalized recommendations and consume a lot of computational resources, making them unable to meet real-time requirements.
By capturing snapshots of the dynamic binary network, filling virtual nodes to maintain consistent node size, capturing explicit and implicit features, and converting them into low-dimensional vector representations, some of which are stored in a cache to save computational resources, a personalized video recommendation scheme is generated.
It improves the accuracy and relevance of recommendations, reduces computational complexity, and meets the real-time requirements of personalized recommendations.
Smart Images

Figure CN120950733A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a video recommendation method and apparatus based on dynamic binary network representation. Background Technology
[0002] Existing network representation learning methods, such as graph-based representations, visualization methods, matrix factorization and eigenvector representations, and random walk and deep learning representations, while performing well in various network analysis tasks, are mostly used to represent static heterogeneous networks. Because the structure of dynamic heterogeneous networks changes over time, these existing static heterogeneous network representation learning methods are no longer applicable. One special type is the dynamic bipartite network, which contains only two types of nodes, with edges existing only between the different types of nodes. Due to its unique properties, many relationships in the real world can be abstracted as dynamic bipartite networks, such as user-product and author-paper relationships. By capturing the features and behaviors in dynamic bipartite networks, we can discover the network's latent characteristics and reduce the complexity of later computations. Therefore, research on representation learning for dynamic bipartite networks has significant meaning and application value.
[0003] While representation learning methods for dynamic heterogeneous networks, based on continuous time and continuous snapshots, can represent most nodes in these networks, their performance is less than ideal due to the unique characteristics of dynamic bipartite networks. Both of these methods focus solely on the relationships between nodes, neglecting the importance of the nodes themselves. This may overlook crucial network information, resulting in incomplete network structure data and ultimately inaccurate application results. Furthermore, the accuracy of personalized recommendations based on dynamic bipartite networks needs further improvement. Summary of the Invention
[0004] This invention provides a video recommendation method and apparatus based on dynamic bipartite network representation to address the problem that the accuracy of the network representation results of dynamic bipartite networks is low, and the accuracy of recommendation results needs to be further improved when used for personalized recommendations.
[0005] In a first aspect, embodiments of the present invention provide a video recommendation method based on dynamic binary network representation, including:
[0006] Within a set time period, capture multiple binary network snapshots at set time intervals, and fill virtual nodes into other binary network snapshots based on the largest binary network snapshot.
[0007] Obtain the processed binary network snapshot and the adjacency matrix composed of user nodes and project nodes, and capture the explicit features of nodes in each binary network snapshot; wherein, the user node is a user account in the social media platform; and the project node is a video in the social media platform.
[0008] Obtain the processed snapshots of each binary network and the explicit features of the corresponding nodes, and capture the implicit features of the nodes in each binary network snapshot.
[0009] The explicit and implicit features of a node are represented by a low-dimensional vector, and the low-dimensional vector is stored in a cache.
[0010] Based on node-based low-dimensional vector representations, a video recommendation scheme is generated for each user.
[0011] In one possible implementation, populating virtual nodes in other binary network snapshots based on the largest binary network snapshot includes:
[0012] Select the snapshot of the binary network with the largest number of nodes as the baseline snapshot;
[0013] Based on the nodes in the benchmark snapshot, the nodes in other binary network snapshots are compared and verified;
[0014] When a node is missing in a binary network snapshot, a virtual node is added to fill the gap, so that the node size of each binary network snapshot is consistent with the node size of the baseline snapshot.
[0015] In one possible implementation, the virtual node and the real node are in an initial or newly added state during network changes;
[0016] When a new node appears, an initial virtual node is converted into a real node in the newly added state;
[0017] When a real node disappears, it is converted into a virtual node in its initial state.
[0018] In one possible implementation, capturing explicit features of nodes in each binary network snapshot includes:
[0019] One-hot encoding is performed on the user nodes and project nodes in each binary network snapshot;
[0020] Concatenate the one-hot encoded vectors of user nodes and project nodes into a complete matrix;
[0021] An adjacency matrix is constructed based on the concatenated matrix, and a meta-path is obtained from the adjacency matrix. If the adjacency matrix is zero, it indicates that there is no connection between the user node and the project node; otherwise, it indicates a connection between the user node and the project node. The meta-path is a path composed of multiple nodes that are connected to the user node and the project node.
[0022] Information about different neighboring nodes is obtained by using metapaths, features of each neighboring node are aggregated, and explicit features of the node are obtained by combining the node's own features.
[0023] In one possible implementation, the meta-path includes user-project, user-project-user, and project-user-project.
[0024] In one possible implementation, the feature aggregation of each neighbor node includes:
[0025] For each node, the average value of the features of the neighboring nodes under each metapath is calculated based on the metapath set. Combined with the node's own initial features and weight matrix, the explicit features of the node are obtained through activation function processing.
[0026] In one possible implementation, capturing the implicit features of nodes in each binary network snapshot includes:
[0027] Each binary network snapshot is processed to obtain a weighted homogeneous graph of user nodes and project nodes;
[0028] The explicit features of the weighted homogeneous graph and nodes are input into the graph attention mechanism.
[0029] The attention coefficients between nodes are calculated, and the explicit features of neighboring nodes are weighted and summed based on the attention coefficients. The implicit features of the nodes are then processed by an activation function.
[0030] In one possible implementation, the weighted homogeneous graph is obtained by exponentiation of the adjacency matrix of a bipartite network snapshot, where the exponentiation of the adjacency matrix represents the number of paths of corresponding length connecting nodes.
[0031] In one possible implementation, representing the explicit and implicit features of a node using a low-dimensional vector includes:
[0032] Position embedding is performed on nodes in a binary network snapshot;
[0033] Concatenate the location embedding results with the implicit node features of the project nodes;
[0034] Map the cascaded results to the query, key, and value space, and calculate the attention weights;
[0035] The values are weighted and summed based on attention weights to obtain a low-dimensional vector representation of the node.
[0036] In one possible implementation, storing the low-dimensional vector portion in the cache includes:
[0037] In each snapshot, the feature vectors of the monitoring node and all its neighboring nodes;
[0038] Calculate the similarity between the feature vectors of a node and its neighboring nodes;
[0039] If the similarity of all neighboring nodes is greater than or equal to the preset threshold within the time window, the node is marked as stable, its representation is not updated, and the information in the cache is reused; if it is unstable, representation learning is performed again and the cache is updated.
[0040] In one possible implementation, the similarity is measured using cosine similarity.
[0041] In one possible implementation, the node-based low-dimensional vector representation generates a video recommendation scheme for the user, including:
[0042] Calculate user preferences for items based on low-dimensional vector representations of user nodes and item nodes;
[0043] Sort the items according to preference level, and select the items with the highest preference to generate personalized recommendation solutions.
[0044] Secondly, embodiments of the present invention provide a video recommendation device based on dynamic binary network representation, characterized in that it includes: a dynamic binary network representation learning model and a recommendation module; the dynamic binary network representation learning model includes: a virtual node filling module, an EPCT module, an IMPT module, a Temporal Self-attention module, and a Nodecache module;
[0045] The virtual node filling module is used to capture multiple binary network snapshots at set time intervals within a set time period, and fill virtual nodes in other binary network snapshots based on the largest binary network snapshot.
[0046] The EPCT module is used to acquire processed binary network snapshots and adjacency matrices composed of user nodes and project nodes, and to capture explicit features of nodes in each binary network snapshot; wherein, the user nodes are user accounts in the social media platform; and the project nodes are videos in the social media platform.
[0047] The IMPT module is used to obtain the processed snapshots of each binary network and the explicit features of the corresponding nodes, and to capture the implicit features of the nodes in each binary network snapshot.
[0048] The Temporal Self-attention module is used to represent the explicit and implicit features of a node using a low-dimensional vector.
[0049] The Node cache module is used to store the low-dimensional vector portion into the cache area;
[0050] The recommendation module is used to generate video recommendation schemes for users based on the low-dimensional vector representation of nodes.
[0051] In this embodiment of the invention, by taking snapshots of the binary network at set time intervals and filling virtual nodes based on the largest snapshot, the node size of the binary network snapshots at different time points can be guaranteed to be consistent, providing a unified foundation for stable model training. Capturing explicit features of nodes can effectively obtain the direct interaction relationship between users and videos, while capturing implicit features can uncover deep connections between users and videos. Converting features into low-dimensional vector representations reduces data complexity, facilitates efficient computation, and storing some in a cache avoids repeated learning of stable nodes, saving computational resources. Finally, a recommendation scheme is generated based on comprehensive and time-related low-dimensional vectors, which can improve the accuracy and targeting of recommendations and better meet users' personalized needs. Attached Figure Description
[0052] Figure 1 This is a flowchart illustrating the implementation of the video recommendation method based on dynamic binary network representation provided in this embodiment of the invention.
[0053] Figure 2 This is a schematic diagram of the data processing flow of the video recommendation method based on dynamic binary network representation provided in an embodiment of the present invention;
[0054] Figure 3 This is a schematic diagram of the data processing flow of low-dimensional vector representation provided in an embodiment of the present invention;
[0055] Figure 4 This is a schematic diagram of the structure of a video recommendation device based on dynamic binary network representation provided in an embodiment of the present invention. Detailed Implementation
[0056] With the development of information technology, network data, such as social networks, knowledge graphs, and biological networks, exhibits characteristics of massive scale and complex structure. Traditional network analysis methods often struggle to efficiently process such large-scale data. Network representation learning, also known as graph embedding or network embedding, is an efficient method for modeling network structures. It aims to map nodes in a network to a low-dimensional vector space, while the resulting vectors store the neighborhood structure and attribute information of the nodes, greatly reducing the complexity of data processing and storage costs. Furthermore, these vectors possess representational and reasoning capabilities in vector space, making them readily applicable to various network analysis tasks, such as recommendation systems, social networks, financial network analysis, urban risk analysis, node classification, and link prediction. Network representation learning not only improves the accuracy of network analysis but also significantly enhances computational efficiency. Therefore, as a key technology for processing and analyzing the characteristics and behaviors of complex networks, network representation learning holds significant background importance in handling large-scale data and improving the accuracy and efficiency of network analysis. With continuous technological advancements and the expansion of application scenarios, network representation learning will play an even more crucial role in the future. However, current applications still have many shortcomings: First, the network contains a large amount of inaccurate, irrelevant, or missing information, which affects the accuracy of the analysis results; second, existing analysis methods suffer from latency when processing large-scale dynamic data and cannot meet the requirements of real-time monitoring. Therefore, it is crucial that the results of network representation learning, especially dynamic network representation learning, guarantee both accuracy and real-time performance.
[0057] In current network representation learning tasks, based on whether the network topology changes and whether the nodes in the network are of a single or multiple types, network representation learning is divided into two main categories: static heterogeneous network representation learning and dynamic heterogeneous network representation learning. Static heterogeneous networks refer to networks containing multiple types of nodes and edges in their structure. Such networks can more accurately reflect the diversity of objects and relationships in the real world, and representation learning of heterogeneous networks is crucial for utilizing the complex structure and semantic information within the network. Currently, many methods for heterogeneous network representation learning exist, including those based on random walks, matrix factorization, and graph neural networks. Dynamic heterogeneous networks differ from static heterogeneous networks in that the network is constantly changing; that is, the nodes and connecting edges in the network appear or disappear over time. Currently, representation learning methods applied to dynamic networks are generally divided into two categories: one is based on continuous time, which assigns a timestamp to each edge of the network and learns the low-dimensional representation vector of the network nodes using these timestamps; the other is based on network snapshots, which divides the dynamic network into multiple static networks at certain time intervals and then performs representation learning on each snapshot network.
[0058] While representation learning methods for dynamic heterogeneous networks, based on continuous time and continuous snapshots, can represent most nodes in these networks, their performance is less than ideal due to the unique characteristics of dynamic bipartite networks. Furthermore, both methods require repeatedly learning the entire network structure each time, while the network's characteristics and behaviors change only slightly over short periods, resulting in high computational costs and resource requirements. Additionally, these methods focus only on the relationships between nodes, neglecting the importance of the nodes themselves. This may overlook crucial network information, leading to incomplete network structure data and inaccurate application results. In personalized recommendations based on dynamic bipartite networks, both speed and accuracy need further improvement.
[0059] Therefore, considering the maximum size of all nodes in the network, designing an efficient dynamic bipartite network representation learning model that minimizes computational resources during the learning process is a worthwhile research problem. This application uses video recommendation as an example for illustration. In specific implementation, the method provided in this application is applicable to recommendation applications in other fields to improve the speed and accuracy of personalized recommendations.
[0060] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0061] Figure 1 This is a flowchart illustrating the video recommendation method based on dynamic bipartite network representation provided in an embodiment of the present invention. Figure 1 As shown, it includes the following steps:
[0062] S101, capture multiple binary network snapshots at set time intervals within a set time period, and fill virtual nodes in other binary network snapshots based on the largest binary network snapshot;
[0063] The execution subject of each embodiment of this application can be a server, processor, microprocessor, or other device with data processing capabilities. In actual implementation, the specific implementation method of the execution subject can be selected according to actual needs. This embodiment does not impose any special restrictions on this, as long as it is a device with data processing capabilities.
[0064] In the overall structure of a dynamic binary search network, both edges and nodes exhibit dynamic characteristics, leading to potential changes in the network's node size at different points in time. When applied to video recommendation, users and videos on social media platforms are considered two different types of nodes, with user-video interactions represented by an edge. These interactions include actions such as liking, saving, sharing, recommending, and disliking. If user u...p I like the video v q Then node u p and v q There are connecting edges between them, and the edge type is labeled "like". Similarly, if user u p I've saved the video v q Then add an edge of type "favorites". For each discrete time point, repeat the above edge construction process to obtain time t. s Binary network snapshot G s All time points G s Arranging them in chronological order can form a dynamic bipartite network.
[0065] To ensure consistency in model training, these network snapshots must be preprocessed. First, the binary network snapshots at all time points are aligned. That is, the snapshot with the largest node size is selected as the baseline from a series of binary network snapshots. Based on the number of nodes and the structural characteristics of the nodes in this network snapshot, nodes missing in other snapshots at the same time are filled by introducing virtual nodes to align the node size in network snapshots at different time points.
[0066] like Figure 2 As shown, multiple binary network snapshots include t1 to t2. i The largest binary network snapshot is t1, which includes four user nodes u1 to u4 and three video nodes v1 to v3. Virtual nodes are then populated in other binary network snapshots based on the largest binary network snapshot t1.
[0067] S102, obtain the processed binary network snapshot and the adjacency matrix composed of user nodes and project nodes, and capture the explicit features of nodes in each binary network snapshot; where user nodes are user accounts in the social media platform; project nodes are videos in the social media platform.
[0068] A binary network consists of two different types of nodes. Although the connecting edges only exist between nodes of different types, there are also paths between nodes of the same type. These two different connection relationships are explicit and implicit. Explicit relationships only exist between nodes of different types. For example, if there is a direct edge connecting two nodes, then these two nodes have an explicit relationship. The larger the weight, the stronger the relationship. Implicit relationships only exist between nodes of the same type. For two nodes of the same type, if there is a path between them, then there should be some implicit relationship between them. Furthermore, the number and length of the paths indicate the strength of the implicit relationship.
[0069] Specifically, when capturing the explicit features of nodes in each binary network snapshot, explicit features are obtained separately for user nodes and project nodes. and
[0070] S103, obtain the processed snapshots of each binary network and the explicit features of the corresponding nodes, and capture the implicit features of the nodes in each binary network snapshot.
[0071] S104 represents the explicit and implicit features of a node using a low-dimensional vector and stores the low-dimensional vector portion in a cache.
[0072] Representing the explicit and implicit features of nodes using low-dimensional vectors aims to capture the temporal changes in relationships and interactions between nodes.
[0073] In one possible implementation, storing the low-dimensional vector portion in a cache includes:
[0074] In each snapshot, the feature vectors of the monitoring node and all its neighboring nodes;
[0075] Calculate the similarity between the feature vectors of a node and its neighboring nodes;
[0076] If the similarity of all neighboring nodes is greater than or equal to the preset threshold within the time window, the node is marked as stable, its representation is not updated, and the information in the cache is reused; if it is unstable, representation learning is performed again and the cache is updated.
[0077] In one possible implementation, similarity is measured using cosine similarity.
[0078] After observing that the states of network nodes are relatively stable over a specific time period, continuously repeating the learning process on static nodes is clearly an inefficient use of computational resources. Therefore, a smart caching strategy needs to be incorporated into the model. This strategy aims to store and reuse the state and representation information of learned nodes, thereby avoiding unnecessary redundant computation. Using cosine similarity to calculate the cosine value of the angle between two feature vectors can effectively capture the directional similarity between feature vectors. Especially in high-dimensional feature spaces, it has strong robustness and can better determine whether a node is in a stable state.
[0079] Define node v q and its neighbor node u p The stability conditions are as follows:
[0080] If the similarity S(v) q ,u p If )≥θ, then node v q It is considered stable, where θ is a pre-set similarity threshold. Under this condition, node v q The representation will not be updated, thus avoiding repeated learning.
[0081] In each time snapshot, node v is continuously monitored. q The feature vectors of node v and all its neighboring nodes are used. After each feature update, the similarity between nodes is calculated. If the similarity of all neighboring nodes remains above a threshold within the time window, node v is considered successful. q If a node is marked as stable, the learning process will be skipped. If it is unstable, the representation learning will be restarted and the node's representation updated.
[0082] In other words, when t s After the network node representation learning task is completed, the current state of these nodes and their corresponding representation data are stored in a cache using a hash table structure for fast access and reuse later. As time progresses to the next time step, the network snapshot maintains its size consistency by adding virtual nodes. In this scenario, the model compares the current network snapshot with the node states in the cache. If it determines that the node or its surrounding environment has not changed significantly, it skips relearning that node; otherwise, if changes are detected, it initiates a new round of learning for that node and updates its state and representation information in the cache.
[0083] S105 generates a video recommendation scheme for users based on node-based low-dimensional vector representation.
[0084] In one possible implementation, a video recommendation scheme for a user is generated based on the low-dimensional vector representation of nodes, including:
[0085] Calculate user preferences for items based on low-dimensional vector representations of user nodes and item nodes;
[0086] Sort the items according to preference level, and select the items with the highest preference to generate personalized recommendation solutions.
[0087] In one possible implementation, the above method flow is as follows:
[0088] First, virtual nodes are filled into each snapshot to ensure consistency in model training. User nodes and video nodes are converted into vectors using one-hot encoding and concatenated into an adjacency matrix. After processing the adjacency matrix, the explicit relationships between user nodes and video nodes are integrated into the initial explicit feature representation vector. Explicit relationships represent direct interactions between users and videos, i.e., edges. Based on the bipartite network snapshots and the initial explicit feature representation vector, implicit relationships between user nodes and video nodes are obtained and integrated into the feature vectors. Implicit relationships represent deep associations. For example, users u1 and u2 have never interacted, but the types of videos u1 frequently watches highly overlap with u2's, suggesting that u1 may also be interested in videos that u2 has liked. Many users watch "camping" videos followed by "astrophotography" videos, but there is no direct tag association between the two, suggesting the existence of an implicit interest chain. Based on the discrete time points of the dynamic bipartite network, temporal information is obtained to obtain the representation vectors of user nodes and video nodes respectively. and Implement intelligent caching to save computing resources.
[0089] Based on the representation vectors of all user nodes and video nodes learned in previous moments, for a given user u p Using cosine similarity to measure user u at the next time step p For video v q Preferences:
[0090]
[0091] in, User u p The representation vector is Video v q The representation vector; finally, calculate the user u p The system analyzes and sorts all video preferences, then selects the k videos with the highest preferences for personalized recommendations.
[0092] In this embodiment, by taking snapshots of the binary network at set time intervals and filling virtual nodes based on the largest snapshot, the node size of the binary network snapshots at different time points can be guaranteed to be consistent, providing a unified foundation for stable model training. Capturing explicit features of nodes can effectively obtain the direct interaction relationship between users and videos, while capturing implicit features can uncover deep connections between users and videos. Converting features into low-dimensional vector representations reduces data complexity, facilitates efficient computation, and storing some in a cache avoids repeated learning of stable nodes, saving computational resources. Finally, a recommendation scheme is generated based on comprehensive and time-related low-dimensional vectors, which improves the accuracy and relevance of recommendations and better meets users' personalized needs.
[0093] In one possible implementation, virtual nodes are populated in other binary network snapshots based on the largest binary network snapshot, including:
[0094] Select the snapshot of the binary network with the largest number of nodes as the baseline snapshot;
[0095] Based on the nodes in the baseline snapshot, the nodes in other binary network snapshots are compared and verified;
[0096] When a node is missing in a binary network snapshot, a virtual node is added to fill the gap, so that the node size of each binary network snapshot is consistent with the node size of the baseline snapshot.
[0097] In practice, the process of adding virtual nodes is described as follows:
[0098] (1) Obtain the maximum number of user nodes M in a snapshot of a binary network within a certain time period. U And the maximum node size M of the project V .
[0099] (2) For the binary network snapshot G i ,like or Add to and The user nodes and project nodes in the image are called virtual nodes. The initial feature values of these virtual nodes are initialized to 0, and their state is initialized to "initial". The "initial" state of the node is: state = "initial". This indicates that the node is only used to populate the binary network snapshot to unify the node size; it does not actually exist, meaning that these virtual nodes are not processed during subsequent model training. As for the U in the snapshot... i and V i The node, called the real node, is initially in the state of "used".
[0100] The above operations are used to split the network snapshot G. i This will transform into an aligned bipartite network snapshot G′ with virtual nodes. i After completing this alignment operation, take the aligned binary network snapshots G′={G′1,G′2,...,G′} at each time point. nThis serves as the foundation for feature extraction. Aligning all binary network snapshots ensures the proper transmission of hidden states and parameters, thereby enabling efficient model training. After alignment, the binary network snapshots at each time step are unified into a structure with the same virtual nodes, mapping the previously different number of nodes and topology at different time points to the same "coordinate system." In this way, the hidden state corresponding to each virtual node in the mask matrix can be directly transmitted between adjacent time points without interruption due to missing nodes, ensuring the continuity of state variables and the effectiveness of parameter sharing, thus achieving smooth evolution of hidden states and efficient training.
[0101] Furthermore, the virtual nodes filled during alignment are isolated nodes that are not connected to other nodes in the network snapshot and do not exist. To improve the model's computational power and ensure the accuracy of the learning results, the model does not need to learn these virtual nodes during the bipartite network representation learning process. Therefore, a mask matrix is introduced to mask these introduced virtual nodes. The mask matrix creates a mask vector for each bipartite network snapshot with added virtual nodes, marking the real node positions as 1 and the virtual node positions as 0.
[0102] In this embodiment, by selecting the snapshot with the largest node size as the benchmark and comparing and verifying other snapshots and filling virtual nodes, the node size of the binary network snapshots at different time points can be accurately aligned. This allows snapshots with different numbers of nodes and topological structures to be mapped to the same "coordinate system," ensuring the effective transmission of hidden states and parameters during model training and avoiding state breaks caused by missing nodes. This provides a unified structural foundation for the accurate capture of subsequent explicit and implicit features, ensuring the continuity and efficiency of model training.
[0103] In one possible implementation, virtual nodes and real nodes are in an initial or newly added state during network changes;
[0104] When a new node appears, an initial virtual node is converted into a real node in the newly added state;
[0105] When a real node disappears, it is converted into a virtual node in its initial state.
[0106] During implementation, virtual nodes and real nodes may be in one of two states: initial or newly added, throughout the entire network change process.
[0107] When a new node appears in the network, a virtual node with the state "initial" is selected as the new node. Attributes and edge relationships are added to this virtual node, and the virtual node is changed from the "initial state" to the "new" state: state = "used". This operation transforms the virtual node into a real node.
[0108] If a real node in a network disappears for some reason, it will be transformed into a virtual node, and its state will change from "used" to "initial".
[0109] In this embodiment, the dynamic state transition mechanism between virtual and real nodes can flexibly handle the addition and disappearance of nodes in the network: when a new node appears, it is quickly incorporated into the network through virtual node state transition; when a real node disappears, it is converted to a virtual node to maintain the network's size. This ensures both the consistency of the node size and accurately reflects the dynamic changes in the network. Simultaneously, virtual nodes in their initial state do not participate in learning, avoiding invalid computation and allowing the model to focus on feature learning of real nodes, thus improving computational efficiency while maintaining dynamism.
[0110] In one possible implementation, explicit characteristics of nodes in each binary network snapshot are captured, including:
[0111] One-hot encoding is performed on the user nodes and project nodes in each binary network snapshot;
[0112] Concatenate the one-hot encoded vectors of user nodes and project nodes into a complete matrix;
[0113] An adjacency matrix is constructed based on the concatenated matrix, and meta-paths are obtained from the adjacency matrix. If the adjacency matrix is zero, it means that there is no connection between the user node and the project node; otherwise, it means that there is a connection between the user node and the project node. The meta-path is a path composed of multiple nodes that are connected to the user node and the project node.
[0114] Information about different neighboring nodes is obtained by using metapaths, features of each neighboring node are aggregated, and explicit features of the node are obtained by combining the node's own features.
[0115] Explicit interaction relationships between different types of nodes, such as users' direct behaviors like liking or saving videos.
[0116] In this embodiment, by performing one-hot encoding on users and video nodes and constructing an adjacency matrix, the connection relationships between nodes can be clearly and accurately represented. By using metapaths to obtain information about different neighboring nodes and performing feature aggregation, the explicit interaction relationships between different types of nodes can be effectively captured. Combined with the node's own features, the explicit features not only include individual node information but also integrate neighborhood associations, thereby more comprehensively reflecting the explicit attributes and connection patterns of nodes in the network, providing reliable basic features for subsequent implicit feature learning and recommendation scheme generation.
[0117] In one possible implementation, the metapath includes user-project, user-project-user, and project-user-project.
[0118] In this embodiment, meta-paths include user-item, user-item-user, and item-user-item. These diverse meta-paths can cover the direct and indirect explicit relationships between different types of nodes: user-item reflects direct interaction, user-item-user reflects the association formed by users through shared videos, and item-user-item reflects the association formed by videos through shared users. The neighborhood information obtained through these meta-paths is richer, enabling the aggregated explicit features to contain multi-dimensional semantic relationships, thus improving the comprehensiveness and semantic expressiveness of the explicit features.
[0119] In one possible implementation, feature aggregation is performed on each neighboring node, including:
[0120] For each node, the average value of the features of the neighboring nodes under each metapath is calculated based on the metapath set. Combined with the node's own initial features and weight matrix, the explicit features of the node are obtained through activation function processing.
[0121] In this embodiment, the average value of the features of neighboring nodes is calculated for each node based on the metapath set, which can balance and integrate the information of different neighborhoods and avoid the excessive influence of a single neighbor. By combining the node's own initial features and weight matrix and processing them with the activation function, important features can be highlighted by adjusting the weights while preserving the node's own characteristics. The activation function increases the non-linear expressive power of the features, making the obtained explicit features more accurately reflect the explicit relationships and attributes of the nodes, thereby improving the discriminativeness and effectiveness of the features.
[0122] The above describes how to obtain the explicit features of nodes in each binary network snapshot. First, create a binary network snapshot G′ = {G′1, G′2, ..., G′} with filled virtual nodes. n Each snapshot of the binary network in} is represented as an adjacency matrix. The adjacency matrix is represented by a one-hot encoded vector of the real nodes.
[0123] One-hot encoding is a method of representing nodes as binary vectors. Assume a binary network snapshot G′. i There are i user nodes, each user node u s one-hot encoded vector It is an i+j dimensional vector where only the s-th element is 1, and the rest are 0. For example, the one-hot encoding of user node u2 is: Each project node v t The one-hot encoded vector is an i+j dimensional vector, and its encoding method is similar to that described above, with the i+t-th element being 1 and the remaining elements being 0. For example, the encoding of project node v3 is: Next, the one-hot encoded vectors of the user node and the project node are concatenated into a complete matrix X. i : The first i elements represent the one-hot encoding of the user node, and the last j elements represent the one-hot encoding of the item node. Finally, G′ i Represented as adjacency matrix A i :
[0124]
[0125] Where 0 is a zero matrix, representing no connection between user nodes or project nodes, X i It is an i×j matrix representing the connections between user nodes and project nodes. It is X i The transpose of represents the connection between project nodes and user nodes.
[0126] Metapaths are paths defined on heterogeneous network patterns, used to connect different node types. They can be viewed as higher-order neighbors with specific semantics, with different metapaths representing different semantic relationships. Therefore, metapaths can be used to effectively obtain information about surrounding nodes. For example, selecting a metapath (uvu) allows information to be passed between the current user and its associated second-order nodes.
[0127] After the above processing, the adjacency matrix A iThe adjacency matrix stores the topological structure information between nodes, while the initial representation of a node provides its characteristic information. Therefore, meta-paths are determined based on the adjacency matrix and the initial representation of a node. Since meta-paths are paths defined on heterogeneous network patterns to connect different node types, they can effectively acquire information about surrounding nodes. Therefore, we first obtain information about different neighboring nodes using different meta-paths, then aggregate the features of these nodes, and finally combine them with the node's own features to obtain the node's representation. In the specific implementation, the same operation is performed for user nodes and project nodes, so the following section only details the calculation steps for project nodes, as shown in the formula:
[0128]
[0129] Where R is the set of metapaths, R = {uv, uvu, vuv} (i.e., user-project, user-project-user, project-user-project). H represents the number of neighboring nodes. U It is the feature matrix of all user nodes. It is node v q The initial feature vector, It is a weight matrix used to weight user node v q The initial features are linearly transformed, and σ is the activation function.
[0130] After the above processing, each binary network snapshot G′ s The initial explicit characteristics of user nodes and project nodes in the code are represented as follows: Where d′ is the dimension.
[0131] In one possible implementation, the implicit features of nodes in each binary network snapshot are captured, including:
[0132] Each binary network snapshot is processed to obtain a weighted homogeneous graph of user nodes and project nodes;
[0133] The explicit features of the weighted homogeneous graph and nodes are input into the graph attention mechanism.
[0134] The attention coefficients between nodes are calculated, and the explicit features of neighboring nodes are weighted and summed based on the attention coefficients. The implicit features of the nodes are then processed by an activation function.
[0135] In this embodiment, by processing snapshots to obtain a weighted homogeneous graph of user and item nodes, the indirect associations between nodes of the same type can be quantified. After inputting the graph attention mechanism, the attention coefficients between nodes can be calculated to highlight the influence of important neighbors. Based on the coefficients, the explicit features are weighted and summed and processed by the activation function, which can effectively capture the deep implicit relationships between nodes of the same type (such as the implicit associations between users with similar interests and related videos). This allows the implicit features to supplement the deep semantics not covered by the explicit features, enriching the dimensions of node features and providing more comprehensive feature support for subsequent accurate recommendations.
[0136] In one possible implementation, the weighted homogeneous graph is obtained by exponentiation of the adjacency matrix of a bipartite network snapshot, where the exponentiation of the adjacency matrix represents the number of paths of corresponding length connecting nodes.
[0137] In this embodiment, a weighted homogeneous graph is obtained by power calculation of the adjacency matrix, where the nth power element value of the adjacency matrix represents the number of paths of length n between nodes, which can quantitatively reflect the strength of indirect association between nodes of the same type (the more paths, the stronger the association). The weighted homogeneous graph constructed based on this provides a quantitative basis for implicit feature capture, enabling the graph attention mechanism to calculate the attention coefficient more accurately, thereby enabling the captured implicit features to accurately reflect the strength difference of deep association between nodes, improving the accuracy and targeting of implicit features.
[0138] Before obtaining the implicit features of nodes in each binary network snapshot, we obtain each binary network snapshot G′. s and the explicit characteristics of nodes and First, regarding network snapshot G′ s Processing is performed to obtain U from the network snapshot. s V s Weighted homogeneous graph of two types of nodes and Then, the processed weighted homogeneous graph and explicit feature vectors are calculated to obtain the implicit feature information of user nodes and item nodes, respectively. Where d ” This refers to the embedding dimension of the output. Since the binary network snapshots are generated from the same dynamic binary network, the same parameters are used when processing different binary network snapshots. The specific process is as follows:
[0139] Due to the unique structure of a binary network, uvu and vuv are the only meta-paths for user nodes and project nodes, respectively, in the binary network. Below, we will use t... s Binary network snapshot at a given moment For example, according to the principle of matrix multiplication, G s Adjacency matrix A sThe nth power is expressed as Among them, element value G represents s Connecting user nodes u p and project node v q The adjacency matrix is used to calculate the number of simple paths of length n. A larger value indicates a greater number of paths between the corresponding two nodes. Next, the user node set U is obtained by exponentiation of the adjacency matrix. s and project node set V s Weighted homogeneous graph and Then, and The inputs are fed into the graph attention mechanism, which performs the same processing on both types of weighted homogeneous graphs. Therefore, the following text only describes the item node v. q The calculation process, and the input homogeneous graph is collectively referred to as G′, the explicit characteristics of the two types of nodes are: and The calculation process is as follows:
[0140]
[0141] in, Indicates time t s Binary network snapshot G s In, node u p and node v q The attention coefficient between nodes represents the degree of contribution of user nodes to project nodes. and For user nodes and project nodes at time t, respectively. s Explicit eigenvectors, It is project node v q The set of neighboring nodes, where || represents the concatenation operation of feature vectors. W represents the weighted sum of the features of the neighboring nodes of node v. s σ is the weight matrix used for the transformation, and σ is a non-linear activation function.
[0142] In one possible implementation, the explicit and implicit features of a node are represented by a low-dimensional vector, including:
[0143] Position embedding is performed on nodes in a binary network snapshot;
[0144] Concatenate the location embedding results with the implicit node features of the project nodes;
[0145] Map the cascaded results to the query, key, and value space, and calculate the attention weights;
[0146] The values are weighted and summed based on attention weights to obtain a low-dimensional vector representation of the node.
[0147] The acquired user and project node feature information (including explicit feature information) and ) and implicit feature information ( and After processing, the results of user nodes and project nodes are obtained. The results of each type of node include not only static feature information but also time-related dynamic features, which are defined as follows: Where D′ is the dimension of the feature representation vector of the output result. The same parameters are used when processing the feature information of user nodes and item nodes; therefore, for ease of explanation, the following calculation process uses item node v as an example. q For example, a detailed explanation will be provided, with the flowchart as follows: Figure 3 As shown, the calculation process is as follows:
[0148] First, in order to distinguish project node information obtained at different times, t s Binary network snapshot G at time point s All project nodes are embedded in a position, defined as follows: Where V s For t s The set of all project nodes in the binary network snapshot at time step j, where j is the number of project nodes in the snapshot. Then, the position embedding result is... and the output results of project nodes Perform cascading: Finally, the cascaded result is passed to the temporal self-attention part, which operates as follows: The project node v... q The representation is packed into a matrix. Then a linear transformation is used to transform the input node v q The vector represents a query, a key, and a value, which are represented by three different spaces: W. q ∈R d″×D′ W k ∈R d″×D′ W v ∈R d″×D′ Finally, the node representations at the current time step and those at previous time steps are merged. The formula is as follows:
[0149]
[0150] in, This is the attention weight, representing the similarity between the query vector of the q-th item node and the key vector of the j-th item node, given a query W and a key K. M∈R T×T It is a mask matrix to ensure that information from future moments does not involve current calculations. M is defined as follows:
[0151]
[0152] In this embodiment, embedding the location of nodes can distinguish node information at different time points, injecting temporal identifiers into the features; cascading the location embedding with implicit features can integrate the static features and temporal information of the nodes; mapping to the query, key, and value space and calculating attention weights can capture the differences in the importance of node features at different time points; the low-dimensional vector representation obtained by weighted summation of values based on the weights can not only retain the core features of the nodes, but also effectively integrate temporal information, so that the vector not only reflects the node attributes, but also contains the interaction rules of the time dimension, providing effective time-related features for the generation of subsequent recommendation schemes, and improving the timeliness and accuracy of recommendations.
[0153] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0154] The following are device embodiments of the present invention. For details not described in detail, please refer to the corresponding method embodiments described above.
[0155] Figure 4 The diagram illustrates the structure of a video recommendation device based on a dynamic binary network representation provided in an embodiment of the present invention. For ease of explanation, only the parts relevant to the embodiment of the present invention are shown, and are described in detail below:
[0156] like Figure 4 As shown, the video recommendation device 4 based on dynamic binary network representation includes: a dynamic binary network representation learning model 401 and a recommendation module 402. The dynamic binary network representation learning model 401 includes: a virtual node filling module 4011, an EPCT module 4012, an IMPT module 4013, a Temporal Self-attention module 4014, and a Node cache module 4015.
[0157] The virtual node filling module 4011 is used to capture multiple binary network snapshots at set time intervals within a set time period, and fill virtual nodes in other binary network snapshots based on the largest binary network snapshot.
[0158] The EPCT module 4012 is used to obtain the processed binary network snapshot and the adjacency matrix composed of user nodes and project nodes, and to capture the explicit features of nodes in each binary network snapshot; where user nodes are user accounts in social media platforms; and project nodes are videos in social media platforms.
[0159] The IMPT module 4013 is used to obtain the explicit features of each processed binary network snapshot and the corresponding nodes, and to capture the implicit features of the nodes in each binary network snapshot.
[0160] The Temporal Self-attention module 4014 is used to represent the explicit and implicit features of a node using a low-dimensional vector.
[0161] The Node cache module 4015 is used to store the low-dimensional vector portion into the cache area.
[0162] Recommendation module 402 is used to generate video recommendation schemes for users based on node-based low-dimensional vector representations.
[0163] In one possible implementation, the virtual node filling module 4011 is specifically used for:
[0164] Select the snapshot of the binary network with the largest number of nodes as the baseline snapshot;
[0165] Based on the nodes in the baseline snapshot, the nodes in other binary network snapshots are compared and verified;
[0166] When a node is missing in a binary network snapshot, a virtual node is added to fill the gap, so that the node size of each binary network snapshot is consistent with the node size of the baseline snapshot.
[0167] In one possible implementation, virtual nodes and real nodes are in an initial or newly added state during network changes;
[0168] When a new node appears, an initial virtual node is converted into a real node in the newly added state;
[0169] When a real node disappears, it is converted into a virtual node in its initial state.
[0170] In one possible implementation, the EPCT module 4012 is specifically used for:
[0171] One-hot encoding is performed on the user nodes and project nodes in each binary network snapshot;
[0172] Concatenate the one-hot encoded vectors of user nodes and project nodes into a complete matrix;
[0173] An adjacency matrix is constructed based on the concatenated matrix, and meta-paths are obtained from the adjacency matrix. If the adjacency matrix is zero, it means that there is no connection between the user node and the project node; otherwise, it means that there is a connection between the user node and the project node. The meta-path is a path composed of multiple nodes that are connected to the user node and the project node.
[0174] Information about different neighboring nodes is obtained by using metapaths, features of each neighboring node are aggregated, and explicit features of the node are obtained by combining the node's own features.
[0175] In one possible implementation, the metapath includes user-project, user-project-user, and project-user-project.
[0176] In one possible implementation, feature aggregation is performed on each neighboring node, including:
[0177] For each node, the average value of the features of the neighboring nodes under each metapath is calculated based on the metapath set. Combined with the node's own initial features and weight matrix, the explicit features of the node are obtained through activation function processing.
[0178] In one possible implementation, the IMPT module 4013 is specifically used for:
[0179] Each binary network snapshot is processed to obtain a weighted homogeneous graph of user nodes and project nodes;
[0180] The explicit features of the weighted homogeneous graph and nodes are input into the graph attention mechanism.
[0181] The attention coefficients between nodes are calculated, and the explicit features of neighboring nodes are weighted and summed based on the attention coefficients. The implicit features of the nodes are then processed by an activation function.
[0182] In one possible implementation, the weighted homogeneous graph is obtained by exponentiation of the adjacency matrix of a bipartite network snapshot, where the exponentiation of the adjacency matrix represents the number of paths of corresponding length connecting nodes.
[0183] In one possible implementation, the Temporal Self-attention module 4014 is specifically used for:
[0184] Position embedding is performed on nodes in a binary network snapshot;
[0185] Concatenate the location embedding results with the implicit node features of the project nodes;
[0186] Map the cascaded results to the query, key, and value space, and calculate the attention weights;
[0187] The values are weighted and summed based on attention weights to obtain a low-dimensional vector representation of the node.
[0188] In one possible implementation, the Node cache module 4015 is specifically used for:
[0189] In each snapshot, the feature vectors of the monitoring node and all its neighboring nodes;
[0190] Calculate the similarity between the feature vectors of a node and its neighboring nodes;
[0191] If the similarity of all neighboring nodes is greater than or equal to the preset threshold within the time window, the node is marked as stable, its representation is not updated, and the information in the cache is reused; if it is unstable, representation learning is performed again and the cache is updated.
[0192] In one possible implementation, similarity is measured using cosine similarity.
[0193] In one possible implementation, module 402 is recommended, specifically for:
[0194] Calculate user preferences for items based on low-dimensional vector representations of user nodes and item nodes;
[0195] Sort the items according to preference level, and select the items with the highest preference to generate personalized recommendation solutions.
[0196] In this embodiment of the invention, by taking snapshots of the binary network at set time intervals and filling virtual nodes based on the largest snapshot, the node size of the binary network snapshots at different time points can be guaranteed to be consistent, providing a unified foundation for stable model training. Capturing explicit features of nodes can effectively obtain the direct interaction relationship between users and videos, while capturing implicit features can uncover deep connections between users and videos. Converting features into low-dimensional vector representations reduces data complexity, facilitates efficient computation, and storing some in a cache avoids repeated learning of stable nodes, saving computational resources. Finally, a recommendation scheme is generated based on comprehensive and time-related low-dimensional vectors, which can improve the accuracy and targeting of recommendations and better meet users' personalized needs.
[0197] This invention also provides a video recommendation device based on a dynamic bipartite network representation, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described in the above method embodiments. Exemplarily, the video recommendation device based on a dynamic bipartite network representation can be a server, a laptop computer, etc., and is not limited thereto.
[0198] In the above embodiments, the descriptions of each embodiment have their own emphasis. Parts not detailed or described in a particular embodiment can be referred to in the relevant descriptions of other embodiments. Unless otherwise specified or in conflict with logic, the terminology and / or descriptions between different embodiments are consistent and can be referenced interchangeably. Technical features in different embodiments can be combined to form new embodiments based on their inherent logical relationships.
[0199] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A video recommendation method based on dynamic binary network representation, characterized in that, include: Within a set time period, capture multiple binary network snapshots at set time intervals, and fill virtual nodes into other binary network snapshots based on the largest binary network snapshot. Obtain the processed binary network snapshot and the adjacency matrix composed of user nodes and project nodes, and capture the explicit features of nodes in each binary network snapshot; wherein, the user node is a user account in the social media platform; and the project node is a video in the social media platform. Obtain the processed snapshots of each binary network and the explicit features of the corresponding nodes, and capture the implicit features of the nodes in each binary network snapshot. The explicit and implicit features of a node are represented by a low-dimensional vector, and the low-dimensional vector is stored in a cache. Based on node-based low-dimensional vector representations, a video recommendation scheme is generated for each user.
2. The video recommendation method based on dynamic binary network representation according to claim 1, characterized in that, The process of filling virtual nodes in other binary network snapshots based on the largest binary network snapshot includes: Select the snapshot of the binary network with the largest number of nodes as the baseline snapshot; Based on the nodes in the benchmark snapshot, the nodes in other binary network snapshots are compared and verified; When a node is missing in a binary network snapshot, a virtual node is added to fill the gap, so that the node size of each binary network snapshot is consistent with the node size of the baseline snapshot.
3. The video recommendation method based on dynamic binary network representation according to claim 2, characterized in that, The virtual nodes and real nodes are either in an initial or newly added state during network changes; When a new node appears, an initial virtual node is converted into a real node in the newly added state; When a real node disappears, it is converted into a virtual node in its initial state.
4. The video recommendation method based on dynamic binary network representation according to claim 1, characterized in that, The capture of explicit features of nodes in each binary network snapshot includes: One-hot encoding is performed on the user nodes and project nodes in each binary network snapshot; Concatenate the one-hot encoded vectors of user nodes and project nodes into a complete matrix; An adjacency matrix is constructed based on the concatenated matrix, and a meta-path is obtained from the adjacency matrix. If the adjacency matrix is zero, it indicates that there is no connection between the user node and the project node; otherwise, it indicates a connection between the user node and the project node. The meta-path is a path composed of multiple nodes that are connected to the user node and the project node. Information about different neighboring nodes is obtained by using metapaths, features of each neighboring node are aggregated, and explicit features of the node are obtained by combining the node's own features.
5. The video recommendation method based on dynamic binary network representation according to claim 4, characterized in that, The metapath includes user-project, user-project-user, and project-user-project.
6. The video recommendation method based on dynamic binary network representation according to claim 4, characterized in that, The feature aggregation of each neighbor node includes: For each node, the average value of the features of the neighboring nodes under each metapath is calculated based on the metapath set. Combined with the node's own initial features and weight matrix, the explicit features of the node are obtained through activation function processing.
7. The video recommendation method based on dynamic binary network representation according to claim 1, characterized in that, The process of capturing the implicit features of nodes in each binary network snapshot includes: Each binary network snapshot is processed to obtain a weighted homogeneous graph of user nodes and project nodes; The explicit features of the weighted homogeneous graph and nodes are input into the graph attention mechanism. The attention coefficients between nodes are calculated, and the explicit features of neighboring nodes are weighted and summed based on the attention coefficients. The implicit features of the nodes are then processed by an activation function.
8. The video recommendation method based on dynamic binary network representation according to claim 7, characterized in that, The weighted homogeneous graph is obtained by exponentiation of the adjacency matrix of the binary network snapshot. The exponentiation of the adjacency matrix represents the number of paths of corresponding length connecting nodes.
9. A video recommendation method based on dynamic binary network representation according to claim 1, characterized in that, The method of representing the explicit and implicit features of nodes using low-dimensional vectors includes: Position embedding is performed on nodes in a binary network snapshot; Concatenate the location embedding results with the implicit node features of the project nodes; Map the cascaded results to the query, key, and value space, and calculate the attention weights; The values are weighted and summed based on attention weights to obtain a low-dimensional vector representation of the node.
10. A video recommendation device based on dynamic binary network representation, characterized in that, include: Dynamic binary network representation learning model and recommendation module; The dynamic binary network representation learning model includes: a virtual node filling module, an EPCT module, an IMPT module, a Temporal Self-attention module, and a Node cache module; The virtual node filling module is used to capture multiple binary network snapshots at set time intervals within a set time period, and fill virtual nodes in other binary network snapshots based on the largest binary network snapshot. The EPCT module is used to acquire processed binary network snapshots and adjacency matrices composed of user nodes and project nodes, and to capture explicit features of nodes in each binary network snapshot; wherein, the user nodes are user accounts on social media platforms; and the project nodes are videos on social media platforms. The IMPT module is used to obtain the processed snapshots of each binary network and the explicit features of the corresponding nodes, and to capture the implicit features of the nodes in each binary network snapshot. The Temporal Self-attention module is used to represent the explicit and implicit features of a node using a low-dimensional vector. The Node cache module is used to store the low-dimensional vector portion into the cache area; The recommendation module is used to generate video recommendation schemes for users based on the low-dimensional vector representation of nodes.