Online learning video resource personalized recommendation method based on graph neural network
By constructing heterogeneous graphs and using multi-layer graph attention networks and multi-layer perception machines, combining user, video and tag information for multi-hop feature aggregation and weight update, the problem of insufficient accuracy of video recommendations in the existing technology is solved, and more efficient personalized recommendations of video resources and the improvement of user actual viewing rates is achieved.
Patent Information
- Application Number
- CN202510254913.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-05
AI Technical Summary
The prior art is difficult to improve the personalized recommendation accuracy of online learning video resources and the actual viewing rate of users when the actual viewing rate of recommended videos is poor.
By constructing heterogeneous graphs, combining the Graph Attention Network (GAT) and the Multi-Layer Perceptron (MLP), users, video and tag information are integrated, multi-hop feature aggregation and weight updates are performed, and videos are accurately recommended. For users with poor actual viewing rates, the sub-graph structure of users with similar but good viewing rates is updated to improve the accuracy of recommended videos.
It effectively improves the accuracy of personalized recommendations of video resources and the actual viewing rate of users, and improves learning experience and efficiency.
Smart Images

Figure CN119760170B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the fields of artificial intelligence and educational technology, and particularly to a personalized recommendation method for online learning video resources based on graph neural networks. Background Art
[0002] With the rapid development of Internet technology, online learning platforms have become an important way for learners to acquire knowledge. However, in the face of a vast amount of learning video resources, how to accurately recommend content that meets the interests and needs of learners is a common problem.
[0003] In the prior art, some patents have explored recommendation methods for short videos and course content, such as the patent "Short Video Recommendation Method and System Based on Heterogeneous Graph Neural Network Integrating Multimodal Data" (CN202411211539.1), and the patent "Course Recommendation Method and System Based on Attention Mechanism Heterogeneous Graph Neural Network" (CN202310588836.7). However, none of these methods mention how to further improve the recommendation accuracy and increase the actual viewing rate when the actual viewing rate of the recommended videos by users is not good. Summary of the Invention
[0004] Embodiments of the present invention provide a personalized recommendation method for online learning video resources based on graph neural networks to solve the above technical problems.
[0005] In a first aspect, embodiments of the present invention provide a personalized recommendation method for online learning video resources based on graph neural networks, including:
[0006] Obtain user information, video resource information, and learning topic tags of an online learning platform, and construct a heterogeneous graph including user nodes, video nodes, and tag nodes according to the obtained content;
[0007] Input the heterogeneous graph into a multi-layer graph attention network, perform multi-hop aggregation on the initial features of each node according to the graph structure, and recommend at least one video to each user according to the final features of each node;
[0008] Divide multiple users into a first group of users with an actual viewing rate greater than a first threshold and a second group of users with an actual viewing rate less than a second threshold according to the actual viewing rate of the recommended videos by the users, where the first threshold is greater than or equal to the second threshold;
[0009] Determine target second users whose change between the final feature and the initial feature is less than a third threshold, and target first users whose initial features are most similar to those of the target second users;
[0010] Update the second sub-graph structure of the target second user in the heterogeneous graph according to the first sub-graph structure of the target first user in the heterogeneous graph, so as to improve the actual viewing rate of the target second user for the recommended videos. The updated heterogeneous graph is used to be re-input into the multi-layer graph attention network and recommend at least one new video to each user respectively.
[0011] In a second aspect, an embodiment of the present invention provides an electronic device, which includes:
[0012] One or more processors;
[0013] A memory for storing one or more programs,
[0014] When the one or more programs are executed by the one or more processors, the one or more processors implement the online learning video resource personalized recommendation method based on a graph neural network according to any embodiment.
[0015] In a third aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the online learning video resource personalized recommendation method based on a graph neural network according to any embodiment.
[0016] In summary, an embodiment of the present invention provides an online learning video resource personalized recommendation method based on a graph neural network. By integrating learner, learning video, and tag information to construct a heterogeneous graph, using a graph attention network to accurately calculate vectors and update weights, deeply extracting features and modeling relationships, it realizes accurate personalized recommendation of video resources, effectively improving the learning experience and efficiency. After the recommendation is completed, record the actual viewing rate of the user for the recommended videos. Select users with a low viewing rate and a small degree of aggregation and transformation of user features by the multi-layer GAT network as the improvement objects, and at the same time use users who are similar to this user but have a better actual viewing rate as the reference objects. Targetedly mine the information data of the improvement objects with reference to their information structure in the isomer graph, capture more information that can accurately reflect the complex rules of users' video viewing, further improve the accuracy of the recommended videos, and increase the actual viewing rate of users. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0018] Figure 1It is a flowchart of a personalized recommendation method for online learning video resources based on a graph neural network provided by an embodiment of the present invention;
[0019] Figure 2 It is a flowchart of another personalized recommendation method for online learning video resources based on a graph neural network provided by an embodiment of the present invention;
[0020] Figure 3 It is a comparison graph of a first sub - graph and a second sub - graph provided by an embodiment of the present invention;
[0021] Figure 4 It is a comparison graph of a simplified first sub - graph and a simplified second sub - graph provided by an embodiment of the present invention;
[0022] Figure 5 It is a comparison graph of a current first sub - graph and a current second sub - graph provided by an embodiment of the present invention;
[0023] Figure 6 It is a comparison graph of the remaining part of a current first sub - graph and a current second sub - graph provided by an embodiment of the present invention;
[0024] Figure 7 It is another comparison graph of the remaining part of a current first sub - graph and a current second sub - graph provided by an embodiment of the present invention;
[0025] Figure 8 It is yet another comparison graph of the remaining part of a current first sub - graph and a current second sub - graph provided by an embodiment of the present invention;
[0026] Figure 9 It is a comparison graph of a simplified first sub - graph and a newly simplified second sub - graph provided by an embodiment of the present invention;
[0027] Figure 10 It is a comparison schematic diagram of a first sub - graph and a new second sub - graph provided by an embodiment of the present invention;
[0028] Figure 11 It is a comparison graph of a new current first sub - graph and a new current second sub - graph provided by an embodiment of the present invention;
[0029] Figure 12 It is a comparison graph of the remaining part of a new current first sub - graph and a new current second sub - graph provided by an embodiment of the present invention;
[0030] Figure 13 It is a comparison graph of the remaining part of a new current first sub - graph and a supplemented newly simplified second sub - graph provided by an embodiment of the present invention;
[0031] Figure 14 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0032] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be described clearly and completely below. Obviously, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope protected by the present invention.
[0033] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and should not be construed as indicating or implying relative importance.
[0034] In the description of the present invention, it should also be noted that unless otherwise clearly defined and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0035] Figure 1 is a flowchart of an online learning video resource personalized recommendation method based on a graph neural network provided by an embodiment of the present invention. This method is executed by an electronic device, such as Figure 1 shown, this method specifically includes:
[0036] S110. Obtain user information, video resource information, and learning topic tags of the online learning platform.
[0037] Among them, the user information includes user age, identity, hobbies, subscribed courses, historical video click behavior, historical viewing behavior, etc. The video resource information includes the title, content, abstract, overview, belonging course system, and hierarchical position in the course of the learning video. The learning topic tags are used to summarize the course content, course knowledge points, etc., and can be extracted from the headings at all levels in the course system.
[0038] This embodiment uses the information data accumulated by the online learning platform as the data source for the entire method. Generally speaking, these information data are relatively large, and some key information can be extracted from them for subsequent operations.
[0039] S120. According to the obtained content, construct a heterogeneous graph including user nodes, video nodes, and tag nodes. Among them, the edge between the user node and the tag node is used to represent the association relationship between the user and the tag, the edge between the video node and the tag node is used to represent the association relationship between the video and the tag, and the edge between the user node and the video node is used to represent the interaction relationship between the user and the video.
[0040] Based on the content obtained in the previous step, this embodiment constructs a heterogeneous graph including user nodes U, video nodes V, and tag nodes T. Among them, each user node corresponds to a user, each video node corresponds to a learning video, and each tag node corresponds to a tag. According to user behaviors such as video clicks and video views, edges can be established between user nodes and video nodes. For example, edges can be established between the videos a user has watched; according to user information, the learning topics preferred by the user can be determined, and edges can be established between the user node and the tag nodes of the preferred learning topics; according to the text similarity between the video resource information and the tag, the learning topic to which the video belongs can be determined, and edges can be established between the user node and the tag nodes of the learning topics.
[0041] Figure 2 An exemplary heterogeneous graph with a tripartite graph structure is shown. Among them, u x , v y and t z respectively represent user nodes, video nodes, and tag nodes. x, y, and z are all node indices. The user-video edge E UV represents the interaction relationship between the user and the video. The user-tag edge E UT represents the association relationship between the user and the tag. The video-tag edge E VT represents the association relationship between the video and the tag. These nodes and edges form a complex heterogeneous graph structure. This heterogeneous graph structure comprehensively captures the rich data information in the online learning platform. Compared with the traditional single-relationship graph, it can better alleviate the data sparsity problem of new users or new videos. Because by introducing tag nodes, even if new users or new videos do not have enough historical interaction data, potential similar users or videos can be found by means of the association with tags, thus providing a basis for subsequent recommendations. At the same time, it perceives and learns the relationships between different types of nodes from multiple dimensions, which helps to deeply mine the potential information of the data and improve the accuracy and diversity of recommendations.
[0042] S130. Input the heterogeneous graph into a multi-layer graph attention network, perform multi-hop aggregation on the initial features of each node according to the graph structure, and recommend at least one video to each user according to the final features of each node.
[0043] In this embodiment, a neural network model based on a multi-layer Graph Attention Network (GAT) and a Multi-Layer Perceptron (MLP) is used to process the above heterogeneous graph. The multi-layer GAT is used to update the node vectors and aggregate the neighbor information to accurately capture the complex relationships between nodes. The MLP is used for the final prediction to further improve the accuracy of video recommendation. The model can be pre-trained based on a large number of heterogeneous graphs. The input of the model is a heterogeneous graph within a period of time (including the initial features and connection relationships of each node in the graph), and the output is the viewing probability of each user for each video in the future period of time. When constructing samples during model training, a heterogeneous graph within a certain historical period can be selected as the input, and the viewing probability can be calculated according to the number of times each user views each video within another historical period after the said historical period. The viewing probability is used as the output to fully train the model.
[0044] In a specific embodiment, the data processing process of the heterogeneous graph in the model may include the following steps:
[0045] Step 1: Generate the initial features of each node according to the obtained content. Optionally, user age, identity, hobbies, subscribed courses, etc. are encoded as the initial features of user nodes; the obtained video information is encoded as the initial features of video nodes, and the label content is encoded as the initial features of label nodes. Among them, the type information can adopt encoding methods such as one-hot, hash, etc., and the text information can adopt the encoding method of text embedding. This embodiment does not make specific limitations. Combined Figure 2 , each node has a corresponding feature vector, and the initial feature encoding methods of the same type of nodes are the same.
[0046] Step 2: Input the initial features and connection relationships of each node into the multi-layer graph attention network. In the multi-layer graph attention network, update the feature vectors of each node according to the feature vectors of adjacent nodes, and update multiple times repeatedly to complete multi-hop aggregation. This step uses the multi-layer GAT to update the node vectors. The core idea of GAT is to dynamically allocate weights to neighbor nodes through the attention mechanism, so as to more accurately capture the complex relationships between nodes. Specifically, in each hop of feature aggregation, for each node i, the node vector is updated according to the following formula:
[0047]
[0048] Among them, is the updated feature vector of node i, N(i) represents the set composed of all neighbor nodes of node i, W is a learnable weight matrix, is the current feature vector of each node in the union of neighbor nodes and node i, is a non - linear activation function; a ij is the attention coefficient or weight assignment matrix, used to measure the importance of different nodes to the target node, and the calculation method is as follows:
[0049]
[0050] wherein, is a learnable attention vector, is the current feature vector of each node in the union of neighbor nodes and node i, || represents the vector concatenation operation, and LeakyReLU( ) is the activation function. The updated node feature vector will be used as the current feature vector of node i in the next - hop aggregation.
[0051] Through multiple - layer GAT, more - hop neighbor information can be gradually aggregated. Assume that the multiple - layer GAT has L layers. After passing through L - layer GAT, the final vector representation of node i is . GAT dynamically assigns weights to each neighbor node through the attention mechanism, enabling the model to more accurately capture the complex relationships between nodes. By gradually aggregating multi - hop neighbor information, the global structural information in the graph can be captured more comprehensively. For online learning video recommendation, this model can better understand the potential interests of users and the comprehensive features of videos, thereby improving the accuracy and coverage of recommendations. In particular, in this embodiment, it is set that the final features output after a node passes through multiple - layer GAT have the same dimension as the initial features input to GAT.
[0052] Step three: Input the final features of each node into a multi - layer perceptron to predict the probability that each user watches each video. In this step, on the basis of the graph neural network, an MLP is introduced for the final prediction. The input of the MLP is the node vector updated through multiple - layer GAT , and the output is the probability that each user watches each prediction . Optionally, the MLP consists of K fully - connected layers, the activation function of each layer is ReLU, and the activation function of the output layer is Sigmoid. Then the prediction score can be expressed as:
[0053]
[0054] where, W k and b k are the weight matrix and bias term of the k - th layer respectively, k = 1, 2, …, K.
[0055] By introducing the MLP, the advantages of the graph neural network in capturing node relationships and the advantages of the MLP in non - linear fitting can be fully utilized to achieve complementary advantages. This combination method not only improves the accuracy of prediction but also enhances the generalization ability of the model.
[0056] Step 4: Recommend at least one video with the highest viewing probability to each user. For each user, recommend the top several videos with the highest predicted viewing probability to that user.
[0057] S140: Divide multiple users into a group of users with an actual viewing rate greater than a certain threshold and another group of users with an actual viewing rate less than the other threshold according to the actual viewing rate of the recommended videos by the users, where the certain threshold is greater than or equal to the other threshold.
[0058] In this embodiment, after recommending a specific video to a user, record the actual viewing rate of each user for the recommended video in a future period of time. Among them, the viewing rate of a certain user for the recommended video = the number of recommended videos actually viewed by the user / the total number of videos recommended to the user by the model. This viewing rate reflects the accuracy of video recommendation. When the test accuracy of the model is relatively stable, this viewing rate also characterizes whether the heterogeneous graph extracts enough effective information from a large amount of information data to learn the complex rules of user video viewing.
[0059] Further, according to the different actual viewing rates, this embodiment divides users into two groups, one group with a relatively high actual viewing rate for the recommended video and the other group with a relatively low actual viewing rate for the recommended video. For the convenience of distinction and description, this embodiment refers to the certain threshold as the first threshold, the other threshold as the second threshold, each user with an actual viewing rate greater than the first threshold as the first user, and each user with an actual viewing rate less than the second threshold as the second user. Optionally, the second threshold can take a value less than the value of, where represents the total number of videos recommended to each user, to screen out the second users with an actual viewing rate of 0. The of each user can be the same or different; at the same time, the first threshold takes a value greater than the second threshold.
[0060] S150: Determine the target second user whose change between the final feature and the initial feature is less than another threshold, and the target first user most similar to the aggregated pre-feature of the target second user.
[0061] The final feature here refers to the aggregated feature output by the last layer in the multi-layer GAT network , and the initial feature refers to the initial feature input into the multi-layer GAT network. In this embodiment, for the same user node, the same dimension is set for the final feature and the initial feature, and the difference between the final feature and the initial feature of this node is used to reflect the degree of mining of the global relationship of this node by the multi-layer GAT network, and at the same time, this difference is also used to reflect whether the heterogeneous graph provides enough effective information for the mining of the relationship of this node.
[0062] In a specific embodiment, first, for each second user, calculate the similarity between the initial feature and the final feature of the current second user respectively. , and use the second user whose value is less than another threshold as the target second user. For the convenience of distinction and description, the said another threshold is called the third threshold. If the change between the initial feature and the final feature of the same target second user is not significant, it indicates that the multi-layer GAT network does not aggregate enough effective information for the initial feature of the user. When the test accuracy of the model is relatively stable, this may be due to the fact that the heterogeneous graph does not contain much effective information about this user (for example, the number of edges or nodes related to this user in the heterogeneous graph is too small, or the number of effective edges or nodes is too small), and the multi-layer GAT network does not aggregate enough information from neighbor nodes, resulting in insufficient actual viewing rate of the recommended videos by the user.
[0063] Then, calculate the similarity between the initial feature of each first user and the initial feature of the target second user respectively, and select the first user with the maximum similarity as the target first user. The fact that the similarity between this user's own initial feature and the initial feature of the target second user is the highest indicates that the two users have the most similar information. Then, this target first user can be used as a reference node for expanding the information related to the target second user in the heterogeneous graph subsequently.
[0064] S160. Update the subgraph structure of the target second user in the heterogeneous graph according to the subgraph structure of the target first user in the heterogeneous graph, so as to improve the actual viewing rate of the target second user for the recommended videos.
[0065] This step expands the interaction information and association information of the target second user in the heterogeneous graph with reference to the interaction information and association information of the target first user in the isomer graph. For the convenience of distinction and description, in this embodiment, the subgraph corresponding to the target first user is called the first subgraph, and the subgraph corresponding to the target second user is called the second subgraph.
[0066] In a specific embodiment, this process may include the following steps:
[0067] Step 1. Respectively extract the first subgraph that spreads outward centered on the target first user and the second subgraph that spreads outward centered on the target second user from the heterogeneous graph. Figure 3An exemplary extracted sub - graph is shown, where U1 represents the target first user, V11, V12, and V13 respectively represent different video nodes in the first sub - graph, U12 represents the user node in the first sub - graph, and T11, T12, and T13 respectively represent different label nodes in the first sub - graph; U2 represents the target second user, V21 represents the video node in the second sub - graph, U22 represents the user node in the second sub - graph, and T21 and T23 respectively represent different label nodes in the second sub - graph.
[0068] Step 2: Simplify the node attributes in each sub - graph into node types to obtain the simplified first sub - graph and the simplified second sub - graph, where the node types include user, video, and label. Through node simplification, nodes of the same type are regarded as the same node. For example, multiple user nodes are all the same U node without distinguishing which user it is. Figure 3 For example, the simplified first sub - graph and the second sub - graph are as Figure 4 shown, where U, V, and T respectively represent user nodes, video nodes, and label nodes. The different colors in the nodes are for convenience in subsequent operations and do not represent differences in the nodes.
[0069] Step 3: Initialize N to 1 and loop through the following operations:
[0070] S1 - 10: Extract the central node and its first N - layer neighbor nodes closest to it, as well as the edges between the extracted nodes from the simplified first sub - graph and the simplified second sub - graph respectively, to form the current first sub - graph and the current second sub - graph. Taking N = 1 as an example, the current first sub - graph and the current second sub - graph extracted from Figure 4 are as Figure 5 shown.
[0071] S1 - 20: Remove the parts in the current first sub - graph that are the same as those in the current second sub - graph. Optionally, first, remove the central node, the first N - 1 layer neighbor nodes, and the first N - 1 layer edges from the current first sub - graph. When N = 1, only the two central nodes are the same nodes, and only the central node needs to be removed. The remaining part is as Figure 6 shown.
[0072] Then, perform the following operations on each neighbor node in the Nth layer of the current second sub - graph in sequence:
[0073] If there is only one node in the remaining part of the current first sub - graph (excluding the parts that have been removed for the previous neighbor nodes) that is the same as the current neighbor node, remove the same node and the edges of the same type as the edges connecting the current neighbor node from the remaining part, where the types of edges include the edges between users and labels, the edges between videos and labels, and the edges between users and videos.Figure 6 For example, the neighbor nodes of the first layer of the current second sub-graph include a V node and a T node. For the V node among them (i.e., the blue V node), there is only one V node in the remaining part of the current first sub-graph. Therefore, the white V node and the user-video edges (i.e., blue edges) of the same type as the edges connecting the white V node and the blue V node (user-video edges) are removed from the remaining part of the current first sub-graph. The two sub-graphs after removal are as Figure 7 shown.
[0074] If there are multiple nodes in the remaining part of the current first sub-graph (excluding the part that has been removed for the previous neighbor nodes) that are the same as the current neighbor node, select the node with the initial vector most similar to the current neighbor node from these identical nodes as the final identical node, and remove the final identical node and the edges of the same type as the edges connecting the current neighbor node from the remaining part. For example Figure 7 in the first layer of the current second sub-graph, there is another neighbor node T (i.e., the green T node). In the remaining part of the corresponding current first sub-graph, there are two T nodes (the red T node and the blue T node). Then calculate the similarity between the initial vector of the red T node and the initial vector of the green T node, and the similarity between the initial vector of the blue T node and the initial vector of the green T node. Select the T node with the highest similarity from the red and blue T nodes, and remove the user-tag edges of the same type as the edges connecting the green T node from Figure 7 it. Assume that the similarity of the red T node is high. The two sub-graphs after removal are as Figure 8 shown.
[0075] If there are no nodes in the remaining part of the current first sub-graph (excluding the part that has been removed for the previous neighbor nodes) that are the same as the current neighbor node, it indicates that the current neighbor node is the part that the current second sub-graph has more than the current first sub-graph. Then mark the current neighbor node as a retrieval failure node.
[0076] S1-30. Respectively determine the removed nodes connected by each remaining node and remaining isolated edge in the current first sub-graph, and extract the nodes determined to be the same as the removed nodes in S1-20 from the current second sub-graph as target nodes; extract the information data of the target nodes from the information data accumulated by the online learning platform, and retrieve the nodes of the same type as each remaining node and / or the relationships the same as each remaining isolated edge from the information data of the target nodes, and supplement them to the simplified second sub-graph. The following will respectively explain S1-30 for each remaining node and remaining isolated edge in the current first sub-graph:
[0077] For each remaining node in the current first sub - graph, perform the following operations respectively: Determine the removed nodes adjacent to the current remaining node, and extract from the current second sub - graph the target nodes determined to be the same as the removed nodes in S1 - 20; Extract the information data of the target nodes from the large amount of information data accumulated by the online learning platform, and retrieve the nodes of the same type as the current remaining node from the information data of the target nodes, supplement them to the simplified second sub - graph, and add corresponding edges between the supplemented nodes and the target nodes. Take the supplemented nodes as the nodes of the same type as the current remaining node in S1 - 20 of the next round of loop. Figure 8 For example, in the remaining part of the current first sub - graph, there is only one remaining node, that is, the blue T node. Then, for this node, determine the removed nodes adjacent to the blue T node, that is, Figure 5 the yellow U node in; Then, extract from the current second sub - graph the target nodes determined to be the same as the yellow U node in S1 - 20 (the same nodes here are nodes of the same type, and different colors are only for easy description and do not represent different nodes), that is, Figure 8 the gray U node in; Retrieve from the large amount of information data of the gray U node the nodes of the same type as the blue T node, that is, search for the T nodes associated with the target second user, and supplement them to the simplified second sub - graph. The result after supplementation is as shown in Figure 9 The purple T node is the supplemented node.
[0078] Finally, if no nodes of the same type as the current remaining node can be retrieved from the large amount of information data, mark the remaining node as a retrieval - failure node in this round of loop. For the sake of easy distinction and description, the retrieval - failure nodes marked in the current second sub - graph in S1 - 20 are called second retrieval - failure nodes, corresponding to the part of the current second sub - graph that is more than the current first sub - graph; The retrieval - failure nodes marked in the current first sub - graph in S1 - 30 are called first retrieval - failure nodes, corresponding to the part of the current first sub - graph that is more than the current second sub - graph but no analogous nodes can be retrieved from the information data of the current second sub - graph.
[0079] For each remaining single - edge in the current first sub - graph, perform the following operations respectively: Determine the two removed nodes connected by the current remaining single - edge, and extract from the current second sub - graph the target nodes determined to be the same as the two removed nodes in S1 - 20; Retrieve the relationship of the same type as the current remaining single - edge from the information of the two target nodes, and supplement it to the simplified second sub - graph. Among them, a single - edge refers to an edge whose two endpoints are both removed. The above operations do not exist in the example of N = 1 shown in Figures 4 - 9 and will be described in the next round of loop when N = 2.
[0080] S1-40. Restore the nodes in the new simplified second subgraph to their complex attributes, and fuse the restored graph into the overall heterogeneous graph. This step performs an operation inverse to the simplification operation in Step 2 on the new simplified second subgraph, that is, further differentiating the same nodes into different nodes of the same type. The new second subgraph after restoration is as shown in Figure 10 . After restoration, perform the fusion operation, adding the newly added nodes and edges in the new second subgraph to the heterogeneous graph. If the newly added node is an existing node in the heterogeneous graph, there is no need to add it repeatedly, and only add an edge between the existing nodes; finally, perform node disambiguation to obtain the fused heterogeneous graph.
[0081] S1-50. Input the fused heterogeneous graph into the first N layers of the multi-layer graph attention network to obtain a new aggregated feature of the target second user. Under the guidance of the target first user, the fused heterogeneous graph re-mines the information data of the target second user and extracts more information; this step then recalculates the aggregated feature of the target second user using the new data to check whether the information mined this time can aggregate more effective information beneficial to video recommendation for the user features. In particular, since each layer of the GAT network corresponds to the aggregation of information of one layer of neighbor nodes, the aggregated feature processed by the first N layers of the multi-layer GAT is taken here.
[0082] S1-60. If the change between the new aggregated feature and the initial feature of the target second user is greater than the third threshold (or another independent threshold, called the fourth threshold), it indicates that the multi-layer GAT network has effectively aggregated the neighbor information of the user, and the initial feature of the user has been fully transformed. Then, the loop from S1-10 to S1-60 can be terminated. Otherwise, if the change between the new aggregated feature and the initial feature of the target second user is less than the third threshold (or the fourth threshold), increase N by 1, and restart the loop from S1-10 to S1-60 according to the simplified first subgraph and the new simplified second subgraph until N reaches the number of layers of the first subgraph or the second subgraph.
[0083] Taking the latter case as an example in this embodiment, referring to Figure 9 , let N = 2, and restart a new round of the loop from S1-10 to S1-60. First, in S1-10, extract the central node and its first N layers of neighbor nodes closest to it, as well as the edges between the extracted nodes, from the simplified first subgraph and the new simplified second subgraph, respectively, to form a new current first subgraph and a new current second subgraph, as shown in Figure 11 .
[0084] In S1-20, the same parts as those in the new current second subgraph are removed from the new current first subgraph. Specifically, first, the central node, the first-layer neighbor nodes, and the first-layer edges are removed from the new current first subgraph. Since, among the first-layer neighbor nodes of the two new current subgraphs, except for the nodes with retrieval failures, the other nodes have been made consistent in the previous round of loop and can all find neighbor nodes identical to themselves, while the nodes with retrieval failures will no longer look for neighbor nodes identical to themselves through subsequent operations, the entire layer can be removed. Then, the following operations are performed for the second-layer neighbor nodes in different cases:
[0085] Case 1: If there are no nodes with retrieval failures in the previous round of loop, the following operations are sequentially performed for the second-layer neighbor nodes in the new current second subgraph: If there is only one node identical to the current neighbor node at the second layer in the new current second subgraph in the remaining part of the new current first subgraph, the identical node and the edges of the same edge type connected to the identical node and connected to the current neighbor node are removed from the remaining part; If there are multiple nodes identical to the current neighbor node in the remaining part of the new current first subgraph, the node with the initial vector most similar to the current neighbor node is selected from these identical nodes as the final identical node, and the final identical node and the edges of the same edge type connected to it and connected to the current neighbor node are removed from the remaining part; If there is no node identical to the current neighbor node in the remaining part of the new current first subgraph, the current neighbor node is marked as the second retrieval failure node of this round of loop. Taking Figure 11 as an example, the remaining part after removal is as Figure 12 shown.
[0086] Case 2: If there are retrieval failure nodes in the previous round of loop, then in the new current first subgraph, mark the nodes in the second-layer neighbor nodes that are only connected to the first retrieval failure node in the previous round of loop as the first retrieval failure nodes of this round of loop and remove them; at the same time, in the new current second subgraph, mark the nodes in the second-layer neighbor nodes that are only connected to the second retrieval failure node in the previous round of loop as the second retrieval failure nodes of this round of loop and remove them; then, perform the following operations on each of the remaining neighbor nodes B in the second-layer neighbor nodes of the new current second subgraph in turn: If there is only one node in the remaining part of the new current first subgraph (excluding the part that has been removed for the previous neighbor nodes and the first retrieval failure nodes of this round of loop, the same below) that is the same as the currently processed neighbor node B, remove the same node, and the edges of the same edge type as the edge connecting the same node and the current neighbor node B, from the remaining part of the new current first subgraph; If there are multiple nodes in the remaining part of the new current first subgraph that are the same as the current neighbor node B, select the node with the initial vector most similar to the current neighbor node B from these same nodes as the final same node, and remove the final same node and the edges of the same edge type as the edge connecting the same node and the current neighbor node B, from the remaining part of the new current first subgraph; If there is no node in the remaining part of the new current first subgraph that is the same as the current neighbor node B, then mark the current neighbor node B as the second retrieval failure node of this round of loop. In this case, for the neighbor nodes of the current layer, first remove the branches extended from the part that was not retrieved to a suitable addition node from the large amount of information data in the previous round of loop in the new current first subgraph, and at the same time remove the part that is more than the new current first subgraph in the new current second subgraph, and then find the same nodes of the remaining neighbor nodes in the new current first subgraph respectively and remove them. The operations of this case are performed in Figure 11 and Figure 12 The example is not shown, but it can be executed in the loops after the second round.
[0087] In S1-30, for the gray V node among the remaining nodes, its adjacent removed node is the red T node in Figure 9 , and the target node corresponding to this node in the new current second subgraph is the green T node. Extract the information data of the green T node (i.e., T21) from the large amount of information data of the online learning platform, and retrieve the nodes of the same type as the gray V node (i.e., the videos associated with T21) from it. For example, the red V node in Figure 13 is extracted and added to the new simplified second subgraph, and the corresponding edges are added. Similarly, for the other remaining green V node, similar operations are performed, and finally the purple V node in Figure 13 is extracted and added to the new simplified second subgraph.
[0088] For the remaining isolated edges (for the sake of distinction, the edge is shown as bold red in Figure 11 and Figure 12 ), determine the two removed nodes connected by the remaining isolated edges (i.e., the blue U node and the red T node in Figure 9 ), extract the target nodes in the current second subgraph that are determined to be the same as the two removed nodes in S1-20 (i.e., the white U node and the green T node in Figure 9 ); from the large amount of information data accumulated by the online learning platform, extract the information data of the two target nodes, and determine the relationship of the same type as the current remaining isolated edge from them, that is, determine whether there is an association relationship between the white U node and the green T node, and supplement it to the new simplified second subgraph, as shown in Figure 13 .
[0089] Then perform the operations of S1-40 to S1-60 again, and so on in a loop until N reaches the maximum neighbor layer number of the first subgraph or the second subgraph.
[0090] S170. Re-input the updated heterogeneous graph into the multi-layer graph attention network, perform multi-hop aggregation on the initial features of each node according to the new graph structure, and recommend at least one new video to each user according to the new final features of each node.
[0091] If there are multiple target second users, the above operations can be performed on each target second user in turn to update the subgraph structure of each target second user in the heterogeneous graph. After the update, re-input the new heterogeneous graph into the multi-layer GAT and MLP for processing, re-predict the probability of each user watching each video, and recommend at least one new video with the highest watching probability to each user. In this recommendation, since more effective information is provided for the target second user, it is at least beneficial to improve the actual viewing rate of the recommended video by the target second user.
[0092] It should be noted that the user information involved in this application is all information and data authorized by the users or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and a corresponding operation entry is provided for users to choose to authorize or refuse.
[0093] In summary, this embodiment provides an online learning video resource personalized recommendation method based on a graph neural network. By integrating learner, learning video, and tag information, a heterogeneous graph is constructed. The graph attention network is used to accurately calculate vectors and update weights, deeply extract features, and model relationships. Finally, a multi-layer perceptron is applied to generate prediction probabilities, realizing accurate personalized recommendation of video resources and effectively improving the learning experience and efficiency. After the recommendation is completed, the actual viewing rate of the recommended videos by the user is recorded. Users with a poor viewing rate are screened, and users whose aggregation and transformation degree of user features by the multi-layer GAT network is not large are selected as improvement targets. At the same time, users who are similar to this user but have a better actual viewing rate are used as reference objects. Targeted mining of the information data of the improvement objects is carried out with reference to their information structure in the isomeric graph, capturing more information that can accurately reflect the complex rules of users' video viewing, further improving the accuracy of the recommended videos, and increasing the actual viewing rate of users.
[0094] Figure 14 The following is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. As Figure 14 shown, the device includes a processor 60, a memory 61, an input device 62, and an output device 63. The number of processors 60 in the device can be one or more. Figure 14 Here, one processor 60 is taken as an example. The processor 60, memory 61, input device 62, and output device 63 in the device can be connected through a bus or other means. Figure 14 Here, connection through a bus is taken as an example.
[0095] The memory 61, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as program instructions / modules corresponding to the online learning video resource personalized recommendation method based on a graph neural network in the embodiment of the present invention. The processor 60 executes various functional applications and data processing of the device by running the software programs, instructions, and modules stored in the memory 61, that is, realizes the above-mentioned online learning video resource personalized recommendation method based on a graph neural network.
[0096] The memory 61 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function. The data storage area can store data created according to the use of the terminal, etc. In addition, the memory 61 can include high-speed random access memory and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 61 can further include a memory remotely set relative to the processor 60, and these remote memories can be connected to the device through a network. Examples of the above network include but are not limited to the Internet, an enterprise internal network, a local area network, a mobile communication network, and their combinations.
[0097] The input device 62 can be used to receive input digital or character information and generate key signal inputs related to the user settings and function controls of the device. The output device 63 can include display devices such as a display screen.
[0098] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the online learning video resource personalized recommendation method based on a graph neural network according to any embodiment.
[0099] The computer storage medium of the embodiment of the present invention can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in combination with an instruction execution system, apparatus, or device.
[0100] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.
[0101] The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0102] Computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the C language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., connected through the Internet using an Internet service provider).
[0103] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.
Claims
1. A personalized recommendation method for online learning video resources based on graph neural network, characterized in that: include: Obtain user information, video resource information, and learning topic labels of the online learning platform, and construct a heterogeneous graph including user nodes, video nodes, and label nodes based on the obtained content; Inputting the heterogeneous graph into a multi-layer graph attention network, performing multi-hop aggregation on the initial features of each node according to the graph structure, and recommending at least one video to each user according to the final features of each node; According to actual viewing rates of the users for the recommended videos, the plurality of users are divided into first users whose actual viewing rates are greater than a first threshold and second users whose actual viewing rates are less than a second threshold, wherein the first threshold is greater than or equal to the second threshold; Determine a target second user whose final feature has a change with the initial feature that is less than a third threshold, and a target first user whose initial feature is most similar to the target second user; According to the first subgraph structure of the target first user in the heterogeneous graph, the second subgraph structure of the target second user in the heterogeneous graph is updated to improve the actual viewing rate of the target second user on the recommended video, and the updated heterogeneous graph is used to re-input the multi-layer graph attention network and recommend at least one new video to each user; specifically, from the heterogeneous graph, a first subgraph diffused outward with the target first user as the center and a second subgraph diffused outward with the target second user as the center are extracted respectively; the node attributes in each subgraph are simplified into node types, wherein the node types include users, videos and labels; N is initialized to 1, and the following operations are performed cyclically, wherein N is the number of layers of neighbor nodes: S1-10, extracting the central node and the first N layers of neighboring nodes closest to it, and the edges between the extracted nodes from the simplified first subgraph and the second subgraph, respectively, to form the current first subgraph and the current second subgraph; S1-20, removing the same portion as the current second sub-image from the current first sub-image; S1-30, respectively determine the removed nodes connected to each remaining node and the remaining isolated edge in the current first subgraph, and extract the target node determined to be the same as the removed node in S1-20 from the current second subgraph, retrieve the nodes of the same type as each remaining node and / or the same relationship as each remaining isolated edge from the information of the target node, and add them to the simplified second subgraph; S1-40, restoring each node in the new simplified second subgraph to a complex attribute, and integrating the restored graph into the overall heterogeneous graph; S1-50, inputting the fused heterogeneous graph into the first N layers of the multi-layer graph attention network to obtain an aggregated feature of the target second user; S1-60. If the change between the aggregated feature and the initial feature of the target second user is greater than the third threshold, terminate the loop from S1-10 to S1-60; otherwise, increase N by 1, and restart the loop from S1-10 to S1-60 according to the new simplified second subgraph until N reaches the maximum number of neighbor node layers of the first subgraph or the second subgraph.
2. The method according to claim 1, characterized in that The step of inputting the heterogeneous graph into a multi-layer graph attention network, performing multi-hop aggregation on the initial features of each node according to the graph structure, and recommending at least one video to each user according to the final features of each node, includes: Generate initial features of each node according to the acquired content; Inputting the initial features and connection relationships of each node into a multi-layer graph attention network, in which the features of each node are updated multiple times according to the features of adjacent nodes to complete multi-hop feature aggregation; The final features of each node are input into the multi-layer perceptron to predict the probability of each user watching each video; At least one video with the highest viewing probability is recommended to each user.
3. The method according to claim 2, characterized in that The method of updating the features of each node multiple times according to the features of adjacent nodes in the multi-layer graph attention network to complete multi-hop feature aggregation includes: In the multi-layer graph attention network, each hop feature aggregation is completed in the following way: For each node i, update the node's feature vector h according to the following formula i : h i ′=σ(∑ j∈N(i)∪{i} a ij Wh j ) Among them, h i ′ is the updated feature vector of node i, N(i) represents the set of all neighboring nodes of node i, W is the learnable weight matrix, h j is the current feature vector of each node in the union of neighbor nodes and node i, σ() is a nonlinear activation function; a ij is the attention coefficient or weight distribution matrix, which is used to measure the importance of different nodes to node i and is calculated as follows: Among them, a is the learnable attention vector, h j is the current feature vector of each node in the union of neighbor nodes and node i, || represents the vector concatenation operation, and LeakyReLU() is the activation function; The updated node feature vector is used as the current feature vector of node i in the next hop aggregation.
4. The method according to claim 2, characterized in that: The step of determining a target second user whose final feature has a change with the initial feature that is less than a third threshold, and a target first user whose initial feature is most similar to the target second user, comprises: For each second user, respectively calculate the similarity between the initial feature and the final feature of the current second user, and take the second user whose difference between 1 and the similarity is less than a third threshold as the target second user; The similarities between the initial features of each first user and the initial features of the target second user are calculated respectively, and the first user with the greatest similarity is selected as the target first user.
5. The method according to claim 1, characterized in that S1-20 includes: Remove the central node and the first N-1 layers of neighboring nodes, and the first N-1 layers of edges from the current first subgraph; For each neighbor node in the Nth layer of the current second subgraph, perform the following operations in sequence: S1-21. If there is only one node identical to the current neighbor node in the remaining part of the current first subgraph, remove the identical node and the edges connected to it that are of the same type as the edges connected to the current neighbor node from the remaining part; S1-22. If there are multiple nodes identical to the current neighbor node in the remaining part of the current first subgraph, select a node whose initial vector is most similar to the current neighbor node from the multiple identical nodes as the final identical node, and remove the final identical node and the edges connected to it that are of the same type as the edges connected to the current neighbor node from the remaining part; The edge types include edges between users and tags, edges between videos and tags, and edges between users and videos.
6. The method according to claim 1, characterized in that S1-30 includes: The following operations are performed for each remaining node: determine the removed nodes adjacent to the current remaining node, extract the target node determined in S1-20 to be the same as the removed node from the current second subgraph; retrieve nodes of the same type as the current remaining node from the information of the target node, add them to the simplified second subgraph, and add corresponding edges between the added nodes and the target node; The following operations are performed for each remaining isolated edge: determine the two removed nodes connected to the current remaining isolated edge, extract the target nodes determined in S1-20 to be the same as the two removed nodes from the current second subgraph; retrieve the relationship of the same type as the current remaining isolated edge from the information of the two target nodes, and add it to the simplified second subgraph.
7. The method according to claim 1, characterized in that The second threshold is less than The value of N x Represents the total number of videos recommended to each user; the first threshold is a value greater than the second threshold.
8. An electronic device, characterized in that: include: one or more processors; a memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the personalized recommendation method for online learning video resources based on graph neural network as described in any one of claims 1-7.
9. A computer-readable storage medium, characterized in that: A computer program is stored thereon, which, when executed by a processor, implements the personalized recommendation method for online learning video resources based on graph neural network as described in any one of claims 1-7.
Citation Information
Patent Citations
Heterogeneous graph neural network course recommendation method and system based on attention mechanism
CN116662651A
Heterogeneous graph neural network-based short video recommendation method and system fused with multi-modal data
CN119089004A
Video content recommendation method and device
CN118233673A
Social network friend recommendation method based on heterogeneous graph neural network clustering
CN118939884A