Course recommendation system of attention network based on user preference
By using a user preference-based attention network to recommend courses, and leveraging knowledge graphs to expand user preferences and capture higher-order structural relevance, the system solves the challenges of user data sparsity and higher-order signal capture, thereby improving the accuracy of online course recommendations.
Patent Information
- Application Number
- CN202511182875.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-11-21
AI Technical Summary
In massive online courses, it becomes difficult for users to find courses that meet their needs from a vast amount of learning resources. Existing technologies struggle to effectively address the issues of user data sparsity and capturing high-order signals.
A course recommendation system based on user preference attention network is adopted. The user preference propagation module expands user preferences in the knowledge graph, and the project neighbor enhancement module captures high-order structural correlations. The prediction module outputs the probability of the user selecting candidate courses.
It effectively reduces the sparsity of user data, captures complex structural information, improves the performance of recommendation systems, and enhances the accuracy of course recommendations.
Smart Images

Figure CN120994911A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of course recommendation, specifically to a course recommendation system based on an attention network of user preferences. Background Technology
[0002] With the rapid development of the internet and the increasing abundance of online educational resources, the number of large-scale online courses has exploded, making it increasingly difficult for users to find courses that meet their actual needs amidst this vast sea of learning resources. Knowledge graphs themselves contain rich semantic information and can discover deep connections between users and items through multi-hop relational reasoning. Therefore, introducing knowledge graphs into recommendation systems can effectively solve the course recommendation problem in MOOC scenarios. Knowledge-aware recommendation models based on graph neural networks are a mainstream method for addressing the challenges of recommendation systems, but they still have shortcomings in capturing high-order signals and solving the problem of user data sparsity. Therefore, this application proposes a course recommendation system based on an attention network of user preferences. Summary of the Invention
[0003] The purpose of this invention is to provide a course recommendation system based on user preference attention networks to solve the problems mentioned in the background art.
[0004] To achieve the above objectives, the present invention provides the following technical solution: a course recommendation system based on user preference attention networks, comprising:
[0005] The user preference propagation module explores courses that users may have potential preferences by propagating different relationships in the knowledge graph, thus enriching the user's preference representation. The project neighbor enhancement module captures the higher-order neighboring projects of the target course, aggregates the correlations between neighboring projects, and generates the final candidate courses. The prediction module outputs a prediction of the user's choice of candidate courses based on the fused user and project representations.
[0006] Preferably, the user preference propagation module specifically includes a user preference propagation set, a path grouping propagation set, and a path preference aggregation, and the user preference propagation set specifically includes an initial set defined as follows:
[0007]
[0008] The initial set is the course entity, and the course entity is the header entity. By exploring different relational chains along the knowledge graph, tail entities are obtained. These tail entities then become the new head entities in the next iteration. This time received The propagation set of entities can be viewed as a natural extension of user interests. The user preference propagation set is defined as follows:
[0009]
[0010] In other words, the number of courses selected by the user is equal to... Define a The set of triples is as follows:
[0011]
[0012] in Each element is The head entity of a set of triples The set finds candidate items by using different relational chains in the knowledge graph to find other related items.
[0013] Preferably, the path grouping propagation set is specifically traversed through relationship chains. The path group propagation set was obtained, as follows:
[0014]
[0015] in It is a set of entity alignments The links in the code expand the user's path grouping propagation set into a set of triple entities, as follows:
[0016] .
[0017] Preferably, the path preference aggregation is specifically for a given user and candidate courses Given a set of groups and different triples The specific definitions are as follows:
[0018]
[0019] in, It means Embedded vector, For the embedded dimension, each We obtain a weighted sum of all tail entities to get a weighted representation of different paths:
[0020]
[0021] in, , The attention function is defined as follows:
[0022]
[0023] Where the function The specific definitions are as follows:
[0024]
[0025] in, It is a non-linear activation function. and These are the trainable linear transformation matrix and the bias term, respectively. This is a concatenation operation between two vectors with different indices. and These represent parameters at different levels, and then the coefficients of the entire relation path are normalized:
[0026]
[0027] By identifying which types of course knowledge learners pay more attention to, and then adding the obtained attention weights to the propagation triples of different paths, the user embedding representation based on the user preference propagation module is obtained as follows:
[0028] .
[0029] Preferably, the project neighbor enhancement module specifically includes a first-order neighborhood information acquisition and iteration layer, and the first-order neighborhood information acquisition specifically involves traversing the knowledge graph to find candidate courses. A set of one or more directly related entities ,gather For the course The first-order neighbor set, each edge is weighted... To indicate:
[0030]
[0031] in, This represents the user vector. yes Connecting the first The relation vector of neighbors, function This is achieved by calculating the correlation score between users and relationships using the inner product. , representing the target user's relationship The degree of preference, that is, after The weight of message passing at the edge, and then the result Perform regularization.
[0032]
[0033] in, Representative candidate courses The first-order neighbor set is used, and finally, the adjacent entities of the project are aggregated and weighted to obtain candidate courses containing user characteristics. Final embedding:
[0034]
[0035] in, It indicates the first The vector of each of the neighbors.
[0036] Preferably, the specific process of the iterative layer is as follows: the neighbor entities of the candidate project can be further represented as follows: ,in , It is a configurable constant; setting it to 2 will be achieved through the summation aggregator function. Candidate courses and candidate course neighborhood set To synthesize a single vector, the aggregation process is as follows:
[0037]
[0038]
[0039] in It is a non-linear activation function. It is a linear transformation matrix. The bias term is used, while the summation aggregator adds the two vectors together. As the number of iterative layers increases, the model extracts information from multiple jump ranges and integrates it into the project. Based on the project neighbor enhancement module, the final course embedding representation is obtained.
[0040] Preferably, the prediction module takes the following specific steps: given the user embedding... and project embedding By seeking and The probability of a user selecting a course is calculated using the inner product of the terms, as follows:
[0041]
[0042] The loss function is defined as follows:
[0043]
[0044] in, The function is represented as cross-entropy loss. This refers to interactions that the user actually participates in or shows interest in. It is represented as a set containing all negative user-course interactions. Following a uniform distribution, the loss of a knowledge graph is defined as follows:
[0045]
[0046] in, Relationships in a knowledge graph corresponding tensor slices, The embedding matrix of the entities is then used, followed by a regularized loss function:
[0047]
[0048] The loss function of the model is then expressed as:
[0049]
[0050] in, and All of these are hyperparameters.
[0051] Compared with the prior art, the beneficial effects of this invention are as follows:
[0052] The system proposed in this invention can be effectively integrated with knowledge graphs. It not only expands user preferences in the knowledge graph through the user preference module, reducing the sparsity problem of user data, but also captures high-order structural correlations between entities through the item neighbor enhancement module. This enables it to capture and utilize more complex structural information and use two different structures to balance and optimize user representation and item representation. Attached Figure Description
[0053] Figure 1 This is a system framework diagram proposed in this invention;
[0054] Figure 2 Example case diagrams for course recommendations for each user in this embodiment of the invention;
[0055] Figure 3 This is a case analysis diagram of an embodiment of the present invention. Detailed Implementation
[0056] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0057] Example
[0058] Please see Figure 1-2 The illustration shows a recommendation course system based on a user preference-based attention network, comprising:
[0059] The user preference propagation module explores courses that users may have potential preferences by propagating different relationships in the knowledge graph, thus enriching the user's preference representation. The project neighbor enhancement module captures the higher-order neighboring projects of the target course, aggregates the correlations between neighboring projects, and generates the final candidate courses. The prediction module outputs a prediction of the user's choice of candidate courses based on the fused user and project representations.
[0060] Similar to the assumptions of traditional collaborative filtering models, this invention assumes that users' preferences are mainly reflected in their historical interaction data. Specifically, users tend to choose courses similar to those they have interacted with in the past. However, this assumption faces many challenges in practical applications, especially in MOOC scenarios, where the number of users usually far exceeds the number of courses, and user behavior is often concentrated on popular courses. This results in extremely sparse historical interaction data, making it difficult to use user information to mine potential media users.
[0061] The entities in a knowledge graph contain rich facts and connections. The similarity between entities is often measured by their connectivity within the graph; the smaller the distance between entities, the more similar they are in the graph. For example, the course "University Physics I (Mechanics, Thermodynamics)" is related to "Tsinghua University" (school), "Maxwell's Rate Distribution Law" (knowledge point), "Chen Xinyi" (teacher), and "Physics" (course type). However, "Chen Xinyi" (teacher) has a deeper connection with the courses he teaches, "University Physics II (Electromagnetism, Optics, and Quantum Physics)" and "University Physics (Prerequisite)." If a user has previously watched Professor Chen Xinyi's courses, he is likely to be interested in physics or the teacher's courses, and thus might click on "University Physics II (Mechanics, Thermodynamics)" or "University Physics (Prerequisite)" to study. Therefore, a user's personality and preferences are defined not only by the courses they have previously watched but also by the movies they might be interested in within the knowledge graph. Therefore, this invention enhances user representation and user-oriented information through user preference sets and knowledge graphs, thereby bringing better recommendation performance. The detailed process is as follows:
[0062] The user preference propagation set is defined in this invention as the initial seed in the knowledge graph, which considers the set of entities that have interacted with the user.
[0063]
[0064] In order to match real-world recommendation scenarios, the courses... Both can be associated with entities in a knowledge graph. One-to-one correspondence, establishing a set of project entity alignments. This makes it more consistent with real-world recommendation scenarios. By aligning courses with knowledge graph entities, it can provide rich semantic information for interactive data. The user-course interaction matrix can be defined as follows: ,in, and These represent the number of users and courses, respectively. In this interaction matrix, if users and courses If an interaction occurs, such as clicking, adding, or purchasing, then... ,otherwise , Represented as user Courses that have been interacted with The set of courses, and the courses in the course entity set are aligned with the set of courses. The project entities in the text correspond one-to-one, and the initial set is the course entities. The course entities are the header entities. By exploring different relational chains along the knowledge graph, tail entities are obtained. These tail entities then become the new head entities in the next iteration. This time received The propagation set of entities can be viewed as a natural extension of user interests. The user preference propagation set is defined as follows:
[0065]
[0066] Knowledge graphs can enrich user preference information in recommendation systems. This invention formally defines them as follows: in and This represents the head and relation of a knowledge triple. The tail represents the knowledge graph, which is composed of many knowledge triples. The knowledge triple is the basic unit of the knowledge graph. These are the courses selected by users, and in terms of quantity, they equal... To facilitate the calculation of the similarity between previously selected courses and candidate courses, this invention defines a... The set of triples is as follows:
[0067]
[0068] and The difference is that, Each element is The head entity of a set of triples The set finds other related items by leveraging different relationship chains in the knowledge graph, thus identifying candidate items. This process is of great significance for MOOC recommendations.
[0069] Path grouping propagation sets, although the rich semantic information of knowledge graphs extends user encoding, Having too many or too few elements in each layer of the set will both reduce the recommendation effect. Too many elements in each layer, for example, a course like "Data Structures" contains 140 concepts, which is a one-to-many relationship, while other relationships, such as course-to-school or course-to-teacher, are one-to-one. Adding 140 concept entities to each layer of the set makes it difficult to determine which concept in "Data Structures" the user is interested in, and it also distracts the user. Too few elements in each layer will result in insufficient knowledge entities for each connection, which will reduce the effectiveness of representing the user's encoding and may prevent iterations from reaching the next level. Layers, due to the different relationship chains in the knowledge graph, can be viewed as extensions of different attributes of the same course for a user, naturally possessing the advantage of grouping. Based on this, this invention utilizes relationship chain traversal. The path group propagation set was obtained, as follows:
[0070]
[0071] Due to knowledge graphs The link itself It has grouping attributes, so we link them according to relationships. For traversal The entity set obtained this time Grouping the entities to obtain ,in, The tail entities correspond one-to-one with the entities in the knowledge graph. It is a set of entity alignments Similarly, for ease of calculation, the user's path grouping propagation set is expanded into a set of triple entities, as follows:
[0072]
[0073] Path preference aggregation, for a given user and candidate courses Tail entities can be passed from different head entities through different relation chains. For example, the course "Bioinformatics" can be obtained from a course offered by Central South University or from a course taught by Tong Jianbin. Therefore, the semantics of tail entities obtained from different head entities and different relation chains are different within different triples. Thus, to more granularly represent the different semantics of tail entities and candidate courses... The relationship between groups, in a given set and different triples This invention requires weighted representation of the derived tail entities to determine the similarity of candidate courses to courses that the user has interacted with across different relation chains, specifically defined as follows:
[0074]
[0075] in, It means Embedded vector, As the embedding dimension, the resulting similarity can be viewed as a set of groups. Assign different The probability of a triple, and then, the present invention will... The weighted summation of all tail entities yields the weighted representations of different paths:
[0076] in, , .
[0077] The different path groups in each propagation layer can be seen as extensions of user interests in a certain direction. In order to reveal the user's interest hotspots, this invention needs to calculate which types of course information the user pays more attention to through attention weights. The attention function is defined as follows:
[0078]
[0079] Where the function Controlling the weight of each path reflects the user's preference for different paths. Inspired by the attention function in existing technologies, the function... The specific definition is as follows:
[0080] in, It is a non-linear activation function. and These are the trainable linear transformation matrix and the bias term, respectively. This is a concatenation operation between two vectors with different indices. and These represent parameters from different layers. This process yields scores for different paths chosen by the user. Subsequently, to convert these attention scores into weights, the coefficients of the entire relational path need to be normalized.
[0081]
[0082] Therefore, it was found that learners pay more attention to which types of course knowledge.
[0083] Finally, the attention weights obtained in the present invention are added to the propagated triples of different paths, and the user embedding representation based on the user preference propagation module is obtained as follows:
[0084]
[0085] Item neighbor enhancement. The present invention believes that the final representation of a candidate item not only depends on its own features, but is also affected by the item neighbors in the knowledge graph. Specifically, first starting from the candidate item, by iterating along different relationship chains in the knowledge graph, an entity set related to the candidate item is constructed. Subsequently, the attention weights are calculated through the information in the entity set, and then the message passing based on the attention mechanism is realized, so as to capture the neighbor information adjacent to the candidate item. Finally, aggregating to the center to form the final item representation, which can not only capture the potential relationships between different entities, but also aggregate the neighborhood information of the candidate item, and can better explain the preference degree of the relationship to the user behavior. The detailed process is as follows:
[0086] Obtaining first-order neighborhood information. By traversing one or more entities directly related to the candidate course in the knowledge graph, the set obtained is called the first-order neighbor set of the course by the present invention. In the process of message aggregation, if the relationship vector is directly used as the weight of message passing, it will blur the influence of different relationships on the user's behavior. For example, for students who have taken the course <<C++ Programming>>, some choose it because they want to take <<Data Structure>> and need the foundation of <<C++ Programming>>; some choose it because of Teacher Zheng Li; and some choose this course because it belongs to Tsinghua University. Therefore, each edge in the present invention is represented by the weight as follows:
[0087]
[0088] where, represents the user vector, is the relationship vector connecting the th neighbor, and the function calculates the correlation score between the user and the relationship through the inner product, so that can be obtained, which represents the preference degree of the target user for the relationship , that is, the weight of message passing when passing through the edge. Subsequently, the obtained is regularized.
[0089]
[0090] in, Representative candidate courses The first-order neighbor set.
[0091] Finally, by aggregating and weighting adjacent entities within the project, we obtain candidate courses that include user characteristics. Final embedding:
[0092]
[0093] in, It indicates the first The vector of each of the neighbors.
[0094] like Figure 3 As shown, inspired by existing graph sampling algorithms, the entire message passing spreads from the outermost layer to the innermost layer, aggregating layer by layer, getting closer to the project. The inner layer accumulates the outer layer project entities, which are then passed to the central node to obtain the final candidate courses. eigenvectors This allows us to capture higher-order semantic information from the knowledge graph.
[0095] The iterative layer, on the one hand, takes into account the real-world knowledge graph... The actual sizes of the candidates vary; on the other hand, if a candidate project accumulates too many neighbors, it will put a lot of pressure on the model's computation. Considering these two points, this invention provides a solution for each candidate course... The number of first-order neighbors is fixed and randomly sampled, resulting in a set that does not contain all neighbor entities. This achieves both a fixed number of neighbors and efficient computation. Specifically, the neighbor entities of a candidate item can be further represented as... ,in , It is a configurable constant, set to 2.
[0096] Through a summation aggregator function Candidate courses and candidate course neighborhood set The aggregation process to synthesize a single vector is as follows:
[0097]
[0098]
[0099] in It is a non-linear activation function. It is a linear transformation matrix. The bias term is used, while the summation aggregator adds the two vectors together. As the number of iterative layers increases, the model can gradually extract information from multiple jump ranges and integrate it into the project. Through the project neighbor-based enhancement module, the final course embedding representation is obtained.
[0100] The prediction module, as shown in the figure, takes the obtained user embedding and candidate course embedding as input vectors and outputs the prediction result. The specific steps are explained below:
[0101] In a given user embedding and project embedding This invention can be achieved by seeking and The probability of a user selecting a course is calculated using the inner product of the terms, as follows:
[0102]
[0103] First, to balance positive and negative samples and improve classification performance, a negative sampling method is usually adopted, randomly generating an equal number of negative samples for each user. Simultaneously, cross-entropy is used to measure the difference between the predicted results and the actual interaction. The loss function is defined as follows:
[0104]
[0105] in, The function is represented as cross-entropy loss. This refers to interactions that the user actually participates in or shows interest in. It is represented as a set containing all negative user-course interactions. It follows a uniform distribution.
[0106] In order to capture as many potential features as possible about the relationships between entities, this invention follows the principle of constraining the embedding of entities and relationships using knowledge graph embedding technology. The loss of the knowledge graph is defined as follows:
[0107]
[0108] in, Relationships in a knowledge graph corresponding tensor slices, is the embedding matrix of the entity.
[0109] To suppress overfitting and reduce noise interference, this invention incorporates a regularized loss function:
[0110]
[0111] Therefore, the loss function of the model can ultimately be expressed as:
[0112]
[0113] in, and These are all hyperparameters, used to prevent overfitting.
[0114] The following experiments demonstrate the effectiveness of the system proposed in this invention. To verify the effectiveness of the proposed model in online course recommendation, experiments were conducted on the MOOCCube public dataset. Furthermore, to ensure the fairness of the experimental results, comparative experiments were also conducted on the Movielens-1M and Book-Crossing public datasets to verify the model's general applicability in recommendation scenarios.
[0115] MOOCCube 1 It is a large-scale and open-source educational dataset that not only contains 706 real courses and nearly 38,000 videos, but also provides additional information for course recommendations.
[0116] MovieLens-1M 2 It is a dataset widely used for movie recommendation, containing a total of 6,036 user ratings (ranging from 1 to 5) for 2,445 movies.
[0117] Book-Crossing 3 It collected reviews (ranging from 0 to 10) from 17,860 readers across the book community for 2,445 books.
[0118] This invention preprocesses the publicly available dataset MOOCCube, with the following specific steps: First, interaction data from July 1, 2017 to October 1, 2017 is selected, and records where learners selected fewer than 10 courses are removed to ensure sufficient interaction information is extracted and to avoid excessive sparsity of the interaction data. Then, this invention extracts all course names from the obtained dataset, preserving the historical order of learners' course clicks. Finally, to construct training data, this invention follows the standard practice of negative sampling, randomly sampling from items the learner has not interacted with to generate a negative sample, thus obtaining user-course interaction data. Furthermore, regarding the construction of the knowledge subgraph, this invention starts from course entities, expanding from different relations into different course knowledge triples until the tail entity reverts to the course entity.
[0119] Since the interaction data of MovieLens-1M and Book-Crossing are explicit feedback, this invention follows the RippleNet approach, transforming them into implicit feedback to obtain user-item interaction records. The threshold for positive samples in MovieLens-1M is 4, while no specific threshold is set for Book-Crossing due to its sparsity. If a user gives a positive rating to an item, the rating is set to 1; if a user has never interacted with an item or gave a very low rating, the rating is set to 0. Furthermore, in constructing the sub-knowledge graph, this invention follows the KGCN approach, using Microsoft's commercial knowledge graph Satori to construct the MovieLens-1M and Book-Crossing datasets. The final sub-knowledge graph is a subset of the original knowledge graph, consisting of triples with a confidence greater than 0.9, all constructed in standard triplet form.
[0120] Table 1. Dataset Information Statistics Table
[0121] Evaluation metrics and experimental setup: To evaluate the performance of the proposed method, this invention predicts whether a user will click based on a trained model, and evaluates the performance of the click-through rate (CTR) prediction task using two commonly used evaluation metrics: AUC and F1 score. AUC (the area under the curve) is an evaluation metric for measuring the performance of classification tasks, ranging from 0 to 1, with higher values indicating higher performance.
[0122] To ensure fairness in the experiment, the embedding size (d) of all models was set to 64, and the batch size was set to 512. The dataset was randomly divided into training, testing, and validation sets in a 6:2:2 ratio to facilitate hyperparameter tuning. The model parameters were initialized using existing techniques, optimized using existing techniques, and other hyperparameters were fine-tuned using grid search to obtain suitable hyperparameters. The specific hyperparameters requiring fine-tuning are as follows: the learning rate was evaluated between {0.0001, 0.001, 0.01, 0.1}. , in {10 -7 10 -6 10 -5 10 -4 10 -3 10 -2 10 -1 Check the regularization coefficient between} The model explores the maximum number of hops L between {1,2,3,4,5} and finds a suitable number of neighbors D between {2,3,4,5,6,7,8}. To save training time, training stops if the model's loss on the validation set does not significantly improve over five consecutive validation epochs, and the training result is the average of the best results obtained from five runs. Table 4 lists the specific hyperparameter settings of the models, while the hyperparameters of other baseline models are based on experience or follow the settings of the original paper to ensure optimal performance for each model. The hardware configuration relied upon for the experiments included a computing platform equipped with an NVIDIA GeForce RTX 4050 graphics card (GPU) with at least 8GB of video memory and an AMD Ryzen 7840HS processor (CPU). All models were implemented in the PyTorch framework and shared the same hardware acceleration.
[0123] To validate the performance of the proposed model KGCN-UP, this invention compares KGCN-UP with five classic recommended methods: CF-based methods (LFM, BPRMF), Embedding-based methods (CKE, MKR), Path-based methods (RippleNet, PER), GNN-based methods (KGCN, KGNN-LS, KGAT, CKAN, KGIN, KGIE, CG-KGR), and CL-based methods (KGCCL, KGRec), which are described in detail below:
[0124] LFM: It is a collaborative filtering method that uses ALS decomposition to obtain latent factors of users and items, constructs an interaction matrix of users and items, and represents the user's rating of the item.
[0125] BPRMF: It is a collaborative filtering-based method that uses pairwise matrix factorization to obtain implicit feedback and optimizes the BPR loss to rank user preferences.
[0126] CKE: It is an embedding-based approach that introduces knowledge graphs into recommendation systems and uses the TransR algorithm to fuse structured knowledge extraction features with collaborative filtering.
[0127] MKR: It is an embedding-based method that combines recommendation prediction tasks with knowledge graph embedding methods and updates cross-compression units synchronously to improve recommendation performance.
[0128] RippleNet: It is a meta-path based method that propagates users' latent preferences through different relationship links in the knowledge graph, resulting in a final item representation that includes the user's preferences.
[0129] PER: It is a meta-path-based method that calculates the path similarity between items to obtain the preference score between users and items, forming a user preference diffusion matrix to represent the connection between users and items.
[0130] KGCN: It is a GNN-based approach that captures high-order semantic information in knowledge graphs through message aggregation and recursively forms the final item entities to enrich the item representation.
[0131] KGAT: It is a GNN-based method that recursively updates the embeddings of neighboring nodes while employing an attention mechanism to dynamically adjust the weights of each neighbor, thereby distinguishing the importance of different neighbors in node embeddings.
[0132] CKAN: It is based on GNN and adopts a new propagation strategy that combines collaborative information with knowledge information. It dynamically adjusts the weights of each neighbor node through an attention mechanism to obtain better recommendations.
[0133] KGIN: It is a GNN-based method that constructs a relational path awareness, mines the rich semantic information contained in the knowledge graph, extracts the key features of user intent, and enriches the embedded representation of users and items.
[0134] KGIE: It's a GNN-based approach that constructs a user-item interaction matrix and then fuses this matrix with item information using GNN. During the recommendation process, it effectively captures user information and preferences while considering user context.
[0135] KGCCL: It is a CL-based approach that combines interaction graphs and knowledge graphs, using noise enhancement to improve the representation of users and items by comparing local and global aspects.
[0136] KGRec: It is a CL-based approach that proposes a novel knowledge masking mechanism to mask high-scoring triples by aligning signals from the knowledge graph and interaction graph to mask the knowledge graph.
[0137] Based on the performance results of all the above models, their performance differences were analyzed, and the observations are shown in Table 2 below:
[0138] Table 2 Performance of different models
[0139]
[0140] The specific observations based on Table 2 are as follows:
[0141] KGCN-UP achieved the best results, consistently outperforming all benchmark models across three different datasets for all evaluation metrics. Specifically, compared to the strongest benchmark model, the KGCN-UP model showed significant improvements in AUC on the Book, Movie, and Course datasets, with increases of 0.14%, 0.24%, and 4.94%, respectively. This invention attributes the improved model performance to the following aspects: (1) the user preference propagation module is better able to capture users' higher-order preferences; and (2) the item neighbor enhancement module can capture higher-order semantic information of the knowledge graph.
[0142] Incorporating knowledge base (KG) information is beneficial to recommendation systems. Compared with traditional recommendation algorithms LFM and BPRMF, the CKE model simply embeds knowledge information into the MF, yet achieves a significant performance improvement, demonstrating the effectiveness of the knowledge information and consistent with previous research findings.
[0143] GNNs exhibit superior model performance. Most graph neural network-based methods outperform embedding- and path-based methods, highlighting the importance of information propagation between graph nodes. This inspired this invention to suggest that enhancing user and item representations can improve model performance when interaction data is sparse.
[0144] The KGCN-UP model traverses all courses a user has historically interacted with, propagating its insights based on different relationship chains within the knowledge graph to identify courses the user might like. This propagation process is known as the user preference propagation module. Similarly, starting with candidate courses, the model traverses related courses adjacent to the selected item to enhance the candidate course representation, forming an item enhancement module. This model effectively mines users' higher-order preferences and fully utilizes the information from the knowledge graph's relationship chains, thus solving the "information overload" problem and improving model performance. Figure 2 This is a visual example of the model.
[0145] In this embodiment, the effectiveness of the proposed KGCN-UP is demonstrated in a CTR prediction task, such as... Figure 3 As shown, to visualize the user preference propagation module of KGCN-UP, a user was randomly selected and a portion of their interaction history was displayed, including courses such as "Data Structures" and "Software Engineering." It can be concluded that the proposed model produces different prediction results for different candidate courses. It is also evident that the user shows a stronger preference for schools offering more courses and a higher interest in the course content. In contrast, although Tsinghua University also offers the "Software Engineering" course, the user's selection probability remains low because the user is more focused on learning "Data Structures" rather than solving engineering project-related content.
[0146] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0147] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A course recommendation system based on user preference attention networks, characterized in that, include: The user preference propagation module uses different relationships in the knowledge graph to propagate courses that explore users' potential preferences, thus enriching the representation of users' preferences. The project neighbor enhancement module captures the higher-order neighboring projects of the target course, aggregates the correlation between neighboring projects, and generates the final candidate courses; the prediction module outputs the prediction of the user's choice of candidate courses based on the fused user and project representations.
2. The recommended course system based on user preference attention network according to claim 1, characterized in that: The user preference propagation module specifically includes a user preference propagation set, a path grouping propagation set, and a path preference aggregation, and the user preference propagation set specifically includes an initial set defined as follows: The initial set is the course entity, and the course entity is the header entity. By exploring different relational chains along the knowledge graph, tail entities are obtained. These tail entities then become the new head entities in the next iteration. This time received The propagation set of entities can be viewed as a natural extension of user interests. The user preference propagation set is defined as follows: In other words, the number of courses selected by the user is equal to... Define a The set of triples is as follows: in Each element is The head entity of a set of triples The set finds candidate items by using different relational chains in the knowledge graph to find other related items.
3. The recommended course system based on user preference attention network according to claim 2, characterized in that: The path grouping propagation set is specifically traversed through the relationship chain. The path group propagation set was obtained, as follows: in It is a set of entity alignments The links in the code expand the user's path grouping propagation set into a set of triple entities, as follows:
4. The recommended course system based on user preference attention network according to claim 3, characterized in that: The path preference aggregation specifically refers to, for a given user and candidate courses Given a set of groups and different triples Specifically defined as follows in, It means Embedded vector, For the embedded dimension, each We obtain a weighted sum of all tail entities to get a weighted representation of different paths: in, , The attention function is defined as follows: Where the function The specific definitions are as follows: in, It is a non-linear activation function. and These are the trainable linear transformation matrix and the bias term, respectively. This is a concatenation operation between two vectors with different indices. and These represent parameters at different levels, and then the coefficients of the entire relation path are normalized: By identifying which types of course knowledge learners pay more attention to, and then adding the obtained attention weights to the propagation triples of different paths, the user embedding representation based on the user preference propagation module is obtained as follows: 。 5. A course recommendation system based on user preference attention networks according to claim 4, characterized in that: The project neighbor enhancement module specifically includes a first-order neighborhood information acquisition and an iteration layer. The first-order neighborhood information acquisition specifically involves traversing the knowledge graph to find candidate courses. A set of one or more directly related entities ,gather For the course The first-order neighbor set, each edge is weighted... To indicate: in, This represents the user vector. yes Connecting the first The relation vector of neighbors, function This is achieved by calculating the correlation score between users and relationships using the inner product. , representing the target user's relationship The degree of preference, that is, after The weight of message passing at the edge, and then the result Perform regularization. in, Representative candidate courses The first-order neighbor set is used, and finally, the adjacent entities of the project are aggregated and weighted to obtain candidate courses containing user characteristics. Final embedding: in, It indicates the first The vector of each of the neighbors.
6. A course recommendation system based on user preference attention networks according to claim 5, characterized in that: The specific process of the iterative layer is as follows, where the neighbor entities of candidate projects can be further represented as follows: ,in , It is a configurable constant; setting it to 2 will be achieved through the summation aggregator function. Candidate courses and candidate course neighborhood set To synthesize a single vector, the aggregation process is as follows: in It is a non-linear activation function. It is a linear transformation matrix. The bias term is used, while the summation aggregator adds the two vectors together. As the number of iterative layers increases, the model extracts information from multiple jump ranges and integrates it into the project. Based on the project neighbor enhancement module, the final course embedding representation is obtained.
7. A course recommendation system based on user preference attention networks according to claim 6, characterized in that: The specific steps of the prediction module are as follows: given the user embedding... and project embedding By seeking and The probability of a user selecting a course is calculated using the inner product of the terms, as follows: The loss function is defined as follows: in, The function is represented as cross-entropy loss. This refers to interactions that the user actually participates in or shows interest in. It is represented as a set containing all negative user-course interactions. Following a uniform distribution, the loss of a knowledge graph is defined as follows: in, Relationships in a knowledge graph corresponding tensor slices, The embedding matrix of the entities is then used, followed by a regularized loss function: The loss function of the model is then expressed as: in, and All of these are hyperparameters.