A method and system for recommending explainable exercises based on relational attributes
Patent Information
- Application Number
- CN202510112077.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-24
Smart Images

Figure CN119557343B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of smart teaching, and specifically relates to a method and system for recommending explainable exercises based on relationship attributes. Background Art
[0002] Intelligent personalized exercise recommendation aims to solve the problem of accurately finding targeted learning resources faced by students when using large-scale online education platforms. Although existing recommendation systems make personalized recommendations based on student interaction data, they usually ignore the rich relationship attributes generated by students during the learning process, such as exercise difficulty, answer accuracy, practice time, etc. These relationship attributes can deeply reflect students' mastery of knowledge points and potential learning difficulties. Therefore, while existing technologies lack in-depth utilization of these attributes, they also fail to effectively improve the interpretability of recommendations, making it difficult for users to understand the basis of recommendations.
[0003] There are several main issues that need to be addressed in the exercise recommendation method in smart education:
[0004] Although existing recommendation methods based on knowledge graphs improve the accuracy of recommendations by mining entity relationships, they lack the ability to explain the recommendation process and results, making it difficult for users to understand the recommendation logic, judge the rationality of recommendations, maintain trust in the system, and obtain guidance on learning directions. This significantly reduces the transparency, rationality, and credibility of the system, greatly affecting the user experience.
[0005] Most methods fail to fully utilize the rich relational attribute information generated by users during the learning process, such as the difficulty of the practice questions, whether hints are used, whether the answers are correct, etc. These relational attribute information can reflect the depth and defects of the user's mastery of knowledge points, and are important information for accurately analyzing user portraits. Ignoring this information will affect the accurate modeling of the user's knowledge status, and thus affect the quality of recommendations.
[0006] Existing reinforcement learning recommendation methods usually adopt static reward and punishment strategies, and cannot dynamically adjust the recommendation strategy according to the user's real-time learning status, knowledge ability level, learning rhythm preference, knowledge point relevance, etc. This lack of flexibility will lead to recommendations that are inconsistent with the user's current needs and abilities, ignore the changes in the user's knowledge status, fail to tailor the appropriate recommendation rhythm for different users, and ignore the inherent relevance between knowledge points.
[0007] Therefore, in order to solve the above problems, the present invention proposes an explainable exercise recommendation method based on relationship attributes. Summary of the invention
[0008] In response to the above problems, the present invention provides an explainable exercise recommendation method based on relational attributes. The entire inventive method is a dynamic relational path reasoning DRPR model, which can significantly improve the explainability of the exercise recommendation system, the accuracy of user modeling and the flexibility of the recommendation strategy.
[0009] A method for explaining exercise recommendation based on relational attributes is as follows:
[0010] Step 1: Process the interaction data between users and exercises in the Junyi Academy Online Learning Activity Dataset, construct triples based on the relationship between the data to build a user-exercise knowledge graph, and initialize the user-exercise knowledge graph;
[0011] Step 2: Input the initialized user-exercise knowledge graph instance into the TransE model for training to obtain the embedding vectors of trained entities and relations;
[0012] Step 3: Recommend user practice exercises on the user-exercise knowledge graph, formalize the user practice exercise recommendation into a Markov decision process, and implement the user practice exercise recommendation by constructing a reinforcement learning system to perform multi-step path reasoning on the user-exercise knowledge graph.
[0013] Step 1 is as follows:
[0014] Step 1.1, load the Junyi Academy Online Learning Activity Dataset, and convert the interaction data between users and exercises in the dataset into user-exercise relationship data. The user-exercise relationship data includes user ID, exercise ID, and exercise completion status. The exercise completion status includes whether it is correct, exercise time, and exercise difficulty;
[0015] Step 1.2: Create a user-exercise knowledge graph based on the user-exercise relationship data. The construction of the user-exercise knowledge graph is as follows:
[0016] User-Exercise Knowledge Graph By entity set and relationship set Constructed triples constitute, ,in Represents the head entity, Indicates relationship, Indicates tail entity, head entity and tail entity Belongs to entity set , head entity and tail entity Contains three types: users, exercises, and concepts. Belongs to a relation set ,relation It includes two types: practice and training. The attributes of the practice relationship include the correctness of the answer. , Practice time and difficulty of exercises Three attributes;
[0017] Step 1.3: User-Exercise Knowledge Graph Initialize and save.
[0018] Step 2 is as follows:
[0019] The initialized user-exercise knowledge graph Loaded into the TransE model, compared with traditional one-hot or rule-based knowledge representation methods, TransE can learn low-dimensional continuous vector representations of entities and relations, which can not only characterize the semantic similarity between entities, but also reveal the structural role of entities in the graph, which facilitates subsequent relational reasoning. In addition, through embedded representation, heterogeneous students, exercises, and knowledge point entities are mapped to the same vector space, so that exercise recommendations can be efficiently implemented based on vector operations. Therefore, the user-exercise knowledge graph Entity Set and relationship set Each entity and relationship is mapped to a dimension In the vector space of and relation embedding vector , entity embedding vector Includes the header entity embedding vector and tail entity embedding vector , then randomly initialize the embedding vectors of entities and relations as the starting point for TransE model training, and then minimize the loss function of TransE and use the user-exercise knowledge graph The initialized embedding vectors are optimized with training data, and finally the optimal entity and relationship embedding vectors are obtained.
[0020] In the process of initializing the embedding vectors of entities and relations through the TransE model, the gradient descent algorithm is used to train the TransE model to continuously optimize the embedding vectors of entities and relations in the TransE model, and to optimize the objective function Optimize the TransE model and optimize the objective function as follows:
[0021] ,
[0022] in, Represents the initialized knowledge graph The set of triples in , represents the triple after random replacement, Represents a set of negative samples generated by randomly replacing the head entity or tail entity in a triple. Negative samples refer to erroneous or unreal triplets generated by randomly replacing the original triples. represents the distance function, express and
[0023] The distance between After replacement and The distance function is calculated by using the L1 or L2 norm. Represents the preset boundary value that controls the distance difference between positive and negative samples. It means taking the positive part;
[0024] The trained TransE model is used to further calculate the score of each triplet. For the positive sample triplet, the trained TransE model makes the score The distance between the head entity and the relation vector and the tail entity vector becomes smaller, so that the positive samples are more closely connected in the vector space. For the negative sample triples, the trained TransE model makes the classification becomes larger, so that the distance between the wrong or untrue triples in the vector space will be enlarged, thereby distinguishing the positive samples from the negative samples. The score of the positive sample triple is expressed as , the score of the negative sample triplet is expressed as ;
[0025] After training is completed, the entity set is obtained and relationship set Each entity in and relationship Optimized entity embedding vector and relation embedding vector .
[0026] Step 3 is as follows:
[0027] Step 3.1: Build a reinforcement learning system:
[0028] The reinforcement learning system uses the user-exercise knowledge graph as the environment and learns the recommendation strategy through the interaction between the reinforcement learning system and the environment;
[0029] The duration of the entire user practice exercise recommendation process in the reinforcement learning system is , where each time step is represented by , 1 ;
[0030] The state of each time step in the process of user practice exercise recommendation is a tuple, expressed as , the initial user state before entering the reinforcement learning system is ( ),in Indicates empty, represents the initial user entity, Indicates that in the system The entity at the time of the step, Indicates A record of entities that the system has passed through before the step. Indicates that the system has reached The combination of relational attribute features at the step time;
[0031] Relationship attribute feature combination The acquisition process is as follows:
[0032] Set the relationship attribute and The corresponding relation embedding vectors are concatenated into a vector The Mamba encoder is sent to the Mamba encoder for feature extraction. The Mamba encoder is a sequence modeling method with linear time complexity. It dynamically selects state subspaces to significantly reduce computational overhead while retaining long-range dependencies. It is particularly suitable for processing long-sequence user interaction data in exercise recommendation scenarios. It can automatically learn the dynamic association rules between different relationship attributes and mine the evolution pattern of students' problem-solving behavior over time. The extracted features are then output mapped through a linear layer to obtain a combination of relationship attribute features. , as the current state Parameters in, relationship attribute feature combination The calculation process is as follows:
[0033] ,
[0034] in, represents the bias of the linear layer, represents a linear layer, Represents the Mamba decoder, Represents the weights of the linear layer.
[0035] Step 3.2: Recommend exercises for users through the reinforcement learning system:
[0036] Step 3.2.1, according to the current system Entity at step time In the user-question knowledge graph The directly related entities and relations in constructing the action space , the action space is represented as , Indicates Step Entity The previous entity, represents the similarity calculation function between the feature vector of the exercise attribute relationship and the student state vector, is the threshold parameter of similarity, the default value is 0.5, Representing Entities Different from the traditional random sampling or similarity-based candidate generation method, the present invention designs an action space construction strategy based on relational attributes. Specifically, the present invention not only considers the structural relationship between the candidate exercises and the current exercises in the knowledge graph (such as whether they are associated with the same knowledge point), but also focuses on evaluating the relational attribute feature vector of the candidate exercises. and the student's current status vector The degree of match between them.
[0037] Then according to the policy network From the current state Action Space Select Action , the calculation process is as follows:
[0038] ,
[0039] Among them, Movement during step Including Entity at step time , No. Entity at step time ,entity With entity The relationship between .
[0040] Step 3.2.2: Execute the selected action After that, the status changes from Transfer to , the parameters in the state are updated as the transfer process progresses, and the calculation is performed given the current state and actions In the case of Probability , the specific process is as follows:
[0041] ,
[0042] Among them, the new state , Indicates The state of the step, represents the initial user entity, Indicates The entity at the time of the step, Indicates A record of entities that the system has passed through before the step. Indicates that the system has reached The combination of relational attribute features at the time of step, Indicates the current The relationship between the first step entity and the next step entity.
[0043] Step 3.2.3. Current Entity When the type is a question entity, a dynamic reward mechanism that considers knowledge point similarity and difficulty adaptability is used. Calculate the current entity Instant Rewards , the calculation process is as follows:
[0044]
[0045] ,
[0046] in, , Indicates Records of entities that the system has passed through before the step The exercise entity in Indicates the relationship type as the difficulty of the exercise The exercise entity, Indicates that the relationship type is the correctness of the answer The exercise entity, Represents the concept of exercise entity, Indicates the current The concept of exercise entity during step time, Indicates the relationship type as the difficulty of the exercise The current Step-by-step exercise entity, is the reward value, and They represent two situations in real recommendations.
[0047] Step 3.2.4: After completing The immediate reward for each time step After calculation, through the scoring function Calculating Terminal Rewards , the calculation process is as follows:
[0048]
[0049] ,
[0050] ,
[0051] in, Indicates The entity at the time of the step, Representing the user-question knowledge graph The set of all exercise entities in , Representing a collection Any exercise entity in Indicates taking the maximum value, express and The relationship between As the head entity in the scoring function, as the tail entity in the scoring function.
[0052] Step 3.3: Output the recommended results of user practice questions:
[0053] The output results include a path set that records the paths traversed by the recommendation system in the knowledge graph. , record the exercise set recommended to the user , record the reward set for each recommended action ;
[0054] in, , represents the initial user entity, Indicates The entity at the time of the step, Indicates Entity at step time; , Indicates The action during the step; , Indicates Entity at step time Corresponding instant rewards.
[0055] The advantages of the present invention are:
[0056] The present invention mainly aims at the problem that it is difficult to accurately locate learning resources in smart education. By performing path reasoning on the knowledge graph, the present invention can not only generate an interpretable path from the user to the recommended content, but also enable the user to understand the recommendation logic, thereby improving the transparency and credibility of the system. In addition, a Mamba-based relationship attribute modeling module is designed, which can efficiently integrate the rich relationship attribute information generated by the user during the practice process, such as the difficulty of the question and the correctness of the answer. Therefore, the present invention can accurately characterize the user's knowledge status, thereby achieving more accurate recommendations.
[0057] In addition, the method also proposes a terminal reward mechanism. The terminal reward quantifies the effect of the entire recommendation sequence as the core indicator for evaluating the quality of the sequence. It not only considers the short-term impact of local recommendation decisions in the recommendation sequence, but also comprehensively considers the contribution of the entire sequence to the user's long-term learning goals. By maximizing the terminal reward, the system can balance local recommendations and global learning goals, avoid simply optimizing short-term benefits while ignoring long-term learning effects, and guide the system to learn to generate recommendation strategies that are more in line with the user's long-term learning goals. The recommendation strategy can then be dynamically adjusted according to the user's exercise adaptability and difficulty adaptability, effectively avoiding the problem of the recommended content being out of touch with the user's actual needs and abilities, and achieving high-quality personalized recommendations. The present invention can not only recommend exercises of appropriate difficulty to promote progressive learning, but also recommend related content based on the relevance of knowledge points, helping users to establish a complete knowledge framework, thereby comprehensively enhancing the pertinence, rationality and user experience of the recommendation system. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.
[0059] Figure 1 It is a schematic diagram of the process of the present invention.
[0060] Figure 2 Schematic diagram of the recommended process of the present invention. DETAILED DESCRIPTION
[0061] In order to more clearly understand the above-mentioned purpose, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.
[0062] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited to the specific embodiments disclosed below.
[0063] Example 1
[0064] Combine the following Figure 1 and Figure 2 A method for recommending explainable exercises based on relationship attributes implemented by the present invention is specifically described.
[0065] A method for recommending explainable exercises based on relationship attributes includes the following steps:
[0066] Step 1: Process the interaction data between users and exercises in the Junyi Academy Online Learning Activity Dataset, construct triples based on the relationship between the data to build a user-exercise knowledge graph, and initialize the user-exercise knowledge graph;
[0067] In addition to using the existing Junyi Academy Online Learning Activity Dataset, you can also choose other existing datasets or collect data yourself to build a dataset;
[0068] Step 2: Input the initialized user-exercise knowledge graph instance into the TransE model for training to obtain the embedding vectors of trained entities and relations;
[0069] Step 3: Recommend user practice exercises on the user-exercise knowledge graph, formalize the user practice exercise recommendation into a Markov decision process, and implement the user practice exercise recommendation by constructing a reinforcement learning system to perform multi-step path reasoning on the user-exercise knowledge graph.
[0070] Step 1 is as follows:
[0071] Step 1.1, load the Junyi Academy Online Learning Activity Dataset, and convert the interaction data between users and exercises in the dataset into user-exercise relationship data. The user-exercise relationship data includes user ID, exercise ID, and exercise completion status. The exercise completion status includes whether it is correct, exercise time, and exercise difficulty;
[0072] Step 1.2: Create a user-exercise knowledge graph based on the user-exercise relationship data. The construction of the user-exercise knowledge graph is as follows:
[0073] User-Exercise Knowledge Graph By entity set and relationship set Constructed triples constitute, ,in Represents the head entity, Indicates relationship, Indicates tail entity, head entity and tail entity Belongs to entity set , head entity and tail entity Contains three types: users, exercises, and concepts. Belongs to a relation set ,relation It includes two types: practice and training. The attributes of the practice relationship include the correctness of the answer. , Practice time and difficulty of exercises Three attributes;
[0074] Step 1.3: User-Exercise Knowledge Graph Initialize and save.
[0075] Step 2 is as follows:
[0076] The initialized user-exercise knowledge graph Loaded into the TransE model, compared with traditional one-hot or rule-based knowledge representation methods, TransE can learn low-dimensional continuous vector representations of entities and relations, which can not only characterize the semantic similarity between entities, but also reveal the structural role of entities in the graph, which facilitates subsequent relational reasoning. In addition, through embedded representation, heterogeneous students, exercises, and knowledge point entities are mapped to the same vector space, so that exercise recommendations can be efficiently implemented based on vector operations. Therefore, the user-exercise knowledge graph Entity Set and relationship set Each entity and relationship is mapped to a dimension In the vector space of and relation embedding vector , entity embedding vector Includes the header entity embedding vector and tail entity embedding vector , then randomly initialize the embedding vectors of entities and relations as the starting point for TransE model training, and then minimize the loss function of TransE and use the user-exercise knowledge graph The initialized embedding vectors are optimized with training data, and finally the optimal entity and relationship embedding vectors are obtained.
[0077] In the process of initializing the embedding vectors of entities and relations through the TransE model, the gradient descent algorithm is used to train the TransE model to continuously optimize the embedding vectors of entities and relations in the TransE model, and to optimize the objective function Optimize the TransE model and optimize the objective function as follows:
[0078] ,
[0079] in, Represents the initialized knowledge graph The set of triples in , represents the triple after random replacement, Represents a set of negative samples generated by randomly replacing the head entity or tail entity in a triple. Negative samples refer to incorrect or untrue triplets generated by randomly replacing the original triplets. Negative samples are used to distinguish correct and incorrect relationships during training. represents the distance function, express and
[0080] The distance between After replacement and The distance function is calculated by using the L1 or L2 norm. Represents the preset boundary value that controls the distance difference between positive and negative samples. It means taking the positive part;
[0081] The trained TransE model is used to further calculate the score of each triplet. For the positive sample triplet, the trained TransE model makes the score The score tends to 0, which reduces the distance between the head entity and the relationship vector and the tail entity vector, making the positive samples more closely connected in the vector space. For the negative sample triples, the trained TransE model makes the score becomes larger, the score tends to 1, so that the distance between the wrong or untrue triples in the vector space will be enlarged, thereby distinguishing the positive samples from the negative samples. The score of the positive sample triple is expressed as , the score of the negative sample triplet is expressed as ;
[0082] After training is completed, the entity set is obtained and relationship set Each entity in and relationship Optimized entity embedding vector and relation embedding vector .
[0083] Step 3 is as follows:
[0084] Step 3.1: Build a reinforcement learning system:
[0085] The reinforcement learning system uses the user-exercise knowledge graph as the environment and learns the recommendation strategy through the interaction between the reinforcement learning system and the environment;
[0086] The duration of the entire user practice exercise recommendation process in the reinforcement learning system is , where each time step is represented by , 1 ;
[0087] The state of each time step in the process of user practice exercise recommendation is a tuple, expressed as , the initial user state before entering the reinforcement learning system is ( ),in Indicates empty, represents the initial user entity, Indicates that in the system The entity at the time of the step, Indicates A record of entities that the system has passed through before the step. Indicates that the system has reached The combination of relational attribute features at the step time;
[0088] Relationship attribute feature combination The acquisition process is as follows:
[0089] Set the relationship attribute and The corresponding relation embedding vectors are concatenated into a vector The Mamba encoder is sent to the Mamba encoder for feature extraction. The Mamba encoder is a sequence modeling method with linear time complexity. It dynamically selects state subspaces to significantly reduce computational overhead while retaining long-range dependencies. It is particularly suitable for processing long-sequence user interaction data in exercise recommendation scenarios. It can automatically learn the dynamic association rules between different relationship attributes and mine the evolution pattern of students' problem-solving behavior over time. The extracted features are then output mapped through a linear layer to obtain a combination of relationship attribute features. , as the current state Parameters in, relationship attribute feature combination The calculation process is as follows:
[0090] ,
[0091] in, represents the bias of the linear layer, represents a linear layer, Represents the Mamba decoder, Represents the weights of the linear layer.
[0092] Step 3.2: Recommend exercises for users through the reinforcement learning system:
[0093] Step 3.2.1, according to the current system Entity at step time In the user-question knowledge graph The directly related entities and relations in constructing the action space , the action space is represented as , Indicates Step Entity The previous entity, represents the similarity calculation function between the feature vector of the exercise attribute relationship and the student state vector, is the threshold parameter of similarity, the default value is 0.5, Representing Entities Different from the traditional random sampling or similarity-based candidate generation method, the present invention designs an action space construction strategy based on relational attributes. Specifically, the present invention not only considers the structural relationship between the candidate exercises and the current exercises in the knowledge graph (such as whether they are associated with the same knowledge point), but also focuses on evaluating the relational attribute feature vector of the candidate exercises. and the student's current status vector The degree of match between them.
[0094] Then according to the policy network From the current state Action Space Select Action , the calculation process is as follows:
[0095] ,
[0096] Among them, Movement during step Including Entity at step time , No. Entity at step time ,entity With entity The relationship between .
[0097] Step 3.2.2: Execute the selected action After that, the status changes from Transfer to , the parameters in the state are updated as the transfer process progresses, and the calculation is performed given the current state and actions In the case of Probability , the specific process is as follows:
[0098] ,
[0099] Among them, the new state , Indicates The state of the step, represents the initial user entity, Indicates The entity at the time of the step, Indicates A record of entities that the system has passed through before the step. Indicates that the system has reached The combination of relational attribute features at the time of step, Indicates the current The relationship between the first step entity and the next step entity.
[0100] Step 3.2.3. Current Entity When the type is a question entity, a dynamic reward mechanism that considers knowledge point similarity and difficulty adaptability is used. Calculate the current entity Instant Rewards , the calculation process is as follows:
[0101]
[0102] ,
[0103] in, , Indicates Records of entities that the system has passed through before the step The exercise entity in Indicates the relationship type as the difficulty of the exercise The exercise entity, Indicates that the relationship type is the correctness of the answer The exercise entity, Represents the concept of exercise entity, Indicates the current The concept of exercise entity during step time, Indicates the relationship type as the difficulty of the exercise The current Step-by-step exercise entity, is the reward value, and They represent two situations in real recommendations.
[0104] Step 3.2.4: After completing The immediate reward for each time step After calculation, through the scoring function Calculating Terminal Rewards , the calculation process is as follows:
[0105]
[0106] ,
[0107] ,
[0108] in, Indicates The entity at the time of the step, Representing the user-question knowledge graph The set of all exercise entities in , Representing a collection Any exercise entity in Indicates taking the maximum value, express and The relationship between As the head entity in the scoring function, as the tail entity in the scoring function.
[0109] Step 3.3: Output the recommended results of user practice questions:
[0110] The output results include a path set that records the paths traversed by the recommendation system in the knowledge graph. , record the exercise set recommended to the user , record the reward set for each recommended action ;
[0111] in, , represents the initial user entity, Indicates The entity at the time of the step, Indicates Entity at step time; , Indicates The action during the step; , Indicates Entity at step time Corresponding instant rewards.
[0112] In actual applications, positive or negative feedback is obtained from users on their mastery of questions of different difficulty levels through instant rewards. If a user makes mistakes when doing exercises of low difficulty, exercises of higher difficulty will not be recommended to the user.
[0113] In order to more intuitively illustrate the interpretability of the recommendation system of the present invention, Figure 2 An example flowchart showing the recommendation process is shown in Figure 2 In the figure, we can see that the recommendation process starts from a user and reasons along the path in the knowledge graph based on the relationship between entities in the knowledge graph. Specifically, after user A completes a certain exercise, the recommendation system recommends other exercises to other users by tracking the knowledge points, concepts and learning paths of other users related to the exercise. This path reasoning process makes the recommendation not just a simple prediction based on the user's historical behavior, but can also generate recommended content through logical deduction through the relationship on the knowledge graph. In this way, the system can provide a clear explanation path for each recommendation, and users can understand why a certain exercise is recommended to them, which greatly enhances the interpretability and transparency of the system.
[0114] In this embodiment, a personalized practice recommendation system is constructed by combining the Mamba modeling module and the dynamic reward strategy: by modeling the rich relationship attributes generated in the student's practice activities, the system can accurately capture the student's performance characteristics and dynamically adjust the recommendation strategy based on these characteristics. Specifically, the method uses knowledge graphs and reinforcement learning techniques to transform the recommendation task into a Markov decision process based on the knowledge graph. The reinforcement learning system learns how to navigate from a given user to potential items and uses the path history as a reason for recommendation, thereby improving the interpretability and accuracy of the recommendation. In this way, the system can not only effectively solve the problem of resource overload, but also provide a reasonable explanation path and increase user trust; in addition, the introduction of a dynamic reward strategy enables the recommendation system to adjust the recommended content in real time, ensuring that students always obtain the practice resources that best suit their current learning status, thereby improving learning efficiency and effectiveness.
[0115] In order to verify the effectiveness of the Mamba modeling module and the dynamic reward strategy, this paper conducted an ablation study on the Mamba modeling module and the dynamic reward strategy, and compared the following four models, as shown in Table 1, which are the reference model A, PGPR (Policy-Guided Path Reasoning) policy-guided path reasoning, model B (removing the Mamba modeling module in DRPR), model C (removing the Reward dynamic reward strategy in DRPR) and the core model D in this invention, DRPR (Dynamic Relational Path Reasoning) dynamic relational path reasoning. The comparison is made in three aspects: NDCG ranking evaluation index, Recall recall rate and Precision accuracy rate. In order to exclude the influence of other factors, the same data set and parameter settings are used uniformly;
[0116] As shown in Table 1, among the four models, the results of Model B are better than those of Model A, which verifies the effectiveness of the Mamba modeling module; the training results of Model C exceed those of Model A, which verifies the effectiveness of the dynamic reward strategy;
[0117] Therefore, after experimental comparison, the following conclusion can be drawn: the DRPR model with the Mamba modeling module and dynamic reward strategy has the best performance, and this method achieves the best performance in multiple evaluation indicators.
[0118] Table 1 Comparison of ablation test performance results
[0119]
[0120] Example 2
[0121] The algorithm for recommending user exercises on the user-exercise knowledge graph is defined as follows:
[0122]
[0123]
[0124]
[0125]
[0126]
[0127]
[0128]
[0129]
[0130]
[0131]
[0132]
[0133]
[0134]
[0135]
[0136]
[0137]
[0138]
[0139]
[0140]
[0141]
[0142]
[0143] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for recommending explainable exercises based on relational attributes, characterized in that: The following steps are involved: Step 1: Process the interaction data between users and exercises in the Junyi Academy Online Learning Activity Dataset, construct triples based on the relationship between the data to build a user-exercise knowledge graph, and initialize the user-exercise knowledge graph; Step 2: Input the initialized user-exercise knowledge graph instance into the TransE model for training to obtain the embedding vectors of trained entities and relations; Step 3: Recommend user practice questions on the user-exercise knowledge graph, formalize the user practice question recommendation into a Markov decision process, and implement the user practice question recommendation by constructing a reinforcement learning system to perform multi-step path reasoning on the user-exercise knowledge graph; Step 3.1: Build a reinforcement learning system: The reinforcement learning system uses the user-exercise knowledge graph as the environment and learns the recommendation strategy through the interaction between the reinforcement learning system and the environment; The duration of the entire user practice exercise recommendation process in the reinforcement learning system is , where each time step is represented by , 1 ; The state of each time step in the process of user practice exercise recommendation is a tuple, expressed as , the initial user state before entering the reinforcement learning system is ( ),in Indicates empty, represents the initial user entity, Indicates that in the system The entity at the time of the step, Indicates A record of the entities that the system has passed through before the step. Indicates that the system has reached The combination of relational attribute features at the step time; Relationship attribute feature combination The acquisition process is as follows: The relationship attribute and The corresponding relation embedding vectors are concatenated into a vector The feature is sent to the Mamba encoder for feature extraction, and then the output is mapped through the linear layer to obtain the relationship attribute feature combination. , as the current state Parameters in, relationship attribute feature combination The calculation process is as follows: , in, represents the bias of the linear layer, represents a linear layer, Represents the Mamba decoder, represents the weight of the linear layer; Step 3.2: Recommend exercises for users through the reinforcement learning system: Step 3.2.1, according to the current system Entity at step time In the user-exercise knowledge graph The directly related entities and relations in constructing the action space , the action space is represented as , Indicates Step Entity Previous entity; Then according to the policy network From the current state Action Space Select Action , the calculation process is as follows: , Among them, Movement during step Including Entity at step time , No. Entity at step time ,entity With entity The relationship between ; Step 3.2.2: Execute the selected action After that, the status changes from Transfer to , the parameters in the state are updated as the transfer process progresses, and the calculation is performed given the current state and actions In the case of Probability , the specific process is as follows: , Among them, the new state , Indicates The state of the step, represents the initial user entity, Indicates The entity at the time of the step, Indicates A record of the entities that the system has passed through before the step. Indicates that the system has reached The combination of relational attribute features at the step time, Indicates the current The relationship between the step entity and the next step entity; Step 3.2.
3. Current Entity When the type is a question entity, a dynamic reward mechanism that considers knowledge point similarity and difficulty adaptability is used. Calculate the current entity Instant Rewards , the calculation process is as follows: , in, , Indicates Records of entities that the system has passed through before the step The exercise entity in Indicates the relationship type as the difficulty of the exercise The exercise entity, Indicates that the relationship type is the correctness of the answer The exercise entity, Represents the concept of exercise entity, Indicates the current The concept of exercise entity during step time, Indicates the relationship type as the difficulty of the exercise The current Step-by-step exercise entity, is the reward value, and Respectively represent two situations in real recommendations; Step 3.2.4: After completing The immediate reward for each time step After calculation, through the scoring function Calculating Terminal Rewards , the calculation process is as follows: , , in, Indicates The entity at the time of the step, Representing the user-question knowledge graph The set of all exercise entities in , Representing a collection Any exercise entity in Indicates taking the maximum value, express and The relationship between As the head entity in the scoring function, as the tail entity in the scoring function.
2. The method for recommending explainable exercises based on relationship attributes according to claim 1, characterized in that: Step 1 is as follows: Step 1.1, load the Junyi Academy Online Learning Activity Dataset, and convert the interaction data between users and exercises in the dataset into user-exercise relationship data. The user-exercise relationship data includes user ID, exercise ID, and exercise completion status. The exercise completion status includes whether it is correct, exercise time, and exercise difficulty; Step 1.2: Create a user-exercise knowledge graph based on the user-exercise relationship data. The construction of the user-exercise knowledge graph is as follows: User-Exercise Knowledge Graph By entity set and relationship set Constructed triples constitute, ,in Represents the head entity, Indicates relationship, Indicates tail entity, head entity and tail entity Belongs to entity set , head entity and tail entity Contains three types: users, exercises, and concepts. Belongs to a relation set ,relation It includes two types: practice and training. The attributes of the practice relationship include the correctness of the answer. , Practice time and difficulty of exercises Three attributes; Step 1.3: User-Exercise Knowledge Graph Initialize and save.
3. The method for recommending explainable exercises based on relationship attributes according to claim 2, characterized in that: Step 2 is as follows: The initialized user-exercise knowledge graph Load into the TransE model and map the user-question knowledge graph Entity Set and relationship set Each entity and relationship is mapped to a dimension In the vector space of and relation embedding vector , entity embedding vector Includes the header entity embedding vector and tail entity embedding vector , then randomly initialize the embedding vectors of entities and relations as the starting point for TransE model training, and then minimize the loss function of TransE and use the user-exercise knowledge graph The initialized embedding vectors are optimized with training data, and finally the optimal entity and relationship embedding vectors are obtained.
4. The method for recommending explainable exercises based on relationship attributes according to claim 3, characterized in that: In the process of initializing the embedding vectors of entities and relations through the TransE model, the gradient descent algorithm is used to train the TransE model to continuously optimize the embedding vectors of entities and relations in the TransE model, and to optimize the objective function Optimize the TransE model and optimize the objective function as follows: , in, Represents the initialized knowledge graph The set of triples in , represents the triple after random replacement, Represents a set of negative samples generated by randomly replacing the head entity or tail entity in a triple. Negative samples refer to erroneous or unreal triplets generated by randomly replacing the original triples. represents the distance function, express and The distance between After replacement and The distance function is calculated by using the L1 or L2 norm. Represents the preset boundary value that controls the distance difference between positive and negative samples. It means taking the positive part; The trained TransE model is used to further calculate the score of each triplet. For the positive sample triplet, the trained TransE model makes the score The distance between the head entity and the relation vector and the tail entity vector becomes smaller, so that the positive samples are more closely connected in the vector space. For the negative sample triples, the trained TransE model makes the classification becomes larger, so that the distance between the wrong or untrue triples in the vector space will be enlarged, thereby distinguishing the positive samples from the negative samples. The score of the positive sample triple is expressed as , the score of the negative sample triplet is expressed as ; After training is completed, the entity set is obtained and relationship set Each entity in and relationship Optimized entity embedding vector and relation embedding vector .
5. The method for recommending explainable exercises based on relationship attributes according to claim 4, characterized in that: Step 3.3: Output the recommended results of user practice questions: The output results include a path set that records the paths traversed by the recommendation system in the knowledge graph. , record the exercise set recommended to the user , record the reward set for each recommended action ; in, , represents the initial user entity, Indicates The entity at the time of the step, Indicates Entity at step time; , Indicates The action during the step; , Indicates Entity at step time Corresponding instant rewards.
Citation Information
Patent Citations
Collaborative recommendation model construction method based on knowledge graph preference propagation
CN113158033A
Explanatable recommendation method based on reinforcement learning and path reasoning
CN117033793A
Individualized exercise recommendation method and recommendation system integrating reinforcement and comparative learning
CN119089030A