A Reinforcement Learning Path Planning Method Based on a Generative Adversarial User Model
Through the reinforcement learning method based on the generation of adversarial user models, the learning path is dynamically adjusted, and the problem of insufficient personalized and real-time recommendation of learning paths in the prior art is solved, and an efficient and personalized learning resource path is provided in the case of changes in user preferences and learning abilities.
Patent Information
- Application Number
- CN202210528946.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-16
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-05-16
AI Technical Summary
The existing learning path recommendation method is difficult to dynamically adjust the learning path when user preferences and learning abilities change rapidly, resulting in poor recommendation results. Especially in course learning, the path sequence algorithm is not applicable, and the path generation algorithm ignores user performance changes.
The reinforcement learning path planning method based on the generation and adversarial user model is adopted, and the learner similarity matrix is constructed, and the user learning behavior type clustering is used to cluster, and the reinforcement learning model of the hierarchical reward function is constructed in combination with the knowledge forest, and the learning resource path is generated using the cascading DQN algorithm.
It realizes that on the basis of considering users' long-term and current learning interests, the learning path is adjusted in real time to adapt to the feedback changes of online learners, improves the personalization and real-timeness of the learning resource path, and improves the computing efficiency.
Smart Images

Figure CN115249072B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for learning resource path planning, and more particularly to a reinforcement learning path planning method based on a generative adversarial user model. Background Art
[0002] Existing learning path recommendation algorithms can be divided into two categories: path generation and path sequence. After determining the characteristics and requirements of the user, path generation algorithms generate the entire learning path in a single recommendation and conduct learning evaluation only after completing the entire path. Kardan proposed a two-stage path generation method. In the first stage, the K-means algorithm is used to group users according to the results of pre-tests. In the second stage, the ant colony optimization method is used to generate a path for each group; Zhan Li generates three types of learning paths, namely deadline-driven paths, goal-driven paths, and sorted paths (considering the user's sorting preferences) based on a graph search algorithm according to given user input constraints such as learning objectives, starting points, and preferred rankings of the output paths; Adorni and Koceva apply an Educational Concept Map (ECM) to generate paths. The user determines the knowledge background, starting point, and end point by selecting a set of topics from the ECM, and uses ENCODE to generate paths. Path sequence algorithms recommend learning paths step by step according to the user's progress in the learning path. Govindarajan applies a parallel particle swarm optimization algorithm to predict the user's dynamic path; Yarandi proposed an ontology knowledge-based model that takes the user's ability, knowledge background, learning style, and preferences as inputs and recommends paths; Salahli uses item response theory to estimate the user's understanding of knowledge and thus plan the path.
[0003] As can be seen from the above literature, in learning path recommendation, accurately profiling the user himself is an important aspect, and it is often necessary to combine the static and dynamic characteristics of the user to establish the best user model. Especially as time goes by, characteristics such as the user's preferences and learning ability will change, and the recommended learning path should also change dynamically. How to accurately model the user when characteristics such as the user's preferences change rapidly is the difficulty of adaptive path recommendation. In existing path planning methods, path sequence algorithms often need to rely on the results of knowledge tracing for cognitive diagnosis and are commonly used for exercise recommendation, but are not suitable for course learning; while path generation algorithms mostly ignore the changes that occur in the user's performance and learning process, which may lead to incorrect recommendations after the user's state changes, and the search speed is relatively slow. Therefore, how to adaptively adjust the path in combination with the user modeling results and recommend a learning path suitable for the learner's learning preferences and learning progress in real time is an urgent problem to be solved. Summary of the Invention
[0004] The purpose of the present invention is to overcome the shortcomings of the above-mentioned prior art and provide a reinforcement learning path planning method based on a generative adversarial user model.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] A reinforcement learning path planning method based on a generative adversarial user model includes the following steps:
[0007] 1) Obtain and construct a learner similarity matrix W according to the user learning log, and use the spectral clustering method to complete the clustering of user learning behavior types on the learner similarity matrix W to obtain N types of user learning behavior types {Cluster ui | ui = 1,…, N}, and each type of learning behavior type corresponding training data set D can be divided according to the user learning behavior type ui ;
[0008] 2) Combine with a knowledge forest to construct a path planning model based on hierarchical reward function reinforcement learning. The reward function in the path planning model based on hierarchical reward function reinforcement learning is a two-level reward function composed of a sequential decision reward and a knowledge point planning reward, and use the user behavior model as the environment of reinforcement learning, and train the path planning model in the form of generative adversarial training;
[0009] 3) Use the user learning behavior type, user historical learning sequence, target knowledge point, learning resource set and course knowledge forest as inputs, and complete the learning resource path planning to the target knowledge point based on the cascaded DQN algorithm, and output the planned path.
[0010] Further, the specific operation of constructing the learner similarity matrix W in step 1) is: obtain the course learning status state of each learner ui,course , the average time-consuming ratio of the completed knowledge points the average centrality of the completed knowledge points the number of completed key knowledge points and the learning status state of the target knowledge point ui,target , and construct a learner scoring vector U i :
[0011]
[0012] Calculate the cosine similarity between the normalized learner scoring vectors, and construct the learner similarity matrix W:
[0013] 3. The reinforcement learning path planning method based on a generative adversarial user model according to claim 2, wherein the specific process of using the spectral clustering method to combine with the similarity matrix W to complete the clustering of N types of user learning behaviors and the division of the data set in step 1) is as follows:
[0014] Construct the degree matrix D and the Laplacian matrix L respectively:
[0015]
[0016] L = D - W (6)
[0017] Use to standardize L, then calculate the eigenvectors of the first N smallest eigenvalues, form an M * N-dimensional matrix with the N eigenvectors, standardize it row by row to obtain the matrix F, take each row in the matrix F as an N-dimensional sample, a total of M samples, and use k-means for clustering to obtain the final N-class classification result, and divide the learners into N different learning behavior types {Cluster ui | ui = 1,..., N}, and accordingly divide the user logs to obtain the training data set D corresponding to each learning behavior type ui .
[0018] Furthermore, in step 2) of constructing the path planning model based on the hierarchical reward function reinforcement learning, the five-tuple M = (s t , A t , P(·|s t , A t ), r(s t , a t ), γ) of the Markov decision process corresponding to the reinforcement learning;
[0019] Among them, the learner is used as the environment, and the state s t represents the historical learning resource sequence of the learner before the t-th moment, and the action a t represents selecting a learning resource from the set L t of candidate learning resources at the t-th moment and recommending it to the learner. The action set A t represents the set of k actions corresponding to the learning resource path of length k recommended to the learner at the t-th moment; the state transition probability P(·|s t , A t ) corresponds to the probability of transferring to the next state s t given the state s t and the action set A t+1 , and at the same time serves as the uniform distribution of the user's actions The reward function r(s t , a t ) and the discount factor γ.
[0020] Furthermore, decompose the reward function r(s t , a t ) into a sequential decision-making reward r seq and a knowledge point planning decision-making reward r c , that is, r = r seq + r c ;
[0021] When calculating the sequential decision-making reward r seq , calculate the sequential level accuracy of the recommended subsequence and the actual interaction subsequence, as shown in Equation (8):
[0022]
[0023] In Equation (8), prec m represents the sequential decision-making accuracy, i t:t+k is the actual interaction subsequence, is the recommended subsequence, p m is a subsequence of length m of the subsequence i t:t+k , and M represents the number of subsequences of length m used;
[0024] When calculating the knowledge point planning decision-making reward function r c , considering whether the difficulty of the recommended learning resource matches the difficulty of the actually clicked learning resource, estimate the difficulty of the learning resource using the learning duration, as follows:
[0025]
[0026] In Equation (9), the actual learning subsequence of the user is i t:t+k , the predicted user learning sequence is c t:t+k is the representation vector used to represent the actual learning sequence of the user, is the representation vector used to represent the predicted user learning sequence, and the sequence representation vector c t:t+k is calculated by taking the mean of the feature vectors of each learning resource in the sequence. c t+i and are respectively used to represent the feature vectors of the i-th learning resource in the actual and predicted user learning resource learning sequences; v t+i represents the feature vector of the i-th learning resource, dur total represents the default learning duration of this learning resource, and dur watch represents the learning duration of the user on this learning resource.
[0027] Furthermore, the method for constructing the user behavior model and joint training in step 2) is as follows:
[0028] For each Clusterui , design a user behavior model Learned strategy As a probability distribution over the action set A t ={a 1 , a 2 ,..., a n}, when calculating the reward function, the current action a t and the user's state s t are both used as inputs to the reward function r(s t , a t ). The strategy adopted by the user will maximize the expected reward r(s t , a t ). When solving, it is regarded as an optimization problem of the strategy in its probability distribution space Δ k-1 for solution;
[0029] On the dataset D ui corresponding to each type of learning behavior, in the form of generative adversarial training, the user behavior model is regarded as the generator, and the reward function r ui is regarded as the discriminator to complete the parameter learning of the path planning model Planer ui to obtain N Planers ui for simulation.
[0030] Furthermore, the optimization problem is solved as follows:
[0031]
[0032] Among them, the regularization term uses negative Shannon entropy, and the parameter η is used to control the strength of regularization.
[0033] Furthermore, the calculation method of the min-max function when implementing generative adversarial training is:
[0034] According to T user action sequences in the historical behavior and the corresponding features of the clicked course resources calculate the state and jointly learn the user behavior model and the reward function r, see Equation (11):
[0035]
[0036] In Equation (11), α represents all the parameters used in the model , and θ represents all the parameters used in the reward function r.
[0037] Further, the specific method for generating a recommended learning resource path using the cascaded DQN algorithm in step 3) is as follows: For the target knowledge point k target , the learning resource set is According to the action decision strategy that maximizes the current Q-function value for each learning resource recommendation, using a cascaded manner, find the optimal action that maximizes each level of the Q-function, and iterate step by step until a learning resource containing the target knowledge point is found, and output the planned path.
[0038] Further, step 3) also includes: If the learner user i has no learning record, then based on the idea of behavior cloning, complete the learning resource path planning based on similar users in the same course with the same major or the same grade in history, specifically:
[0039] Given the target knowledge point k target , according to the learner's grade, school, and major information, perform similarity matching among users with existing learning histories, find users with the same major or the same grade in the same course in the historical records, and use the learning history of the similar users to generate a path to the target knowledge point for users without learning histories.
[0040] Compared with the prior art, the present invention has the following beneficial effects:
[0041] The reinforcement learning path planning method based on the generative adversarial user model of the present invention, compared with the existing path planning methods, the reinforcement learning path planning method of the present invention can, while considering the inherent knowledge structure of the learning resources, consider the user's long-term learning interests and current learning interests, and combine user feedback to provide the user with a learning resource path to the target knowledge point; the model proposed by the present invention can handle the situation where the feedback of online learners changes in real time, and uses the form of combining the user behavior model and the reinforcement learning path planning model to provide real-time path planning results for the learners; the reinforcement learning path planning method proposed by the present invention belongs to the model-based reinforcement learning method, which can learn good recommendation strategies with less user interaction and can quickly learn new user dynamics; the cascaded DQN algorithm used in the reinforcement learning model of the present invention is used to obtain a combined recommendation strategy, which can find the best learning resource subset from a large number of candidates, and the time complexity of this algorithm is only linearly related to the number of candidate objects, which can greatly improve the model calculation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is the flowchart of the learning resource path planning of the present invention;
[0043] Figure 2 is the schematic diagram of the reinforcement learning model framework using the combined user generation model;
[0044] Figure 3 It is a framework diagram of the cascaded DQN algorithm model. Specific implementation manners
[0045] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0046] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0047] Different from the mainstream learning resource recommendations which are mainly point-level resource recommendations based on resource popularity, professional categories, user similarity, etc., in the present invention, a learning path consists of a sequence of learning resources. The learning path planning is applicable to many scenarios. For example, when newly learning a certain course, a learning path for the course knowledge needs to be planned; another example is that when self-studying a new knowledge point, a learning path from the latest learned knowledge point to the target knowledge point needs to be planned. These scenarios require planning the learning resource sequence at the path level according to the user's learning goals, learning preferences, etc., that is, recommending personalized learning paths.
[0048] The present invention will be further described in detail below in conjunction with the accompanying drawings:
[0049] See Figure 1 , Figure 1 which is a flowchart of the present invention. The learning resource path planning method based on reinforcement learning of the present invention includes the following steps:
[0050] Step 1: Division of user groups and training data sets driven by big data
[0051] Obtain the course learning status state of each learner ui,course , the average time-consuming ratio of the completed knowledge points Average centrality of completed knowledge points Number of completed key knowledge points And the learning status state of the target knowledge point ui,target , calculate the similarity matrix W of learners according to the above indicators, and use the spectral clustering method to complete the clustering of user learning behavior types on the similarity matrix W, and N types of user learning behavior types {Cluster ui | ui = 1,..., N} can be obtained, and the training data set D corresponding to each learning behavior type can be obtained accordingly ui , specifically:
[0052] 101) Analyze the learning logs of learners. For each learner user i Obtain their course status state ui,course , average time-consuming ratio of completed knowledge points Average centrality of completed knowledge points Number of completed key knowledge points num ui And the learning status state of the target knowledge point ui,tarqet ; among them, for the course learning status state ui,course , the completed course status is recorded as 0, and the uncompleted course status is recorded as 1; for the calculation of the average time-consuming ratio of completed knowledge points , the knowledge point time-consuming ratio is the ratio of the average learning duration to the original duration of itself, as shown in formula (1). In formula (1), dur sum Represents the total learning duration of knowledge point i, frequency sum Represents the total learning frequency of knowledge point i, dur i Represents the original duration of knowledge point i; the knowledge point centrality degree i Is defined as the degree centrality of the node. The higher the degree d i Of the knowledge point, the higher its importance. The calculation is shown in formula (2). In formula (2), n i Represents the degree of the node, and n represents the number of nodes in the graph; the number of completed key knowledge points num ui Is the number of knowledge points with a knowledge point centrality greater than 0.2 in the historical learning of learner user i ; if the learner does not specify the target knowledge point, the last knowledge point of this course is regarded as the target knowledge point, and the learning status state of the target knowledge point ui,target Is expressed as: uncompleted is expressed as 0, and completed is expressed as 1
[0053]
[0054]
[0055] 102) Divide different learner types using spectral clustering
[0056] According to the course state state of the learner ui,course , average time consumption ratio of completed knowledge points Average centrality of completed knowledge points Number of completed key knowledge points num′ ui and the learning state state of the target knowledge point ui,target , construct the learner score vector U i :
[0057]
[0058] Calculate the cosine similarity between the normalized learner score vectors, and construct the learner similarity matrix W:
[0059] According to the learner similarity matrix W, successively construct the diagonal matrix D and the Laplacian matrix L:
[0060]
[0061] L = D - W (6)
[0062] Normalize the Laplacian matrix L, that is , to obtain Subsequently, calculate the matrix eigenvalues of, sort them in ascending order according to the numerical values of the solved eigenvalues, obtain the eigenvectors of the first N smallest eigenvalues, form an M * N-dimensional matrix with the N eigenvectors, normalize each row of the matrix to obtain the matrix F, take each row in the matrix F as an N-dimensional sample, a total of M samples, and use k-means for clustering to obtain the final N-class classification result, and divide the learners into N different learning behavior types.
[0063] Step 2: Path planning model Planer based on hierarchical reward function reinforcement learning ui Offline training
[0064] Combine with the knowledge forest to construct a reinforcement learning framework for learning resource path planning:
[0065] 201) The main idea of constructing the reinforcement learning framework is to regard it as a Markov decision process, as Figure 2 shown, the five-tuple M = (s t , A t , P(·|s t , A t ), r(s t , a t) is defined as follows: regarding the learner as the environment E, the state s t is defined as the historical learning resource sequence of the learner before time t, and the action a t is defined as selecting a learning resource from the candidate learning resource set L t at time t and recommending it to the learner. The action set A t represents the k action sets corresponding to the learning resource path of length k recommended to the learner at time t. The state transition probability P(·|s t , A t ) corresponds to the probability of transitioning to the next state s t given the state s t and the action set A t+1 and can be regarded as the uniform distribution of user actions The recommendation strategy corresponds to the selection of the action set at time t, A t ~π(s t , L t ), indicating the probability of obtaining the action set A t by selecting learning resources. For a user in state s t , the probability of selecting from the candidate learning resource set L t is denoted as r(s t , a t ). The design of the reward function takes into account both the sequence-level features of the overall path obtained from path planning and the knowledge-level features of individual learning resources. The reward function is decomposed into a sequence decision reward r seq and a knowledge point planning decision reward r c , that is, r = r seq +r c , and the discount factor is denoted as γ.
[0066] 202) Calculate the learning resource feature vector: According to the knowledge forest KG corresponding to the course, use the TransE model to calculate the feature vector v of each learning resource in the learning resource set t . The objective function used is designed as follows:
[0067] min∑ (h,r,t)∈KG ∑ (h′,r′,t′)∈KG , [dis + distance(h + r, t), -distance(h′ + r′, t′)] + (7)
[0068] In Equation (7), h represents the vector of the head entity in the knowledge graph KG, t represents the vector of the tail entity in the knowledge graph KG, r represents the vector of the relationship in the knowledge graph KG, (h, r, t) represents the correct triple in the knowledge graph KG, (h′, r′, t′) represents the wrong triple, dis represents the distance between the positive sample and the negative sample, which is a constant, and [x] + represents taking max(0, x), and the distance calculation method uses the Euclidean distance.
[0069] The obtained learning resource feature vector will be used for the calculation of the user state representation vector s t and the calculation of the reward function r(s t , a t ).
[0070] 203) Calculate the user state representation s t and the action representation a t : All click histories before the user's t-th click are denoted as s t , s t : = h(F 1:t-1 : = [f 1 ,..., f t-1 ), where f t represents the feature vector of each clicked learning resource, and the function h(·) is used to calculate the embedding representation of the sequence F 1:t-1 containing (t - 1) historical click features, and the calculation of this sequence embedding representation is implemented using an LSTM network.
[0071] 204) Implement the sequence decision reward function: Refer to the method of measuring sequence similarity by BLEU in machine learning, and calculate the sequence-level accuracy as the sequence decision reward. The specific formula is as follows:
[0072]
[0073] In Equation (8), prec m represents the sequence decision accuracy, i t:t+k is the actual interaction subsequence, is the recommended subsequence, p m is a subsequence of length m of the subsequence i t:t+k , and M represents the number of subsequences of length m used; it can be seen that the calculation method of the reward function makes the generated recommendation sequence converge in the direction of containing more consistent subsequences, that is, when generating the sequence, not only the performance of each step is considered, but also whether the overall performance of the sequence is optimal.
[0074] 205) Implement the knowledge point planning decision reward function: not only pay attention to whether the specific learning resources recommended match the learning resources actually clicked by the user, but also consider whether the difficulty, learning duration, and resource type of the knowledge points contained in the learning resources are consistent; therefore, calculate the knowledge point planning decision reward function r c When calculating, considering whether the difficulty of the recommended learning resources matches the difficulty of the actually clicked learning resources, use the learning duration to estimate the difficulty of the learning resources. Among them, first use the above-mentioned feature vector v of the learning resources calculated on the course knowledge forest KG according to the TransE model t Then, combined with the learning resource difficulty weight calculated using attributes such as learning duration, obtain the learning resource representation after difficulty weighting; the knowledge point planning decision reward function is realized by calculating the cosine similarity of the vector representations of the actual and predicted learning sequences, and the calculation formula is as follows:
[0075]
[0076] In formula (9), the actual learning subsequence of the user is i t:t+k , and the predicted user learning sequence is c t:t+k is the representation vector used to represent the actual learning sequence of the user, is the representation vector used to represent the predicted user learning sequence, and the sequence representation vector c t:t+k is obtained by taking the mean of the feature vectors of each learning resource in the sequence, and c t+i and are respectively used to represent the feature vectors of the i-th learning resource in the actual and predicted user learning resource learning sequences; v t+i represents the feature vector of the i-th learning resource, dur total represents the default learning duration of this learning resource, and dur watch represents the learning duration of the user on this learning resource.
[0077] 206) For each Cluster ui , use the user behavior model ui trained with this learning behavior type dataset D as the simulation environment of reinforcement learning in Planer ui . For the user user i , the user behavior model of this Cluster ui is used to utilize the similarity of user types to simulate and explore the recommendation strategy suitable for user i : For each Cluster ui construct a user behavior model as the simulation environment of reinforcement learning, and for the user user i, this Cluster ui 's user behavior model is used to utilize the similarity of user types to simulate and explore a recommendation strategy suitable for the user i . It can simulate the sequential decision-making of learners on learning resources during the course learning process, and give the state and action of the learner at a certain moment t (s t , a t ), where the state s t at a certain moment t corresponds to the historical learning resource sequence s t of the learner before moment t: = h(F 1:t-1 : = [f 1 ,..., f t-1 ), and the action a t at a certain moment t represents learning a certain learning resource.
[0078] Use the generative adversarial learning formula to simulate the behavioral dynamics of learners. While taking into account both the learning resources a t clicked by the user (i.e., the action of the user) and the historical click sequence s t of the user (i.e., the state of the user), maximize the reward function r(s t , a t ). Imitate the process that when the user selects among the recommended learning resource paths of length k, the user always learns the learning resource that maximizes their own benefits. Considering that different users' evaluations of learning resources may vary according to personal experiences, the model believes that the reward here is not only related to the user's current choice but also related to the user's learning history; the learned strategy can be regarded as a probability distribution on the action set A t = {a 1 , a 2 ,..., a n}. When calculating the reward function, both the current action a t and the state s t of the user are used as inputs to the reward function r(s t , a t ). The strategy adopted by the user will maximize the expected reward r(s t , a t ). Therefore, when solving, it can be regarded as an optimization problem of the strategy on the probability simplex Δ k-1 . The specific formula is as follows, where the regularization term uses negative Shannon entropy, and the parameter η is used to control the strength of regularization.
[0079]
[0080] 207) In the form of generative adversarial training, utilize the training dataset D corresponding to the learning behaviorui , regard the user behavior model corresponding to the learning behavior type as the generator, and the reward function r ui as the discriminator, complete the training and learning of model parameters, and obtain N Planers ui for simulation. According to the obtained user behavior model and the obtained reward function is r(s t , a t ), the user behavior model is used to simulate the user's true behavior sequence that can maximize the reward function r(s t , a t ). The user takes actions to maximize the reward function r(s t , a t ). Similar to the idea of the generative adversarial network, the training and learning process of the model can be analogized to the generative adversarial network GAN, so that is used as the generator to generate the user's next action based on the user's history, and r is used to distinguish the user's true action and the action a generated by the user model t . Using the minimax function, according to the T user action sequences in the historical behavior and the corresponding characteristics of the clicked course resources to calculate the state and jointly learn the user behavior model and the reward function r, as shown in the following formula. In formula (11), α represents all the parameters used in the model , and θ represents all the parameters used in the reward function r.
[0081]
[0082] Step 3: Complete path planning based on the cascaded DQN algorithm
[0083] For each learner user i , if the learner user i has a learning history, use its learning history to calculate the learning behavior type to which the learner belongs and thus call the path planning model of the corresponding learning type. Using the cascaded DQN algorithm, complete the learning resource path planning for it: For the target knowledge point k target , the learning resource set is . According to the action decision strategy that each step of learning resource recommendation should maximize the current Q function value, use the cascaded method to find the optimal action that maximizes each level of Q function, and iterate step by step until the learning resource containing the target knowledge point is found, and output the planned path; if the learner user iIf there is no learning record, based on the idea of behavior cloning, the learning resource path planning is completed based on similar users in the same major / grade of the same course in history.
[0084] 301) Implement the cascaded DQN algorithm: The implementation framework of the cascaded DQN algorithm is as Figure 3 shown. Use the Q function to find the optimal action at each step in the search space, and the learned optimal action-value function Q * (s t , A t ) satisfies the condition a t ∈A t ; After learning the action-value function Q * (s t , A t ), the recommendation policy function π * (s t , L t ) can be obtained through , where represents the set of learning resource candidates for recommendation at time t. Use the cascaded Q function network to solve the optimal action policy for each step on the path, and the calculation method is as follows.
[0085]
[0086] 302) Recommend the learning resource path according to the policy function learned by the DQN algorithm: For the target knowledge point k target , the learning resource set is Use the algorithm in Table 1. According to the Q function, find the learning resources recommended by each level of the Q function, and iterate step by step until the learning resources containing the target knowledge point are found, and the learning resource path is obtained:
[0087] Table 1 Algorithm for generating recommended learning resource paths using cascaded Q functions
[0088]
[0089] 303) The specific operation of completing the learning resource path planning for users without learning history based on the idea of behavior cloning and similar users in the same major / grade of the same course in history in step 3) is as follows: Given the target knowledge point k target , according to the learner's grade, school, and major information, perform similarity matching among users with existing learning history, find users in the same major / grade of the same course in the historical records, and use the learning history of these similar users to generate a path to the target knowledge point for users without learning history.
[0090] Embodiment
[0091] The method proposed in the present invention was experimented on the online learning log data of the data structure and algorithm course on the Douge practical teaching platform. This dataset contains 61,506 interaction records of 18,093 users. The experimental comparison was made between the method proposed in the present invention and the classical sequential recommendation methods including GRU4Rec, SHAN, NARM, STAMP, and SASRec. The evaluation metrics used were MRR@10 and NDCG@10. As shown in Table 2, it can be seen that the method proposed in the present invention can achieve the optimal recommendation results.
[0092] Table 2 Evaluation Metrics of the Embodiment
[0093]
[0094]
[0095] The above content is only to illustrate the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. Any changes made on the basis of the technical solution according to the technical idea proposed in the present invention shall fall within the protection scope of the claims of the present invention.
Claims
1. A reinforcement learning path planning method based on a generative adversarial user model, characterized in that, it includes the following steps: 1) Obtain and construct a learner similarity matrix based on the user's learning log , and use the spectral clustering method to complete the clustering of user learning behavior types on the learner similarity matrix , obtaining types of user learning behavior types . Based on the user learning behavior types, the training data set corresponding to each learning behavior type can be divided ; 2) Combine a knowledge forest to construct a path planning model based on hierarchical reward function reinforcement learning. The reward function in the path planning model based on hierarchical reward function reinforcement learning is a two-level reward function composed of sequence decision reward and knowledge point planning reward. Use the user behavior model as the environment of reinforcement learning, and train the path planning model in the form of generative adversarial training; 3) Use the user learning behavior type, user historical learning sequence, target knowledge point, learning resource set and course knowledge forest as inputs, complete the learning resource path planning to the target knowledge point based on the cascaded DQN algorithm, and output the planned path; In step 2), in the path planning model constructed based on reinforcement learning with a hierarchical reward function, the five-tuple of the Markov decision process corresponding to reinforcement learning ; Among them, the learner is regarded as the environment, and the state represents the historical learning resource sequence of the learner before the moment, and the action represents selecting a learning resource from the set of candidate learning resources at the moment and recommending it to the learner. The action set then represents the action set corresponding to the learning resource path of length recommended to the learner at the moment ; The state transition probability corresponds to the probability of transferring to the next state given the state and the action set , and at the same time serves as the uniform distribution of user actions , the reward function and the discount factor ; Decompose the reward function into a sequential decision reward and a knowledge point planning decision reward , that is ; When calculating the reward for sequential decision-making calculate the sequence-level accuracy of the recommended subsequence and the actual interaction subsequence, as shown in Equation (8): (8) In formula (8), represents the sequence decision accuracy,[[]]END]] is the actual interaction subsequence,[[]]END]] is the recommended subsequence,[[]]END]] is the subsequence[[]]END]] with a length of[[]]END]] for the subsequence,[[]]END]] represents the number of subsequences with a length of[[]]END]] used;[[]]END]] When calculating the reward function for knowledge point planning and decision-making Considering whether the recommended learning resources match the difficulty of the actually clicked learning resources, the learning duration is used to estimate the difficulty of the learning resources, as follows: (9) In formula (9), the actual learning subsequence of the user is , and the predicted user learning sequence is , is the representation vector used to represent the actual learning sequence of the user, is the representation vector used to represent the predicted user learning sequence, and the sequence representation vector is calculated by taking the mean of the feature vectors of each learning resource in the sequence. and are respectively used to represent the feature vectors of the th learning resource in the actual and predicted user learning resource learning sequences; represents the feature vector of the th learning resource, represents the default learning duration of this learning resource, and represents the learning duration of the user on this learning resource.
2. The reinforcement learning path planning method based on the generative adversarial user model according to claim 1, characterized in that, Construct the learner similarity matrix in step 1) The specific operation is as follows: Obtain the course learning status of each learner , the average time-consuming ratio of completed knowledge points , the average centrality of completed knowledge points , the number of completed key knowledge points and the learning status of target knowledge points , and construct a learner scoring vector : (3) Calculate the cosine similarity between the normalized learner rating vectors and construct a learner similarity matrix : (4)。 3. The reinforcement learning path planning method based on the generative adversarial user model according to claim 2, characterized in that, In step 1), the spectral clustering method is used in combination with the similarity matrix to complete the specific processes of clustering the types of user learning behaviors and partitioning the data set are as follows: Construct the degree matrix and the Laplacian matrix respectively and the Laplacian matrix : (5) (6) Utilize to standardize L, and then calculate the eigenvectors of the first minimum eigenvalues. Combine eigenvectors to form a - dimensional matrix, and standardize it row - by - row to obtain matrix . For each row in matrix as a - dimensional sample, there are a total of samples. Use k - means for clustering to obtain the final N - class classification result, and divide the learners into N different learning behavior types . Based on this, divide the user logs to obtain the training data set corresponding to each learning behavior type .
4. The reinforcement learning path planning method based on the generative adversarial user model according to claim 1, characterized in that, The method for constructing the user behavior model and joint training in step 2) is: For each , design a user behavior model , and the learned strategy is used as a probability distribution over the action set . When calculating the reward function, the current action and the user's state are both used as inputs to the reward function . The strategy adopted by the user will maximize the expected reward . When solving, it is regarded as an optimization problem of the strategy in its probability distribution space and solved; On the dataset corresponding to each type of learning behavior in the form of generative adversarial training, the user behavior model is regarded as the generator, and the reward function is regarded as the discriminator to complete the parameter learning of the path planning model and obtain a for simulation.
5. The reinforcement learning path planning method based on the generative adversarial user model according to claim 4, characterized in that, The solution to the optimization problem is as follows: (10) Among them, the regularization term adopts the negative Shannon entropy, and the parameter is used to control the strength of regularization.
6. The reinforcement learning path planning method based on the generative adversarial user model according to claim 4, characterized in that, The calculation method of the minimax function when implementing generative adversarial training is: According to the user action sequences in the historical behavior and the corresponding characteristics of the clicked course resources calculate the state , jointly learn the user behavior model and the reward function , see Equation (11): (11) In formula (11), represents all the parameters used in the model , and represents all the parameters used in the medium reward function .
7. The reinforcement learning path planning method based on the generative adversarial user model according to claim 1, characterized in that, The specific method of generating a recommended learning resource path using the cascaded DQN algorithm in step 3) is as follows: For the target knowledge point , the learning resource set is . According to the action decision strategy that maximizes the current function value for each learning resource recommendation, using a cascaded method, find the optimal action that maximizes the function at each level, and iterate step by step until a learning resource containing the target knowledge point is found, and then output the planned path.
8. The reinforcement learning path planning method based on the generative adversarial user model according to claim 1, characterized in that, Step 3) further includes: if the learner has no learning record, then based on the idea of behavior cloning, complete the learning resource path planning based on similar users in the same major or grade under the same course, specifically: Known target knowledge points , according to the learner's grade, school and major information, perform similarity matching among users with existing learning histories to find users with the same major or the same grade in the same course in the historical records, and use the learning histories of the similar users to generate a path to the target knowledge points for users without learning histories.
Citation Information
Patent Citations
Adaptive learning path planning system based on reinforcement learning
CN110569443A
Reinforcement learning method and system in adaptive learning path recommendation
CN113434563A