Learning Recommendation Method, Device and Medium Based on Artificial Intelligence and Knowledge Tracking

By combining artificial intelligence and knowledge tracking in smart learning technology, building a reinforcement learning framework and using large language models, the problem of insufficient perception and decision-making capabilities in the existing technology is solved, and more accurate learning recommendation results are achieved.

CN119293223BActive Publication Date: 2025-05-27GUANGZHOU QIANJING TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411200678.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-29
Publication Date
2025-05-27
Estimated Expiration
2044-08-29

AI Technical Summary

Technical Problem

In existing smart learning technologies, deep knowledge tracking lacks decision-making ability, while reinforcement learning lacks the ability to perceive learning state, resulting in the inability to effectively improve the accuracy of learning recommendations.

Method used

A learning recommendation method based on artificial intelligence and knowledge tracking is proposed. By constructing a set of knowledge of to-choice exercises and a reinforcement learning framework, combining large language models and pre-trained difficulty matching knowledge tracking models, we generate recommended exercises and update the recommended exercises set.

Benefits of technology

Effectively combine the characteristics of artificial intelligence feedback and knowledge tracking to improve the accuracy of learning recommendations and realize personalized learning path recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119293223B_ABST
    Figure CN119293223B_ABST
Patent Text Reader

Abstract

The present application discloses a learning recommendation method, device and medium based on artificial intelligence and knowledge tracing, which can be applied to the field of intelligent learning technology. After constructing a set of candidate exercise knowledge based on exercise knowledge documents and constructing a reinforcement learning framework, the present application performs novelty retrieval in the set of candidate exercise knowledge through the reinforcement learning framework to generate a prompt word template, inputs the prompt word template into a large language model to generate recommended exercises, the knowledge points of the recommended exercises and the exercise difficulty, and predicts the first knowledge mastery level of the student for the knowledge points. Then, the learning state of the current student, the response information corresponding to the recommended exercises, the knowledge points of the recommended exercises and the exercise difficulty are input into a pre-trained difficulty matching knowledge tracing model to obtain the second knowledge mastery level of the current student, and the recommended exercise set is updated in combination with the first knowledge mastery level, so as to effectively combine the characteristics of artificial intelligence feedback and knowledge tracing for learning recommendation, thereby improving the accuracy of learning recommendation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent learning technologies, and particularly to a learning recommendation method, device, and medium based on artificial intelligence and knowledge tracing. Background Art

[0002] In related technologies, deep knowledge tracing has strong perception ability to perceive the current learning state of students, but lacks certain decision-making ability; while reinforcement learning has decision-making ability, but lacks the perception ability of learning state. By combining the perception ability of deep knowledge tracing and the decision-making ability of reinforcement learning, deep knowledge tracing perceives the learning state of students, and reinforcement learning makes decisions based on the perceived learning state with the goal of obtaining the best learning efficiency and effect, and can obtain the optimal performance of learning path recommendation effect. Currently, the technological development in the field of reinforcement learning shows a trend: reinforcement learning based on artificial intelligence feedback and reinforcement learning based on knowledge tracing feedback. These two intelligent learning methods have developed independently and relatively maturely. It can be seen that although artificial intelligence feedback and knowledge tracing each show significant advantages in improving model performance, however, there are relatively few reinforcement learning technologies that deeply integrate artificial intelligence and knowledge tracing, resulting in the inability of existing methods to further improve the accuracy of learning recommendation by utilizing the characteristics of artificial intelligence and knowledge tracing.

[0003] In summary, the technical problems existing in related technologies need to be improved. Summary of the Invention

[0004] The main purpose of the embodiments of this application is to propose a learning recommendation method, device, and medium based on artificial intelligence and knowledge tracing, which can effectively improve the accuracy of learning recommendation.

[0005] To achieve the above object, on the one hand, an embodiment of this application proposes a learning recommendation method based on artificial intelligence and knowledge tracing, and the method includes the following steps:

[0006] Construct a set of candidate exercise knowledge based on the exercise knowledge document;

[0007] Construct a reinforcement learning framework;

[0008] Perform novelty search in the set of candidate exercise knowledge through the reinforcement learning framework to generate a prompt template;

[0009] Input the prompt template into a large language model to generate recommended exercises, the knowledge points of the recommended exercises, and the exercise difficulty, and predict the first knowledge mastery level of the student for the knowledge points;

[0010] Input the learning status of the current student, the response information corresponding to the recommended exercise, the knowledge points of the recommended exercise, and the exercise difficulty into a pre-trained difficulty matching knowledge tracing model to obtain the second knowledge mastery level of the current student;

[0011] Update the recommended exercise set according to the first knowledge mastery level and the second knowledge mastery level.

[0012] In some embodiments, the constructing the set of candidate exercise knowledge from the exercise knowledge document includes:

[0013] Collect the exercise knowledge document;

[0014] Segment the exercise knowledge document into multiple text blocks of a preset size;

[0015] Convert the multiple text blocks into exercise vectors through an embedding model;

[0016] Store the exercise vectors in a vector database as the set of candidate exercise knowledge.

[0017] In some embodiments, the reinforcement learning framework includes an evaluation network and a policy network; the constructing the reinforcement learning framework includes:

[0018] Initialize the first learning rate and the first network parameters of the evaluation network, and initialize the second learning rate and the second network parameters of the policy network;

[0019] Loop through the following steps until the number of loops reaches a preset number:

[0020] Interact the agent with the environment and collect a first sequence, the first sequence including states, actions, and rewards;

[0021] Calculate a first loss function of the evaluation network and a second loss function of the policy network according to the first sequence;

[0022] Update the first network parameters according to the first loss function, and update the second network parameters according to the second loss function.

[0023] In some embodiments, the performing novelty search and retrieval in the set of candidate exercise knowledge through the reinforcement learning framework to generate a prompt template includes:

[0024] Obtain an exercise extraction request of the agent in the reinforcement learning framework;

[0025] Convert the exercise extraction request into a query vector;

[0026] Search in the vector database for knowledge texts or historical conversation record information similar to the query vector;

[0027] Construct the prompt template by extracting the knowledge text or historical conversation record information corresponding to the exercise extraction request of the current step and the exercise extraction request of the previous step.

[0028] In some embodiments, inputting the prompt template into a large language model to generate recommended exercises, the knowledge points of the recommended exercises, and the difficulty level of the exercises, and predicting the first knowledge mastery level of the student for the knowledge points includes:

[0029] Input the prompt template into the large language model for prompt parsing to obtain the recommended exercises;

[0030] Analyze the keywords, phrases, and sentence structures in the recommended exercises;

[0031] Determine the first knowledge point according to the keywords, the phrases, and the sentence structures;

[0032] Match the first knowledge point with the knowledge points in the set of candidate exercise knowledge to determine the knowledge points corresponding to the recommended exercises;

[0033] Obtain the historical answer record information of each exercise in the set of candidate exercise knowledge;

[0034] Determine the difficulty level of the recommended exercises according to the historical answer record information;

[0035] Input the exercise data of the current time step and the historical answer record information of the exercise corresponding to the current time step into the large language model, and predict the learning effect of the next time step as the first knowledge mastery level of the student for the knowledge points.

[0036] In some embodiments, inputting the learning state of the current student, the answer information corresponding to the recommended exercises, the knowledge points of the recommended exercises, and the difficulty level of the exercises into a pre-trained difficulty matching knowledge tracking model to obtain the second knowledge mastery level of the current student includes:

[0037] Perform multi-layer perception on the learning state of the current student, the answer information corresponding to the recommended exercises, the knowledge points of the recommended exercises, and the difficulty level of the exercises to obtain the first difficulty-enhanced problem embedding at the current time step;

[0038] Input the first difficulty-enhanced problem embedding into an adaptive sequence neural network to obtain the first knowledge state at the current time step;

[0039] Obtain the second difficulty-enhanced problem embedding at the next time step;

[0040] Take the inner product of the first knowledge state and the second difficulty-enhanced problem embedding to obtain the second knowledge mastery level of the current student.

[0041] In some embodiments, updating the recommended exercise set according to the first knowledge mastery level and the second knowledge mastery level includes:

[0042] Performing a weighted average on the first knowledge mastery level and the second knowledge mastery level to obtain a predicted value of the comprehensive level of the current student;

[0043] Obtaining a knowledge target status value;

[0044] Calculating the recommended probability of the exercises in the set of candidate exercise knowledge according to the predicted value of the comprehensive level and the knowledge target status value;

[0045] Updating the exercises in the recommended exercise set according to the recommended probability and the recommended threshold.

[0046] To achieve the above object, on the other hand, an embodiment of the present application proposes a learning recommendation device based on artificial intelligence and knowledge tracing, and the device includes:

[0047] A first module for constructing a set of candidate exercise knowledge according to an exercise knowledge document;

[0048] A second module for constructing a reinforcement learning framework;

[0049] A third module for performing novelty search and retrieval in the set of candidate exercise knowledge through the reinforcement learning framework to generate a prompt word template;

[0050] A fourth module for inputting the prompt word template into a large language model to generate recommended exercises, the knowledge points of the recommended exercises, and the exercise difficulty, and predicting the first knowledge mastery level of the student for the knowledge points;

[0051] A fifth module for inputting the learning status of the current student, the response information corresponding to the recommended exercises, the knowledge points of the recommended exercises, and the exercise difficulty into a pre-trained difficulty matching knowledge tracing model to obtain the second knowledge mastery level of the current student;

[0052] A sixth module for updating the recommended exercise set according to the first knowledge mastery level and the second knowledge mastery level.

[0053] To achieve the above object, on the other hand, an embodiment of the present application proposes a computer device, including:

[0054] At least one processor;

[0055] At least one memory for storing at least one program;

[0056] When the at least one program is executed by the at least one processor, the at least one processor implements the above method.

[0057] To achieve the above object, on the other hand, an embodiment of the present application proposes a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above method.

[0058] The embodiments of the present application at least include the following beneficial effects: The present application provides a learning recommendation method, device, and medium based on artificial intelligence and knowledge tracing. After constructing a set of candidate exercise knowledge based on exercise knowledge documents and constructing a reinforcement learning framework, the reinforcement learning framework performs novelty retrieval in the set of candidate exercise knowledge to generate a prompt word template, and the prompt word template is input into a large language model to generate recommended exercises, knowledge points of the recommended exercises, and exercise difficulty, and predict the first knowledge mastery level of the student for the knowledge points. Then, the learning state of the current student, the response information corresponding to the recommended exercise, the knowledge points of the recommended exercise, and the exercise difficulty are input into a pre-trained difficulty matching knowledge tracing model to obtain the second knowledge mastery level of the current student. Then, the recommended exercise set is updated according to the first knowledge mastery level and the second knowledge mastery level, so that the characteristics of artificial intelligence feedback and knowledge tracing can be effectively combined for learning recommendation, and thus the accuracy of learning recommendation can be effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 is a flowchart of the learning recommendation method based on artificial intelligence and knowledge tracing provided by an embodiment of the present application;

[0060] Figure 2 is a flowchart of RAG provided by an embodiment of the present application;

[0061] Figure 3 is a schematic diagram of the CBOW model provided by an embodiment of the present application;

[0062] Figure 4 is a schematic diagram of the architecture of the reinforcement learning framework provided by an embodiment of the present application;

[0063] Figure 5 is a schematic diagram of data processing of the DIMKT model provided by an embodiment of the present application;

[0064] Figure 6 is a complete flowchart of the learning recommendation method based on artificial intelligence and knowledge tracing provided by an embodiment of the present application;

[0065] Figure 7 is a schematic diagram of the structure of the learning recommendation device based on artificial intelligence and knowledge tracing provided by an embodiment of the present application;

[0066] Figure 8 is a schematic diagram of the hardware structure of the computer device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0067] In order to make the objectives, technical solutions, and advantages of this application clearer and more understandable, the following further details this application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application. When the following description involves the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of this application. They are only examples of devices and methods that are consistent with some aspects of the embodiments of this application.

[0068] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0069] Before elaborating on the embodiments of this application in detail, first explain some nouns and terms involved in the embodiments of this application. The nouns and terms involved in the embodiments of this application are applicable to the following explanations:

[0070] Reinforcement learning is a machine learning technique that learns how to make optimal decisions by interacting with the environment. In the field of intelligent education, reinforcement learning is used to explore the interaction relationship between students and learning resources and learn how to generate personalized learning recommendations for students.

[0071] Retrieval-Augmented Generation (RAG) refers to optimizing the output of a large language model so that it can reference an authoritative knowledge base outside the training data source before generating a response.

[0072] Large language models (LLMs) are trained with massive amounts of data and use billions of parameters to generate raw outputs for tasks such as answering questions, translating languages, and completing sentences. Based on the powerful capabilities of LLMs, RAG extends them to be able to access the internal knowledge bases of specific domains or organizations.

[0073] The Actor-Critic method in reinforcement learning is an algorithm that combines policy learning and value function learning. Among them, the Actor uses the policy function, is responsible for generating actions and interacting with the environment, while the Critic uses the value function, is responsible for evaluating the performance of the Actor, and guides the actions of the Actor in the next stage. In a personalized learning recommendation system, the Actor is responsible for generating recommendation actions based on the student's state and interacting with the environment, while the Critic evaluates the effects of these recommendation actions, guides the Actor to optimize the policy through feedback, and jointly promotes the personalization and accuracy of the recommendation system.

[0074] In an embodiment of the present application, a learning recommendation method, device, and medium based on artificial intelligence and knowledge tracing are provided. In the present application, after constructing a set of candidate exercise knowledge and a reinforcement learning framework according to exercise knowledge documents, a novelty search and retrieval are performed in the set of candidate exercise knowledge through the reinforcement learning framework to generate a prompt word template. The prompt word template is input into a large language model to generate recommended exercises, the knowledge points of the recommended exercises, and the difficulty of the exercises, and the first knowledge mastery level of the student for the knowledge points is predicted. Then, the learning state of the current student, the response information corresponding to the recommended exercises, the knowledge points of the recommended exercises, and the difficulty of the exercises are input into a pre-trained difficulty matching knowledge tracing model to obtain the second knowledge mastery level of the current student. The recommended exercise set is updated according to the first knowledge mastery level and the second knowledge mastery level, so that the characteristics of artificial intelligence feedback and knowledge tracing can be effectively combined for learning recommendation, and thus the accuracy of learning recommendation can be effectively improved.

[0075] The learning recommendation method based on artificial intelligence and knowledge tracing provided by the embodiments of the present application relates to the field of intelligent learning technology. The learning recommendation method based on artificial intelligence and knowledge tracing provided by the embodiments of the present application can be applied to a terminal, can also be applied to a server, or can be software running on a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto; the server side can be configured as an independent physical server, can also be configured as a server cluster or a distributed system composed of multiple physical servers, or can be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application implementing the learning recommendation method based on artificial intelligence and knowledge tracing, etc., but is not limited to the above forms.

[0076] This application can be used in numerous general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0077] Figure 1 is an optional flowchart of the learning recommendation method based on artificial intelligence and knowledge tracing provided by an embodiment of this application. Figure 1 The method in may include but is not limited to steps S110 to S160:

[0078] Step S110, construct a set of candidate exercise knowledge based on the exercise knowledge document;

[0079] Step S120, construct a reinforcement learning framework;

[0080] Step S130, perform novelty search and retrieval in the set of candidate exercise knowledge through the reinforcement learning framework to generate a prompt word template;

[0081] Step S140, input the prompt word template into a large language model to generate recommended exercises, the knowledge points of the recommended exercises, and the exercise difficulty, and predict the student's first knowledge mastery level of the knowledge points;

[0082] Step S150, input the current student's learning status, the response information corresponding to the recommended exercises, the knowledge points of the recommended exercises, and the exercise difficulty into a pre-trained difficulty matching knowledge tracing model to obtain the current student's second knowledge mastery level;

[0083] Step S160, update the recommended exercise set according to the first knowledge mastery level and the second knowledge mastery level.

[0084] In the embodiment of this application, this embodiment uses retrieval-augmented generation (RAG) technology to optimize the output of the large language model so that it can reference authoritative knowledge bases outside the training data sources before generating responses. Among them, the flowchart of RAG is as Figure 2As shown. It can be understood that when constructing the set of candidate exercise knowledge in this embodiment, after collecting exercise knowledge documents, the exercise knowledge documents can be segmented into multiple text blocks of a preset size to process and retrieve information more efficiently. Then, through an embedding model, the multiple text blocks are converted into exercise vectors, and the exercise vectors are stored in a vector database as the set of candidate exercise knowledge. Specifically, the embedding model in this embodiment can adopt the Word2Vec model. Among them, Word2Vec is a neural network-based word embedding method used to map words into a low-dimensional vector space, and there are two main training methods: CBOW (Continuous Bag of Words) and Skip-gram. This embodiment uses the CBOW model. The CBOW model is suitable for smaller datasets and common vocabulary. Its training input is the word vectors corresponding to the context-related words of a certain feature word, and the output is the word vector of this specific word. The CBOW model predicts the central word through the context words, that is, uses , to predict . It predicts the central word by summing the context word vectors and inputting them into a neural network. The CBOW model diagram is as Figure 3 shown. In this embodiment, by embedding the exercise text blocks into the CBOW model, all the generated vectors are stored in the vector database as the set of candidate exercise knowledge, providing an exercise knowledge base for subsequent recommendations and optimizing the output of the large language model LLM.

[0085] In the embodiment of the present application, the reinforcement learning framework can adopt the actor-critic framework. Among them, the actor-critic algorithm is a type of reinforcement learning method that combines policy gradient and value function estimation, training a policy and a value function simultaneously. The structure diagram is as Figure 4 shown. From Figure 4 , it can be seen that the reinforcement learning framework in this embodiment includes an evaluation network (Critic) and a policy network (Actor). The actor-critic framework selects recommended actions through the actor, and the critic evaluates the quality of the recommended actions and provides feedback to improve the recommendation strategy, ultimately realizing the recommendation of personalized exercise resources for students.

[0086] It can be understood that the process of constructing the reinforcement learning framework in this embodiment includes, but is not limited to, the following steps:

[0087] Initializing the first learning rate of the evaluation network α 1 and the first network parameters θ q , and initializing the second learning rate of the policy networkα 2 and the second network parameter θ π ;

[0088] Then loop through the following steps until the number of loops reaches the preset number:

[0089] Interact the agent with the environment for n steps and collect the first sequence , where the first sequence includes states, actions, and rewards;

[0090] Calculate the first loss function of the evaluation network Critic based on the first sequence and the second loss function of the policy network Actor :

[0091] ;

[0092] ;

[0093] Update the first network parameter according to the first loss function θ q , and update the second network parameter according to the second loss function θ π :

[0094] ;

[0095] ;

[0096] In the embodiment of the present application, the agent in this embodiment completes the interaction with the environment by extracting exercises from the above-mentioned set of candidate exercise knowledge. Each time, one exercise is extracted and then given to the large language model LLM in the form of a prompt template. It can be understood that the exercise extraction request of the agent will be regarded as a user query, and the exercise extraction request is input into the embedding model for vectorization processing to obtain a query vector. Then, search for knowledge texts or historical conversation records that are semantically similar to the query vector in the vector database. Finally, combine the exercise extraction request of the current step with the information retrieved in the previous step (knowledge text or historical conversation record information) to construct a prompt template and input it into the large language model to wait for the model to output an exercise.

[0097] In the embodiment of the present application, this embodiment obtains the knowledge points and difficulties of the exercises by using the large language model and predicts the learning effect to obtain the first knowledge mastery level. Specifically, in time step t, this embodiment inputs the prompt template into the large language model for prompt word parsing to obtain recommended exercises q tAfter that, analyze the keywords, phrases, and sentence structures in the recommended exercises, determine the first knowledge point based on the keywords, phrases, and sentence structures, and then match the first knowledge point with the knowledge points in the knowledge set of the exercises to be selected to determine the knowledge point corresponding to the recommended exercises k m ; At the same time, obtain the historical answer record information of each exercise in the knowledge set of the exercises to be selected, determine the exercise difficulty of the recommended exercises according to the historical answer record information, and then input the exercise data (content q t , knowledge concepts k m , question difficulty QS t , knowledge concept difficulty KC t ) and the historical answer record information of the corresponding exercise at the current time step t into the large language model, and predict the learning effect at the next time step t + 1 as the student's first knowledge mastery level of the knowledge point y t+1 .

[0098] It can be understood that the exercise difficulty includes question difficulty and knowledge concept difficulty.

[0099] In the embodiments of the present application, the current knowledge state of the student is tracked through the difficulty-matched knowledge tracing model, and the learning effect is predicted. Specifically, the difficulty-matched knowledge tracing (DIMKT) model is a knowledge tracing model that measures the question difficulty effect by establishing the relationship between the student's knowledge state and the question difficulty level, uses the inner product of the student's knowledge state and the question embedding to simulate the process of the student applying knowledge to answer questions, and predicts the student's performance. As Figure 5 shown, in this embodiment, the DIMKT model is used to establish the relationship between the student's knowledge state and the question difficulty level, evaluate the student's dynamic knowledge state, and predict the learning effect.

[0100] It can be understood that in this embodiment, the question difficulty and the knowledge concept difficulty are used to enhance the question embedding. Specifically, at time step t, the recommended exercises obtained from the above steps q t , the current learning state of the student, the answer information corresponding to the recommended exercises, the knowledge points of the recommended exercises k m , exercise difficulty QS t , knowledge concept difficulty KC t are input into the multi-layer perceptron (MLP) to obtain the first difficulty-enhanced question embedding at the current time step x t , as shown in the following formula:

[0101] ;

[0102] In the formula, ⊕ is the concatenation operation, is the weight matrix, is the bias term, is the dimension.

[0103] Then, the first difficulty-enhanced problem is embedded into the input adaptive sequence neural network to obtain the first knowledge state at the current time step . Specifically, the adaptive sequence neural network is used to capture the connection between the knowledge state and the problem difficulty in DIMKT, including the subjective difficulty perception module before practice, the personalized knowledge acquisition module during practice, and the knowledge state update module after practice:

[0104] Subjective difficulty perception module before practice: To calculate the students' subjective perception of problem difficulty, consider the difference between the difficulty-enhanced problem embedding and the students' previous knowledge state, that is, if the students' knowledge state cannot meet the requirements of the problem, they will feel difficult. The subjective perception can be obtained as shown in the following formula:

[0105]

[0106] In the formula, is the non-linear activation function, is the sigmoid activation function, is the weight matrix. Here, is the slope, is the direct output of the subjective perception, and further use to select and retain important features, and finally obtain the students' subjective difficulty perception .

[0107] Personalized knowledge acquisition module during practice: Combine the students' subjective difficulty perception and the students' answers ans t to obtain knowledge in a similar way to the subjective difficulty perception module, as shown in the following formula:

[0108]

[0109] In the formula, is the weight matrix, is the slope, is the direct output of the individual knowledge, and further use to select and retain important features, and finally obtain the students' individual knowledge .

[0110] Knowledge state update module after practice: Define the knowledge index , according to the student's previous knowledge state , answer and the difficulty of the exercise , to update the student's knowledge state , as shown in the following formula:

[0111]

[0112] In the formula, is the weight matrix, is the slope.

[0113] Then obtain the second difficulty-enhanced question embedding at the next time step, and take the knowledge state at the time step t and the second difficulty-enhanced question embedding at time step and time step t+ 1 +1 to perform an inner product, simulating the process of the student applying the learned knowledge to answer the exercise. Infer the student's performance at time step t +1 through the sigmoid activation function from the inner product result y t+1 as the second knowledge mastery level, as shown in the following formula:

[0114] .

[0115] In the embodiment of the present application, after obtaining the first knowledge mastery level predicted by the LLM and y 1 , the second knowledge mastery level predicted by DIMKT is y 2 , then the first knowledge mastery level and the second knowledge mastery level are weighted and averaged to obtain the predicted value of the student's comprehensive knowledge mastery level y , as shown in the following formula:

[0116] + ;

[0117] In the formula, w 1 and w 2 are the weights of the predicted values of the LLM and DIMKT respectively, and w 1 + w 2 = 1.

[0118] Set a threshold s t(The value range can be set between 0.5 and 1) is the knowledge point target status value, which is used to define the reward value Reward in the actor-critic framework.

[0119] Calculate the recommendation probability of the exercises in the set of candidate exercise knowledge based on the comprehensive level prediction value and the knowledge target status value. Specifically, determine the reward feedback Reward as: the comprehensive knowledge mastery level of the student after performing a certain recommended action from the current state has reached the knowledge point target status value s t , the reward value Reward is 1, otherwise Reward is 0; the specific formula is as follows:

[0120] ;

[0121] The critic network judges whether the current knowledge point mastery degree reaches the target value and evaluates the behavior score of the actor; the actor network modifies the recommendation probability based on the critic score p .

[0122] Then, set a recommended behavior threshold rec , rec The value range of can be set between 0.5 and 1. If p≥ rec , put the corresponding exercise into the recommended exercise set, and then continue with step S130, where the agent extracts the next exercise from the exercise knowledge base; if p < rec , directly go to step S130.

[0123] In some embodiments, the complete implementation process of the method provided by the embodiments of the present application, as Figure 6 shown, includes but is not limited to the following steps:

[0124] Step S610, construct an exercise knowledge base and embed the model to form a vector database as a set of candidate exercises;

[0125] Step S620, construct a reinforcement learning actor-critic framework, where the agent extracts exercises from the set of candidate exercises and then gives them to the large language model LLM in the form of a prompt template.

[0126] Step S630, the large language model LLM generates a recommended exercise e and outputs the knowledge points and difficulty of exercise e, as well as the predicted value y of the student's mastery level of the current knowledge point 1 .

[0127] Step S640: For a certain student, input the learning status s, the answer to exercise e, the knowledge points and difficulty level of the exercise into the trained DIMKT model, and the output is the predicted value y of the mastery level of the current knowledge point. 2 And update the student's learning status.

[0128] Step S650: The critic network determines whether the mastery level of the current knowledge point reaches the target value and evaluates the behavior score of the actor; based on the critic score, the actor network modifies the recommended behavior probability, puts the exercises corresponding to the recommended behavior probability exceeding the threshold into the recommended exercise set, and then continues to execute step S620.

[0129] In summary, the method of the embodiment of the present application provides an innovative reinforcement learning framework, which can integrate artificial intelligence feedback and knowledge tracing technology to provide highly personalized learning resource recommendations for students.

[0130] Referring to Figure 7 The embodiment of the present application provides a learning recommendation device based on artificial intelligence and knowledge tracing, and the device includes:

[0131] The first module 710 is used to construct a set of candidate exercise knowledge according to the exercise knowledge document;

[0132] The second module 720 is used to construct a reinforcement learning framework;

[0133] The third module 730 is used to perform novelty search and retrieval in the set of candidate exercise knowledge through the reinforcement learning framework to generate a prompt word template;

[0134] The fourth module 740 is used to input the prompt word template into the large language model to generate recommended exercises, the knowledge points of the recommended exercises and the exercise difficulty, and predict the first knowledge mastery level of the student for the knowledge points;

[0135] The fifth module 750 is used to input the learning status of the current student, the answer information corresponding to the recommended exercise, the knowledge points of the recommended exercise and the exercise difficulty into the pre-trained difficulty matching knowledge tracing model to obtain the second knowledge mastery level of the current student;

[0136] The sixth module 760 is used to update the recommended exercise set according to the first knowledge mastery level and the second knowledge mastery level.

[0137] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present application. The functions specifically implemented by the device embodiments of the present application are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0138] An embodiment of the present application also provides a computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above learning recommendation method based on artificial intelligence and knowledge tracking is implemented. The computer device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0139] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present application. The functions specifically implemented by the device embodiments of the present application are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0140] Please refer to Figure 8 , Figure 8 which schematically shows the hardware structure of a computer device in another embodiment. The computer device includes:

[0141] A processor 810, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;

[0142] A memory 820, which can be implemented in forms such as a Read Only Memory (ROM), a static storage device, a dynamic storage device, or a Random Access Memory (RAM). The memory 820 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 820, and the processor 810 is used to call and execute the learning recommendation method based on artificial intelligence and knowledge tracking in the embodiments of the present application;

[0143] An input / output interface 830, which is used to implement information input and output;

[0144] A communication interface 840, which is used to implement communication interaction between the device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0145] A bus 850, which transmits information between various components of the device (such as the processor 810, the memory 820, the input / output interface 830, and the communication interface 840);

[0146] Among them, the processor 810, the memory 820, the input / output interface 830, and the communication interface 840 are communicatively connected to each other inside the device through the bus 850.

[0147] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned learning recommendation method based on artificial intelligence and knowledge tracing.

[0148] It can be understood that the content in the above method embodiments is applicable to the embodiments of this storage medium. The functions specifically implemented by the embodiments of this storage medium are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0149] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely provided with respect to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0150] The embodiments described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0151] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.

[0152] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solutions of this embodiment.

[0153] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0154] In the description of this application and the above-mentioned accompanying drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0155] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist simultaneously. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0156] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the above-mentioned division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0157] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0158] In addition, in each embodiment of the present application, the functional units may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0159] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present application. The foregoing storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0160] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall fall within the scope of the rights of the embodiments of the present application.

Claims

1. A learning recommendation method based on artificial intelligence and knowledge tracing, characterized in that: The method comprises the following steps: Construct a knowledge set of exercises to be selected according to the exercise knowledge document; Build a reinforcement learning framework; Performing a novel search in the candidate exercise knowledge set through the reinforcement learning framework to generate a prompt word template; Input the prompt word template into the large language model, generate recommended exercises, the knowledge points and exercise difficulty of the recommended exercises, and predict the student's first knowledge mastery level of the knowledge points; Perform multi-layer perception on the current student's learning status, the answer information corresponding to the recommended exercises, the knowledge points of the recommended exercises, and the difficulty of the exercises to obtain the first difficulty enhanced question embedding of the current time step; Embedding the first difficulty-enhanced problem into an input adaptive sequence neural network to obtain a first knowledge state at a current time step; Get the second difficulty enhanced question embedding for the next time step; Performing an inner product of the first knowledge state and the second difficulty-enhanced problem embedding to obtain a second knowledge mastery level of the current student; The recommended exercise set is updated according to the first knowledge mastering level and the second knowledge mastering level.

2. The method according to claim 1, characterized in that The step of constructing a set of to-be-selected exercise knowledge according to the exercise knowledge document includes: Collect exercise knowledge documents; Splitting the exercise knowledge document into a plurality of text blocks of preset sizes; Converting the plurality of text blocks into question vectors through an embedding model; The exercise vectors are stored in a vector database as the knowledge set of exercises to be selected.

3. The method according to claim 2, characterized in that The reinforcement learning framework includes an evaluation network and a strategy network; the construction of the reinforcement learning framework includes: Initializing a first learning rate and a first network parameter of the evaluation network, and initializing a second learning rate and a second network parameter of the policy network; The following steps are executed repeatedly until the number of cycles reaches the preset number: The agent interacts with the environment and collects a first sequence, the first sequence including a state, an action, and a reward; Calculate a first loss function of the evaluation network and a second loss function of the policy network according to the first sequence; The first network parameters are updated according to the first loss function, and the second network parameters are updated according to the second loss function.

4. The method according to claim 3, characterized in that The process of performing a novel search in the knowledge set of the selected exercises through the reinforcement learning framework to generate a prompt word template includes: Obtaining an exercise extraction request of an agent in the reinforcement learning framework; Converting the exercise extraction request into a query vector; Searching the vector database for knowledge text or historical conversation record information similar to the query vector; The prompt word template is constructed according to the knowledge text or historical dialogue record information corresponding to the current step exercise extraction request and the previous step exercise extraction request.

5. The method according to claim 1, characterized in that The step of inputting the prompt word template into a large language model, generating recommended exercises, knowledge points and exercise difficulty of the recommended exercises, and predicting the student's first knowledge mastery level of the knowledge points includes: Inputting the prompt word template into a large language model to perform prompt word parsing to obtain the recommended exercises; Analyzing key words, phrases and sentence structures in the recommended exercises; Determine a first knowledge point according to the keyword, the phrase and the sentence structure; Matching the first knowledge point with the knowledge points in the knowledge set of the to-be-selected exercises to determine the knowledge point corresponding to the recommended exercises; Obtaining historical answer record information of each exercise in the knowledge set of exercises to be selected; Determining the difficulty of the recommended exercises according to the historical answer record information; The exercise data of the current time step and the historical answer record information of the exercise corresponding to the current time step are input into the large language model, and the learning effect of the next time step is predicted as the student's first knowledge mastery level of the knowledge point.

6. The method according to claim 1, characterized in that The updating of the recommended exercise set according to the first knowledge mastering level and the second knowledge mastering level includes: Taking a weighted average of the first knowledge mastery level and the second knowledge mastery level to obtain a comprehensive level prediction value of the current student; Get the knowledge target state value; Calculating the recommendation probability of the exercises in the knowledge set of the to-be-selected exercises according to the comprehensive level prediction value and the knowledge target state value; The exercises in the recommended exercise set are updated according to the recommendation probability and the recommendation threshold.

7. A learning recommendation device based on artificial intelligence and knowledge tracking, characterized in that: The device comprises: The first module is used to construct a set of knowledge of exercises to be selected according to the knowledge documents of exercises; The second module is used to build a reinforcement learning framework; The third module is used to perform a novel search in the candidate exercise knowledge set through the reinforcement learning framework to generate a prompt word template; The fourth module is used to input the prompt word template into the large language model, generate recommended exercises, the knowledge points and exercise difficulty of the recommended exercises, and predict the student's first knowledge mastery level of the knowledge points; The fifth module is used to perform multi-layer perception on the current student's learning state, the answer information corresponding to the recommended exercises, the knowledge points of the recommended exercises and the difficulty of the exercises to obtain the first difficulty enhanced problem embedding of the current time step; input the first difficulty enhanced problem embedding into the adaptive sequence neural network to obtain the first knowledge state of the current time step; obtain the second difficulty enhanced problem embedding of the next time step; perform inner product between the first knowledge state and the second difficulty enhanced problem embedding to obtain the second knowledge mastery level of the current student; The sixth module is used to update the recommended exercise set according to the first knowledge mastering level and the second knowledge mastering level.

8. A computer device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Exercise generation method and device based on language model and medium

    CN116561260A

  • Intelligent teaching method and system based on deep knowledge tracking and large language model

    CN118484520A