Learning Path Demonstration Method, Device, Equipment and Medium for Deep Reinforcement Learning

Through the learning path demonstration method of deep reinforcement learning, the reinforcement learning factors of the target user's learning path are determined and reinforcement learning is carried out, which solves the problem of unsatisfactory learning effect demonstration in the existing technology, and achieves a more intuitive and intelligent learning effect demonstration.

CN113094495BActive Publication Date: 2025-07-01SHANGHAI MIYUE ARTIFICIAL INTELLIGENCE INFORMATION TECH CO LTD

Patent Information

Application Number
CN202110431018.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-21
Publication Date
2025-07-01
Estimated Expiration
2041-04-21

AI Technical Summary

Technical Problem

The existing intelligent adaptation learning demonstration system cannot intuitively demonstrate how the intelligent adaptation learning system is like a famous teacher with many years of teaching experience implementing personalized teaching, resulting in unsatisfactory learning effect demonstration.

Method used

Through the learning path demonstration method of deep reinforcement learning, the learning path demonstration instructions of the target user are received, the reinforcement learning factors of the target user's learning path, including state space, action space and learning evaluation indicators, reinforcement learning is carried out, the path generation process of the target user's learning path is obtained, and the path generation process is visually demonstrated.

Benefits of technology

It enriches the demonstration function of the intelligent adaptive learning demonstration system for learning effects, and improves the intuitiveness and intelligence of the intelligent adaptive learning demonstration system for learning learning effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113094495B_ABST
    Figure CN113094495B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses a learning path demonstration method, device, equipment and medium for deep reinforcement learning. The method includes: receiving a learning path demonstration instruction of a target user; determining reinforcement learning factors of the learning path of the target user according to the learning path demonstration instruction; wherein the reinforcement learning factors include an agent, a learning environment, a state space, an action space and a learning evaluation index; in response to the learning path demonstration instruction, performing reinforcement learning on the learning path of the target user according to the reinforcement learning factors to obtain a path generation process of the learning path of the target user; and intuitively demonstrating the path generation process. The technical solution of the embodiment of the present invention can enrich the demonstration function of the learning effect of the intelligent adaptive learning demonstration system, thereby improving the intuitiveness and intelligence of the intelligent adaptive learning demonstration system in demonstrating the learning effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the technical field of artificial intelligence online education, and in particular, to a learning path demonstration method, device, equipment and medium for deep reinforcement learning. Background Art

[0002] An intelligent adaptive learning system can "personally customize" a learning mode and learning courses according to each student's respective learning strengths and weaknesses. When a student user enters the intelligent adaptive learning system, a round of tests is required to detect the weak points of the student user's current level. The intelligent adaptive learning system combines a nanoscale knowledge graph to test the knowledge points consistent with the dynamic learning objectives in the least amount of time, and dynamically establishes a dynamic user profile for each student user based on the mastery status of the knowledge points after the student's learning, understands the learning status and anomaly warnings of each student user, and timely adjusts the learning path and learning content of the student user to obtain the most suitable and personalized learning path and learning content for the student user among numerous learning contents.

[0003] There are many influencing factors for the knowledge point learning strategy of intelligent adaptive learning. Personalized learning materials for student users can be formulated according to learning strategies such as recommendation based on a logical graph, recommendation based on a sequential graph, tracing the origin, and strategic priority. However, the learning effect of each student user is the result of long-term accumulation. Therefore, it is necessary for the intelligent adaptive learning demonstration system to display the maximum learning effect to the student user within a certain period of time to demonstrate the intelligent recommendation characteristics of the intelligent adaptive learning system, so that parents and students can intuitively understand the intelligence of the knowledge points recommended by the intelligent adaptive learning system.

[0004] Currently, the existing intelligent adaptive learning demonstration systems display the maximum learning effect in two ways. The first way is to push a two-dimensional demonstration form through a dynamic display processor to display the maximum learning effect through the pushed demonstration form. This way of demonstrating the learning effect is difficult to intuitively show the complex factors considered in the selection of knowledge points for learning when the artificial intelligence system makes decisions on teaching strategies. The second way is to configure different ratios and influencing factors of each user benefit item into a preset database for storage and then output for demonstration. In this way of demonstrating the learning effect, the knowledge point recommendation algorithm of the intelligent adaptive learning system will dynamically make recommendations according to the learning status of the student user and multiple attributes of the knowledge points in the knowledge graph, such as position, front and rear relationship, difficulty, and examination frequency. However, the method of configuring influencing factors will result in a high maintenance cost and it is difficult to achieve data-driven, and it is also difficult to intuitively show the complex factors considered in the selection of knowledge points for learning when the artificial intelligence system makes decisions on teaching strategies. Thus, it can be seen that the existing intelligent adaptive learning demonstration systems cannot intuitively demonstrate how the intelligent adaptive learning system implements personalized teaching like a famous teacher with many years of teaching experience in a teaching scenario, resulting in an unsatisfactory demonstration of the learning effect. Summary of the Invention

[0005] An embodiment of the present invention provides a learning path demonstration method, device, equipment and medium for deep reinforcement learning, which can enrich the demonstration function of the intelligent adaptive learning demonstration system for learning effects, thereby improving the intuitiveness and intelligence of the intelligent adaptive learning demonstration system in demonstrating learning effects.

[0006] In a first aspect, an embodiment of the present invention provides a learning path demonstration method for deep reinforcement learning, including:

[0007] Receiving a learning path demonstration instruction of a target user;

[0008] Determining the reinforcement learning factors of the target user's learning path according to the learning path demonstration instruction; wherein, the reinforcement learning factors include a state space, an action space, and a learning evaluation index;

[0009] In response to the learning path demonstration instruction, performing reinforcement learning on the target user's learning path according to the reinforcement learning factors to obtain a path generation process of the target user's learning path;

[0010] Intuitively demonstrating the path generation process.

[0011] In a second aspect, an embodiment of the present invention further provides a learning path demonstration device for deep reinforcement learning, including:

[0012] A learning path demonstration instruction receiving module, configured to receive a learning path demonstration instruction of a target user;

[0013] A reinforcement learning factor determining module, configured to determine the reinforcement learning factors of the target user's learning path according to the learning path demonstration instruction; wherein, the reinforcement learning factors include a state space, an action space, and a learning evaluation index;

[0014] A reinforcement learning module, configured to perform reinforcement learning on the target user's learning path according to the reinforcement learning factors in response to the learning path demonstration instruction to obtain a path generation process of the target user's learning path;

[0015] A path generation process demonstration module, configured to intuitively demonstrate the path generation process.

[0016] In a third aspect, an embodiment of the present invention further provides an electronic device, where the electronic device includes:

[0017] One or more processors;

[0018] A storage device, configured to store one or more programs;

[0019] When the one or more programs are executed by the one or more processors, the one or more processors implement the learning path demonstration method of deep reinforcement learning provided in any embodiment of the present invention.

[0020] In a fourth aspect, an embodiment of the present invention further provides a computer storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the learning path demonstration method of deep reinforcement learning provided in any embodiment of the present invention.

[0021] In the embodiment of the present invention, the intelligent adaptive learning demonstration system determines reinforcement learning factors such as the state space, action space, and learning evaluation index of the target user's learning path according to the received learning path demonstration instruction, so as to respond to the learning path demonstration instruction, perform reinforcement learning on the target user's learning path according to the determined reinforcement learning factors, obtain the path generation process of the target user's learning path, and intuitively demonstrate the path generation process, solving the problem that the existing intelligent adaptive learning demonstration system has a poor demonstration effect on the learning effect, enriching the demonstration function of the intelligent adaptive learning demonstration system for the learning effect, and improving the intuitiveness and intelligence of the intelligent adaptive learning demonstration system in demonstrating the learning effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is a flowchart of a learning path demonstration method of deep reinforcement learning provided in Embodiment 1 of the present invention;

[0023] Figure 2 is a flowchart of a learning path demonstration method of deep reinforcement learning provided in Embodiment 2 of the present invention;

[0024] Figure 3 is a schematic flowchart of the reinforcement learning in the prior art;

[0025] Figure 4 is a schematic execution flowchart of the reinforcement learning in the prior art;

[0026] Figure 5 is a schematic structural diagram of each functional module included in an intelligent adaptive learning demonstration system provided in Embodiment 2 of the present invention;

[0027] Figure 6 is a schematic flowchart of the reinforcement learning of an agent provided in Embodiment 2 of the present invention;

[0028] Figure 7 is a schematic diagram of the correlation effect between some knowledge points of the knowledge graph provided in Embodiment 2 of the present invention;

[0029] Figure 8 is a schematic flowchart of the demonstration process of a single-person interaction mode provided in Embodiment 2 of the present invention;

[0030] Figure 9 It is a schematic diagram of a learning path demonstration device for deep reinforcement learning provided in Embodiment 3 of the present invention;

[0031] Figure 10 It is a schematic diagram of the structure of an electronic device provided in Embodiment 4 of the present invention. Detailed implementation manners

[0032] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention.

[0033] In addition, it should be noted that for the sake of convenience of description, only parts related to the present invention are shown in the accompanying drawings rather than all the content. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of the operations can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the operations can be rearranged. When the operations are completed, the process can be terminated, but there can also be additional steps not included in the drawings. The process can correspond to a method, a function, a procedure, a subroutine, a subprogram, etc.

[0034] The terms "first" and "second" etc. in the description, claims and drawings of the embodiments of the present invention are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may include steps or units not listed.

[0035] Embodiment 1

[0036] Figure 1 It is a flowchart of a learning path demonstration method for deep reinforcement learning provided in Embodiment 1 of the present invention. This embodiment is applicable to the situation of intuitively and intelligently demonstrating the learning path to the user. This method can be executed by a learning path demonstration device for deep reinforcement learning. The device can be implemented in software and / or hardware and is generally integrated in an electronic device. The electronic device can be a device capable of running an intelligent adaptive learning demonstration system. Correspondingly, as Figure 1 shown, the method includes the following operations:

[0037] S110. Receive a learning path demonstration instruction from a target user.

[0038] Among them, the target user can be the learning user for whom the intelligent adaptive learning system generates a learning path. The learning path demonstration instruction can be input by the operating user to the intelligent adaptive learning demonstration system, and is an instruction for requesting the intelligent adaptive learning demonstration system to demonstrate the learning path of the target user. The operating user can be the user who operates the intelligent adaptive learning demonstration system, which can be a student user or a teacher user, etc. The embodiments of the present invention do not limit the specific type of the operating user.

[0039] In the embodiments of the present invention, when the operating user needs to preview the learning path of the target user through the intelligent adaptive learning demonstration system, the operating user can input the learning path demonstration instruction of the target user to the intelligent adaptive learning demonstration system. Optionally, the intelligent adaptive learning demonstration system can interact with the intelligent adaptive learning system as an independent system to demonstrate the decision-making process of the intelligent adaptive learning system to the operating user in real time. Or, the intelligent adaptive learning demonstration system can also be integrated inside the intelligent adaptive learning system and used as a subsystem of the intelligent adaptive learning system to directly output the learning path decision-making process of the intelligent adaptive learning system. The embodiments of the present invention do not limit this.

[0040] Exemplarily, when the operating user is a student user, the operating user can also be the target user. Then, the target user can input the learning path demonstration instruction of himself / herself to the intelligent adaptive learning demonstration system to request the intelligent adaptive learning demonstration system to demonstrate the learning path of himself / herself. When the operating user is a teacher user, the operating user can input the learning path demonstration instruction of a certain student user to the intelligent adaptive learning demonstration system to request the intelligent adaptive learning demonstration system to demonstrate the learning path of the student user.

[0041] S120. Determine the reinforcement learning factors of the target user's learning path according to the learning path demonstration instruction; among them, the reinforcement learning factors include an intelligent agent, a learning environment, a state space, an action space, and a learning evaluation index.

[0042] Among them, the target user's learning path is also the learning path of the target user. It can be understood that the learning path can include the learning process and learning content of the target user for each knowledge point. Exemplarily, the learning path of target user A can be: the concept of quadratic radicals - the effective conditions of quadratic radicals - the simplification of quadratic radicals - rationalization of the denominator - multiplication of quadratic radicals - division of quadratic radicals - multiplication and division of quadratic radicals. The learning paths of different target users can be the same or different, and need to be determined specifically according to relevant factors such as the learning ability of the target user. The reinforcement learning factors can be the relevant factors of reinforcement learning, and can include but are not limited to an intelligent agent, a learning environment, a state space, an action space, and a learning evaluation index, etc.

[0043] Correspondingly, after the intelligent adaptive learning demonstration system receives the learning path demonstration instruction of the target user, it can determine the state space, action space, learning evaluation index and other reinforcement learning factors of the target user's learning path according to the learning path demonstration instruction, and realize the initialization configuration of the reinforcement learning for the learning path.

[0044] S130. In response to the learning path demonstration instruction, perform reinforcement learning on the target user's learning path according to the reinforcement learning factors to obtain the path generation process of the target user's learning path.

[0045] Among them, the path generation process can be the planning and generation process of a complete learning path.

[0046] After the intelligent adaptive learning demonstration system realizes the initialization configuration of the reinforcement learning for the learning path, it can start to respond to the learning path demonstration instruction, perform reinforcement learning on the target user's learning path according to the configured reinforcement learning factors in combination with the intelligent adaptive learning system and the reinforcement learning model, and obtain the path generation process of the target user's learning path.

[0047] It should be noted that the path generation process in the embodiments of the present invention can reflect the decision-making process and effect of the intelligent adaptive learning system for dynamically recommending each knowledge point. The entire path generation process can intuitively reflect the change of the complete learning state of the target user and the intelligent decision-making inference method of the intelligent adaptive learning system for the real-time learning state change process.

[0048] Optionally, the reinforcement learning model in the embodiments of the present invention can be a deep reinforcement learning model, and the embodiments of the present invention do not limit the type of the reinforcement learning model. It should be noted that the reinforcement learning model can be integrated inside the intelligent adaptive learning demonstration system to be directly scheduled by the intelligent adaptive learning demonstration system for reinforcement learning to generate the path generation process of the target user's learning path. Or, the reinforcement learning model can also be executed independently of the intelligent adaptive learning demonstration system. The intelligent adaptive learning demonstration system can send an instruction to the system or device where the reinforcement learning model is located to schedule the reinforcement learning model for reinforcement learning to generate the path generation process of the target user's learning path. The embodiments of the present invention do not limit the integration method between the reinforcement learning model and the intelligent adaptive learning demonstration system and the method for the intelligent adaptive learning demonstration system to schedule the reinforcement learning model.

[0049] S140. Intuitively demonstrate the path generation process.

[0050] Correspondingly, after obtaining the path generation process, the intelligent adaptive learning demonstration system can demonstrate the path generation process of the entire target user's learning path in real time and intuitively, that is, intuitively demonstrate the learning effect of the target user, so that the operating user can intuitively understand the personalized learning dynamic process suitable for the target user.

[0051] It should be noted that in the embodiments of the present invention, in addition to being able to specify a target user for the intelligent adaptive learning demonstration system using the learning path demonstration instruction, the operating user can also use the learning path demonstration instruction to indicate to the intelligent adaptive learning demonstration system the path generation process of different types of learning paths. Optionally, the intelligent adaptive learning demonstration system can intuitively demonstrate the path generation process in the form of a knowledge graph or a roadmap, etc. The embodiments of the present invention do not limit the demonstration method of the intelligent adaptive learning demonstration system.

[0052] For example, the operating user can use the learning path demonstration instruction to specify that the intelligent adaptive learning demonstration system uses the automatic mode for demonstration, that is, there is no human-computer interaction throughout the path generation process demonstration, and the intelligent adaptive learning system automatically and intelligently determines the content that the target user needs to learn, so as to show the best learning effect. The operating user can also use the learning path demonstration instruction to specify that the intelligent adaptive learning demonstration system uses the single-person interaction mode for demonstration, that is, the operating user can use the learning path demonstration instruction to specify the knowledge point for the target user to start learning. Specifically, the intelligent adaptive learning system initially determines the knowledge points that the target user needs to learn, and the operating user can choose to accept or not. When the operating user chooses to accept, the intelligent adaptive learning system continues to automatically and intelligently determine the learning path of the target user. To further improve the user experience, the operating user can also use the learning path demonstration instruction to specify that the intelligent adaptive learning demonstration system uses the multi-person interaction mode for demonstration, that is, the operating user can use the learning path demonstration instruction to specify that the operating user selects the next knowledge point for the target user to learn, so as to observe the situation of the knowledge points that can be mastered by the learning path arranged by the operating user for the target user without the intervention of the intelligent adaptive learning system.

[0053] It can be seen that the learning path demonstration method of deep reinforcement learning provided by the intelligent adaptive learning demonstration system in the embodiments of the present invention can not only allow users to intuitively understand the dynamic intelligent decision-making process of the intelligent adaptive learning system for the entire learning path of the target user, improve the intuitiveness and intelligence of the learning effect demonstration of the intelligent adaptive learning demonstration system, but also provide a variety of different types of interactive demonstration methods, further enriching the demonstration function of the intelligent adaptive learning demonstration system for the learning effect.

[0054] In an embodiment of the present invention, an intelligent adaptive learning demonstration system determines reinforcement learning factors such as the state space, action space, and learning evaluation metrics of the learning path of a target user according to a received learning path demonstration instruction. In response to the learning path demonstration instruction, reinforcement learning is performed on the learning path of the target user according to the determined reinforcement learning factors to obtain the path generation process of the learning path of the target user, and the path generation process is visually demonstrated, solving the problem that the existing intelligent adaptive learning demonstration system has a poor demonstration effect on the learning effect, enriching the demonstration function of the intelligent adaptive learning demonstration system for the learning effect, and improving the intuitiveness and intelligence of the intelligent adaptive learning demonstration system in demonstrating the learning effect.

[0055] Embodiment 2

[0056] Figure 2 FIG. is a flowchart of a method for demonstrating a learning path of deep reinforcement learning provided in Embodiment 2 of the present invention. This embodiment is specific based on the above embodiment. In this embodiment, various specific and optional implementation manners are given for determining the reinforcement learning factors of the learning path of the target user according to the learning path demonstration instruction, performing reinforcement learning on the learning path of the target user according to the reinforcement learning factors to obtain the path generation process of the learning path of the target user, and visually demonstrating the path generation process. Correspondingly, as Figure 2 shown, the method of this embodiment may include:

[0057] S210. Receive a learning path demonstration instruction of a target user.

[0058] Reinforcement learning belongs to a machine learning method. Figure 3 FIG. is a schematic flowchart of reinforcement learning in the prior art. As Figure 3 shown, the reinforcement learning algorithm includes five major elements: Agent, Environment, Action, State, and Reward. The Agent interacts with the environment in real time. After observing the state of the environment, the Agent outputs an action according to the policy model, and the action acts on the environment and then affects the state of the environment. In addition, the environment gives the Agent a reward according to the quality of the action and the state, and the Agent updates its policy model for selecting actions according to the action state and the reward. By continuously trying in the environment, the maximum reward is obtained, and the mapping from the state to the action is learned. This mapping is the policy model, or simply the model, which is represented by a parameterized deep neural network.

[0059] Figure 4 FIG. is a schematic execution flowchart of reinforcement learning in the prior art. As Figure 4As shown in the figure, the existing reinforcement learning execution process specifically includes: The Agent observes the Environment and obtains the state, makes an action based on its Policy for the state, at which time a reward can be obtained, and the Environment has changed, so the Agent will obtain a new state and continue to execute until the learning is successful.

[0060] In the embodiments of the present invention, reinforcement learning is applied to the application scenario of learning path deduction. When determining the reinforcement learning factors of the target user's learning path, the five major elements included in the reinforcement learning algorithm are respectively configured to obtain the agent, learning environment (i.e., environment), state space, action space, and learning evaluation index (i.e., reward return function) corresponding to the intelligent adaptive learning demonstration system. The process of generating the path of the target user's learning path by the intelligent adaptive learning demonstration system using the reinforcement learning method is specifically referred to the following operations.

[0061] S220. Determine the knowledge point recommendation environment in the reinforcement learning model as the intelligent adaptive learning system.

[0062] S230. Determine the type of knowledge point recommendation environment matched by the state space according to the type of the learning path demonstration instruction.

[0063] Specifically, the intelligent adaptive learning system can be determined as the knowledge point recommendation environment in the reinforcement learning model. That is, the agent in the reinforcement learning model remains unchanged, the intelligent adaptive learning system is set as the learning environment factor in the reinforcement learning model, and the type of knowledge point recommendation environment matched by the state space of the reinforcement learning is determined according to the type of the learning path demonstration instruction. It can be understood that different types of learning path demonstration instructions can specify different types of knowledge point recommendation environment types, and each knowledge point recommendation environment type can correspond to a reinforcement learning environment.

[0064] In an optional embodiment of the present invention, the type of the knowledge point recommendation environment may include at least one of an adaptive engine recommended knowledge point environment and an operator user feedback recommended knowledge point environment; the adaptive engine recommended knowledge point environment is used to recommend knowledge points to the target user by using an adaptive engine; the operator user feedback recommended knowledge point environment is used to recommend knowledge points to the target user according to the feedback information of the operator user; the operator user includes a first operator user or a second operator user; the feedback information of the first operator user is used to confirm whether to accept the knowledge points recommended by the adaptive engine; the feedback information of the second operator user is used to independently recommend knowledge points to the target user.

[0065] Among them, the adaptive engine is also the intelligent learning engine of the intelligent adaptive learning system. The environment for the adaptive engine to recommend knowledge points can be an environment where the adaptive engine recommends knowledge points to the target user. The environment for the operating user to feedback and recommend knowledge points can be an environment where knowledge points are recommended to the target user according to the feedback information of the operating user. The first operating user can be an operating user who interacts with the intelligent adaptive learning demonstration system in a single-person interaction mode. For example, the target user himself / herself can feedback information to the intelligent adaptive learning demonstration system to confirm whether to accept the knowledge points recommended by the adaptive engine. The second operating user can be an operating user who interacts with the intelligent adaptive learning demonstration system in a multi-person interaction mode. For example, a teacher user can also feedback information to the intelligent adaptive learning demonstration system to independently determine the knowledge points recommended to the target user by using the intelligent adaptive learning demonstration system.

[0066] That is to say, in the embodiments of the present invention, the intelligent adaptive learning demonstration system can simulate the generation methods of three different types of learning paths. By configuring the environment for the adaptive engine to recommend knowledge points, the generation of learning paths in an automatic mode can be achieved. That is, there is no human-computer interaction throughout the process of the path generation demonstration. The intelligent adaptive learning system automatically and intelligently determines the content that the target user needs to learn, so as to show the maximum learning effect. By configuring the environment for the operating user to feedback and recommend knowledge points, the generation of learning paths in a human-computer interaction mode can be achieved, which can implement both the single-person interaction mode and the multi-person interaction mode. In the single-person interaction mode, the first operating user can use the learning path demonstration instruction to specify the knowledge point for the target user to start learning. Specifically, the intelligent adaptive learning system initially determines the knowledge points that the target user needs to learn, and the operating user can choose to accept or not accept. When the operating user chooses to accept, the intelligent adaptive learning system continues to automatically and intelligently determine the learning path of the target user. In the multi-person interaction mode, the second operating user can use the learning path demonstration instruction to specify that the second operating user independently selects the next knowledge point for the target user to learn, so as to observe the situation of the knowledge points that can be mastered by the learning path arranged for the target user by the operating user without the intervention of the intelligent adaptive learning system. The multi-person interaction mode can enable the operating user to make guidance according to his / her own experience for the current learning state of the target user.

[0067] It can be seen that the learning path demonstration method of deep reinforcement learning in the human-computer interaction mode can enable the operating user to participate in the demonstration process, improving the interactivity and scalability of the learning path demonstration. When the teacher user, as the second operating user, makes guidance according to his / her own experience for the current learning state of the target user and determines the learning path of the target user, it can also be compared with the intelligent learning strategy of the intelligent adaptive learning system, so that the teacher user can understand the differences between a real teacher and the intelligent adaptive learning system in planning the learning path for students under the same conditions, thereby highlighting the advantages of the intelligent adaptive learning system.

[0068] S240. Determine the state space according to the student user attributes, the knowledge point attributes of the knowledge graph, the learning situation attributes, and the demonstration order.

[0069] Among them, the student user attributes may include the learning ability of the target user; the knowledge point attributes of the knowledge graph may include the logical relationship between knowledge points, the examination frequency of knowledge points, the importance degree of knowledge points, and the difficulty of knowledge points; the learning situation attributes may include the knowledge point mastery status; the demonstration attributes may include the number of mastered knowledge points, the demonstration duration of the learning path, or the learning scope of knowledge points.

[0070] In an embodiment of the present invention, the state space of reinforcement learning can be determined according to the student user attributes, the knowledge point attributes of the knowledge graph, and the learning situation attributes. Optionally, if it is necessary to set the demonstration method of the learning path, such as setting the demonstration duration or the demonstration learning scope, the demonstration attributes can also be added to the state space.

[0071] Optionally, the student user attributes can be the student ability of the target user, which can come from the set value provided by the intelligent adaptive learning demonstration system to the target user. For example, the student user attributes of the target user can be excellent in learning, medium in learning, poor in learning, etc. The student user attributes can also be determined by the intelligent adaptive learning system saving the historical learning data of the user, that is, multiple ability value intervals divided according to the ability value of the user at each moment obtained by the Item Response Theory (IRT). Such as excellent in learning (ability value in the range of 0.7 - 1), medium in learning (ability value in the range of 0.36 - 0.69), and poor in learning (ability value in the range of 0.35 - 0). The operating user can use the learning path demonstration instruction to set the initial student user attribute level of the target user. When the intelligent adaptive learning system simulates the target user's learning of knowledge points, the ability value of the target user in the intelligent adaptive learning system can be used as the updated student user attribute. The knowledge point attributes of the knowledge graph can be the characteristics possessed in the knowledge point dimension, which can include but are not limited to the logical relationship between knowledge points (such as: there is a pre - post relationship between knowledge points), the examination frequency of knowledge points, the importance degree of knowledge points, and the difficulty of knowledge points. The learning situation attributes can include the knowledge point mastery status of the target user, that is, in the case of the student user attributes of the target user, whether the target user masters the knowledge point. The demonstration attributes can be the conditions provided to the operating user to set the demonstration requirements, including but not limited to the maximum number of mastered knowledge points and the demonstration duration, that is, the duration for the agent to achieve the task goal, or the learning scope to be demonstrated.

[0072] S250. Determine the action space according to the recommended learning knowledge points of the target user.

[0073] Among them, the recommended learning knowledge points are the knowledge points recommended for the target user to learn.

[0074] Specifically, the knowledge point that the target user will learn next can be set as the action output by the intelligent agent.

[0075] S260. In response to the learning path demonstration instruction, perform reinforcement learning on the learning path of the target user according to the reinforcement learning factor.

[0076] After the above reinforcement learning factor is configured, the intelligent adaptive learning demonstration system can, in response to the learning path demonstration instruction, perform reinforcement learning on the learning path of the target user according to the reinforcement learning factor.

[0077] Figure 5 It is a schematic structural diagram of each functional module included in an intelligent adaptive learning demonstration system provided in the second embodiment of the present invention. In a specific example, such as Figure 5As shown in the figure, the functional modules of the intelligent adaptive learning demonstration system based on reinforcement learning can be subdivided into a state module, a decision-making module, an interaction module, a recommendation module, a learning simulation module, and a demonstration module. Among them, the state module can provide the state attributes required by the agent, including student user attributes, knowledge point attributes of the knowledge graph, learning situation attributes, and demonstration attributes, etc. The decision-making module can display the actions output by the agent, that is, the next knowledge point for the student to learn. Specifically, the decision-making module can establish a deep reinforcement learning model for the agent, set the state space of the agent in the environment, the behavior space that the agent can decide, and the behavior rewards of the environment for the agent, and use a deep neural network to approximate the mapping function from state to action. The agent makes behavior decisions by observing the dynamic environment states such as the mastery states of each knowledge point on the knowledge graph and the learning level of the target student, that is, the dynamic programming of the knowledge point recommendation of the agent. The interaction module can provide various interaction modes of the intelligent adaptive learning demonstration system. The operating user can set the demonstration attributes required by the state module through the interaction module and select the interaction modes to be adopted, which can include an automatic mode, a single-person interaction mode, and a multi-person interaction mode, etc. The recommendation module can access the knowledge point recommendation algorithm of the intelligent adaptive learning system and provide the required data through the knowledge point recommendation algorithm interface, such as the exercises associated with the current knowledge point. When the intelligent adaptive learning system can infer the next recommended knowledge point for the target user based on the mastery state of the target user for the current knowledge point, the recommendation module can continue to provide the required data according to the next recommended knowledge point using the knowledge point recommendation algorithm. This process is repeated until the intelligent adaptive learning system completes the generation of the learning path and obtains the complete learning path of the target user. Optionally, when there are multiple question-pushing strategies for the knowledge point recommendation algorithm, the name of the question-pushing strategy can also be demonstrated in real time during the process of demonstrating the learning path of the intelligent adaptive learning demonstration system, so that the operator can intuitively understand the decision-making process of the intelligent adaptive learning system. The learning simulation module can access the question-pushing algorithm of the intelligent adaptive learning system, receive the learning data sent by the question-pushing algorithm interface, such as various knowledge point exercises, etc., and perform simulation learning according to the student user attributes of the target user to obtain the state value of whether the target user can master the knowledge point under the established learning situation state. The learning simulation module can use the question-pushing algorithm to simulate the ability values of multiple student users, determine the possibility of answering questions correctly or incorrectly under multiple question difficulties, and judge whether the target user has mastered the knowledge point according to the mastery conditions of the intelligent adaptive learning system. The demonstration module can demonstrate the learning path of the target user in the intelligent adaptive learning system, usually in the form of a knowledge graph or a roadmap, and can be a computer client, a web page, a television, a mobile terminal, an intelligent terminal, or various demonstration screens. The embodiments of the present invention do not limit this. The content demonstrated by the demonstration module can be designed according to user needs, such as including user attributes, knowledge point attributes, and learning strategies, etc. The embodiments of the present invention also do not limit this.

[0078] Figure 6 It is a schematic flowchart of an agent reinforcement learning provided in the second embodiment of the present invention. In a specific example, as Figure 6 shown, step S260 may specifically include the following operations.

[0079] S261. The agent in the reinforcement learning model observes the knowledge point recommendation environment to obtain a multi-dimensional vector state.

[0080] Among them, the multi-dimensional vector state may be the observation result of the agent's environmental state represented by a high-dimensional vector. Exemplarily, the multi-dimensional vector state may include student user attributes, knowledge point attributes of the knowledge graph, learning situation attributes, demonstration attributes, etc.

[0081] In an optional embodiment of the present invention, if the knowledge point recommendation environment type includes an adaptive engine recommendation knowledge point environment, then the agent observes the knowledge point recommendation environment to obtain a multi-dimensional vector state, which may include: the adaptive engine of the intelligent adaptive learning system determines adaptive recommended knowledge points according to the knowledge point recommendation algorithm; the agent determines the multi-dimensional vector state of the adaptive engine recommendation knowledge point environment according to the adaptive recommended knowledge points.

[0082] Among them, the adaptive recommended knowledge points may be the knowledge points automatically recommended by the adaptive engine according to the knowledge point recommendation algorithm.

[0083] Optionally, if the operating user indicates in the learning path demonstration instruction that the adaptive engine recommendation knowledge point environment is the environment type of reinforcement learning, the adaptive engine of the intelligent adaptive learning system may determine adaptive recommended knowledge points according to the knowledge point recommendation algorithm. Further, the agent may determine the multi-dimensional vector state of the adaptive engine recommendation knowledge point environment according to the adaptive recommended knowledge points.

[0084] In an optional embodiment of the present invention, if the knowledge point recommendation environment type includes an operating user feedback recommendation knowledge point environment, then the agent observes the knowledge point recommendation environment to obtain a multi-dimensional vector state, which may include: the intelligent adaptive learning system receives the feedback recommended knowledge points determined by the operating user; the agent determines the multi-dimensional vector state of the operating user feedback recommendation knowledge point environment according to the feedback recommended knowledge points; among them, if the operating user is the first operating user, the feedback recommended knowledge points are the target adaptive recommended knowledge points selected by the first operating user according to the adaptive recommended knowledge points determined by the adaptive engine; if the operating user is the second operating user, the feedback recommended knowledge points are the self-recommended knowledge points determined by the second operating user (according to teaching experience).

[0085] Among them, the feedback recommended knowledge points can be the knowledge points feedback by the operating user to the intelligent adaptive learning demonstration system. The target adaptive recommended knowledge points can be one of the recommended knowledge points selected by the first operating user according to the adaptive recommended knowledge points determined by the adaptive engine. The independent recommended knowledge points can be the recommended knowledge points selected by the second operating user according to their own experience.

[0086] Optionally, if the operating user indicates in the learning path demonstration instruction that the feedback recommended knowledge points environment is used as the environment type of reinforcement learning, the intelligent adaptive learning system can receive the feedback recommended knowledge points determined by the operating user. Further, the agent can determine the multi-dimensional vector state of the operating user's feedback recommended knowledge points environment according to the feedback recommended knowledge points.

[0087] Optionally, if the operating user is the first operating user, the feedback recommended knowledge points can be the target adaptive recommended knowledge points selected by the first operating user according to the adaptive recommended knowledge points determined by the adaptive engine; if the operating user is the second operating user, the feedback recommended knowledge points are the independent recommended knowledge points determined by the second operating user.

[0088] In an optional embodiment of the present invention, the obtaining of the multi-dimensional vector state by observing the knowledge point recommendation environment by the agent may include: obtaining simulated recommended exercises through the knowledge points recommended by the intelligent adaptive learning system according to the knowledge point recommendation environment; automatically simulating the answering results of the target user by the intelligent adaptive learning system according to the student user attributes and learning situation attributes of the target user, and determining the knowledge point mastery state of the target user according to the answering results; the agent receiving the knowledge point mastery state of the target user and determining the multi-dimensional vector state according to the knowledge point mastery state.

[0089] Among them, the simulated recommended exercises can be the exercises recommended by the recommendation module of the intelligent adaptive learning system according to the current knowledge points. The knowledge point mastery state can represent whether the target user has mastered the knowledge points.

[0090] Specifically, the intelligent adaptive learning system can automatically obtain simulated recommended exercises through the recommendation module according to the knowledge points recommended by the knowledge point recommendation environment. After obtaining the simulated recommended exercises, the intelligent adaptive learning system automatically simulates the answering results of the target user according to the student user attributes and learning situation attributes of the target user. That is, during the entire demonstration process, the target user itself does not need to participate in the actual answering process, and the entire learning process can be automatically simulated by the intelligent adaptive learning system. Correspondingly, the intelligent adaptive learning system can automatically simulate the answering results of the target user according to the student user attributes and learning situation attributes of the target user, so as to determine the knowledge point mastery state of whether the target user has mastered the knowledge points according to the answering results. Correspondingly, the agent can observe the intelligent adaptive learning system, and then obtain the multi-dimensional vector state corresponding to the target user.

[0091] S262. Determine the agent action through the agent according to the action selection policy model and the multi-dimensional vector state.

[0092] S263. Execute the agent action through the agent to update the state of the knowledge point recommendation environment according to the execution result of the agent action, and obtain the updated multi-dimensional vector state.

[0093] Among them, the execution result of the agent action, that is, the execution result of the agent action, can act on the current knowledge point recommendation environment, so that the knowledge point recommendation environment updates the current state. The updated multi-dimensional vector state can be the state after the agent executes the agent action and updates the knowledge point recommendation environment.

[0094] S264. Receive, through the agent, the reward value determined by the knowledge point recommendation environment according to the updated multi-dimensional vector state and the agent action.

[0095] S265. Determine, through the agent, an updated agent action according to the reward value, the updated multi-dimensional vector state, and the learning evaluation index.

[0096] Among them, the learning evaluation index includes the knowledge point mastery goal and the teaching rule determination rule of the reward value.

[0097] Among them, the updated agent action can be a new action determined by the agent according to the action selection policy. The knowledge point mastery goal can be to enable the target user to master the most knowledge points as soon as possible during the specified demonstration period. The teaching rule determination rule can be a rule for determining the reward value according to the teaching rules or logic.

[0098] Figure 7 It is a schematic diagram of the effect of the association relationship between some knowledge points of the knowledge graph provided in the second embodiment of the present invention. As Figure 7 shown, taking the study of the knowledge points of quadratic radicals in junior high school mathematics as an example for specific illustration, Figure 7 The list of the knowledge graph shown is the association relationship between a part of the nanoscale knowledge points related to quadratic radicals in the constructed knowledge graph.

[0099] In Figure 7Among them, the third column of the list is the name of the nano-level knowledge point, the second column of the list is the label number of this knowledge point, and the fourth column of the list is the label number of the prerequisite knowledge point of this knowledge point. Usually, the difficulty of subsequent knowledge points is higher than that of prerequisite knowledge points. That is, the more subsequent the knowledge point is, the higher the difficulty. It can be understood that usually, if the current knowledge point is not mastered, it is more reasonable to recommend the prerequisite knowledge point for learning, and it is unreasonable to recommend the subsequent knowledge point. Taking the knowledge point marked as c090201 as an example, its subsequent knowledge points include: c090301, and its prerequisite knowledge points are c090203, c090204, and c090103. After the knowledge point of c090201 is learned excellently, it is in line with the teaching law for the intelligent adaptive learning system to recommend the subsequent knowledge point c090301. However, if the knowledge point of c090201 is learned poorly, recommending the subsequent knowledge point c090301 violates the teaching law. Therefore, the learning evaluation index can specifically determine the rule according to the knowledge point mastery target and the teaching law of the reward value to determine the reward value. When the recommended knowledge point conforms to the teaching law and the target user has mastered the knowledge point, a certain reward can be given; when the recommended knowledge point violates the teaching law and / or the target user has not mastered the knowledge point, a certain punishment can be given.

[0100] S266. Determine whether the learning termination condition of the reinforcement learning model is satisfied through the intelligent agent. If so, execute S267; otherwise, return to execute S261.

[0101] S267. Terminate the reinforcement learning process to obtain the path generation process of the target user's learning path.

[0102] Specifically, the interaction process between the intelligent agent and the environment includes three stages: the environmental observation perceived by the intelligent agent, the action of the intelligent agent, and the environmental feedback. Among them, the environmental observation perceived by the intelligent agent uses a high-dimensional vector to represent the observation result of the intelligent agent on the environmental state. This high-dimensional vector can include the set of information obtained from the intelligent agent. The action of the intelligent agent represents which knowledge point the target user will learn next. The environmental feedback refers to the feedback of the environment to the intelligent agent in the form of a numerical reward. At each time step t, the intelligent agent receives the state information S t ∈S, where S is the set of possible states, and S t represents the state at time t; based on this state, the intelligent agent selects an action A t ∈A(S t ), where A(S t ) is the set of all actions in state S t , and A t represents the action at time t. After one time step, the intelligent agent receives a numerical reward R t+1(Reward at time (t+1)) ∈ R, as the reward for this action, and at the same time observe a new environmental state S t+1 , and thus enter the loop process of the next interaction.

[0103] Optionally, the decision-making module in the intelligent adaptive learning demonstration system can establish a deep reinforcement learning model for the intelligent agent, set the state space of the intelligent agent in the environment, the behavior space that the intelligent agent can decide, and the behavior reward of the environment for the intelligent agent.

[0104] In the learning range of the learning path that the intelligent agent wants to plan, each node corresponds to a nanoscale knowledge point in the intelligent adaptive learning system, and the connection between knowledge points corresponds to the logical relationship between knowledge points in the intelligent adaptive learning system. When the intelligent agent establishes a model, it can obtain the required input from the state module of the intelligent adaptive learning system to achieve the observation of the environment: the state can take the state of the current knowledge point mastered by the target user at the current learning level as the observation value, denoted as (learning level, mastery state), such as: (excellent in learning, learned and mastered).

[0105] Specifically, each knowledge point can have the state mastered by the target user and the learning level of the target student. Among them, the state can include: not yet learned, learned and mastered, not learned, and not determined, etc.; the learning level can be divided into several categories according to user settings, such as: excellent in learning, medium in learning, poor in learning, etc., or according to multiple ability value intervals corresponding to the ability value of the student at each moment obtained by the intelligent adaptive learning system through IRT, such as excellent in learning (ability value in the range of 0.7-1), medium in learning (ability value in the range of 0.36-0.69), poor in learning (ability value in the range of 0.35-0). The operating user can only set the initial learning level of the target student. Once the target user has learned a knowledge point, the learning level is based on the ability value of the target user in the intelligent adaptive learning system. Considering that the target student may re-learn the knowledge points that were not mastered when the ability value was low before, there may be multiple learning level states for the same target user at the same knowledge point.

[0106] Optionally, when performing reinforcement learning on the target user's learning path, the Q-learning algorithm can be used. In the framework of the Q-learning algorithm, Q is Q(s,a), which represents the expected reward obtained by taking action a (a∈A) in state s at a certain moment (s∈S). The knowledge point recommendation environment will feedback the corresponding reward r according to the agent's action. That is to say, it includes an agent, a state set S representing its state in the environment, and an action set A that can be executed in each state. The agent is in the starting state s and selects and executes an action a, a∈A, specifically by randomly selecting a knowledge point from the learning scope corresponding to the target user. This knowledge point can be selected by the adaptive engine or by the operating user. In the interaction with the environment, the agent will transfer from the current state s to the next state s', and will receive an immediate reward r from the environment, and modify the Q value according to the update rule. It can be understood that the update rule can be adaptively set according to the different action selection policy models and specific learning algorithms, and the embodiments of the present invention do not limit this. The purpose of the agent's learning is to maximize the cumulative reward obtained from the environment, that is, to execute the action that obtains the maximum reward in each state. Correspondingly, the method for updating the Q value is as follows:

[0107] Q(s, a) ← Q(s, a) + α[r + γmax a′ Q(s′, a′) - Q(s, a)]

[0108] Among them, α represents the learning rate. The learning rate α∈[0, 1] affects the proportion of the newly learned value replacing the original value in the future. If α = 0, it means that the agent cannot learn new knowledge; if α = 1, it means that the learned knowledge is not stored and all is replaced with new knowledge. γ represents the discount factor. The discount factor γ∈[0, 1] represents the foresight of the agent. Its size affects the weight of the predicted reward of future actions. γ approaching 0 means that the agent only values the reward of the current action and will often execute the behavior that maximizes the current immediate reward; when γ approaches 1, the agent will consider future rewards more; when γ∈[0, 1], it means that the earlier actions have a greater impact, while the later actions have a smaller impact and can even be ignored.

[0109] To solve the problem of the overly large state space (i.e., the curse of dimensionality), Q(s,a) can be represented by a function instead of a Q-table. That is to say, for a given state, the Q value obtained by selecting which action can be calculated by a deep neural network. Optionally, a Deep Q Network (DQN) can be used. After the network is trained, it can be calculated when needed without storing the Q value. The model of the DQN network is as follows:

[0110]

[0111]

[0112] where q π (s, a) represents the discounted return starting from state S, taking action A, and following policy π, which is the effect of taking a specific action from a specific state. E represents expectation, and G t represents the discounted return at time t, w represents the weight, and it can be solved using supervised learning algorithms in machine learning algorithms (such as linear regression, decision tree, neural network, etc.). A suitable function is fitted to extract features of the input state as the input, and the value function is calculated as the output through the Monte Carlo method (MC) or Temporal Difference (TD) learning. Then, the function parameters are trained until convergence. The DQN algorithm continuously learns knowledge during the training process, but what is learned is not the Q value stored in the table, but the learning of the neural network parameters. Through the deep reinforcement learning method, it can be known which action to select to maximize the sum of future rewards obtained.

[0113] In the environmental feedback stage, a discrete reward function R (Reward Function) can be used as the learning evaluation index. The reward function is one of the important elements of the environmental information feedback to the agent. The reward function is to tell the agent the goal it hopes to achieve, rather than how to achieve the goal. The goal of the reward function can be to tell the agent to let the target user master the most knowledge points as soon as possible during the demonstration period, give rewards when the target user masters knowledge points at each time step, and impose certain penalties when the target user does not master knowledge points. At the same time, the design of the reward function also needs to conform to the laws and logic of the knowledge point learning order, consider the diversity and process complexity of the course, and impose certain penalties for recommendations that violate the teaching laws.

[0114] In an alternative embodiment, in the single - user interaction mode of the intelligent adaptive learning demonstration system, the intelligent adaptive learning demonstration system can be externally connected to a human - machine interaction device. The devices available for human - machine interaction mainly include, but are not limited to, keyboards, mice, joysticks, and various pattern recognition devices (such as gesture recognition, action recognition, and speech recognition), etc. Figure 8 is a schematic diagram of the demonstration process of a single - user interaction mode provided by the second embodiment of the present invention. As Figure 8As shown, the knowledge point recommendation module of the intelligent adaptive learning system can send knowledge point data to the interface of the recommendation module and wait for feedback from the target user (referred to as the user). The user can interact with the intelligent adaptive learning demonstration system through a human-computer interaction device. When the user feedback confirms the learning of this knowledge point, the intelligent adaptive learning demonstration system proceeds to the next process and conducts a simulation of knowledge point learning. When the user feedback skips the learning of this knowledge point, it returns to the knowledge point recommendation process of the intelligent adaptive learning system. The intelligent adaptive learning system recommends the next knowledge point and waits for the user's feedback again.

[0115] In an optional embodiment, in the multi-person interaction mode of the intelligent adaptive learning demonstration system, the interaction process between the agent and the environment does not require the knowledge point recommendation module of the intelligent adaptive learning system to recommend knowledge points for the target user. Instead, it receives the action of the operating user to select the next knowledge point to learn and becomes the next state. One time step later, the agent receives a numerical reward as the result of this action and simultaneously observes a new environmental state S t+1 , thus entering the loop process of the next interaction. In this mode, the intelligent adaptive learning demonstration system can demonstrate the difference in the results of the number of knowledge points mastered by the agent and the operating user when recommending the learning strategy of the next knowledge point for the same target user in two ways.

[0116] S270. Intuitively demonstrate the path generation process.

[0117] Correspondingly, step S270 can specifically include the following operations:

[0118] S271. Determine the demonstration state attribute of the path generation process.

[0119] Among them, the demonstration state attribute can be the demonstration attribute specified by the operating user through the learning path demonstration instruction.

[0120] In the embodiments of the present invention, the operating user can also specify the demonstration state attribute through the learning path demonstration instruction. For example, the operating user can specify the demonstration duration or the learning range of the demonstration, and the embodiments of the present invention do not limit this. Correspondingly, the intelligent adaptive learning demonstration system can determine the demonstration state attribute of the path generation process according to the learning path demonstration instruction.

[0121] S272. Intuitively demonstrate the path generation process according to the demonstration state attribute.

[0122] Correspondingly, after the intelligent adaptive learning demonstration system determines the demonstration state attribute of the path generation process according to the learning path demonstration instruction, it can intuitively demonstrate the path generation process according to the demonstration method specified by the operating user.

[0123] Exemplarily, when the demonstration status attribute is a 5-minute demonstration duration, the intelligent adaptive learning demonstration system needs to intuitively demonstrate the path generation process of the learning path completed by the target user under the condition of the largest number of mastered knowledge points within 5 minutes.

[0124] It should be noted that the intelligent adaptive learning demonstration system can display in real time every time a knowledge point content decision is generated during the reinforcement learning process, or can display the complete path generation process in sequence after the reinforcement learning ends. The embodiments of the present invention do not limit this.

[0125] In summary, in the embodiments of the present invention, the intelligent agent of the intelligent adaptive learning demonstration system automatically makes decisions according to the state of the environment, so that the operating user can intuitively understand the decision-making process and effect of the intelligent adaptive learning system's dynamically recommended knowledge points during the specified demonstration period, enriching the demonstration function of the intelligent adaptive learning demonstration system for learning effects, and improving the intuitiveness and intelligence of the intelligent adaptive learning demonstration system in demonstrating learning effects.

[0126] It should be noted that any permutation and combination of the technical features in the above embodiments also belong to the protection scope of the present invention.

[0127] Embodiment III

[0128] Figure 9 is a schematic diagram of a learning path demonstration device for deep reinforcement learning provided by Embodiment III of the present invention. As Figure 9 shown, the device includes: a learning path demonstration instruction receiving module 310, a reinforcement learning factor determination module 320, a reinforcement learning module 330, and a path generation process demonstration module 340, where:

[0129] The learning path demonstration instruction receiving module 310 is used to receive the learning path demonstration instruction of the target user;

[0130] The reinforcement learning factor determination module 320 is used to determine the reinforcement learning factors of the target user's learning path according to the learning path demonstration instruction; wherein, the reinforcement learning factors include a state space, an action space, and a learning evaluation index;

[0131] The reinforcement learning module 330 is used to perform reinforcement learning on the target user's learning path according to the reinforcement learning factors in response to the learning path demonstration instruction, and obtain the path generation process of the target user's learning path;

[0132] The path generation process demonstration module 340 is used to intuitively demonstrate the path generation process.

[0133] In an embodiment of the present invention, an intelligent adaptive learning demonstration system determines reinforcement learning factors such as the state space, action space, and learning evaluation metrics of the learning path of the target user according to the received learning path demonstration instruction. In response to the learning path demonstration instruction, reinforcement learning is performed on the learning path of the target user according to the determined reinforcement learning factors to obtain the path generation process of the learning path of the target user, and the path generation process is visually demonstrated, solving the problem that the existing intelligent adaptive learning demonstration system has a poor demonstration effect on the learning effect, enriching the demonstration function of the intelligent adaptive learning demonstration system for the learning effect, and improving the intuitiveness and intelligence of the intelligent adaptive learning demonstration system in demonstrating the learning effect.

[0134] Optionally, the reinforcement learning factor determination module 320 is specifically configured to: determine the knowledge point recommendation environment in the intelligent adaptive learning system as the knowledge point recommendation environment in the reinforcement learning model; determine the type of knowledge point recommendation environment matched with the state space according to the type of the learning path demonstration instruction; determine the state space according to the student user attributes, the knowledge point attributes of the knowledge graph, the learning situation attributes, and the demonstration attributes; wherein, the student user attributes include the learning ability of the target user; the knowledge point attributes of the knowledge graph include the logical relationship between knowledge points, the knowledge point examination frequency, the importance degree of knowledge points, and the difficulty of knowledge points; the learning situation attributes include the knowledge point mastery state; the demonstration attributes include the number of mastered knowledge points, the learning path demonstration duration, or the knowledge point learning scope; determine the action space according to the recommended learning knowledge points of the target user.

[0135] Optionally, the type of the knowledge point recommendation environment includes at least one of an adaptive engine recommended knowledge point environment and an operator user feedback recommended knowledge point environment; the adaptive engine recommended knowledge point environment is used to recommend knowledge points to the target user by using an adaptive engine; the operator user feedback recommended knowledge point environment is used to recommend knowledge points to the target user according to the feedback information of the operator user; the operator user includes a first operator user or a second operator user; the feedback information of the first operator user is used to confirm whether to accept the knowledge points recommended by the adaptive engine; the feedback information of the second operator user is used to independently recommend knowledge points to the target user.

[0136] Optionally, the reinforcement learning module 330 is specifically configured to: observe the knowledge point recommendation environment through the agent in the reinforcement learning model to obtain a multi-dimensional vector state; determine the agent action by the agent according to the action selection policy model and the multi-dimensional vector state; execute the agent action by the agent to update the state of the knowledge point recommendation environment according to the execution result of the agent action to obtain an updated multi-dimensional vector state; receive, by the agent, a reward value determined by the knowledge point recommendation environment according to the updated multi-dimensional vector state and the agent action; determine an updated agent action by the agent according to the reward value, the updated multi-dimensional vector state, and the learning evaluation index; wherein the learning evaluation index includes a knowledge point mastery target and a teaching rule determination rule for the reward value; return, by the agent, an operation of observing the multi-dimensional vector state obtained by the agent observing the knowledge point recommendation environment until it is determined that the learning termination condition of the reinforcement learning model is satisfied.

[0137] Optionally, if the knowledge point recommendation environment type includes an adaptive engine recommended knowledge point environment, the reinforcement learning module 330 is specifically configured to: determine an adaptive recommended knowledge point through the adaptive engine of the intelligent adaptive learning system according to the knowledge point recommendation algorithm; determine the multi-dimensional vector state of the adaptive engine recommended knowledge point environment by the agent according to the adaptive recommended knowledge point.

[0138] Optionally, if the knowledge point recommendation environment type includes an operation user feedback recommended knowledge point environment, the reinforcement learning module 330 is specifically configured to: receive, by the intelligent adaptive learning system, a feedback recommended knowledge point determined by the operation user; determine the multi-dimensional vector state of the operation user feedback recommended knowledge point environment by the agent according to the feedback recommended knowledge point; wherein, if the operation user is the first operation user, the feedback recommended knowledge point is the target adaptive recommended knowledge point selected by the first operation user according to the adaptive recommended knowledge point determined by the adaptive engine; if the operation user is the second operation user, the feedback recommended knowledge point is the self-recommended knowledge point determined by the second operation user.

[0139] Optionally, the reinforcement learning module 330 is specifically configured to: obtain a simulated recommended exercise through the intelligent adaptive learning system according to the knowledge point recommended by the knowledge point recommendation environment; automatically simulate the answering result of the target user by the intelligent adaptive learning system according to the student user attribute and the learning situation attribute of the target user, and determine the knowledge point mastery state of the target user according to the answering result; receive, by the agent, the knowledge point mastery state of the target user, and determine the multi-dimensional vector state according to the knowledge point mastery state.

[0140] Optionally, the path generation process demonstration module 340 is specifically configured to: determine the demonstration status attribute of the path generation process; and perform an intuitive demonstration of the path generation process according to the demonstration status attribute.

[0141] The above learning path demonstration device for deep reinforcement learning can execute the learning path demonstration method for deep reinforcement learning provided in any embodiment of the present invention, and has functional modules and beneficial effects corresponding to the execution of the method. For technical details not described in detail in this embodiment, reference may be made to the learning path demonstration method for deep reinforcement learning provided in any embodiment of the present invention.

[0142] Since the above-introduced learning path demonstration device for deep reinforcement learning is a device that can execute the learning path demonstration method for deep reinforcement learning in the embodiments of the present invention, based on the learning path demonstration method for deep reinforcement learning introduced in the embodiments of the present invention, those skilled in the art can understand the specific implementation manners and various variations of the learning path demonstration device for deep reinforcement learning in this embodiment. Therefore, the implementation of how the learning path demonstration device for deep reinforcement learning realizes the learning path demonstration method for deep reinforcement learning in the embodiments of the present invention will not be described in detail here. As long as the device adopted by those skilled in the art to implement the learning path demonstration method for deep reinforcement learning in the embodiments of the present invention belongs to the scope to be protected by this application.

[0143] Embodiment 4

[0144] Figure 10 FIG. is a schematic structural diagram of an electronic device provided in Embodiment 4 of the present invention. Figure 10 FIG. shows a block diagram of an exemplary electronic device 12 suitable for implementing the embodiments of the present invention. Figure 10 The shown electronic device 12 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.

[0145] As Figure 10 shown, the electronic device 12 is presented in the form of a general-purpose computing device. The components of the electronic device 12 may include, but are not limited to: one or more processors 16, a memory 28, and a bus 18 connecting different system components (including the memory 28 and the processor 16).

[0146] Bus 18 represents one or more of several types of bus architectures, including a memory bus or memory controller, a peripheral bus, an Accelerated Graphics Port, a processor, or a local bus using any of the various bus architectures. By way of example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.

[0147] Electronic device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by electronic device 12, including volatile and nonvolatile media, removable and non-removable media.

[0148] Memory 28 may include computer system readable media in the form of volatile memory, such as Random Access Memory (RAM) 30 and / or cache memory 32. Electronic device 12 may further include other removable / non-removable, volatile / nonvolatile computer system storage media. By way of example only, storage system 34 can be used for reading and writing on non-removable, nonvolatile magnetic media ( Figure 10 not shown, typically referred to as a "hard disk drive"). Although Figure 10 not shown in, a disk drive for reading and writing on a removable nonvolatile disk (such as a "floppy disk") and an optical disk drive for reading and writing on a removable nonvolatile optical disk (such as a Compact Disc-Read Only Memory (CD-ROM), Digital Video Disc-Read Only Memory (DVD-ROM), or other optical media) can be provided. In these cases, each drive can be connected to bus 18 through one or more data media interfaces. Memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present invention.

[0149] A program / utilities 40 having a set (at least one) of program modules 42 can be stored, for example, in a memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. The program modules 42 generally execute the functions and / or methods in the embodiments described in the present invention.

[0150] The electronic device 12 can also communicate with one or more external devices 14 (such as a keyboard, a pointing device, a display 24, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 12, and / or communicate with any device that enables the electronic device 12 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface 22. Moreover, the electronic device 12 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 20. As shown in the figure, the network adapter 20 communicates with other modules of the electronic device 12 through a bus 18. It should be understood that although Figure 10 not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, (Redundant Arrays of Independent Disks, RAID) systems, tape drives, and data backup storage systems, etc.

[0151] The processor 16 executes various functional applications and data processing by running programs stored in the memory 28, and implements the learning path demonstration method of deep reinforcement learning provided by the embodiments of the present invention: receiving a learning path demonstration instruction of a target user; determining reinforcement learning factors of the learning path of the target user according to the learning path demonstration instruction; wherein, the reinforcement learning factors include an agent, a learning environment, a state space, an action space, and a learning evaluation index; in response to the learning path demonstration instruction, performing reinforcement learning on the learning path of the target user according to the reinforcement learning factors to obtain a path generation process of the learning path of the target user; and intuitively demonstrating the path generation process.

[0152] Embodiment 5

[0153] Embodiment 5 of the present invention further provides a computer storage medium storing a computer program, where the computer program is used to execute the learning path demonstration method of deep reinforcement learning according to any one of the above embodiments of the present invention when executed by a computer processor: receiving a learning path demonstration instruction of a target user; determining reinforcement learning factors of the learning path of the target user according to the learning path demonstration instruction; where the reinforcement learning factors include an agent, a learning environment, a state space, an action space, and a learning evaluation index; in response to the learning path demonstration instruction, performing reinforcement learning on the learning path of the target user according to the reinforcement learning factors to obtain a path generation process of the learning path of the target user; and visually demonstrating the path generation process.

[0154] The computer storage medium of the embodiments of the present invention may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM) or a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device.

[0155] The computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, where the data signal carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium may send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.

[0156] The program code contained on the computer-readable medium may be transmitted by any appropriate medium, including but not limited to wireless, wire, optical cable, radio frequency (RF), etc., or any suitable combination of the above.

[0157] Computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0158] Note that the above is only the preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments may be included, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A learning path demonstration method for deep reinforcement learning, characterized in that Including: Receiving a learning path demonstration instruction of a target user; Determining reinforcement learning factors of the learning path of the target user according to the learning path demonstration instruction; wherein, the reinforcement learning factors include an agent, a learning environment, a state space, an action space, and a learning evaluation index; In response to the learning path demonstration instruction, performing reinforcement learning on the learning path of the target user according to the reinforcement learning factors to obtain a path generation process of the learning path of the target user; Intuitively demonstrating the path generation process; The determining the reinforcement learning factors of the learning path of the target user according to the learning path demonstration instruction includes: Determining an intelligent adaptive learning system as a knowledge point recommendation environment in the reinforcement learning model; Determining a knowledge point recommendation environment type matched with the state space according to the type of the learning path demonstration instruction; Determining the state space according to student user attributes, knowledge point attributes of a knowledge graph, learning situation attributes, and demonstration attributes; wherein, the student user attributes include the learning ability of the target user; the knowledge point attributes of the knowledge graph include the logical relationship between knowledge points, the examination frequency of knowledge points, the importance degree of knowledge points, and the difficulty of knowledge points; the learning situation attributes include the mastery state of knowledge points; the demonstration attributes include the number of mastered knowledge points, the learning path demonstration duration, or the knowledge point learning scope; Determining the action space according to the recommended learning knowledge points of the target user; Wherein, the knowledge point recommendation environment type includes at least one of an adaptive engine recommended knowledge point environment and an operator feedback recommended knowledge point environment; The adaptive engine recommended knowledge point environment is used to recommend knowledge points to the target user by using an adaptive engine; The operator feedback recommended knowledge point environment is used to recommend knowledge points to the target user according to the feedback information of an operator; the operator includes a first operator or a second operator; the feedback information of the first operator is used to confirm whether to accept the knowledge points recommended by the adaptive engine; the feedback information of the second operator is used to independently recommend knowledge points to the target user.

2. The method according to claim 1, characterized in that, The performing reinforcement learning on the learning path of the target user according to the reinforcement learning factors includes: Observing a knowledge point recommendation environment through an agent in the reinforcement learning model to obtain a multi-dimensional vector state; Determining an agent action by the agent according to an action selection policy model and the multi-dimensional vector state; Executing the agent action by the agent to update the state of the knowledge point recommendation environment according to the execution result of the agent action to obtain an updated multi-dimensional vector state; Receiving, by the agent, a reward value determined by the knowledge point recommendation environment according to the updated multi-dimensional vector state and the agent action; Determining an updated agent action by the agent according to the reward value, the updated multi-dimensional vector state, and the learning evaluation index; wherein, the learning evaluation index includes a knowledge point mastery target and a teaching rule determination rule of the reward value; Returning, by the agent, to execute the operation of obtaining the multi-dimensional vector state by observing the knowledge point recommendation environment through the agent until it is determined that the learning termination condition of the reinforcement learning model is satisfied.

3. The method according to claim 2, characterized in that, If the knowledge point recommendation environment type includes an adaptive engine recommendation knowledge point environment, the multi-dimensional vector state is obtained by observing the knowledge point recommendation environment through the agent, including: Determining the adaptive recommended knowledge points through the adaptive engine of the adaptive learning system according to the knowledge point recommendation algorithm; Determining the multi-dimensional vector state of the adaptive engine recommendation knowledge point environment through the agent according to the adaptive recommended knowledge points.

4. The method according to claim 2, wherein If the knowledge point recommendation environment type includes an operating user feedback recommendation knowledge point environment, the multi-dimensional vector state is obtained by observing the knowledge point recommendation environment through the agent, including: Receiving the feedback recommended knowledge points determined by the operating user through the adaptive learning system; Determining the multi-dimensional vector state of the operating user feedback recommendation knowledge point environment through the agent according to the feedback recommended knowledge points; Wherein, if the operating user is the first operating user, the feedback recommended knowledge points are the target adaptive recommended knowledge points selected by the first operating user according to the adaptive recommended knowledge points determined by the adaptive engine; If the operating user is the second operating user, the feedback recommended knowledge points are the independent recommended knowledge points determined by the second operating user.

5. The method according to claim 3 or 4, characterized in that The obtaining of the multi-dimensional vector state by observing the knowledge point recommendation environment through the agent includes: Obtaining simulated recommended exercises according to the knowledge points recommended by the knowledge point recommendation environment through the adaptive learning system; Automatically simulating the answering results of the target user according to the student user attributes and learning situation attributes of the target user through the adaptive learning system, and determining the knowledge point mastery state of the target user according to the answering results; Receiving the knowledge point mastery state of the target user through the agent, and determining the multi-dimensional vector state according to the knowledge point mastery state.

6. The method according to claim 1, wherein The intuitive demonstration of the path generation process includes: Determining the demonstration state attribute of the path generation process; Intuitively demonstrating the path generation process according to the demonstration state attribute.

7. A learning path demonstration device for deep reinforcement learning, characterized in that, Including: A learning path demonstration instruction receiving module, configured to receive a learning path demonstration instruction of a target user; A reinforcement learning factor determining module, configured to determine the reinforcement learning factors of the target user's learning path according to the learning path demonstration instruction; wherein, the reinforcement learning factors include a state space, an action space, and a learning evaluation index; A reinforcement learning module, configured to respond to the learning path demonstration instruction, perform reinforcement learning on the target user's learning path according to the reinforcement learning factors, and obtain the path generation process of the target user's learning path; A path generation process demonstration module, configured to intuitively demonstrate the path generation process; The reinforcement learning factor determination module is specifically configured to: determine the intelligent adaptive learning system as the knowledge point recommendation environment in the reinforcement learning model; determine the type of knowledge point recommendation environment matched with the state space according to the type of the learning path demonstration instruction; determine the state space according to the student user attributes, the knowledge point attributes of the knowledge graph, the learning situation attributes, and the demonstration attributes, where the student user attributes include the learning ability of the target user; the knowledge point attributes of the knowledge graph include the logical relationship between knowledge points, the knowledge point examination frequency, the importance degree of knowledge points, and the difficulty of knowledge points; the learning situation attributes include the knowledge point mastery state; the demonstration attributes include the number of mastered knowledge points, the learning path demonstration duration, or the knowledge point learning scope; determine the action space according to the recommended learning knowledge points of the target user. Among them, the type of the knowledge point recommendation environment includes at least one of an adaptive engine recommended knowledge point environment and an operator user feedback recommended knowledge point environment; the adaptive engine recommended knowledge point environment is used to recommend knowledge points to the target user by using an adaptive engine; the operator user feedback recommended knowledge point environment is used to recommend knowledge points to the target user according to the feedback information of the operator user; the operator user includes a first operator user or a second operator user; the feedback information of the first operator user is used to confirm whether to accept the knowledge points recommended by the adaptive engine; the feedback information of the second operator user is used to independently recommend knowledge points to the target user.

8. An electronic device, characterized in that, The electronic device includes: One or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the learning path demonstration method of deep reinforcement learning as described in any one of claims 1-6.

9. A computer storage medium, on which a computer program is stored, characterized in that, When the program is executed by the processor, it implements the learning path demonstration method of deep reinforcement learning as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Data processing method and device, medium and electronic device

    CN109886848A

  • A self-adaptive learning path planning system based on reinforcement learning

    CN109948054A

Cited By

  • Learning path generation method based on reinforcement learning

    CN121981299A