Data Processing Method and Apparatus, Electronic Device, and Computer-Readable Storage Medium
Through the interaction between the two-person game model and multiple subject models, the subject state of the learning subject is dynamically adjusted, and the problem of insufficient adaptability of learners in traditional machine teaching models is solved, achieving more efficient learning personalization and adaptability.
Patent Information
- Application Number
- CN202411705764.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-26
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-11-26
AI Technical Summary
Due to fixed induction bias, traditional machine teaching models have caused learners to lack adaptability when facing new tasks and lack generalization skills, making it difficult for them to learn and transfer knowledge from different types of tasks.
The two-person game model is adopted, based on the interaction between the education subject and the learning subject, an educational strategy is generated, and the subject state of the learning subject is dynamically adjusted. Through the interaction between the education subject model, the learning subject model and the subject state transformation model, the learning strategy and subject state are optimized.
By dynamically adjusting the subject status of the learning subject, the personalization and adaptability of the learning subject is improved, the learning subject can better adapt to the learning tasks, and the independent learning ability of the learning subject is enhanced.
Smart Images

Figure CN119624713B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the field of education, and particularly relates to a data processing method and apparatus, an electronic device, and a computer-readable storage medium. Background Art
[0002] The education system can be regarded as an intelligent system, in which an agent - teacher transmits information to another agent - learner. As a computational model of the education system, machine teaching aims to solve the problem of how to discover the best training data to guide the learner to reach the target model with minimal effort.
[0003] Traditional machine teaching is limited to a class of restrictive learners and hyperparameters with fixed inductive biases. Such learners cannot update their inductive biases during the learning process. For example, in the teacher - student model, which involves the process of the teacher imparting knowledge to the student, the teacher generates learning strategies and guides the learner to learn, thereby achieving better learning efficiency. However, due to the fixed inductive bias of the teacher - student model, the learner lacks adaptability when facing new tasks. The learner is unable to adjust its learning strategy according to the changes in the task, resulting in insufficient generalization ability and difficulty in learning and transferring knowledge from different types of tasks.
[0004] The information disclosed in this background art section is only intended to increase the understanding of the overall background of the present disclosure and should not be regarded as an admission or any form of suggestion that this information constitutes prior art already known to those of ordinary skill in the art. Summary of the Invention
[0005] The present disclosure provides a data processing method and apparatus, an electronic device, and a computer-readable storage medium.
[0006] According to a first aspect, a data processing method is provided. The method includes: generating an education strategy based on a two-player game model and teaching task data, where the two-player game model is a model pre-constructed for an education subject and a learning subject; inputting the education strategy into an education subject model pre-constructed for the education subject to obtain an action sequence output by the education subject model, where the education subject model is used to represent the correspondence between the education strategy and the action sequence; inputting the action sequence and the teaching task data into a learning subject model pre-constructed for the learning subject to obtain a learning strategy output by the learning subject model, where the learning subject model is used to represent the correspondence between the action sequence and the teaching task data and the learning strategy; inputting the learning strategy into a pre-constructed subject state transformation model to obtain a current subject state output by the subject state transformation model, where the subject state transformation model is used to represent the correspondence between the subject state and the learning strategy; and evaluating the performance of the learning subject model on the teaching task data based on the current subject state to obtain performance data.
[0007] According to a second aspect, a data processing device is provided. The device includes: a generation unit configured to generate an education strategy based on a two-player game model and teaching task data, where the two-player game model is a model pre-constructed for an education entity and a learning entity; an action obtaining unit configured to input the education strategy into an education entity model pre-constructed for the education entity to obtain an action sequence output by the education entity model, where the education entity model is used to represent the correspondence between the education strategy and the action sequence; a strategy obtaining unit configured to input the action sequence and the teaching task data into a learning entity model pre-constructed for the learning entity to obtain a learning strategy output by the learning entity model, where the learning entity model is used to represent the correspondence between both the action sequence and the teaching task data and the learning strategy; a transformation unit configured to input the learning strategy into a pre-constructed entity state transformation model to obtain a current entity state output by the entity state transformation model, where the entity state transformation model is used to represent the correspondence between the entity state and the learning strategy; and an evaluation unit configured to evaluate the performance of the learning entity model on the teaching task data based on the current entity state to obtain performance data.
[0008] According to a third aspect, an electronic device is provided. The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor, where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in any implementation manner of the first aspect.
[0009] According to a fourth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, where the computer instructions are used to cause a computer to execute the method described in any implementation manner of the first aspect.
[0010] The data processing method and device provided by the embodiments of the present disclosure first generate an education strategy based on a two-player game model and teaching task data. The two-player game model is a model pre-constructed for the education subject and the learning subject. Secondly, the education strategy is input into the education subject model pre-constructed for the education subject to obtain an action sequence output by the education subject model. The education subject model is used to represent the correspondence between the education strategy and the action sequence. Thirdly, the action sequence and the teaching task data are input into the learning subject model pre-constructed for the learning subject to obtain a learning strategy output by the learning subject model. The learning subject model is used to represent the correspondence between the action sequence and the teaching task data and the learning strategy. Then, the learning strategy is input into the pre-constructed subject state transformation model to obtain the current subject state output by the subject state transformation model. The subject state transformation model is used to represent the correspondence between the subject state and the learning strategy. Finally, based on the current subject state, the performance of the learning subject model on the teaching task data is evaluated to obtain performance data. Thus, through the two-player game model, the subject state of the learning subject can change dynamically with the change of the action sequence of the education subject, strengthening the adjustment and evolution of the subject state of the learning subject, helping the learning subject better adapt to the learning task, and improving the learning personalization and adaptability of the learning subject.
[0011] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0013] Figure 1 is a flowchart according to an embodiment of the data processing method of the present disclosure;
[0014] Figure 2 is a schematic structural diagram according to an embodiment of the data processing device of the present disclosure;
[0015] Figure 3 is a block diagram of an electronic device for implementing the data processing method of the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] Unless otherwise explicitly stated, in the whole specification and claims, the term "comprise" or its variations such as "comprises" or "including" etc. will be understood to include the stated elements or components, without excluding other elements or other components.
[0017] The technical solutions of the present disclosure are illustrated below through specific embodiments. It should be understood that one or more steps mentioned in the present disclosure do not exclude the existence of other methods and steps before and after the combined steps, or other methods and steps can be inserted between these explicitly mentioned steps. It should also be understood that these examples are only used to illustrate the present disclosure and not to limit the scope of the present disclosure. Unless otherwise specified, the numbers of the method steps are only for the purpose of identifying the method steps, rather than limiting the arrangement order of each method or the implementation scope of the present disclosure. The change or adjustment of their relative relationship can also be regarded as the implementable scope of the present disclosure under the condition of no substantial change in technical content.
[0018] There are no specific restrictions on the sources of the raw materials and instruments used in the embodiments, and they can be purchased in the market or prepared according to the conventional methods well-known to those skilled in the art.
[0019] Aiming at the deficiencies in the traditional technology, the present disclosure proposes a data processing method. Through the two-player game model, the educational entity can guide the learning entity to continuously adjust its own entity state in practice, form an inductive bias for future tasks, and improve the learning personalization and adaptability of the learning entity. Figure 1 Flow 100 of an embodiment of the data processing method according to the present disclosure is shown. The above data processing method includes the following steps:
[0020] Step 101, generating an educational strategy based on the two-player game model and teaching task data.
[0021] In this embodiment, the two-player game model is a mathematical model pre-constructed for the educational entity and the learning entity, used to describe the interaction between the educational entity and the learning entity. Since the educational entity and the learning entity have multiple-dimensional states, in order to effectively quantify the educational entity and the learning entity, an educational entity model and a learning entity model can be constructed for the educational entity and the learning entity respectively.
[0022] In this embodiment, the learning entity model and the educational entity model are constructed separately first, and then the two-player game model is constructed, allowing for focusing on understanding the independent behaviors and characteristics of each entity. Once a clear understanding of the learning entity and the educational entity is obtained, they can be put into the two-player game model to consider the strategic interaction between the learning entity model and the educational entity model.
[0023] In this embodiment, the educational agent model is a model of the behavioral strategy constructed for the educational agent, which can reflect at least one aspect of the educational agent; the learning agent model is a model of the behavioral strategy constructed for the learning agent, which can reflect at least one aspect of the learning agent; the educational agent model and the learning agent model can be two key components in the two-player game model. The educational agent model represents the behavioral strategy of the educational agent, and the learning agent model represents the behavioral strategy of the learning agent. The educational agent model and the learning agent model cooperate with each other in the two-player game model to jointly promote the learning process. The actual game process between the educational agent and the learning agent is also a process of adjusting the parameters of the educational agent model and the learning agent model. The educational agent model guides the learning agent to learn through the action sequence and educational strategy, while the learning agent model continuously adjusts the learning strategy according to the guidance of the educational agent and its own learning state, and finally achieves the learning goal.
[0024] Optionally, the agent states of the educational agent model and the learning agent can also be components of the two-player game model. The two-player game model can also characterize the corresponding relationship between the educational strategy of the educational agent and the agent state of the learning agent. Through the two-player game model, the adjustment and evolution of the learner's agent state can be emphasized. The strategy spaces of the learning agent and the educational agent are defined respectively in the two-player game model. The two-player game model can enable the educational agent model to dynamically adjust its teaching strategy according to the current agent state of the learner, thereby optimizing the learning process of the learning agent.
[0025] In this embodiment, the agent state is a multi-dimensional state space, where the dimensions represent the characteristics, knowledge level, and confidence level of the learning agent. The initial value of the agent state of the learning agent can be determined through preliminary evaluation or basic knowledge tests.
[0026] In this embodiment, the educational task data is teaching data related to the learning agent and the educational agent respectively. The agent state of the learning agent will change with the teaching task data and the learning process of the learning agent. For example, the educational task data includes: learning knowledge data, and the agent state of the learning agent will become richer and more perfect due to the learning knowledge data. The educational task data also includes important factors that shape the agent state of the learning agent. For example, the educational task data includes: the knowledge level and learning preferences of the learning agent. The educational agent can formulate better educational strategies by understanding the knowledge level and learning preferences of the learning agent. The educational task data can provide a basis for parameter setting of the two-player game model. For example, through the educational task data, the knowledge level and learning preferences of the learning agent can be determined, so that better teaching strategies can be specified; the internal state of the learning agent will change with the change of the educational task data and the learning process; the payoff functions of the educational agent model and the learning agent model are both related to the educational task data.
[0027] In this embodiment, an educational strategy is a strategy for guiding a learning subject to learn. Through the educational strategy, the educational subject can be guided to perform a multi-step action sequence. Each action in the action sequence has its own execution order, and each action is aimed at realizing the educational strategy. The actions of the educational subject can be in various forms, such as giving specific feedback, asking questions, guiding discussions, or providing examples, etc. The educational strategy is a rule established through a two-player game model, used to describe the impact of each action of the educational subject on the subject state of the learning subject. For example, when the educational subject gives positive feedback, the confidence level of the learning subject may increase; after the educational subject points out the mistakes, the learning subject may need to adjust its knowledge state.
[0028] In this embodiment, step 101 above includes: determining the subject state of the learning subject based on the teaching task data; inputting the subject state into the two-player game model to obtain the educational strategy output by the two-player game model. The two-player game model is a model reflecting the corresponding relationship between the subject state of the learning subject and the educational strategy of the educational subject. The learning subject model of the learning subject is used to generate the subject state of the learning subject, and the educational subject model of the educational subject is used to generate the educational strategy.
[0029] Optionally, the educational task data further includes: educational indicators and learning indicators. The educational indicators are the indicators of the profit function of the educational subject model in the two-player game model, including the learning achievements and learning efficiency of the learning subject. The learning indicators are the indicators of the profit function of the learning subject model in the two-player game model, including: learning satisfaction and knowledge mastery level.
[0030] Optionally, step 101 above further includes: detecting whether the educational strategy meets the indicator requirements of the educational indicators in the educational task data, where the indicator requirements refer to that the educational indicators meet the preset thresholds. If it meets the requirements, execute step 102; otherwise, adjust the parameters of the educational subject model or the learning subject model in the two-player game model to make the educational strategy meet the indicator requirements of the educational indicators.
[0031] In this embodiment, the educational subject is a subject that provides knowledge and learning directions to the learning subject based on the teaching task data. Among them, the behavior strategy of the educational subject can be reflected by the educational subject model constructed for the educational subject, and the educational subject model can generate an action sequence, and the action sequence can act on the learning subject. For example, the educational subject is a virtual person such as a digital teacher or a digital master, or the educational subject is an actual person; the learning subject is a subject that receives the information of the educational subject and makes a learning state. Among them, the behavior strategy of the learning subject can be reflected by the learning subject model constructed for the learning subject. For example, the learning subject is a student, an apprentice, etc.
[0032] Step 102: Input the education strategy into the education agent model pre-constructed for the education agent to obtain the action sequence output by the education agent model.
[0033] In this embodiment, the education agent model is used to represent the correspondence between the education strategy and the action sequence. Among them, both the education strategy and the action sequence are information involved in the education agent in the two-player game model. The education strategy is the strategy for the education agent to educate the learning agent, and the action sequence is the action sequence when the education agent specifically educates the learning agent.
[0034] In this embodiment, the education agent model is a machine learning model. When constructing the education agent model, the general steps include: First, a functional relationship needs to be defined, which maps the input (such as the data corresponding to the education strategy) to the output (such as the data corresponding to the action sequence of the education agent). Then a cost function (also called a loss function) is defined, which measures the difference between the predicted value and the actual value of the model. The cost function is used to evaluate the quality of the education agent model. The parameters in the functional relationship are adjusted through an optimization process (such as the gradient descent algorithm) to minimize the cost function. This process aims to find a set of parameters that make the functional relationship perform best on the given data set. After multiple iterations of optimization, the finally obtained functional relationship should be the optimal or most suitable education agent model for the current teaching task data.
[0035] Step 103: Input the action sequence and the teaching task data into the learning agent model pre-constructed for the learning agent to obtain the learning strategy output by the learning agent model.
[0036] In this embodiment, the learning agent model is used to represent the correspondence between the action sequence and the teaching task data and the learning strategy. Among them, the learning strategy is the information involved in the learning agent in the two-player game model, and the teaching task data is the information involved in the learning agent and the education agent in the two-player game model. The learning strategy is the strategy for the learning agent to learn the knowledge corresponding to the teaching task data under the guidance of the action sequence of the education agent, and the action sequence is the action sequence when the education agent specifically educates the learning agent.
[0037] In this embodiment, the learning agent model is a machine learning model. When constructing the learning agent model, the general steps include: First, a functional relationship needs to be defined, which maps the input (such as learning behavior data) to the output (such as learning outcomes or teaching effects). Next, a cost function is defined, which measures the difference between the predicted value and the actual value of the learning agent model. The cost function is used to evaluate the quality of the learning agent model. The parameters in the functional relationship are adjusted through an optimization process (such as the gradient descent algorithm) to minimize the cost function. This process aims to find a set of parameters that make the functional relationship perform best on the given data set. After multiple iterations of optimization, the finally obtained functional relationship should be the optimal or most suitable learning agent model for the current teaching task data.
[0038] Step 104: Input the learning strategy into the pre-constructed agent state transformation model to obtain the current agent state output by the agent state transformation model.
[0039] In this embodiment, the agent state transformation model is used to represent the correspondence between the agent state and the learning strategy. The agent state transformation model aims to describe the state changes of the learning agent under different learning strategies, rather than the changes of the learning strategy itself.
[0040] In this embodiment, the agent state is a multi-dimensional state space of the learning agent. The dimensions represent the knowledge level, characteristics, confidence, etc. of the learning agent under the current teaching task data. The agent state transformation model is a reflection of the agent state of the learning agent under the current learning strategy.
[0041] In this embodiment, through the agent state transformation model, the educational agent can monitor the changes in the agent state of the learning agent in real time during each educational interaction. By collecting the agent state of the learning agent, it breaks through the limitation of traditional machine teaching that only targets static learning agents, and endows the learning agent with higher adaptability and autonomous learning ability.
[0042] Step 105: Based on the current agent state, evaluate the performance of the learning agent model on the teaching task data to obtain performance data.
[0043] In this embodiment, the performance data is an embodiment of the specific task completion degree of the learning agent model on the data task data. The learning agent model outputs a learning strategy, which passes through the agent state transformation model to obtain the current agent state of the learning agent. The current agent state of the learning agent can reflect the dimension values of the learning agent or the learning agent model from multiple dimensions. The dimension values of each dimension represented by the agent state are compared with the standard values under the corresponding dimensions of the teaching task data to obtain the performance data of the learning agent model on the teaching task data.
[0044] Specifically, step 105 above includes: determining the dimensional values of each dimension of the current subject state; extracting the standard values related to each dimension from the teaching task data; for each dimension, comparing the dimensional value of this dimension with the standard value of this dimension to obtain the comparison data for each dimension; performing weighted summation on the comparison data of all dimensions in the current subject state to obtain the performance data.
[0045] Optionally, step 105 above includes: evaluating the performance of the learning subject model on the current task using multiple evaluation metrics according to the current subject state and the learning task data to obtain the performance data, where the evaluation metrics can be a loss function, an F1 score, etc., and the performance data can be the accuracy rate, generalization ability, etc. of the learning subject model.
[0046] Optionally, in order to better adapt to the changes in the subject state of the learning subject, the present disclosure can also introduce a collaborative learning algorithm. Specifically, this data processing method includes: obtaining the experience data of the reference subject, and in response to the data value of the performance data being less than the data value of the experience data, adjusting the parameters of the learning subject model. This collaborative learning algorithm allows the experiences of multiple learning subjects to learn from each other, and at the same time provides dynamic resources and policy guidance based on the state of the learning subject. For example, when a learning subject needs repeated guidance on a certain concept, the system can identify its learning obstacles and meet its needs through different teaching strategies (such as examples, analogies, or enhanced feedback).
[0047] The data processing method provided in this embodiment combines digital technology and intelligent technology, and through means such as computer processing, big data analysis, and artificial intelligence, realizes the highly integrated, optimized, and personalized push of resources, so as to achieve the purpose of improving efficiency, enhancing interactivity, and personalized experience. In the field of education, digital intelligence usually involves converting traditional educational resources into digital forms and using intelligent algorithms to adapt to the needs of learning subjects and provide adaptive and personalized learning experiences.
[0048] The present disclosure proposes a new teaching framework. If the initial bias of the learning agent does not suit the current task and cannot be changed, the teaching agent may have to hide a part of the training data to teach them a good model. More precisely, by hiding part of the data, the learning agent can be taught a better model than by providing the entire data set. For teaching machine learning algorithms, this teaching strategy is similar to data poisoning, while for teaching humans, it is an undesirable behavior that attempts to manipulate the learning agent. However, considering that the bias of the learning agent can be changed and influenced by the teaching agent, this gives rise to a completely new teaching strategy: helping the learning agent adjust their agent state, essentially teaching them better inductive biases. This approach not only teaches the learning agent to perform better during the learning phase but also helps them perform better in future tasks without the teaching agent, thus empowering the learning agent.
[0049] The present disclosure proposes the concept of "teaching for learning", and the concept of "teaching for learning" (TtL) emphasizes that the teaching agent not only imparts knowledge but also cultivates the learning ability of the learning agent, enabling it to select the best model without the intervention of the teaching agent and apply it to future similar tasks. The present disclosure formalizes the concept of "teaching for learning" as a two-player game between the learning agent and the teaching agent, unifies the inductive biases in machine learning algorithms (such as model families, initializations, etc.) and the inductive biases in human learning agents (such as prior task knowledge and meta-knowledge, etc.), and models both cases as the potential agent state of the learning agent. The main difference from traditional machine teaching is that the agent state of the learning agent in the new framework of the present disclosure changes dynamically with the behavior of the teaching agent. Therefore, in the new framework, the task of the teaching agent not only includes guiding the learning agent to approach the best model but also includes ensuring that the learning agent can autonomously learn a good model without the supervision of the teaching agent in future similar tasks. To achieve this goal, the teaching agent needs to guide the learning agent to reach an agent state that contains inductive biases suitable for current and future tasks.
[0050] The data processing method provided by the embodiments of the present disclosure first generates an education strategy based on a two-player game model and teaching task data. The two-player game model is a model pre-constructed for an education subject and a learning subject. Secondly, the education strategy is input into the education subject model pre-constructed for the education subject to obtain an action sequence output by the education subject model. The education subject model is used to represent the correspondence between the education strategy and the action sequence. Thirdly, the action sequence and the teaching task data are input into the learning subject model pre-constructed for the learning subject to obtain a learning strategy output by the learning subject model. The learning subject model is used to represent the correspondence between the action sequence and the teaching task data and the learning strategy. Then, the learning strategy is input into the pre-constructed subject state transformation model to obtain the current subject state output by the subject state transformation model. The subject state transformation model is used to represent the correspondence between the subject state and the learning strategy. Finally, based on the current subject state, the performance of the learning subject model on the teaching task data is evaluated to obtain performance data. Thus, through the two-player game model, the subject state of the learning subject can change dynamically with the change of the action sequence of the education subject, strengthening the adjustment and evolution of the subject state of the learning subject, helping the learning subject better adapt to the learning task, and improving the learning personalization and adaptability of the learning subject.
[0051] In some embodiments of the present disclosure, the above method further includes: obtaining a new education strategy based on the performance data; inputting the new education strategy into the education subject model to obtain a new action sequence output by the education subject model; inputting the new action sequence and the teaching task data into the learning subject model to obtain a new learning strategy output by the learning subject model.
[0052] In this embodiment, the performance data is the performance result of the learning subject obtained by analyzing the learning subject model. The performance data includes: learning bottleneck data, learning strategy effectiveness data, and learning mode. Among them, the learning bottleneck can represent the aspects where the learning subject has difficulties or deficiencies. For example, if the learning subject scores low on a certain specific concept, it indicates that they may have obstacles in understanding this concept. The learning strategy effectiveness data is data that helps evaluate the effectiveness of the existing education strategy. For example, if the learning subject does not show obvious improvement after receiving a certain teaching method, it indicates that this strategy may need to be adjusted or improved. The learning mode is the learning mode and regular data of the learning subject. For example, some learning subjects may be more suitable for visual learning, while others may be more suitable for auditory learning.
[0053] In this embodiment, obtaining a new educational strategy based on the performance data includes: updating the parameters in the educational subject model based on the performance data to make it more accurately reflect the learning situation of the learning subject. For example, according to the performance of the learning subject on different tasks, the parameters related to the learning speed or learning style of the learning subject in the model can be adjusted. The educational subject model can generate a new educational strategy based on the updated parameters. For example, if the model finds that the learning subject has difficulties in a certain concept, it can suggest that the educational subject adopt different teaching methods or provide more practice opportunities.
[0054] In this embodiment, the educational subject can monitor the state changes of the learning subject in real time during each teaching interaction. By collecting information such as the reactions, learning behaviors, guesses, and mistakes of the learning subject, the educational subject can adjust the teaching strategy in a timely manner based on the current state and performance of the learning subject. This can be achieved with the help of a feedback loop to ensure that the state of the learning subject can flexibly respond to the guidance of the educational subject and accurately adjust the state of the learning subject.
[0055] This disclosure introduces an evaluation mechanism to quantify the state of the learning subject. This mechanism tracks the learning progress and understanding level of the learning subject through regular assessments and feedback, using indicators such as confidence index and mastery level. Machine learning algorithms can be used to analyze this data and predict changes in the adaptability and future performance of the learning subject. Based on the state of the learning subject, the educational subject can dynamically adjust the difficulty of the learning content and tasks to provide a personalized learning path. Through reinforcement learning algorithms, the educational subject can learn the best guiding strategies, thereby gradually guiding them to access more challenging content without affecting the confidence and learning motivation of the learning subject.
[0056] The data processing method provided in this embodiment obtains a new educational strategy based on the performance data; inputs the new educational strategy into the educational subject model to obtain a new action sequence output by the educational subject model; inputs the new action sequence and the teaching task data into the learning subject model to obtain a new learning strategy output by the learning subject model, and can dynamically adjust the teaching strategy according to the performance of the learning subject to adapt to the changing learning state of the learning subject over time. This design of real-time feedback enhances the interaction between the learning subject and the educational subject, resulting in more effective learning.
[0057] In some alternative implementations of the present disclosure, the above-mentioned performance data includes: learning style and learning level, mastered knowledge points, and feedback information. Based on the above-mentioned performance data, obtaining a new educational strategy includes at least one of the following: determining a learning plan strategy and a teaching content strategy based on the learning style and learning level, and using the learning plan strategy and the teaching content strategy as the new educational strategy; removing information related to the mastered knowledge points from the educational strategy to obtain a new educational strategy; generating learning suggestion information based on the feedback information, and using the learning suggestion information as the new educational strategy.
[0058] In this alternative implementation, the learning path includes a learning plan strategy and a teaching content strategy. The learning path can be optimized through performance data. For example, different learning plan strategies and teaching content strategies can be designed for learning subjects with different learning styles or different learning levels.
[0059] In this alternative implementation, adjusting the teaching content includes removing or adding information, and dynamically adjusting the teaching content and teaching difficulty according to the performance data of the learning subject. For example, if the learning subject has mastered a certain concept, the relevant content can be skipped and the next learning stage can be entered.
[0060] In this alternative implementation, the learning suggestion information includes personalized learning suggestions or strategies for solving learning problems, feedback, and guidance: The educational strategy can include content for providing feedback and guidance to the learning subject. For example, the educational subject can provide personalized learning suggestions or strategies for solving learning problems based on the performance data of the learning subject.
[0061] The method for obtaining a new educational strategy provided by this alternative implementation determines a learning plan strategy and a teaching content strategy based on the learning style and learning level, and uses the learning plan strategy and the teaching content strategy as the new educational strategy; removes information related to the mastered knowledge points from the educational strategy to obtain a new educational strategy; generates learning suggestion information based on the feedback information, and uses the learning suggestion information as the new educational strategy, which provides a reliable implementation for obtaining a variety of new educational strategies and improves the reliability of obtaining new educational strategies.
[0062] In some alternative implementations of the present disclosure, the above-mentioned learning subject model is constructed through the following steps: constructing a dynamic conversion model that represents the change of the subject state of the learning subject over time under the teaching task data; constructing a task selection model, which is used to represent the correspondence relationship between the action sequence, the learning task data, and the learning strategy under the current subject state, and the current subject state is provided by the dynamic conversion model; using the dynamic conversion model and the task selection model as the learning subject model, and optimizing the parameters of the learning subject model through the cost function of the learning subject model.
[0063] In this embodiment, the dynamic conversion model is responsible for describing the law of the subject state of the learning subject changing over time. It reflects the state evolution process of the learning subject under the influence of different learning tasks and teaching strategies. The task selection model, based on the dynamic conversion model, links the action sequence, learning task data, and the learning subject strategy, guiding the educational subject model to select the most appropriate teaching strategy and task to promote the positive change of the subject state of the learning subject and ultimately achieve the teaching goal. That is to say, the dynamic conversion model provides a theoretical basis for the state change of the learning subject for the task selection model, while the task selection model is the application of the dynamic conversion model in actual teaching.
[0064] In this alternative implementation, the cost function is a function used to measure the gap between the model prediction and the actual data. By calculating the sum of the squared differences between the predicted value and the target value and then dividing by two times the number of samples, the total cost is obtained. The cost function of the learning subject model includes: the cost function of the task selection model and the cost function of the dynamic conversion model. By calculating the actual values of the cost functions of both, the purpose of optimizing the parameters of the learning subject model can be achieved.
[0065] The method for constructing the learning subject model provided by the present disclosure constructs a dynamic conversion model that characterizes the change of the subject state of the learning subject over time under teaching task data; constructs a task selection model, which is used to characterize the correspondence between the action sequence, learning task data, and the learning strategy under the current subject state, and the current subject state is provided by the dynamic conversion model; takes the dynamic conversion model and the task selection model as the learning subject model, and optimizes the parameters of the learning subject model through the cost function of the learning subject model, providing a reliable implementation basis for the construction of the learning subject model in machine learning and improving the accuracy obtained by the learning subject model.
[0066] In some alternative implementations of the present disclosure, the above-mentioned construction of the dynamic conversion model that characterizes the change of the subject state of the learning subject over time under teaching task data includes: determining the first recognition inductive bias in the machine learning algorithm and the second recognition inductive bias of the learning subject; based on the first recognition inductive bias and the second recognition inductive bias, determining the subject state of the learning subject; constructing a dynamic conversion model that characterizes the change of the subject state over time, and optimizing the parameters of the dynamic conversion model through the cost function of the dynamic conversion model.
[0067] In this alternative implementation, the above-mentioned determination of the first recognition inductive bias in the machine learning algorithm and the second recognition inductive bias of the learning subject includes: identifying the inductive bias in the machine learning algorithm, such as model architecture, hyperparameters, initialization, etc., to obtain the first recognition inductive bias; identifying the inductive bias from the human learning subject, such as prior knowledge, meta-knowledge, learning style, etc., to obtain the second recognition inductive bias.
[0068] In this alternative implementation, determining the subject state of the learning subject based on the first recognition inductive bias and the second recognition inductive bias includes: abstracting the first recognition inductive bias and the second recognition inductive bias into parameters of the subject state of the learning subject, such as the weights of the modeling preference function, the parameters of the learning algorithm, etc.
[0069] In this alternative implementation, a dynamic transition model representing the change of the subject state over time is established according to the actions of the educational subject and the learning strategy. For example, the actions of the tutor may affect the modeling preference of the learning subject, and the variable recommendation action may affect the model selection of the learning subject.
[0070] In this alternative implementation, a Markov switching model can be used to describe the change of the subject state of the learning subject. The actions of the educational subject will cause the subject state of the learning subject to switch from "not understanding collinearity" to "understanding collinearity", and the variable recommendation action will affect the weights of the modeling preference function of the learning subject.
[0071] The method for constructing a dynamic transition model provided in this alternative implementation integrates the concept of inductive bias in machine learning and human learning, including model families, initialization, prior knowledge, and meta-knowledge, etc., providing a more comprehensive description of the subject state of the learning subject. This unity will optimize the teaching process and make it more targeted.
[0072] In some alternative implementations of the present disclosure, the above-mentioned subject state transformation model is constructed through the following steps: identifying at least two key factors affecting the educational subject and the learning subject; determining the quantization dimension indicators based on the key factors; constructing a state transition function based on the key factors and the quantization dimension indicators, where the state transition function is used to describe the probability of the learning subject transitioning from one state to another; setting the parameters of the state transition function based on the influence data of the educational subject on the learning subject to obtain the learning subject transition function; establishing a correspondence relationship between the learning subject transition function and the pre-constructed learning strategy to obtain the subject state transformation model, and optimizing the parameters of the subject state transformation model through the cost function of the subject state transformation model.
[0073] In this alternative implementation, the key factors are the factors affecting the cognition and behavior of the learning subject, such as prior knowledge, learning style, etc. The quantization dimension indicators are the dimensions selected from the key factors as the subject state and the indicators obtained after transforming the indicators of this dimension. The state transition function is a state transition rule.
[0074] In this alternative implementation, the specific construction steps of the subject state transformation model are shown as Step1 to Step4 below:
[0075] Step1: Define the subject state dimension.
[0076] Identify influencing factors: Analyze the factors that may affect the learning subject's cognition and behavior during the learning process, such as:
[0077] a Prior knowledge: The relevant knowledge and experience that the learning subject already has before learning a specific knowledge or skill.
[0078] b Learning style: The preferred learning method of the learning subject, such as visual, auditory, or kinesthetic learning.
[0079] c Motivation: The internal and external motivations of the learning subject to learn a specific knowledge or skill.
[0080] d Metacognition: The learning subject's ability to recognize and monitor their own learning process.
[0081] e Emotional state: The emotional reactions of the learning subject during the learning process, such as anxiety, interest, or frustration.
[0082] f Cognitive load: The cognitive burden borne by the learning subject during the learning process.
[0083] Select dimensions: Based on the characteristics of the learning task and the characteristics of the learning subject, select the most critical influencing factors as the dimensions of the subject state. For example, for mathematics learning, "mathematical knowledge level" and "learning style" can be selected as dimensions.
[0084] Quantify dimensions: Convert each dimension into quantifiable indicators. For example, use a scale or test to evaluate the prior knowledge level of the learning subject, or use observation methods to record the learning style of the learning subject.
[0085] Step2: Establish state transition rules.
[0086] Analyze the relationships between influencing factors: Study how different dimensions interact with each other. For example, the prior knowledge level will affect the learning motivation and cognitive load of the learning subject.
[0087] Establish a state transition function: Based on the relationships between influencing factors, establish a state transition function to describe the probability of the learning subject transitioning from one state to another. For example, if the learning subject has a thorough understanding of a certain concept, then they are more likely to enter the "understanding" state.
[0088] Consider the influence of the educator's behavior: The educator's behavior will affect the subject state of the learning subject. For example, the educator's encouragement can increase the learning motivation of the learning subject, and the educator's guidance can help the learning subject overcome cognitive obstacles. In the state transition function, the influence of the educator's behavior on the state transition probability needs to be considered.
[0089] Step3: Design learning strategies.
[0090] Determine teaching objectives: Clearly define the objectives that teaching hopes to achieve, such as improving the knowledge level of the learning subject, cultivating the learning interest of the learning subject, etc.
[0091] Select teaching activities: According to the current state of the learning subject and the teaching objectives, select appropriate teaching activities, such as explanations, demonstrations, exercises, discussions, etc.
[0092] Adjust learning strategies: Dynamically adjust learning strategies according to changes in the state of the learning subject. For example, if the cognitive load of the learning subject is too high, the educational subject can reduce the learning difficulty or provide more support.
[0093] Step4: Implementation and evaluation
[0094] Implement teaching strategies: Implement teaching strategies according to the design plan, and record the state changes of the learning subject and the teaching effects.
[0095] Evaluate teaching effects: Use the evaluation index system to evaluate teaching effects, such as the grades, learning attitudes, learning motivations, etc. of the learning subject.
[0096] Optimize teaching strategies: Optimize teaching strategies according to the evaluation results, such as adjusting teaching activities, improving teaching materials, etc.
[0097] The method for constructing the subject state conversion model provided by this optional implementation method first obtains the learning subject conversion function, then establishes the correspondence between the learning subject conversion function and the learning strategy, and finally optimizes the parameters of the subject state conversion model through the cost function, providing a reliable implementation basis for the construction of the subject state conversion model and improving the reliability of the subject state conversion model obtained.
[0098] A specific technical solution example of the subject state conversion model is as follows:
[0099] Learning task: Cultivate the mathematical problem-solving ability of the learning subject.
[0100] Subject state dimensions: Mathematical knowledge level, use a mathematical test to evaluate the mathematical knowledge level of the learning subject; Problem-solving strategies, observe the strategies used by the learning subject when solving problems, such as the trial-and-error method, the analytical method, etc.
[0101] State conversion rules: If the learning subject has a high level of mathematical knowledge, they are more likely to choose the analytical method to solve problems; if the learning subject has a low level of mathematical knowledge, they are more likely to choose the trial-and-error method to solve problems; the guidance of the educational subject can help the learning subject transition from the trial-and-error method to the analytical method.
[0102] Learning strategies: For learning subjects with a relatively high level of mathematical knowledge, the educational subject can provide more complex math problems and encourage the learning subject to use the analytical method to solve problems. For learning subjects with a relatively low level of mathematical knowledge, the educational subject can first provide some simple math problems, guide the learning subject to use the trial-and-error method to solve problems, and then gradually transition to the analytical method.
[0103] Evaluation metrics: The accuracy rate of the learning subject in solving math problems, the frequency of the learning subject using different problem-solving strategies, and the learning subject's interest in math learning.
[0104] In some alternative implementation manners of the present disclosure, the above-mentioned two-player game model is constructed through the following steps: determining the game elements including the educational subject and the learning subject; based on the game elements, defining the strategy spaces and the payoff functions of the educational subject and the learning subject; the strategy space of the educational subject includes: the subject state and the sub-model, and the sub-model is used to represent the learning subject's understanding and prediction of data and transform the corresponding subject state; the strategy space of the educational subject includes teaching strategies; based on the strategy spaces, constructing a two-player game model, which is used to represent the correspondence relationship among the subject state of the learning subject, the action sequence of the educational subject, and the sub-model of the learning subject under the teaching task data, and optimizing the parameters of the two-player game model through the cost function of the two-player game model.
[0105] In this alternative implementation manner, the steps for constructing the two-player game model are specifically shown as Step_1 to Step_5:
[0106] Step_1: Determine the game elements - the learning subject and the educational subject:
[0107] a The learning subject: Represents the learning subject, which can be a student, a machine learning algorithm, etc.
[0108] b The educational subject: Represents the educational subject, which can be an educational entity, an AI assistant, etc.
[0109] Learning task: Define the learning objective and the data set.
[0110] Model space: The set of models that the learning subject can choose from.
[0111] Subject state space: The set of the subject states of the learning subject, such as prior knowledge, learning strategies, etc.
[0112] Action space: The set of actions that the educational subject can perform, such as providing data, explaining concepts, etc.
[0113] State transition kernel: Describes how the learning subject's state changes according to the educational subject's actions and its own strategies.
[0114] Observation set: The set of information observable by the educational entity, such as the model selected by the learning entity, learning progress, etc.
[0115] Conditional observation probability: Describes the relationship between the information observed by the educational entity and the state of the learning entity.
[0116] Cost function: A function for evaluating teaching effectiveness, such as model error, manipulation level, etc.
[0117] Step_2: Define the entity state of the learning entity, which is used to describe the learning entity's preference for different models. This will affect its choice of the learning entity model and learning strategy, thus influencing the two-player game process between the educational entity and the learning entity.
[0118] Modeling preference: Define the preference function of the learning entity for the model, such as the weights of factors like model complexity and interpretability.
[0119] Learning algorithm: Define the algorithm used by the learning entity to construct the learning entity model, such as linear regression, neural network, etc.
[0120] Entity state evolution: Define how the entity state of the learning entity changes over time, such as through methods like Bayesian update and reinforcement learning.
[0121] Step_3: Construct the framework of the educational entity model, through which to describe how the educational entity observes the entity state of the learning entity, what educational strategies can be taken, and how to make action sequence decisions based on the observation results and its own educational strategies.
[0122] State space: The set of states observed by the educational entity, including the model selected by the learning entity and the estimation of the entity state of the learning entity.
[0123] Action space: The set of actions executable by the educational entity, such as providing data, explaining concepts, etc.
[0124] State transition kernel: Describes how the entity state of the learning entity changes according to the actions of the educational entity and its own educational strategies.
[0125] Observation set: The set of information observable by the educational entity, such as the learning entity model, learning progress, etc.
[0126] Conditional observation probability: Describes the relationship between the information observed by the educational entity and the entity state of the learning entity.
[0127] Cost function: A function for evaluating teaching effectiveness, such as model error, manipulation level, etc.
[0128] Step_4: Select a solution method to enable the educational entity to learn the optimal teaching strategy, thereby influencing the two-player game process between the educational entity and the learning entity.
[0129] Fully Observable Markov Decision Process (MDP): When the educational agent can fully observe the state of the learning agent, methods such as dynamic programming can be used to solve for the optimal teaching strategy.
[0130] Partially Observable Markov Decision Process (POMDP): When the educational agent cannot fully observe the state of the learning agent, approximate methods such as value iteration, Monte Carlo planning, etc. can be used.
[0131] Reinforcement Learning: The educational agent can learn the optimal teaching strategy through trial and error based on the feedback and reward signals of the learning agent.
[0132] Step_5: Implementation and Optimization. It can be achieved through code implementation and experimental verification to further optimize the educational agent model and learning strategy, making the game process between the educational agent and the learning agent more in line with the actual situation.
[0133] Code Implementation: Implement the game model and solution algorithm using a programming language.
[0134] Parameter Setting: Set the model parameters according to the specific task and the characteristics of the learning agent.
[0135] Experimental Verification: Evaluate the teaching effect through experiments and optimize the model.
[0136] Specific example for constructing a two-player game model: Variable selection task. Learning task: Select variables in a linear regression model. Model space: A d-dimensional binary vector space representing whether each variable is included. Agent state space: Modeling preference: Weights for factors such as variable correlation and collinearity. Learning algorithm: Linear regression algorithm. Agent state evolution: Through Bayesian update, update the modeling preference based on the information provided by the educational agent and the observations of the learning agent itself. Action space: Select a variable for interpretation or recommendation. State transition kernel: The learning agent decides whether to accept the suggestion of the educational agent according to the modeling preference. Observation set: The model selected by the learning agent. Conditional observation probability: The agent state can be estimated through the model selected by the learning agent. Cost function: Model error, manipulation level, action cost of the educational agent, etc.
[0137] Further reference Figure 2 , as an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a data processing device. This device embodiment corresponds to Figure 1 the method embodiment shown, and this device can be specifically applied to various electronic devices.
[0138] As Figure 2As shown in the figure, the data processing device 200 provided in this embodiment includes: a generation unit 201, an action obtaining unit 202, a strategy obtaining unit 203, a transformation unit 204, and an evaluation unit 205. Among them, the above-mentioned generation unit 201 can be configured to generate an education strategy based on a two-player game model and teaching task data, and the two-player game model is a model pre-constructed for an education subject and a learning subject. The above-mentioned action obtaining unit 202 can be configured to input the education strategy into an education subject model pre-constructed for the education subject to obtain an action sequence output by the education subject model, and the education subject model is used to represent the correspondence between the education strategy and the action sequence. The above-mentioned strategy obtaining unit 203 can be configured to input the action sequence and the teaching task data into a learning subject model pre-constructed for the learning subject to obtain a learning strategy output by the learning subject model, and the learning subject model is used to represent the correspondence between the action sequence and the teaching task data and the learning strategy. The above-mentioned transformation unit 204 can be configured to input the learning strategy into a pre-constructed subject state transformation model to obtain the current subject state output by the subject state transformation model, and the subject state transformation model is used to represent the correspondence between the subject state and the learning strategy. The above-mentioned evaluation unit 205 can be configured to evaluate the performance of the learning subject model on the teaching task data based on the current subject state to obtain performance data.
[0139] In this embodiment, in the data processing device 200: the specific processing of the generation unit 201, the action obtaining unit 202, the strategy obtaining unit 203, the transformation unit 204, and the evaluation unit 205 and the technical effects brought by them can respectively refer to Figure 1 The relevant descriptions of steps 101, 102, 103, 104, and 105 in the corresponding embodiment will not be elaborated here.
[0140] In some embodiments of the present disclosure, the above-mentioned data processing device 200 further includes: an adjustment unit (not shown in the figure), and the adjustment unit is configured to: obtain a new education strategy based on the performance data; input the new education strategy into the education subject model to obtain a new action sequence output by the education subject model; input the new action sequence and the teaching task data into the learning subject model to obtain a new learning strategy output by the learning subject model.
[0141] In some embodiments of the present disclosure, the above performance data includes: learning style and learning level, mastered knowledge points, and feedback information. The adjustment unit is further configured to at least one of the following: determine a learning plan strategy and a teaching content strategy based on the learning style and the learning level, and use the learning plan strategy and the teaching content strategy as new educational strategies; remove information related to the mastered knowledge points from the educational strategies to obtain new educational strategies; generate learning suggestion information based on the feedback information, and use the learning suggestion information as new educational strategies.
[0142] In some embodiments of the present disclosure, the learning subject model is constructed by a learning construction unit (not shown in the figure). The learning construction unit is configured to: construct a dynamic conversion model representing the change of the subject state of the learning subject over time under teaching task data; construct a task selection model, which is used to represent the correspondence between the action sequence, the learning task data and the learning strategy under the current subject state, and the current subject state is provided by the dynamic conversion model; use the dynamic conversion model and the task selection model as the learning subject model, and optimize the parameters of the learning subject model through the cost function of the learning subject model.
[0143] In some embodiments of the present disclosure, the learning construction unit is further configured to: determine the first recognition inductive bias in the machine learning algorithm and the second recognition inductive bias of the learning subject; determine the subject state of the learning subject based on the first recognition inductive bias and the second recognition inductive bias; construct a dynamic conversion model representing the change of the subject state over time, and optimize the parameters of the dynamic conversion model through the cost function of the dynamic conversion model.
[0144] In some embodiments of the present disclosure, the subject state conversion model is constructed by an internal construction unit (not shown in the figure). The internal construction unit is configured to: identify at least two key factors affecting the educational subject and the learning subject; determine the quantization dimension index based on the key factors; construct a state transition function based on the key factors and the quantization dimension index, and the state transition function is used to describe the probability of the learning subject transferring from one state to another state; perform parameter setting on the state transition function based on the influence data of the educational subject on the learning subject to obtain a learning subject transition function; establish a correspondence between the learning subject transition function and a pre-constructed learning strategy to obtain a subject state conversion model, and optimize the parameters of the subject state conversion model through the cost function of the subject state conversion model.
[0145] In some embodiments of the present disclosure, the two-player game model is constructed by a game construction unit (not shown in the figure). The game construction unit is configured to: determine game elements including an educational entity and a learning entity; define the strategy spaces of the educational entity and the learning entity based on the game elements; the strategy space of the educational entity includes: the entity state and a sub-model, where the sub-model is used to represent the learning entity's understanding and prediction of data and transform the corresponding entity state; the strategy space of the educational entity includes teaching strategies; construct a two-player game model based on the strategy spaces, where the two-player game model is used to represent the correspondence between the entity state of the learning entity, the action sequence of the educational entity, and the sub-model of the learning entity under teaching task data, and optimize the parameters of the two-player game model through the cost function of the two-player game model.
[0146] The data processing device provided by the embodiments of the present disclosure, first, the generation unit 201 generates an educational strategy based on the two-player game model and teaching task data, and the two-player game model is a model pre-constructed for the educational entity and the learning entity; the action obtaining unit 202 inputs the educational strategy into the educational entity model pre-constructed for the educational entity to obtain the action sequence output by the educational entity model, and the educational entity model is used to represent the correspondence between the educational strategy and the action sequence; the strategy obtaining unit 203 inputs the action sequence and the teaching task data into the learning entity model pre-constructed for the learning entity to obtain the learning strategy output by the learning entity model, and the learning entity model is used to represent the correspondence between the action sequence and the teaching task data and the learning strategy; the transformation unit 204 is configured to input the learning strategy into the pre-constructed entity state transformation model to obtain the current entity state output by the entity state transformation model, and the entity state transformation model is used to represent the correspondence between the learning strategy entity state and the learning strategy; the evaluation unit 205 evaluates the performance of the learning entity model on the teaching task data based on the current entity state to obtain performance data. Thus, through the two-player game model, the entity state of the learning entity can change dynamically with the change of the action sequence of the educational entity, strengthening the adjustment and evolution of the entity state of the learning entity, helping the learning entity better adapt to the learning task, and improving the learning personalization and adaptability of the learning entity.
[0147] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0148] Figure 3FIG. 0 shows a schematic block diagram of an example electronic device 300 that may be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their manners of operation are meant to be examples only and are not intended to limit the implementations of the present disclosure described and / or claimed herein.
[0149] As Figure 3 shown, the device 300 includes a computing unit 301 that may perform various appropriate actions and processes in accordance with a computer program stored in a read only memory (ROM) 302 or a computer program loaded from a storage unit 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the device 300 may also be stored. The computing unit 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0150] A plurality of components in the device 300 are connected to the I / O interface 305, including: an input unit 306, such as, for example, a keyboard, a mouse, etc.; an output unit 307, such as, for example, various types of displays, speakers, etc.; a storage unit 308, such as, for example, a magnetic disk, an optical disk, etc.; and a communication unit 309, such as, for example, a network card, a modem, a wireless communication transceiver, etc. The communication unit 309 allows the device 300 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0151] The computing unit 301 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 301 executes the various methods and processes described above, such as the data processing method. For example, in some embodiments, the data processing method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 308. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 300 via the ROM 302 and / or the communication unit 309. When the computer program is loaded into the RAM 303 and executed by the computing unit 301, one or more steps of the data processing method described above can be executed. Alternatively, in other embodiments, the computing unit 301 can be configured to execute the data processing method in any other suitable manner (e.g., by means of firmware).
[0152] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-a-chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0153] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program code is executed by the processor or controller, the patterns / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0154] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0155] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0156] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0157] It should be understood that the various forms of the processes shown above can be reordered, added, or removed steps. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0158] The foregoing description of specific exemplary embodiments of the present disclosure is for purposes of illustration and exemplification. These descriptions are not intended to limit the present disclosure to the precise forms disclosed, and it is apparent that many changes and variations are possible in light of the above teachings. The purpose of selecting and describing the exemplary embodiments is to explain the specific principles of the present disclosure and its practical applications, so that those skilled in the art can implement and utilize the various different exemplary embodiments of the present disclosure, as well as various different selections and changes. The scope of the present disclosure is intended to be defined by the claims and their equivalents.
Claims
1. A data processing method, the method comprising: Generate an educational strategy based on a two-person game model and teaching task data. The two-person game model is a mathematical model pre-constructed for an educational subject and a learning subject, and is used to describe the interaction between the educational subject and the learning subject. Through the two-person game model, the actual game process between the learning subject and the educational subject is a process of adjusting parameters between the educational subject model and the learning subject model. The educational subject model guides the learning subject to learn through action sequences and educational strategies, and the learning subject model adjusts the learning strategy according to the guidance of the educational subject and its own learning status, and finally achieves the learning goal. Inputting the educational strategy into an educational subject model pre-constructed for the educational subject to obtain an action sequence output by the educational subject model, wherein the educational subject model is used to characterize the corresponding relationship between the educational strategy and the action sequence; Inputting the action sequence and the teaching task data into a learning subject model pre-constructed for the learning subject, and obtaining a learning strategy output by the learning subject model, wherein the learning subject model is used to characterize the corresponding relationship between the action sequence and the teaching task data and the learning strategy; Inputting the learning strategy into a pre-built subject state transformation model to obtain a current subject state output by the subject state transformation model, wherein the subject state transformation model is used to characterize the corresponding relationship between the subject state and the learning strategy; Based on the current subject state, the performance of the learning subject model on the teaching task data is evaluated to obtain performance data.
2. The method according to claim 1, further comprising: deriving new educational strategies based on the performance data; Inputting the new educational strategy into the educational subject model to obtain a new action sequence output by the educational subject model; The new action sequence and the teaching task data are input into the learning subject model to obtain a new learning strategy output by the learning subject model.
3. The method according to claim 2, wherein: The performance data includes: learning style and learning level, mastered knowledge points and feedback information. The new education strategy based on the performance data includes at least one of the following: Determine a learning plan strategy and a teaching content strategy based on the learning style and the learning level, and use the learning plan strategy and the teaching content strategy as a new education strategy; Removing information related to the mastered knowledge points from the educational strategy to obtain a new educational strategy; Based on the feedback information, learning suggestion information is generated, and the learning suggestion information is used as a new educational strategy.
4. The method according to claim 1, wherein: The learning subject model is constructed by the following steps: Construct a dynamic transformation model to characterize the change of the learning subject's state over time under the teaching task data; Constructing a task selection model, wherein the task selection model is used to characterize the correspondence between the action sequence, the learning task data and the learning strategy under the current subject state, wherein the current subject state is provided by the dynamic conversion model; The dynamic conversion model and the task selection model are used as learning subject models, and the parameters of the learning subject model are optimized through the cost function of the learning subject model.
5. The method according to claim 4, wherein: The construction of a dynamic conversion model for representing the change of the subject state of the learning subject over time under the teaching task data includes: Identify first-identification inductive biases in machine learning algorithms and second-identification inductive biases in learning agents; determining a subject state of the learning subject based on the first identification inductive bias and the second identification inductive bias; A dynamic conversion model is constructed to characterize the change of the subject state over time, and the parameters of the dynamic conversion model are optimized through the cost function of the dynamic conversion model.
6. The method according to claim 1, wherein: The subject state transformation model is constructed by the following steps: Identify at least two key factors that influence the subject of education and the subject of learning; Based on the key factors, determine the quantitative dimension indicators; Based on the key factors and the quantitative dimension indicators, a state transition function is constructed, where the state transition function is used to describe the probability of a learning subject transferring from one state to another state; Based on the influence data of the education subject on the learning subject, setting parameters of the state transition function to obtain a learning subject transition function; A corresponding relationship between the learning subject conversion function and a pre-constructed learning strategy is established to obtain a subject state conversion model, and the parameters of the subject state conversion model are optimized through the cost function of the subject state conversion model.
7. The method according to claim 1, wherein: The two-player game model is constructed through the following steps: Determine the game elements including the education subject and the learning subject; Based on the game elements, define the strategy space of the education subject and the learning subject; The strategy space of the education subject includes: subject state and sub-model, the sub-model is used to characterize the understanding and prediction of the learning subject on the data and convert the corresponding subject state; the strategy space of the education subject includes teaching strategy; Based on the strategy space, a two-player game model is constructed. The two-player game model is used to characterize the correspondence between the subject state of the learning subject, the action sequence of the teaching subject and the sub-model of the learning subject under the teaching task data, and the parameters of the two-player game model are optimized through the cost function of the two-player game model.
8. A data processing device, comprising: A generating unit is configured to generate an educational strategy based on a two-person game model and teaching task data, wherein the two-person game model is a mathematical model pre-constructed for an educational subject and a learning subject, and is used to describe the interaction between the educational subject and the learning subject. Through the two-person game model, the actual game process between the learning subject and the educational subject is a process of adjusting parameters between the educational subject model and the learning subject model. The educational subject model guides the learning subject to learn through an action sequence and an educational strategy, and the learning subject model adjusts the learning strategy according to the guidance of the educational subject and its own learning state, and finally achieves the learning goal; an action obtaining unit, configured to input the educational strategy into an educational subject model pre-constructed for the educational subject, and obtain an action sequence output by the educational subject model, wherein the educational subject model is used to characterize the corresponding relationship between the educational strategy and the action sequence; A strategy obtaining unit is configured to input the action sequence and the teaching task data into a learning subject model pre-constructed for the learning subject, and obtain a learning strategy output by the learning subject model, wherein the learning subject model is used to characterize the corresponding relationship between the action sequence and the teaching task data and the learning strategy; a conversion unit configured to input the learning strategy into a pre-built subject state conversion model to obtain a current subject state output by the subject state conversion model, wherein the subject state conversion model is used to characterize the corresponding relationship between the subject state and the learning strategy; The evaluation unit is configured to evaluate the performance of the learning subject model on the teaching task data based on the current subject state to obtain performance data.
9. An electronic device, characterized in that: include: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable the computer to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Machine learning system for a training model of an adaptive trainer
US10552764B1