Cognitive Diagnosis Method, System, Device and Medium for Sparse Data Scenarios
By employing Bayesian hierarchical modeling and data-driven meta-learning, the method addresses the challenge of sparse data in smart education systems, enabling accurate and interpretable cognitive state estimation for new learners, thus improving personalized educational services.
Patent Information
- Application Number
- CN202310123264.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-06
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-02-06
AI Technical Summary
Existing cognitive diagnosis methods in smart education systems fail to accurately assess the cognitive states of new learners in sparse data scenarios, lacking both diagnostic accuracy and interpretability of uncertainty, which hampers personalized educational services.
A method utilizing Bayesian hierarchical modeling and data-driven meta-learning to construct a global prior distribution from historical learner data, optimizing meta-parameters to predict cognitive states in new learners using a variational inference approach.
Enables accurate cognitive state estimation in sparse data conditions, providing reliable and interpretable results for personalized educational services such as resource recommendation and path planning.
Smart Images

Figure CN116051329B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent education technology, and in particular, to a cognitive diagnosis method, system, device and medium for sparse data scenarios. Background Art
[0002] In an intelligent education system, learners will generate learning records during the interaction process. Evaluating and diagnosing the cognitive state of learners (such as knowledge mastery level, ability level, etc.) based on the learning records is a basic and key task in intelligent education (hereinafter referred to as the "cognitive diagnosis" task).
[0003] When the interaction data of learners is sufficient, existing cognitive diagnosis methods can accurately evaluate the cognitive state of learners, and then help the intelligent education system better provide personalized services suitable for the learning situation for learners, such as learning resource recommendation, learning path planning, etc.
[0004] However, in reality, intelligent education systems must handle sparse data scenarios in many cases, such as the cold start scenario for new learners. However, existing cognitive diagnosis methods not only cannot achieve sufficient diagnostic accuracy in sparse data scenarios, but also do not have interpretability for the uncertainty of diagnostic results. In other words, in the current intelligent education system, when the interaction data of students is sparse, not only can accurate cognitive diagnosis results not be obtained, but even the accuracy and reliability of the results cannot be perceived, which is harmful to subsequent educational services (learning resource recommendation, learning path planning, etc.). Summary of the Invention
[0005] The purpose of the present invention is to provide a cognitive diagnosis method, system, device and medium for sparse data scenarios, which can accurately infer the cognitive state of new learners, and then accurately provide subsequent educational services (learning resource recommendation, learning path planning, etc.) for new users.
[0006] The purpose of the present invention is achieved through the following technical solutions:
[0007] A cognitive diagnosis method for sparse data scenarios includes:
[0008] Collecting the historical answering records of multiple learners and dividing them into a training set and a validation set; the historical answering records of learners include: question information and corresponding answering results;
[0009] Model the cognitive state of learners using a training set, construct a Bayesian hierarchical model to establish the relationship between the cognitive state of learners and the global prior distribution shared by all learners, introduce data-driven meta-learning techniques, use the relevant parameters in the global prior distribution as meta-parameters, based on the Bayesian hierarchical model, use the training set to infer the posterior distribution of the learners' cognitive state, and predict the learners' answer results on a validation set. Construct a loss function by combining the answer results included in the validation set and the predicted answer results, optimize the meta-parameters, and obtain an optimized global prior distribution;
[0010] For new learners, use the historical answer records and the optimized global prior distribution to infer the posterior distribution of the new learners' cognitive state, and the inferred posterior distribution of the cognitive state can be used to predict the answer results of the new learners on each question.
[0011] A cognitive diagnosis system for sparse data scenarios, including:
[0012] A data collection unit for collecting the historical answer records of multiple learners and dividing them into a training set and a validation set; the historical answer records of learners include: question information and corresponding answer results;
[0013] A model establishment and training unit for modeling the cognitive state of learners using a training set, constructing a Bayesian hierarchical model to establish the relationship between the cognitive state of learners and the global prior distribution shared by all learners, introducing data-driven meta-learning techniques, using the relevant parameters in the global prior distribution as meta-parameters, based on the Bayesian hierarchical model, using the training set to infer the posterior distribution of the learners' cognitive state, and predicting the learners' answer results on a validation set. Construct a loss function by combining the answer results included in the validation set and the predicted answer results, optimize the meta-parameters, and obtain an optimized global prior distribution;
[0014] A speculation unit for, for new learners, using the historical answer records and the optimized global prior distribution to infer the posterior distribution of the new learners' cognitive state, and the inferred posterior distribution of the cognitive state can be used to predict the answer results of the new learners on each question.
[0015] A processing device, including: one or more processors; a memory for storing one or more programs;
[0016] Wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the foregoing method.
[0017] A readable storage medium storing a computer program, which implements the foregoing method when the computer program is executed by a processor.
[0018] As can be seen from the technical solution provided by the present invention above, the cognitive state of a new learner can be accurately evaluated based on the limited learning records of the new learner. Specifically, through Bayesian hierarchical modeling and data-driven meta-learning techniques, a cognitive diagnosis result containing probability distribution information (i.e., the posterior distribution mentioned later) and meta-knowledge information (i.e., the global prior distribution ψ mentioned later) can be obtained, which can provide certain technical support for the practice of intelligent education systems in sparse data scenarios such as user cold start, and further help intelligent education systems better provide personalized services suitable for the learning situation of learners. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1 Flowchart of a cognitive diagnosis method for sparse data scenarios provided by an embodiment of the present invention;
[0021] Figure 2 Description diagram of the Bayesian hierarchical model of the learner's cognitive state provided by an embodiment of the present invention;
[0022] Figure 3 Description diagram of updating the global prior distribution in a data-driven manner provided by an embodiment of the present invention;
[0023] Figure 4 Schematic diagram of a cognitive diagnosis system for sparse data scenarios provided by an embodiment of the present invention;
[0024] Figure 5 Schematic diagram of a processing device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0026] First, the following explanations are given for the terms that may be used in this article:
[0027] Descriptions using terms such as "including", "comprising", "containing", "having" or other similar semantics should be construed as non-exclusive inclusion. For example, including a technical feature element (such as raw materials, components, ingredients, carriers, dosage forms, materials, dimensions, parts, components, mechanisms, devices, steps, processes, methods, reaction conditions, processing conditions, parameters, algorithms, signals, data, products or articles, etc.) should be construed as not only including the specifically listed technical feature element, but also including other technical feature elements known in the art that are not specifically listed.
[0028] The following provides a detailed description of a cognitive diagnosis method, system, device, and medium for sparse data scenarios provided by the present invention. Content not described in detail in the embodiments of the present invention belongs to the prior art well-known to those skilled in the art. Conditions not specified in the embodiments of the present invention are carried out according to conventional conditions in the art or conditions recommended by the manufacturer.
[0029] Embodiment 1
[0030] The embodiment of the present invention provides a cognitive diagnosis method for sparse data scenarios, as Figure 1 shown, which mainly includes:
[0031] Step 1: Collect the historical answering records of multiple learners and divide them into a training set and a validation set; the historical answering records of the learners include: question information and corresponding answering results.
[0032] Step 2: Use the training set to model the cognitive state of the learners, construct a Bayesian hierarchical model to establish the relationship between the cognitive state of the learners and the global prior distribution shared by all learners, introduce data-driven meta-learning technology, take the relevant parameters in the global prior distribution as meta-parameters, based on the Bayesian hierarchical model, use the training set to infer the posterior distribution of the learners' cognitive state, and predict the answering results of the learners on the validation set. Combine the answering results included in the validation set with the predicted answering results to construct a loss function, optimize the meta-parameters, and obtain an optimized global prior distribution.
[0033] Step 3: For new learners, use the historical answering records and the optimized global prior distribution to calculate the posterior distribution of the cognitive state of the new learners, and the calculated posterior distribution of the cognitive state can be used to predict the answering results of the new learners on each question.
[0034] In addition, during the processes of the above Step 2 and Step 3, a method for inferring the posterior distribution of the cognitive state based on variational inference is also designed. The variational distribution is used to approximate the posterior distribution, and by minimizing the distance between the variational distribution and the posterior distribution, a gradient solving algorithm for the variational distribution is derived to improve the operation efficiency.
[0035] The above solution provided by the embodiments of the present invention uses Bayesian hierarchical modeling and data-driven meta-learning techniques to obtain a cognitive diagnosis result containing probability distribution information and meta-knowledge information, which can provide certain technical support for the practice of intelligent education systems in sparse data scenarios such as user cold start, and further help intelligent education systems better provide personalized services suitable for the learning situation for learners.
[0036] To more clearly show the technical solution provided by the present invention and the resulting technical effects, the following uses specific embodiments to describe in detail a cognitive diagnosis method for sparse data scenarios provided by the embodiments of the present invention.
[0037] I. Solution Overview.
[0038] The present invention provides a cognitive diagnosis method for sparse data scenarios, which uses a Bayesian meta-learned cognitive diagnosis framework (BETA-CD). Generally speaking, BETA-CD is based on meta-learning technology and Bayesian modeling technology in the field of machine learning, aiming to input as little learner historical interaction data as possible and output the probability distribution of the learner's cognitive state, so that the expected value of this distribution is as close as possible to the true value of the cognitive state, and the uncertainty of this distribution (such as variance, entropy, etc.) can reflect the uncertainty of the diagnosis result itself as much as possible. Specifically, BETA-CD consists of two parts. (1) A Bayesian hierarchical model for the student's cognitive diagnosis state. BETA-CD believes that there is a prior distribution for the learner's cognitive state, and this prior distribution contains the meta-knowledge of the entire learner population. According to the prior distribution and the learner's learning data, we can infer the posterior distribution of its cognitive state according to the Bayesian rule as its cognitive diagnosis result. (2) Data-driven Bayesian prior probability learning and variational posterior probability inference. Different from most existing Bayesian-based cognitive diagnosis methods, the meta-knowledge in BETA-CD (that is, the prior distribution of the cognitive state) is not manually specified, but in a data-driven mode, automatically learned through meta-learning; at the same time, in order to improve the training and inference speed to adapt to application scenarios with a large number of learners, BETA-CD uses variational inference to approximately infer the posterior distribution of each learner individual.
[0039] Finally, BETA-CD outputs the inferred posterior distribution as the cognitive diagnosis result of the learner. The output of the cognitive diagnosis result is the cognitive state of the learner, which can reflect the learner's knowledge mastery, ability level, etc. The intelligent education system can use the expected value of this distribution as a reference for judging the learner's cognitive state, and can use the uncertainty of this distribution as a reference for the reliability of this result. At the same time, it can better provide personalized services (learning resource recommendation, learning path planning, etc.) suitable for the learning situation for the learner.
[0040] II. Problem Definition and Formalization.
[0041] Suppose there are M learners in an intelligent education system and N questions Each learner interacts with the system by completing questions, and the result of answering the questions is recorded by the system in the form of a triple (s i , q j , r ij ), where r ij ∈{1, 0} indicates that the learner s i correctly (or incorrectly) answers the question q j . The set of all questions completed by the learner s i and their answering results are denoted as Q i and R i respectively. According to these learning data, a cognitive model can be used to model the cognitive state of the student to obtain the cognitive state of the corresponding learner. Abstractly speaking, a cognitive model contains two sets of parameters: the set of parameters θ representing the personalized cognitive state of the learner, and the parameters question features Φ, such as question difficulty, knowledge points, etc., are generally obtained through pre-training in all relevant learning data, or can be directly obtained through expert manual verification. Given the cognitive state parameters θ i of the learner s i , the cognitive model can predict its answering result p(r j =1|q ij , θ j ) for any question q i ∈Q.
[0042] The problem definition of the present invention is as follows: Given an intelligent education system, a cognitive model with formalized parameters θ, learners and the corresponding learning data For any new learner the goal is to obtain an estimate of its personalized cognitive state parameters θ * from its small number of learning data samples.
[0043] II. Data Collection and Preprocessing.
[0044] 1. Data collection.
[0045] In the embodiment of the present invention, the test data recorded when the learner interacts with the intelligent education system is used as the input data set, and the data includes the corresponding question information and answer results of the learner. Such data samples include open source data sets (ASSISTment), etc. In addition, the input data set can also be obtained by web crawling, providing support from the education platform, or collecting the homework or examination information of junior and senior high school students offline.
[0046] 2. Data preprocessing.
[0047] Before building a model, the collected data needs to be preprocessed to ensure the effectiveness of the model. Preprocessing mainly includes the following:
[0048] 1) Data filtering.
[0049] To ensure the stability of training and the reliability of experimental results, learners who answered less than 20 questions and questions with less than 10 answers in the dataset were filtered out.
[0050] 2) Data division.
[0051] In each data set, learners are randomly divided into training set, validation set and test set according to a certain ratio. For example, the ratio of the above three data sets can be 6:2:2. The test set is used to test the effect of the model. After the test passes, predictions are made for new learners.
[0052] 3. Model establishment and training.
[0053] The embodiment of the present invention establishes a Bayesian meta-learning cognitive diagnosis framework, which is introduced from three aspects: the Bayesian hierarchical model of cognitive state, data-driven meta-knowledge learning based on meta-learning, and variational posterior inference of cognitive state distribution.
[0054] 1. Bayesian hierarchical model of cognitive states.
[0055] In the embodiment of the present invention, a Bayesian hierarchical model is constructed for the cognitive state of the learner. For each learner, two distributions of his cognitive state are defined: (1) Global prior distribution: the distribution shared by all learners, representing the meta-knowledge of the entire learner group, such as the overall ability level distribution. (2) Posterior distribution: the personalized state distribution unique to each learner, inferred from the prior distribution and the learner's historical learning data, representing the diagnosis result of the individual learner.
[0056] Figure 2A Bayesian hierarchical model description diagram showing the cognitive state of learners, including: the global prior distribution ψ shared by all learners and the personalized variable θ of learner s i (i = 1,..., M). Specifically, θ i represents the cognitive state of learner s i The global prior distribution ψ defines a parameterized prior distribution p(θ i |ψ) for the cognitive state of learner s, which can be regarded as a representation of the prior knowledge about the cognitive state of learners. For example, assuming that the cognitive state of learners follows a Gaussian distribution, ψ contains the mean vector and covariance matrix used to parameterize the Gaussian distribution. The specific form of the personalized variable θ i depends on the cognitive model adopted, and then θ i can be substituted into the model to predict the response result p(r i = 1|q i , θ i ) of the learner on any question q j . ij Part (a) of shows the relationship between ψ and θ j , and the response result can be predicted using θ i ; Figure 2 Part (b) of shows that after calculating the posterior distribution of the cognitive state using ψ and the historical response records in the training set, that is, the i mentioned later i Then, the effect of predicting the response result using θ Figure 2 can be verified through the validation set . i
[0057] After completing meta-knowledge learning through the method introduced later, for a new learner s * , combined with the optimized global prior distribution obtained from meta-knowledge learning and the historical response records (Q * , R * ) of this learner, the posterior distribution of its cognitive state θ * can be inferred through Bayes' rule:
[0058]
[0059] where Q * is the question information, R * is the corresponding response result, θ * represents the cognitive state, which is a random variable, represents the optimized global prior distribution is the parameterized prior distribution defined for the cognitive state θ * ; Denote the given question information Q * , the cognitive state θ * and the optimized global prior distribution , infer the answering result R * of the probability distribution; Denote the historical answering records of the given learner (Q * , R * ) and the optimized global prior distribution , infer the cognitive state θ * of the probability distribution; Denote that the cognitive state θ * satisfies the posterior distribution
[0060] Taking this posterior distribution as the result of the new learner cognitive modeling to replace the traditional deterministic result has two advantages. First, by introducing a prior distribution that includes external prior knowledge, the overfitting problem caused by limited data quantity can be alleviated. Second, the cognitive modeling in probability form contains the uncertainty information of the modeling process itself, which can provide better interpretability and risk resistance for downstream educational services. Specifically: Since the cognitive modeling in probability form contains the uncertainty of the modeling process itself, when the cognitive diagnosis result is indeed inaccurate, downstream educational services can perceive the uncertainty contained therein and thus decide whether to trust and adopt the result by themselves. Therefore, downstream educational services have better interpretability of cognitive diagnosis results and risk resistance in service scenarios.
[0061] 2. Meta-learning data-driven meta-knowledge learning.
[0062] Generally speaking, first define the relevant parameters in the global prior distribution (Global Prior) as meta-parameters (Meta-parameter). Infer the parameterized posterior distribution of the learner from the prior distribution based on the training set, and this parameter is called the local parameter (Local parameter). Then take the expected value of the posterior distribution as the cognitive diagnosis result, predict the answering behavior on the validation set, and calculate the loss function. Finally, aiming to minimize the loss function, update the meta-parameters by the gradient descent method until convergence. Finally, output the distribution corresponding to the meta-parameters as the optimized global prior distribution (UpdatedGlobal Prior). Specifically:
[0063] After obtaining the Bayesian hierarchical model of the learner's cognitive state, it is necessary to find a suitable global prior distribution. Different from the traditional method of artificially specifying the prior distribution, in the embodiments of the present invention, prior knowledge is automatically mined from the learning data accumulated by a large number of historical users in the intelligent education system, such as Figure 3As shown. Generally speaking, the relevant parameters of the global prior distribution are used as meta-parameters to construct a data-driven meta-learning objective. For the learner s i ∈S, formalize its cognitive modeling process as a "task", denoted as where respectively represent a training set / validation set partition of Q i , respectively represent a training set / validation set partition of R i . As the cognitive modeling processes for different learners, these tasks have an inherent similar structure, so there is potential meta-knowledge that can be learned. This meta-knowledge can help the cognitive model to adapt to each learner more quickly in a personalized manner, that is, infer the posterior distribution of the cognitive state of a learner s i such that its prediction on the validation set is good .
[0064] Among them, based on the Bayesian hierarchical model, use the training set to infer the posterior distribution of the learner's cognitive state denoted as:
[0065]
[0066] where is the historical answer record corresponding to the learner s i in the training set, represents the question information corresponding to the learner s i in the training set, represents the answer result corresponding to the learner s i in the training set; represents the probability distribution of inferring the answer result i corresponding to the learner s i under the condition of the cognitive state θ i , the question information corresponding to the learner s i in the training set and the global prior distribution ψ; represents the probability distribution of inferring the cognitive state θ of the learner s i corresponding to the historical answer record in the training set and the global prior distribution ψ, that is, the posterior distribution of the cognitive state θ i of the learner s i . i That is, the posterior distribution of the cognitive state θ i of the learner s
[0067] Then, combine the answer results included in the validation set with the predicted answer results to construct a loss function and optimize the global prior distribution. The loss function is denoted as:
[0068]
[0069] Among them, M represents the number of learners, represents the loss function of learner s i of represents the loss functions of all learners, represents the historical answer records corresponding to learner s in the given training set, and the question information corresponding to learner s in the validation set i i corresponding question information Under the condition of i and the global prior distribution ψ, infer the probability distribution of the answer results corresponding to learner s in the validation set; represents the question information corresponding to learner s in the given validation set i and the cognitive state θ of learner s i i Under the condition of i infer the probability distribution of the answer results corresponding to learner s in the validation set; E represents the expectation; represents θ i satisfying the posterior distribution of the cognitive state θ of learner s i i
[0070] By minimizing the above loss function: optimize the meta-parameters to obtain an optimized global prior distribution
[0071] Since the meta-parameters involved in the above meta-knowledge learning process are the parameters in the global prior distribution, and the purpose of optimizing the global prior distribution is achieved by optimizing the meta-parameters, therefore, the symbol of the meta-parameters is omitted and the symbol of the global prior distribution is directly used.
[0072] For the following two reasons, it can be assumed that the global prior distribution ψ is deterministic and does not have randomness: First, in a general intelligent education system, the number of historical users is much larger than the number of data samples of each learner, so it can be considered that the uncertainty of estimating the global prior distribution ψ is small; Second, in downstream personalized education services, compared with the personalized cognitive state θ of learners, ψ, as a latent variable, has less practical significance of uncertainty.
[0073] 3. Gradient-based variational inference.
[0074] Based on the meta-learning objective proposed above, in principle, the optimal global prior distribution can be learned. However, when the dimension of the cognitive state parameters is high, the posterior distribution It is likely to become uncomputable, and the posterior distribution has an excessive computational cost in the actual process. To solve this problem, in the embodiments of the present invention, a posterior distribution inference method for cognitive states based on variational inference is used, and the variational distribution obtained by the inference for the learner is used as the posterior distribution of the learner's cognitive state.
[0075] Specifically: during the meta-parameter optimization process, for learner s i a variational distribution q(θ i ; λ i ) is used as its corresponding posterior distribution of the cognitive state During the calculation process of the cognitive state of a new learner, the variational distribution q(θ * ; λ * ) is used as its corresponding posterior distribution of the cognitive state where λ i and λ * are both variational parameters in the corresponding variational distribution, and they have the same structure and meaning as the global prior distribution ψ; the corresponding variational parameters are obtained by minimizing the KL divergence between the variational distribution and the corresponding posterior distribution of the cognitive state, and the variational parameters are optimized through the local loss.
[0076] Taking the variational parameter λ i as an example, its calculation process is introduced, and it is related to ψ.
[0077] In variational inference, the variational parameter λ i ; λ i ) and the corresponding posterior distribution of the cognitive state is obtained by minimizing the KL divergence between them, expressed as: i , expressed as:
[0078]
[0079] where the cognitive state θ i of learner s i is sampled from the variational distribution corresponding to the variational parameter λ i , denotes denotes the probability distribution of inferring the answering result i corresponding to learner s i in the training set, given the cognitive state θ i of learner s and the question information i corresponding to learner s in the training set, and the global prior distribution ψ; Denote the learner s in the given training set i The corresponding question information Given the learner s i Cognitive state θ i Under the condition of, infer the learner s in the training set i The corresponding answer result The probability distribution of; p(θ i |ψ) represents that the global prior distribution ψ is the cognitive state θ of the learner s i A parameterized prior distribution defined; E represents the expectation; θ i ~q(θ i ; λ i ) means that θ i Satisfies the variational distribution q(θ i ; λ i ) i )
[0080] Intuitively, the meaning of the first term in the formula is to maximize the expected log-likelihood; the second term is the regularization term to prevent the gap between the approximate posterior and the prior distribution from being too large; the last term is a constant term for the variational parameter λ i And can be ignored during the optimization process. Based on this, define a variational distribution q(θ i ; λ i ) for the learner s i ) and minimize the local loss:
[0081]
[0082] Among them, η is an empirical weight parameter, which is beneficial to the stability of training; Represents minimizing the local loss
[0083] Minimize the local loss by gradient descent and optimize λ i :
[0084]
[0085] Among them, Represents the k-step stochastic gradient descent operation, α is the learning rate, and ← is the assignment symbol. The expected term during the training process is calculated by Monte Carlo sampling. This gradient-based variational inference method has two advantages. First, compared with directly inferring the posterior distribution through the Bayesian rule originally, this method is far more computationally feasible and efficient. Second, this method establishes an association between the prior distribution and the posterior distribution through the gradient, and further calculates the meta-learning loss function through the posterior distribution, so as to optimize the meta-learning objective using the gradient descent method
[0086] λ iAfter optimization, the optimized variational distribution q(θ i ; λ i ) is obtained and used as the posterior distribution of the corresponding cognitive state for the calculation of the loss function . Finally, the optimized global prior distribution variational distribution q(θ * ; λ * ) is obtained. The optimization process of the variational distribution q(θ i ; λ i ) is the same as that of the variational distribution q(θ i ; λ i ), so it will not be elaborated here.
[0087] In the above calculation process, the main purpose is to calculate the variational parameter λ i . The practical meaning of the variational parameter λ i is to determine the parameters of a certain distribution. For example, when the variational distribution is a normal distribution, the variational parameter is the mean and variance in the normal distribution. Therefore, the process of calculating the variational parameter λ i can be regarded as the process of determining the variational distribution. The cognitive state θ of the learner s i i used in the calculation can be understood as the random variable in the variational distribution corresponding to the variational parameter λ i . In specific calculations, θ i can be directly sampled from the corresponding variational distribution and then substituted into the formula for calculation, and then continuously optimized to finally obtain a more appropriate variational parameter λ i . Similarly, the variational parameter λ * is also in a similar way.
[0088] IV. Conduct cognitive diagnosis and related applications for new learners.
[0089] In the embodiments of the present invention, for new learners, the cognitive state distribution of new learners is inferred by using historical answer records and the optimized global prior distribution. At this time, the optimized variational distribution q(θ * * ; λ * ) corresponding to the new learner s * can be obtained through the above scheme and used as the posterior distribution of the corresponding cognitive state Then, based on the posterior distribution of the cognitive state of the new learner, personalized services can be provided for them. For example, the learning resource recommendation and learning path planning mentioned above.
[0090] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiments can be implemented through software, or can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions of the above embodiments can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present invention.
[0091] Embodiment II
[0092] The present invention also provides a cognitive diagnosis system for sparse data scenarios, which is mainly implemented based on the method provided in the foregoing embodiments, as Figure 4 shown, the system mainly includes:
[0093] A data collection unit, configured to collect historical answer records of multiple learners and divide them into a training set and a validation set; the historical answer records of the learners include: question information and corresponding answer results;
[0094] A model establishment and training unit, configured to model the cognitive state of the learners using the training set, construct a Bayesian hierarchical model to establish the relationship between the cognitive state of the learners and the global prior distribution shared by all learners, introduce data-driven meta-learning technology, use the relevant parameters in the global prior distribution as meta-parameters, based on the Bayesian hierarchical model, infer the posterior distribution of the learners' cognitive state using the training set, and predict the answer results of the learners on the validation set, combine the answer results included in the validation set with the predicted answer results to construct a loss function, optimize the meta-parameters, and obtain an optimized global prior distribution;
[0095] A speculation unit, configured to, for a new learner, use the historical answer records and the optimized global prior distribution to calculate the posterior distribution of the cognitive state of the new learner, and the calculated posterior distribution of the cognitive state can be used to predict the answer results of the new learner on each question.
[0096] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the system is divided into different functional modules to complete all or part of the functions described above.
[0097] Embodiment III
[0098] The present invention also provides a processing device, as Figure 5As shown in the figure, it mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in the foregoing embodiments.
[0099] Furthermore, the processing device further includes at least one input device and at least one output device; in the processing device, the processor, the memory, the input device, and the output device are connected through a bus.
[0100] In the embodiments of the present invention, the specific types of the memory, the input device, and the output device are not limited; for example:
[0101] The input device can be a touch screen, an image acquisition device, a physical button, or a mouse, etc.
[0102] The output device can be a display terminal;
[0103] The memory can be a Random Access Memory (RAM), or a non-volatile memory, such as a disk memory.
[0104] Embodiment Four
[0105] The present invention also provides a readable storage medium storing a computer program, which implements the method provided in the foregoing embodiments when the computer program is executed by a processor.
[0106] In the embodiments of the present invention, the readable storage medium as a computer-readable storage medium can be disposed in the foregoing processing device, for example, as the memory in the processing device. In addition, the readable storage medium can also be various media that can store program codes, such as a USB flash drive, a mobile hard disk, a Read-Only Memory (ROM), a magnetic disk, or an optical disc.
[0107] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A cognitive diagnosis method for sparse data scenarios, characterized in that including: Collecting the historical answer records of multiple learners and dividing them into a training set and a validation set; The historical answer records of learners include: question information and corresponding answer results; Using the training set to model the cognitive state of learners, constructing a Bayesian hierarchical model to establish the relationship between the cognitive state of learners and the global prior distribution shared by all learners, introducing data-driven meta-learning techniques, taking the relevant parameters in the global prior distribution as meta-parameters, based on the Bayesian hierarchical model, using the training set to infer the posterior distribution of the cognitive state of learners, predicting the answer results of learners on the validation set, constructing a loss function by combining the answer results included in the validation set and the predicted answer results, optimizing the meta-parameters, and obtaining an optimized global prior distribution; For new learners, using the historical answer records and the optimized global prior distribution to calculate the posterior distribution of the cognitive state of new learners, and the calculated posterior distribution of the cognitive state can be used to predict the answer results of new learners on each question.
2. The cognitive diagnosis method for sparse data scenarios according to claim 1, characterized in that The construction of the Bayesian hierarchical model to establish the relationship between the cognitive state of learners and the global prior distribution shared by all learners includes: The learners i The cognitive state is represented by θ i , the global prior distribution shared by all learners is denoted as ψ, and the global prior distribution ψ is the learner s i The cognitive state θ i Define a parameterized prior distribution p(θ i |ψ), as a representation of the prior knowledge about the learner’s cognitive state.
3. The cognitive diagnosis method for sparse data scenarios according to claim 2, characterized in that Based on the Bayesian hierarchical model, the inference of the posterior distribution of the cognitive state of learners using the training set is expressed as: Among them, is the historical answering record corresponding to learner s in the training set i ; represents the question information corresponding to learner s in the training set i ; represents the answering result corresponding to learner s in the training set i ; represents the cognitive state θ of the given learner s i , the question information corresponding to learner s in the training set i and i under the condition of the global prior distribution ψ, the probability distribution of the answering result corresponding to learner s in the training set i is inferred; ; represents the probability distribution of the cognitive state θ of learner s inferred under the condition of the historical answering record corresponding to learner s in the given training set i and the global prior distribution ψ, that is, the posterior distribution of the cognitive state θ of learner s i i i i . 4. A cognitive diagnosis method for sparse data scenarios according to claim 3, characterized in that The construction of the loss function by combining the answer results included in the validation set and the predicted answer results and optimizing the global prior distribution includes: Constructing a loss function by combining the answer results included in the validation set and the predicted answer results: Among them, M represents the number of learners, represents the loss function of learner s i ; represents the loss functions of all learners, represents the historical answer records corresponding to learner s i in the given training set, and the question information corresponding to learner s i in the validation set Under the condition of i and the global prior distribution ψ, the probability distribution of the answer results corresponding to learner s in the validation set is inferred; represents the question information corresponding to learner s i in the given validation set and the cognitive state θ i of learner s i Under the condition of i the probability distribution of the answer results corresponding to learner s in the validation set is inferred; E represents the expectation; represents that θ i satisfies the posterior distribution of the cognitive state θ i of learner s i ; By minimizing the above loss function: Optimize the meta-parameters to obtain an optimized global prior distribution 5. A cognitive diagnosis method for sparse data scenarios according to claim 1, characterized in that For new learners, using the historical answer records and the optimized global prior distribution to calculate the cognitive state of new learners includes: For a new learner s * , combined with an optimized global prior distribution and historical answer records (Q * , R * ), infer the posterior distribution of its cognitive state θ * : Among them, Q * is the question information, R * is the corresponding answer result, θ * represents the cognitive state, which is a random variable, represents the optimized global prior distribution for the cognitive state θ * defined parametric prior distribution; represents the probability distribution of inferring the answer result R * under the conditions of the given question information Q * , the cognitive state θ and the optimized global prior distribution; * represents the probability distribution of inferring the cognitive state θ ' under the conditions of the given historical answer record (Q * , R * ) of the learner and the optimized global prior distribution; * represents that the cognitive state θ * satisfies the posterior distribution * 6. A cognitive diagnosis method for sparse data scenarios according to claim 1 or 3 or 5, characterized in that, The method further includes: using an inference method for the posterior distribution of the cognitive state based on variational inference, and taking the variational distribution obtained by inference for the learner as the posterior distribution of the learner's cognitive state; during the meta-parameter optimization process, for learner s i using a variational distribution q(θ i ; λ i ) as the corresponding posterior distribution of the cognitive state During the calculation process of the cognitive state of a new learner, using the variational distribution q(θ * ; λ * ) as the corresponding posterior distribution of the cognitive state where λ i and λ * are both variational parameters, which have the same structure and meaning as the global prior distribution ψ; obtaining the corresponding variational parameters by minimizing the KL divergence between the variational distribution and the corresponding posterior distribution of the cognitive state, and optimizing the variational parameters through the local loss; For the variational distribution q(θ i ; λ i ), the variational parameter λ i ; λ i ) is obtained by minimizing the KL divergence between the variational distribution q(θ ) and the posterior distribution of the corresponding epistemic state i , expressed as: Among them, learner s i 's cognitive state θ i is obtained by sampling from the variational distribution corresponding to the variational parameter λ i and represents represents the probability distribution of inferring the answering result corresponding to learner s i 's cognitive state θ i in the training set, given the cognitive state θ i of learner s in the training set, the corresponding item information and the global prior distribution ψ; i represents the probability distribution of inferring the answering result corresponding to learner s in the training set, given the corresponding item information i of learner s in the training set and the cognitive state θ i of learner s; i p(θ i |ψ) represents a parameterized prior distribution defined by the global prior distribution ψ for the cognitive state θ of learner s; E represents the expectation; θ i ~q(θ i ; λ i ) means that θ i satisfies the variational distribution q(θ i ; λ i ); i satisfies the variational distribution q(θ i ; λ i ); Optimize the variational parameter λ through local loss i After the optimization of λ i After optimization, obtain the corresponding variational distribution.
7. A cognitive diagnosis method for sparse data scenarios according to claim 6, characterized in that The optimization of the variational parameter λ by local loss i includes: For learners s i Define a local loss for the variational distribution q(θ i ; λ i ) and minimize it: where η is an empirical weight parameter, denotes minimizing the local loss Minimize the local loss by gradient descent and optimize λ i : Among them represents the random gradient descent operation for K steps, α is the learning rate, and ← is the assignment symbol.
8. A cognitive diagnosis system for sparse data scenarios, characterized in that, Implemented based on the method according to any one of claims 1 to 7, the system includes: A data collection unit for collecting the historical answer records of multiple learners and dividing them into a training set and a validation set; the historical answer records of learners include: question information and corresponding answer results; A model establishment and training unit for using the training set to model the cognitive state of learners, constructing a Bayesian hierarchical model to establish the relationship between the cognitive state of learners and the global prior distribution shared by all learners, introducing data-driven meta-learning techniques, taking the relevant parameters in the global prior distribution as meta-parameters, based on the Bayesian hierarchical model, using the training set to infer the posterior distribution of the cognitive state of learners, predicting the answer results of learners on the validation set, constructing a loss function by combining the answer results included in the validation set and the predicted answer results, optimizing the meta-parameters, and obtaining an optimized global prior distribution; A speculation unit for, for new learners, using the historical answer records and the optimized global prior distribution to calculate the posterior distribution of the cognitive state of new learners, and the calculated posterior distribution of the cognitive state can be used to predict the answer results of new learners on each question.
9. A processing device, characterized in that, including: One or more processors; A memory for storing one or more programs; Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.
10. A readable storage medium stores a computer program, characterized in that, When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Multivariate intelligent education method and system
CN110473123A
Student score prediction method and device based on fuzzy cloud cognitive diagnosis model
CN113674116A