A method for generating a student portrait based on a classification neural network
By adopting classified neural networks and distributed training architecture in student portrait generation technology, students' data and question data are classified and trained in parallel, solving the accuracy and training efficiency problems caused by data dispersion in the existing technology, and achieving higher prediction accuracy and training efficiency.
Patent Information
- Application Number
- CN202211266946.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-17
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-10-17
AI Technical Summary
The existing student portrait generation technology based on deep neural networks has problems in accuracy and training efficiency, especially the inaccurate prediction results caused by data dispersion, and the problem of the degradation of the timeliness of deep learning frameworks under the large amount of data.
The student portrait generation method based on classification neural network is adopted. By classifying student data and question data, different categories of models are trained in parallel using a distributed training architecture, training data is centralized, data distribution is automatically adjusted, distributed deployment is simplified, and model training is accelerated using distributed technology.
It improves the accuracy of model prediction, improves training efficiency, enhances the availability of image generation technology in online learning systems, and effectively solves the problem of poor accuracy caused by data dispersion.
Smart Images

Figure CN115577107B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of neural networks and machine learning, and particularly to a method for generating a student portrait based on a classification neural network. Background Art
[0002] With the sharp increase in the amount of student learning record data in online learning systems, the technology for generating student portraits based on neural networks has developed rapidly. Such methods are based on the historical learning data of students, track the learning progress of students, and then depict the knowledge mastery level of students, etc. For online learning systems, portrait generation needs to meet high accuracy while meeting high real-time performance to improve the user experience. How to generate student portraits faster and more accurately remains a highly challenging problem.
[0003] Currently, the knowledge tracing technology based on deep neural networks has achieved better prediction performance than the traditional BKT method and can capture the temporal and semantic features of exercise texts at a deeper level. Therefore, the technology based on deep neural networks has been widely studied and discussed. Most works are committed to improving the accuracy of student portraits by improving the accuracy of predicting whether students answer questions correctly or not. For example, the model is improved by combining educational data features: Cheng et al. introduced a mistake and guess factor in DKT to better simulate the real question-solving situation of students. Nagatani et al. introduced a time interval to simulate the forgetting behavior of students and studied the influence of the time interval on the prediction accuracy; the attention mechanism was introduced: Su et al. proposed to use cosine similarity to calculate the similarity of exercise questions when the model predicts the output, effectively capturing the long-term dependencies in the exercise sequence; interpretable knowledge tracing: EKT expands the knowledge state vector of each student into a knowledge state matrix updated over time, where each vector represents the degree of mastery of a certain concept by the student, effectively solving the problem that it is difficult for traditional DKT to determine which concepts students are good at or not familiar with. However, the current work still has problems in terms of accuracy and training efficiency. 1) Most studies ignore the negative impact of the dispersion of training data on the prediction results, that is, the answer data of all students to all questions are uniformly trained, which may reduce the prediction accuracy. 2) At the same time, with the sharp increase in the amount of data, the timeliness of deep learning frameworks is constantly decreasing. In real online learning scenarios, the answer data of students increases sharply every day. Only by updating the model in a timely manner can the accuracy of the portrait be guaranteed. In order to apply this technology in real online learning systems, it is extremely important to improve the training timeliness of the model. The current neural networks are optimized by increasing the training data dimension and model depth, etc. Although the prediction accuracy is improved, the training delay is greatly increased, which reduces the applicability in real scenarios.
[0004] Today, cloud servers have become one of the mainstream training platforms. In most deep learning-based knowledge tracing techniques, a single processor uses a complete dataset to train a model. To process the huge cloud data, network transmission latency and training latency have become the main bottlenecks. For online learning scenarios with high real-time requirements, this high-latency problem has greatly reduced the learning progress of learners and affected the user experience. To accelerate model training, distributed model training techniques have been widely studied. However, distributed architecture design requires analyzing data distribution. It is very inefficient for developers to manually adjust the network structure to obtain a robust network structure without knowing the specific data distribution. Summary of the Invention
[0005] The object of the present invention is to propose a method for generating a student portrait based on a classification neural network in view of the deficiencies of the prior art. By adopting a distributed training architecture, student data and question data are classified, and different category models are trained and predicted in parallel, and the training data is centralized, thereby improving the model prediction accuracy. By automatically adjusting the data distribution, the model can be simply and quickly deployed distributively, and then the distributed technology is used to accelerate model training, enhancing the high availability of the portrait generation technology in the online learning system, and effectively solving the problem that data dispersion may lead to poor accuracy.
[0006] The object of the present invention is achieved as follows: A method for generating a student portrait based on a classification neural network, characterized in that by classifying student data and question data, based on a distributed training architecture, different category models are trained and predicted in parallel, thereby improving the model prediction accuracy. The specific process for efficiently generating a student portrait based on a classification neural network is as follows:
[0007] S100: K devices participating in knowledge tracing training send the device identification number and the distribution information of local data to the central server.
[0008] S200: Use a mathematical model to obtain the performance functions of students and questions, and generate the embedded representation vectors of students and questions.
[0009] S300: The central server receives the information sent by the devices, and statistically analyzes the data distribution for each proxy server. Using a controllable classification algorithm, the representation vectors of students and questions generated in S100 are classified, so that the knowledge points distribution after integrating the datasets within the equivalent classes is as concentrated as possible.
[0010] S400: The central server remaps the classified students and questions and sends them together with the weight parameters W 0 and b 0 to the corresponding proxy servers.
[0011] S500: The models are trained in parallel among the proxy servers, and the internal devices of the proxy servers simultaneously use the dataset sent by the central server to train the models. The proxy servers will, according to the weight parameters W * and b * sent to the central server. The central server will, according to the weight parameters W * and b * restore the knowledge tracing model and predict the knowledge states of such students; meanwhile, all proxy servers will use their respective updated weight parameters W * and b * to continue training.
[0012] S600: The central server updates the knowledge states of the students through the federated averaging algorithm and visualizes the updated knowledge states to construct the user portraits of the students.
[0013] S700: Repeat the above steps S500 - S600 until the model converges and the training ends to obtain the final user portraits.
[0014] In step S100, sending the device identification number K to the central server is a very crucial step. The value of K is not arbitrary. The size of K is closely related to the accuracy of knowledge tracing prediction and the number of user portraits. Assuming the number of user portraits is M and the number of question features is N, then K = M * N;
[0015] In step S200, the mathematical model is represented by the following formulas (a) and (b):
[0016]
[0017] Q = DNN(L, N,..., Time) (b)
[0018] where X represents the input layer vector, which is composed of the text information T, the knowledge point vector K, and the secondary compression vector. The secondary compression information is jointly trained by secondary information such as difficulty, number of hints, and training time.
[0019] In step S300, the classification algorithm adopts K - means, which is a controllable clustering algorithm and is calculated by the following formulas (c) and (d):
[0020]
[0021]
[0022] where c (i) and μ jrespectively represent the category to which student i belongs and the centroid of class j.
[0023] In step S400, the following (e) and (f) are used for calculation in the remapping technique:
[0024]
[0025]
[0026] In the formula: Q is the query variable, and through this variable, important information for the Query is found from all K; K is the address pointing to the Value of the information; V is the content of the information.
[0027] The above formulas (e) and (f) are the core formulas for the remapping of question information and also the core formulas of the Bert model. Using the Bert model for remapping can retain as much text information of the question as possible with digital information.
[0028] In step S500, the iteration of the neural network and the weight parameters W * and b * The update formulas are represented by the following (g) to (l):
[0029] f t = σ(W f ·[h t-1 , x t +b f ) (g);
[0030] i t = σ(W i ·[h t-1 , x t +b i ) (h);
[0031]
[0032]
[0033] o t = σ(W o ·[h t-1 , x t +b o ) (k);
[0034] h t = o t *tanh(C t ) (l);
[0035] Among them, in the formula: σ is the sigmoid activation function; h t-1 and x trespectively represent the hidden layer and the input layer of the neural network, and iteratively train the data of the next hidden layer according to the information of the previous hidden layer and the input layer; f t is the forget gate, acting as C t-1 's coefficient is used to calculate C t ; i t is the input gate; C t is the information of the current neuron; is the cell state update value; o t is the output gate, determining the information of the hidden layer; h t is the hidden layer; W f is the weight vector of the forget gate cell state; W C is the weight vector for calculating the information of the current neuron; W i is the weight vector of the input cell state; W o is the weight value vector of the output cell state; b f is the bias term of the forget gate cell state; b i is the bias term of the input cell state; b C is the bias term for calculating the information of the current neuron; b o is the bias term of the output cell state.
[0036] In step S600, the joint average algorithm of the student knowledge state is represented by the following formula (m):
[0037]
[0038] where, M i represents the student knowledge state under each batch training in the N-class question set.
[0039] Compared with the prior art, the present invention has the following remarkable technical progress and beneficial technical effects:
[0040] 1) Build models for different categories of students and questions in a divide-and-conquer manner, reducing the problem of low model prediction accuracy caused by the dispersion of data distribution, and further improving the prediction accuracy;
[0041] 2) Quickly achieve distributed deployment in a data classification manner, improving the model training efficiency, and thus improving the usability of the online portrait generation technology in the online learning system. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is the flowchart of the present invention;
[0043] Figure 2 is the schematic diagram of the mathematical model;
[0044] Figure 3 is the schematic diagram of the data classification result;
[0045] Figure 4 Schematic diagram of a neural network architecture under a single training model;
[0046] Figure 5 Schematic diagram of a student user profile. Detailed implementation manners
[0047] It is assumed that a large amount of students' answering data is collected from a remote online learning platform and stored in a central server; it is assumed that information such as the device number corresponding to the proxy server is known and can be normally sent to the central server. Taking the answering data of advanced mathematics in a university as an example (this part of the data consists of 2,500 students answering 200 questions, and a total of 18 knowledge points are included), the present invention will be further described below in conjunction with the accompanying drawings and specific implementations.
[0048] Example 1
[0049] Refer to Figure 1 , the student profile is created in the following steps in this example:
[0050] Step S100: After the corresponding device information is sent by the proxy server, the central server will select the corresponding device according to the data information and relevant requirements. It is assumed that it is planned to divide these 2,500 students into 3 categories and generate corresponding student profiles; it is assumed that dividing the question data into 3 categories can best reflect the relationship between the data. At this time, K = 3 * 3 = 9 proxy servers will be selected to participate in distributed training.
[0051] Step 200: After the corresponding proxy server is selected, the central server will vectorize the students and questions using a mathematical model. For the students, they will be vectorized according to their answering situations for each question; for the questions, they will be vectorized according to information such as the text information of this question, the average answering time, and the average number of prompts.
[0052] Refer to Figure 2 , use a mathematical model to vectorize the students and questions, and use the method of compressing secondary information and splicing important information to retain the question information while compressing the question dimension to accelerate future model training. Each block represents a kind of information dimension, and through the compression of multiple information dimensions, high-density information storage is achieved.
[0053] Step 300: After the students and questions are vectorized, the K-means clustering algorithm is used to perform unsupervised learning on the questions and students. The pseudo-code of the K-means clustering algorithm is shown in Table 1 below:
[0054] Table 1 Pseudo-code of the K-means clustering algorithm
[0055]
[0056] This algorithm has a certain degree of controllability, and the number of classifications can be controlled, which is also one of the methods for choosing this algorithm. Because when generating student portraits, the number of student portraits can be freely controlled, which is very crucial. The students are divided into 3 categories for generating three types of student portraits. The questions are divided into 3 categories because through a large number of experiments, it is found that dividing the questions into 3 categories can best reflect the relationships among the data in this dataset.
[0057] See Figure 3 a, Visualization graph of the question classification results obtained by using the K-means clustering algorithm for questions. Among them, different symbols represent their respective categories (where "+", "*", and "o" respectively represent the three equivalent classes divided by K-means), and the numbers represent question IDs. From this graph, it can be seen that the K-means clustering algorithm can well classify the question vectors after extracting the hidden information in the text.
[0058] See Figure 3 b, Visualization graph of the student classification results obtained by using the K-means clustering algorithm for students. Among them, different symbols represent their respective categories (where "+", "*", and "o" respectively represent the three equivalent classes divided by K-means), and the numbers represent student IDs. From this graph, it can be seen that the K-means clustering algorithm can well classify the student vectors.
[0059] Step 400: To prevent information leakage, remap the student IDs and question IDs before sending them to the proxy server. The student IDs use a one-to-one mapping; the question IDs are mapped using the Bert model.
[0060] Step 500: The proxy server updates the neural network weight parameters W * and b * , and use the LSTM model for training because this model has certain advantages in sequence modeling problems, has long-term memory function, and solves the problems of gradient disappearance and gradient explosion existing in the training process of long sequences.
[0061] See Figure 4 , This figure shows the neural network framework and model training details under a single training model. The specific search for this neural network architecture includes three parts: defining the search space, defining the search algorithm, and the model performance evaluation strategy. The search space is continuously differentiable, and the weight parameters W and b can be searched jointly through the gradient descent method; the search algorithm uses the long short-term memory model LSTM; the model performance evaluation strategy is that the central server integrates the model with each joint average algorithm as the evaluation object of the network architecture.
[0062] Step 600: The proxy server will send the parameter training results to the central server at regular intervals. After receiving the parameters, the central server will restore the knowledge tracing model and predict the student knowledge state. Previously, the answering data of each type of student was divided into three categories, and each type of student would generate N types of student portraits. The central server will use the joint averaging algorithm to summarize these N types of portraits and finally generate an overall student portrait.
[0063] See Figure 5 , after visualization according to the knowledge state vector, a student user portrait is generated, a 18-sided radar chart, where the distance from the center of the angle represents the student's knowledge mastery level. The closer to the center, the less firmly the knowledge point is mastered.
[0064] See Figure 5 a, The user portrait N1 shows the mastery status of some knowledge points of the Nth type of student. This knowledge state is obtained by training the neural network model on the data of the first type of questions. From this figure, it can be known that the mastery levels of the 0-17 knowledge points are 0.9%, 1.1%,..., 45.7% respectively.
[0065] See Figure 5 b, The user portrait N2 shows the mastery status of some knowledge points of the Nth type of student. This knowledge state is obtained by training the neural network model on the data of the second type of questions. From this figure, it can be known that the mastery levels of the 0-17 knowledge points are 99.9%, 99.9%,..., 41.7% respectively.
[0066] See Figure 5 c, The user portrait N3 shows the mastery status of some knowledge points of the Nth type of student. This knowledge state is obtained by training the neural network model on the data of the third type of questions. From this figure, it can be known that the mastery levels of the 0-17 knowledge points are 100.0%, 100.0%,..., 34.8% respectively.
[0067] See Figure 5 d, The overall user portrait shows the overall knowledge mastery status of the Nth type of student. This knowledge state is calculated by the joint averaging algorithm from the first three knowledge states. From this figure, it can be known that the mastery levels of the 0-17 knowledge points are 66.6%, 66.6%,..., 40.0% respectively.
[0068] The present invention starts from the data level. Aiming at the problem that data dispersion may lead to poor accuracy, the training data is centralized to improve the model prediction accuracy and effectively solve the problem that data dispersion may lead to poor accuracy. At the same time, by automatically adjusting the data distribution, the model can be simply and quickly deployed in a distributed manner, and then the distributed technology is used to accelerate the model training, enhancing the high availability of the portrait generation technology in the online learning system.
[0069] The above is only a further description of the present invention, and is not intended to limit this patent. All equivalent implementations of the present invention should be included within the scope of the claims of this patent.
Claims
1. A method for generating a student portrait based on a classification neural network, characterized in that, the method specifically includes the following steps: S100: K devices participating in knowledge tracing training send device identification numbers, local data labels, and distribution information of the quantity of each label to the central server; S200: The central server uses a mathematical model to generate student and question embedding vectors; S300: The central server receives the information sent by the devices, statistically analyzes the data distribution for each proxy server, classifies the generated student and question embedding vectors using a controllable classification method, constructs models for different categories of students and questions, and performs distributed deployment of the data; S400: The central server remaps the classified student and question information and sends it, together with the weight parameters W 0 and b 0 of the initialized neural network model, to the corresponding proxy server; S500: Each proxy server uses the dataset sent by the central server to train the model in parallel. The proxy server updates the weight parameter W obtained from this round of training for this group. * and b * , and then all proxy servers will use the self-updated weight parameter W * and b * to continue training. During the training process of the proxy server, the weight parameter W * and b * are sent to the central server. The central server restores the knowledge tracing model according to the weight parameter W * and b * sent by the proxy server and predicts the knowledge state of this type of students. S600: The central server updates the knowledge state of the students through the joint average algorithm, and visualizes the updated knowledge state to construct a user portrait of the students; S700: Repeat the above steps S500 - S600 until the model converges and the training ends to obtain the final student portrait; The controllable classification method in the step S300 is to classify the generated student and question embedding vectors using the K-means algorithm represented by the following formulas (c) - (d): where: c () is the class that student i is closest to among k classes; μ is the centroid of class j; x () is the n-dimensional vector of student i.
2. The method for generating a student portrait based on a classification neural network according to claim 1, characterized in that, K in the step S100 = M * N, where M is the number of user portraits; N is the number of question features.
3. The method for generating a student portrait based on a classification neural network according to claim 1 above, characterized in that, the mathematical model in the step S200 is calculated by the following formulas (a) - (b): X = T ⊕ K ⊕ Q (a); Q = DNN(L, N,..., Time) (b); In the formulas: X is the input layer vector, and is composed of the text information T, the knowledge point vector K, and the vector of secondary compression information spliced together; T is the text information; K is the knowledge point vector; q is the Q matrix, representing the relationship between the question and the knowledge point; L is the question difficulty; N is the average number of hints for the student to answer questions; Time is the average time for the student to answer questions; The secondary compression information is jointly trained and compressed by the secondary information of difficulty, hint times, and training time.
4. The method for generating a student portrait based on a classification neural network according to claim 1, characterized in that, the remapping in the step S400 is the attention mechanism in the Bert neural network model, and the attention mechanism is calculated by the following formulas (e) - (f): In the formulas: Q is the query variable, and through this variable, important information for the Query is found from all K; K is the address pointing to the Value of the information; V is the content of the information.
5. The method for generating a student portrait based on a classification neural network according to claim 1, characterized in that, The weight parameters W * and b * in step S500 are updated by self - mixing according to the following equations (g) to (l): f A = σ(W f · [h A-1 , x A + b f ) (g); i A = σ(W i · [h A-1 , x A + b i ) (h); o A = σ(W o · [h A-1 , x A + b o ) (k); h A = o A *tanh (C A ) (1); Where: σ is the sigmoid activation function; h t-1 and x t represent the hidden layer and the input layer of the neural network respectively, and the data of the next hidden layer is iteratively trained according to the information of the previous hidden layer and the input layer; f A is the forget gate, which is used as the coefficient of C t-1 to calculate C t ; i A is the input gate; C A is the information of the current neuron; is the unit state update value; o A is the output gate, which determines the information of the hidden layer; h A is the hidden layer; W f is the weight vector of the forget gate unit state; W C is the weight vector for calculating the information of the current neuron; W is the weight vector of the input unit state; W o is the weight vector of the output unit state; b f is the bias term of the forget gate unit state; b is the bias term of the input unit state; b C is the bias term for calculating the information of the current neuron; b o Is the bias term for the output unit state.
6. The method for generating a student portrait based on a classification neural network according to claim 1, characterized in that, the knowledge state in the step S600 is updated by the joint average algorithm of the following formula (m): In the formula: N represents the student knowledge state under each batch training in the N-class question set.
Citation Information
Patent Citations
Self-balancing model training method based on federal learning
CN113962359A
User portrait construction method based on microblog heterogeneous information
WO2022206103A1