MOOC Dropout Prediction Method and System for Graph Node Classification Assisted by Large Language Model
Through the large language model auxiliary graph node classification method, discussion graphs are constructed and pseudo-labels are generated using discussion area data, which solves the problem of insufficient accuracy of dropout prediction in the MOOC platform, and achieves more accurate dropout risk assessment and prediction.
Patent Information
- Application Number
- CN202411881677.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-12-19
AI Technical Summary
The existing MOOC platform's dropout prediction method fails to effectively utilize discussion area data, especially learner speeches and interaction records, resulting in insufficient accuracy of dropout prediction.
The large language model assisted graph node classification method is adopted, and the pre-trained model is fine-tuned by randomized singular value decomposition, discussion graphs are constructed and pseudo-labels are generated. Combined with learners' cognitive hierarchy analysis and course behavior type judgment, the MOOC dropout prediction model is trained.
It improves the accuracy of MOOC dropout prediction, can more comprehensively evaluate students' dropout risks, provide more accurate predictions and intervention measures, and improve course completion rates.
Smart Images

Figure CN119323293B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of MOOC dropout prediction, and particularly relates to a MOOC dropout prediction method and system assisted by a large language model for graph node classification. Background Art
[0002] Massive Open Online Courses (MOOCs) are large-scale online education courses that integrate thousands of carefully designed online courses through the network, providing learners with a rich selection of courses. MOOC platforms are favored by students for their flexible learning atmosphere and low pressure, but this free and open environment is also accompanied by a high dropout rate problem. Due to the large number of online education groups, manual dropout prediction is difficult to achieve, and automated dropout prediction methods based on machine learning have been widely used in the field of MOOC dropout prediction.
[0003] Current prediction methods often only focus on students' personal information and course behavior information, and use some simple machine learning methods to predict students as separate individuals, while the information contained in the discussion area is ignored. There is currently no use of students' posts and interaction records in the discussion area for dropout prediction. First, the posts and replies of students in the discussion area often contain a large amount of semantic information, which can be used to judge the learning situation of students in this course; second, the interaction of students in the comment area enables the mutual reference of the learning situation assessment of different students. The difficulty in using the data in the discussion area lies in that although the data in the discussion area is suitable for graph processing, the levels of the participants in the discussion are often uneven, and the methods generally used for processing graph data may not have good effects; the massive text data in the comment area also brings difficulties to data processing. Summary of the Invention
[0004] The purpose of the present invention is to solve the problems existing in the prior art and provide a MOOC dropout prediction method and system assisted by a large language model for graph node classification.
[0005] In order to achieve the above invention purpose, the present invention specifically adopts the following technical solutions:
[0006] In the first aspect, the present invention provides a MOOC dropout prediction method assisted by a large language model for graph node classification, which includes the following steps:
[0007] S1: Obtain the course behavior information data of each learner in the target course from the MOOC platform, and encode the preprocessed course behavior information data to obtain the course behavior feature data of the learners;
[0008] S2: Take each learner as a node in the discussion graph, take the interaction relationships between learners in the discussion area as the edges connecting the nodes in the discussion graph, and take the course behavior characteristic data of the learners as the node features to obtain an initial discussion graph, and process the initial discussion graph to obtain a final discussion graph;
[0009] S3: Use the randomized singular value decomposition method to replace the random initialization process of the LoRA parameter matrix in the O-LoRA method, and fine-tune the pre-trained large language model based on the LoRA parameter matrix after randomized singular value decomposition by the O-LoRA method to obtain a fine-tuned large language model;
[0010] S4: Input a first prompt word consisting of a first example containing questions and answers and a first additional question into the fine-tuned large language model. The fine-tuned large language model outputs the corresponding cognitive level classification results of the learners under the target chapter according to the answers in the first example, and obtains a cognitive feature matrix based on the cognitive level classification results;
[0011] S5: Input a second prompt word consisting of a second example containing questions and answers and a second additional question into the fine-tuned large language model. The fine-tuned large language model judges whether the course behavior types of two learners are the same, outputs the course behavior type judgment result according to the answers in the second example, encodes the course behavior type judgment result, and uses the course behavior type encoding result as the pseudo-label of each edge in the final discussion graph;
[0012] S6: Use the cognitive feature matrix of the learners and the final discussion graph as the input of the MOOC dropout prediction model to train the MOOC dropout prediction model;
[0013] S7: Input the course behavior information data of the learner to be predicted into the trained MOOC dropout prediction model, and output the prediction probability that the learner belongs to a dropout.
[0014] Based on the above solutions, each step can be implemented in the following preferred specific ways.
[0015] As a preference of the above first aspect, in step S1, each course behavior information data includes N video teaching videos for learners to watch, N work homework completion situations, N test test situations of tests, the total number of times N visit that the learner accesses the target course, and the text data composed of all the speeches of the learner in the target course discussion area, and preprocess the text data to form preprocessed course behavior information data;
[0016] In step S1, the specific process of encoding the preprocessed course behavior information data is:
[0017] S11: For each teaching video of each learner, the ratio between the viewing time of the learner and the total viewing time of the corresponding teaching video is used as the viewing time coding value of the teaching video, and all viewing time coding values of the learner constitute a corresponding viewing time coding vector;
[0018] S12: The total number of times the learner visits the target course is used as the access frequency coding vector corresponding to the learner;
[0019] S13: For each assignment completion status of each learner, if the learner completes the assignment, the ratio between the number of correct answers of the learner in each assignment and the corresponding total number of questions is used as the assignment submission status coding value; if the learner does not complete the assignment, the assignment submission status coding value is 0. Finally, all the assignment submission status coding values of the learner constitute the corresponding assignment submission status coding vector;
[0020] S14: For each test case of each learner, if the learner completes the test, the ratio between the learner's test score and the corresponding test total score is used as the test score coding value; if the learner does not complete the test, the test score coding value is 0, and all the test score coding values of the learner constitute the corresponding test score coding vector;
[0021] S15: for each learner's preprocessed text data, the preprocessed text data is sent to the encoding visible large language model to obtain the learner's discussion content encoding vector;
[0022] S16: Concatenate the viewing time coding vector, the access frequency coding vector, the homework submission status coding vector, the test score coding vector and the discussion content coding vector as the learner's course behavior characteristic data.
[0023] As a preferred embodiment of the above-mentioned first aspect, in step S2, the specific process of obtaining the final discussion graph is: determine whether there is an isolated node in the initial discussion graph: if so, traverse each isolated node in the initial discussion graph, calculate the cosine similarity between the node feature corresponding to an isolated node and the node features of the remaining nodes, and use the K nodes with the largest cosine similarity as the most similar nodes to the isolated node, connect the isolated node and its corresponding most similar nodes to form K edges and add them to the initial discussion graph to obtain the final discussion graph; if not, use the initial discussion graph as the final discussion graph.
[0024] As a preferred embodiment of the first aspect, in step S3, the specific process of obtaining the LoRA parameter matrix after randomized singular value decomposition by using the randomized singular value decomposition method is:
[0025] S31: Randomly sample from a normal distribution as the column vectors of a random matrix to generate a complete random matrix;
[0026] S32: Use the random matrix to project the parameter matrix of the large language model into a low-dimensional space to obtain a projection matrix; perform an orthogonal triangular decomposition on the projection matrix to obtain a first orthogonal matrix and an upper triangular matrix; multiply the transposed first orthogonal matrix by the parameter matrix of the large language model to obtain a low-dimensional matrix; perform a singular value decomposition on the low-dimensional matrix to obtain a second orthogonal matrix, a third orthogonal matrix, and a diagonal matrix;
[0027] S33: Use the indices corresponding to the largest r singular values in the diagonal matrix as the initialization indices. In the second orthogonal matrix, use the column vectors corresponding to the initialization indices as the initialized first LoRA parameter matrix. In the third orthogonal matrix, use the column vectors corresponding to the initialization indices as the initialized second LoRA parameter matrix. The initialized first LoRA parameter matrix and the initialized second LoRA parameter matrix form a set of LoRA parameter matrices after randomized singular value decomposition.
[0028] As a preference of the first aspect above, in step S4, both the questions in the first example and the first additional questions include all the speeches of the learner in the target chapter of the target course discussion area, the name of the target course, and the name of the target chapter, the cognitive level to which the learner belongs judged according to the learner's speech, the cognitive levels of the online community discussion, and the judgment criteria corresponding to each cognitive level; design the answers in the first example in the way of thinking chain prompting, and the answers in the first example include the classification results of the learner's cognitive levels and the logical reasoning process of judging the cognitive level to which the learner belongs;
[0029] In step S4, the generation process of the cognitive feature matrix is as follows: perform one-hot encoding on the classification results of the cognitive levels to obtain the cognitive level encoding results. The cognitive level encoding results of all chapters form the cognitive feature vector of the learner. Use the cognitive feature vector of one learner as a row in the cognitive feature matrix, and the cognitive feature vectors of all learners form a complete cognitive feature matrix.
[0030] As a preference of the first aspect above, in step S5, both the questions in the second example and the second additional questions include the course behavior information data of two learners and judge whether the course behavior types of the two learners are the same. Design the answers in the second example in the way of thinking chain prompting, and the answers in the second example include the judgment results of the course behavior types and the logical reasoning process of judging whether the course behavior types are the same; among them, when both learners belong to dropouts or both belong to non-dropouts, it is considered that the course behavior types of the two learners are the same.
[0031] Preferably, in the first aspect above, in step S6, during the training process of the MOOC dropout prediction model, a course behavior information matrix is composed of the course behavior information data of all learners. After splicing the course behavior information matrix and the cognitive feature matrix of the learners, the learner feature matrix is obtained. Two multi-layer perceptrons are used to output the assortativity probability of each edge according to the edge set in the final discussion graph. The assortativity probability is adjusted using the pseudo-labels of each edge in the final discussion graph to obtain the adjusted assortativity probability. The adjusted assortativity probability is used as the element in the low-pass adjacency matrix to construct the low-pass adjacency matrix. The learner feature matrix is input into the simplified graph convolutional network. First, the learner feature matrix is processed by the third multi-layer perceptron to output the initial low-pass feature matrix. The initial low-pass feature matrix is processed through multiple low-pass filtering processes. In each low-pass filtering process, the result of the previous low-pass filtering process is multiplied by the low-pass adjacency matrix to obtain the current low-pass filtering result. The result of the last low-pass filtering process is used as the learner low-pass feature matrix. The result of subtracting 1 from the adjusted assortativity probability is used as the element in the high-pass adjacency matrix to construct the high-pass adjacency matrix. The processed high-pass adjacency matrix is obtained by subtracting the identity matrix from the weighted high-pass adjacency matrix. The learner feature matrix is input into the Lap-SGC model. First, the learner feature matrix is processed by the fourth multi-layer perceptron to output the initial high-pass feature matrix. The initial high-pass feature matrix is processed through multiple high-pass filtering processes. In each high-pass filtering process, the result of the previous high-pass filtering process is multiplied by the processed high-pass adjacency matrix to obtain the current high-pass filtering result. The result of the last high-pass filtering process is used as the learner high-pass feature matrix. The learner low-pass feature matrix and the learner high-pass feature matrix are spliced to form the learner comprehensive representation matrix. After performing position space embedding on the final discussion graph, the nearest neighbor graph is obtained. The nearest neighbor graph is input into the graph convolutional network to obtain the learner encoding matrix. The learner comprehensive representation matrix and the learner encoding matrix are spliced and input into the fifth multi-layer perceptron to output the prediction probability that each learner belongs to a dropout student. Based on the prediction probability that each learner belongs to a dropout student and the true course behavior label of the learner, the cross-entropy loss between the two is calculated. Based on minimizing the cross-entropy loss, the parameters of the MOOC dropout prediction model are updated until the preset iteration round threshold is reached, and the MOOC dropout prediction model converges to obtain the trained MOOC dropout prediction model.
[0032] Preferably, as the first aspect described above, the specific process of using two multi-layer perceptrons to output the assortativity probability of each edge is as follows: The learner feature vectors of the two nodes corresponding to each edge pass through the first multi-layer perceptron respectively, and the corresponding outputs are the first intermediate representation vector and the second intermediate representation vector representing the two learners respectively. After concatenating the first intermediate representation vector and the second intermediate representation vector and passing them through the second multi-layer perceptron, a third intermediate representation vector is obtained. After concatenating the second intermediate representation vector and the first intermediate representation vector and passing them through the second multi-layer perceptron, a fourth intermediate representation vector is obtained. The average of the third intermediate representation vector and the fourth intermediate representation vector is calculated to obtain the assortativity probability corresponding to each edge; wherein, the learner feature vector is a row in the learner feature matrix.
[0033] The specific process of adjusting the assortativity probability is as follows: When the course behavior types of the two learners are different, take the minimum value between and 0 as the adjusted assortativity probability; when the course behavior types of the two learners are the same, take the maximum value between and 0 as the adjusted assortativity probability; is the assortativity probability before adjustment; is a preset hyperparameter.
[0034] Preferably, as the first aspect described above, the specific process of obtaining the nearest neighbor graph after performing position space embedding on the final discussion graph is as follows: First, perform an embedding operation on each node in the final discussion graph, map it to an implicit dimension of a position space, take the average of the learner feature vectors of all the first-order neighbor nodes of each node in the final discussion graph as the position encoding result of the node, calculate the L2 distance between the position encoding result of each node and the position encoding results of the remaining nodes in the final discussion graph, take the M nodes with the smallest L2 distance as the nearest neighbor nodes of each node, connect each node in the final discussion graph with its nearest neighbor nodes, and use the learner feature matrix as the node feature to construct the nearest neighbor graph.
[0035] Furthermore, for the node in the final discussion graph, the specific form of its position encoding result is as follows:
[0036]
[0037] In the formula, represents the position encoding result of the node in the final discussion graph; represents the degree of the node with a self-loop; represents the set of first-order neighbors of the node with a self-loop; represents a first-order neighbor node in the set of first-order neighbors; represents the first-order neighbor node The corresponding learner feature vector.
[0038] In a second aspect, the present invention provides a massive open online course (MOOC) dropout warning system assisted by a large language model for graph node classification, including:
[0039] A data acquisition module, configured to acquire the course behavior information data of each learner in the target course from the MOOC platform, encode the preprocessed course behavior information data, and obtain the course behavior feature data of the learners;
[0040] A discussion graph construction module, configured to use each learner as a node in the discussion graph, use the interaction relationship between learners in the discussion area as the edge connecting the nodes in the discussion graph, and use the course behavior feature data of the learners as the node features to obtain an initial discussion graph, and process the initial discussion graph to obtain a final discussion graph;
[0041] A model fine-tuning module, configured to replace the random initialization process of the LoRA parameter matrix in the O-LoRA method with the randomized singular value decomposition method, and fine-tune the pre-trained large language model by the O-LoRA method based on the LoRA parameter matrix after the randomized singular value decomposition to obtain a fine-tuned large language model;
[0042] A cognitive level analysis module, configured to input a first prompt word composed of a first example including questions and answers and a first additional question into the fine-tuned large language model, and the fine-tuned large language model outputs the corresponding cognitive level classification result of the learner under the target chapter according to the answers in the first example, and obtains a cognitive feature matrix based on the cognitive level classification result;
[0043] A pseudo-label acquisition module, configured to input a second prompt word composed of a second example including questions and answers and a second additional question into the fine-tuned large language model, and the fine-tuned large language model determines whether the course behavior types of two learners are the same, and outputs a course behavior type judgment result according to the answers in the second example, and encodes the course behavior type judgment result, and uses the course behavior type encoding result as the pseudo-label of each edge in the final discussion graph;
[0044] A model training module, configured to jointly use the cognitive feature matrix of the learners and the final discussion graph as the input of the MOOC dropout prediction model to train the MOOC dropout prediction model;
[0045] A result acquisition module, configured to input the course behavior information data of the learner to be predicted into the trained MOOC dropout prediction model, and output the prediction probability that the learner belongs to a dropout;
[0046] An early warning prompt module, configured to send a dropout early warning message to the teacher user or management user of the MOOC platform, or send a reminder message to urge the learner belonging to a dropout to log in and study.
[0047] Compared with traditional machine learning-based MOOC dropout prediction methods, the present invention has the following beneficial effects:
[0048] Through in-depth semantic analysis of students' posts and replies in the discussion area by the large language model, the present invention can more accurately identify students' understanding of course knowledge points and learning difficulties, so as to more comprehensively evaluate their dropout risks. By constructing a discussion graph to combine course behavior data and discussion area interaction records, the MOOC dropout prediction model of the present invention can learn the interaction behaviors among students, effectively capture the implicit learning motivation, participation, and collaboration of students in the discussion area, and thus provide more accurate dropout prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 is a flowchart of the steps of the method of the present invention;
[0050] Figure 2 is a schematic diagram of a final discussion graph provided by an embodiment of the present invention;
[0051] Figure 3 is a schematic diagram of the process of training the MOOC dropout prediction model of the present invention;
[0052] Figure 4 is a system block diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention will be given with reference to the accompanying drawings. Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below. The technical features in the various embodiments of the present invention can be combined correspondingly without conflict.
[0054] In the description of the present invention, it should be understood that the terms "first" and "second" are only used for the purpose of distinguishing descriptions, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features.
[0055] As Figure 1 shown, in a preferred implementation manner of the present invention, the above-mentioned MOOC dropout prediction method assisted by large language model for graph node classification includes the following steps S1 to S7. The following will separately describe the specific implementation processes thereof in detail.
[0056] S1: Obtain the course behavior information data of each learner in the target course from the MOOC platform, and encode the preprocessed course behavior information data to obtain the course behavior feature data of the learner.
[0057] It should be noted that in step S1 of the present invention, each course behavior information data includes N video teaching videos for learners to watch, the completion status of N work assignments, the test status of N test tests, the total number of times N visit that the learner accesses the target course, and the text data composed of all the speeches of the learner in the target course discussion area. After preprocessing the text data, the preprocessed course behavior information data is formed.
[0058] In this embodiment, when preprocessing the text data, the text data is cleaned, specifically by removing irrelevant information in the text data, such as advertisements, emoticons, irrelevant links, etc., to ensure the quality of the text data for analysis. Then, the text data is subjected to text normalization processing, including converting to a unified character encoding, removing redundant spaces, and making the punctuation usage consistent.
[0059] It should be noted that in step S1 of the present invention, for each learner in the target course, the present invention involves encoding the detailed behavior information of their activities in the target course, such as video watching, assignment submission, and test participation, etc., to reflect the viewing situation of each teaching video by the learner, the completion status of the assignments, and the test scores, so as to facilitate the analysis of their learning behavior patterns and participation, and thus provide data support for further learning analysis and personalized feedback.
[0060] In step S1 of this embodiment, the specific process of encoding the preprocessed course behavior information data is as follows:
[0061] S11: For each teaching video of each learner, use the ratio between the viewing duration of the learner and the total duration of the corresponding teaching video as the viewing duration encoding value of the teaching video, and form the corresponding viewing duration encoding vector from all the viewing duration encoding values of the learner. ;
[0062] S12: Use the total number of times the learner accesses the target course as the access frequency encoding vector corresponding to the learner. ;
[0063] S13: For each learner's completion status of each assignment, if the learner has completed the assignment, the ratio between the number of correct problem-solving for each assignment of the learner and the total number of problems corresponding to the assignment is used as the assignment submission status coding value. If the learner has not completed the assignment, the assignment submission status coding value is 0. Finally, the corresponding assignment submission status coding vector is formed by all the assignment submission status coding values of the learner. ;
[0064] S14: For each test situation of each learner, if the learner has completed the test, the ratio between the test score of the learner and the total test score corresponding to the test is used as the test score coding value. If the learner has not completed the test, the test score coding value is 0. The corresponding test score coding vector is formed by all the test score coding values of the learner. ;
[0065] S15: For the preprocessed text data of each learner, the preprocessed text data is fed into a large language model with visible encoding (such as Bert) to obtain the discussion content coding vector of the learner. ;
[0066] S16: The viewing duration coding vector, access frequency coding vector, assignment submission status coding vector, test score coding vector, and discussion content coding vector are concatenated as the course behavior feature data of the learner. 。
[0067] S2: Each learner is regarded as a node in the discussion graph, the interaction relationship between learners in the discussion area is regarded as the edge connecting the nodes in the discussion graph, and the course behavior feature data of the learner is regarded as the node feature to obtain the initial discussion graph, and the initial discussion graph is processed to obtain the final discussion graph.
[0068] It should be noted that in step S2 of the present invention, in the process of analyzing the discussion area data, the method of constructing a discussion graph is particularly crucial. This method can effectively record and measure the communication patterns and behavioral characteristics between learners, and at the same time make a flexible response to the continuously evolving information in the discussion area. In such a graph, learners are regarded as nodes, and the interactions between learners are reflected by the edges connecting these nodes. The construction of the discussion graph enables the application of advanced analysis techniques such as graph neural networks to deeply explore and learn the complex patterns of students' behaviors. Thus, the present invention constructs a final discussion graph to capture the interactions of learners in the discussion area.
[0069] In step S2 of this embodiment, the specific process of obtaining the final discussion graph is as follows: Determine whether there are isolated nodes in the initial discussion graph. If there are, traverse each isolated node in the initial discussion graph, calculate the cosine similarity between the node features corresponding to an isolated node and the node features of the other nodes, and use the K nodes with the largest cosine similarity as the most similar nodes of the isolated node. Connect the isolated node to its corresponding most similar nodes, form K edges and add them to the initial discussion graph to obtain the final discussion graph. If not, use the initial discussion graph as the final discussion graph.
[0070] As Figure 2 shown, if learner A initiates a discussion, learners B and C reply to it, and learner D responds to learner B's reply, then edges {A, B}, {A, C}, and {B, D} are respectively constructed in the discussion graph.
[0071] S3: Replace the random initialization process of the LoRA parameter matrix in the O-LoRA method with the randomized singular value decomposition (SVD) method. Based on the LoRA parameter matrix after the randomized singular value decomposition, the O-LoRA method fine-tunes the pre-trained large language model to obtain the fine-tuned large language model.
[0072] It should be noted that the O-LoRA method is a method for fine-tuning a large language model, and its implementation belongs to the prior art. The large language model fine-tuned by the O-LoRA method can be used to implement different tasks. The O-LoRA method first randomly initializes the LoRA parameter matrix. In the present invention, the randomized singular value decomposition (SVD) method is used to replace the original random initialization process of the LoRA parameter matrix in the O-LoRA method to obtain the LoRA parameter matrix after the randomized singular value decomposition. This randomized singular value decomposition method usually has a lower computational cost than the complete singular value decomposition method because it does not require calculating the complete singular value decomposition process but approximates it through random sampling. Moreover, by initializing with the randomized singular value decomposition method, the model can be promoted to approach the effective solution space more quickly. Then, it is processed according to the original steps of the O-LoRA method, and finally, the fine-tuned large language model is obtained. Thus, the present invention efficiently processes two different tasks using the large language model while ensuring that the large language model does not forget the valuable knowledge it has obtained during the pre-training stage during the continuous learning process, optimizing the performance of the large language model.
[0073] In step S3 of the present invention, the specific process of obtaining the LoRA parameter matrix after the randomized singular value decomposition using the randomized singular value decomposition method is as follows:
[0074] S31: Randomly sample from the normal distribution as the column vectors of the random matrix to generate a complete random matrix.
[0075] In this embodiment, first, a random matrix S is created, with a dimension of k×(r + p), where k is the number of columns of the parameter matrix of the large language model; r is a positive integer that is much smaller than the number of columns k and the number of rows d of the parameter matrix of the large language model; p is the oversampling parameter, which is a value slightly larger than r. The column vectors of the random matrix are randomly drawn from a normal distribution, aiming to capture the key features of the column space of the parameter matrix of the large language model. This strategy can perform an approximate singular value decomposition of the parameter matrix of the large language model with higher computational efficiency.
[0076] S32: Project the parameter matrix W of the large language model into a low-dimensional space using the random matrix S to obtain a projection matrix T = WS; perform an orthogonal triangular (QR) decomposition on the projection matrix T to obtain a first orthogonal matrix Q and an upper triangular matrix R; multiply the transpose of the first orthogonal matrix by the parameter matrix W of the large language model to obtain a low-dimensional matrix D; perform a singular value decomposition (SVD) on the low-dimensional matrix D to obtain a second orthogonal matrix , a third orthogonal matrix and a diagonal matrix ;
[0077] S33: Use the indices corresponding to the largest singular values in the diagonal matrix as the initialization indices, use the column vectors corresponding to the initialization indices in the second orthogonal matrix as the initialized first LoRA parameter matrix, use the column vectors corresponding to the initialization indices in the third orthogonal matrix as the initialized second LoRA parameter matrix, and form a set of LoRA parameter matrices after randomized singular value decomposition from the initialized first LoRA parameter matrix and the initialized second LoRA parameter matrix.
[0078] In step S3, for the first cognitive level analysis task, the present invention first generates a random matrix corresponding to the first task and executes the above S31 - S33 processes to obtain a set of LoRA parameter matrices after randomized singular value decomposition corresponding to the first task , where are respectively the initialized first LoRA parameter matrix and the second LoRA parameter matrix corresponding to the first task. For the second edge pseudo-label discrimination, the present invention also first generates a random matrix corresponding to the second task and executes the above S31 - S33 processes to obtain a set of LoRA parameter matrices after randomized singular value decomposition corresponding to the second task , where are respectively the initialized first LoRA parameter matrix and the second LoRA parameter matrix corresponding to the second task.
[0079] S4: Input a first prompt consisting of a first example containing questions and answers and a first additional question into the fine-tuned large language model. The fine-tuned large language model outputs the corresponding cognitive level classification results of the learner under the target chapter according to the answers in the first example, and obtains a cognitive feature matrix based on the cognitive level classification results.
[0080] It should be noted that in step S4 of the present invention, both the questions in the first example and the first additional question include all the speeches of the learner under the target chapter in the target course discussion area, the name of the target course and the name of the target chapter, the cognitive level to which the learner belongs judged according to the learner's speech, the cognitive level of online community discussion, and the judgment criteria corresponding to each cognitive level; the answer in the first example is designed in the way of thinking chain prompt, and the answer in the first example includes the cognitive level classification results of the learner and the logical reasoning process for judging the cognitive level to which the learner belongs.
[0081] In this embodiment, considering that large language models have a series of significant advantages when performing cognitive level analysis tasks for online discussions, such as their excellent automation capabilities, which can quickly process and analyze large-scale discussion data, significantly improving work efficiency. Through deep learning technology training, large language models can understand the complexity of language, capture the nuances and underlying meanings in the discussion content. In addition, large language models have strong generalization capabilities and can adapt to different topics and diverse expression styles, ensuring the accuracy and universality of the analysis results. Large language models can also reduce biases caused by human factors and provide more objective and consistent analysis results. Therefore, the present invention designs the prompt of the large language model in the way of chain of thought prompting, and steps the logical reasoning process as much as possible to teach the large language model to think step by step. Even in the case of only a small number of examples, the large language model can demonstrate complex reasoning capabilities and ultimately help complete the cognitive level analysis of the discussion content. In addition to the chain of thought prompting method, this embodiment also designs the cognitive levels of online community discussions and the corresponding judgment criteria for each cognitive level into the prompt of the large language model. Specifically, the online community discussion cognitive level evaluation form proposed by Henri in 1992 is used, and the five cognitive levels into which the online discussion content of the discussants is divided in this evaluation form and the judgment criteria for each cognitive level are added to the prompt of the large language model. These five cognitive levels are basic identification, in-depth identification, reasoning, judgment, and strategy. Among them, the judgment criterion for basic identification is: in the cognitive process, first, by observing and studying the problem, identify its basic elements, and explore the connections between these elements, so as to construct a basic understanding of the problem; the judgment criterion for in-depth identification is: on the basis of the basic understanding, further analyze the problem and deeply understand the values, beliefs, and assumptions behind it. This level requires a deeper description of the problem to form a more comprehensive understanding; the judgment criterion for reasoning is: adopt logical reasoning strategies, adopt existing viewpoints or put forward new viewpoints. In this process, the recognized propositions are used as the basis for constructing new connections and arguments, reflecting the logical reasoning ability in the cognitive process; the judgment criterion for judgment is: judgment involves various forms such as decision-making, description, appreciation, and criticism. The learner's judgment is based on certain values or rules, reflecting the individual's ability to evaluate and select information based on specific criteria; the judgment criterion for strategy is: on the basis of understanding and judgment, propose strategies for coordinated actions to implement the plan, or make a choice among multiple options to make a decision. This level reflects the learner's ability to formulate effective strategies to achieve goals when facing problems.
[0082] In this embodiment, the designed first prompt is as follows: "Question: The following are the criteria for judging the cognitive levels of online discussions, which are divided into five levels. Basic identification: <Basic identification criteria>; In-depth identification: <In-depth identification criteria>; Reasoning: <Reasoning criteria>; Judgment: <Judgment criteria>; Strategy: <Strategy criteria>; Learner A posted a discussion under the "Learning Theory" chapter of the "Introduction to Modern Psychology" course: "I believe that the reinforcement learning theory can explain how rewards promote the frequent occurrence of specific behavior patterns." Please analyze the cognitive level of A's understanding of the knowledge points included in this chapter. Answer: Learner A's statement directly quotes the concept of "reinforcement learning theory" and applies it to explain the change in behavior frequency. This indicates that A can identify and use the basic terms and concepts in the course. A's statement does not show an exploration of the deeper meaning behind the reinforcement learning theory, nor does it provide additional analysis or critical thinking. Since the statement mainly focuses on the basic identification of the concept without further reasoning or in-depth analysis, it is judged to be at the basic identification level. Question: The following are the criteria for judging the cognitive levels of online discussions, which are divided into five levels. Basic identification: <Basic identification criteria>; In-depth identification: <In-depth identification criteria>; Reasoning: <Reasoning criteria>; Judgment: <Judgment criteria>; Strategy: <Strategy criteria>; Please analyze the cognitive level of the following discussion content according to this criterion and the course and chapter names I will input next, and tell me the cognitive level and your reason for judgment. Course name: <Course name> Chapter name: <Chapter name> Discussion content: <Specific discussion content>."
[0083] It should be noted that in step S4 of the present invention, the generation process of the cognitive feature matrix is specifically as follows: perform one-hot encoding on the cognitive level classification results to obtain the cognitive level encoding results, form the cognitive feature vector of the learner from the cognitive level encoding results of all chapters, use the cognitive feature vector of one learner as a row in the cognitive feature matrix, and form a complete cognitive feature matrix from the cognitive feature vectors of all learners.
[0084] S5: Input a second prompt word group containing questions and answers and second additional questions into the fine-tuned large language model. The fine-tuned large language model judges whether the course behavior types of two learners are the same, outputs the course behavior type judgment result according to the answer in the second prompt word group, encodes the course behavior type judgment result, and uses the course behavior type encoding result as the pseudo-label of each edge in the final discussion graph.
[0085] It should be noted that in step S5 of the present invention, both the questions in the second example and the second additional question contain the course behavior information data of two learners and determine whether the course behavior types of the two learners are the same. The answer in the second example is designed in the way of the chain of thought prompt, and the answer in the second example contains the judgment result of the course behavior type and the logical reasoning process of judging whether the course behavior types are the same; among them, when both learners are dropouts or both learners are non-dropouts, it is considered that the course behavior types of the two learners are the same.
[0086] In this embodiment, the designed second prompt is as follows: "Question: The following are the behavior information of two learners in a MOOC course. Please help me determine whether they belong to the same type of learners, that is, whether they are both dropouts or non-dropouts, and give your reasons for judgment. Learner A: Total viewing duration of course videos: 300 minutes, Total duration of course videos: 500 minutes; Course access frequency: 47 times; Assignment submission status: 1, 1, 1, 1; Test situation: 87, 95, 90; Learner B: Total viewing duration of course videos: 200 minutes; Total duration of course videos: 500 minutes; Course access frequency: 16 times; Assignment submission status: 0, 1, 0, 1; Test situation: 54, 0, 63. Answer: Video viewing situation: Learner A watched 60% of the total course duration (300 / 500), while Learner B only watched 40% (200 / 500). This reflects that Learner A has better learning participation. Course access frequency: The access frequency of Learner A is 3 times per week, while that of Learner B is 1 time per week. More frequent access behavior may indicate stronger learning motivation. Assignment submission status: Learner A submitted all assignments, while the assignment submission status of Learner B is unstable, with records of non-submission. Submitting assignments is an important indicator of participation. Test scores: The test scores of Learner A are generally excellent, which implies better learning effects. The test scores of Learner B fluctuate greatly, with records of not taking the exam, which may be a sign of insufficient learning participation. Conclusion: Based on the above analysis, Learner A shows a high level of learning participation and learning effects and is unlikely to be a dropout. While Learner B is unstable in terms of video viewing, course access frequency, assignment submission, and test scores, showing a low level of learning participation and may belong to the category of dropouts. Therefore, they do not belong to the same type of learners. Question: The following are the behavior information of two learners in a MOOC course. Please help me determine whether they belong to the same type of learners, that is, whether they are both dropouts or non-dropouts, and give your reasons for judgment. Learner A: Total viewing duration of course videos: <Total viewing duration>, Total duration of course videos: <Total duration>, Course access frequency: <Access frequency>, Assignment submission status: 0 indicates non-submission, 1 indicates submission, and the assignment submission sequence is <1, 1...1>, <Test situation>: 0 indicates not taking the exam, and the exam score sequence is <87, 95...90>; Learner B: Total viewing duration of course videos: <Total viewing duration>, Total duration of course videos: <Total duration>, Course access frequency: <Access frequency>, Assignment submission status: 0 indicates non-submission, 1 indicates submission, and the assignment submission sequence is <0, 1...0>, <Test situation>: 0 indicates not taking the exam, and the exam score sequence is <54, 0...63>."
[0087] It should be noted that in step S5 of the present invention, when encoding the judgment result of the course behavior type, when the course behavior types of two learners are the same, it is encoded as 1, and when the course behavior types of two learners are different, it is encoded as 0, thereby obtaining the course behavior type encoding result.
[0088] S6: Use the cognitive feature matrix of the learner and the final discussion graph together as the input of the MOOC dropout prediction model to train the MOOC dropout prediction model.
[0089] In step S6, during the training process of the MOOC dropout prediction model, a course behavior information matrix is constructed from the course behavior information data of all learners. After splicing the course behavior information matrix and the cognitive feature matrix of the learner, it is used as the learner feature matrix. Two multi-layer perceptrons are used to output the assortativity probability of each edge according to the edge set in the final discussion graph, and the assortativity probability is adjusted using the pseudo-label of each edge in the final discussion graph to obtain the adjusted assortativity probability. The adjusted assortativity probability is used as an element in the low-pass adjacency matrix to construct the low-pass adjacency matrix. The learner feature matrix is input into the simplified graph convolutional network. First, the learner feature matrix is processed by the third multi-layer perceptron to output the initial low-pass feature matrix. The initial low-pass feature matrix is subjected to multiple low-pass filtering processes. In each low-pass filtering process, the result of the previous low-pass filtering process is multiplied by the low-pass adjacency matrix to obtain the current low-pass filtering result. The result of the last low-pass filtering process is used as the learner low-pass feature matrix. The result of subtracting 1 from the adjusted assortativity probability is used as an element in the high-pass adjacency matrix to construct the high-pass adjacency matrix. The identity matrix is subtracted from the weighted high-pass adjacency matrix to obtain the processed high-pass adjacency matrix. The learner feature matrix is input into the Lap-SGC model. First, the learner feature matrix is processed by the fourth multi-layer perceptron to output the initial high-pass feature matrix. The initial high-pass feature matrix is subjected to multiple high-pass filtering processes. In each high-pass filtering process, the result of the previous high-pass filtering process is multiplied by the processed high-pass adjacency matrix to obtain the current high-pass filtering result. The result of the last high-pass filtering process is used as the learner high-pass feature matrix. The learner low-pass feature matrix and the learner high-pass feature matrix are spliced together to form the learner comprehensive representation matrix. After performing position space embedding on the final discussion graph, the nearest neighbor graph is obtained. The nearest neighbor graph is input into the graph convolutional network to obtain the learner encoding matrix. The learner comprehensive representation matrix and the learner encoding matrix are spliced together and input into the fifth multi-layer perceptron to output the prediction probability that each learner belongs to a dropout student. Based on the prediction probability that each learner belongs to a dropout student and the true course behavior label of the learner, the cross-entropy loss between the two is calculated. Based on minimizing the cross-entropy loss, the parameters of the MOOC dropout prediction model are updated until the preset iteration round threshold is reached, and the MOOC dropout prediction model converges to obtain the trained MOOC dropout prediction model.
[0090] It should be noted that in the present invention, the specific process of using two multi-layer perceptrons to output the assortativity probability of each edge is as follows: The learner feature vectors of the two nodes corresponding to each edge are respectively passed through the first multi-layer perceptron (Multilayer Perceptron, MLP), and the corresponding outputs are the first intermediate representation vector and the second intermediate representation vector representing the two learners respectively. After concatenating the first intermediate representation vector and the second intermediate representation vector and passing them through the second multi-layer perceptron, a third intermediate representation vector is obtained. After concatenating the second intermediate representation vector and the first intermediate representation vector and passing them through the second multi-layer perceptron, a fourth intermediate representation vector is obtained. The average of the third intermediate representation vector and the fourth intermediate representation vector is calculated to obtain the assortativity probability corresponding to each edge; wherein, the learner feature vector is a row in the learner feature matrix.
[0091] In this embodiment, for the edge connecting learner a and learner b in the final discussion graph , the assortativity probability of the edge is estimated according to the following formula:
[0092]
[0093]
[0094]
[0095] wherein, represents the matrix concatenation operation; and respectively represent the feature vectors of learner a and learner b, and both are from a row in the learner feature matrix; and respectively represent the intermediate representation vectors obtained after the feature vectors of learner a and learner b are processed by embedding; and represent two multi-layer perceptrons. The present invention represents the undirectedness of the edges in the discussion graph data by performing feature embedding on and together. The finally obtained assortativity probability of the edge is the probability that the two learners connected by the edge in the final discussion graph belong to the same type.
[0096] It should be noted that in the present invention, after obtaining the assortativity probability of each edge, the assortativity probability is further adjusted. When the course behavior types of the two learners are different, the minimum value between and 0 is taken as the adjusted assortativity probability; when the course behavior types of the two learners are the same, then between Take the maximum value between and 0 as the adjusted assortativity probability; is the assortativity probability before adjustment. The above specific process is expressed by the formula:
[0097]
[0098] where is the hyperparameter used to adjust the assortativity probability of the edge. In this embodiment, takes values between 0 and 0.5; is the adjusted assortativity probability; represents the pseudo-label of the edge corresponding to node (learner) a and node (learner) b. When it means that the course behavior types of the two learners are the same. When it means that the course behavior types of the two learners are different; and represent the maximum function and the minimum function respectively.
[0099] In this embodiment, when the large language model recognizes an edge connecting learners with the same course behavior type, the assortativity probability value of the edge needs to be increased accordingly to strengthen the model's recognition of the consistency of learners' course behavior types. On the contrary, if the edge connects learners with different course behavior types, the assortativity probability value of this edge will be decreased to reflect the model's recognition of the difference in learners' course behavior types.
[0100] It should be noted that in the present invention, according to the adjusted edge assortativity probability, the present invention performs adaptive encoding on the final discussion graph. Specifically, using a basic low-pass graph neural network (GNN), namely a simplified graph convolution network (Simplified Graph Convolution, abbreviated as SGC), a low-pass filtering process is performed on the learner feature matrix This process aims to capture and retain the similarity feature information between adjacent nodes. The low-pass filtering operation of the simplified graph convolution network can be expressed as:
[0101]
[0102]
[0103] where is the low-pass adjacency matrix. The element in the a-th row and b-th column of the low-pass adjacency matrix is the adjusted assortativity probability of the edge connecting learner a and learner b, that is ; represents the initial low-pass feature matrix; represents the third multi-layer perceptron; and respectively represent the low-pass filtering results obtained at the -th and the -th times; is the filtering times index; the low-pass filtering result obtained at the -th time is the learner's low-pass feature matrix.
[0104] After completing the low-pass filtering, the present invention uses a Lap-SGC model to perform high-pass filtering on the learner feature matrix, so as to capture the differential feature information among learners. First, construct a high-pass adjacency matrix . The element in the a-th row and b-th column of the high-pass adjacency matrix is the result of subtracting 1 from the adjusted assortativity probability of the edge connecting learner a and learner b, that is . Then start to perform times of high-pass filtering process.
[0105]
[0106]
[0107]
[0108] Among them, is the processed high-pass adjacency matrix; is the hyperparameter for controlling the filtering intensity; represents the initial high-pass feature matrix; represents the fourth multi-layer perceptron; and respectively represent the high-pass filtering results obtained at the -th and the -th times; the high-pass filtering result obtained at the -th time is the learner's high-pass feature matrix.
[0109] It should be noted that in the present invention, the homophilic nodes of the learner nodes are searched within the entire graph to supplement the comprehensive representation of the learners. This process can, to a certain extent, overcome the heterophily of the discussion graph data. Therefore, the present invention first performs a positional space embedding on the final discussion graph to obtain the nearest neighbor graph. The specific process is as follows: First, an embedding operation is performed on each node in the final discussion graph, mapping it into an implicit dimension of a positional space. The average value of the learner feature vectors of all the first-order neighbor nodes of each node in the final discussion graph is used as the positional encoding result of that node. The L2 distance between the positional encoding result of each node and the positional encoding results of the remaining nodes in the final discussion graph is calculated. The M nodes with the smallest L2 distance are used as the nearest neighbor nodes of each node. Each node in the final discussion graph is connected to its respective nearest neighbor nodes, and the learner feature matrix is used as the node features to construct the nearest neighbor graph.
[0110] In this embodiment, for the nodes in the final discussion graph , the specific form of its positional encoding result is as follows:
[0111]
[0112] In the formula, represents the positional encoding result of the node in the final discussion graph; represents the degree of the node with a self-loop; represents the set of first-order neighbors of the node with a self-loop; represents a first-order neighbor node in the set of first-order neighbors; represents the first-order neighbor node corresponding to the learner feature vector.
[0113] In this embodiment, based on the final discussion graph, a nearest neighbor graph is constructed from the positional encoding results of each node, that is, each node is connected to M other nodes with the closest L2 distance. Each node in the nearest neighbor graph is only connected to the nodes with a closer positional encoding to itself, which, to a certain extent, overcomes the heterophily of the final discussion graph.
[0114] Generally speaking, in the present invention, as Figure 3As shown, due to the disassortativity presented in the structure and content of the final discussion graph, general graph data processing methods may not perform well. Therefore, the present invention first estimates the assortativity degree of the edges in the final discussion graph data, and then adaptively encodes the nodes in the final discussion graph according to the assortativity probability of the edges. The cognitive feature matrix obtained in S4 is used as a supplement to the learner node features in the final discussion graph, and the pseudo-labels obtained in S5 are used to adjust the node adaptive encoding. After splicing the course behavior information matrix and the learner's cognitive feature matrix as the learner feature matrix, through the above process, the learner comprehensive representation matrix and the learner encoding matrix are obtained respectively. After splicing these two matrices and inputting them into a multi-layer perceptron, the prediction analysis of learner dropout is carried out, and finally the prediction probability that each learner belongs to a dropout is output. Then, the present invention measures the deviation between the prediction probability and the learner's true course behavior label through cross-entropy loss. For each learner, there corresponds a learner's true course behavior label, whose value of 1 indicates that the learner has dropped out of school, while its value of 0 indicates that the learner has not dropped out of school, that is, the learner has successfully completed the study of this course. Then the functional form of the above cross-entropy loss is as follows:
[0115]
[0116] wherein, represents the number of learners; represents the th prediction probability that the learner belongs to a dropout; represents the th true course behavior label of the learner.
[0117] S7: Input the course behavior information data of the learner to be predicted into the trained MOOC dropout prediction model, and output the prediction probability that the learner belongs to a dropout.
[0118] Next, the present invention will use a specific example to demonstrate the application effect of the MOOC dropout prediction method for large language model assisted graph node classification described in S1~S7 in the above embodiments on a specific data set, so as to facilitate the understanding of the essence of the present invention.
[0119] Embodiment
[0120] The specific implementation process of the MOOC dropout prediction method for large language model assisted graph node classification adopted in this embodiment is as described above and will not be elaborated.
[0121] To demonstrate the technical effects of the method proposed in the present invention, the method of the present invention is verified in an actual MOOC scenario. The data selected in the present invention is an educational dataset, and the data is collected from a real online course learning area. Students learn courses through videos in this course area, complete the homework after each section and the unit tests after each unit, and students need to post and reply posts in the course discussion area to obtain the usual course scores. Based on this data, in this embodiment, the course behavior information of students is encoded and used as the features of the learner nodes in the graph. Taking learners as nodes, edges connecting learners are constructed according to the post and reply records of students in the discussion area, that is, each edge represents the communication situation of two students under the same post, and a learner discussion graph is constructed in this way. The accuracy is selected as the index to evaluate the model effect, and the value range of the accuracy is from 0 to 1. The larger the accuracy value, the better the model effect. The calculation formula of the above accuracy is as follows:
[0122]
[0123] Among them, True Positive (TP) represents the number of positive classes predicted as positive classes, True Negative (TN) represents the number of negative classes predicted as negative classes, False Positive (FP) represents the number of negative classes predicted as positive classes, and False Negative (FN) represents the number of positive classes predicted as negative classes.
[0124] The comparison between the method proposed in the present invention and the benchmark machine learning model is shown in Table 1. For fairness, each method is experimented 5 times, and the average value of the results is taken as the final dropout prediction result. In Table 1, the implementation processes of the three methods, namely the method Attn introducing the attention mechanism and conditional random field, the method EMIF based on multi-view interaction, and the method HMM based on the hidden Markov model, all belong to the prior art and will not be elaborated here. The experimental results in Table 1 show that the method proposed in the present invention has achieved the optimal result compared with the benchmark method. A main reason is that the method proposed in the present invention utilizes the rich information in the course discussion area.
[0125] Table 1. Dropout prediction results of different methods on the educational dataset
[0126] Method ACC (%) Attn 90.83 EMIF 89.55 HMM 91.34 The present invention 93.32
[0127] In addition, an ablation study was also conducted in this embodiment to better demonstrate the contribution of the innovative points proposed by the present invention. The experimental results are shown in Table 2. In Table 2, "W / O R.A." represents the process of obtaining the cognitive feature matrix in ablation step S4, and only the course behavior information matrix composed of the course behavior feature data of the learner is used as the learner feature matrix; "W / O E.D." represents the process of obtaining the pseudo-labels of each edge in the final discussion graph in ablation step S5, and during the training process of the MOOC dropout prediction model, the assortativity probability of each edge in the final discussion graph is not adjusted, and only the unadjusted assortativity probability is used for low-pass and high-pass filtering; "W / O H.W." represents the process of obtaining the nearest neighbor graph by performing positional space embedding on the final discussion graph, and only the classical graph neural network method GCN is used for graph representation learning and node classification.
[0128] Table 2. Ablation Experiment Results of the Innovative Modules of the Present Invention on the Educational Dataset
[0129] Method ACC (%) Without R.A. 88.23 Without E.D. 91.83 Without H.W. 85.46 The present invention 93.32
[0130] The ablation experiment results show that all three innovative parts proposed by the present invention have significant effects. The main reason is that the cognitive level analysis and pseudo-label discrimination processes can utilize the excellent text reading and analysis capabilities of the large language model to supplement the learner node features and the structural information in the final discussion graph; the process of obtaining the nearest neighbor graph can effectively handle the disassortativity problem caused by the uneven levels among the learners who communicate with each other in the discussion area. In summary, the method of the present invention can provide an efficient and accurate student dropout detection solution for the online education scenario.
[0131] It should be further noted that the above-mentioned MOOC dropout prediction method using the large language model to assist graph node classification can essentially be executed by a computer program or module. Therefore, similarly, based on the same inventive concept, another preferred embodiment of the present invention also provides a large language model-assisted graph node classification-based MOOC dropout warning system corresponding to the above-mentioned large language model-assisted graph node classification-based MOOC dropout prediction method, as Figure 4 shown, which includes:
[0132] A data acquisition module, configured to obtain the course behavior information data of each learner in the target course from the MOOC platform, encode the preprocessed course behavior information data to obtain the course behavior feature data of the learner;
[0133] A discussion graph construction module, configured to use each learner as a node in the discussion graph, use the interaction relationship between learners in the discussion area as the edge connecting the nodes in the discussion graph, and use the course behavior feature data of the learner as the node feature to obtain an initial discussion graph, and process the initial discussion graph to obtain a final discussion graph;
[0134] A model fine-tuning module, which is used to replace the random initialization process of the LoRA parameter matrix in the O-LoRA method with the randomized singular value decomposition method, and fine-tune the pre-trained large language model based on the LoRA parameter matrix after the randomized singular value decomposition by the O-LoRA method to obtain the fine-tuned large language model;
[0135] An cognitive level analysis module, which is used to input a first prompt word composed of a set of first examples including questions and answers and a first additional question into the fine-tuned large language model, and the fine-tuned large language model outputs the corresponding cognitive level classification result of the learner under the target chapter according to the answers in the first examples, and obtains the cognitive feature matrix based on the cognitive level classification result;
[0136] A pseudo-label acquisition module, which is used to input a second prompt word composed of a set of second examples including questions and answers and a second additional question into the fine-tuned large language model, and the fine-tuned large language model judges whether the course behavior types of two learners are the same, and outputs the course behavior type judgment result according to the answers in the second examples, and encodes the course behavior type judgment result, and uses the course behavior type encoding result as the pseudo-label of each edge in the final discussion graph;
[0137] A model training module, which is used to jointly use the learner's cognitive feature matrix and the final discussion graph as the input of the MOOC dropout prediction model to train the MOOC dropout prediction model;
[0138] A result acquisition module, which is used to input the course behavior information data of the learner to be predicted into the trained MOOC dropout prediction model and output the prediction probability that the learner belongs to a dropout;
[0139] A warning prompt module, which is used to send dropout warning information to the teacher users or management users of the MOOC platform, or send reminder information to the learners who belong to dropouts to urge them to log in and study.
[0140] It should be noted that in the warning prompt module, the predicted probability corresponding to the learner can be used to judge whether each learner belongs to a dropout. When the predicted probability is greater than the preset probability threshold, it is considered that the learner belongs to a dropout or may belong to a dropout. When the predicted probability is less than or equal to the preset probability threshold, it is considered that the learner does not belong to a dropout. Therefore, warning prompts can be made according to the predicted probability output by the result acquisition module. Specifically, dropout warning information can be sent to teacher users or management users to inform them that there may be learners dropping out of the course. Reminder information can also be sent to the learners who belong to dropouts to urge them to log in and study, so as to reduce the number of dropouts and improve the completion rate of the course and the overall effect of online education.
[0141] In addition, it should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the system described above can refer to the corresponding process in the foregoing method embodiments, and will not be elaborated here. In the embodiments provided in the present application, the division of steps or modules in the system and method is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or steps can be combined or integrated together, and a module or step can also be split.
[0142] In addition, it should be noted that the course behavior information data of the learners involved in the present invention is obtained with full consent and authorization, and the collection, use, and processing of relevant information need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0143] The above-described embodiments are only a preferred solution of the present invention, but they are not intended to limit the present invention. Those of ordinary skill in the relevant technical field can still make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all technical solutions obtained by means of equivalent replacement or equivalent transformation fall within the protection scope of the present invention.
Claims
1. A massive open online course (MOOC) dropout prediction method for graph node classification assisted by large language models, characterized in that, It includes the following steps: S1: Obtain the course behavior information data of each learner in the target course from the MOOC platform, encode the preprocessed course behavior information data, and obtain the course behavior feature data of the learners; S2: Take each learner as a node in the discussion graph, take the interaction relationship between learners in the discussion area as the edge connecting the nodes in the discussion graph, and take the course behavior feature data of the learners as the node features to obtain the initial discussion graph, and process the initial discussion graph to obtain the final discussion graph; S3: Use the randomized singular value decomposition method to replace the random initialization process of the LoRA parameter matrix in the O-LoRA method, and fine-tune the pre-trained large language model based on the LoRA parameter matrix after randomized singular value decomposition by the O-LoRA method to obtain the fine-tuned large language model; S4: Input a first prompt word composed of a first example containing questions and answers and a first additional question into the fine-tuned large language model. The fine-tuned large language model outputs the corresponding cognitive level classification results of the learners under the target chapter according to the answers in the first example, and obtains the cognitive feature matrix based on the cognitive level classification results; S5: Input a second prompt word composed of a second example containing questions and answers and a second additional question into the fine-tuned large language model. The fine-tuned large language model judges whether the course behavior types of two learners are the same, outputs the course behavior type judgment result according to the answers in the second example, encodes the course behavior type judgment result, and uses the course behavior type encoding result as the pseudo label of each edge in the final discussion graph; S6: Use the cognitive feature matrix of the learners and the final discussion graph as the input of the MOOC dropout prediction model, and train the MOOC dropout prediction model; S7: Input the course behavior information data of the learner to be predicted into the trained MOOC dropout prediction model, and output the prediction probability that the learner belongs to a dropout student; In step S6, during the training process of the MOOC dropout prediction model, a course behavior information matrix is formed from the course behavior information data of all learners. The course behavior information matrix is concatenated with the cognitive feature matrix of the learners to obtain the learner feature matrix. Two multi-layer perceptrons are used to output the assortativity probability of each edge according to the edge set in the final discussion graph. The assortativity probability is adjusted using the pseudo-labels of each edge in the final discussion graph to obtain the adjusted assortativity probability. The adjusted assortativity probability is used as an element in the low-pass adjacency matrix to construct the low-pass adjacency matrix. The learner feature matrix is input into the simplified graph convolutional network. First, the learner feature matrix is processed by the third multi-layer perceptron to output the initial low-pass feature matrix. The initial low-pass feature matrix is processed through multiple low-pass filtering operations. In each low-pass filtering operation, the result of the previous low-pass filtering operation is multiplied by the low-pass adjacency matrix to obtain the current low-pass filtering result. The result of the last low-pass filtering operation is used as the learner low-pass feature matrix. The result of subtracting 1 from the adjusted assortativity probability is used as an element in the high-pass adjacency matrix to construct the high-pass adjacency matrix. The identity matrix is subtracted from the weighted high-pass adjacency matrix to obtain the processed high-pass adjacency matrix. The learner feature matrix is input into the Lap-SGC model. First, the learner feature matrix is processed by the fourth multi-layer perceptron to output the initial high-pass feature matrix. The initial high-pass feature matrix is processed through multiple high-pass filtering operations. In each high-pass filtering operation, the result of the previous high-pass filtering operation is multiplied by the processed high-pass adjacency matrix to obtain the current high-pass filtering result. The result of the last high-pass filtering operation is used as the learner high-pass feature matrix. The learner low-pass feature matrix and the learner high-pass feature matrix are concatenated to obtain the learner comprehensive representation matrix. The final discussion graph is embedded in the position space to obtain the nearest neighbor graph. The nearest neighbor graph is input into the graph convolutional network to obtain the learner encoding matrix. The learner comprehensive representation matrix and the learner encoding matrix are concatenated and input into the fifth multi-layer perceptron to output the prediction probability that each learner belongs to a dropout student. The cross-entropy loss between the prediction probability that each learner belongs to a dropout student and the true course behavior label of the learner is calculated. The parameters of the MOOC dropout prediction model are updated based on minimizing the cross-entropy loss until the preset iteration round threshold is reached and the MOOC dropout prediction model converges, obtaining the trained MOOC dropout prediction model; The specific process of using two multi-layer perceptrons to output the assortativity probability of each edge is as follows: The learner feature vectors of the two nodes corresponding to each edge are respectively passed through the first multi-layer perceptron, and the corresponding outputs are the first intermediate representation vector and the second intermediate representation vector representing the two learners. After concatenating the first intermediate representation vector and the second intermediate representation vector and passing them through the second multi-layer perceptron, a third intermediate representation vector is obtained. After concatenating the second intermediate representation vector and the first intermediate representation vector and passing them through the second multi-layer perceptron, a fourth intermediate representation vector is obtained. The average of the third intermediate representation vector and the fourth intermediate representation vector is calculated to obtain the assortativity probability corresponding to each edge; among them, the learner feature vector is a row in the learner feature matrix; The specific process of adjusting the assortativity probability is as follows: when the course behavior types of two learners are different, take the minimum value between and 0 as the adjusted assortativity probability; when the course behavior types of two learners are the same, take the maximum value between and 0 as the adjusted assortativity probability; is the assortativity probability before adjustment; is a preset hyperparameter.
2. The massive open online course dropout prediction method for graph node classification assisted by large language models according to claim 1, wherein In step S1, each course behavior information data contains N video teaching videos for learners to watch, N work assignment completion statuses, N test test statuses of tests, the total number of times N visit the learner accesses the target course, and text data composed of all the learner's speeches in the target course discussion area. After preprocessing the text data, preprocessed course behavior information data is formed; In step S1, the specific process of encoding the preprocessed course behavior information data is as follows: S11: For each teaching video of each learner, the ratio between the viewing duration of the learner and the total duration of the corresponding teaching video is used as the viewing duration encoding value of the teaching video, and all the viewing duration encoding values of the learner form the corresponding viewing duration encoding vector; S12: The total number of times a learner accesses the target course is used as the access frequency encoding vector corresponding to the learner; S13: For each homework completion situation of each learner, if the learner completes the homework, the ratio between the number of correctly solved problems and the total number of problems under each homework of the learner is used as the homework submission situation encoding value. If the learner does not complete the homework, the homework submission situation encoding value is 0. Finally, all the homework submission situation encoding values of the learner form the corresponding homework submission situation encoding vector; S14: For each test situation of each learner, if the learner completes the test, the ratio between the test score of the learner and the total score of the corresponding test is used as the test score encoding value. If the learner does not complete the test, the test score encoding value is 0. All the test score encoding values of the learner form the corresponding test score encoding vector; S15: For the preprocessed text data of each learner, the preprocessed text data is fed into the large language model with visible encoding to obtain the discussion content encoding vector of the learner; S16: The viewing duration encoding vector, the access frequency encoding vector, the homework submission situation encoding vector, the test score encoding vector, and the discussion content encoding vector are concatenated as the course behavior feature data of the learner.
3. The massive open online course dropout prediction method for graph node classification assisted by large language models according to claim 1, characterized in that, In step S2, the specific process of obtaining the final discussion graph is as follows: Determine whether there are isolated nodes in the initial discussion graph: If so, traverse each isolated node in the initial discussion graph, calculate the cosine similarity between the node feature corresponding to an isolated node and the node features of the other nodes, and use the K nodes with the largest cosine similarity as the most similar nodes of the isolated node. Connect the isolated node to its corresponding most similar nodes, form K edges and add them to the initial discussion graph to obtain the final discussion graph; If not, the initial discussion graph is used as the final discussion graph.
4. The MOOC dropout prediction method for large language model-assisted graph node classification according to claim 1, wherein In step S3, the specific process of obtaining the LoRA parameter matrix after randomized singular value decomposition using the randomized singular value decomposition method is as follows: S31: Randomly sample from a normal distribution as the column vectors of a random matrix to generate a complete random matrix; S32: Use the random matrix to project the parameter matrix of the large language model into a low-dimensional space to obtain a projection matrix; perform an orthogonal triangular decomposition on the projection matrix to obtain a first orthogonal matrix and an upper triangular matrix; multiply the transpose of the first orthogonal matrix by the parameter matrix of the large language model to obtain a low-dimensional matrix; Perform a singular value decomposition on the low-dimensional matrix to obtain a second orthogonal matrix, a third orthogonal matrix, and a diagonal matrix; S33: Use the indices corresponding to the largest r singular values in the diagonal matrix as the initialization indices, use the column vectors corresponding to the initialization indices in the second orthogonal matrix as the initialized first LoRA parameter matrix, use the column vectors corresponding to the initialization indices in the third orthogonal matrix as the initialized second LoRA parameter matrix, and form a set of LoRA parameter matrices after randomized singular value decomposition from the initialized first LoRA parameter matrix and the initialized second LoRA parameter matrix.
5. The method for predicting the dropout of massive open online courses by using a large language model to assist in graph node classification according to claim 1, wherein In step S4, both the questions in the first example and the first additional question include all the speeches of the learner in the target chapter of the target course discussion area, the name of the target course, and the name of the target chapter, judge the cognitive level to which the learner belongs according to the learner's speech, the cognitive level of the online community discussion, and the judgment criteria corresponding to each cognitive level; design the answer in the first example in the way of chain-of-thought prompting, and the answer in the first example includes the cognitive level classification result of the learner and the logical reasoning process for judging the cognitive level to which the learner belongs; In step S4, the generation process of the cognitive feature matrix is specifically as follows: perform one-hot encoding on the cognitive level classification result to obtain a cognitive level encoding result, form the cognitive feature vector of the learner from the cognitive level encoding results of all chapters, use the cognitive feature vector of one learner as a row in the cognitive feature matrix, and form a complete cognitive feature matrix from the cognitive feature vectors of all learners.
6. The massive open online course dropout prediction method for graph node classification assisted by large language models according to claim 1, characterized in that In step S5, both the questions in the second example and the second additional question include the course behavior information data of two learners and judge whether the course behavior types of the two learners are the same, design the answer in the second example in the way of chain-of-thought prompting, and the answer in the second example includes the judgment result of the course behavior type and the logical reasoning process for judging whether the course behavior types are the same; among them, when both learners belong to dropouts or both belong to non-dropouts, it is considered that the course behavior types of the two learners are the same.
7. The massive open online course dropout prediction method for large language model-assisted graph node classification according to claim 1, wherein, The specific process of obtaining the nearest neighbor graph after performing position space embedding on the final discussion graph is as follows: First, perform an embedding operation on each node in the final discussion graph, map it to an implicit dimension of the position space, take the average value of the learner feature vectors of all first-order neighbor nodes of each node in the final discussion graph as the position encoding result of the node, calculate the L2 distance between the position encoding result of each node and the position encoding results of the remaining nodes in the final discussion graph, take the M nodes with the smallest L2 distance as the nearest neighbor nodes of each node, connect each node in the final discussion graph with its nearest neighbor nodes, and use the learner feature matrix as the node feature to construct the nearest neighbor graph.
8. A massive open online course (MOOC) dropout warning system assisted by a large language model for graph node classification, characterized in that, Including: A data acquisition module, which is used to obtain the course behavior information data of each learner in the target course from the MOOC platform, encode the preprocessed course behavior information data to obtain the course behavior feature data of the learners; A discussion graph construction module, which is used to take each learner as a node in the discussion graph, take the interaction relationship between learners in the discussion area as the edge connecting nodes in the discussion graph, take the course behavior feature data of the learners as the node feature to obtain the initial discussion graph, and process the initial discussion graph to obtain the final discussion graph; A model fine-tuning module, which is used to replace the random initialization process of the LoRA parameter matrix in the O-LoRA method with the randomized singular value decomposition method, and fine-tune the pre-trained large language model based on the LoRA parameter matrix after the randomized singular value decomposition by the O-LoRA method to obtain the fine-tuned large language model; A cognitive level analysis module, which is used to input a first prompt word consisting of a first example containing questions and answers and a first additional question into the fine-tuned large language model, and the fine-tuned large language model outputs the corresponding cognitive level classification result of the learner under the target chapter according to the answer in the first example, and obtains the cognitive feature matrix based on the cognitive level classification result; A pseudo-label acquisition module, which is used to input a second prompt word consisting of a second example containing questions and answers and a second additional question into the fine-tuned large language model, and the fine-tuned large language model judges whether the course behavior types of two learners are the same, outputs the course behavior type judgment result according to the answer in the second example, encodes the course behavior type judgment result, and takes the course behavior type encoding result as the pseudo-label of each edge in the final discussion graph; A model training module, which is used to jointly use the cognitive feature matrix of the learners and the final discussion graph as the input of the MOOC dropout prediction model to train the MOOC dropout prediction model; A result acquisition module, which is used to input the course behavior information data of the learner to be predicted into the trained MOOC dropout prediction model and output the prediction probability that the learner belongs to a dropout; An early warning prompt module, which is used to send a dropout early warning message to the teacher user or management user of the MOOC platform, or send a reminder message to urge the learner who belongs to a dropout to log in and study; In step S6, during the training process of the MOOC dropout prediction model, a course behavior information matrix is formed from the course behavior information data of all learners. After concatenating the course behavior information matrix with the cognitive feature matrix of the learners, the result is used as the learner feature matrix. Two multi-layer perceptrons are used to output the assortativity probability of each edge according to the edge set in the final discussion graph. The assortativity probability is adjusted using the pseudo-labels of each edge in the final discussion graph to obtain the adjusted assortativity probability. The adjusted assortativity probability is used as the element in the low-pass adjacency matrix to construct the low-pass adjacency matrix. The learner feature matrix is input into the simplified graph convolutional network. First, the learner feature matrix is processed by the third multi-layer perceptron to output the initial low-pass feature matrix. The initial low-pass feature matrix is processed through multiple low-pass filtering operations. In each low-pass filtering operation, the result of the previous low-pass filtering operation is multiplied by the low-pass adjacency matrix to obtain the current low-pass filtering result. The result of the last low-pass filtering operation is used as the learner low-pass feature matrix. The result of subtracting 1 from the adjusted assortativity probability is used as the element in the high-pass adjacency matrix to construct the high-pass adjacency matrix. The identity matrix is subtracted from the weighted high-pass adjacency matrix to obtain the processed high-pass adjacency matrix. The learner feature matrix is input into the Lap-SGC model. First, the learner feature matrix is processed by the fourth multi-layer perceptron to output the initial high-pass feature matrix. The initial high-pass feature matrix is processed through multiple high-pass filtering operations. In each high-pass filtering operation, the result of the previous high-pass filtering operation is multiplied by the processed high-pass adjacency matrix to obtain the current high-pass filtering result. The result of the last high-pass filtering operation is used as the learner high-pass feature matrix. The learner low-pass feature matrix and the learner high-pass feature matrix are concatenated to form the learner comprehensive representation matrix. The final discussion graph is embedded in the position space to obtain the nearest neighbor graph. The nearest neighbor graph is input into the graph convolutional network to obtain the learner encoding matrix. The learner comprehensive representation matrix and the learner encoding matrix are concatenated and input into the fifth multi-layer perceptron to output the prediction probability that each learner belongs to a dropout student. Based on the prediction probability that each learner belongs to a dropout student and the true course behavior label of the learner, the cross-entropy loss between the two is calculated. The parameters of the MOOC dropout prediction model are updated based on minimizing the cross-entropy loss until the preset iteration round threshold is reached and the MOOC dropout prediction model converges, obtaining the trained MOOC dropout prediction model; The specific process of using two multi-layer perceptrons to output the assortativity probability of each edge is as follows: The learner feature vectors of the two nodes corresponding to each edge are respectively passed through the first multi-layer perceptron, and the corresponding outputs are the first intermediate representation vector and the second intermediate representation vector representing the two learners. After concatenating the first intermediate representation vector and the second intermediate representation vector and passing them through the second multi-layer perceptron, a third intermediate representation vector is obtained. After concatenating the second intermediate representation vector and the first intermediate representation vector and passing them through the second multi-layer perceptron, a fourth intermediate representation vector is obtained. The third intermediate representation vector and the fourth intermediate representation vector are averaged to obtain the assortativity probability corresponding to each edge; among them, the learner feature vector is a row in the learner feature matrix; The specific process of adjusting the assortativity probability is as follows: when the course behavior types of two learners are different, take the minimum value between and 0 as the adjusted assortativity probability; when the course behavior types of two learners are the same, take the maximum value between and 0 as the adjusted assortativity probability; is the assortativity probability before adjustment; is a preset hyperparameter.
Citation Information
Patent Citations
Mooc learner learning effect prediction method based on multi-level enhanced contrast learning
CN118227791A
Large language model acceleration method and device
CN118569324A