Online learning early warning attribution method and system based on large language model and conditional residual causal inference
By combining large language models with conditional residual causal inference, an online learning early warning system was constructed, which solved the problems of difficulty in identifying causal relationships and low accuracy in dropout prediction in online learning. It provides interpretable dropout risk early warning and intervention strategies, and improves the effect of reducing dropout rates.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGNAN UNIV
- Filing Date
- 2026-05-11
- Publication Date
- 2026-07-31
AI Technical Summary
Existing online learning early warning and attribution technologies struggle to balance semantic understanding depth with online reasoning efficiency. They are susceptible to confounding factors due to correlation analysis, lack interpretable causal transmission paths, and fail to effectively integrate prior educational causal knowledge into data-driven modeling, resulting in low dropout prediction accuracy and intervention efficiency.
By employing a large language model to extract causal prior knowledge and combining it with conditional residual causal inference, a knowledge-driven prior causal graph and a data-driven residual causal graph are constructed. By integrating causal graph search, root cause localization and intervention targets are achieved, providing interpretable dropout risk warnings.
It enables the introduction of expert knowledge in the field of education at low cost, effectively distinguishes between causal relationships and spurious correlations, enhances the confidence of causal margins, provides an end-to-end explainable transmission path of dropout risk, and optimizes the allocation of educational resources.
Smart Images

Figure CN122491499A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and educational information technology, and in particular to an online learning early warning attribution method and system based on large language models and conditional residual causal inference. Background Technology
[0002] With the rapid proliferation of massive online open courses (MOOCs) and online learning platforms, learning management systems have accumulated multimodal learning behavior data, including login logs, video viewing, assignment submissions, forum interactions, and academic performance. This data provides a crucial foundation for assessing learners' status and enabling early warning of dropout risks. However, faced with high dropout rates and increasingly complex online learning behavior patterns, existing early warning and attribution technologies still have several fundamental limitations, making it difficult to simultaneously meet the multiple demands for predictive accuracy, causal interpretability, and deployment efficiency.
[0003] First, existing models struggle to strike a balance between semantic understanding depth and online inference efficiency. To capture deep semantic features from high-dimensional behaviors such as forum text and video interactions, practice increasingly relies on deep learning or large language models with massive parameter sets. However, the inference overhead of these models is too high, making them difficult to deploy on resource-constrained edge servers or in high-concurrency, real-time online teaching scenarios. While lightweight networks or traditional machine learning methods meet the response speed requirements, they often lose key semantic connections and causal relationships in behavioral sequences due to a lack of knowledge representation capabilities compared to large models, resulting in limited accuracy in predicting learning risks.
[0004] Secondly, statistical correlation-based analysis methods are generally prone to attribution bias, failing to distinguish between causal relationships and spurious correlations. Current dropout prediction models are mostly built on correlation modeling, making them susceptible to confounding factors. For example, a positive correlation may exist between reduced video viewing time and dropout tendency, but this association could stem from a common cause—a surge in homework difficulty—rather than viewing time directly leading to dropout. Existing technologies lack effective means to correct causal bias; misclassifying spurious correlations caused by confounding factors as causative agents leads to misleading attribution conclusions for teachers and administrators, causing interventions to deviate from the true source of risk.
[0005] Third, the "black box" nature of the models hinders the discovery of risk transmission mechanisms and the construction of precise intervention pathways. Deep predictive models typically only provide the final dropout probability, failing to reveal how behavioral characteristics ultimately influence learning outcomes through intermediate links. Educators cannot determine whether learners drop out due to decreased perceived usefulness, interaction fatigue, or weak knowledge base. Due to the lack of explicit modeling of mediating variables, warning signals remain superficial, making it difficult to prioritize and provide targeted guidance for numerous risk characteristics, resulting in inaccurate allocation of intervention resources.
[0006] Fourth, existing multimodal feature fusion methods lack adaptive causal knowledge alignment mechanisms. Risk factors that have a critical impact on learners often differ at different learning stages (such as the course introduction period and the final sprint period), but existing technologies mostly employ static weights or simple attention mechanisms, failing to dynamically align prior causal structures in the education field or accurately identify the dominant causal edges among hierarchical behavioral variables. This leads to significant fluctuations in attribution results across different time periods, making it difficult to guarantee the confidence level of causal directions, further weakening the practical value of early warning systems.
[0007] In summary, existing online learning behavior early warning and attribution technologies face four major bottlenecks: an imbalance between semantic depth and efficiency, difficulty in bridging the gap between relevance and causality, lack of interpretability, and insufficient integration of prior knowledge. Therefore, there is an urgent need for a method that can systematically introduce domain-specific prior causal knowledge into the data-driven modeling process, and on this basis, achieve accurate biased attribution and causal path discovery. This would support educators in implementing effective interventions from the root causes and substantially reducing online course dropout rates. Summary of the Invention
[0008] To address these issues, this invention provides an online learning early warning attribution method and system based on a large language model and conditional residual causal inference. This method addresses four key technical bottlenecks in existing technologies: difficulty in balancing semantic understanding depth and online inference efficiency; susceptibility to confounding factors leading to attribution bias in correlation analysis; lack of interpretable causal transmission paths in early warning models; and the inability to effectively integrate causal knowledge from prior education into data-driven modeling.
[0009] To address the aforementioned technical problems, this invention provides an online learning early warning attribution method based on a large language model and conditional residual causal inference. This method includes the following steps: Step S1: Collect learning behavior data generated by learners on the online course platform to obtain... 1 variable and A dataset consisting of 10 samples The variables include behavioral indicators and early warning result indicators; Step S2: Apply the preset large language model from the... Causal prior knowledge is extracted from the semantic information of each variable. A knowledge-driven prior causal graph is constructed through variable semantic hierarchy division, inter-hierarchical causal inference, and prior strength assignment. ; Step S3: Based on conditional residual causal inference, from the dataset In adaptive learning, the causal direction and strength between variables are determined, and a data-driven residual causal graph is constructed. ; Step S4: The prior cause-effect graph With the residual cause-effect graph Perform fusion to obtain a fusion causal graph. and in the fusion causal graph The system performs a causal chain search from the root cause node to the early warning target node, and outputs online learning early warning attribution results containing root cause location, transmission path and intervention target.
[0010] Preferably, in step S2, the application of a preset large language model from the... Causal prior knowledge is extracted from the semantic information of each variable. A knowledge-driven prior causal graph is constructed through variable semantic hierarchy division, inter-hierarchical causal inference, and prior strength assignment. Specifically, it includes: Step S21: Construct suitable large language model prompt words, designate the large language model to play the role of causal inference expert, and automatically complete the variable hierarchy division based on the semantic features of each variable, and simultaneously determine the causal direction and prior strength between different levels. Step S22: Input the variable set, the semantic description of each variable, and the prompt words into the large language model, and the large language model will then process... The variables are automatically divided into semantic levels There is no overlap between the levels and all variables are covered; Step S23: For any two different semantic levels , The large language model infers the causal direction between the two and outputs a direction indicator function. And based on this, directed causal edges are established for all variable pairs between levels; Step S24: Assign prior causal strength values. For causal edges between levels, the large language model generates the prior strength between each pair of variables. For variable pairs belonging to the same level, a pre-defined prior strength is uniformly assigned. ; Step S25: with the aforementioned Using each variable as a node and the causal relationships between and within all levels as directed edges, the knowledge-driven prior causal graph is constructed. ,in For a set of nodes, Let be the causal prior strength matrix.
[0011] Preferably, the direction indication function in step S23 Defined as: like Then the hierarchy For the reason, For the result, for any and All are established by point to Causal edge; like Then the hierarchy For the result, For the reason, establish by point to Causal edge; like Then no directed causal edge is established between the two layers.
[0012] Preferably, the preset prior strength mentioned in step S24 Causal strength values uniformly assigned among variables at the same level; prior strength of variable pairs between levels. Generated by a large language model that integrates variable semantics, hierarchical attribution, and domain knowledge, with a value range of [missing information]. .
[0013] Preferably, in step S3, the conditional residual causal inference is derived from the dataset. In adaptive learning, the causal direction and strength between variables are determined, and a data-driven residual causal graph is constructed. Specifically, it includes: Step S31: For any two variables , Calculate the Pearson correlation coefficient. and significance probability value ,like In an undirected graph with correlations... Set the corresponding edge to 1. Preset significance level; Step S32: For the undirected graph of correlations Each variable pair that exists in K-nearest neighbor regression is used to estimate the given... hour Conditional expectations ; Step S33: Calculation , Relative conditions between ; Step S34: Calculate the squared conditional residuals based on the full sample ,in Indicates the first A sample about Because of, It is the RCR value of the fruit; Step S35: For each pair of variables, compare and The magnitude of the two factors determines the direction of cause and effect; Step S36: For each causal edge that is determined to exist, calculate its confidence level. ; Step S37: Using all variables as nodes, and the causal direction determined in step S35 and the confidence level calculated in step S36 as edge weights, construct the data-driven residual causal graph. ,in For a set of nodes, For a data-driven causal strength matrix, if but Otherwise, it is 0.
[0014] Preferably, the causal direction determination rule in step S35 is: like Then the direction of causality is determined as follows: ; like Then the direction of causality is determined as follows: ; like Then it is marked as a bidirectional edge. .
[0015] Preferably, the confidence level mentioned in step S36 The calculation formula is: ; This ensures that the confidence level is 0.5 when the squares of the residuals in both directions are equal, and the confidence level is higher when the difference is greater, with an upper limit of 0.9.
[0016] Preferably, in step S4, the prior cause-effect graph is... With the residual cause-effect graph Perform fusion to obtain a fusion causal graph. and in the fusion causal graph The system performs a causal chain search from the root cause node to the early warning target node, and outputs online learning early warning attribution results containing root cause location, transmission path, and intervention target, specifically including: Step S41: For each pair of variables Define the causal strength of fusion If both the prior graph and the data graph have causal edges in the same direction and both have a strength greater than 0, then the maximum value is taken; if only one has a strength greater than 0, then the strength of that side is taken; if the two graphs conflict in direction, then the one with the higher confidence score is taken; otherwise, the value is 0. This yields the fused directed weight matrix. and fusion causal diagram ; Step S42: In the fusion causal graph In this context, the root cause node set is constructed using variables whose incoming edge weights are all zero. ; Step S43: For each root cause node The algorithm recursively traverses its successor nodes using a depth-first search. When the current node is a preset warning target node, it is recorded. The complete path to the target node; Step S44: Organize all paths from root causes to target nodes to form a set of causal chains.
[0017] Preferably, in step S4, the fusion causal strength The specific definition is: ,when ,in For the strength of prior causality, Data-driven causal strength; ,when and ; ,when and ; Other situations; When there is a directional conflict, the edge with the greater strength is selected as the edge to be retained and the fusion strength.
[0018] This invention also provides an online learning early warning attribution system based on large language models and conditional residual causal inference. This system is used to implement the aforementioned online learning early warning attribution method based on large language models and conditional residual causal inference, specifically including: The data collection module is used to collect learning behavior data generated by learners on the online course platform, and obtain data from... 1 variable and A dataset consisting of 10 samples The variables include behavioral indicators and early warning result indicators; The large-scale model knowledge distillation and prior causal graph construction module is used to apply a pre-defined large language model to the knowledge distillation and prior causal graph construction module. Causal prior knowledge is extracted from the semantic information of each variable. A knowledge-driven prior causal graph is constructed through variable semantic hierarchy division, inter-hierarchical causal inference, and prior strength assignment. ; The conditional residual causal inference and data causal graph construction module is used to perform conditional residual causal inference from the dataset. In adaptive learning, the causal direction and strength between variables are determined, and a data-driven residual causal graph is constructed. ; The causal graph fusion and early warning attribution output module is used to process the prior causal graph. With the residual cause-effect graph Perform fusion to obtain a fusion causal graph. and in the fusion causal graph The system performs a causal chain search from the root cause node to the early warning target node, and outputs online learning early warning attribution results containing root cause location, transmission path and intervention target.
[0019] As can be seen from the above technical solutions, this invention application has the following beneficial effects: First, by utilizing the pre-trained knowledge of large language models, we can automatically segment the semantic hierarchy of variables and distill prior causal knowledge, thereby introducing expert knowledge in the field of education at a lower cost. Second, by adaptively learning the causal direction from the observed data through conditional residual causal inference, the interference of confounding factors can be effectively removed, and the causal relationship in the correlation can be distinguished from the spurious correlation. Third, by using a selective fusion strategy to integrate prior causal graphs and data-driven causal graphs, the complementarity of knowledge-driven and data-driven approaches was achieved, thereby enhancing the confidence of causal edges. Fourth, depth-first search based on fused causal graphs can discover the complete transmission path from root cause to early warning target, providing interpretable end-to-end causal evidence for educational intervention. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly described below. Referring to the drawings will make the features and advantages of the present invention clearer. The drawings are illustrative and should not be construed as limiting the present invention in any way. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein: Figure 1 This is a flowchart of an online learning early warning attribution method based on large language models and conditional residual causal inference provided by the present invention; Figure 2 This is a block diagram of an online learning early warning attribution system based on a large language model and conditional residual causal inference, provided by the present invention. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] Example 1: To address the four major technical bottlenecks in existing technologies—namely, the difficulty in balancing semantic understanding depth with online inference efficiency, the susceptibility of correlation analysis to confounding factors leading to attribution bias, the lack of interpretable causal transmission paths in early warning models, and the inability to effectively integrate causal knowledge from prior education into data-driven modeling—such as... Figure 1 As shown, this invention proposes an online learning early warning attribution method based on large language models and conditional residual causal inference, which includes the following steps S1 to S4.
[0023] Step S1: Collect learning behavior data generated by learners on the online course platform to obtain... 1 variable and A dataset consisting of 10 samples The variables include behavioral indicators and early warning result indicators.
[0024] Specifically, it involves collecting learning behavior data generated by learners on online course platforms to obtain data from... 1 variable and A dataset consisting of 10 samples The variables include behavioral indicators and early warning result indicators. Let the variable set be... This step provides the basic data input for all subsequent analyses.
[0025] In this embodiment, a learner behavior dataset from an online course platform was used, collecting learning behavior records from 171 learners. The original dataset contains 22 variables, as shown in Table 1 below, specifically including: number of logins, morning login time (08-12), afternoon login time (12-18), evening login time (18-24), early morning login time (24-08), video viewing duration, number of chapter studies, chapter test scores, number of completed assignments, number of times academic procrastination occurred, grades for regular assignment 1, grades for regular assignment 2, grades for regular assignment 3, number of interactive discussions, number of posts published, number of replies to posts, number of interactive topics, number of likes received for interactions, learning satisfaction, final assignment score, and summative evaluation (where T represents pass, F represents fail / failing, serving as an early warning indicator).
[0026] Table 1 Learning Behavior Record
[0027] Step S2: Apply the preset large language model from the... Causal prior knowledge is extracted from the semantic information of each variable. A knowledge-driven prior causal graph is constructed through variable semantic hierarchy division, inter-hierarchical causal inference, and prior strength assignment. .
[0028] This step utilizes the pre-trained knowledge of the large language model to achieve variable semantic understanding and prior causal knowledge distillation, providing a domain prior causal structure from the large model for step S4. Step S2 specifically includes the following sub-steps S21 to S25. Step S21: Constructing adapted prompt words for the large language model to achieve variable semantic understanding and causal prior knowledge extraction. The design of the prompt words follows the following principles: (1) The large language model is designated to play the role of causal inference expert and must be familiar with the business logic and intrinsic relationship of learning behavior in the online education field; (2) The large language model is required to automatically complete the variable hierarchy division based on the semantic features of each learning behavior variable, so that each level contains variables with similar semantics or at the same level of abstraction, and simultaneously determine the causal direction and corresponding prior strength between different levels.
[0029] Step S22: Set the variables The semantic descriptions and prompts of each variable are input into the large language model. The large language model then uses the semantic similarity of the variables and the business logic of online education to... The variables are automatically divided into There are several semantic levels, denoted as the level set. . satisfy: ,and Each of these levels It represents a set of variables with similar semantics or at the same business abstraction level. The level name is automatically generated by the large language model based on the semantics of the variables.
[0030] In this embodiment, the large language model successfully divided the 22 variables into 5 semantic levels, as shown in the following example: level (Learning participation behavior): includes login times, login times in the morning, afternoon, evening, and early morning, video viewing time, and number of chapters studied; level (Academic performance): includes chapter test scores, number of assignments completed, grades for Homework 1, Homework 2, Homework 3, and final assignment; level (Study Procrastination and Time Management): Includes the number of times academic procrastination occurred; level (Interactive Participation): Includes the number of interactive discussions, the number of posts published, the number of replies to posts, the number of interactive topics, and the number of likes received for interactive activities; level (Learning experience and results): This includes learning satisfaction and summative evaluation.
[0031] Step S23: Inter-level causal reasoning. For any two different semantic levels... The large language model infers the causal direction between common sense in the online education field and pre-trained business logic, and outputs a causal direction indicator function. The indicator function is defined as follows: (1) If This indicates a semantic level. For the reason, For the result; at this time and All of them exist from point to The total number of causal edges is . ,in Representing hierarchy The number of variables included; (2) If This indicates a semantic level. For the result, For the reason; at this time and All of them exist from point to Causal edge; (3) If This indicates a semantic level. and There is no prior causal relationship between them, at this time and No directed causal edges are constructed between any two pairs of elements.
[0032] Step S24: Assigning a priori causal strength. This includes two aspects: (1) Assigning causal strength between levels. For those belonging to the cause level... variables and belong to the fruit level variables The causal prior strength between the two is generated by integrating variable semantics, hierarchical attribution, and domain common sense through a large language model. .
[0033] (2) Assigning causal strength within the same hierarchical level. For those belonging to the same semantic level Any two distinct variables inside The causal prior strength is uniformly set as ,in This refers to the pre-defined a priori causal strength values between variables at the same level. .
[0034] Step S25: The set of learning behavior variables Using the set of nodes and the set of directed edges derived from all the causal relationships between variables determined in steps S22 to S24, a knowledge-driven prior causal graph is constructed. .in, This is a causal prior strength matrix, where all elements are filled with the causal prior strengths between variables generated in step S24. If the variables... and If there is no prior causal relationship between them, then the corresponding .
[0035] Step S3: Based on conditional residual causal inference, from the dataset In adaptive learning, the causal direction and strength between variables are determined, and a data-driven residual causal graph is constructed. .
[0036] In this embodiment, the large model knowledge distillation successfully extracted 5 semantic levels and 18 prior causal edges.
[0037] This step uses the dataset from step S1. As input, causal relationships are identified from the observed data and confounding biases are removed through conditional residual causal inference, providing fully data-driven evidence of causal structure for step S4. Step S3 specifically includes the following sub-steps S31 to S37.
[0038] Step S31: Use the Pearson correlation coefficient test to initially screen the correlation between variables and construct an undirected graph of the correlation between variables. This reduces the computational cost of causal discovery. Initialize to A matrix consisting entirely of zeros. For any two variables... ,calculate , Pearson correlation coefficient between and their corresponding significance probability values .in The correlation coefficient is used to measure the statistical significance of a correlation; the smaller the value, the more reliable the correlation. Then the variable is considered to be , There is a statistically significant linear correlation between them. middle The edge represented is set to 1. For significance level, take .
[0039] Step S32: For any pair of variables (That is, step S31 determines the pairs of variables that are significantly correlated), and the K-nearest neighbor regression method is used to estimate the given... Down Conditional expectations The specific calculation formula is as follows: ; in, Indicates and The set of the K nearest samples, where K is the number of nearest neighbors. for The corresponding sample The value of .
[0040] Step S33: Calculation , Relative conditions between The specific formula is as follows: ; The residual measure is when... When used as a "cause", The actual value and based on The relative deviation between the conditional expectations of the K-nearest neighbor regression and the residuals. The smaller the residual, the better. In the given The smaller the variation, right The stronger the explanatory power, that is... The direction is more likely to be the true causal direction.
[0041] Step S34: Calculate the squared conditional residuals based on the full sample The specific formula is as follows: ; in, Indicates the first A sample about Because of, The RCRP is the RCR value of the result. The RCRP is an overall measure of the degree of RCR fluctuation across the entire sample. The smaller the RCRP value, the higher the stability of the causal explanation in that direction.
[0042] Step S35: Determine the causal direction. For variables... Calculate separately and Judge based on the size of the two. The direction of the causal edge. The rules for determining the causal direction are as follows: like Then the direction of causality is determined as follows: ; like Then the direction of causality is determined as follows: ; like Then it is marked as a bidirectional edge. .
[0043] This judgment rule is based on the principle of asymmetric causal discovery: when yes When the time comes, use predict The squared residuals are usually less than that obtained by using predict The square of the residual, i.e. Therefore, the direction corresponding to the smaller RCRP value is determined to be the causal direction.
[0044] Step S36: For the causal edges that were determined to exist in step S35 Calculate its confidence level The specific formula is as follows: ; The confidence formula ensures that when the RCRPs of the two directions are equal, the confidence level is 0.5, indicating the highest uncertainty in direction determination; as the difference increases, the confidence level increases, with an upper limit of 0.9, to avoid the absoluteness of the confidence level reaching 1.
[0045] Step S37: Data-driven residual causal graph construction. Using all learning behavior variables... Using the node set and the causal relationships determined in steps S31 to S36 as directed edges, construct a data-driven residual causal graph. .in, The data-driven causal strength matrix is defined as follows: .
[0046] In this embodiment, the conditional residual causal inference method discovered 45 data-driven causal edges from the data.
[0047] Step S4: The prior cause-effect graph With the residual cause-effect graph Perform fusion to obtain a fusion causal graph. and in the fusion causal graph The system performs a causal chain search from the root cause node to the early warning target node, and outputs online learning early warning attribution results containing root cause location, transmission path and intervention target.
[0048] This step integrates the knowledge causal graph from step S2 with the data causal graph from step S3 using a strength-optimized approach to achieve complementary fusion of knowledge and data, and provides an interpretable end-to-end risk attribution path. Step S4 specifically includes the following sub-steps S41 to S44.
[0049] Step S41: Cause-effect graph fusion. For each variable pair... Define the causal strength after fusion. for: ; The above fusion rules mean the following: when both the prior graph and the data graph have causal edges in the same direction (i.e., both have strengths greater than 0), the maximum value of the two strengths is taken to enhance the confidence of that causal edge; when only one graph has a causal edge, the strength of that graph is retained; when the two graphs conflict in direction, the graph with higher confidence (i.e., greater strength) is selected as the edge to be retained and the fusion strength; when neither graph has any edges, the strength is 0. This yields the fused directed weight matrix. With fusion causal graph .
[0050] In this embodiment, the graph fusion yields 52 fused causal edges.
[0051] Step S42: Find the root node. (In the fused causal graph) In this context, the root cause node is a node with no incoming edges, which satisfies the following condition for a node: There are no nodes. Make All nodes that satisfy this condition form the root cause node set. .Right now: ; The root cause node represents the variable at the very top of the causal transmission chain and is the root cause of learning risk.
[0052] Step S43: For each root cause node The depth-first search algorithm is executed, recursively traversing the successor nodes. The specific steps are as follows: a. From the current node Set off and find all that meet the requirements. successor node ; b. If the current node is the target node (i.e., summative evaluation / dropout warning), then the current path is recorded as a complete causal chain.
[0053] Step S44: Organize all paths from the root cause node to the target node, forming a set of causal chains. The warning attribution results contain the following three types of information: (1) Root cause localization: Extract the starting node of each causal chain as the root cause of learning risk; (2) Transmission path: Fully demonstrate the causal transmission chain from the root cause to the early warning, and identify key mediating variables; (3) Intervention targets: Based on the role of each node in the causal chain, the key links that can be intervened are identified. The final output is an online learning early warning attribution result based on a large language model and conditional residual causal inference, realizing end-to-end interpretable attribution from behavioral data to early warning risks. This can help teachers prioritize the most influential risk features, optimize the allocation efficiency of educational resources, and fundamentally reduce the dropout rate of online courses.
[0054] In this embodiment, a total of 5 multi-level causal chains were found, and the exemplary intervention target hierarchy is as follows; The first level (root cause) is to improve learning satisfaction and reduce procrastination. The second layer (intermediary) increases the number of assignments completed and the number of times chapters are studied; The third layer (outcomes) focuses on students whose test scores and final exam scores are below the threshold.
[0055] As can be seen from the above technical solution, the online learning early warning attribution method based on large language models and conditional residual causal inference provided by this invention has at least the following beneficial effects: First, it utilizes the pre-trained knowledge of large language models to achieve automatic segmentation of variable semantic levels and distillation of prior causal knowledge, introducing expert knowledge in the field of education at a lower cost; Second, it adaptively learns the causal direction from the observation data through conditional residual causal inference, effectively removing the interference of confounding factors and distinguishing between causal relationships and spurious correlations in correlations; Third, it integrates prior causal graphs and data-driven causal graphs with an optimal fusion strategy, achieving complementarity between knowledge-driven and data-driven approaches and enhancing the confidence of causal edges; Fourth, based on the depth-first search of the fused causal graph, it can discover the complete transmission path from root cause to early warning target, providing interpretable end-to-end causal evidence for educational intervention.
[0056] Example 2: like Figure 2 As shown, this invention provides an online learning early warning attribution system based on large language models and conditional residual causal inference. This system is used to implement the online learning early warning attribution method based on large language models and conditional residual causal inference described in Embodiment 1 above, and specifically includes the following four modules.
[0057] Data collection module 100 is used to collect learning behavior data generated by learners on the online course platform, and obtain data from... 1 variable and A dataset consisting of 10 samples The variables include behavioral indicators and early warning result indicators. This module corresponds to the implementation of step S1 in Embodiment 1.
[0058] The large-scale model knowledge distillation and prior causal graph construction module 200 is used to apply a preset large language model to the knowledge distillation and prior causal graph construction module 200. Causal prior knowledge is extracted from the semantic information of each variable. A knowledge-driven prior causal graph is constructed through variable semantic hierarchy division, inter-hierarchical causal inference, and prior strength assignment. This module corresponds to the implementation of step S2 and its sub-steps S21 to S25 in Embodiment 1.
[0059] Specifically: Construct suitable prompt words for the large language model (step S21); input the variable set and semantic description into the large language model to complete the automatic division of K semantic levels (step S22); for any two different semantic levels, infer the causal direction according to the direction indicator function. Establish directed causal edges between levels (step S23); assign prior causal strength values between and within levels (step S24); generate a knowledge-driven prior causal graph and a causal prior strength matrix with all variables as nodes and all causal relationships as directed edges (step S25).
[0060] The conditional residual causal inference and data causal graph construction module 300 is used to perform conditional residual causal inference from the dataset. In adaptive learning, the causal direction and strength between variables are determined, and a data-driven residual causal graph is constructed. This module corresponds to the implementation of step S3 and its sub-steps S31 to S37 in Embodiment 1.
[0061] Specifically: through Pearson correlation coefficient test and significance level Constructing an undirected graph of correlations for initial screening (step S31); estimating conditional expectation using K-nearest neighbor regression (step S32); calculating the conditional residuals RCR (step S33); calculating the squared conditional residuals RCRP based on the full sample (step S34); comparing the RCRP values of the two directions to determine the causal direction (step S35); calculating the confidence of the causal edge. (Step S36); Finally, construct the data-driven residual causal graph and the data-driven causal strength matrix (Step S37).
[0062] The causal graph fusion and early warning attribution output module 400 is used to process the prior causal graph. With the residual cause-effect graph Perform fusion to obtain a fusion causal graph. and in the fusion causal graph The module performs a causal chain search from the root cause node to the early warning target node, and outputs online learning early warning attribution results containing root cause location, transmission path, and intervention target. This module corresponds to the implementation of step S4 and its sub-steps S41 to S44 in Embodiment 1.
[0063] Specifically: calculate the fusion causal intensity matrix according to the optimal fusion rule and generate the fusion causal graph (step S41); identify all root cause nodes with zero incoming edges (step S42); recursively traverse successor nodes from each root cause node using a depth-first search algorithm and record the complete path to the warning target node (step S43); organize the causal chain set and output three types of attribution information (step S44).
[0064] This embodiment presents an online learning early warning attribution system based on a large language model and conditional residual causal inference. It is used to implement the aforementioned online learning early warning attribution method based on a large language model and conditional residual causal inference. Therefore, the specific implementation of the system can be found in the description of the method embodiment section above. The data collection module 100, the large model knowledge distillation and prior causal graph construction module 200, the conditional residual causal inference and data causal graph construction module 300, and the causal graph fusion and early warning attribution output module 400 are used to implement steps S1, S2, S3, and S4 in the above method, respectively, and will not be described again here.
[0065] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0066] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0067] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0068] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.
Claims
1. An online learning early warning attribution method based on large language models and conditional residual causal inference, characterized in that, Includes the following steps: Step S1: Collect learning behavior data generated by learners on the online course platform to obtain... 1 variable and A dataset consisting of 10 samples The variables include behavioral indicators and early warning result indicators; Step S2: Apply the preset large language model from the... Causal prior knowledge is extracted from the semantic information of each variable. A knowledge-driven prior causal graph is constructed through variable semantic hierarchy division, inter-hierarchical causal inference, and prior strength assignment. ; Step S3: Based on conditional residual causal inference, from the dataset In adaptive learning, the causal direction and strength between variables are determined, and a data-driven residual causal graph is constructed. ; Step S4: The prior cause-effect graph With the residual cause-effect graph Perform fusion to obtain a fusion causal graph. and in the fusion causal graph The system performs a causal chain search from the root cause node to the early warning target node, and outputs online learning early warning attribution results containing root cause location, transmission path and intervention target.
2. The online learning early warning attribution method based on large language models and conditional residual causal inference according to claim 1, characterized in that, In step S2, the application of a preset large language model from the... Causal prior knowledge is extracted from the semantic information of each variable. A knowledge-driven prior causal graph is constructed through variable semantic hierarchy division, inter-hierarchical causal inference, and prior strength assignment. Specifically, it includes: Step S21: Construct suitable large language model prompt words, designate the large language model to play the role of causal inference expert, and automatically complete the variable hierarchy division based on the semantic features of each variable, and simultaneously determine the causal direction and prior strength between different levels. Step S22: Input the variable set, the semantic description of each variable, and the prompt words into the large language model, and the large language model will then process... The variables are automatically divided into semantic levels There is no overlap between the levels and all variables are covered; Step S23: For any two different semantic levels , The large language model infers the causal direction between the two and outputs a direction indicator function. And based on this, directed causal edges are established for all variable pairs between levels; Step S24: Assign prior causal strength values. For causal edges between levels, the large language model generates the prior strength between each pair of variables. For variable pairs belonging to the same level, a pre-defined prior strength is uniformly assigned. ; Step S25: with the aforementioned Using each variable as a node and the causal relationships between and within all levels as directed edges, the knowledge-driven prior causal graph is constructed. ,in For a set of nodes, Let be the causal prior strength matrix.
3. The online learning early warning attribution method based on large language models and conditional residual causal inference according to claim 2, characterized in that, The direction indication function described in step S23 Defined as: like Then the hierarchy For the reason, For the result, for any and All established by point to Causal edge; like Then the hierarchy For the result, For the reason, establish by point to Causal edge; like Then no directed causal edge is established between the two layers.
4. The online learning early warning attribution method based on large language models and conditional residual causal inference according to claim 2, characterized in that, The preset prior strength mentioned in step S24 Causal strength values uniformly assigned among variables at the same level; prior strength of variable pairs between levels. Generated by a large language model that integrates variable semantics, hierarchical attribution, and domain knowledge, with a value range of [missing information]. .
5. The online learning early warning attribution method based on large language models and conditional residual causal inference according to claim 1, characterized in that, In step S3, the conditional residual causal inference is performed from the dataset. In adaptive learning, the causal direction and strength between variables are determined, and a data-driven residual causal graph is constructed. Specifically, it includes: Step S31: For any two variables , Calculate the Pearson correlation coefficient. and significance probability value ,like In an undirected graph with correlations... Set the corresponding edge to 1. Preset significance level; Step S32: For the undirected graph of correlations Each variable pair that exists in K-nearest neighbor regression is used to estimate the given... hour Conditional expectation ; Step S33: Calculation , Relative conditions between ; Step S34: Calculate the squared conditional residuals based on the full sample ,in Indicates the first A sample about Because of, It is the RCR value of the fruit; Step S35: For each pair of variables, compare and The magnitude of the two factors determines the direction of cause and effect; Step S36: For each causal edge that is determined to exist, calculate its confidence level. ; Step S37: Using all variables as nodes, and the causal direction determined in step S35 and the confidence level calculated in step S36 as edge weights, construct the data-driven residual causal graph. ,in For a set of nodes, For a data-driven causal strength matrix, if but Otherwise, it is 0.
6. The online learning early warning attribution method based on large language models and conditional residual causal inference according to claim 5, characterized in that, The causal direction determination rule mentioned in step S35 is as follows: like Then the direction of causality is determined as follows: ; like Then the direction of causality is determined as follows: ; like Then it is marked as a bidirectional edge. .
7. The online learning early warning attribution method based on large language models and conditional residual causal inference according to claim 5, characterized in that, The confidence level mentioned in step S36 The calculation formula is: ; This ensures that the confidence level is 0.5 when the squares of the residuals in both directions are equal, and the confidence level is higher when the difference is greater, with an upper limit of 0.
9.
8. The online learning early warning attribution method based on large language models and conditional residual causal inference according to claim 1, characterized in that, In step S4, the prior cause-effect graph is... With the residual cause-effect graph Perform fusion to obtain a fusion causal graph. and in the fusion causal graph The system performs a causal chain search from the root cause node to the early warning target node, and outputs online learning early warning attribution results containing root cause location, transmission path, and intervention target, specifically including: Step S41: For each pair of variables Define the causal strength of fusion If both the prior graph and the data graph have causal edges in the same direction and both have a strength greater than 0, then the maximum value is taken; if only one has a strength greater than 0, then the strength of that side is taken; if the two graphs conflict in direction, then the one with the higher confidence score is taken; otherwise, the value is 0. This yields the fused directed weight matrix. and fusion causal diagram ; Step S42: In the fusion causal graph In this context, the root cause node set is constructed using variables whose incoming edge weights are all zero. ; Step S43: For each root cause node The algorithm recursively traverses its successor nodes using a depth-first search. When the current node is a preset warning target node, it is recorded. The complete path to the target node; Step S44: Organize all paths from root causes to target nodes to form a set of causal chains.
9. The online learning early warning attribution method based on large language models and conditional residual causal inference according to claim 8, characterized in that, In step S4, the fusion causal strength The specific definition is: ,when ,in For the strength of prior causality, Data-driven causal strength; ,when and ; ,when and ; Other situations; When there is a directional conflict, the edge with the greater strength is selected as the edge to be retained and the fusion strength.
10. An online learning early warning attribution system based on large language models and conditional residual causal inference, characterized in that, The system is used to implement the online learning early warning attribution method based on large language models and conditional residual causal inference as described in any one of claims 1 to 9, specifically including: The data collection module is used to collect learning behavior data generated by learners on the online course platform, and obtain data from... 1 variable and A dataset consisting of 10 samples The variables include behavioral indicators and early warning result indicators; The large-scale model knowledge distillation and prior causal graph construction module is used to apply a pre-defined large language model to the knowledge distillation and prior causal graph construction module. Causal prior knowledge is extracted from the semantic information of each variable. A knowledge-driven prior causal graph is constructed through variable semantic hierarchy division, inter-hierarchical causal inference, and prior strength assignment. ; The conditional residual causal inference and data causal graph construction module is used to perform conditional residual causal inference from the dataset. In adaptive learning, the causal direction and strength between variables are determined, and a data-driven residual causal graph is constructed. ; The causal graph fusion and early warning attribution output module is used to process the prior causal graph. With the residual cause-effect graph Perform fusion to obtain a fusion causal graph. and in the fusion causal graph The system performs a causal chain search from the root cause node to the early warning target node, and outputs online learning early warning attribution results containing root cause location, transmission path and intervention target.