Method and system for predicting graph attention knowledge tracking model based on causal inference

By introducing causal inference and graph attention mechanisms into the knowledge tracking model, building a causal structure chart and eliminating short-term factors, the shortcomings of the existing models in distinguishing causal relationships and capturing learning changes are solved, and higher prediction accuracy and robustness are achieved.

CN120012888APending Publication Date: 2025-05-16SHANDONG NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510093523.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing knowledge tracking model based on graph neural networks has difficulties in distinguishing causality and non-causality, and fails to effectively capture important changes in the learning process, resulting in a decrease in prediction accuracy.

Method used

A graph attention knowledge tracking model based on causal inference is proposed. By constructing a causal structure chart, calculating causal weights, eliminating short-term factors, combining time convolution networks and multi-head attention mechanisms, a more comprehensive knowledge state feature is extracted, and the graph features are decoupled into causal features and shortcut features for prediction.

Benefits of technology

Effectively distinguish causal characteristics, improve classification accuracy, improve model robustness and generalization ability, and significantly improve prediction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012888A_ABST
    Figure CN120012888A_ABST
Patent Text Reader

Abstract

The invention discloses a causal inference-based graph attention knowledge tracking model prediction method and system, and the method comprises the steps: inputting data containing students, concepts and answer sequences into a pre-trained improved time convolution network for feature extraction, and obtaining comprehensive knowledge state features; constructing a causal structure diagram based on the data, and eliminating short-term factors by calculating causal weights of node neighborhoods in the causal structure diagram to obtain data after the short-term factors are eliminated; inputting the data after the short-term factors are eliminated and the knowledge state features into a pre-trained knowledge tracking model based on a graph neural network for feature extraction to obtain graph features; and decoupling the graph features into causal features and shortcut features, and performing prediction based on the causal features and the shortcut features to obtain a prediction result. Causal features can be effectively distinguished, and the classification accuracy is improved. And a back door path is cut off by utilizing a do operator, so that the influence of non-causal factors can be reduced, the robustness of the model is improved, and the performance of the model is not influenced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of knowledge tracking technology, and in particular to a prediction method and system for a graph attention knowledge tracking model based on causal inference. Background Art

[0002] The statements in this section merely provide background information related to the present disclosure and do not necessarily constitute prior art.

[0003] In the field of knowledge tracking, most graph neural networks (GNNs) follow the paradigm of "learning participation", which makes GNNs extensive and non-selective in dealing with the relationship between features and targets, and thus may not be able to clearly distinguish between causal and non-causal relationships. They often mix all features together and are easily affected by short-term factors during classification, which makes the GNN classifier's classification ability poor. In addition, existing graph neural network-based knowledge tracking models (GKT) usually rely on static feature aggregation and fail to consider the changes in neighboring nodes when a node is disturbed. This makes it impossible for the model to effectively capture important changes in the learning process when dealing with complex learning environments, resulting in reduced prediction accuracy. The limitations of this feature aggregation cannot meet the needs of personalized learning and limit its effectiveness in practical applications. Summary of the invention

[0004] In order to overcome the shortcomings of the above-mentioned prior art, the present invention provides a prediction method and system for a graph attention knowledge tracking model based on causal inference, which can effectively distinguish causal features and improve classification accuracy. At the same time, when a node in the graph is disturbed, the backdoor path is cut off by using the do operator, which can reduce the influence of non-causal factors, improve the robustness of the model, and keep its performance unaffected.

[0005] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:

[0006] In a first aspect, the present invention provides a prediction method for a graph attention knowledge tracking model based on causal inference, comprising:

[0007] Acquire data including student, concept and answer sequences, input the data into a pre-trained improved temporal convolutional network for feature extraction, and obtain comprehensive knowledge state features;

[0008] Building a causal structure graph based on data including students, concepts, and answer sequences, eliminating short-term factors by calculating causal weights of node neighborhoods in the causal structure graph, and obtaining data after eliminating short-term factors;

[0009] Inputting the data and knowledge state features after eliminating short-term factors into a pre-trained knowledge tracking model based on a graph neural network for feature extraction to obtain graph features;

[0010] The graph features are decoupled into causal features and shortcut features, predictions are performed based on the causal features and shortcut features, and finally a prediction result is obtained.

[0011] A further technical solution is to input the students’ historical answers into the improved temporal convolutional network to obtain a more comprehensive knowledge state feature, which can be expressed as:

[0012] s t =TCN(x 0 ,x 1 ,…x t )

[0013] Among them, s t Represented as knowledge state feature, x 0 ,x 1 ,…x t Indicates the results of students' history test.

[0014] A further technical solution is to improve the temporal convolutional network by adding residual blocks between convolutional layers.

[0015] According to a further technical solution, the residual block is composed of an expanded causal convolution, a weight normalization, a ReLU activation function and a Dropout layer in sequence.

[0016] A further technical solution is to eliminate short-term factors through the do operator, which can be expressed as:

[0017]

[0018] Among them, C represents the causal feature, Y represents the prediction, do represents the do operator, and P m (YC,s) represents the probability of given causal feature C and confounding factor s, P m (s) represents the prior probability of the confounding factor.

[0019] A further technical solution is that the knowledge tracking model is a knowledge tracking model based on graph neural network that introduces a multi-head attention mechanism.

[0020] A further technical solution uses a multi-head attention mechanism to extract graph features, which can be expressed as:

[0021]

[0022] f self (h k ' t )=MLP(h k ' t )

[0023]

[0024] in, represents the graph features after the multi-head attention mechanism, f self represents a multilayer perceptron, h k ' t represents the updated knowledge state, f attention represents the attention module, K represents the number of heads, represents the weight of the kth attention head between node i and node j.

[0025] In a second aspect, the present invention provides a prediction system for a graph attention knowledge tracking model based on causal inference, comprising:

[0026] A first feature extraction module is configured to: obtain data including student, concept and answer sequence, input the data into a pre-trained improved temporal convolutional network for feature extraction, and obtain a comprehensive knowledge state feature;

[0027] A short-term factor elimination module is configured to: construct a causal structure graph based on data including students, concepts and answer sequences, eliminate short-term factors by calculating causal weights of node neighborhoods in the causal structure graph, and obtain data after eliminating short-term factors;

[0028] The second feature extraction module is configured to: input the data and knowledge state features after eliminating short-term factors into a pre-trained knowledge tracking model based on a graph neural network to extract features and obtain graph features;

[0029] The prediction module is configured to: decouple the graph features into causal features and shortcut features, perform prediction based on the causal features and shortcut features, and finally obtain a prediction result.

[0030] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the prediction method of the graph attention knowledge tracking model based on causal inference as described in the first aspect.

[0031] In a fourth aspect, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in the prediction method of the graph attention knowledge tracking model based on causal inference as described in the first aspect are implemented.

[0032] One or more of the above technical solutions have the following beneficial effects:

[0033] The present invention proposes a graph attention knowledge tracking model (CAGKT) based on causal inference. This model effectively identifies causal features such as the order and correlation of students' answers by constructing a causal structure, and processes the causal feature C through the do operator to eliminate the influence of short-term factors such as students' learning behavior at a certain point in time, the performance of the most recent answer, and the most recent learning activities on the students' knowledge status. At the same time, the convolution operation and dilated convolution of the temporal convolutional network (TCN) effectively extract the long-term dependency information in the sequence, which can cover more contextual information, obtain more comprehensive knowledge state features, and combine the attention mechanism to aggregate the features of adjacent nodes, thereby improving the classification ability and robustness of the model. Finally, by decomposing the graph features into two independent parts, it is actually to make the model more flexible and accurate in processing different types of information. Traditional knowledge tracking models often mix all features together, which makes it difficult for the model to distinguish the role of different features when facing students' complex learning behaviors. Through decoupling, CAGKT can process and utilize causal relationships and shortcut information separately, so that each part can be processed and learned independently, thereby improving the effect and interpretability of the model.

[0034] The graph attention knowledge tracking model (CAGKT) of the present invention can extract students' answer information more comprehensively through two methods: causal features and multi-head attention mechanism. This can better predict whether students will answer the next question correctly, and at the same time more accurately understand students' learning status, and provide more suitable practice and learning plans.

[0035] The present invention can effectively distinguish causal features and improve classification accuracy, which is 16.55% higher than the AUC of the GKT model. At the same time, when a node in the graph is disturbed, the backdoor path is cut off by the do operator, which can reduce the influence of non-causal factors, improve the robustness of the model, and keep its performance unaffected. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0037] Figure 1 is a residual block structure diagram of a temporal convolutional network according to an embodiment of the present invention;

[0038] Figure 2 It is a structural diagram of a graph attention knowledge tracking model based on causal inference in an embodiment of the present invention;

[0039] Figure 3 It is a cause-effect structure diagram of an embodiment of the present invention. DETAILED DESCRIPTION

[0040] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.

[0041] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.

[0042] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.

[0043] Existing graph neural network models (GNN) have difficulty distinguishing between causal and non-causal relationships, and often mix all features together, which is easily affected by short-term factors during classification, resulting in poor classification performance. The traditional graph neural network-based knowledge tracking model (GKT) fails to consider that when one node is disturbed, it will affect the features of other nodes and reduce prediction accuracy during feature aggregation, so it has limitations in complex learning environments.

[0044] The present invention proposes a graph attention knowledge tracking model (CAGKT) based on causal inference. The model constructs a causal structure to effectively identify causal features, and processes the causal feature C through the do operator to eliminate the influence of short-term factors such as students' learning behavior at a certain point in time, the performance of the most recent answering question, and the most recent learning activities on the students' knowledge state; a temporal convolutional network (TCN) is used to extract more comprehensive knowledge state features, and the attention mechanism is combined to aggregate the features of adjacent nodes; by decoupling the graph features into two parts, each part can independently process and learn causal information and shortcut information, and make predictions based on these two features, thereby improving the robustness and generalization ability of the model.

[0045] Embodiment 1

[0046] like Figure 1 As shown, this embodiment discloses a prediction method of a graph attention knowledge tracking model based on causal inference, and the method includes the following steps:

[0047] S1: Obtain data including student, concept and answer sequences, input the data into a pre-trained improved temporal convolutional network for feature extraction, and obtain comprehensive knowledge state features;

[0048] In this embodiment, data including students, concepts and answer sequences, i.e., historical answer data, is collected and converted into a graph structure. Specifically, a graph structure G = (V, E) including students, concepts and answer sequences is constructed, where the node V represents the concept and the edge E represents the relationship between the concepts.

[0049] The answer sequence is 1 when the student answers correctly and 0 when the student answers incorrectly, denoted by x t .

[0050] The collected student ID, answer concept and answer sequence are all one-to-one corresponding. One corresponding data represents one answer data. Collect multiple historical answer data to train the improved temporal convolutional network to obtain the trained improved temporal convolutional network. At each time step t, it is assumed that a student has a temporal knowledge state for each concept independently (knowledge state refers to the degree of mastery of a series of concepts by the student at the current time point). The student's historical answer data is input into the improved temporal convolutional network (TCN). TCN effectively extracts long-term dependency information in the sequence through convolution operations and dilated convolutions, which can cover more context information and extract more comprehensive knowledge state features.

[0051] Inputting students’ historical answer data into the improved temporal convolutional network, we can obtain more comprehensive knowledge state features, which can be expressed as:

[0052] s t =TCN(x 0 ,x 1 ,…x t )

[0053] Among them, s t represents the knowledge state characteristics, x 0 ,x 1 ,…x t Represents student historical answer data.

[0054] When students solve exercises related to a certain concept, aggregate the students' knowledge status of the concept itself and its related concepts, and then update the knowledge status.

[0055] To avoid network degradation in deep models, the improved temporal convolutional network (TCN) adds residual connections between convolutional layers. Figure 2 As shown in the figure, the residual block is composed of dilated causal convolution Dilated Causal Conv, weight normalization WeightNorm, ReLU activation function and Dropout layer in sequence.

[0056] S2: constructing a causal structure graph based on the data including students, concepts and answer sequences, eliminating short-term factors by calculating the causal weights of node neighborhoods in the causal structure graph, and obtaining data after eliminating short-term factors;

[0057] In this embodiment, the causal structure diagram is constructed based on original data, such as the assistance2015 data set, which includes student IDs, answer sequences, concept sequences, correct or incorrect answer sequences, etc.

[0058] like Figure 3 As shown in the figure, the causal structure graph (SCM) describes the causal relationship between variables, which is explained as follows: D (data): graph data, usually used as input information of the model; C (causal feature): by learning the intrinsic pattern D of the data, the causal features reflecting the intrinsic connection between the data are obtained; S (shortcut feature): shortcut features obtained by learning the data D. These features are usually caused by factors such as data deviation, noise interference and disturbance of adjacent nodes in the graph data; R (representation): the model aggregates and extracts information from the data D, and learns all extractable features (including causal features C and shortcut features S). Then, the graph representation information R is obtained through the causal features C and shortcut features S; Y (prediction): the model uses the graph representation information R to predict the task, which is the ultimate goal of the task.

[0059] The causal feature C is processed by the do operator to eliminate the influence of short-term factors on the student's knowledge status, such as the student's learning behavior at a certain point in time, the performance of the most recent answer, the most recent learning activities, etc. These features are usually intuitive and directly reflect the student's current learning status. They will be helpful in extracting the student's knowledge status, but if all rely on these features, the accuracy will be reduced. At the same time, the causal features are effectively extracted. The specific steps are: first construct a causal structure diagram and learn the causal weights, and then eliminate the influence of short-term factors through the do operator to obtain an accurate causal estimation effect, which can be expressed as:

[0060]

[0061] Among them, C refers to the causal characteristics that reflect the internal connection between data by learning the internal pattern of data, such as the order and correlation of students' answers; do represents the do operator, P m (YC,s) represents the probability of given causal feature C and confounding factor s, P m (s) represents the prior probability of the confounding factor.

[0062] S3: Inputting the data and knowledge state features after eliminating short-term factors into a pre-trained knowledge tracking model based on a graph neural network for feature extraction to obtain graph features;

[0063] In this embodiment, the knowledge tracking model is a graph neural network-based knowledge tracking model that introduces a multi-head attention mechanism. The multi-head attention mechanism is used to improve GKT to aggregate adjacent node features, thereby improving the classification ability and robustness of the model.

[0064] First, the knowledge tracking model (GKT) based on graph neural network is used to obtain the current knowledge state of students, which aggregates the historical answer sequence and the current answer sequence information, and obtains the characteristics of neighbor nodes to obtain a global feature that includes the characteristics of all nodes, which can be expressed as:

[0065]

[0066] Among them, h k ' t represents the updated knowledge state, t represents the current time point, k represents the neighbor node, represents the current knowledge state, i represents the current node, p t The input vector {0,1} represents the correctness of the concept answer at time step t (obtained based on the student's historical answer data), s t Represents the knowledge state characteristics, E x Matrix embeddings representing concept indices and corresponding answers.

[0067] Then, a multi-head attention mechanism is used to learn different functions and different features. The attention mechanism enables the model to focus on important features and relationships in the knowledge graph, thereby effectively integrating relevant information into the knowledge state. The multi-head attention mechanism is used to obtain aggregate features, also known as graph features. The final knowledge state is obtained by updating the graph features and the knowledge graph structure, which can be expressed as:

[0068]

[0069]

[0070]

[0071] in, represents the graph features after the multi-head attention mechanism, f self represents a multi-layer perceptron (MLP), f attention represents the attention module, which uses a multi-head attention mechanism to obtain the weight of the edge; K represents the number of multi-heads, represents the weight of the kth attention head between node i and node j.

[0072] Finally, the knowledge hidden state is updated by using erase-add gates and gated recurrent units. The gating mechanism is mainly used to solve the gradient vanishing problem in recurrent neural networks. It can be expressed as:

[0073]

[0074] in, Indicates the intermediate process, G ea Represented as an erase-add gate, Represents the final knowledge state, which is updated by graph features and knowledge graph structure G = (V, E); G gru Represents a gated recurrent unit.

[0075] S4: Decoupling the graph features into causal features and shortcut features, performing predictions based on the causal features and shortcut features, and finally obtaining prediction results.

[0076] In this embodiment, the graph features (expressed as graph structures, i.e., the final graph structures after aggregation and feature update) are separated into causal features and shortcut features. The Attention module is used to decouple the graph features. Specifically, the node representation can be obtained based on the GNN encoder and the graph G, and then two MLPs are used. node () and MLP edge () Estimate the attention scores from the perspective of nodes and edges, and build soft masks based on the attention scores. Finally, decouple the graph into two independent parts (causal features and shortcut features, also known as causal graphs and shortcut graphs, and make predictions based on these two graphs), so that the model can process different types of information more flexibly and accurately, so that each part can independently process and learn causal information and shortcut information, and make predictions based on these two features, thereby improving the robustness and generalization ability of the model. It can be expressed as:

[0077] α c ,α s =f softmax (MLP(h i ))

[0078] Among them, α c represents the attention score of the node, α s represents the attention score of the edge, f softmax represents the activation function and MLP() represents the multi-layer perceptron, which is used to derive the attention score.

[0079] The GNN encoder is used to obtain the representation of the graph and f readout The function and classifier combine causal features and shortcut features for prediction. The readout function obtains the feature representation of the entire image by aggregating features, which can be expressed as:

[0080]

[0081] Among them, σ represents the activation function, W represents the weight, represents the graph causal knowledge state, b represents the bias term, y G′ represents the predicted probability, φ represents the classifier, Represents the graph shortcut knowledge state.

[0082] The invention makes predictions based on students' historical answer data and outputs the probability of answering the next question correctly or incorrectly. Finally, it conducts empirical verification on multiple open data sets to evaluate the effect of the model on student knowledge prediction and its interpretability, significantly improving the accuracy and having better performance than traditional methods.

[0083] Embodiment 2

[0084] This embodiment discloses a prediction system of a graph attention knowledge tracking model based on causal inference, including:

[0085] A first feature extraction module is configured to: obtain data including student, concept and answer sequence, input the data into a pre-trained improved temporal convolutional network for feature extraction, and obtain a comprehensive knowledge state feature;

[0086] A short-term factor elimination module is configured to: construct a causal structure graph based on data including students, concepts and answer sequences, eliminate short-term factors by calculating causal weights of node neighborhoods in the causal structure graph, and obtain data after eliminating short-term factors;

[0087] The second feature extraction module is configured to: input the data and knowledge state features after eliminating short-term factors into a pre-trained knowledge tracking model based on a graph neural network to extract features and obtain graph features;

[0088] The prediction module is configured to: decouple the graph features into causal features and shortcut features, perform prediction based on the causal features and shortcut features, and finally obtain a prediction result.

[0089] Embodiment 3

[0090] The purpose of this embodiment is to provide a computing device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method of embodiment 1 when executing the program.

[0091] Embodiment 4

[0092] The purpose of this embodiment is to provide a computer-readable storage medium, a computer-readable storage medium having a computer program stored thereon, and when the program is executed by a processor, the steps of the method of embodiment 1 are performed.

[0093] The steps involved in the apparatus of the above embodiments 3 and 4 correspond to the method embodiment 1, and the specific implementation method can refer to the relevant description part of embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.

[0094] Those skilled in the art should understand that the modules or steps of the present invention described above can be implemented by a general-purpose computer device, or alternatively, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0095] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

[0096] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.

Claims

1. A prediction method based on a graph attention knowledge tracking model based on causal inference, characterized in that: include: Acquire data including student, concept and answer sequences, input the data into a pre-trained improved temporal convolutional network for feature extraction, and obtain comprehensive knowledge state features; Building a causal structure graph based on data including students, concepts, and answer sequences, eliminating short-term factors by calculating causal weights of node neighborhoods in the causal structure graph, and obtaining data after eliminating short-term factors; Inputting the data and knowledge state features after eliminating short-term factors into a pre-trained knowledge tracking model based on a graph neural network for feature extraction to obtain graph features; The graph features are decoupled into causal features and shortcut features, predictions are performed based on the causal features and shortcut features, and finally a prediction result is obtained.

2. The prediction method of the graph attention knowledge tracking model based on causal inference as claimed in claim 1 is characterized in that: Inputting the students’ historical answers into the improved temporal convolutional network, we can obtain a more comprehensive knowledge state feature, which can be expressed as: s t =TCN(x 0 ,x 1 ,…x t ) Among them, s t Represented as knowledge state feature, x 0 ,x 1 ,…x t Indicates the results of students' history test.

3. The prediction method of the graph attention knowledge tracking model based on causal inference as claimed in claim 1 is characterized in that: The improved temporal convolutional network adds residual blocks between convolutional layers.

4. The prediction method of the graph attention knowledge tracking model based on causal inference as claimed in claim 3 is characterized in that: The residual block is composed of dilated causal convolution, weight normalization, ReLU activation function and Dropout layer in sequence.

5. The prediction method of the graph attention knowledge tracking model based on causal inference as claimed in claim 1, characterized in that: By eliminating short-term factors through the do operator, it can be expressed as: Among them, C represents the causal feature, Y represents the prediction, do represents the do operator, and P m (YC,s) represents the probability of given causal feature C and confounding factor s, P m (s) represents the prior probability of the confounding factor.

6. The prediction method of the graph attention knowledge tracking model based on causal inference as claimed in claim 1, characterized in that: The knowledge tracking model is a knowledge tracking model based on graph neural network that introduces a multi-head attention mechanism.

7. The prediction method of the graph attention knowledge tracking model based on causal inference as claimed in claim 1, characterized in that: The multi-head attention mechanism is used to extract graph features, which can be expressed as: in, represents the graph features after the multi-head attention mechanism, f self represents a multilayer perceptron, represents the updated knowledge state, f attention represents the attention module, K represents the number of heads, represents the weight of the kth attention head between node i and node j.

8. A prediction system based on a graph attention knowledge tracking model with causal inference, characterized in that: include: A first feature extraction module is configured to: obtain data including student, concept and answer sequence, input the data into a pre-trained improved temporal convolutional network for feature extraction, and obtain a comprehensive knowledge state feature; A short-term factor elimination module is configured to: construct a causal structure graph based on data including students, concepts and answer sequences, eliminate short-term factors by calculating causal weights of node neighborhoods in the causal structure graph, and obtain data after eliminating short-term factors; The second feature extraction module is configured to: input the data and knowledge state features after eliminating short-term factors into a pre-trained knowledge tracking model based on a graph neural network to extract features and obtain graph features; The prediction module is configured to: decouple the graph features into causal features and shortcut features, perform prediction based on the causal features and shortcut features, and finally obtain a prediction result.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, it implements the steps in the prediction method of the graph attention knowledge tracking model based on causal inference as described in any one of claims 1-7.

10. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, it implements the steps in the prediction method of the graph attention knowledge tracking model based on causal inference as described in any one of claims 1-7.