A Deep Learning-Based Method and System for Constructing a Behavioral Model of Exam Question Setters
By constructing a deep learning-based model of exam question setter behavior, the problems of question setters' overall grasp of the knowledge system and dynamic changes in preferences were solved, enabling accurate characterization and anomaly detection of question setting behavior, thereby improving the quality and fairness of question setting.
Patent Information
- Application Number
- CN202511324315.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-17
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-09-17
AI Technical Summary
Existing technologies in intelligent question-setting systems fail to effectively characterize the question setter's overall grasp of the knowledge system, ignore the transfer and correlation between knowledge points, fail to consider the dynamic changes in the question setter's preferences over time, and lack anomaly monitoring mechanisms, resulting in insufficient adaptability and low modeling accuracy of the question-setting behavior model.
We construct a deep learning-based exam question setter behavior model. Through knowledge association analysis, dynamic preference analysis, and anomaly detection modules, we comprehensively consider knowledge point transfer association, dynamic preference changes, and abnormal behavior. We also optimize the extraction and identification of question-setting behavior features by utilizing knowledge application graphs and transfer paths.
It enables precise characterization and anomaly monitoring of question-setting behavior, improves the quality and fairness of exam question setting, can promptly detect deviations and irregularities, and enhances the scientific rigor and standardization of the question-setting process.
Smart Images

Figure CN120822160B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to artificial intelligence and educational assessment technology, and in particular to a method and system for constructing a behavior model of exam question setters based on deep learning. Background Technology
[0002] With the development of educational informatization, intelligent question-setting systems are widely used in the generation of exam questions. In the intelligent question-setting process, the question-setting behavior characteristics of the question setter have a significant impact on the quality of the questions. Existing technologies mainly analyze question-setting behavior from two dimensions: question difficulty distribution and knowledge point coverage. However, these methods have the following shortcomings: existing methods only focus on the frequency of use of individual knowledge points, ignoring the transfer relationships between knowledge points, and cannot effectively characterize the question setter's overall grasp of the knowledge system; existing technologies treat question-setting behavior as a static feature, failing to consider the dynamic changes in question setter preferences over time, resulting in insufficient adaptability of the question-setting behavior model; and there is a lack of anomaly monitoring mechanisms for question-setting behavior, making it impossible to promptly detect and intervene in inappropriate question-setting behaviors. Furthermore, existing technologies often analyze knowledge point selection and difficulty setting separately, failing to reveal the inherent relationship between the two, reducing the accuracy of question-setting behavior modeling. Therefore, a question-setting behavior modeling method that can comprehensively consider multiple dimensions such as knowledge point transfer relationships, dynamic preference changes, and anomaly behavior monitoring is needed to improve the quality control level of intelligent question-setting systems. Summary of the Invention
[0003] This invention provides a method and system for constructing a behavior model of exam question setters based on deep learning, which can solve the problems in the prior art.
[0004] A first aspect of this invention provides a method for constructing a deep learning-based model of exam question setter behavior, comprising:
[0005] The test question setter's question setting records are obtained, and a knowledge association analysis module is constructed. The knowledge association analysis module calculates the frequency of use based on the knowledge point tags in the question setting records, constructs a knowledge point mapping network based on a preset knowledge system, obtains the transfer weight by calculating the conditional probability between knowledge points, combines the knowledge point mapping network and the transfer weight to generate a knowledge application graph, and outputs a knowledge application vector based on the knowledge application graph.
[0006] A dynamic preference analysis module is constructed based on knowledge application vectors. The dynamic preference analysis module groups the question-setting behavior into time windows according to the time information in the question-setting record, extracts features from the knowledge point selection sequence in each time window, and generates a question preference vector by combining the difficulty information.
[0007] A proposition behavior recognition network is trained based on proposition preference vectors. The proposition behavior recognition network models the knowledge point selection sequence and difficulty information respectively. By optimizing the migration path and difficulty distribution in the knowledge application graph, proposition behavior features are obtained.
[0008] The question-setting behavior characteristics are input into the anomaly detection module. The anomaly detection module calculates the warning threshold based on the migration path and difficulty distribution. When the question-setting behavior characteristics deviate from the warning threshold, the module outputs warning information and the cause of the anomaly.
[0009] A knowledge association analysis module is constructed, which calculates the frequency of use based on the knowledge point tags in the proposition records and constructs a knowledge point mapping network based on a preset knowledge system, including:
[0010] The proposition text of the proposition record is converted into a word vector sequence. Bidirectional attention is calculated by query vector and word vector sequence, and candidate fragments of knowledge points are identified based on attention distribution.
[0011] The candidate knowledge point fragments are semantically vectorized, and the semantic similarity between the semantically vectorized representation and the standard knowledge points in the preset knowledge system is calculated. The corresponding knowledge point tags are extracted based on the similarity matching results.
[0012] The frequency of direct use of knowledge point tags in the question record is calculated. At the same time, based on the hierarchical structure of the preset knowledge system, the frequency of use of the parent node and child node corresponding to each knowledge point tag in the question record is counted and weighted to obtain the hierarchical usage frequency of knowledge points containing hierarchical relationships.
[0013] Using knowledge point tags as nodes and hierarchical usage frequency as the initial weight of nodes, node connections are established based on the dependencies between knowledge points in the preset knowledge system to construct an initial knowledge point mapping network; the order of knowledge point selection in the question-setting process is extracted from the question-setting records, and the transfer rules between knowledge points are analyzed based on the order of knowledge point selection to calculate the degree of correlation between nodes based on the transfer rules.
[0014] The connection strength between nodes in the initial knowledge point mapping network is updated based on the degree of relevance. Features are extracted from the updated node weights, and finally, a knowledge point mapping network representing the relationship between knowledge points is constructed.
[0015] Transfer weights are obtained by calculating the conditional probabilities between knowledge points. The knowledge point mapping network and transfer weights are combined to generate a knowledge application graph. Based on the knowledge application graph, the knowledge application vector is output, including:
[0016] Obtain the knowledge point selection sequence from the proposition record, set multiple sliding time windows of different scales for the knowledge point selection sequence, count the number of times the knowledge point appears and the number of consecutive occurrences of knowledge point pairs in each sliding time window, calculate the conditional probability of the knowledge point in the current window, and weight the conditional probabilities of the knowledge points in different sliding time windows to obtain the transfer weight between knowledge points.
[0017] The knowledge application graph is generated by combining the transfer weights between knowledge points with the knowledge point mapping network. The knowledge application graph includes nodes, the connection relationships between nodes, and the transfer weights corresponding to the connection relationships.
[0018] The node feature vectors in the knowledge application graph are transformed by a transformation matrix to obtain the node's transformed feature vector. The transformed feature vectors of adjacent nodes are concatenated and processed by an activation function. The attention coefficients between nodes are obtained through normalization. Based on the attention coefficients, the transformed feature vectors of adjacent nodes are weighted and aggregated to obtain an updated feature vector that integrates the feature information of adjacent nodes.
[0019] Average pooling is performed on the updated feature vectors of all nodes in the knowledge application graph to generate a knowledge application vector, which is used to represent the knowledge point selection pattern and transfer rule.
[0020] Based on the time information in the question-setting records, question-setting behaviors are grouped by time windows. Feature extraction is performed on the knowledge point selection sequence within each time window, and combined with difficulty information to generate a question-setting preference vector, including:
[0021] A time series is constructed from the time information in the proposition records. Kernel density estimation is performed on the time series to obtain the proposition activity density curve. The time window size is dynamically calculated based on the proposition activity density curve to obtain the time window division of proposition behavior.
[0022] Based on the time window division, the knowledge point selection sequence within each time window is extracted from the question record, a knowledge point word vector mapping table is constructed, the knowledge point selection sequence within each time window is mapped to a word vector sequence, and relative position encoding information is incorporated to obtain a position-aware vector sequence.
[0023] The location-aware vector sequence is input into a bidirectional recurrent neural network, and the context-dependent features of knowledge point selection are extracted to obtain the hidden state sequence.
[0024] The pre-acquired knowledge application vector is transformed into a query vector through nonlinear transformation. The similarity between the query vector and the hidden state sequence is calculated to obtain the attention score. The attention score is then normalized to obtain the attention weight.
[0025] The context vector is obtained by weighted summation of the hidden state sequence based on the attention weights. The context vector is then fused with the query vector and subjected to nonlinear transformation to generate a question preference vector that represents the question setter's preference for selecting knowledge points within the time window.
[0026] A proposition behavior recognition network is trained based on proposition preference vectors. This network models the knowledge point selection sequence and difficulty information separately. By optimizing the transfer path and difficulty distribution in the knowledge application graph, proposition behavior features are obtained, including:
[0027] A migration path identification module is constructed based on the proposition preference vector. The proposition preference vector is input into the feature extraction network to obtain initial features. A feature propagation matrix is generated based on the migration weights between nodes in the knowledge application graph. The initial features are aggregated using the feature propagation matrix to obtain the migration path features of the knowledge point selection sequence.
[0028] The difficulty information is encoded into a difficulty vector, and the correlation score between the difficulty vector and the migration path features is calculated. Based on the correlation score, the migration pattern of the knowledge point selection sequence under each difficulty level is identified.
[0029] Based on the migration pattern of the knowledge point selection sequence, a difficulty distribution weight is constructed. The difficulty distribution weight and the migration path feature are weighted and calculated to obtain the difficulty adaptive migration path feature. At the same time, the difficulty vector is corrected based on the difficulty adaptive migration path feature to obtain the corrected difficulty vector.
[0030] The difficulty-adaptive migration path features and the modified difficulty vector are fused to generate propositional behavior features that reflect the synergistic relationship between the knowledge point selection sequence and difficulty information.
[0031] The question-setting behavior characteristics are input into the anomaly detection module. The anomaly detection module calculates a warning threshold based on the migration path and difficulty distribution. When the question-setting behavior characteristics deviate from the warning threshold, it outputs warning information and the cause of the anomaly, including:
[0032] The proposition behavior features are input into the anomaly detection module. The anomaly detection module extracts migration path information and difficulty distribution information. Based on the migration path information, it calculates the migration probability between knowledge points to obtain a positive coupling matrix. Based on the difficulty distribution information, it calculates the difficulty constraint strength of knowledge point pairs to obtain a negative coupling matrix. The positive coupling matrix and the negative coupling matrix are combined to construct a bidirectional coupling representation.
[0033] The bidirectional coupling representation is used to forward propagate migration path information to obtain path features, and backward propagate difficulty distribution information to obtain difficulty features. The path features and difficulty features are then interactively calculated to obtain the coupling anomaly score.
[0034] The stability index of knowledge point transfer is calculated based on the bidirectional coupling representation. The distribution of coupling anomaly scores is dynamically corrected using the stability index to obtain an early warning threshold. When the coupling anomaly score exceeds the early warning threshold, anomaly detection is triggered.
[0035] An anomaly propagation network is constructed based on the bidirectional coupling representation. In the anomaly propagation network, the knowledge point migration sequence that causes the anomaly is located along the forward path, and the knowledge point combination that violates the difficulty constraint is identified along the reverse path. The located anomaly knowledge point migration sequence and the knowledge point combination that violates the difficulty constraint are used as the cause of the anomaly to output early warning information.
[0036] Based on the migration path information, the migration probability between knowledge points is calculated to obtain a forward coupling matrix. Based on the difficulty distribution information, the difficulty constraint strength of knowledge point pairs is calculated to obtain a reverse coupling matrix. Combining the forward coupling matrix and the reverse coupling matrix to construct a bidirectional coupling representation includes:
[0037] Based on the migration sequence and temporal features in the migration path information, the number of migrations between knowledge points is counted according to the time window. The migration statistics results under different time windows are weighted and combined to calculate the migration probability between knowledge points and obtain the positive coupling matrix. The positive coupling matrix represents the migration probability from the source knowledge point to the target knowledge point.
[0038] A hierarchical difficulty constraint graph is constructed based on the difficulty distribution information. In the hierarchical difficulty constraint graph, knowledge points are organized into layers according to their difficulty values. The difficulty gradient of knowledge point pairs within a layer and the cross-layer constraints of knowledge point pairs between layers are calculated to obtain the inverse coupling matrix. The inverse coupling matrix represents the strength of the difficulty constraint between knowledge point pairs.
[0039] The forward coupling matrix and the reverse coupling matrix are normalized, and the weighted combination coefficient of the migration probability and the difficulty constraint strength is calculated. Based on the weighted combination coefficient, the migration probability and the difficulty constraint strength are nonlinearly combined to construct a bidirectional coupling representation, which includes a migration probability vector, a constraint strength vector and coupling weight coefficients.
[0040] A second aspect of this invention provides a system for constructing a deep learning-based model of exam question setter behavior, comprising:
[0041] The first unit is used to obtain the question-setting records of the exam question setters, construct a knowledge association analysis module, calculate the usage frequency based on the knowledge point tags in the question-setting records, construct a knowledge point mapping network based on a preset knowledge system, obtain the transfer weight by calculating the conditional probability between knowledge points, combine the knowledge point mapping network and the transfer weight to generate a knowledge application graph, and output a knowledge application vector based on the knowledge application graph.
[0042] The second unit is used to construct a dynamic preference analysis module based on knowledge application vectors. The dynamic preference analysis module groups the question-setting behavior into time windows according to the time information in the question-setting record, extracts features from the knowledge point selection sequence in each time window, and generates a question preference vector by combining the difficulty information.
[0043] The third unit is used to train a proposition behavior recognition network based on proposition preference vectors. The proposition behavior recognition network models the knowledge point selection sequence and difficulty information respectively. By optimizing the migration path and difficulty distribution in the knowledge application graph, proposition behavior features are obtained.
[0044] The fourth unit is used to input the proposition behavior characteristics into the anomaly detection module. The anomaly detection module calculates the warning threshold based on the migration path and difficulty distribution. When the proposition behavior characteristics deviate from the warning threshold, it outputs warning information and the cause of the anomaly.
[0045] A third aspect of the present invention provides an electronic device, comprising:
[0046] processor;
[0047] Memory used to store processor-executable instructions;
[0048] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0049] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0050] In this embodiment, by constructing a deep learning-based model of exam question setter behavior, precise characterization and anomaly monitoring of question-setting behavior are achieved, effectively improving the quality and fairness of exam question setting. The introduction of knowledge application graphs and dynamic preference analysis captures the behavioral patterns of question setters in knowledge point selection and difficulty setting, making the extraction of question-setting behavior features more comprehensive and accurate, and helping to promptly identify deviations and non-standard behaviors in the question-setting process. Optimizing knowledge transfer paths and difficulty distribution using deep learning methods not only automatically identifies question-setting behavior characteristics but also proactively issues warning information based on warning thresholds and analyzes the causes of anomalies, providing exam management departments with a powerful quality monitoring tool and improving the scientific rigor and standardization of exam question setting. Attached Figure Description
[0051] Figure 1 This is a flowchart illustrating the method for constructing a deep learning-based exam question setter behavior model according to an embodiment of the present invention.
[0052] Figure 2This is a flowchart of the proposition behavior anomaly detection and analysis in an embodiment of the present invention. Detailed Implementation
[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0054] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0055] Figure 1 This is a flowchart illustrating the method for constructing a deep learning-based exam question setter behavior model according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0056] The test question setter's question setting records are obtained, and a knowledge association analysis module is constructed. The knowledge association analysis module calculates the frequency of use based on the knowledge point tags in the question setting records, constructs a knowledge point mapping network based on a preset knowledge system, obtains the transfer weight by calculating the conditional probability between knowledge points, combines the knowledge point mapping network and the transfer weight to generate a knowledge application graph, and outputs a knowledge application vector based on the knowledge application graph.
[0057] A dynamic preference analysis module is constructed based on knowledge application vectors. The dynamic preference analysis module groups the question-setting behavior into time windows according to the time information in the question-setting record, extracts features from the knowledge point selection sequence in each time window, and generates a question preference vector by combining the difficulty information.
[0058] A proposition behavior recognition network is trained based on proposition preference vectors. The proposition behavior recognition network models the knowledge point selection sequence and difficulty information respectively. By optimizing the migration path and difficulty distribution in the knowledge application graph, proposition behavior features are obtained.
[0059] The question-setting behavior characteristics are input into the anomaly detection module. The anomaly detection module calculates the warning threshold based on the migration path and difficulty distribution. When the question-setting behavior characteristics deviate from the warning threshold, the module outputs warning information and the cause of the anomaly.
[0060] In one optional implementation, a knowledge association analysis module is constructed. This module calculates the usage frequency based on the knowledge point tags in the proposition records and constructs a knowledge point mapping network based on a preset knowledge system, including:
[0061] The proposition text of the proposition record is converted into a word vector sequence. Bidirectional attention is calculated by query vector and word vector sequence, and candidate fragments of knowledge points are identified based on attention distribution.
[0062] The candidate knowledge point fragments are semantically vectorized, and the semantic similarity between the semantically vectorized representation and the standard knowledge points in the preset knowledge system is calculated. The corresponding knowledge point tags are extracted based on the similarity matching results.
[0063] The frequency of direct use of knowledge point tags in the question record is calculated. At the same time, based on the hierarchical structure of the preset knowledge system, the frequency of use of the parent node and child node corresponding to each knowledge point tag in the question record is counted and weighted to obtain the hierarchical usage frequency of knowledge points containing hierarchical relationships.
[0064] Using knowledge point tags as nodes and hierarchical usage frequency as the initial weight of nodes, node connections are established based on the dependencies between knowledge points in the preset knowledge system to construct an initial knowledge point mapping network; the order of knowledge point selection in the question-setting process is extracted from the question-setting records, and the transfer rules between knowledge points are analyzed based on the order of knowledge point selection to calculate the degree of correlation between nodes based on the transfer rules.
[0065] The connection strength between nodes in the initial knowledge point mapping network is updated based on the degree of relevance. Features are extracted from the updated node weights, and finally, a knowledge point mapping network representing the relationship between knowledge points is constructed.
[0066] In this embodiment, a dataset containing multiple proposition records is first received. Each proposition record includes information such as the proposition text, the proposition time, and the person who set the proposition. For each proposition record, its proposition text is converted into a sequence of word vectors. Specifically, a pre-trained word embedding model (such as Word2Vec) is used to convert each word in the text into a 300-dimensional word vector, forming a word vector sequence. Simultaneously, key knowledge points are extracted from a pre-defined knowledge system as query vectors.
[0067] Bidirectional attention is used to identify candidate knowledge point segments by performing a two-way attention calculation between the query vector and the word vector sequence. In this mechanism, the similarity score between the query vector and each word vector in the word vector sequence is calculated using the vector dot product. For example, the similarity distribution between the query vector "quadratic function" and the proposition text "analyze the properties of a quadratic function" is [0.1, 0.2, 0.8, 0.7, 0.3]. Based on this distribution, consecutive text segments with scores higher than a threshold of 0.5 are identified as candidate knowledge point segments; in this example, "quadratic function" is identified as a candidate segment.
[0068] The identified knowledge point candidate fragments are semantically vectorized. A text encoder (such as BERT) is used to encode the candidate fragments, obtaining 768-dimensional semantic vectors. The semantic similarity between this semantic vector and the standard knowledge points in the preset knowledge system is calculated. For example, the cosine similarity between the semantic vector of the candidate fragment "quadratic function" and the semantic vectors of standard knowledge points such as "quadratic function," "function," and "linear function" in the knowledge system is calculated, yielding similarities of 0.92, 0.75, and 0.68, respectively. Based on the similarity matching results, the standard knowledge point "quadratic function" with the highest similarity is extracted as the knowledge point label for this candidate fragment. The frequency of use of the knowledge point label in the question records is calculated. Direct usage frequency is obtained by counting the number of times each knowledge point label appears in all question records. For example, in 1000 question records, "quadratic function" appears 120 times, and its direct usage frequency is 120. At the same time, based on the hierarchical structure of the preset knowledge system, the usage frequency of the parent and child nodes corresponding to each knowledge point label in the question records is counted. For example, the parent node "Function" of "Quadratic Function" appears 350 times, while the child nodes "Properties of Quadratic Function" and "Graph of Quadratic Function" appear 50 and 40 times respectively. The hierarchical frequency of use for the knowledge point "Quadratic Function" is calculated as follows: the direct usage frequency of 120 plus the parent node usage frequency of 350 multiplied by an upward weight of 0.3, plus the child node usage frequency of (50+40) multiplied by a downward weight of 0.5, resulting in a hierarchical usage frequency of 120 + 350 × 0.3 + (50+40) × 0.5 = 270.
[0069] When constructing the knowledge point mapping network, knowledge point labels are used as nodes, and hierarchical usage frequency is used as the initial weight of the nodes. Node connections are established based on the dependencies between knowledge points in the predefined knowledge system to construct the initial knowledge point mapping network. For example, in the initial network, a connection is established between "function" and "quadratic function," with an initial connection strength set to 0.7, indicating a strong dependency between the two.
[0070] Extract the order of knowledge point selection during the question-setting process from the question-setting records. For example, in one question-setting process, the teacher first selected "function," then "quadratic function," and finally "properties of quadratic function," forming a knowledge point selection sequence ["function," "quadratic function," "properties of quadratic function"]. By analyzing the knowledge point selection sequences in multiple question-setting processes, calculate the transfer probability between knowledge points. For example, the probability of transferring from "function" to "quadratic function" is 0.4, and the probability of transferring from "quadratic function" to "properties of quadratic function" is 0.6.
[0071] The correlation between nodes is calculated based on migration patterns. For a node pair (A, B), the correlation is calculated by combining the direct migration probability and the indirect migration path probability. For example, the correlation between "function" and "properties of quadratic functions" is the direct migration probability of 0.1 plus the probability of the indirect path "function → quadratic function → properties of quadratic functions" of 0.4 × 0.6 = 0.24, for a total correlation of 0.34. The connection strength between nodes in the initial knowledge point mapping network is updated based on the calculated correlation. For example, the connection strength between "function" and "quadratic function" is updated from the initial value of 0.7 to 0.7 × 0.8 + 0.4 × 0.2 = 0.64, where 0.8 and 0.2 are the weights of the initial connection strength and the migration probability, respectively.
[0072] Finally, feature extraction is performed on the updated network, including calculating network features such as the centrality and clustering coefficient of each node. For example, the degree centrality of the "quadratic function" node is 0.85, indicating that this knowledge point has high importance in the entire network. Through these features, a knowledge point mapping network representing the relationships between knowledge points is finally constructed.
[0073] In this embodiment, by introducing deep semantic understanding and knowledge graph modeling, the accurate identification and mapping of implicit knowledge points in the proposition text is achieved, improving the accuracy of knowledge point tag extraction. Simultaneously, by integrating the hierarchical structure of the knowledge system and the proposition order, a graph structure reflecting the frequency of knowledge point usage and transfer relationships is constructed, which helps to depict the actual knowledge application patterns of the question setter. The updated knowledge point mapping network not only reflects the semantic and logical dependencies between knowledge points but also their evolutionary paths in the actual question-setting process, providing higher-quality structured input for subsequent behavior modeling and preference analysis.
[0074] In one optional implementation, transfer weights are obtained by calculating the conditional probabilities between knowledge points. The knowledge point mapping network and the transfer weights are combined to generate a knowledge application graph. Based on the knowledge application graph, a knowledge application vector is output, including:
[0075] Obtain the knowledge point selection sequence from the proposition record, set multiple sliding time windows of different scales for the knowledge point selection sequence, count the number of times the knowledge point appears and the number of consecutive occurrences of knowledge point pairs in each sliding time window, calculate the conditional probability of the knowledge point in the current window, and weight the conditional probabilities of the knowledge points in different sliding time windows to obtain the transfer weight between knowledge points.
[0076] The knowledge application graph is generated by combining the transfer weights between knowledge points with the knowledge point mapping network. The knowledge application graph includes nodes, the connection relationships between nodes, and the transfer weights corresponding to the connection relationships.
[0077] The node feature vectors in the knowledge application graph are transformed by a transformation matrix to obtain the node's transformed feature vector. The transformed feature vectors of adjacent nodes are concatenated and processed by an activation function. The attention coefficients between nodes are obtained through normalization. Based on the attention coefficients, the transformed feature vectors of adjacent nodes are weighted and aggregated to obtain an updated feature vector that integrates the feature information of adjacent nodes.
[0078] Average pooling is performed on the updated feature vectors of all nodes in the knowledge application graph to generate a knowledge application vector, which is used to represent the knowledge point selection pattern and transfer rule.
[0079] In one embodiment, the user's question-asking records during the learning process are first obtained, and the knowledge point selection sequence is extracted from them. For example, for a student learning mathematics, their knowledge point selection sequence might be: [trigonometric functions, vectors, derivatives, integrals, differential equations, probability theory]. Multiple sliding time windows of different scales are set for this sequence, such as four windows: 3 days, 7 days, 14 days, and 30 days. Within each window, the frequency of each knowledge point and the frequency of consecutive occurrences of knowledge point pairs are counted. Taking the 7-day window as an example, if "trigonometric functions" appears 5 times, "vectors" appears 3 times, and "trigonometric functions" is immediately followed by "vectors" 2 times, then the conditional probability of "trigonometric functions" leading to "vectors" is 2 / 5 = 0.4.
[0080] For each sliding time window, the conditional probabilities between knowledge points are calculated. For example, the conditional probability of "trigonometric functions" to "vector" is 0.3 in a 3-day window, 0.4 in a 7-day window, 0.45 in a 14-day window, and 0.5 in a 30-day window. These conditional probabilities are then weighted and combined, with weights set to [0.1, 0.2, 0.3, 0.4], corresponding to the importance of different time windows. By weighted summation, the transfer weight from "trigonometric functions" to "vector" is obtained as 0.3×0.1 + 0.4×0.2 + 0.45×0.3 + 0.5×0.4 = 0.44. Similar calculations are performed for all knowledge point pairs to obtain the complete transfer weight matrix.
[0081] The calculated transfer weights are combined with a pre-constructed knowledge point mapping network to generate a knowledge application graph. The knowledge point mapping network is a graph structure built upon the intrinsic connections between knowledge points, where nodes represent knowledge points and edges represent the relationships between them. For example, in the mathematics knowledge point mapping network, "derivative" is directly connected to knowledge points such as "differential" and "integral." Transfer weights are assigned to these connections to form a weighted knowledge application graph. In this graph, the connection weight from "trigonometric functions" to "vector" is 0.44, representing the probability strength of a learner's transfer from "trigonometric functions" to "vector."
[0082] To fully utilize the information in the knowledge application graph, the node features are processed. Assume each knowledge point node initially has a 64-dimensional feature vector representing its attribute information. A 64×32 transformation matrix is used to transform the node feature vector, resulting in a 32-dimensional transformed feature vector. Taking the "derivative" node as an example, its initial feature vector is transformed to obtain a new 32-dimensional transformed feature vector.
[0083] For each node in the graph, its transformation feature vector is concatenated with that of its neighboring nodes. For example, the "derivative" node is concatenated with the transformation feature vector of its neighboring "differential" node to form a 64-dimensional vector. This concatenated vector is processed using the LeakyReLU activation function and then normalized to obtain the attention coefficients between nodes. For example, the attention coefficient from "derivative" to "differential" is 0.6, and the attention coefficient from "derivative" to "integral" is 0.4. Based on the calculated attention coefficients, the transformation feature vectors of neighboring nodes are weighted and aggregated. For the "derivative" node, its updated feature vector is the result of a weighted sum of the transformation feature vectors of its neighboring nodes "differential" and "integral" with weights of 0.6 and 0.4, respectively. This process ensures that the features of each node not only contain its own information but also incorporate information from neighboring nodes, reflecting the correlation between knowledge points.
[0084] A mean pooling operation is performed on the updated feature vectors of all nodes in the knowledge application graph. This involves summing the updated feature vectors of all nodes and dividing by the number of nodes to generate a unified knowledge application vector. For example, for a knowledge application graph containing 100 knowledge point nodes, the average of these 100 32-dimensional updated feature vectors yields a 32-dimensional knowledge application vector. This vector comprehensively reflects the learner's knowledge point selection patterns and knowledge transfer rules, and can be used for subsequent applications such as learning path recommendation and learning effect prediction.
[0085] In practical applications, knowledge application maps and knowledge utilization vectors can be updated regularly to adapt to changes in learners' knowledge mastery. For example, updates can be made weekly based on the latest learning records to ensure an accurate reflection of learners' current knowledge status and learning patterns. This approach provides learners with more personalized and effective learning support.
[0086] In this embodiment, a multi-scale sliding window is used to capture the co-occurrence and transfer patterns of knowledge points at different time granularities, improving the modeling accuracy of knowledge transfer patterns in question-setting behavior. The knowledge application graph generated by combining conditional probability calculation and a knowledge point mapping network comprehensively reflects the dynamic associations and semantic relationships between knowledge points. Under the graph neural network mechanism, an attention mechanism is introduced to weighted aggregate node features, effectively enhancing the model's ability to identify key knowledge transfer paths. The final generated knowledge application vector can represent the behavioral preferences and logical patterns of question setters in the knowledge selection and transfer process in a high-dimensional way, providing key feature support for subsequent behavior recognition and anomaly detection.
[0087] In one optional implementation, the question-setting behavior is grouped into time windows based on the time information in the question-setting record. Feature extraction is performed on the knowledge point selection sequence within each time window, and a question preference vector is generated by combining it with difficulty information, including:
[0088] A time series is constructed from the time information in the proposition records. Kernel density estimation is performed on the time series to obtain the proposition activity density curve. The time window size is dynamically calculated based on the proposition activity density curve to obtain the time window division of proposition behavior.
[0089] Based on the time window division, the knowledge point selection sequence within each time window is extracted from the question record, a knowledge point word vector mapping table is constructed, the knowledge point selection sequence within each time window is mapped to a word vector sequence, and relative position encoding information is incorporated to obtain a position-aware vector sequence.
[0090] The location-aware vector sequence is input into a bidirectional recurrent neural network, and the context-dependent features of knowledge point selection are extracted to obtain the hidden state sequence.
[0091] The pre-acquired knowledge application vector is transformed into a query vector through nonlinear transformation. The similarity between the query vector and the hidden state sequence is calculated to obtain the attention score. The attention score is then normalized to obtain the attention weight.
[0092] The context vector is obtained by weighted summation of the hidden state sequence based on the attention weights. The context vector is then fused with the query vector and subjected to nonlinear transformation to generate a question preference vector that represents the question setter's preference for selecting knowledge points within the time window.
[0093] In this embodiment, a time series is constructed from the time information in the question-setting records. This time series contains a set of time points when the question setter conducted question-setting activities. For example, for a question setter's question-setting records over the past three months, the timestamp information of each question-setting activity can be extracted to form a time series {t1, t2, ..., t...}. n}
[0094] Kernel density estimation is performed on the time series to obtain the proposition activity density curve. The kernel density estimation process uses a Gaussian kernel function, and the bandwidth parameter is determined through cross-validation. For the example above, a bandwidth of 0.5 days can be selected. The generated density curve represents the distribution of proposition activities on the time axis, with the curve peaks corresponding to periods of frequent proposition activities. The time window size is dynamically calculated based on the proposition activity density curve. Specifically, local minima in the density curve are identified as window boundaries to ensure that proposition activities are relatively concentrated within each window and that activity patterns differ between windows. For the example data, five time windows may be identified, with window sizes of 12 days, 18 days, 25 days, 16 days, and 20 days.
[0095] Based on time window division, the knowledge point selection sequence within each time window is extracted from the question-setting records. For example, in the first window, the knowledge point sequence selected by the question setter might be {"Linear Algebra", "Calculus", "Probability Theory", "Linear Algebra", "Numerical Analysis"}. To process these textual knowledge points, a knowledge point word vector mapping table is constructed. This mapping table is generated using a pre-trained word embedding model, mapping knowledge point concepts to a 64-dimensional vector space. For common subject knowledge points, a word embedding model specifically fine-tuned for the education field can be used to ensure that the vectors accurately reflect the semantic relationships between knowledge points. The knowledge point selection sequence within each time window is mapped to a word vector sequence. For the above example, this results in a sequence of five 64-dimensional vectors. To capture the order information of knowledge point selection, relative position encoding information is incorporated to obtain a position-aware vector sequence. The position encoding uses a sine-cosine position encoding method, encoding the position information into a vector with the same dimension as the knowledge point vector. Then, the position encoding and the knowledge point vector are fused through an addition operation to obtain a position-aware knowledge point representation.
[0096] The position-aware vector sequence is input into a bidirectional recurrent neural network (BRNN), and context-dependent features of knowledge point selection are extracted to obtain the hidden state sequence. The BRNN adopts a gated recurrent unit structure with a hidden layer dimension of 128 and contains two hidden layers. For each knowledge point, forward propagation captures information about previously selected knowledge points, and back propagation captures information about subsequently selected knowledge points; the two are combined to form a complete contextual representation. For example, for "probability theory" in the example sequence, its hidden state will include the influence of "linear algebra" and "calculus," as well as the influence on subsequent selections of "linear algebra" and "numerical analysis," generating a 128-dimensional vector representing the semantic and positional importance of this knowledge point in the overall selection sequence.
[0097] The pre-acquired knowledge application vector is transformed into a query vector through a nonlinear transformation. The knowledge application vector, extracted from the question record, reflects the question setter's application of knowledge points and has a dimension of 32. It is expanded to 128 dimensions through a fully connected layer transformation, and the activation function is a hyperbolic tangent function, resulting in a query vector that matches the hidden layer state dimension. The similarity between the query vector and each state vector in the hidden layer state sequence is calculated to obtain an attention score, using a dot product operation. The attention scores are then Softmax normalized to obtain attention weights, which reflect the importance of each knowledge point in shaping question preferences.
[0098] The context vector is obtained by weighting and summing the hidden state sequences based on attention weights. For example, if the attention weights of five knowledge points within a window are [0.15, 0.25, 0.3, 0.2, 0.1], the corresponding hidden state vectors are weighted and summed according to these weights to generate a context vector that reflects the knowledge point selection preference within the entire window. The context vector and query vector are then fused through concatenation and transformed nonlinearly by a fully connected layer. The activation function is ReLU, and the output dimension is 96, generating a question-setting preference vector that represents the question setter's knowledge point selection preference within the time window. This vector contains both the question setter's knowledge point selection pattern information and the characteristics of knowledge application, comprehensively reflecting the question setter's question-setting preference within a specific time window.
[0099] In practical applications, such as analyzing a math exam setter's exam question creation records over a three-month period, the above method can yield a question preference vector for each time window. These vectors can be further used for exam question creation behavior analysis, anomaly detection, or personalized exam question recommendations. For example, cluster analysis reveals a clear pattern in the exam setter's preferences across different time windows: a preference for basic concept questions at the beginning of the month, a preference for calculation application questions in the middle of the month, and a preference for comprehensive application questions at the end of the month. This provides data support for understanding the exam setter's behavioral patterns.
[0100] In this embodiment, by estimating the kernel density of the question-setting time information, dynamic time window segmentation of the question-setting behavior is achieved, which can more accurately capture the phased changes in the question setter's behavior. A position-aware vector sequence is constructed by combining the knowledge point selection order, and a bidirectional recurrent neural network is introduced to effectively model the contextual dependencies of knowledge points, improving the ability to characterize the question setter's thought process. Simultaneously, by utilizing the attention mechanism associated with the knowledge application vector, a weighted extraction of the question setter's key interests is achieved, thereby generating a question preference vector that better reflects question-setting preferences and knowledge transfer tendencies, providing a more discriminative feature representation for subsequent question-setting behavior analysis and personalized identification.
[0101] In one optional implementation, a proposition behavior recognition network is trained based on proposition preference vectors. This network models the knowledge point selection sequence and difficulty information, respectively. By optimizing the migration path and difficulty distribution in the knowledge application graph, proposition behavior features are obtained, including:
[0102] A migration path identification module is constructed based on the proposition preference vector. The proposition preference vector is input into the feature extraction network to obtain initial features. A feature propagation matrix is generated based on the migration weights between nodes in the knowledge application graph. The initial features are aggregated using the feature propagation matrix to obtain the migration path features of the knowledge point selection sequence.
[0103] The difficulty information is encoded into a difficulty vector, and the correlation score between the difficulty vector and the migration path features is calculated. Based on the correlation score, the migration pattern of the knowledge point selection sequence under each difficulty level is identified.
[0104] Based on the migration pattern of the knowledge point selection sequence, a difficulty distribution weight is constructed. The difficulty distribution weight and the migration path feature are weighted and calculated to obtain the difficulty adaptive migration path feature. At the same time, the difficulty vector is corrected based on the difficulty adaptive migration path feature to obtain the corrected difficulty vector.
[0105] The difficulty-adaptive migration path features and the modified difficulty vector are fused to generate propositional behavior features that reflect the synergistic relationship between the knowledge point selection sequence and difficulty information.
[0106] In this implementation, a transfer path recognition module is first constructed based on the proposition preference vector. The core function of this module is to capture the transfer patterns of the question setter during the knowledge point selection process. Specifically, the previously generated 96-dimensional proposition preference vector is input into a feature extraction network to obtain initial features. The feature extraction network adopts a multilayer perceptron structure, containing three fully connected layers with 128, 64, and 32 neurons in each layer. The ReLU activation function is used, and batch normalization layers are used to improve training stability. Taking a question setter in a certain mathematics subject as an example, their proposition preference vector is processed by the feature extraction network to obtain a 32-dimensional initial feature vector. A feature propagation matrix is generated based on the transfer weights between nodes in the knowledge application graph. The knowledge application graph is a pre-constructed directed graph reflecting the relationships between knowledge points, where nodes represent knowledge points and edges represent transfer relationships. For example, in mathematics, there is a significant transfer relationship from "differentiation of univariate functions" to "differentiation of multivariate functions," with a transfer weight of 0.75, while the transfer relationship from "probability theory" to "linear algebra" is weaker, with a weight of 0.15. For a knowledge application graph consisting of n core knowledge points, an n×n dimensional feature propagation matrix is generated, where the element values are determined by the transfer weights between corresponding knowledge points. The initial features are then aggregated using the feature propagation matrix to obtain the transfer path features of the knowledge point selection sequence. The aggregation process employs a graph convolutional network mechanism; for each knowledge point node, information is aggregated based on its connections with neighboring nodes. After two iterations, a 32-dimensional transfer path feature vector containing knowledge point transfer pattern information is generated.
[0107] Difficulty information is encoded into a difficulty vector, derived from the difficulty level labels in the question record. Assuming five difficulty levels—basic, easy, medium, difficult, and hard—one-hot encoding is used to convert each level into a 5-dimensional difficulty vector. For question records containing multiple questions, the distribution ratio of each difficulty level is statistically analyzed. For example, if a question setter's difficulty distribution is [0.1, 0.2, 0.4, 0.2, 0.1], it indicates that basic questions account for 10%, easy questions account for 20%, and so on. The correlation score between the difficulty vector and the transfer path features is calculated. Specifically, the difficulty vector is mapped to a 32-dimensional space identical to the transfer path features through a fully connected layer, and then the cosine similarity between the two vectors is calculated to obtain the correlation score between each of the five difficulty levels and the transfer path features. Based on the correlation score, the transfer pattern of the knowledge point selection sequence under each difficulty level is identified. A high correlation score indicates a high degree of association between that difficulty level and a specific transfer pattern. In practical applications, it may be observed that when setting "relatively difficult" and "hard" level questions, the question setter tends to choose the transition path from "differential equations" to "complex functions," with correlation scores of 0.78 and 0.82, respectively. However, when setting "basic" and "easy" level questions, the question setter prefers the transition path from "linear algebra" to "analytic geometry," with correlation scores of 0.85 and 0.76, respectively.
[0108] The difficulty distribution weights are constructed based on the knowledge point selection sequence transfer pattern. Specifically, the relevance scores of each difficulty level are normalized using a Softmax function to obtain a weight vector reflecting the importance of each difficulty level. The difficulty distribution weights are then weighted with the transfer path features to obtain adaptive transfer path features. The weighting calculation uses element-wise multiplication, where each dimension of the transfer path features is multiplied by its corresponding difficulty distribution weight, generating a 32-dimensional feature vector that adaptively reflects the characteristics of the transfer pattern under different difficulty levels. Simultaneously, the difficulty vector is corrected based on the adaptive transfer path features to obtain a modified difficulty vector. The correction process uses an attention mechanism, employing the adaptive transfer path features as the query vector and the difficulty vector as the key-value vector, to generate a 5-dimensional modified difficulty vector that better reflects the actual difficulty preferences of the test creators. For example, although a test setter claims to prefer medium-difficulty questions, actual test-setting behavior analysis shows that the difficulty of the questions on the "probability and statistics" knowledge point is generally higher than self-assessment. The revised difficulty vector may be adjusted from [0.1, 0.2, 0.4, 0.2, 0.1] to [0.05, 0.15, 0.35, 0.3, 0.15].
[0109] This paper fuses the difficulty-adaptive migration path features and the modified difficulty vector to generate a question-setting behavior feature that reflects the synergistic relationship between knowledge point selection sequence and difficulty information. The feature fusion employs a concatenation-then-transformation strategy, concatenating the 32-dimensional difficulty-adaptive migration path features with the 5-dimensional modified difficulty vector to form a 37-dimensional vector. This 37-dimensional vector is then transformed into a final 48-dimensional question-setting behavior feature vector through a two-layer fully connected network. The first fully connected layer outputs 64 dimensions using the Leaky ReLU activation function; the second layer outputs 48 dimensions using the Tanh activation function to ensure that the feature values are distributed within the range [-1, 1]. The resulting question-setting behavior feature vector comprehensively captures the behavioral characteristics of question setters in both knowledge point selection and difficulty control, providing an effective feature representation for subsequent question-setting behavior analysis, personalized question recommendation, and question quality assessment.
[0110] In a case study analyzing the question-setting behavior of a mathematics professor at a university, the question-setting behavior features extracted using the above method successfully identified that the professor preferred a question-setting pattern that progressed from simple to complex and of increasing difficulty when setting questions for the probability theory unit, while tending to use a question-setting strategy that mixed difficulty and highlighted key knowledge points in the numerical analysis unit. This personalized behavior feature provides an important reference for intelligent question-setting assistance systems.
[0111] In this embodiment, by constructing a migration path identification module and combining it with question preference vectors, migration features in the knowledge point selection sequence are effectively extracted, achieving accurate modeling of the question setter's knowledge application path. Simultaneously, correlation modeling is performed between difficulty information and migration paths to identify knowledge point migration patterns at different difficulty levels, enabling the model to possess a fine-grained understanding of question-setting strategies. By fusing difficulty distribution and migration features, an adaptive difficulty representation is generated, and the difficulty vector is corrected, improving the model's accuracy in recognizing the relationship between question structure and cognitive load. The final output question-setting behavior features accurately reflect the question setter's collaborative behavior in knowledge transfer and difficulty control, providing a high-quality foundation for question-setting behavior prediction and anomaly detection.
[0112] like Figure 2 The diagram illustrates the process for detecting and analyzing anomalies in propositional behavior in this embodiment.
[0113] In one optional implementation, the proposition behavior characteristics are input into an anomaly detection module, which calculates a warning threshold based on the migration path and difficulty distribution. When the proposition behavior characteristics deviate from the warning threshold, the module outputs warning information and the cause of the anomaly, including:
[0114] The proposition behavior features are input into the anomaly detection module. The anomaly detection module extracts migration path information and difficulty distribution information. Based on the migration path information, it calculates the migration probability between knowledge points to obtain a positive coupling matrix. Based on the difficulty distribution information, it calculates the difficulty constraint strength of knowledge point pairs to obtain a negative coupling matrix. The positive coupling matrix and the negative coupling matrix are combined to construct a bidirectional coupling representation.
[0115] The bidirectional coupling representation is used to forward propagate migration path information to obtain path features, and backward propagate difficulty distribution information to obtain difficulty features. The path features and difficulty features are then interactively calculated to obtain the coupling anomaly score.
[0116] The stability index of knowledge point transfer is calculated based on the bidirectional coupling representation. The distribution of coupling anomaly scores is dynamically corrected using the stability index to obtain an early warning threshold. When the coupling anomaly score exceeds the early warning threshold, anomaly detection is triggered.
[0117] An anomaly propagation network is constructed based on the bidirectional coupling representation. In the anomaly propagation network, the knowledge point migration sequence that causes the anomaly is located along the forward path, and the knowledge point combination that violates the difficulty constraint is identified along the reverse path. The located anomaly knowledge point migration sequence and the knowledge point combination that violates the difficulty constraint are used as the cause of the anomaly to output early warning information.
[0118] For example, the anomaly detection module first extracts migration path information and difficulty distribution information from the question-setting behavior characteristics. Migration path information records the transformation relationships between knowledge points during the question-setting process, such as the migration frequency from "quadratic function" to "function extrema". Difficulty distribution information includes the difficulty value of each knowledge point and its combinations. Based on the migration path information, the migration probability between knowledge points is calculated, and a positive coupling matrix is constructed. Specifically, for the migration from knowledge point A to knowledge point B, the migration probability is obtained by dividing the frequency of A followed by B in historical question data by the total frequency of A occurrences. For example, if the frequency of "linear equation" followed by "system of two linear equations" is 80 times, and "linear equation" appears a total of 120 times, then the migration probability is 0.67.
[0119] The difficulty constraint strength of knowledge point pairs is calculated based on difficulty distribution information, and a reverse coupling matrix is constructed. The difficulty constraint strength reflects the degree of correlation between two knowledge points in terms of difficulty. In implementation, the distribution of difficulty differences between knowledge point pairs in historical data is analyzed, and the reciprocal of the standard deviation of the difficulty difference is calculated as the constraint strength. For example, if the standard deviation of the difficulty difference between the knowledge points "derivative" and "differential" in historical data is 0.2, then the difficulty constraint strength between them is 5.
[0120] A bidirectional coupling representation is constructed by combining the forward coupling matrix and the reverse coupling matrix. This representation contains two parts of information: the transfer relationship between knowledge points and the difficulty constraint relationship. In practical implementation, the bidirectional coupling representation can be stored in a three-dimensional data structure, where each pair of knowledge points has both a transfer probability value and a difficulty constraint strength value. Using the bidirectional coupling representation, the transfer path information is forward propagated to obtain path features. During the forward propagation process, the reasonableness score of the path is accumulated and calculated along the knowledge point transfer sequence in the proposition. For example, for the transfer sequence "function → derivative → integral", the overall reasonableness of the path is calculated based on the corresponding transfer probabilities in the forward coupling matrix. If the transfer probability of "function → derivative" is 0.8 and the transfer probability of "derivative → integral" is 0.7, then the path feature value may be 0.56.
[0121] Simultaneously, the difficulty distribution information is backpropagated to obtain difficulty characteristics. During the backpropagation process, it checks whether the difficulty differences between adjacent knowledge points in the question conform to historical statistical patterns. For example, if the difficulty difference between "vector" and "matrix" in historical data is usually around 0.2, while this difference is 0.5 in the current question, the system will calculate a high difficulty outlier.
[0122] The coupling anomaly score is obtained by interactively calculating path features and difficulty features. The interactive calculation takes into account the combined effects of path rationality and difficulty rationality. For example, by using a weighted average, the path features are assigned a weight of 0.6 and the difficulty features a weight of 0.4 to obtain the final coupling anomaly score.
[0123] Based on the bidirectional coupling representation, a stability index for knowledge point transfer is calculated. This stability index reflects the consistency of knowledge point transfer patterns in historical data. In implementation, the entropy value of the transfer probabilities in the forward coupling matrix can be calculated; a lower entropy value indicates a more stable transfer pattern. For example, if the outgoing edge transfer probabilities of the knowledge point "equation" are concentrated and the entropy value is 0.3, then its stability index is 0.7.
[0124] A dynamic correction method is used to adjust the distribution of coupling anomaly scores using a stability index to obtain an early warning threshold. The stringency of the early warning threshold is adjusted according to the stability index, with stricter thresholds used for knowledge point transfers with higher stability. For example, for knowledge point transfers with a stability index of 0.8, the early warning threshold might be set to 0.7; while for knowledge point transfers with a stability index of 0.4, the early warning threshold might be relaxed to 0.85. Anomaly detection is triggered when the coupling anomaly score exceeds the early warning threshold.
[0125] An anomaly propagation network is constructed based on bidirectional coupling representation. This network simulates the propagation process of anomalous states in the knowledge point graph, helping to locate the source of the anomaly. In the anomaly propagation network, the knowledge point migration sequence leading to the anomaly is located along the forward path. For example, if an anomalous proposition contains a migration sequence of "geometric proof → algebraic operation → calculus," and the forward coupling matrix shows that the migration probability of "geometric proof → algebraic operation" is extremely low (e.g., 0.05), then this migration will be identified as one of the causes of the anomaly. Simultaneously, combinations of knowledge points that violate difficulty constraints are identified along the reverse path. For example, if the difficulty of "simple function" in a proposition is set to 0.8, while the difficulty of "complex integral" is set to 0.6, violating the constraint that the difficulty should increase step by step, then this combination will be identified as the cause of the anomaly.
[0126] Finally, the identified abnormal knowledge point migration sequences and knowledge points violating difficulty constraints are combined as the causes of the anomalies to generate warning messages. These warning messages include the anomaly type, location, and severity. For example: "An abnormal migration of knowledge points from 'Linear Algebra' to 'Probability Theory' was detected, with a migration probability of only 0.03; the question's logic should be re-examined," or "The difficulty setting of knowledge points 'Elementary Functions' (difficulty 0.9) and 'Function Graphs' (difficulty 0.4) is found to be unreasonable, violating the principle of progressive difficulty." These warning messages help teachers improve their question-setting practices and enhance the quality of test questions.
[0127] In this embodiment, by constructing a bidirectional coupled representation that integrates knowledge transfer probability and difficulty constraint strength, the intrinsic relationship between knowledge point selection paths and difficulty control in question-setting behavior can be comprehensively characterized. During anomaly detection, a coupled anomaly score is generated by the interactive calculation of path features and difficulty features, and dynamic threshold correction is performed in conjunction with the knowledge point transfer stability index, effectively improving the sensitivity and robustness of anomaly detection. Furthermore, the introduction of an anomaly propagation network can accurately locate the migration sequence and difficulty conflict point that triggers the anomaly, providing a clear explanation of the anomaly cause, thereby achieving timely and explainable early warning of anomalies in question-setting behavior and improving the ability to monitor question quality and manage risks.
[0128] In one optional implementation, a forward coupling matrix is obtained by calculating the migration probability between knowledge points based on the migration path information, and a reverse coupling matrix is obtained by calculating the difficulty constraint strength of knowledge point pairs based on the difficulty distribution information. Combining the forward coupling matrix and the reverse coupling matrix to construct a bidirectional coupling representation includes:
[0129] Based on the migration sequence and temporal features in the migration path information, the number of migrations between knowledge points is counted according to the time window. The migration statistics results under different time windows are weighted and combined to calculate the migration probability between knowledge points and obtain the positive coupling matrix. The positive coupling matrix represents the migration probability from the source knowledge point to the target knowledge point.
[0130] A hierarchical difficulty constraint graph is constructed based on the difficulty distribution information. In the hierarchical difficulty constraint graph, knowledge points are organized into layers according to their difficulty values. The difficulty gradient of knowledge point pairs within a layer and the cross-layer constraints of knowledge point pairs between layers are calculated to obtain the inverse coupling matrix. The inverse coupling matrix represents the strength of the difficulty constraint between knowledge point pairs.
[0131] The forward coupling matrix and the reverse coupling matrix are normalized, and the weighted combination coefficient of the migration probability and the difficulty constraint strength is calculated. Based on the weighted combination coefficient, the migration probability and the difficulty constraint strength are nonlinearly combined to construct a bidirectional coupling representation, which includes a migration probability vector, a constraint strength vector and coupling weight coefficients.
[0132] In this implementation, firstly, based on the migration sequence and temporal characteristics in the migration path information, the number of migrations between knowledge points is counted according to time windows. Specifically, the question setter's question-setting records are grouped into time windows according to the aforementioned method. Within each time window, the sequence of knowledge points selected by the question setter is extracted. For example, a math teacher might select the following knowledge point sequences in five time windows within a month: Window 1: ["Limits", "Derivatives", "Integrals", "Differential Equations"]; Window 2: ["Linear Algebra", "Matrices", "Eigenvalues"]; Window 3: ["Probability Theory", "Random Variables", "Mathematical Statistics"]; Window 4: ["Differential Equations", "Numerical Analysis", "Interpolation"]; Window 5: ["Complex Functions", "Series Expansion", "Residue Theorem"]. For each knowledge point sequence within a window, the number of migrations between adjacent knowledge points is counted; for example, a migration from "Limits" to "Derivatives" is counted as 1 time, and a migration from "Derivatives" to "Integrals" is counted as 1 time. The migration statistics from different time windows are weighted and combined, with migration data from more recent time windows assigned higher weights. Assuming the weights of the five windows are 0.1, 0.15, 0.2, 0.25, and 0.3, the migration counts between each pair of knowledge points are multiplied by their corresponding window weights and then summed. The migration probabilities between knowledge points are calculated to obtain a positive coupling matrix, which represents the migration probability from the source knowledge point to the target knowledge point. For n knowledge points, an n×n matrix is constructed, where the matrix element values represent the probability of migration from a row-indexed knowledge point to a column-indexed knowledge point. For example, the migration probability from "differential equations" to "numerical analysis" might be 0.75, meaning that if the test setter selects "differential equations," there is a 75% probability that they will choose "numerical analysis" as the next knowledge point.
[0133] A hierarchical difficulty constraint graph is constructed based on difficulty distribution information. The difficulty distribution information comes from the statistical analysis of the difficulty values of each knowledge point in the question record. First, the average difficulty value of each knowledge point is calculated; for example, the average difficulty of "limit" is 3.2, "derivative" is 3.5, and "integral" is 4.1 (assuming the difficulty value range is 1-5). In the hierarchical difficulty constraint graph, knowledge points are organized into layers according to their difficulty values. For example, difficulty values of 1-1.9 are divided into the first layer, 2-2.9 into the second layer, and so on, for a total of five layers. The inverse coupling matrix is obtained by calculating the difficulty gradient of knowledge point pairs within a layer and the cross-layer constraints of knowledge point pairs between layers. The method for calculating the intra-layer difficulty gradient is to calculate the absolute value of the difference in difficulty values between knowledge point pairs within the same layer; the smaller the difference, the stronger the constraint. For example, if "matrix" and "eigenvalue" are both located in the third layer, with difficulty values of 3.3 and 3.4 respectively, then the difficulty gradient between them is 0.1, and the corresponding constraint strength may be 0.9. The method for calculating inter-layer cross-layer constraints involves calculating the constraint strength for pairs of knowledge points from different layers based on the difference in their layer level and difficulty value. The larger the difference in layer level and difficulty value, the smaller the constraint strength. For example, the layer level difference between "limit" (third layer) and "residue theorem" (fifth layer) is 2, and the difficulty value difference is 1.7, so the corresponding constraint strength might be 0.3. The reverse coupling matrix is also an n×n dimensional matrix, where the matrix element values represent the difficulty constraint strength between knowledge points at row and column indices.
[0134] The forward and reverse coupling matrices are normalized to ensure that the matrix element values are distributed within the range of 0-1 and that the sum of the elements in each row is 1. Row-level normalization is used, where each row element is divided by the sum of all elements in that row. A weighted combination coefficient of transfer probability and difficulty constraint strength is calculated, which determines the relative importance of transfer probability and difficulty constraint strength in the bidirectional coupling representation. The calculation of the combination coefficient considers the characteristics of the test setter's test-setting behavior, such as preferences for test coherence and difficulty control. Specifically, the combination coefficient is calculated based on the components related to coherence and difficulty control in the test-setting behavior feature vector. For example, for a test setter who emphasizes knowledge point coherence, the weight of transfer probability might be 0.7, and the weight of difficulty constraint strength might be 0.3; while for a test setter who emphasizes a balanced difficulty level, the weight of transfer probability might be 0.4, and the weight of difficulty constraint strength might be 0.6.
[0135] A bidirectional coupling representation is constructed by nonlinearly combining the transfer probability and difficulty constraint strength using weighted combination coefficients. The nonlinear combination employs a weighted harmonic average method, whereby for each pair of knowledge points, the bidirectional coupling strength is calculated based on the transfer probability value, difficulty constraint strength value, and combination coefficients. The bidirectional coupling representation includes a transfer probability vector, a constraint strength vector, and coupling weight coefficients. For n knowledge points, the transfer probability vector and constraint strength vector each have n×n elements, corresponding to the elements of the forward coupling matrix and the reverse coupling matrix, respectively; the coupling weight coefficient is a scalar value representing the relative importance of the transfer probability and the difficulty constraint strength.
[0136] This bidirectional coupling effectively captures the complex relationship patterns between knowledge points, which are influenced by both transfer logic and difficulty control. This provides an important basis for building a question recommendation system that better reflects the actual behavior of question setters. For example, for teachers who prefer to set questions progressively from basic to applied knowledge, the system will prioritize recommending combinations of knowledge points with high transfer probability and moderate difficulty gradients; while for teachers who prefer challenging questions, the system will recommend combinations of knowledge points with a greater difficulty range but still maintaining a certain transfer logic.
[0137] In this embodiment, by fusing migration path information and difficulty distribution information, a bidirectional coupled representation is constructed to accurately model the synergistic relationship between logical migration of knowledge points and difficulty control in propositional behavior. The forward coupling matrix characterizes the migration tendency between knowledge points, while the reverse coupling matrix reflects the strength of difficulty constraints. Both are normalized and a weighted combination mechanism is introduced, enabling the coupled representation to adapt to knowledge migration patterns while also considering difficulty gradient control, effectively improving the accuracy and discriminative ability of behavior modeling. This method can provide a richer and more interpretable semantic feature foundation for subsequent anomaly detection and behavior recognition, thereby enhancing the system's ability to model and analyze complex propositional behaviors.
[0138] A second aspect of this invention provides a system for constructing a deep learning-based model of exam question setter behavior, the system comprising:
[0139] The first unit is used to obtain the question-setting records of the exam question setters, construct a knowledge association analysis module, calculate the usage frequency based on the knowledge point tags in the question-setting records, construct a knowledge point mapping network based on a preset knowledge system, obtain the transfer weight by calculating the conditional probability between knowledge points, combine the knowledge point mapping network and the transfer weight to generate a knowledge application graph, and output a knowledge application vector based on the knowledge application graph.
[0140] The second unit is used to construct a dynamic preference analysis module based on knowledge application vectors. The dynamic preference analysis module groups the question-setting behavior into time windows according to the time information in the question-setting record, extracts features from the knowledge point selection sequence in each time window, and generates a question preference vector by combining the difficulty information.
[0141] The third unit is used to train a proposition behavior recognition network based on proposition preference vectors. The proposition behavior recognition network models the knowledge point selection sequence and difficulty information respectively. By optimizing the migration path and difficulty distribution in the knowledge application graph, proposition behavior features are obtained.
[0142] The fourth unit is used to input the proposition behavior characteristics into the anomaly detection module. The anomaly detection module calculates the warning threshold based on the migration path and difficulty distribution. When the proposition behavior characteristics deviate from the warning threshold, it outputs warning information and the cause of the anomaly.
[0143] A third aspect of the present invention provides an electronic device, comprising:
[0144] processor;
[0145] Memory used to store processor-executable instructions;
[0146] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0147] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0148] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.
[0149] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for constructing a deep learning-based model of exam question setter behavior, characterized in that, include: The test question setter's question setting records are obtained, and a knowledge association analysis module is constructed. The knowledge association analysis module calculates the frequency of use based on the knowledge point tags in the question setting records, constructs a knowledge point mapping network based on a preset knowledge system, obtains the transfer weight by calculating the conditional probability between knowledge points, combines the knowledge point mapping network and the transfer weight to generate a knowledge application graph, and outputs a knowledge application vector based on the knowledge application graph. A dynamic preference analysis module is constructed based on knowledge application vectors. The dynamic preference analysis module groups the question-setting behavior into time windows according to the time information in the question-setting record, extracts features from the knowledge point selection sequence in each time window, and generates a question preference vector by combining the difficulty information. A proposition behavior recognition network is trained based on proposition preference vectors. The proposition behavior recognition network models the knowledge point selection sequence and difficulty information respectively. By optimizing the migration path and difficulty distribution in the knowledge application graph, proposition behavior features are obtained. The question-setting behavior characteristics are input into the anomaly detection module. The anomaly detection module calculates the warning threshold based on the migration path and difficulty distribution. When the question-setting behavior characteristics deviate from the warning threshold, the warning information and the cause of the anomaly are output. A knowledge association analysis module is constructed, which calculates the frequency of use based on the knowledge point tags in the proposition records and constructs a knowledge point mapping network based on a preset knowledge system, including: The proposition text of the proposition record is converted into a word vector sequence. Bidirectional attention is calculated by query vector and word vector sequence, and candidate fragments of knowledge points are identified based on attention distribution. The candidate knowledge point fragments are semantically vectorized, and the semantic similarity between the semantically vectorized representation and the standard knowledge points in the preset knowledge system is calculated. The corresponding knowledge point tags are extracted based on the similarity matching results. The frequency of direct use of knowledge point tags in the question record is calculated. At the same time, based on the hierarchical structure of the preset knowledge system, the frequency of use of the parent node and child node corresponding to each knowledge point tag in the question record is counted and weighted to obtain the hierarchical usage frequency of knowledge points containing hierarchical relationships. Using knowledge point tags as nodes and hierarchical usage frequency as the initial weight of nodes, node connections are established based on the dependencies between knowledge points in the preset knowledge system to construct an initial knowledge point mapping network; the order of knowledge point selection in the question-setting process is extracted from the question-setting records, and the transfer rules between knowledge points are analyzed based on the order of knowledge point selection to calculate the degree of correlation between nodes based on the transfer rules. The connection strength between nodes in the initial knowledge point mapping network is updated based on the degree of relevance. Features are extracted from the updated node weights, and finally, a knowledge point mapping network representing the relationship between knowledge points is constructed.
2. The method according to claim 1, characterized in that, Transfer weights are obtained by calculating the conditional probabilities between knowledge points. The knowledge point mapping network and transfer weights are combined to generate a knowledge application graph. Based on the knowledge application graph, the knowledge application vector is output, including: Obtain the knowledge point selection sequence from the proposition record, set multiple sliding time windows of different scales for the knowledge point selection sequence, count the number of times the knowledge point appears and the number of consecutive occurrences of knowledge point pairs in each sliding time window, calculate the conditional probability of the knowledge point in the current window, and weight the conditional probabilities of the knowledge points in different sliding time windows to obtain the transfer weight between knowledge points. The knowledge application graph is generated by combining the transfer weights between knowledge points with the knowledge point mapping network. The knowledge application graph includes nodes, the connection relationships between nodes, and the transfer weights corresponding to the connection relationships. The node feature vectors in the knowledge application graph are transformed by a transformation matrix to obtain the node's transformed feature vector. The transformed feature vectors of adjacent nodes are concatenated and processed by an activation function. The attention coefficients between nodes are obtained through normalization. Based on the attention coefficients, the transformed feature vectors of adjacent nodes are weighted and aggregated to obtain an updated feature vector that integrates the feature information of adjacent nodes. Average pooling is performed on the updated feature vectors of all nodes in the knowledge application graph to generate a knowledge application vector, which is used to represent the knowledge point selection pattern and transfer rule.
3. The method according to claim 1, characterized in that, Based on the time information in the question-setting records, question-setting behaviors are grouped by time windows. Feature extraction is performed on the knowledge point selection sequence within each time window, and combined with difficulty information to generate a question-setting preference vector, including: A time series is constructed from the time information in the proposition records. Kernel density estimation is performed on the time series to obtain the proposition activity density curve. The time window size is dynamically calculated based on the proposition activity density curve to obtain the time window division of proposition behavior. Based on the time window division, the knowledge point selection sequence within each time window is extracted from the question record, a knowledge point word vector mapping table is constructed, the knowledge point selection sequence within each time window is mapped to a word vector sequence, and relative position encoding information is incorporated to obtain a position-aware vector sequence. The location-aware vector sequence is input into a bidirectional recurrent neural network, and the context-dependent features of knowledge point selection are extracted to obtain the hidden state sequence. The pre-acquired knowledge application vector is transformed into a query vector through nonlinear transformation. The similarity between the query vector and the hidden state sequence is calculated to obtain the attention score. The attention score is then normalized to obtain the attention weight. The context vector is obtained by weighted summation of the hidden state sequence based on the attention weights. The context vector is then fused with the query vector and subjected to nonlinear transformation to generate a question preference vector that represents the question setter's preference for selecting knowledge points within the time window.
4. The method according to claim 1, characterized in that, A proposition behavior recognition network is trained based on proposition preference vectors. This network models the knowledge point selection sequence and difficulty information separately. By optimizing the transfer path and difficulty distribution in the knowledge application graph, proposition behavior features are obtained, including: A migration path identification module is constructed based on the proposition preference vector. The proposition preference vector is input into the feature extraction network to obtain initial features. A feature propagation matrix is generated based on the migration weights between nodes in the knowledge application graph. The initial features are aggregated using the feature propagation matrix to obtain the migration path features of the knowledge point selection sequence. The difficulty information is encoded into a difficulty vector, and the correlation score between the difficulty vector and the migration path features is calculated. Based on the correlation score, the migration pattern of the knowledge point selection sequence under each difficulty level is identified. Based on the migration pattern of the knowledge point selection sequence, a difficulty distribution weight is constructed. The difficulty distribution weight and the migration path feature are weighted and calculated to obtain the difficulty adaptive migration path feature. At the same time, the difficulty vector is corrected based on the difficulty adaptive migration path feature to obtain the corrected difficulty vector. The difficulty-adaptive migration path features and the modified difficulty vector are fused to generate propositional behavior features that reflect the synergistic relationship between the knowledge point selection sequence and difficulty information.
5. The method according to claim 1, characterized in that, The question-setting behavior characteristics are input into the anomaly detection module. The anomaly detection module calculates a warning threshold based on the migration path and difficulty distribution. When the question-setting behavior characteristics deviate from the warning threshold, it outputs warning information and the cause of the anomaly, including: The proposition behavior features are input into the anomaly detection module. The anomaly detection module extracts migration path information and difficulty distribution information. Based on the migration path information, it calculates the migration probability between knowledge points to obtain a positive coupling matrix. Based on the difficulty distribution information, it calculates the difficulty constraint strength of knowledge point pairs to obtain a negative coupling matrix. The positive coupling matrix and the negative coupling matrix are combined to construct a bidirectional coupling representation. The bidirectional coupling representation is used to forward propagate migration path information to obtain path features, and backward propagate difficulty distribution information to obtain difficulty features. The path features and difficulty features are then interactively calculated to obtain the coupling anomaly score. The stability index of knowledge point transfer is calculated based on the bidirectional coupling representation. The distribution of coupling anomaly scores is dynamically corrected using the stability index to obtain an early warning threshold. When the coupling anomaly score exceeds the early warning threshold, anomaly detection is triggered. An anomaly propagation network is constructed based on the bidirectional coupling representation. In the anomaly propagation network, the knowledge point migration sequence that causes the anomaly is located along the forward path, and the knowledge point combination that violates the difficulty constraint is identified along the reverse path. The located anomaly knowledge point migration sequence and the knowledge point combination that violates the difficulty constraint are used as the cause of the anomaly to output early warning information.
6. The method according to claim 5, characterized in that, Based on the migration path information, the migration probability between knowledge points is calculated to obtain a forward coupling matrix. Based on the difficulty distribution information, the difficulty constraint strength of knowledge point pairs is calculated to obtain a reverse coupling matrix. Combining the forward coupling matrix and the reverse coupling matrix to construct a bidirectional coupling representation includes: Based on the migration sequence and temporal features in the migration path information, the number of migrations between knowledge points is counted according to the time window. The migration statistics results under different time windows are weighted and combined to calculate the migration probability between knowledge points and obtain the positive coupling matrix. The positive coupling matrix represents the migration probability from the source knowledge point to the target knowledge point. A hierarchical difficulty constraint graph is constructed based on the difficulty distribution information. In the hierarchical difficulty constraint graph, knowledge points are organized into layers according to their difficulty values. The difficulty gradient of knowledge point pairs within a layer and the cross-layer constraints of knowledge point pairs between layers are calculated to obtain the inverse coupling matrix. The inverse coupling matrix represents the strength of the difficulty constraint between knowledge point pairs. The forward coupling matrix and the reverse coupling matrix are normalized, and the weighted combination coefficient of the migration probability and the difficulty constraint strength is calculated. Based on the weighted combination coefficient, the migration probability and the difficulty constraint strength are nonlinearly combined to construct a bidirectional coupling representation, which includes a migration probability vector, a constraint strength vector and coupling weight coefficients.
7. A deep learning-based system for constructing a model of exam question setter behavior, used to implement the method of any one of claims 1-6, characterized in that, include: The first unit is used to obtain the question-setting records of the exam question setters, construct a knowledge association analysis module for the question-setting records, calculate the usage frequency based on the knowledge point tags in the question-setting records, construct a knowledge point mapping network based on a preset knowledge system, obtain the transfer weight by calculating the conditional probability between knowledge points, combine the knowledge point mapping network and the transfer weight to generate a knowledge application graph, and output a knowledge application vector based on the knowledge application graph. The second unit is used to construct a dynamic preference analysis module based on knowledge application vectors. The dynamic preference analysis module groups the question-setting behavior into time windows according to the time information in the question-setting record, extracts features from the knowledge point selection sequence in each time window, and generates a question preference vector by combining the difficulty information. The third unit is used to train a proposition behavior recognition network based on proposition preference vectors. The proposition behavior recognition network models the knowledge point selection sequence and difficulty information respectively. By optimizing the migration path and difficulty distribution in the knowledge application graph, proposition behavior features are obtained. The fourth unit is used to input the proposition behavior characteristics into the anomaly detection module. The anomaly detection module calculates the warning threshold based on the migration path and difficulty distribution. When the proposition behavior characteristics deviate from the warning threshold, it outputs warning information and the cause of the anomaly.
8. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 6.
Citation Information
Patent Citations
Knowledge tracking method integrating learning process and difficulty features of question knowledge points
CN114781710A
Personalized dynamic question setting method and system based on large language model
CN119903160A