Teaching strategy optimization model and grammar error early warning method based on data mining

By constructing multidimensional feature vectors and dynamic knowledge graphs, combined with the improved Apriori algorithm and decision tree, a systematic correlation analysis and personalized intervention for high-frequency syntax errors in programming education were achieved. This solved the problem of inaccurate resource allocation in traditional teaching and improved teaching efficiency and response speed.

CN120911680APending Publication Date: 2025-11-07CHONGQING VOCATIONAL COLLEGE OF IND & INFORMATION TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511040223.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

In traditional programming education, the lack of systematic correlation analysis of frequently repeated syntax errors leads to the inaccurate allocation of teaching resources, the inability to promptly prevent the formation of incorrect habits, and the lack of data-driven automated optimization of teaching strategies, making it difficult to adapt to students' dynamically changing learning needs.

Method used

By using a data mining-based teaching strategy optimization model and a syntax error early warning method, student code data is acquired in real time, multi-dimensional feature vectors and dynamic knowledge graphs are constructed, causal relationships between errors are identified, and individual or group early warnings are generated using an improved Apriori algorithm and dynamic optimization formula. This enables tiered intervention and customized teaching resource delivery, and teacher-side suggestions are generated by combining strategy optimization decision trees.

Benefits of technology

It enables dynamic suggestions for resolving errors by individual and group students, improving the accuracy and responsiveness of teaching resources, enhancing error identification accuracy and personalized experience, adapting to students' dynamically changing learning needs, reducing ineffective practice, and improving teaching efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120911680A_ABST
    Figure CN120911680A_ABST
Patent Text Reader

Abstract

The invention relates to the field of wisdom education, and particularly discloses a teaching strategy optimization model and grammar error early warning method based on data mining, and the method comprises the following steps: S1, obtaining original code data in a student code library in real time, and preprocessing the original code data to obtain a structured data matrix; s2, for the structured data matrix, generating a grammar parse tree through a grammar parser, positioning error nodes and extracting context features by using a node traversal algorithm, and constructing a single-error multi-dimensional feature vector; meanwhile, an error association rule base is constructed based on historical error data, a causal relationship across error types is identified, and dynamic knowledge graph data is formed. According to the technical scheme, association analysis and deep mining can be carried out on the multi-source learning data, dynamic solution suggestions for student individual errors or group generality errors are formed, and teachers are assisted to quickly adapt to dynamically changing learning requirements of the students.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of intelligent education, and in particular to a teaching strategy optimization model based on data mining and a syntax error early warning method. BACKGROUND

[0002] In the field of programming education technology, accurate mastery of syntax rules is a basic link in programming ability training. In the teaching process of Python indentation rules, Java semicolon syntax, and other basic syntax points, there is a technical problem of high-frequency repeated errors. According to statistics, more than 80% of programming beginners will make the same type of syntax error multiple times during basic learning, such as the confusion of indentation levels in Python code due to the absence of colons, and the omission of semicolons in Java method definition.

[0003] The traditional teaching mode relies on teachers to manually analyze homework, answer question records, and other fragmented data, and lacks systematic correlation analysis and deep mining of multi-source learning data. Although the existing learning management system (LMS) can record error type statistical information, it can only realize the frequency statistics of a single error type (such as displaying the proportion of syntax errors or counting the number of syntax errors), and cannot build a causal correlation model between error types. For example, for the "for loop colon omission" error, the above technology cannot identify the relevance between it and the subsequent "indentation level error" and "variable scope confusion", and it is even more difficult to predict the high-frequency co-occurrence of error combinations (such as the co-occurrence probability of "indentation error" and "variable undefined error") through data mining technology. Due to the lack of dynamic correlation analysis capability, the positioning of group knowledge blind spots is significantly lagging, and the formation of error habits cannot be timely blocked.

[0004] The existing LMS adopts a unified rule triggering mode and does not establish a hierarchical response mechanism based on error characteristics. For individual student errors or group common errors, teachers still need to manually analyze data and adjust teaching plans. This extensive early warning mode makes it difficult to accurately allocate teaching resources, reducing error intervention efficiency.

[0005] The traditional teaching strategy is designed based on a static knowledge graph and lacks data-driven automated optimization algorithms. The adjustment of teaching content difficulty, exercise intensity, and explanation focus mainly relies on teacher experience. For example, teachers need to manually modify teaching plans after time-consuming data analysis, and students are continuously exposed to inefficient teaching content during this period. During this period (teacher data analysis period), training focus cannot be adjusted according to real-time error data, making it difficult to adapt to students' dynamically changing learning needs.

[0006] Therefore, there is an urgent need for a teaching strategy optimization model and grammar error early warning method based on data mining, which can perform correlation analysis and deep mining on multi-source learning data, form dynamic solution suggestions for individual errors or group common errors of students, and assist teachers in quickly adapting to the dynamically changing learning needs of students. SUMMARY

[0007] The application provides a teaching strategy optimization model and grammar error early warning method based on data mining, which can perform correlation analysis and deep mining on multi-source learning data, form dynamic solution suggestions for individual errors or group common errors of students, and assist teachers in quickly adapting to the dynamically changing learning needs of students.

[0008] To solve the above technical problems, the application provides the following technical solutions: The teaching strategy optimization model and grammar error early warning method based on data mining comprises the following steps: S1, real-time acquisition of original code data in a student code library, and preprocessing of the original code data to obtain a structured data matrix; S2, generating a syntax analysis tree for the structured data matrix through a syntax analyzer, positioning error nodes and extracting context features using a node traversal algorithm, and constructing a single-error multi-dimensional feature vector; at the same time, constructing an error association rule library based on historical error data, identifying the causal relationship between cross-error types, and forming a dynamic knowledge graph data; S3, inputting the single-error multi-dimensional feature vector and the dynamic knowledge graph data into a common difficulty mining model, and processing based on an improved Apriori algorithm and a dynamic optimization formula; when a single student makes the same type of error for more than a preset number of times within a preset time, it is defined as individual early warning, or when more than a preset percentage of students in a class make the same error within a preset time, it is defined as group early warning, and the associated knowledge points are extracted from the single-error multi-dimensional feature vector and the knowledge graph to generate an early warning feature vector containing error types, severity, and associated knowledge points; For high-frequency co-occurrence error combinations with support exceeding a preset value, a teaching strategy dynamic optimization formula is used to generate strategy adjustment parameters including teaching content difficulty, exercise intensity, and explanation emphasis; S4, inputting the early warning feature vector into a real-time early warning module, and implementing hierarchical intervention according to the error impact range; if it is individual early warning, customized micro-lessons are matched and pushed from a teaching resource library; if it is group early warning, a reinforcement training module containing associated knowledge point special question groups and code editing interface real-time error marking is generated to form instant feedback on the student side; S5, inputting the strategy adjustment parameters into a strategy optimization decision tree, combining the knowledge point mapping relationship in the dynamic knowledge graph, and generating teaching strategy suggestions for the teacher side.

[0009] The basic scheme principle and benefits are as follows: the present application realizes intelligent diagnosis of students' programming errors and optimization of teaching strategies through multi-module cooperation, specifically, S1 converts multi-source heterogeneous data such as original code data, compilation error logs, learning behavior records, etc. into a structured matrix, providing a unified data format for subsequent analysis; the matrix structure not only contains code syntax features (such as indentation level, variable declaration mode), but also contains time and space dimension information (such as error occurrence time, modification frequency), supporting multi-dimensional correlation analysis.

[0010] S2 converts code errors into multi-dimensional feature vectors through syntax parsing tree and node traversal algorithm, realizing quantitative representation of errors; a dynamic knowledge graph is constructed, not only storing single error features, but also mining the causal relationship between errors (such as "variable undefined" often leading to "type error") through historical data, forming a knowledge network.

[0011] S3 uses an improved Apriori algorithm to identify high-frequency co-occurrence error combinations and capture group common difficulties; through a pre-set threshold, individual early warning and group early warning are distinguished, realizing differentiated response to "a small number of students repeatedly making the same error" and "a large number of students simultaneously making a certain type of error".

[0012] S4 implements hierarchical intervention based on early warning feature vectors, converting abstract associations in the knowledge graph into executable teaching resources (such as customized micro-lessons, special topic groups); real-time error labeling function through the linkage of code editing interface and knowledge graph, abstract knowledge is directly mapped to the current code of students, improving the timeliness of feedback.

[0013] S5 generates teaching strategy suggestions through decision trees combined with dynamic knowledge graph, converting student error data into teacher-operable adjustment parameters (such as exercise intensity, explanation emphasis); the generation of strategy adjustment parameters is based on dynamic optimization formula, ensuring that the suggestions can adapt to the real-time changes of students' knowledge state.

[0014] Due to the real-time data acquisition and preprocessing mechanism of S1, the response time of the system to new errors is shortened; the incremental updating mechanism of the dynamic knowledge graph (S2) can quickly adapt to new error patterns (such as new syntax errors caused by updating of programming language version), and the knowledge updating delay is reduced.

[0015] The combination of the multi-dimensional feature vector and the knowledge graph improves the recognition accuracy of the application to complex errors (such as cross-file reference errors); the improved Apriori algorithm effectively filters incidental errors through space-time constraints (S3), and improves the discovery accuracy of high-frequency co-occurrence error combinations. The individual early warning mechanism (S3) can identify special error patterns (such as logic errors caused by specific programming habits) that only affect individual students, and the matching degree of the customized micro-lessons pushed to the students is improved; the hierarchical intervention strategy (S4) enables students of different learning levels to receive differentiated support, such as providing more example codes for students with weak foundation and algorithm optimization suggestions for high-level students.

[0016] The dynamic adjustment of the group early warning threshold (S3) realizes automatic adaptation to the changes in course difficulty, such as lowering the group early warning trigger threshold in difficult chapters such as "recursive functions"; the mining of high-frequency co-occurrence error combinations (S3) helps teachers discover hidden teaching blind spots, for example, most students simultaneously make "loop nesting" and "variable scope" errors, prompting the need to strengthen the explanation of basic concepts. The strategy optimization decision tree (S5) improves the relevance of teaching resource recommendations through knowledge point mapping relationships, reducing the amount of invalid practice; the dynamic optimization formula (S3) automatically matches the practice intensity with the current level of the students, shortening the average practice time of the students while maintaining the learning effect. The association of strategy suggestions with the knowledge graph (S5) enables teachers to trace the basis of each suggestion, improving the decision-making confidence; The severity index in the early warning feature vector (S3) (combined with error type, frequency, and associated knowledge points) provides quantitative reference for teachers, improving the efficiency of intervention priority judgment. The real-time updating of the dynamic knowledge graph (S2) can learn new error patterns and associated relationships, continuously optimizing the early warning accuracy in continuous use; the closed-loop feedback mechanism of teaching strategy suggestions (S5) can adjust the decision tree parameters according to the actual teaching effect, improving the effectiveness of the strategy.

[0017] The application realizes fast data acquisition (S1) and deep feature mining (S2) through multi-module cooperation, enabling the application to quickly respond to new errors and accurately judge error types and severity; individual early warning (S3) and hierarchical intervention (S4) ensure that each student receives customized support, while the group early warning mechanism prevents teachers from being overwhelmed by individual problems, achieving optimal allocation of teaching resources; the dynamic knowledge graph (S2) provides prior knowledge for data mining (S3), and the data mining results in turn update the knowledge graph, forming a virtuous cycle.

[0018] In summary, the application realizes the correlation analysis and deep mining of multi-source learning data, forms dynamic solutions for individual student errors or group common errors, and assists teachers in quickly adapting to the dynamic changes in students' learning needs, achieving technical breakthroughs in response speed, recognition accuracy, and personalized experience in multiple dimensions.

[0019] Further, the preprocessing in S1 includes performing lexical analysis on the original code data to extract a code token sequence, identifying natural language descriptions in code comments, and establishing a mapping relationship between code snippets and their functional descriptions through a pre-trained code-text alignment model; synchronously collecting error stack information in the compilation error log, extracting error types, error line numbers, and compiler prompt texts, and constructing error feature triplets <error type, error location, error description>; extracting student code editing behavior data, including code modification time intervals, cursor dwell position distribution, and auto-completion request frequency, and quantifying them into a time-series behavior feature vector; The original code data is segmented and aggregated using a sliding time window. For each set of code modifications within a time window, the code change density is calculated, defined as the proportion of modified code lines to the total number of lines in a unit of time; a code dependency graph is constructed to identify the reference relationships between variables, functions, and classes, and the modification propagation impact value of each code element is calculated.

[0020] Further, the S1 further includes: counting the recurrence frequency of the same type of error within the time window, combining the preset error severity to generate an error weight sequence; performing tensor product operation on the extracted code features, error features, and behavior features to generate a three-dimensional structured data matrix, the first dimension being the error coordinate space, including file path, code line number, and code block level; the second dimension being the error type space, including syntax error, logic error, and semantic error classification labels; the third dimension being the knowledge point mapping space, establishing a many-to-many mapping relationship between errors and knowledge points through a pre-defined knowledge point-error type association rule; The matrix element value is the product of the error frequency and the error weight, representing the importance score of the error under a specific knowledge point; the generated structured data matrix is updated continuously using an incremental learning strategy, when a new error pattern is detected, a new error type label is automatically generated through clustering analysis, and the knowledge point mapping rule is updated; the support and confidence of the error pattern are calculated regularly, and the outdated error patterns with support lower than the preset threshold are decayed; a time decay mechanism for matrix elements is established to make the influence of historical error data decrease exponentially over time, maintaining the sensitivity of the matrix to the latest error patterns.

[0021] Further, the S2 further includes: for each node of the syntax parse tree, extracting its structural features in the abstract syntax tree, including node depth, parent node type, child node number, and sibling node relationship, and constructing a local structure feature vector; extracting the code in a preset number of lines before and after the error node, converting it into a semantic context vector through a word vector model to capture the local semantic information of the code snippet; combining the code change history data to calculate the edit distance and modification type of the current error node relative to the last correct submission version, generating an evolution feature vector to represent the evolution process of the error; Based on the dynamic time window aggregation of error data, the co-occurrence frequency of different error types of the same student in adjacent time windows is calculated, an error timing transition table is constructed to record the evolution probability between error types; For the error data of all students in the class, the propagation centrality of each error type is calculated, and the propagation source node of high-frequency errors is identified, that is, a key node that may lead to subsequent chain errors; Through the pre-trained code semantic model, the semantic similarity between different error code fragments is calculated, and a semantic association network between error types is established; When a new error type or association relationship is detected, the nodes and edges of the knowledge graph are dynamically expanded; when the newly established error association rule conflicts with the existing rule, the support, confidence and time sequence are compared to select the optimal rule; Set the timeliness factor of each edge in the knowledge graph, automatically decay the weight of low-frequency association edges over time, trigger the subgraph reconstruction mechanism when a certain type of error has not appeared for a long time; Map the single-error multi-dimensional feature vector to the node in the knowledge graph; Convert the error association rule into an edge in the knowledge graph; Through the attention mechanism to fuse node features and edge relationships, the representation of each error node in the knowledge graph contains its own features and association information, forming the final dynamic knowledge graph data.

[0022] Further, the step of generating strategy adjustment parameters based on the improved Apriori algorithm in S3 includes: In the candidate set generation process of the Apriori algorithm, time window constraints and spatial constraints are introduced, and only error combinations that frequently co-occur in a specific spatiotemporal range are retained; For each frequent error combination, calculate its spatiotemporal cohesion degree, that is, the proximity of errors in time and the distance in code space, and form a spatiotemporal association rule library; Based on the constructed dynamic knowledge graph, for each high-frequency co-occurrence error combination, identify its shortest propagation path in the knowledge graph; Assign an error weight to each error node on the path , set the value range to 0-1 according to the impact on program running; and count the error frequency , that is, the number of error occurrences per unit time; Combined with the knowledge point mastery rate , that is, the mastery percentage calculated through test data, to provide parameters for subsequent formula calculation; According to the course progress, dynamically adjust the support threshold value, set the threshold value to the first preset key value in the basic syntax stage, and set the second preset key value in the advanced programming stage, and link with the adjustment coefficient in the teaching strategy dynamic optimization formula, when the threshold value decreases Increase the value to enhance the sensitivity of strategy adjustments; Through formula Strengthen strategic intervention; The process of generating parameters for adjusting teaching strategies is transformed into a bi-objective optimization based on the aforementioned formula, so as to... Maximizing this is the primary goal, achieved by increasing the practice intensity of high-weight, error-related knowledge points; and by focusing on each knowledge point... Minimizing variance is the second objective, as stated in the formula. The differentiated assignment balances the focus of the explanation; the generated strategy adjustment parameters are directly reflected in the formula calculation results, including the difficulty coefficient of the teaching content, the proportion of practice intensity, and the allocation of explanation time. Values ​​are dynamically tilted.

[0023] Furthermore, S3 also includes: For single-error multidimensional feature vectors and dynamic knowledge graph data, a multi-head attention mechanism is used to calculate the association weights of each error feature with knowledge points, and the knowledge points with the highest weights are selected as the ones used in the formula. The basis for the calculation; Automatically extract the error weights corresponding to these knowledge points Error frequency and mastery rate This forms the formula input parameter set. The severity of errors is assessed from three dimensions: impact on execution, complexity of remediation, and potential for propagation. The weighted sum of these dimensions is then mapped to an error weight. Adjustment coefficients are set based on severity levels. This enables differentiated responses to strategy adjustments; Based on the student's knowledge mastery state vector Through formula Calculate the exercise intensity adjustment value.

[0024] Furthermore, the implementation of tiered intervention in S4 includes: Construct a three-dimensional impact range assessment system based on knowledge point mastery rate. The study assesses the chain reaction of errors on related knowledge systems; analyzes the number of program crashes, debugging time, and code modifications caused by errors; counts the frequency of discussions about the error and the number of times students seek help from teachers within the class; and inputs the three-dimensional assessment results into a fuzzy decision tree to generate a five-level impact range classification. For individual early warnings, a reinforcement learning-based strategy is used to dynamically adjust the micro-lesson delivery, thereby increasing the student's knowledge acquisition vector. , error feature vector; micro-lesson type, push timing, presentation form; with error recurrence rate reduction, micro-lesson completion rate, subsequent test score improvement as indicators; through Q-learning algorithm iteration optimization of push strategy, so that micro-lesson click rate is improved; In the real-time error marking of the code editing interface, in addition to highlighting the error line, the following is included: Based on dynamic knowledge graph, mark other potential risk points associated with the current error; Show the shortest repair path from the current error to the normal operation of the program, including 3-5 progressive repair steps; Link to the knowledge point explanation micro-lesson directly related to the error in the teaching resource library.

[0025] Further, the content of generating a reinforcement training module in S4 includes: Construct a special topic group generation model based on GAN, generate a three-dimensional topic group structure including basic questions, advanced questions and comprehensive questions according to the associated knowledge points in the early warning feature vector, and each type of question is accompanied by an expected error mode; Based on historical student answer data, judge the difficulty adaptability, knowledge coverage and error inducibility of the generated topic group; Update the knowledge state vector in real time during the student completes the exercise , correct answer through the Bayes update formula to improve the mastery rate of the corresponding knowledge point ; Error analysis of the matching degree between actual error and expected error mode, if matched, reinforce the practice of the knowledge point, if not matched, trigger the associated analysis of the knowledge graph; When the correct rate of 3 consecutive questions exceeds the preset value, automatically reduce the difficulty of the knowledge point practice.

[0026] Further, the content of generating a teacher-end teaching strategy suggestion in S5 includes: The internal node splitting of the strategy optimization decision tree is based on error frequency , error weight and knowledge point mastery rate Realize, prefer to choose the feature with the highest information gain as the splitting basis, where the information gain is calculated by And The association degree; The strategy suggestion of the split child node needs to meet the diversity constraint to avoid outputting homogeneous strategies; The interpretability of the splitting rule is verified by Calculation link, ensuring that each strategy suggestion can be traced back to the specific , and Value; By combining the constructed dynamic knowledge graph, a deep association between strategy suggestions and knowledge points is achieved. The error propagation path in the knowledge graph is transformed into the reasoning rules of the decision tree. When an error is detected at the starting point of the path, preventive strategies for subsequent errors are recommended in advance. The node association weights of the knowledge graph are used to prioritize the strategies generated by the decision tree. The higher the association weight, the higher the ranking of the corresponding strategy suggestion. Through the hierarchical relationship of knowledge points in the knowledge graph, it is ensured that the strategy suggestions cover knowledge points at both the upper and lower levels. Based on the output parameters of the teaching strategy optimization formula, a strategy effectiveness evaluation system is constructed. and The change in the difference is used to assess the error frequency after the strategy is implemented. The extent of the decrease; tracking the mastery rate of knowledge points. The upward curve, when for two consecutive weeks When the growth is lower than the preset value, the decision tree parameters are updated; the evaluation results are fed back to the dynamic knowledge graph, and the weight of erroneous associations is adjusted through an incremental update mechanism so that subsequent strategy suggestions are more in line with the actual teaching effect. Attached Figure Description

[0027] Figure 1 This is a flowchart illustrating an implementation example of a data mining-based teaching strategy optimization model and a grammar error early warning method. Detailed Implementation

[0028] The following detailed description illustrates the specific implementation method: Teaching strategy optimization model and grammar error early warning method based on data mining (e.g.) Figure 1 (As shown), including the following steps: S1: Real-time acquisition of raw code data from the student code library, and preprocessing of the raw code data to obtain a structured data matrix; S2, for the structured data matrix, a syntax parser is used to generate a syntax parse tree, a node traversal algorithm is used to locate error nodes and extract context features, and a single error multidimensional feature vector is constructed; at the same time, an error association rule base is constructed based on historical error data to identify causal relationships across error types and form dynamic knowledge graph data; S3, input the single-error multidimensional feature vector and dynamic knowledge graph data into the common difficulty mining model, and process it based on the improved Apriori algorithm and dynamic optimization formula; when a single student makes the same error a preset number of times or more within a preset time, it is defined as an individual warning, or when a preset percentage or more of the students in the class make the same error within a preset time, it is defined as a group warning. Extract related knowledge points from the single-error multidimensional feature vector and knowledge graph to generate a warning feature vector containing error type, severity, and related knowledge points; For high-frequency co-occurrence error combinations with support exceeding the preset value, the formula is dynamically optimized through teaching strategies to generate strategy adjustment parameters including teaching content difficulty, practice intensity, and explanation focus. S4, input the early warning feature vector into the real-time early warning module, implement hierarchical intervention according to error impact range, if it is individual early warning, match and push customized micro-lessons from the teaching resource library; if it is group early warning, generate a reinforcement training module containing related knowledge point special topic group and code editing interface real-time error marking, forming student end instant feedback; S5, input the strategy adjustment parameters into the strategy optimization decision tree, generate teaching strategy suggestions for the teacher end combined with the knowledge point mapping relationship in the dynamic knowledge graph.

[0029] In specific use, taking the Web front-end development course and 25 students learning JavaScript loop statement unit as an example. In S1 preprocessing, real-time capture of.js files and browser console error logs submitted by students through online programming platform, filter redundant spaces and comments, extract for / while / do-while keyword sequences, standardize error types (such as infinite loop, loop variable not updated), occurrence line number, related knowledge points (such as loop termination condition), submission time, student ID, etc. Construct a 80x5 structured data matrix (80 error records, 5 feature dimensions).

[0030] In S2 feature vector and knowledge graph, for the loop variable not updated error, construct a single error multi-dimensional feature vector: [logical error, line number 8, loop level 1, error frequency 2, involved student ID]; based on the error data of the previous 3 courses, construct a dynamic knowledge graph: containing loop variable not updated→infinite loop (weight 0.7), infinite loop→page crash (weight 0.9) and other causal edges, node attributes cover error repair difficulty (1-5 stars).

[0031] In S3 early warning and parameter generation, the improved Apriori algorithm identifies high-frequency combinations (co-occurrence times 28) of loop variable not updated + conditional judgment error under the condition of support ≥ 50%; student ID appears 4 times of loop variable not updated error within 18 hours (triggering individual early warning, preset threshold is 3 times); 40% of students in the class (10 people) appear the same error (triggering group early warning, preset percentage is 30%), generate early warning feature vector: [loop variable not updated, severity medium, related knowledge point = loop control flow]; Through the dynamic optimization formula of teaching strategy, it is calculated that the teaching content difficulty coefficient remains 1.0 (current basic difficulty), the practice intensity is improved by 50%, and the explanation focus is allocated as 60% for loop variable update rule and 40% for conditional expression.

[0032] In S4 layered intervention, individual early warning is to push the customized micro-lesson (including animation demonstration i++ execution process, 2 minutes and 30 seconds long) to the student ID, with 3 targeted exercises; Group early warning is to generate a special question group (8 questions in total) containing loop variable tracking table filling, error loop code repair, etc. Real-time highlight error line in code editing interface, note that the variable is not updated, which may cause infinite loop.

[0033] In S5 teaching strategy suggestion, the strategy optimization decision tree outputs the following suggestions for teachers: add a 10-minute classroom demonstration of loop variable debugging skills (combined with breakpoint debugging tools), assign 2 comprehensive application questions on loop nesting and variable updating, and prioritize the explanation of the differences between for loop and while loop in variable management.

[0034] Specifically, the preprocessing in S1 includes lexical analysis of the original code data, extraction of code token sequences, and identification of natural language descriptions in code comments. The mapping relationship between code fragments and their function descriptions is established through a pre-trained code-text alignment model. Simultaneously, error stack information in the compilation error log is collected, including error type, error line number, and compiler prompt text, to construct error feature triplets <error type, error location, error description>. Student code editing behavior data is extracted, including code modification time interval, cursor dwell position distribution, and automatic completion request frequency, which are quantified as time series behavior feature vectors. The original code data is segmented and aggregated using a sliding time window. For each code modification set within a time window, the code change density is calculated, defined as the ratio of modified code lines to total lines per unit time. A code dependency graph is constructed to identify the reference relationships between variables, functions, and classes, and the modification propagation impact value of each code element is calculated.

[0035] In actual use, in the S1 preprocessing step, the code fragment for(vari=0;i<5;){console.log(i)} is extracted as the token sequence: [for, var, i, =, 0, ;, i <, 5, ;, ), {, console.log, (, i, ),}], and the missing i++ token anomaly is identified. The pre-trained CodeBERT model is used to calculate the functional matching degree between the annotation implementation 5 times loop output and the code fragment, which is 0.85, and is stored in the annotation-function mapping table. Error feature triplets: <RangeError, line number 8, "Maximum call stack size exceeded"> (caused by infinite loop) are extracted from the browser log; the student editing behavior vector is [25 seconds (modification interval), line number 8 (cursor dwell position), 3 times (automatic completion request)], with the cursor dwell time in the for loop brackets accounting for 60%.

[0036] With a 6-hour sliding window (step 3 hours), the code modification set within the window is calculated, and the code change density = 0.35 (modify 7 lines / total code 20 lines); the variable is referenced in the loop body, conditional judgment, output statement, and the modification propagation impact value = 0.6 (modify i will affect 3 codes). The matrix dimension is expanded to 80x10, and new features such as comment-function matching degree, behavior vector similarity, and propagation impact value are added to provide multi-modal data support for subsequent analysis.

[0037] The preprocessed data reduces noise and improves the recognition of error features, providing high signal-to-noise ratio input for feature extraction in S2; through behavior vector analysis, it is found that the matching rate of high-frequency cursor residence position and error occurrence position is high, which can predict potential error points in advance; automated preprocessing liberates teachers from manual error data sorting and improves work efficiency.

[0038] Perform lexical analysis on the original code data, extract code token sequences, and identify natural language descriptions in code comments. Through a pre-trained code-text alignment model, establish a mapping relationship between code fragments and their function descriptions; simultaneously collect error stack information from compilation error logs, extract error types, error line numbers, and compiler prompt text, and construct error feature triplets <error type, error location, error description>; extract student code editing behavior data, including code modification time interval, cursor residence position distribution, and automatic completion request frequency, and quantify them as time-series behavior feature vectors; Segment and aggregate the original code data using a sliding time window. For each code modification set within a time window, calculate the code change density, defined as the ratio of modified code lines to total lines per unit time; construct a code dependency graph to identify the reference relationship between variables, functions, and classes, and calculate the modification propagation impact value of each code element.

[0039] In specific use, for example, for data index error (such as KeyErroriloc / loc mixed use), an 8-hour sliding window (step 4 hours) is used, which contains 120 code modification records within the window; in data cleaning tasks, the average code change density = 0.5 (modify 10 lines / total code 20 lines), which is significantly higher than that of data reading tasks (0.2); variable dependency analysis: DataFrame object df is referenced by 5 functions, modifying the column name of df causes 3 functions to report errors, and the propagation impact value = 0.8.

[0040] Sort the wrong time stamps in the window [t08:15, t09:30, t10:05], calculate the mean of the time interval of adjacent errors = 45 minutes, standard deviation = 15 minutes, and incorporate the time correlation feature into the structured matrix; find that iloc index errors often occur in the first hour of student submission (accounting for 65%), and supplement this rule to the matrix as a time-sensitive feature.

[0041] Through window aggregation, the accuracy of error time correlation analysis is improved, such as identifying that the index error occurs frequently between 9-10 am, and online tutoring can be arranged during this period; errors with a propagation impact value ≥0.7 are marked as high-risk modifications, which can be focused on by teachers, and can reduce the incidence of chain reactions of such errors.

[0042] The S2 further comprises: for each node of the syntax parsing tree, extracting its structural features in the abstract syntax tree, including node depth, parent node type, child node number, and sibling node relationship, to construct a local structure feature vector; extracting the code of a preset number of lines before and after the error node, and converting it into a semantic context vector through a word vector model to capture the local semantic information of the code fragment; combining the code change history data to calculate the edit distance and modification type of the current error node relative to the last correct submission version, and generating an evolution feature vector to represent the evolution process of the error; Based on the error data aggregated by the dynamic time window, the co-occurrence frequency of different error types of the same student in adjacent time windows is calculated, an error time transfer table is constructed, and the evolution probability between error types is recorded; For the error data of all students in a class, the propagation centrality of each error type is calculated to identify the propagation source node of high-frequency errors, i.e., a key node that can cause subsequent chain errors; Through a pre-trained code semantic model, the semantic similarity between different error code fragments is calculated to establish a semantic association network between error types; When a new error type or association relationship is detected, the nodes and edges of the knowledge graph are dynamically expanded; when the newly established error association rule conflicts with the existing rule, the optimal rule is selected by comparing the support, confidence, and time sequence; Set the timeliness factor of each edge in the knowledge graph, automatically decay the weight of low-frequency association edges over time, trigger the subgraph reconstruction mechanism when a certain type of error has not appeared for a long time; map the single-error multi-dimensional feature vector to the node in the knowledge graph; convert the error association rule into an edge in the knowledge graph; Through the attention mechanism, the node features and edge relationships are fused, so that the representation of each error node in the knowledge graph contains its own features and association information, forming the final dynamic knowledge graph data.

[0043] In specific use, the local structural feature vector is, for example, for a dependency injection failure error, extracting [node depth 4, parent node type BeanDefinition, number of child nodes 3, brother node 2] (based on abstract syntax tree analysis); the semantic context vector is, for example, by converting 8 lines of code before and after the error node (containing the @Autowired annotation) through the Word2Vec model, capturing the semantic association of component scanning and dependency injection (cosine similarity 0.82); The evolution feature vector is compared with the last correct submission version, the edit distance = 4 (add @Service annotation, modify 2 package paths, delete redundant import), and the modification type = annotation configuration error (code 110).

[0044] The co-occurrence frequency of dependency injection failure -> Bean creation exception in the adjacent window is calculated to be 0.8, recorded in the error timing transition table, and the evolution probability is 0.75; the propagation centrality value of the component scanning path error is calculated to be 0.85 (affecting 4 types of downstream errors) through the PageRank algorithm, and it is determined as a propagation source node; the semantic similarity between Spring dependency injection error and Dagger dependency injection error is calculated to be 0.76 through CodeBERT, and a cross-framework association edge is established; the @SpringBootApplication scanning range change error node specific to SpringBoot3.0 is added, and the edge weight with the traditional @ComponentScan is 0.7; the timeliness factor of the edge is set to 0.9 / week, and the @ConfigurationProperties binding error specific to Spring2.x version has not appeared for 3 weeks, and the weight decays from 0.6 to 0.2, triggering subgraph reconstruction; through the attention mechanism, the node features (such as error type) and edge relationships (such as causal weight) are fused, so that the representation of the dependency injection failure node contains both its own features and component scanning association information.

[0045] Dynamic updating delays the identification of new error types by <12 hours, improving adaptability to framework version iterations; after fusing node and edge features, the accuracy of error association is improved, especially for cross-version and cross-framework semantic similar error identification; the subgraph reconstruction mechanism reduces the storage space occupancy of the knowledge graph while maintaining most of the key association information; based on the knowledge graph, the error propagation path prediction accuracy is high, and the chain error can be warned 2-3 steps in advance.

[0046] The steps of S3 for high-frequency co-occurrence error combination mining and strategy adjustment parameter generation based on the improved Apriori algorithm include: In the process of generating candidate sets for the Apriori algorithm, time window constraints and spatial constraints are introduced to retain only error combinations that frequently co-occur within a specific spatiotemporal range. For each frequent error combination, its spatiotemporal cohesion is calculated, which is the temporal proximity of the error and its distance in the code space, forming a spatiotemporal association rule base. Based on the constructed dynamic knowledge graph, for each frequently co-occurring error combination, its shortest propagation path in the knowledge graph is identified; Assign an error weight to each error node on the path. The value is set according to its impact on program execution, ranging from 0 to 1; and the error frequency is statistically analyzed. This refers to the number of errors occurring per unit of time; combined with the knowledge mastery rate. That is, the percentage of mastery calculated from test data, which provides parameters for subsequent formula calculations; The support threshold is dynamically adjusted according to the course progress. The threshold is set as the first preset key value for the basic grammar stage and as the second preset key value for the advanced programming stage, and is also adjusted in accordance with the adjustment coefficient in the dynamic optimization formula of the teaching strategy. Linkage, threshold decrease Increase the value to enhance the sensitivity of strategy adjustments; Through formula Strengthen strategic intervention; The process of generating parameters for adjusting teaching strategies is transformed into a bi-objective optimization based on the aforementioned formula, so as to... Maximizing this is the primary goal, achieved by increasing the practice intensity of high-weight, error-related knowledge points; and by focusing on each knowledge point... Minimizing variance is the second objective, achieved through the formula... The differentiated assignment balances the focus of the explanation; the generated strategy adjustment parameters are directly reflected in the formula calculation results, including the difficulty coefficient of the teaching content, the proportion of practice intensity, and the allocation of explanation time. Values ​​are dynamically tilted.

[0047] In practical use: a 48-hour time window (covering weekend practice periods) and function definition-call code block space constraints are introduced, retaining only the co-occurring error combinations within this range; combinations with 55% support (parameter type mismatch + missing return value) are identified, with a spatiotemporal cohesion of 0.8 (error occurrence interval < 2 hours, code location distance < 10 lines), forming a spatiotemporal association rule base; the shortest path is identified in the knowledge graph: parameter type mismatch → compilation failure (probability 0.9) → function call interruption (probability 0.8).

[0048] The formula parameters and strategy are generated as follows, where the error weight is represented as: parameter type mismatch. =0.7 (affects compilation), return value missing = 0.6 (influence logic); error frequency is represented as = 6 times / day (average daily occurrence in class); knowledge point mastery rate is represented as = 0.45 (unit test score); The basic syntax stage is set to the first preset key value of 50%, which is linked with the adjustment coefficient α (threshold value 50%→α=0.7); The formula is calculated as = + 0.7×[(0.7×6+0.6×4) / 0.45]=S_use+0.7×(6.6 / 0.45)= + 10.27, that is, the strategy adjustment parameter is the exercise intensity increase of 102.7% (rounded to 100%); Double-target optimization, minimizing error frequency is to increase parameter type conversion exercise intensity (60%); maximizing the balance of mastery is to allocate 40% of the exercise to the return value definition, balancing knowledge point coverage; The generated explanation duration allocation is the parameter passing rule of 15 minutes, and the return value type is 10 minutes.

[0049] The spatiotemporal constraint reduces the false positive rate of high-frequency error combinations, avoiding the identification of occasional co-occurring error pairs; the double-target optimization generated parameters reduce error frequency while reducing knowledge point mastery variance, avoiding teaching imbalance of favoring one over the other; through formula quantitative calculation, the pertinence of strategy adjustment is improved, such as automatically doubling the exercise intensity for knowledge points with a mastery rate of less than 50%; teachers do not need to manually calculate the exercise allocation proportion, and the strategy generation time is shortened from 2 hours to 5 minutes, with a 24-fold improvement in decision-making efficiency.

[0050] The S3 also includes: for single-error multi-dimensional feature vectors and dynamic knowledge graph data, the correlation weight of each error feature to the knowledge point is calculated through a multi-head attention mechanism, and the knowledge points with the top-ranked weights are selected as the basis for calculating in the formula; The error weights , error frequencies and mastery rates of these knowledge points are automatically extracted to form the formula input parameter set ; the error severity is evaluated from three dimensions of execution influence, repair complexity, and propagation potential, and the weighted sum result is mapped as the error weight ; the formula adjustment coefficient is set based on the severity classification to realize differentiated response of strategy adjustment; According to the student knowledge mastery state vector , the exercise intensity adjustment value is calculated through the formula .

[0051] Specifically, the multi-head attention mechanism is used to calculate the correlation weight of the feature vector of the character collision detection failure error, including the coordinate range event trigger motion instruction, etc. The correlation weight is calculated by 3 heads of attention mechanism: coordinate system weight = 0.92, collision event weight = 0.85, and motion step weight = 0.6; Top2 knowledge points are selected as the basis for formula calculation (weight ≥ 0.8).

[0052] extracting a parameter set, = 0.7 (coordinate system error weight), = 5 (average number of occurrences per day), = 0.3 (test score); The severity assessment is that the execution impact is 7 points (the game logic is affected but not crashed); the repair complexity is 4 points (2 parameters need to be adjusted); the propagation potential is 6 points (may cause score calculation error); Weighted sum = 7*0.4 + 4*0.3 + 6*0.3 = 5.8 points → mapping = 0.58 (rounded to 0.6); Severity 5.8 points → adjustment coefficient a = 0.65, substitute into the formula to get exercise intensity adjustment value = +65%.

[0053] The multi-head attention mechanism improves the accuracy of positioning core knowledge points and avoids including secondary knowledge points in strategy calculation; the accuracy of multi-dimensional severity assessment is improved compared to single-dimensional assessment (such as only looking at execution impact); the matching degree of exercise intensity generated based on the parameter set with the actual needs of students is improved, and the amount of invalid exercises is reduced; targeted parameter adjustment improves students' mastery of knowledge points and learning efficiency.

[0054] The content of implementing hierarchical intervention in S4 includes: Building a three-dimensional impact range assessment system based on knowledge point mastery rate , assessing the chain impact of errors on related knowledge systems; analyzing the number of program crashes caused by errors, debugging time, and code modification volume; counting the frequency of discussing the error within the class and the number of times seeking help from teachers; inputting the three-dimensional assessment results into a fuzzy decision tree to generate a five-level impact range classification; Early warning for individuals, using reinforcement learning to dynamically adjust micro-lesson push strategies, student knowledge mastery vector , error feature vector; micro-lesson type, push timing, and presentation form; using error recurrence rate reduction, micro-lesson completion rate, and subsequent test score improvement as indicators; iteratively optimizing the push strategy through Q-learning algorithm to improve micro-lesson click rate; In real-time error annotation in the code editing interface, in addition to highlighting the error line, it also includes: Based on dynamic knowledge graph, other potential risk points associated with the current error are annotated; Show the shortest repair path from the current error to the normal operation of the program, including 3-5 progressive repair steps; Link to the knowledge point explanation micro course in the teaching resource library directly related to the error.

[0055] In specific use, take the learning function definition and call unit as an example to further illustrate the high-frequency problems such as parameter passing error and return value missing. First, for the error of mixing positional parameters and keyword parameters, calculate the mastery rate of the associated knowledge points through unit test data: parameter passing rule mastery rate 42%, function signature mastery rate 38%. Use the analytic hierarchy process to calculate the chain influence degree: 0.42×0.6+0.38×0.4=0.404 (0-1 interval, the higher the value, the greater the impact).

[0056] Statistical program crash frequency (average 1.2 times / student), debugging time (average 4.5 minutes), code modification amount (average 3 lines) caused by this error, and through normalization processing (0-10 points) to get the score: crash frequency 6 points, debugging time 7 points, modification amount 5 points, weighted sum (weights are 0.5, 0.3, 0.2) 5.9 points.

[0057] Monitor the online discussion area to find that the daily discussion frequency of this error is 8 times, and the number of students asking for help from teachers is 6 times / day, which is converted into social influence score 6.5 points (0-10 points).

[0058] Input the three-dimensional evaluation results (cognition 0.404, operation 5.9, social 6.5) into the fuzzy decision tree, and after 3 layers of node splitting, the output classification result is three-level impact (medium risk), triggering medium-intensity intervention.

[0059] The state space includes student knowledge mastery vector (such as [0.42, 0.38] corresponding to two associated knowledge points) and error feature vector ([parameter type error, occurrence frequency 3, severity medium]); The action space includes 3 types of micro course (animation demonstration, code comparison, step-by-step debugging), 3 push times (immediately after the error occurs, 5 minutes later, next login), and 2 presentation forms (video + text, interactive exercises).

[0060] Take the error recurrence rate reduction (weight 0.4), micro course completion rate (weight 0.3), and subsequent test score improvement (weight 0.3) as the reward function. For example, student A first appears an error and pushes a video + text micro course (immediately), the reward value is 0.6; the second time is adjusted to an interactive exercise (pushed 5 minutes later), the reward value is 0.82, and the action value matrix is updated through the Q-learning algorithm. After 5 iterations, the optimal strategy is determined. After implementation, the individual micro course click rate is improved to, and the error recurrence rate is reduced.

[0061] When the student calls calc(b=2,3) after writing defcalc(a,b):returna+b, the interface highlights the error line and prompts: keyword argument appears after positional argument, which may cause subsequent 'default argument value' parsing error (associated probability 0.72) ''.

[0062] Provide a 3-step repair path: 1. Adjust the position parameter 3 to the keyword parameter a=3→2. Check the parameter order in the function definition→3. Execute calc(a=3,b=2) to verify the result, the path length is calculated based on the shortest path algorithm (Dijkstra algorithm) of the knowledge graph.

[0063] The annotation area is accompanied by parameter passing rules micro-lesson entrance, linked to the 2-minute-30-second video in the teaching resource library, which is filtered by the cosine similarity between the error feature vector and the resource label (similarity 0.89).

[0064] Three-dimensional evaluation improves the accuracy of intervention; reinforcement learning pushes improve individual learning efficiency; real-time annotation shortens the time-consuming of error repair, and improves the students' ability to solve problems independently.

[0065] The content of the reinforcement training module generated in S4 includes: Build a GAN-based special topic group generation model, generate a three-dimensional topic group structure containing basic questions, advanced questions, and comprehensive questions according to the associated knowledge points in the early warning feature vector, and each type of question is accompanied by an expected error mode; Based on historical student answer data, judge the difficulty adaptability, knowledge coverage and error inducibility of the generated topic group; Update the knowledge state vector in real time during the student's exercise process , correct answer, improve the mastery rate of the corresponding knowledge point through the Bayes update formula ; Error analysis when answering, if the matching degree of the actual error and the expected error mode is matched, the knowledge point practice is strengthened, if not matched, the associated analysis of the knowledge graph is triggered; When the correct rate of 3 consecutive questions exceeds the preset value, automatically reduce the difficulty of the knowledge point practice.

[0066] In specific use, take the Java course exception handling unit learning as an example, some students have problems such as improper use of try-catch block and incorrect order of exception type capture, and need to generate special training topic groups for intensive practice.

[0067] According to the associated knowledge point abnormal inheritance system try-with-resources syntax in the early warning feature vector, generate a three-dimensional topic group structure.

[0068] The basic questions (3) are expected to complete the try-catch block to capture NullPointerException, and the expected error pattern is to miss the catch block; the advanced questions (2) are to repair the capture order of IOException and FileNotFoundException, and the expected error pattern is that the parent class exception is captured before the child class; the comprehensive question (1) is to optimize the file operation code using try-with-resources and handle exceptions, and the expected error pattern is that the resource is not released, causing memory leakage.

[0069] With the answer data of the past 3 students (including error pattern distribution and score rate) as the training set, the discriminator judges the generated question group by the binary classification loss function, and the score rate in the 40%-60% interval (medium difficulty) is determined as valid; covering more than 80% of the core knowledge points of the exception handling unit; the probability of triggering the expected error pattern is ≥70%. After 50 rounds of adversarial training, the discrimination accuracy of the generated question group is improved.

[0070] After the students correctly complete the advanced question of exception capture order, the knowledge point mastery rate is adjusted through the Bayes update formula, assuming that the prior probability P(mastery)=0.4 and the discrimination of the question is 0.7, the updated P(mastery)=(0.4×0.7) / (0.4×0.7+0.6×0.3)=0.59, and the corresponding dimension value in the knowledge state vector is increased from 0.4 to 0.59.

[0071] If the student makes an error in the comprehensive question of not using try-with-resources, and the matching degree with the expected error pattern of not releasing resources is 85% (calculated by semantic similarity), the knowledge point practice is strengthened: 1 more variant question combining resource closing and exception handling is added. If the error pattern is an unexpected custom exception that is not serialized, trigger the knowledge graph association analysis, find the hidden association with the object serialization knowledge point (association weight 0.65), and supplement related basic questions.

[0072] When the student's correct rate of 3 basic questions in a row is ≥90%, the difficulty of the exercise is automatically increased from medium to difficult, such as from single exception handling to multiple exception nested handling; on the contrary, if the correct rate of 2 questions in a row is <50%, the difficulty is reduced to simple, focusing on basic grammar correction.

[0073] The generated question group has a high degree of matching with the teaching goal, and the error inducing effect is improved, effectively exposing the student's knowledge blind spot; the real-time update of the knowledge state improves the efficiency of the exercise, and the number of knowledge points mastered by the student in the same time is increased; dynamic difficulty adjustment reduces invalid exercises and reduces resource consumption.

[0074] The content of the teaching strategy suggestion generated in S5 for the teacher end includes: The internal node splitting of the strategy optimization decision tree is based on error frequency , error weight and knowledge point mastery rate . The feature with the highest information gain is selected as the splitting basis, where the information gain is calculated by the association degree of and ; The strategy suggestion of the split child node needs to meet the diversity constraint to avoid homogenization of the output strategy; The interpretability of the splitting rule is verified by the calculation chain , ensuring that each strategy suggestion can be traced back to the specific , and value; Combined with the constructed dynamic knowledge graph, the deep association between the strategy suggestion and the knowledge point is realized, the error propagation path in the knowledge graph is converted into the inference rule of the decision tree, and when the starting point error is detected, the preventive strategy for the subsequent error is recommended in advance. The node association weight of the knowledge graph is used to prioritize the strategies generated by the decision tree, and the higher the association weight, the higher the ranking of the corresponding strategy suggestion. Through the hierarchical relationship of the knowledge points in the knowledge graph, the strategy suggestion is ensured to cover the upper and lower knowledge points; Based on the output parameters of the teaching strategy optimization formula, a strategy effect evaluation system is constructed, and the reduction amplitude of the error frequency after the implementation of the strategy is evaluated by the difference change of and . The improvement curve of the knowledge point mastery rate is tracked, and when the growth is lower than the preset value for two consecutive weeks, the decision tree parameter update is triggered. The evaluation results are fed back to the dynamic knowledge graph, and the error association weight is adjusted through the incremental update mechanism to make the subsequent strategy suggestion more in line with the actual teaching effect.

[0075] In specific use, taking the Web front-end course JavaScript asynchronous programming unit learning as an example, the teacher needs to obtain teaching strategy suggestions for problems such as callback hell Promise usage errors, etc.

[0076] The strategy optimization decision tree construction and splitting rule, the internal node splitting is characterized by error frequency (8 times per day for callback hell errors), error weight (0.8, affecting the asynchronous execution logic of the program), and knowledge point mastery rate (35%), and the information gain is calculated: information gain IG( )=0.92, IG( )=0.78, IG( )=0.85, and the feature with the highest information gain is selected As the split feature (highest information gain). The split threshold is set to 6 times / day, the left child node contains samples with ≥6, and the right child node contains samples with <6.

[0077] The diversity constraint implementation is the left child node (high frequency of errors) to generate 3-class differentiated strategies: the teaching method is adjusted to increase the callback function and the Promise comparison classroom demonstration (accounting for 20%); the exercise design is to arrange the practice of rewriting the callback nested code into the Promise chain call (the intensity is increased by 40%); the resource push is to recommend the use guide of the asynchronous programming visualization tool to avoid strategy homogenization.

[0078] There is an error propagation path in the dynamic knowledge graph: callback nesting is too deep → Promise rejection is not handled → asynchronous process blockage (path weight 0.8), and the decision tree generates preventive strategies accordingly: when explaining callback hell, introduce the Promise error capture knowledge point in advance (1 class hour earlier than the original plan).

[0079] Based on the association weight of the knowledge graph nodes, the generated strategies are sorted: Promise chain call exercises (association weight 0.9, directly related to core knowledge points); async / await syntax simplifies asynchronous code (association weight 0.7, an advanced solution); asynchronous process debugging tool usage (association weight 0.5, auxiliary tool class knowledge).

[0080] Ensure that the strategy covers the upper and lower knowledge points of asynchronous programming, the upper (JavaScript event loop) arranges 10 minutes of principle explanation, and the lower (Promise.all() method) designs 2 exercises, forming a knowledge closed loop.

[0081] According to the dynamic optimization formula of the teaching strategy, calculate the parameter difference before and after the strategy adjustment - =0.7( is the new strategy parameter, is the original strategy parameter), corresponding to the reduction of error frequency.

[0082] Track the improvement curve of the knowledge point mastery rate , the weekly growth is ≥15%, and there is no need to trigger the decision tree update. If the continuous two-week growth is <10%, adjust the decision tree split threshold (e.g. the split threshold is lowered).

[0083] Feedback the evaluation results (Promise teaching strategy is effective) to the knowledge graph, increase the weight of the Promise usage and callback hell solution association edge (from 0.7 to 0.85), and enhance the accuracy of subsequent strategy recommendations.

[0084] The decision tree splitting improves the targeting of the strategy; the knowledge graph driven reasoning makes the preventive strategy intervene in advance, and the incidence of chain errors is reduced; the explainable rules improve the adoption rate of the teacher strategy, and the effect evaluation and iteration mechanism continuously optimizes.

[0085] The above is only an embodiment of the application, and the application is not limited to the field involved in this embodiment. Common knowledge such as specific structures and characteristics in the scheme is not described in detail here. The ordinary skilled person in the art knows all the ordinary technical knowledge in the technical field of the application before the filing date or the priority date, can know all the prior art in the field, and has the ability to apply conventional experimental means before that date. The ordinary skilled person in the art can perfect and implement the present scheme under the guidance of the present application, combining their own ability. Some typical known structures or known methods should not be an obstacle for the ordinary skilled person in the art to implement the present application. It should be noted that, for those skilled in the art, without departing from the structure of the present application, a number of modifications and improvements can be made, which should also be considered as the protection scope of the present application. These will not affect the effect and practicality of the implementation of the present application. The protection scope of the present application should be subject to the content of its claims. The specific implementation mode and the like in the specification can be used to explain the content of the claims.

Claims

1.A teaching strategy optimization model and a grammar error early warning method based on data mining, characterized in that, The method comprises the following steps: S1, real-time acquisition of original code data in the student code library, and preprocessing of the original code data to obtain a structured data matrix; S2, generating a syntax analysis tree through a syntax analyzer for the structured data matrix, positioning error nodes and extracting context features by using a node traversal algorithm, and constructing a single-error multi-dimensional feature vector; at the same time, constructing an error association rule library based on historical error data, identifying the cause-effect relationship of cross-error types, and forming dynamic knowledge graph data; S3, inputting the single-error multi-dimensional feature vector and the dynamic knowledge graph data into a common difficulty mining model, and processing based on an improved Apriori algorithm and a dynamic optimization formula; when a single student makes the same type of error for more than a preset number of times within a preset time, it is defined as individual early warning, or when more than a preset percentage of students in a class make the same error within a preset time, it is defined as group early warning, and associated knowledge points are extracted from the single-error multi-dimensional feature vector and the knowledge graph to generate an early warning feature vector containing error types, severity, and associated knowledge points; For high-frequency co-occurring error combinations with support exceeding a preset value, a teaching strategy dynamic optimization formula is used to generate strategy adjustment parameters including teaching content difficulty, practice intensity, and explanation focus; S4, inputting the early warning feature vector into a real-time early warning module, and implementing hierarchical intervention according to the error impact range; if it is individual early warning, customized micro-lessons are matched and pushed from a teaching resource library; If it is group early warning, an intensive training module containing associated knowledge point special question groups and a code editing interface real-time error marking is generated to form instant feedback on the student side; S5, inputting the strategy adjustment parameters into a strategy optimization decision tree, and generating teaching strategy suggestions for the teacher side in combination with the knowledge point mapping relationship in the dynamic knowledge graph. 2.The data mining based teaching strategy optimization model and syntax error early warning method according to claim 1, characterized in that, The preprocessing in S1 includes lexical analysis of the original code data, extraction of code token sequences, and identification of natural language descriptions in code comments, and a pre-trained code-text alignment model is used to establish a mapping relationship between code fragments and their function descriptions; Error stack information in the compilation error log is synchronously collected, error types, error line numbers, and compiler prompt texts are extracted, and an error feature triple <error type, error location, error description> is constructed; Student code editing behavior data is extracted, including code modification time intervals, cursor dwell position distribution, and automatic completion request frequency, which are quantified as time series behavior feature vectors; A sliding time window is used to segment and aggregate the original code data, the code change density is calculated for each code modification set in each time window, which is defined as the proportion of modified code lines to total lines per unit time; a code dependency graph is constructed to identify the reference relationship between variables, functions, and classes, and the modification propagation impact value of each code element is calculated. 3.The data mining based teaching strategy optimization model and syntax error early warning method according to claim 2, characterized in that, The preprocessing in S1 includes lexical analysis of the original code data, extraction of code token sequences, and identification of natural language descriptions in code comments, and a pre-trained code-text alignment model is used to establish a mapping relationship between code fragments and their function descriptions; Synchronization acquisition of error stack information in the compilation error log, extraction of error type, error line number and compiler prompt text, construction of error feature triplets <error type, error location, error description>; Extraction of student code editing behavior data, including code modification time interval, cursor dwell position distribution, and auto-completion request frequency, quantification into time series behavior feature vectors; Segmented aggregation of original code data using a sliding time window, calculation of code change density for each code modification set within the time window, definition as the proportion of modified code lines to total lines per unit time; construction of a code dependency graph, identification of reference relationships between variables, functions and classes, and calculation of modification propagation impact values for each code element. 4.The data mining based teaching strategy optimization model and syntax error early warning method according to claim 3, characterized in that, For each node of the syntax parse tree, the S2 further includes: extraction of its structural features in the abstract syntax tree, including node depth, parent node type, number of child nodes, and sibling node relationship, to construct a local structure feature vector; extraction of code within a preset number of lines before and after the error node, conversion into semantic context vectors through a word vector model to capture local semantic information of the code snippet; combination of code change history data to calculate the edit distance and modification type of the current error node relative to the last correct submission version, to generate an evolution feature vector representing the evolution process of the error; Based on the error data aggregated using a dynamic time window, calculation of the co-occurrence frequency of different error types of the same student within adjacent time windows, construction of an error time series transition table to record the evolution probability between error types; For error data of all students in a class, calculation of the propagation centrality of each error type to identify the propagation source node of high-frequency errors, i.e., a key node that may lead to subsequent chain errors; Through a pre-trained code semantic model, calculation of the semantic similarity between different error code snippets to establish a semantic association network between error types; When a new error type or association relationship is detected, dynamically expand the nodes and edges of the knowledge graph; when the newly established error association rule conflicts with existing rules, select the optimal rule by comparing support, confidence and time sequence; Set the timeliness factor of each edge in the knowledge graph to automatically decay the weight of low-frequency association edges over time; when a certain type of error has not appeared for a long time, trigger the subgraph reconstruction mechanism; map single-error multi-dimensional feature vectors to nodes in the knowledge graph; convert error association rules into edges in the knowledge graph; Through attention mechanism, integrate node features and edge relationships to make the representation of each error node in the knowledge graph contain both its own features and association information, forming the final dynamic knowledge graph data. 5.The data mining based teaching strategy optimization model and syntax error early warning method according to claim 4, characterized in that, The step of S3 based on the improved Apriori algorithm for high-frequency co-occurring error combination mining and strategy adjustment parameter generation includes: In the candidate set generation process of the Apriori algorithm, introduce time window constraints and spatial constraints to retain only error combinations that frequently co-occur within a specific spatio-temporal range; for each frequent error combination, calculate its spatio-temporal cohesion degree, i.e., the proximity in time and distance in code space, to form a spatio-temporal association rule library; Based on the constructed dynamic knowledge graph, for each high-frequency co-occurrence error combination, identify the shortest propagation path in the knowledge graph; Assigning error weights to each error node on the path , according to the influence degree of program running, the value range is 0-1; and statistics error frequency , that is, the number of error occurrences per unit time; combined with knowledge point mastery rate , that is, the percentage of mastery calculated through test data, which provides parameters for subsequent formula calculation; According to the course progress, the support threshold is dynamically adjusted, the threshold value in the basic grammar stage is set as a first preset key value, the threshold value in the advanced programming stage is set as a second preset key value, and the adjustment coefficient in the teaching strategy dynamic optimization formula Linkage, when the threshold value is reduced The value is improved, and the sensitivity of strategy adjustment is enhanced. By formula Strengthen the policy intervention efforts; The teaching strategy adjustment parameter generation process is converted into a double-target optimization based on the formula, to maximize as the first target, by increasing the exercise intensity of the high-weight error-related knowledge points; and to minimize the variance of each knowledge point as the second target, by balancing the focus of explanation through the differentiated assignment in the formula The generated strategy adjustment parameters are directly reflected in the formula calculation results, including the teaching content difficulty coefficient, the exercise intensity proportion, and the explanation time length allocation, which dynamically tilt with the value. 6.The data mining based teaching strategy optimization model and syntax error early warning method according to claim 5, characterized in that, The S3 also includes: For single error multi-dimensional feature vector and dynamic knowledge graph data, the correlation weight of each error feature to the knowledge point is calculated through multi-head attention mechanism, and the knowledge points with high weight are selected as the calculation basis of formula ​ Automatically extract the error weight corresponding to these knowledge points , error frequency and mastery rate , form formula input parameter set ; evaluate the error severity from three dimensions of execution influence, repair complexity and propagation potential, and map the weighted sum result to the error weight ; set formula adjustment coefficient based on severity classification , realize differentiated response of strategy adjustment; According to the student knowledge mastery state vector , the exercise intensity adjustment value is calculated by the formula . 7.The data mining based teaching strategy optimization model and syntax error early warning method according to claim 6, characterized in that, The content of the S4 includes the implementation of hierarchical intervention: Construct a three-dimensional impact range evaluation system based on knowledge point mastery rate , evaluate the chain effect of errors on related knowledge systems; analyze the number of program crashes, debugging time, and code modification caused by errors; count the frequency of discussing the error within the class and the number of times seeking help from teachers; input the three-dimensional evaluation results into a fuzzy decision tree to generate a five-level impact range classification; Early warning to individuals, using reinforcement learning to dynamically adjust the micro-lesson push strategy, student knowledge mastery vector , error feature vector; micro-lesson type, push timing, presentation form; with error recurrence rate reduction, micro-lesson completion rate, subsequent test score improvement as indicators; through Q-learning algorithm iteration optimization push strategy, micro-lesson click rate is improved; In the real-time error annotation of the code editing interface, in addition to highlighting the error line, it also includes: Based on the dynamic knowledge graph, mark other potential risk points associated with the current error; Show the shortest repair path from the current error to the normal operation of the program, including 3-5 progressive repair steps; Link to the knowledge point explanation micro-lesson in the teaching resource library directly related to the error. 8.The data mining based teaching strategy optimization model and syntax error early warning method according to claim 7, characterized in that, The content of the S4 includes the generation of reinforcement training modules: Construct a special topic group generation model based on GAN, generate a three-dimensional topic group structure including basic questions, advanced questions, and comprehensive questions according to the associated knowledge points in the early warning feature vector, and each type of question is accompanied by an expected error mode; Based on historical student answer data, judge the difficulty adaptability, knowledge coverage, and error inducement of the generated question group; The knowledge state vector is updated in real time during the student completes the exercise process , the mastery rate of the corresponding knowledge point is improved through the Bayesian update formula when the answer is correct ; when the answer is wrong, the matching degree of the actual error and the expected error mode is analyzed, if it matches, the knowledge point exercise is strengthened, if it does not match, the association analysis of the knowledge graph is triggered; when the correct rate of three consecutive questions exceeds the preset value, the difficulty of the knowledge point exercise is automatically reduced. 9.The data mining based teaching strategy optimization model and syntax error early warning method according to claim 8, characterized in that, The content of the S5 includes the generation of teacher-end teaching strategy suggestions: Splitting of internal nodes of the policy optimization decision tree is based on error frequency , error weight , and knowledge point mastery rate , the feature with the highest information gain is selected as the splitting basis, wherein the information gain is calculated by and correlation degree The strategy suggestions of the split child nodes need to meet the diversity constraint to avoid homogeneous strategy output; The explainability of the splitting rules is achieved by Computing link verifications, ensuring that each policy recommendation can be traced back to a specific , and value; Combined with the constructed dynamic knowledge graph, realize the deep association between strategy suggestions and knowledge points, convert the error propagation path in the knowledge graph into the reasoning rules of the decision tree, and recommend preventive strategies for subsequent errors when detecting the starting point error; Use the node association weight of the knowledge graph to prioritize the strategies generated by the decision tree, the higher the association weight, the higher the ranking of the corresponding strategy suggestion; Through the hierarchical relationship of knowledge points in the knowledge graph, ensure that the strategy suggestions cover the upper and lower knowledge points; Based on the output parameters of the teaching strategy optimization formula, a strategy effectiveness evaluation system is constructed. and The change in the difference is used to assess the error frequency after the strategy is implemented. The extent of the decrease; tracking the mastery rate of knowledge points The upward curve, when for two consecutive weeks When the growth is lower than the preset value, the decision tree parameters are updated; the evaluation results are fed back to the dynamic knowledge graph, and the weight of erroneous associations is adjusted through an incremental update mechanism so that subsequent strategy suggestions are more in line with the actual teaching effect.

Citation Information

Cited By

  • Classroom instant interaction feedback analysis method and system based on portable multimedia display equipment

    CN121354398A

  • Preposed knowledge point prediction method and device, equipment, medium and program product

    CN121503822A

  • Knowledge graph dynamic construction and optimization method based on group behavior and state propagation

    CN121809640A