Programming problem difficulty prediction method and system based on XGBoost
Through the XGBoost-based programming problem difficulty prediction method, combined with dense and sparse characteristics, the problem of the difficulty prediction method in the existing technology lacks the integration of problem characteristics and learners' cognitive state, achieving efficient and accurate difficulty prediction, and enhancing the interpretability of the model.
Patent Information
- Application Number
- CN202510185294.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-13
AI Technical Summary
The existing programming problem difficulty prediction methods are difficult to effectively integrate the characteristics of the question with the learner's cognitive state, and lack high computing efficiency and interpretability, resulting in limited accuracy and comprehensiveness of difficulty prediction.
The difficulty prediction method of programming problem based on XGBoost is used, and dense feature extraction and sparse feature extraction are used, combined with text features and knowledge graph features, and input them into the XGBoost classification model to perform model fusion to generate the final difficulty prediction value.
It improves the accuracy and computing efficiency of programming problem difficulty prediction, enhances the interpretability of the model, and provides more accurate support for personalized programming education.
Smart Images

Figure CN119989174A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of programming education and artificial intelligence, and in particular to a method and system for predicting the difficulty of programming problems based on XGBoost. Background Art
[0002] With the rapid development of information technology and artificial intelligence, online programming education is gradually moving towards intelligence and personalization. As an important tool for programming skill assessment and learning, the online question-judging platform (OJ platform) provides a rich question bank and automated assessment system to help learners with programming training. However, how to intelligently recommend suitable programming questions based on learners' abilities, progress, and needs is still a challenge faced by online platforms.
[0003] At present, the main methods for evaluating the difficulty of programming questions include expert annotation and statistical analysis based on student-submitted data. Expert annotation is inefficient and easily affected by subjective bias, while statistical methods rely on large amounts of data and are difficult to meet the personalized needs of different learners.
[0004] In recent years, deep learning-based prediction of programming problem difficulty has gradually attracted attention, especially the deep knowledge tracking (DKT) model. Although such models have advantages in tracking learners' cognitive states, they often lack interpretability and usually require a large amount of data for training, which is particularly problematic in programming education. In addition, existing models focus more on the relationship between knowledge points and ignore the characteristics of the questions themselves, such as algorithm complexity, input and output requirements, etc., which limits the accuracy and comprehensiveness of difficulty prediction.
[0005] Therefore, there is an urgent need for a new method for predicting the difficulty of programming problems, which can effectively integrate the characteristics of the questions with the learners' cognitive state and have high computational efficiency and interpretability, so as to provide more accurate support for personalized programming education. Summary of the invention
[0006] The purpose of the present invention is to provide a method and system for predicting the difficulty of programming problems based on XGBoost, so as to solve the above-mentioned problems existing in the prior art.
[0007] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0008] A programming problem difficulty prediction method based on XGBoost, the method comprising:
[0009] S1, dense feature extraction, extracts dense features from programming problem descriptions;
[0010] The specific steps of the dense feature extraction include:
[0011] S11, extracting basic text features from the problem description; the basic text features include text description length, average sentence length, frequency of high-frequency words and number of formula symbols; the high-frequency words are domain-specific terms that frequently appear in the problem description;
[0012] S12, encodes the knowledge component of each question into a one-hot vector and combines it with the main data structure and algorithm group to create multiple features to indicate the presence of a specific group of knowledge components;
[0013] S13, extracting features from interaction samples including input and output examples provided to the user in the question text and actual test cases, including input data type, output data type, number of test cases, and maximum data range in the test cases;
[0014] S14, based on the programming constraints of time constraints, space constraints, data sample quantity constraints, and data range constraints, further construct combined features, and combine time-related constraints with knowledge points that are highly dependent on time to strengthen their relevance to task difficulty;
[0015] S2, sparse feature extraction, extracts sparse features from programming problem descriptions;
[0016] The specific steps of the sparse feature extraction include:
[0017] S21, use Doc2Vec algorithm to extract features from the problem statement text and generate a low-dimensional vector representation of the problem description;
[0018] S22, build a knowledge graph in the field of programming, including three types of entities: "data structure", "algorithm" and "problem", and define four edge types: "contain", "solve", "implement" and "predecessor";
[0019] S23, use the TranE algorithm to embed the entities and their relationships in the knowledge graph and generate corresponding vector representations;
[0020] S24, by calculating the similarity between entities, the knowledge graph features related to the specific problem are extracted, thereby effectively representing the domain knowledge and its application in problem solving;
[0021] S25, by combining text features with knowledge graph features, provides the model with richer sparse feature representation, thus significantly improving the ability of question parsing and prediction;
[0022] S3, introducing a classification model based on XGBoost; inputting the dense features and the sparse features into the classification model based on XGBoost, and finally outputting a prediction result after model fusion;
[0023] The model uses XGBoost as the basic model and a tree-based gradient boosting algorithm. Compared with the traditional gradient boosting decision tree, XGBoost has inherent advantages in the generalization ability of the model and can effectively improve the prediction performance of the difficulty of programming tasks. The specific steps of step S3 include:
[0024] S31, dense feature input, inputting the dense features designed by prior knowledge into the first XGBoost model; the dense features have a strong direct correlation with the difficulty of the programming task;
[0025] S32, sparse feature input, inputs the sparse features extracted from the problem description and knowledge graph by Doc2Vec and TranE algorithms into the second XGBoost model; the sparse features provide more extensive background information for the problem and can capture the contextual information of the problem from the text and domain knowledge;
[0026] S33, model fusion, obtains the final difficulty prediction value by weighted fusion of the prediction results of the first and second XGBoost models; the fusion process uses a weighted coefficient α, which is calculated as follows:
[0027]
[0028] in, is the prediction result of the XGBoost model based on dense features, is the prediction result of the XGBoost model based on sparse features, α is a coefficient between 0 and 1, which is used to adjust the relative importance of dense features and sparse features in the final prediction; the weighted coefficient α is optimized through cross-validation;
[0029] S34, output prediction, generate the final programming task difficulty prediction value according to the weighted result obtained in step S33, the prediction value can reflect the difficulty of the task, and provide a basis for the classification and difficulty prediction of programming problems.
[0030] A programming problem difficulty prediction system based on XGBoost, comprising:
[0031] A dense feature extraction module, the sparse feature extraction module and a classification model based on XGBoost;
[0032] The dense feature extraction module is used to extract dense features from the programming problem description; specifically, to:
[0033] Extracting basic text features from the problem description; the basic text features include text description length, average sentence length, frequency of high-frequency words and number of formula symbols; the high-frequency words are domain-specific terms that frequently appear in the problem description;
[0034] Encode each problem’s knowledge component as a one-hot vector and combine it with the main data structure and algorithm group to create multiple features that indicate the presence of a specific group of knowledge components;
[0035] Extract features from interaction samples including input and output examples provided to the user in the question text and actual test cases, including input data type, output data type, number of test cases, and maximum data range in the test cases;
[0036] According to the programming constraints of time limit, space limit, data sample quantity limit and data range limit, we further construct combined features and combine time-related constraints with knowledge points that are highly dependent on time to strengthen their relevance to task difficulty;
[0037] The sparse feature extraction module is used to extract sparse features from the programming problem description; specifically, it is used to:
[0038] Use the Doc2Vec algorithm to extract features from the problem statement text and generate a low-dimensional vector representation of the problem description;
[0039] Build a knowledge graph in the field of programming, including three types of entities: "data structure", "algorithm" and "problem", and define four edge types: "contain", "solve", "implement" and "predecessor";
[0040] Use the TranE algorithm to embed entities and their relationships in the knowledge graph and generate corresponding vector representations;
[0041] By calculating the similarity between entities, the knowledge graph features related to specific problems are extracted, thereby effectively representing domain knowledge and its application in problem solving;
[0042] By combining text features with knowledge graph features, the model is provided with richer sparse feature representations, thus significantly improving the ability to parse and predict problems.
[0043] The XGBoost-based classification model is used to receive the dense features and the sparse features and output prediction results;
[0044] The model uses XGBoost as the basic model and a tree-based gradient boosting algorithm. Compared with the traditional gradient boosting decision tree, XGBoost has inherent advantages in the generalization ability of the model and can effectively improve the prediction performance of the difficulty of programming tasks. The process of receiving the dense features and sparse features and outputting the prediction results is specifically as follows:
[0045] Dense feature input, inputting dense features designed by prior knowledge into the first XGBoost model; the dense features have a strong direct correlation with the difficulty of the programming task;
[0046] Sparse feature input, which is the input of sparse features extracted from the question description and knowledge graph by Doc2Vec and TranE algorithms into the second XGBoost model; the sparse features provide more extensive background information for the question and can capture the contextual information of the question from the text and domain knowledge;
[0047] Model fusion, by weighted fusion of the prediction results of the first and second XGBoost models, to obtain the final difficulty prediction value; the fusion process uses a weighted coefficient α, which is calculated as follows:
[0048]
[0049] in, is the prediction result of the XGBoost model based on dense features, is the prediction result of the XGBoost model based on sparse features, α is a coefficient between 0 and 1, which is used to adjust the relative importance of dense features and sparse features in the final prediction; the weighted coefficient α is optimized through cross-validation;
[0050] Output prediction: Based on the weighted results, the final prediction value of programming task difficulty is generated. This prediction value can reflect the difficulty of the task and provide a basis for the classification and difficulty prediction of programming problems.
[0051] The beneficial effects of the present invention are:
[0052] 1. High efficiency: Compared with traditional deep learning methods, the present invention adopts the XGBoost model, which can complete model training and prediction in a shorter time, significantly improving the prediction efficiency.
[0053] 2. Accuracy: By combining dense features and sparse features and using an integrated learning model, the present invention can better capture the complexity of programming problems and provide more accurate difficulty predictions.
[0054] 3. Explainability: The model of the present invention has strong explainability and can provide teachers or platform administrators with detailed information about which features have a greater impact on the difficulty of the problem, helping to understand and improve the recommendation algorithm.
[0055] 4. Wide application: The present invention is applicable to various online programming platforms, and can provide the platforms with efficient difficulty prediction functions, optimize the question recommendation system, and improve the learning effect of learners. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 It is a flow chart of the programming problem difficulty prediction method based on XGBoost of the present invention.
[0057] Figure 2 It is a schematic diagram of the implementation flow of the programming problem difficulty prediction method based on XGBoost of the present invention. DETAILED DESCRIPTION
[0058] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the following Figure 1 It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0059] The present invention proposes a programming problem difficulty prediction method based on XGBoost, which combines dense features and sparse features and improves prediction accuracy, computational efficiency and model interpretability through ensemble learning. The present invention is described in detail below in conjunction with specific implementation steps.
[0060] Step 1: Dataset and preprocessing: Use public datasets from Codeforces and CodeChef platforms, crawl difficulty labels from the platforms, collect source code, and preprocess the datasets, including noise removal and word segmentation.
[0061] Step 2: Extract dense features. Dense features are mainly obtained by directly extracting or combining basic attributes. Feature sources include:
[0062] Problem description: extract basic text features from the problem description, including text description length, average sentence length, frequency of high-frequency words, and number of formula symbols;
[0063] Knowledge components, which represent the knowledge components of each question through one-hot encoding and are combined with the main data structure and algorithm groups to generate 8 features that represent the presence of a specific knowledge component group;
[0064] Interaction samples, including input and output examples and actual test cases provided to users in the question text, from which four features including input data type, output data type, number of test cases, and maximum data range in the test cases are extracted;
[0065] Programming constraints, including time constraints, space constraints, and data range constraints. Time constraints and space constraints are officially provided classification features. In addition, combined features such as time constraints, space constraints, number of data samples, and data range are constructed, and time-related constraints are combined with highly time-dependent knowledge points to deepen their association with difficulty.
[0066] Step 3: Extract sparse features and use Doc2Vec technology to generate a vector representation of the problem description to extract text features from the text of the problem description and build a knowledge graph about the programming field, including three entities: data structure, algorithm, and problem, and four edge types: "include", "solve", "implement", and "precede";
[0067] Step 4: XGBoost model training. Dense features and sparse features are input into two different XGBoost models to capture different types of information. The models are trained using a 5-fold cross-validation method to ensure that the models can be effectively generalized. XGBoost's hyperparameters include learning rate (0.1), maximum tree depth (6), subsample ratio (0.8), and number of estimated trees (100). During the training process, XGBoost is trained using dense features and sparse features, and the final difficulty prediction is generated based on the contribution weights of different features.
[0068] Step 5: Ensemble learning, combining the XGBoost model based on dense features with the XGBoost model based on sparse features to obtain the final prediction result. Specifically, the XGBoost model based on dense features and the XGBoost model based on sparse features make predictions independently, and then perform weighted averaging according to the preset weights to obtain the final difficulty prediction. The weighted calculation formula is as follows:
[0069]
[0070] Step 6: Model launch. After model training is completed, the difficulty prediction model is deployed online to automatically predict questions without difficulty labels and store the results in the database, supporting functions such as question bank management and difficulty adaptation. After launch, the system monitors the model prediction effect in real time and dynamically optimizes based on new data to ensure that the model continues to adapt to changes and maintains prediction accuracy.
[0071] The present invention also proposes a programming problem difficulty prediction system based on XGBoost, comprising:
[0072] A dense feature extraction module, the sparse feature extraction module and a classification model based on XGBoost;
[0073] The dense feature extraction module is used to extract dense features from the programming problem description; specifically, to:
[0074] Extracting basic text features from the problem description; the basic text features include text description length, average sentence length, frequency of high-frequency words and number of formula symbols; the high-frequency words are domain-specific terms that frequently appear in the problem description;
[0075] Encode each problem’s knowledge component as a one-hot vector and combine it with the main data structure and algorithm group to create multiple features that indicate the presence of a specific group of knowledge components;
[0076] Extract features from interaction samples including input and output examples provided to the user in the question text and actual test cases, including input data type, output data type, number of test cases, and maximum data range in the test cases;
[0077] According to the programming constraints of time limit, space limit, data sample quantity limit and data range limit, we further construct combined features and combine time-related constraints with knowledge points that are highly dependent on time to strengthen their relevance to task difficulty;
[0078] The sparse feature extraction module is used to extract sparse features from the programming problem description; specifically, it is used to:
[0079] Use the Doc2Vec algorithm to extract features from the problem statement text and generate a low-dimensional vector representation of the problem description;
[0080] Build a knowledge graph in the field of programming, including three types of entities: "data structure", "algorithm" and "problem", and define four edge types: "contain", "solve", "implement" and "predecessor";
[0081] Use the TranE algorithm to embed entities and their relationships in the knowledge graph and generate corresponding vector representations;
[0082] By calculating the similarity between entities, the knowledge graph features related to specific problems are extracted, thereby effectively representing domain knowledge and its application in problem solving;
[0083] By combining text features with knowledge graph features, the model is provided with richer sparse feature representations, thus significantly improving the ability to parse and predict problems.
[0084] The XGBoost-based classification model is used to receive the dense features and the sparse features and output prediction results;
[0085] The model uses XGBoost as the basic model and a tree-based gradient boosting algorithm. Compared with the traditional gradient boosting decision tree, XGBoost has inherent advantages in the generalization ability of the model and can effectively improve the prediction performance of the difficulty of programming tasks. The process of receiving the dense features and sparse features and outputting the prediction results is specifically as follows:
[0086] Dense feature input, inputting dense features designed by prior knowledge into the first XGBoost model; the dense features have a strong direct correlation with the difficulty of the programming task;
[0087] Sparse feature input, which is the input of sparse features extracted from the question description and knowledge graph by Doc2Vec and TranE algorithms into the second XGBoost model; the sparse features provide more extensive background information for the question and can capture the contextual information of the question from the text and domain knowledge;
[0088] Model fusion, by weighted fusion of the prediction results of the first and second XGBoost models, to obtain the final difficulty prediction value; the fusion process uses a weighted coefficient α, which is calculated as follows:
[0089]
[0090] in, is the prediction result of the XGBoost model based on dense features, is the prediction result of the XGBoost model based on sparse features, α is a coefficient between 0 and 1, which is used to adjust the relative importance of dense features and sparse features in the final prediction; the weighted coefficient α is optimized through cross-validation;
[0091] Output prediction: Based on the weighted results, the final prediction value of programming task difficulty is generated. This prediction value can reflect the difficulty of the task and provide a basis for the classification and difficulty prediction of programming problems.
[0092] By adopting the above technical solution disclosed in the present invention, the following beneficial effects are obtained:
[0093] 1. High efficiency: Compared with traditional deep learning methods, the present invention adopts the XGBoost model, which can complete model training and prediction in a shorter time, significantly improving the prediction efficiency.
[0094] 2. Accuracy: By combining dense features and sparse features and using an integrated learning model, the present invention can better capture the complexity of programming problems and provide more accurate difficulty predictions.
[0095] 3. Explainability: The model of the present invention has strong explainability and can provide teachers or platform administrators with detailed information about which features have a greater impact on the difficulty of the problem, helping to understand and improve the recommendation algorithm.
[0096] 4. Wide application: The present invention is applicable to various online programming platforms, and can provide the platforms with efficient difficulty prediction functions, optimize the question recommendation system, and improve the learning effect of learners.
[0097] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be considered as the scope of protection of the present invention.
Claims
1. A programming problem difficulty prediction method based on XGBoost, characterized in that: The method comprises: S1, dense feature extraction, extracts dense features from programming problem descriptions; The specific steps of the dense feature extraction include: S11, extracting basic text features from the problem description; the basic text features include text description length, average sentence length, frequency of high-frequency words and number of formula symbols; the high-frequency words are domain-specific terms that frequently appear in the problem description; S12, encodes the knowledge component of each question into a one-hot vector and combines it with the main data structure and algorithm group to create multiple features to indicate the presence of a specific group of knowledge components; S13, extracting features from interaction samples including input and output examples provided to the user in the question text and actual test cases, including input data type, output data type, number of test cases, and maximum data range in the test cases; S14, based on the programming constraints of time constraints, space constraints, data sample quantity constraints, and data range constraints, further construct combined features, and combine time-related constraints with knowledge points that are highly dependent on time to strengthen their relevance to task difficulty; S2, sparse feature extraction, extracts sparse features from programming problem descriptions; The specific steps of the sparse feature extraction include: S21, use Doc2Vec algorithm to extract features from the problem statement text and generate a low-dimensional vector representation of the problem description; S22, build a knowledge graph in the field of programming, including three types of entities: "data structure", "algorithm" and "problem", and define four edge types: "include", "solve", "implement" and "predetermine"; S23, use the TranE algorithm to embed the entities and their relationships in the knowledge graph and generate corresponding vector representations; S24, by calculating the similarity between entities, the knowledge graph features related to the specific problem are extracted, thereby effectively representing the domain knowledge and its application in problem solving; S25, by combining text features with knowledge graph features, provides the model with richer sparse feature representation, thus significantly improving the ability of question parsing and prediction; S3, introducing a classification model based on XGBoost; inputting the dense features and the sparse features into the classification model based on XGBoost, and finally outputting a prediction result after model fusion; The model uses XGBoost as the basic model and a tree-based gradient boosting algorithm. Compared with the traditional gradient boosting decision tree, XGBoost has inherent advantages in the generalization ability of the model and can effectively improve the prediction performance of the difficulty of programming tasks. The specific steps of step S3 include: S31, dense feature input, inputting the dense features designed by prior knowledge into the first XGBoost model; the dense features have a strong direct correlation with the difficulty of the programming task; S32, sparse feature input, inputs the sparse features extracted from the problem description and knowledge graph by Doc2Vec and TranE algorithms into the second XGBoost model; the sparse features provide more extensive background information for the problem and can capture the contextual information of the problem from the text and domain knowledge; S33, model fusion, obtains the final difficulty prediction value by weighted fusion of the prediction results of the first and second XGBoost models; the fusion process uses a weighted coefficient α, which is calculated as follows: , in, is the prediction result of the XGBoost model based on dense features, is the prediction result of the XGBoost model based on sparse features, α is a coefficient between 0 and 1, which is used to adjust the relative importance of dense features and sparse features in the final prediction; the weighted coefficient α is optimized through cross-validation; S34, output prediction, generate the final programming task difficulty prediction value according to the weighted result obtained in step S33, the prediction value can reflect the difficulty of the task, and provide a basis for the classification and difficulty prediction of programming problems.
2. A programming problem difficulty prediction system based on XGBoost, characterized in that: include: A dense feature extraction module, the sparse feature extraction module and a classification model based on XGBoost; The dense feature extraction module is used to extract dense features from the programming problem description; specifically, to: Extracting basic text features from the problem description; the basic text features include text description length, average sentence length, frequency of high-frequency words and number of formula symbols; the high-frequency words are domain-specific terms that frequently appear in the problem description; Encode each problem’s knowledge component as a one-hot vector and combine it with the main data structure and algorithm group to create multiple features that indicate the presence of a specific group of knowledge components; Extract features from interaction samples including input and output examples provided to the user in the question text and actual test cases, including input data type, output data type, number of test cases, and maximum data range in the test cases; According to the programming constraints of time limit, space limit, data sample quantity limit and data range limit, we further construct combined features and combine time-related constraints with knowledge points that are highly dependent on time to strengthen their relevance to task difficulty; The sparse feature extraction module is used to extract sparse features from the programming problem description; Specifically used for: Use the Doc2Vec algorithm to extract features from the problem statement text and generate a low-dimensional vector representation of the problem description; Build a knowledge graph in the field of programming, including three types of entities: "data structure", "algorithm" and "problem", and define four edge types: "include", "solve", "implement" and "predecessor"; Use the TranE algorithm to embed entities and their relationships in the knowledge graph and generate corresponding vector representations; By calculating the similarity between entities, the knowledge graph features related to specific problems are extracted, thereby effectively representing domain knowledge and its application in problem solving; By combining text features with knowledge graph features, the model is provided with richer sparse feature representations, thus significantly improving the ability to parse and predict problems. The XGBoost-based classification model is used to receive the dense features and the sparse features and output prediction results; The model uses XGBoost as the basic model and a tree-based gradient boosting algorithm. Compared with the traditional gradient boosting decision tree, XGBoost has inherent advantages in the generalization ability of the model and can effectively improve the prediction performance of the difficulty of programming tasks. The process of receiving the dense features and sparse features and outputting the prediction results is specifically as follows: Dense feature input, inputting dense features designed by prior knowledge into the first XGBoost model; the dense features have a strong direct correlation with the difficulty of the programming task; Sparse feature input, which is the input of sparse features extracted from the question description and knowledge graph by Doc2Vec and TranE algorithms into the second XGBoost model; the sparse features provide more extensive background information for the question and can capture the contextual information of the question from the text and domain knowledge; Model fusion, by weighted fusion of the prediction results of the first and second XGBoost models, to obtain the final difficulty prediction value; the fusion process uses a weighted coefficient α, which is calculated as follows: , in, is the prediction result of the XGBoost model based on dense features, is the prediction result of the XGBoost model based on sparse features, α is a coefficient between 0 and 1, which is used to adjust the relative importance of dense features and sparse features in the final prediction; the weighted coefficient α is optimized through cross-validation; Output prediction: Based on the weighted results, the final prediction value of programming task difficulty is generated. This prediction value can reflect the difficulty of the task and provide a basis for the classification and difficulty prediction of programming problems.