A proposition examination question rating method and system based on big data deep mining
By integrating and analyzing educational data to generate standard problem-solving logic trees, and combining big data mining and neural network models, the problems of data integration and rating accuracy in existing test question rating technologies have been solved, enabling rapid and scientific rating of test questions and improving the quality of education.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2026-03-24
AI Technical Summary
Existing test item rating technologies suffer from insufficient data integration and processing capabilities, shallow feature mining, and inappropriate rating methods, making it difficult to generate accurate test item rating indicators and affecting the rating capabilities of the rating model.
By integrating historical exam data, detailed student answer data, and teaching resource data, a standard problem-solving logic tree is generated. Big data mining algorithms are used to analyze answer records and score distribution. Combined with the academic ability differentiation base, a neural network model is constructed to grade exam questions.
It enables rapid and scientific grading of test questions, improves the accuracy and efficiency of grading, and provides educators with tools to optimize teaching and examination arrangements.
Smart Images

Figure CN120632083B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of test question assessment technology, and in particular to a test question rating method and system based on deep data mining. Background Technology
[0002] In the field of education, the quality of exam questions directly impacts the accuracy and effectiveness of teaching evaluation. Accurately grading exam questions helps teachers understand students' learning progress, optimize teaching content and methods, and provides a basis for the rational allocation of educational resources. Traditional exam question grading methods rely heavily on teachers' experience and judgment, which is highly subjective and makes it difficult to comprehensively and objectively assess various indicators of exam questions. With the rapid development of information technology, big data is increasingly widely used in education. Exam question grading methods and systems based on deep data mining have emerged. Utilizing massive amounts of historical exam data, detailed student response data, and teaching resource data, through data processing, feature extraction, and algorithm analysis, they can more scientifically and comprehensively evaluate key indicators such as difficulty, discrimination, reliability, and validity of exam questions. This not only improves the accuracy and objectivity of exam question grading but also provides more targeted decision support for education and teaching, promoting the development of a scientific and intelligent education evaluation system, and has broad application prospects.
[0003] However, existing test item rating technologies have several problems. Firstly, their data integration and processing capabilities are inadequate, failing to effectively integrate and process historical exam data, student responses, and teaching resource data, making it difficult to form large-scale datasets and leaving a lack of reliable data foundation for subsequent analysis. Secondly, their analysis of test item characteristics is superficial, failing to generate standard problem-solving logic trees to obtain accurate test item features. Thirdly, their rating methods are inappropriate, failing to utilize big data algorithms for comprehensive analysis and consider differences in academic ability, making it difficult to derive accurate rating indicators, thus affecting the rating capability of the constructed test item rating model and hindering the scientific evaluation of new test items.
[0004] Therefore, this invention proposes a method and system for evaluating exam questions based on deep data mining. Summary of the Invention
[0005] This invention provides a method and system for evaluating exam questions based on deep data mining. The method includes: in the data processing stage, integrating historical exam data, detailed student answer data, and teaching resource data, and performing cleaning and normalization to form a massive dataset. This not only ensures the quality and standardization of the data but also comprehensively covers various key information related to the exam questions, providing a solid data foundation for subsequent in-depth analysis and making the evaluation results more reliable and comprehensive. In the process of generating question features, a standard problem-solving logic tree is generated based on the exam question text and knowledge point tags, and an application difficulty matrix and an application degree matrix are derived accordingly. This deeply mines the internal knowledge structure and application characteristics of the exam questions, characterizing the essential features of the questions from a question-setting perspective, which helps to more accurately understand and evaluate the questions, providing an important internal basis for subsequent evaluation. When obtaining the exam question evaluation results, big data mining algorithms are used to analyze the answer records and score distribution, taking into account the academic ability differential base. This approach fully integrates students' actual responses and comprehensively considers the differences in abilities among students. This results in ratings such as difficulty coefficients, discrimination indices, reliability values, and validity values that are more closely aligned with actual exam scenarios, truly reflecting the effectiveness of the exam questions in assessing students and providing more valuable references for teaching and exam evaluation. A neural network model is used to learn the characteristics of question design and rating, resulting in a question rating model. This model construction process leverages advanced machine learning technology, automatically learning potential patterns from large amounts of data and possessing strong generalization capabilities. It can effectively rate new questions, improving the efficiency and accuracy of rating. The rating results are obtained based on the standard problem-solving logic tree and the question rating model for the questions to be rated, enabling rapid and scientific rating of new questions. This provides educators with a powerful tool for question design, selection, and evaluation of teaching effectiveness, helping to optimize teaching and exam arrangements and improve the quality of education.
[0006] This invention provides a method for evaluating exam questions based on deep data mining, comprising:
[0007] S1: Integrate historical exam data, detailed student answer data, and teaching resource data, and clean and normalize them to form a massive dataset containing exam question text, answer records, score distribution, and knowledge point tags;
[0008] S2: Generate a standard problem-solving logic tree based on the question text and knowledge point tags of each historical question in the ultra-large-scale dataset. Generate the application difficulty matrix and the association application degree matrix of each question based on the structural features and structural correlation of all knowledge points in the standard problem-solving logic tree as the proposition features of each historical question.
[0009] S3: Utilize big data mining algorithms to analyze the answer records and score distribution of each historical exam question in a massive dataset, and consider the base of academic ability differentiation among the answerers to obtain the difficulty coefficient, discrimination index, reliability value, and validity value of each historical exam question as the exam question rating result;
[0010] S4: Use a neural network model to learn the question-setting features and question-rating features of each historical exam question in a massive dataset to obtain an exam question rating model;
[0011] S5: Obtain the test question rating result based on the standard problem-solving logic tree and test question rating model of the test question to be rated.
[0012] Optionally, S2: Generate a standard problem-solving logic tree based on the question text and knowledge point tags of each historical question in the ultra-large-scale dataset. Generate an application difficulty matrix and an application degree matrix for each question based on the structural features and structural relevance of all knowledge points in the standard problem-solving logic tree, as the propositional features of each historical question, including:
[0013] Based on the question text and standard solution approach for each historical question in the ultra-large-scale dataset, at least one standard solution approach for each historical question is determined.
[0014] A standard problem-solving logic tree is generated based on the standard problem-solving approach for each historical exam question and the knowledge point tags of all knowledge points involved in the standard problem-solving approach;
[0015] Based on the level depth of each knowledge point involved in each standard solution approach for each historical exam question and the structural complexity of the corresponding standard solution logic tree, the application difficulty of the corresponding knowledge point is assessed.
[0016] A difficulty matrix for each historical exam question is generated based on the application difficulty of all knowledge points involved in each standard problem-solving approach.
[0017] Based on the structural correlation analysis of the different knowledge points involved in each standard solution approach of each historical exam question in the corresponding standard solution logic tree, the degree of correlation and application between the different knowledge points involved in each standard solution approach of each historical exam question is analyzed.
[0018] A correlation application degree matrix for each historical exam question is generated based on the correlation application degree between different knowledge points involved in each standard problem-solving approach.
[0019] The application difficulty matrix and the correlation application degree matrix of each test question are used as the question-setting features of each historical test question.
[0020] Optionally, based on the hierarchical depth of each knowledge point involved in each standard solution approach for each historical exam question within its respective standard solution logic tree and the structural complexity of the corresponding standard solution logic tree, the application difficulty of the corresponding knowledge point is assessed, including:
[0021] The structural complexity of the standard problem-solving logic tree is calculated based on the tree depth, molecular factor, node density, and cycle degree of the standard logic tree.
[0022] Based on the level of each knowledge point involved in each standard problem-solving approach of each historical exam question in the knowledge point mastery expansion tree of the knowledge point section of the corresponding historical exam question and the level of the corresponding standard logic tree, the weight value of the corresponding knowledge point is determined.
[0023] Based on the level depth and corresponding weight value of each knowledge point involved in each standard solution approach for each historical exam question in the corresponding standard solution logic tree, as well as the structural complexity of the corresponding standard solution logic tree, the application difficulty of the corresponding knowledge point is assessed.
[0024] Optionally, based on the structural relevance analysis of the different knowledge points involved in each standard problem-solving approach of each historical exam question within the corresponding standard problem-solving logic tree, the degree of correlation and application among the different knowledge points involved in each standard problem-solving approach of each historical exam question is analyzed, including:
[0025] Based on the minimum number of edges between each pair of knowledge points involved in each standard solution of each historical exam question, the minimum number of edges between the nearest common ancestor node and the root node, and the overlap between the corresponding subtree structures of the two knowledge points in the standard solution logic tree, a problem-solving logic structure correlation vector for the two knowledge points is generated.
[0026] Based on the minimum number of edges between each pair of knowledge points involved in each standard solution to each historical exam question in the knowledge point mastery extension tree of the smallest range of knowledge point blocks involved in the corresponding historical exam question, the minimum number of edges between the nearest common ancestor node and the root node, and the overlap between the corresponding subtree structures of the two knowledge points in the corresponding knowledge point mastery extension tree, a knowledge point mastery extension structure correlation vector of the two knowledge points is generated.
[0027] For each standard solution approach in each historical exam question, the nearest common ancestor node of each pair of knowledge points in the corresponding standard solution logic tree and the nearest common ancestor node of the knowledge point mastery extension tree in the smallest knowledge point block involved in the corresponding historical exam question are taken as the structural superordinate mapping knowledge point combination of the two nodes, and the knowledge point mastery extension structure correlation vector of the corresponding structural superordinate mapping knowledge point combination is generated as the superordinate knowledge point mastery extension structure correlation vector of the two knowledge points.
[0028] Based on the correlation vectors of the logical structure of each pair of knowledge points involved in each standard problem-solving approach for each historical exam question, the correlation vectors of the knowledge point mastery and extension structure, and the correlation vectors of the mastery and extension structure of the superior knowledge points, the degree of correlation and application between different knowledge points involved in each standard problem-solving approach for each historical exam question is analyzed.
[0029] Optionally, based on the correlation vector of the logical structure of each pair of knowledge points involved in each standard problem-solving approach for each historical exam question, the correlation vector of the knowledge point mastery and extension structure, and the correlation vector of the mastery and extension structure of higher-level knowledge points, the degree of correlation and application between different knowledge points involved in each standard problem-solving approach for each historical exam question is analyzed, including:
[0030] Based on the level of each pair of knowledge points involved in each standard problem-solving approach of each historical exam question in the corresponding standard problem-solving logic tree, the level of knowledge point mastery extension tree in the knowledge point mastery of the smallest knowledge point section involved in the corresponding historical exam question, and the level of the corresponding structural superordinate mapping knowledge point combination in the knowledge point mastery extension tree in the knowledge point mastery of the smallest knowledge point section involved in the corresponding historical exam question, determine the weight of the problem-solving logic structure relevance vector of the two knowledge points, the weight of the knowledge point mastery extension structure relevance vector, and the weight of the superordinate knowledge point mastery extension structure relevance vector.
[0031] Based on the weights of the problem-solving logic structure relevance vector, the knowledge point mastery extension structure relevance vector, and the higher-level knowledge point mastery extension structure relevance vector involved in each standard problem-solving approach for each historical exam question, the problem-solving logic structure relevance vector, knowledge point mastery extension structure relevance vector, and higher-level knowledge point mastery extension structure relevance vector for corresponding two knowledge points are weighted and summed to obtain the degree of correlation and application between different knowledge points involved in each standard problem-solving approach for each historical exam question.
[0032] Optionally, S3: Utilize big data mining algorithms to analyze the answer records and score distribution of each historical exam question in the massive dataset, and consider the academic ability disparity among the respondents to obtain the difficulty coefficient, discrimination index, reliability value, and validity value of each historical exam question as the exam question rating result, including:
[0033] By using an academic ability stratification model, we can analyze the academic ability-related characteristics of all respondents to all historical exam questions in a massive dataset and obtain the academic ability stratification label for each respondent.
[0034] Based on the academic ability stratification model, the academic ability-related characteristics of all respondents to all historical exam questions in the ultra-large-scale dataset are analyzed, and the academic ability base of each respondent is calculated.
[0035] Based on the academic ability base of all respondents with the same academic ability stratification label among all respondents of all historical exam questions in the ultra-large-scale dataset, the horizontal alienation index of academic ability for each respondent is calculated.
[0036] Based on the academic ability base of all respondents to all historical exam questions in the ultra-large-scale dataset, the maximum and minimum academic ability base in each academic ability stratification label, the longitudinal alienation index of academic ability for each respondent is calculated;
[0037] The horizontal and vertical academic ability differentiation indices of all respondents were used as the base for academic ability differentiation among respondents. Combined with the answer records and score distribution of each historical test question in the ultra-large-scale dataset, the difficulty coefficient, discrimination index, reliability value and validity value of each historical test question were obtained as the test question rating result.
[0038] Optionally, the horizontal and vertical academic ability differentiation indices of all respondents are used as the base for academic ability differentiation among respondents. Combined with the answer records and score distribution of each historical exam question in the large-scale dataset, the difficulty coefficient, discrimination index, reliability value, and validity value of each historical exam question are obtained as the exam question rating results for each historical exam question, including:
[0039] Based on the answer records of each historical exam question in the ultra-large-scale dataset, the ratio of the number of correct answers to the total number of people in each academic ability level label is determined as the original difficulty of the corresponding historical exam question under each academic ability level label. Based on the horizontal and vertical academic ability differentiation indices of all answerers under each academic ability level label, the original difficulty of the corresponding historical exam question under the corresponding academic ability level label is differentiated and corrected to obtain the difficulty coefficient of each historical exam question.
[0040] The discrimination index of each historical exam question is calculated based on the score distribution of each historical exam question in the ultra-large-scale dataset and the horizontal alienation index of academic ability of all respondents.
[0041] The reliability value of each historical exam question was calculated based on the score distribution of each historical exam question in the ultra-large-scale dataset and the cross-academic alienation index of all respondents.
[0042] Based on the score distribution of each historical test question in the ultra-large-scale dataset, a knowledge point-differentiation basis matrix is constructed for each knowledge point. Based on the knowledge point-differentiation basis matrix of each knowledge point, the mastery of each knowledge point in different differentiation intervals is calculated. Based on the mastery of each knowledge point in different differentiation intervals, the validity value of each historical test question is calculated.
[0043] Optionally, the discrimination index of each historical exam question is calculated based on the score distribution of each historical exam question in the ultra-large-scale dataset and the cross-academic differentiation index of all respondents, including:
[0044] Based on the score distribution analysis of each historical test question in the ultra-large-scale dataset, the correlation coefficient between the horizontal differentiation index of academic ability for each academic ability stratification label and the score of the corresponding historical test question is used as the intra-stratification discrimination of each historical test question.
[0045] Based on the score distribution of each historical test question in the ultra-large-scale dataset, the average score and standard deviation of all respondents with the highest academic ability stratum label and the average score and standard deviation of all respondents with the lowest academic ability stratum label are determined. Based on the average score and standard deviation of all respondents with the highest academic ability stratum label and the average score and standard deviation of all respondents with the lowest academic ability stratum label, the cross-stratum discrimination of each historical test question is calculated.
[0046] The discrimination index of each historical test question is calculated based on the intra-layer discrimination and cross-layer discrimination of each historical test question in the ultra-large-scale dataset.
[0047] Optionally, the reliability value of each historical exam question is calculated based on the score distribution of each historical exam question in the ultra-large-scale dataset and the cross-academic alienation index of all respondents, including:
[0048] Based on the variance of the scores of all respondents to the corresponding historical exam questions for each academic ability stratification label in the ultra-large-scale dataset and the variance of the total score of the corresponding historical exam questions to which all respondents belong for the corresponding academic ability stratification labels, the stratified Cronbach's coefficient for each academic ability stratification label is calculated.
[0049] The stratified Cronbach's coefficient of each academic ability stratification label is modified by using the standard deviation of the horizontal alienation index of all respondents for each academic ability stratification label.
[0050] The minimum value among the modified stratified Cronbach's coefficients of all academic ability stratification labels is used as the reliability value for each historical test question.
[0051] This invention provides a test question rating system based on deep data mining, comprising:
[0052] The data integration module is used to integrate historical exam data, detailed student answer data, and teaching resource data, and to clean and normalize them to form a massive dataset containing exam question text, answer records, score distribution, and knowledge point tags.
[0053] The question feature analysis module is used to generate a standard problem-solving logic tree based on the question text and knowledge point tags of each historical question in the ultra-large-scale dataset. Based on the structural features and structural relevance of all knowledge points in the standard problem-solving logic tree, it generates the application difficulty matrix and the correlation application degree matrix of each question as the question features of each historical question.
[0054] The historical exam question rating module uses big data mining algorithms to analyze the answer records and score distribution of each historical exam question in a massive dataset, and takes into account the academic ability differentiation base among the respondents to obtain the difficulty coefficient, discrimination index, reliability value and validity value of each historical exam question as the exam question rating result.
[0055] The rating model building module is used to learn the question-setting features and rating features of each historical exam question in a large-scale dataset using a neural network model, and obtain the exam question rating model.
[0056] The model output rating module is used to obtain the test rating result of the test questions based on the standard problem-solving logic tree and test question rating model.
[0057] The beneficial effects of this invention compared to existing technologies are as follows: In the data processing stage, by integrating historical exam data, detailed student answer data, and teaching resource data, and performing cleaning and normalization processes, a massive dataset is formed. This not only ensures the quality and standardization of the data but also comprehensively covers various key information related to the exam questions, providing a solid data foundation for subsequent in-depth analysis and making the rating results more reliable and comprehensive. In the process of generating question features, a standard problem-solving logic tree is generated based on the exam question text and knowledge point tags, and based on this, an application difficulty matrix and an association application degree matrix are derived. This deeply explores the internal knowledge structure and application characteristics of the exam questions, characterizing the essential features of the exam questions from a question-setting perspective, which helps to more accurately understand and evaluate the exam questions, providing an important internal basis for subsequent rating. When obtaining the exam question rating results, big data mining algorithms are used to analyze the answer records and score distribution, taking into account the academic ability differentiation base. This approach fully integrates students' actual responses and comprehensively considers the differences in abilities among students. This results in ratings such as difficulty coefficients, discrimination indices, reliability values, and validity values that are more closely aligned with actual exam scenarios, truly reflecting the effectiveness of the exam questions in assessing students and providing more valuable references for teaching and exam evaluation. A neural network model is used to learn the characteristics of question design and rating, resulting in a question rating model. This model construction process leverages advanced machine learning technology, automatically learning potential patterns from large amounts of data and possessing strong generalization capabilities. It can effectively rate new questions, improving the efficiency and accuracy of rating. The rating results are obtained based on the standard problem-solving logic tree and the question rating model for the questions to be rated, enabling rapid and scientific rating of new questions. This provides educators with a powerful tool for question design, selection, and evaluation of teaching effectiveness, helping to optimize teaching and exam arrangements and improve the quality of education.
[0058] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in this application.
[0059] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0060] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0061] Figure 1 This is a flowchart of a question-based assessment method for exam questions based on deep data mining, as described in an embodiment of the present invention.
[0062] Figure 2 This is a distributed computing architecture design diagram for the question-setting and rating method based on big data deep mining in this embodiment of the invention. Detailed Implementation
[0063] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0064] refer to Figure 1 This invention provides an implementation method for a test question rating method based on deep data mining, including:
[0065] S1: Integrating historical exam data, detailed student answer data, and teaching resource data, and performing cleaning and normalization processes to form a massive dataset containing exam question text, answer records, score distribution, and knowledge point labels. This integration of historical exam data, detailed student answer data, and teaching resource data comes from a wide range of sources, covering various information from past exams, specific student answer details, and related teaching materials. Cleaning removes noise, errors, or duplicate information from the data, such as deleting obviously incorrect answers from answer records and correcting typos in the exam question text. Normalization transforms data of different magnitudes or units into a unified scale, such as standardizing scores from different exams to a range of 0-100 for subsequent analysis. After these operations, a massive dataset is formed, containing exam question text (i.e., question content), answer records (students' responses to each question), score distribution (the distribution of students across different score ranges), and knowledge point labels (the tags indicating the knowledge points involved in each question). For example, data from a history math exam might include the text of a geometry question, students' responses to the question, the distribution of scores across all students, and tags related to the concepts of triangles and similarities.
[0066] S2: A standard problem-solving logic tree is generated based on the exam text and knowledge point tags of each historical exam question in the massive dataset. The application difficulty matrix and relevance matrix of each exam question are generated based on the structural features and relevance of all knowledge points within the standard problem-solving logic tree, serving as the propositional features of each historical exam question. This is like constructing a logical framework for solving a problem, showing the steps required from the given conditions to arriving at the answer, as well as the logical relationships between the involved knowledge points. For example, in a physics mechanics problem, the standard problem-solving logic tree might start by analyzing the forces acting on the object, gradually deriving the calculations using Newton's laws, clearly presenting the position and role of each knowledge point in the problem-solving process.
[0067] Then, based on the structural characteristics of all knowledge points in the standard problem-solving logic tree (such as the level of the knowledge point and its connection with other knowledge points) and structural relevance (the degree of closeness and mutual influence between different knowledge points), an application difficulty matrix and a correlation application degree matrix are generated for each exam question, serving as the question-setting characteristics for each historical exam question. The application difficulty matrix reflects the ease or difficulty of applying each knowledge point in the problem-solving process. For example, if a knowledge point is located at a deeper level in the logic tree, it may mean that its application difficulty is higher, and the corresponding value in the matrix will reflect this characteristic. The correlation application degree matrix reflects the degree of mutual application of different knowledge points in problem-solving. For example, if two knowledge points are closely connected in the logic tree and are frequently used together in the problem-solving process, then their corresponding values in the correlation application degree matrix will be higher.
[0068] S3: Using big data mining algorithms, the answer records and score distribution of each historical exam question in a massive dataset are analyzed. Considering the academic ability disparity among test takers, the difficulty coefficient, discrimination index, reliability value, and validity value of each historical exam question are obtained as the exam question rating result. The academic ability disparity reflects the differences among students in learning ability and knowledge mastery. For example, some students have a solid foundation and strong learning ability, while others may be relatively weak; this difference is reflected through the academic ability disparity. Through comprehensive analysis, the difficulty coefficient (measuring the difficulty level of the question for students), discrimination index (the ability to differentiate between students of different academic ability levels), reliability value (the reliability of the test results), and validity value (whether the question accurately tests the expected knowledge and ability) of each historical exam question are derived. These values serve as the exam question rating result for each historical exam question.
[0069] S4: Utilize a neural network model to learn the question-setting features and rating features of each historical exam question in a massive dataset, obtaining an exam question rating model. This involves using a neural network model to learn the question-setting features (i.e., the features represented by the previously generated application difficulty matrix and association application degree matrix) and rating features (features reflected by difficulty coefficient, discrimination index, reliability value, and validity value) of each historical exam question in the massive dataset. Neural network models possess powerful learning capabilities, automatically extracting patterns and regularities from large amounts of data. By learning these features, an exam question rating model is constructed. This model acts like an intelligent tool, providing corresponding ratings based on the input exam question information.
[0070] S5: Obtain the question rating result for the question to be rated based on the standard solution logic tree and the question rating model. For each question to be rated, its standard solution logic tree is first generated, and then this logic tree is combined with the previously obtained question rating model. The question rating model analyzes the question based on the information contained in the standard solution logic tree and the rules learned previously, ultimately obtaining the question rating result, thus achieving rapid and scientific rating of new questions. For example, for a newly released chemistry question, by generating its standard solution logic tree and inputting it into the model, the model can provide rating results such as the question's difficulty coefficient and discrimination index, helping educators understand the quality and applicability of the question.
[0071] In an alternative implementation, S2: Generate a standard problem-solving logic tree based on the question text and knowledge point tags of each historical question in the ultra-large-scale dataset. Generate an application difficulty matrix and an association application degree matrix for each question based on the structural features and structural relevance of all knowledge points in the standard problem-solving logic tree, as the propositional features of each historical question, including:
[0072] Based on the exam text and standard solution approaches for each historical exam question in a massive dataset, at least one standard solution approach is identified for each historical exam question. For example, a mathematical proof problem may have a conventional solution and a clever, simplified solution, thus forming at least two standard solution approaches. This step forms the basis for subsequently constructing a standard solution logic tree, ensuring a clear understanding of the solution methods for each exam question.
[0073] A standard problem-solving logic tree is generated based on the standard problem-solving approach for each historical exam question and the knowledge point tags involved in the standard problem-solving approach. For example, when solving a physics circuit problem, the problem-solving approach might be to first analyze the circuit connection method, and then use knowledge points such as Ohm's Law to calculate the current and voltage of each part. Based on this approach and the relevant knowledge point tags such as "circuit connection" and "Ohm's Law", the constructed logic tree will display the order and interrelationship of these knowledge points in the problem-solving process in a hierarchical structure. For example, "circuit connection" is the upper-level node, which leads to the calculation steps of the lower-level nodes such as "Ohm's Law".
[0074] Based on the level depth of each knowledge point involved in each standard solution approach for each historical exam question within its corresponding standard solution logic tree and the structural complexity of the corresponding standard solution logic tree, the application difficulty of the corresponding knowledge point is assessed. Level depth refers to the number of levels of the knowledge point from the root node in the logic tree. A higher level may mean greater application difficulty. For example, a knowledge point at the bottom level of a complex logic tree may require mastering multiple levels of knowledge points before it can be applied. Structural complexity considers various factors of the logic tree, such as tree depth (the maximum number of levels in the tree), numerator factors (which may be related to the level of detail of the knowledge points or the number of branches), node density (the ratio of the number of nodes to the size of the tree), and cyclicity (whether there are circular references to knowledge points in the problem-solving process). For example, if a standard solution logic tree has a complex structure, a large tree depth, and a high node density, and a certain knowledge point is at a deep level, then the application difficulty of this knowledge point may be relatively high.
[0075] An application difficulty matrix is generated for each historical exam question based on the application difficulty of all knowledge points involved in each standard problem-solving approach. This matrix presents the distribution of application difficulty of each knowledge point in each exam question in a structured way, facilitating the analysis of the overall difficulty characteristics of the exam questions. For example, rows of the matrix can represent different knowledge points, columns can represent different standard problem-solving approaches, and matrix elements are the application difficulty values of the corresponding knowledge points in the corresponding problem-solving approaches.
[0076] Based on the structural correlation analysis of the different knowledge points involved in each standard problem-solving approach of each historical exam question within the corresponding standard problem-solving logic tree, the degree of correlation and application between the different knowledge points involved in each standard problem-solving approach of each historical exam question is analyzed. In this way, the positional relationship and structural characteristics of knowledge points in the logic tree are comprehensively considered, and the closeness of their interrelationship and application in problem-solving is determined.
[0077] A correlation application degree matrix is generated for each historical exam question based on the degree of correlation between different knowledge points involved in each standard problem-solving approach. This matrix can intuitively show the degree of correlation between knowledge points in each exam question, which helps to further understand the interrelationships between knowledge points within the question. For example, the higher the element value corresponding to two knowledge points in the matrix, the closer their correlation in the problem-solving process.
[0078] The application difficulty matrix and the relevance application degree matrix of each exam question are used as the question-setting characteristics of each historical exam question. These two matrices characterize the inherent features of the exam questions from two important aspects: difficulty and knowledge point relevance, providing a key basis for subsequent comprehensive and in-depth analysis and rating of the exam questions.
[0079] In an alternative implementation, the application difficulty of the corresponding knowledge point is assessed based on the hierarchical depth of each knowledge point involved in each standard solution approach for each historical exam question within its respective standard solution logic tree and the structural complexity of the corresponding standard solution logic tree. This includes:
[0080] The structural complexity of a standard problem-solving logic tree is calculated based on its tree depth, numerator factor, node density, and cycle degree. The path length from the root node to the farthest leaf node (i.e., the longest reasoning chain) is also considered. For example, in the case of root → A → B → C, the tree depth H = 3. A greater tree depth generally indicates a more complex problem-solving process. The numerator factor is the average number of child nodes per node (B = total number of child nodes / total number of nodes). Node density is the ratio of the total number of nodes to the tree depth; denser nodes indicate more complex relationships between knowledge points. Cycle degree is the number of loops in the tree (C = 0 for directed acyclic trees, but repeated paths need to be calculated if multiple applications of knowledge points are allowed). Cycle degree increases the complexity of the problem-solving logic tree. By combining these factors, a quantitative value for the structural complexity of a standard problem-solving logic tree can be obtained. For example, for a complex mathematical proof problem, a large tree depth, a numerator factor indicating detailed knowledge points and many branches, high node density, and a small number of circular references all contribute to a higher structural complexity value.
[0081] Based on the level of each knowledge point involved in each standard problem-solving approach of each historical exam question within the knowledge point mastery expansion tree of the corresponding knowledge point module and its level in the corresponding standard logic tree, the weight value of each knowledge point is determined. The level of a knowledge point in these two structures reflects its relative importance in the overall knowledge system and specific problem-solving logic. For example, if a knowledge point is at a high level in the knowledge point mastery expansion tree, it indicates that it is relatively basic or key in the entire knowledge module; at the same time, it is also at a high level in the standard logic tree, indicating that it plays an important guiding role in the specific problem-solving process. By combining these two levels, the weight value of the knowledge point can be determined. If a knowledge point is at a high level in both structures, its weight value is likely to be large. For example, the weight calculation method is: Weight value = (Level of knowledge point mastery expansion tree + Level of standard logic tree) ÷ 2.
[0082] Based on the level depth and corresponding weight value of each knowledge point involved in each standard problem-solving approach for each historical exam question within its respective standard problem-solving logic tree, as well as the structural complexity of the corresponding standard problem-solving logic tree, the application difficulty of the corresponding knowledge point is assessed. Level depth reflects the depth of the knowledge point in the problem-solving logic; a deeper level may mean that multiple levels of knowledge points must be mastered before it can be applied. The weight value reflects the relative importance of the knowledge point; the higher the importance, the greater the impact of its application difficulty on the overall problem-solving difficulty. Structural complexity reflects the overall complexity of the problem-solving logic; the more complex the structure, the greater the application difficulty of each knowledge point may be. For example, if a knowledge point is at a deeper level in a standard problem-solving logic tree with high structural complexity and a large weight value, then its application difficulty will be higher. This comprehensive evaluation method can accurately measure the application difficulty of each knowledge point in the problem-solving process. For example, the calculation formula is: Application Difficulty = Level Depth × Weight Value + (1 - Weight Value) × Structural Complexity.
[0083] In an alternative implementation, the degree of correlation and application among the different knowledge points involved in each standard solution approach for each historical exam question is analyzed based on the structural correlation of these knowledge points within the corresponding standard solution logic tree. This includes:
[0084] Based on the minimum number of edges between each pair of knowledge points involved in each standard solution of each historical exam question, the minimum number of edges between the nearest common ancestor node and the root node, and the overlap between the corresponding subtree structures of the two knowledge points in the standard solution logic tree, a problem-solving logic structure correlation vector for the two knowledge points is generated.
[0085] Based on the minimum number of edges between each pair of knowledge points involved in each standard solution to each historical exam question in the knowledge point mastery extension tree of the smallest range of knowledge point blocks involved in the corresponding historical exam question, the minimum number of edges between the nearest common ancestor node and the root node, and the overlap between the corresponding subtree structures of the two knowledge points in the corresponding knowledge point mastery extension tree, a knowledge point mastery extension structure correlation vector of the two knowledge points is generated.
[0086] For each standard solution approach in each historical exam question, the nearest common ancestor node of each pair of knowledge points in the corresponding standard solution logic tree and the nearest common ancestor node of the knowledge point mastery extension tree in the smallest knowledge point block involved in the corresponding historical exam question are taken as the structural superordinate mapping knowledge point combination of the two nodes, and the knowledge point mastery extension structure correlation vector of the corresponding structural superordinate mapping knowledge point combination is generated as the superordinate knowledge point mastery extension structure correlation vector of the two knowledge points.
[0087] Based on the correlation vectors of the logical structure of each pair of knowledge points involved in each standard problem-solving approach for each historical exam question, the correlation vectors of the knowledge point mastery and extension structure, and the correlation vectors of the mastery and extension structure of the superior knowledge points, the degree of correlation and application between different knowledge points involved in each standard problem-solving approach for each historical exam question is analyzed.
[0088] In an alternative implementation, based on the correlation vector of the problem-solving logic structure of every two knowledge points involved in each standard problem-solving approach for each historical exam question, the correlation vector of the knowledge point mastery and extension structure, and the correlation vector of the mastery and extension structure of higher-level knowledge points, the degree of correlation application between different knowledge points involved in each standard problem-solving approach for each historical exam question is analyzed, including:
[0089] Based on the level of each pair of knowledge points involved in each standard problem-solving approach of each historical exam question in the corresponding standard problem-solving logic tree, the level of knowledge point mastery extension tree in the knowledge point mastery of the smallest knowledge point section involved in the corresponding historical exam question, and the level of the corresponding structural superordinate mapping knowledge point combination in the knowledge point mastery extension tree in the knowledge point mastery of the smallest knowledge point section involved in the corresponding historical exam question, determine the weight of the problem-solving logic structure relevance vector of the two knowledge points, the weight of the knowledge point mastery extension structure relevance vector, and the weight of the superordinate knowledge point mastery extension structure relevance vector.
[0090] Based on the weights of the problem-solving logic structure relevance vector, the knowledge point mastery extension structure relevance vector, and the higher-level knowledge point mastery extension structure relevance vector involved in each standard problem-solving approach for each historical exam question, the problem-solving logic structure relevance vector, knowledge point mastery extension structure relevance vector, and higher-level knowledge point mastery extension structure relevance vector for corresponding two knowledge points are weighted and summed to obtain the degree of correlation and application between different knowledge points involved in each standard problem-solving approach for each historical exam question.
[0091] In this embodiment, a problem-solving logic structure relevance vector is generated:
[0092] Take a physics circuit analysis problem as an example. Suppose that its standard solution logic tree contains two knowledge points: "Ohm's Law" and "characteristics of series circuits".
[0093] Minimum number of edges for a gap: In a standard problem-solving logic tree, connecting the "Ohm's Law" node to the "Characteristics of Series Circuits" node along the tree's edges requires at least 3 edges. Therefore, the minimum number of edges for a gap is 3. Edges represent the logical derivation relationship between knowledge points; the fewer the edges, the closer the two knowledge points are in terms of problem-solving logic.
[0094] The minimum number of edges between the nearest common ancestor node and the root node: Assuming their nearest common ancestor node is "Basic circuit concepts", the minimum number of edges from "Basic circuit concepts" to the root node (such as "Known conditions of the problem") is 2. This value reflects the relative depth of these two knowledge points in the entire problem-solving logic hierarchy.
[0095] The overlap between the subtree structures corresponding to the two knowledge points in their respective standard problem-solving logic trees is analyzed: the subtree structure corresponding to "Ohm's Law" contains some derivation steps related to the calculation of current, voltage, and resistance, while the subtree structure corresponding to "Characteristics of Series Circuits" contains content such as the current and voltage rules in series circuits. Analysis reveals that they overlap to some extent in calculating total resistance and current distribution. Assuming that the overlap is 0.6 (e.g., by comparing the number of nodes in the overlapping part with the total number of nodes in both subtrees), this overlap can be quantified.
[0096] To generate a problem-solving logic structure relevance vector, we can combine these three values into a vector, such as [3, 2, 0.6]. This vector describes the structural relevance of two knowledge points in the standard problem-solving logic tree from different perspectives, providing a foundation for subsequent analysis of the degree of application of the correlation.
[0097] Similarly, for the two knowledge points "Ohm's Law" and "Characteristics of Series Circuits", the analysis is conducted within the knowledge point mastery expansion tree of the smallest knowledge point section (such as the "Circuit Knowledge Section") that corresponds to the historical exam questions.
[0098] Minimum number of edges for a gap: In the knowledge point mastery expansion tree, the distance from "Ohm's Law" to "characteristics of series circuits" requires at least 4 edges. This tree shows the relationship between knowledge points from the perspective of knowledge system expansion, and the number of edges reflects the distance between two knowledge points in the knowledge expansion context.
[0099] The minimum number of edges between the nearest common ancestor node and the root node: Let's assume the nearest common ancestor node in the knowledge point mastery extension tree is "Electricity Fundamentals." The distance from "Electricity Fundamentals" to the root node (e.g., "Basic Concepts of Physics") requires at least 3 edges. This reflects the hierarchical depth of these two knowledge points within a broader knowledge system.
[0100] The overlap between the corresponding subtree structures in the knowledge point mastery extension tree for two knowledge points is as follows: the subtree of "Ohm's Law" in the knowledge point mastery extension tree may include its application extension in different circuit scenarios, and the subtree of "Characteristics of Series Circuits" includes extension content such as the comparison of series circuits with other circuit types. The analysis shows that the overlap is 0.5.
[0101] Then, a knowledge point mastery and expansion structure relevance vector is generated, such as [4,3,0.5]. This vector describes the relevance between two knowledge points from the perspective of knowledge expansion, which helps to fully understand the relationship between knowledge points.
[0102] For "Ohm's Law" and "Characteristics of Series Circuits", the nearest common ancestor node in the standard problem-solving logic tree is "Basic Concepts of Circuits", and the nearest common ancestor node in the knowledge point mastery extension tree is "Electrical Fundamentals". Combining these two nearest common ancestor nodes into a structural superordinate mapping knowledge point combination, namely, Basic Concepts of Circuits and Electrical Fundamentals.
[0103] For this combination, a knowledge point mastery extension structure relevance vector is generated. For example, by analyzing some structural features of "Basic Circuit Concepts" and "Electrical Fundamentals" in the knowledge point mastery extension tree (such as the overlap of their subtree structures, their positional relationship in the tree, etc.), a vector [2,0.7] is assumed to be obtained (the specific meaning of the vector elements and the generation method are determined according to the specific analysis rules). This vector is the higher-level knowledge point mastery extension structure relevance vector corresponding to the two knowledge points. It reflects the connection between the two knowledge points from a more macro-level knowledge structure perspective.
[0104] Analyze the degree of relevance and applicability between different knowledge points:
[0105] We can analyze the correlation and applicability between the two knowledge points, "Ohm's Law" and "characteristics of series circuits," by combining the three vectors generated earlier.
[0106] Suppose we calculate the relevance application degree using a simple weighted summation method. Let the weight of the problem-solving logic structure relevance vector be w1=0.4, the weight of the knowledge point mastery extension structure relevance vector be w2=0.3, and the weight of the higher-level knowledge point mastery extension structure relevance vector be w3=0.3.
[0107] First, the problem-solving logic structure relevance vector [3,2,0.6] is normalized (assuming it becomes [0.3,0.2,0.6] after normalization; the normalization method is determined according to specific needs). The knowledge point mastery extension structure relevance vector [4,3,0.5] is normalized to [0.4,0.3,0.5], and the higher-level knowledge point mastery extension structure relevance vector [2,0.7] is normalized to [0.2,0.7] (this is just an example; the actual normalization method should ensure that each element of the vector is within a reasonable range and is comparable).
[0108] Association application degree = w1×(0.3×a+0.2×b+0.6×c)+w2×(0.4×a+0.3×b+0.5×c)+w3×(0.2×a+0.7×b) (where a, b, and c are coefficients set according to the influence of vector elements on association application degree, assuming a=0.3, b=0.4, c=0.3).
[0109] Substitute into the calculation:
[0110] The correlation vector part of the problem-solving logic structure = 0.4 × (0.3 × 0.3 + 0.2 × 0.4 + 0.6 × 0.3) = 0.4 × (0.09 + 0.08 + 0.18) = 0.4 × 0.35 = 0.14;
[0111] The knowledge point is extended by the structural correlation vector part = 0.3×(0.4×0.3+0.3×0.4+0.5×0.3)=0.3×(0.12+0.12+0.15)=0.3×0.39=0.117;
[0112] The extended structure of the knowledge points is related to the vector part = 0.3 × (0.2 × 0.3 + 0.7 × 0.4) = 0.3 × (0.06 + 0.28) = 0.3 × 0.34 = 0.102;
[0113] The degree of relevance application = 0.14 + 0.117 + 0.102 = 0.359;
[0114] By comprehensively considering the structural relevance represented by different vectors in this way, the degree of application of the connection between two knowledge points can be obtained, thereby gaining a more comprehensive understanding of their close connection in problem-solving and the knowledge system.
[0115] In an alternative implementation, refer to Figure 2 S3: Utilize big data mining algorithms to analyze the answer records and score distribution of each historical exam question in a massive dataset, and consider the academic ability disparity among the respondents to obtain the difficulty coefficient, discrimination index, reliability value, and validity value of each historical exam question as the exam question rating result, including:
[0116] The academic ability stratification model is used to analyze the academic ability-related features of all respondents to all historical exam questions in a large-scale dataset to obtain the academic ability stratification label for each respondent. It is assumed that the large-scale dataset contains a large amount of students' answer information for multiple historical exam questions, and the academic ability-related features include data indicators such as students' past exam scores, homework completion, and classroom performance.
[0117] The academic ability stratification model analyzes this information to categorize students into different levels. For example, the model might categorize students with scores of 90 or above as "high-level academic ability," those with scores between 60 and 89 as "middle-level academic ability," and those with scores below 60 as "low-level academic ability." In this way, each respondent receives an academic ability stratification label that roughly reflects their learning ability level.
[0118] Based on the academic ability stratification model, the academic ability-related characteristics of all respondents to all historical exam questions in the ultra-large-scale dataset are analyzed to calculate the academic ability base of each respondent. Still based on the aforementioned academic ability-related characteristics, the academic ability stratification model further calculates the academic ability base of each respondent. For example, the academic ability base can be calculated using a comprehensive formula, assumed to be: Academic Ability Base = Average Past Exam Score × 0.6 + Homework Completion Quality Score × 0.3 + Classroom Activity Score × 0.1 (the weights here are only examples; actual calculations may follow more complex and reasonable rules).
[0119] Taking a student as an example, whose average past exam score is 85, homework completion quality score is 90, and class participation score is 80, then this student's academic ability base score = 85 × 0.6 + 90 × 0.3 + 80 × 0.1 = 51 + 27 + 8 = 86. The academic ability base score is a quantifiable value that more accurately reflects each student's learning ability.
[0120] Based on the academic ability base of all respondents with the same academic ability stratification label among all respondents to all historical exam questions in the ultra-large-scale dataset, the horizontal alienation index of academic ability for each respondent is calculated. For all respondents to all historical exam questions, the analysis is conducted within the student group with the same academic ability stratification label. For example, in the "high academic ability stratum", there are students A, B, and C, whose respective academic ability bases are 92, 90, and 88.
[0121] Calculate the horizontal differentiation index of academic ability for each student. Assume the calculation method used is to use the average academic ability base of the group as a benchmark to calculate the degree of difference between each student's academic ability base and the average. The average academic ability base of students in the "high academic ability level" is (92+90+88)÷3=90.
[0122] Student A's horizontal differentiation index = |92−90| = 2, Student B's horizontal differentiation index = |90−90| = 0, and Student C's horizontal differentiation index = |88−90| = 2. The horizontal differentiation index reflects the differences in learning abilities among students at the same academic level.
[0123] Based on the academic ability baseline of all respondents to all historical exam questions in the ultra-large-scale dataset, and the maximum and minimum academic ability baselines for each academic ability stratification label, the longitudinal differentiation index of academic ability for each respondent is calculated. The academic ability baseline of all respondents to all historical exam questions, as well as the maximum and minimum academic ability baselines for each academic ability stratification label, are considered. For example, the maximum academic ability baseline for the "high academic ability stratum" is 95, and the minimum is 85; the maximum is 84, and the minimum is 65; and the maximum is 64, and the minimum is 50.
[0124] Taking student A (belonging to the "high academic ability level", with an academic ability base of 92) as an example, calculate their academic ability longitudinal differentiation index. Assume the calculation method is: Academic Ability Longitudinal Differentiation Index = (∣92−95∣ + ∣92−85∣) / (95−50) (here the denominator is the difference between the maximum and minimum academic ability bases among all academic ability levels, used for normalization).
[0125] Therefore, student A's longitudinal differentiation index of academic ability = 453 + 7 ≈ 0.22. The longitudinal differentiation index of academic ability reflects the differences in learning ability among students at different academic levels.
[0126] The horizontal and vertical academic ability differentiation indices of all respondents are used as the base for academic ability differentiation among respondents. Combined with the answer records of each historical test question in the ultra-large-scale dataset (such as how many students answered correctly and incorrectly) and the score distribution (the number of students in each score range), the difficulty coefficient, discrimination index, reliability value and validity value of each historical test question are obtained as the test question rating result for each historical test question.
[0127] In this embodiment, the difficulty coefficient is as follows: For example, for a certain history exam question, 80% of students at the "high academic ability level" answer correctly, 50% at the "middle academic ability level" answer correctly, and 20% at the "low academic ability level" answer correctly. Considering the base of academic ability differentiation among students at different academic ability levels, these correct answer ratios are adjusted (assuming the adjustment method is based on a weighted average of horizontal and vertical academic ability differentiation indices) to finally derive the difficulty coefficient of the exam question. For example, after complex calculations (simplified description here), a difficulty coefficient of 0.6 is obtained, indicating that the question is of moderate difficulty.
[0128] Discrimination Index: This analyzes the relationship between the scores of students at different academic levels and the horizontal differentiation index of academic ability. For example, students with a higher horizontal differentiation index in the "higher academic level" score higher on this question, indicating that the question has a better discrimination effect on students at the higher academic level. A discrimination index is calculated by combining the results of each academic level. If it is 0.7, it indicates that the question can effectively distinguish students with different academic levels.
[0129] Reliability score: The reliability score is calculated by analyzing the stability of scores on this question among students at the same academic level, and by adjusting for the horizontal divergence index of academic ability (for example, if a student at a certain academic level has a large horizontal divergence index, their score will also fluctuate greatly, affecting the reliability calculation). Assuming a final reliability score of 0.8, this indicates that the measurement results for this question are relatively reliable.
[0130] In an alternative implementation, the horizontal and vertical academic ability differentiation indices of all respondents are used as the base for academic ability differentiation among respondents. Combined with the response records and score distribution of each historical exam question in the large-scale dataset, the difficulty coefficient, discrimination index, reliability value, and validity value of each historical exam question are obtained as the exam question rating result for each historical exam question, including:
[0131] Based on the answer records of each historical exam question in a massive dataset, the ratio of the number of students who answered correctly to the total number of students in each academic ability stratum is determined as the original difficulty of the historical exam question for each academic ability stratum. Taking a history math exam question as an example, assuming the answer records for this question in the massive dataset show that 200 students answered in the "high academic ability stratum," with 160 answering correctly; 300 students answered in the "middle academic ability stratum," with 150 answering correctly; and 250 students answered in the "low academic ability stratum," with 50 answering correctly. Then, the original difficulty for the "high academic ability stratum" is 160 ÷ 200 = 0.8; the original difficulty for the "middle academic ability stratum" is 150 ÷ 300 = 0.5; and the original difficulty for the "low academic ability stratum" is 50 ÷ 250 = 0.2. The original difficulty reflects the proportion of students who answered the question correctly under each academic ability stratum.
[0132] Based on the horizontal and vertical divergence indices of academic ability for all respondents under each academic ability stratum label, the original difficulty of the corresponding historical exam questions under the corresponding academic ability stratum label is adjusted to obtain the difficulty coefficient of each historical exam question. Assume that in the "high academic ability stratum," the average horizontal divergence index of all respondents is 3, and the average vertical divergence index is 0.2. Assume a divergence adjustment formula as: Adjustment coefficient = 1 + Average horizontal divergence index × 0.1 + Average vertical divergence index × 0.5. Then the adjustment coefficient for the "high academic ability stratum" is 1 + 3 × 0.1 + 0.2 × 0.5 = 1 + 0.3 + 0.1 = 1.4. The adjusted difficulty coefficient for this stratum is 0.8 × 1.4 = 1.12 (in practical applications, the result may be normalized to a reasonable range, such as converting 1.12 to a value between 0 and 1 using certain rules; this is only an example calculation process). The same method was used to modify the "secondary academic level" and "lower academic level" levels, and the final difficulty coefficient of each history exam question was obtained by combining the situation of each level. This modification takes into account the differences in learning ability among students within the same academic level (horizontal differentiation) and the differences in learning ability among students between different academic levels (vertical differentiation), so that the difficulty coefficient can more accurately reflect the actual difficulty of the exam questions for different students.
[0133] The discrimination index of each historical exam question is calculated based on the score distribution of each historical exam question in the ultra-large-scale dataset and the horizontal alienation index of academic ability of all respondents.
[0134] The reliability value of each historical exam question was calculated based on the score distribution of each historical exam question in the ultra-large-scale dataset and the cross-academic alienation index of all respondents.
[0135] Based on the score distribution of each historical test question in the ultra-large-scale dataset, a knowledge point-differentiation basis matrix is constructed for each knowledge point. Based on the knowledge point-differentiation basis matrix of each knowledge point, the mastery of each knowledge point in different differentiation intervals is calculated. Based on the mastery of each knowledge point in different differentiation intervals, the validity value of each historical test question is calculated.
[0136] For example, in the matrix of knowledge points on "ancient political systems", there are 30 people in the "high-level academic ability" range in the low alienation range, 50 people in the medium alienation range, and 20 people in the high alienation range; there are 20 people in the "medium-level academic ability" range in the low alienation range, 40 people in the medium alienation range, and 40 people in the high alienation range; and there are 10 people in the "low-level academic ability" range in the low alienation range, 30 people in the medium alienation range, and 60 people in the high alienation range.
[0137] Calculate the mastery level of each knowledge point in different differentiation intervals. For example, for the knowledge point of "ancient political systems" in the low differentiation interval of the "high academic level", assuming the calculation method is the number of people who answered correctly in this interval divided by the total number of people in this interval (assuming the number of people who answered correctly in the low differentiation interval is 25), then the mastery level is 25÷30≈0.83. Similarly, calculate the mastery level in other intervals and other academic levels.
[0138] Suppose a formula for calculating validity is as follows: the validity value is the weighted sum of the mastery levels across all divergence intervals for all academic ability strata, where i represents the academic ability stratum, j represents the divergence interval, and wij is the weight (e.g., w11=0.2, w12=0.3, w13=0.1, w21=0.1, w22=0.2, w23=0.1, w31=0.05, w32=0.05, w33=0.1). Substituting these values into the formula yields the validity value. In this way, validity is calculated based on the mastery levels of knowledge points across different divergence intervals to measure whether the test questions accurately assess the expected knowledge and abilities.
[0139] In an alternative implementation, the discrimination index for each historical test question is calculated based on the score distribution of each historical test question in the ultra-large-scale dataset and the cross-academic differentiation index of all respondents, including:
[0140] Based on the score distribution analysis of each historical test question in the ultra-large-scale dataset, the correlation coefficient between the horizontal differentiation index of academic ability for each academic ability stratification label and the score of the corresponding historical test question is used as the intra-stratification discrimination of each historical test question. Assuming that in the ultra-large-scale dataset, for a Chinese history test question, we stratify students according to their academic ability into "high academic ability stratum", "middle academic ability stratum" and "low academic ability stratum".
[0141] For the "higher academic ability level," the correlation coefficient between the horizontal differentiation index of academic ability and the score distribution of students in this level on this exam question is calculated using statistical methods (such as the Pearson correlation coefficient method). For example, calculations show that students in the "higher academic ability level" with higher horizontal differentiation indices also score relatively higher on this exam question, resulting in a correlation coefficient of r1 = 0.7. This r1 represents the intra-level discrimination of the exam question within the "higher academic ability level." It reflects how well students within the same academic ability level differentiate themselves on this exam question due to differences in learning ability (reflected by the horizontal differentiation index). A higher value indicates a better differentiation effect of the exam question on students with different learning abilities within that level.
[0142] Using the same method, calculate the intra-stratum discrimination index r2 for the "high-achieving level" and r3 for the "low-achieving level". Assume that r2 = 0.5 for the "high-achieving level" and r3 = 0.4 for the "low-achieving level".
[0143] Based on the score distribution of each historical test question in the ultra-large-scale dataset, the average score and standard deviation of all respondents with the highest academic ability stratum label and the average score and standard deviation of all respondents with the lowest academic ability stratum label are determined. Then, based on the average score and standard deviation of all respondents with the highest academic ability stratum label and the average score and standard deviation of all respondents with the lowest academic ability stratum label, the cross-stratum discrimination of each historical test question is calculated. Taking this Chinese test question as an example, the average score xhigh and standard deviation shigh of all respondents with the highest academic ability stratum label (let's say "high academic ability stratum") and the average score xlow and standard deviation slow of all respondents with the lowest academic ability stratum label (let's say "low academic ability stratum") are determined from the score distribution.
[0144] Assume the average score of the "high-achieving level" is xhigh=85 points, and the standard deviation of the score is shigh=5; the average score of the "low-achieving level" is xlow=45 points, and the standard deviation of the score is slow=8.
[0145] Cross-stratum discrimination can be calculated using some common formulas, such as: Cross-stratum discrimination D = |xhigh−xlow| / (shigh) 2 +slow 2 ) 1 / 2 .
[0146] Substituting the numerical values: D≈4.25 (in practical applications, the result may be further normalized or adjusted to ensure it falls within a suitable numerical range). Cross-level discrimination measures the difference in scores between students at the highest and lowest academic levels on this test question, reflecting the test question's ability to differentiate students at different academic levels.
[0147] The discrimination index of each historical test question is calculated based on the intra-layer discrimination and cross-layer discrimination of each historical test question in the ultra-large-scale dataset.
[0148] After obtaining the intra-level discrimination (r1, r2, r3) and cross-level discrimination D for each historical test question, they are combined in a certain way to calculate the discrimination index.
[0149] Assuming a weighted average method is used, let the weight of intra-layer discrimination be w1=0.6 and the weight of cross-layer discrimination be w2=0.4. First, calculate the average intra-layer discrimination r=(r1+r2+r3) / 3=30.7+0.5+0.4=0.53.
[0150] Then, the discrimination index DI is calculated as follows: DI = w1 × average value r + w2 × D = 0.6 × 0.53 + 0.4 × 4.25 = 0.318 + 1.7 = 2.018 (In practice, this result may be further processed to bring it into a reasonable range, such as normalizing it to between 0 and 1 to better represent the degree of discrimination). The discrimination index comprehensively considers the differences in test scores among students within the same academic level and between different academic levels, fully reflecting the test questions' ability to distinguish students at different academic levels.
[0151] In an alternative implementation, the reliability value of each historical test item is calculated based on the score distribution of each historical test item in the ultra-large-scale dataset and the cross-sectional heterogeneity index of academic ability of all respondents, including:
[0152] Based on the variance of the scores of all respondents to the corresponding historical exam question for each academic ability stratification label in the ultra-large-scale dataset and the variance of the total score of the corresponding historical exam question for all respondents to the corresponding academic ability stratification label, the stratified Cronbach's coefficient for each academic ability stratification label is calculated. Taking a math history exam question as an example, it is assumed that in the ultra-large-scale dataset, students are divided into "high academic ability level", "medium academic ability level" and "low academic ability level" according to their academic ability.
[0153] For the "higher academic level", let the scores of all respondents in this level for this history exam question be x11, x12, ..., x1n, and the total scores of the corresponding history exam paper be y11, y12, ..., y1n (n is the number of students in this level).
[0154] First, calculate the variance of the score for each history question, sx12, and the variance of the total score for the corresponding paper, sy12.
[0155] According to the Cronbach's coefficient formula α = 1 - (the quotient of the sum of the variances of the scores of all history test questions sx12 and the variance of the total score of the test paper sy12).
[0156] Using the same method, calculate the stratified Cronbach's coefficients α2 and α3 for the "high-level academic ability level" and the "low-level academic ability level" respectively. Assume that α2 is calculated to be 0.78 for the "high-level academic ability level" and α3 is calculated to be 0.75 for the "low-level academic ability level".
[0157] The stratified Cronbach's coefficient of each academic ability stratification label is modified by using the standard deviation of the horizontal alienation index of all respondents for each academic ability stratification label.
[0158] Continuing with the example of the "high academic level", let the horizontal alienation index of academic ability of all respondents in this level be z11, z12, ..., z1n.
[0159] First, calculate the standard deviation of the horizontal alienation index of academic ability;
[0160] The alienation correction formula is: Corrected stratified Cronbach coefficient α1′=α1×(1−σ1 / 10) (Dividing by 10 here is to make the correction range reasonable, it is only an example and can be adjusted in practice).
[0161] In the same manner, the stratified Cronbach's coefficients for the "high-level academic level" and the "low-level academic level" are corrected.
[0162] The minimum value among the modified stratified Cronbach's coefficients of all academic ability stratification labels is used as the reliability value for each historical test question.
[0163] This invention also provides an implementation method for a test question rating system based on deep data mining, including:
[0164] The data integration module is used to integrate historical exam data, detailed student answer data, and teaching resource data, and to clean and normalize them to form a massive dataset containing exam question text, answer records, score distribution, and knowledge point tags.
[0165] The question feature analysis module is used to generate a standard problem-solving logic tree based on the question text and knowledge point tags of each historical question in the ultra-large-scale dataset. Based on the structural features and structural relevance of all knowledge points in the standard problem-solving logic tree, it generates the application difficulty matrix and the correlation application degree matrix of each question as the question features of each historical question.
[0166] The historical exam question rating module uses big data mining algorithms to analyze the answer records and score distribution of each historical exam question in a massive dataset, and takes into account the academic ability differentiation base among the respondents to obtain the difficulty coefficient, discrimination index, reliability value and validity value of each historical exam question as the exam question rating result.
[0167] The rating model building module is used to learn the question-setting features and rating features of each historical exam question in a large-scale dataset using a neural network model, and obtain the exam question rating model.
[0168] The model output rating module is used to obtain the test rating result of the test questions based on the standard problem-solving logic tree and test question rating model.
[0169] This test question rating system based on deep data mining offers numerous advantages. First, the data integration module integrates, cleans, and normalizes historical exam data, detailed student responses, and teaching resource data to form a comprehensive and standardized large-scale dataset. This makes the data used for subsequent analysis more accurate and comprehensive, covering information from the test questions themselves to student responses, laying a solid foundation for accurate test question rating. Second, the question feature analysis module generates standard problem-solving logic trees and derives application difficulty and correlation application matrices, deeply analyzing the characteristics of each historical test question. This analysis from the perspective of knowledge point structure and relevance allows educators to clearly understand the internal logic and difficulty composition of the test questions, helping to optimize question-setting strategies. Third, the historical test question rating module uses big data mining algorithms to comprehensively consider response records, score distribution, and the base of academic ability differentiation, deriving test question rating results such as difficulty coefficient, discrimination index, reliability value, and validity value. This process closely integrates with students' actual situations, making the rating results more reflective of the test questions' effectiveness in actual exams, and providing targeted references for teaching improvement. Fourth, the rating model building module utilizes a neural network model to learn the characteristics of test questions and their rating features, constructing a test question rating model. The powerful learning ability of the neural network model enables it to extract key information from massive amounts of data, possesses good generalization ability, and can adapt to the rating needs of different types of test questions. Fifth, the model output rating module derives the rating results based on the standard problem-solving logic tree of the test questions to be rated and the test question rating model, achieving rapid and scientific evaluation of the test questions to be rated. This provides educators with an efficient and reliable tool in their daily work of test question design and exam arrangement, helping to improve the quality of education and teaching and optimize the examination and assessment system.
[0170] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of this invention and its equivalents, this invention also intends to include these modifications and variations.
Claims
1. A method for evaluating exam questions based on deep data mining, characterized in that, include: S1: Integrate historical exam data, detailed student answer data, and teaching resource data, and clean and normalize them to form a massive dataset containing exam question text, answer records, score distribution, and knowledge point tags; S2: Generate a standard problem-solving logic tree based on the question text and knowledge point tags of each historical question in the ultra-large-scale dataset. Generate the application difficulty matrix and the association application degree matrix of each question based on the structural features and structural correlation of all knowledge points in the standard problem-solving logic tree as the proposition features of each historical question. S3: Utilize big data mining algorithms to analyze the answer records and score distribution of each historical exam question in a massive dataset, and consider the base of academic ability differentiation among the answerers to obtain the difficulty coefficient, discrimination index, reliability value, and validity value of each historical exam question as the exam question rating result; S4: Use a neural network model to learn the question-setting features and question-rating features of each historical exam question in a massive dataset to obtain an exam question rating model; S5: Obtain the test question rating result based on the standard problem-solving logic tree and test question rating model of the test question to be rated; S2: Generates a standard problem-solving logic tree based on the question text and knowledge point tags of each historical question in the ultra-large-scale dataset. Based on the structural features and structural relevance of all knowledge points in the standard problem-solving logic tree, it generates an application difficulty matrix and an application degree matrix for each question as the propositional features of each historical question, including: Based on the question text and standard solution approach for each historical question in the ultra-large-scale dataset, at least one standard solution approach for each historical question is determined. A standard problem-solving logic tree is generated based on the standard problem-solving approach for each historical exam question and the knowledge point tags of all knowledge points involved in the standard problem-solving approach; Based on the level depth of each knowledge point involved in each standard solution approach for each historical exam question and the structural complexity of the corresponding standard solution logic tree, the application difficulty of the corresponding knowledge point is assessed. A difficulty matrix for each historical exam question is generated based on the application difficulty of all knowledge points involved in each standard problem-solving approach. Based on the structural correlation analysis of the different knowledge points involved in each standard solution approach of each historical exam question in the corresponding standard solution logic tree, the degree of correlation and application between the different knowledge points involved in each standard solution approach of each historical exam question is analyzed. A correlation application degree matrix for each historical exam question is generated based on the correlation application degree between different knowledge points involved in each standard problem-solving approach. The application difficulty matrix and the correlation application degree matrix of each test question are used as the question-setting features of each historical test question.
2. The method for evaluating exam questions based on deep data mining according to claim 1, characterized in that, Based on the hierarchical depth of each knowledge point involved in each standard solution approach for each historical exam question within its corresponding standard solution logic tree and the structural complexity of the corresponding standard solution logic tree, the application difficulty of the corresponding knowledge point is assessed, including: The structural complexity of the standard problem-solving logic tree is calculated based on the tree depth, molecular factor, node density, and cycle degree of the standard logic tree. Based on the level of each knowledge point involved in each standard problem-solving approach of each historical exam question in the knowledge point mastery expansion tree of the knowledge point section of the corresponding historical exam question and the level of the corresponding standard logic tree, the weight value of the corresponding knowledge point is determined. Based on the level depth and corresponding weight value of each knowledge point involved in each standard solution approach for each historical exam question in the corresponding standard solution logic tree, as well as the structural complexity of the corresponding standard solution logic tree, the application difficulty of the corresponding knowledge point is assessed.
3. The method for evaluating exam questions based on deep data mining according to claim 1, characterized in that, Based on the structural correlation analysis of the different knowledge points involved in each standard problem-solving approach for each historical exam question within their respective standard problem-solving logic trees, the degree of application of the correlation between the different knowledge points involved in each standard problem-solving approach for each historical exam question is analyzed, including: Based on the minimum number of edges between each pair of knowledge points involved in each standard solution of each historical exam question, the minimum number of edges between the nearest common ancestor node and the root node, and the overlap between the corresponding subtree structures of the two knowledge points in the standard solution logic tree, a problem-solving logic structure correlation vector for the two knowledge points is generated. Based on the minimum number of edges between each pair of knowledge points involved in each standard solution to each historical exam question in the knowledge point mastery extension tree of the smallest range of knowledge point blocks involved in the corresponding historical exam question, the minimum number of edges between the nearest common ancestor node and the root node, and the overlap between the corresponding subtree structures of the two knowledge points in the corresponding knowledge point mastery extension tree, a knowledge point mastery extension structure correlation vector of the two knowledge points is generated. For each standard solution approach in each historical exam question, the nearest common ancestor node of each pair of knowledge points in the corresponding standard solution logic tree and the nearest common ancestor node of the knowledge point mastery extension tree in the smallest knowledge point block involved in the corresponding historical exam question are taken as the structural superordinate mapping knowledge point combination of the two nodes, and the knowledge point mastery extension structure correlation vector of the corresponding structural superordinate mapping knowledge point combination is generated as the superordinate knowledge point mastery extension structure correlation vector of the two knowledge points. Based on the correlation vectors of the logical structure of each pair of knowledge points involved in each standard problem-solving approach for each historical exam question, the correlation vectors of the knowledge point mastery and extension structure, and the correlation vectors of the mastery and extension structure of the superior knowledge points, the degree of correlation and application between different knowledge points involved in each standard problem-solving approach for each historical exam question is analyzed.
4. The method for evaluating exam questions based on deep data mining according to claim 3, characterized in that, Based on the correlation vectors of the logical structure of each pair of knowledge points involved in each standard problem-solving approach for each historical exam question, the correlation vectors of the knowledge point mastery and extension structure, and the correlation vectors of the mastery and extension structure of higher-level knowledge points, the degree of correlation and application between different knowledge points involved in each standard problem-solving approach for each historical exam question is analyzed, including: Based on the level of each pair of knowledge points involved in each standard problem-solving approach of each historical exam question in the corresponding standard problem-solving logic tree, the level of knowledge point mastery extension tree in the knowledge point mastery of the smallest knowledge point section involved in the corresponding historical exam question, and the level of the corresponding structural superordinate mapping knowledge point combination in the knowledge point mastery extension tree in the knowledge point mastery of the smallest knowledge point section involved in the corresponding historical exam question, determine the weight of the problem-solving logic structure relevance vector of the two knowledge points, the weight of the knowledge point mastery extension structure relevance vector, and the weight of the superordinate knowledge point mastery extension structure relevance vector. Based on the weights of the problem-solving logic structure relevance vector, the knowledge point mastery extension structure relevance vector, and the higher-level knowledge point mastery extension structure relevance vector involved in each standard problem-solving approach for each historical exam question, the problem-solving logic structure relevance vector, knowledge point mastery extension structure relevance vector, and higher-level knowledge point mastery extension structure relevance vector for corresponding two knowledge points are weighted and summed to obtain the degree of correlation and application between different knowledge points involved in each standard problem-solving approach for each historical exam question.
5. The method for evaluating exam questions based on deep data mining according to claim 1, characterized in that, S3: Utilize big data mining algorithms to analyze the answer records and score distribution of each historical exam question in a massive dataset, and consider the academic ability disparity among respondents to obtain the difficulty coefficient, discrimination index, reliability value, and validity value of each historical exam question as the exam question rating result, including: By using an academic ability stratification model, we can analyze the academic ability-related characteristics of all respondents to all historical exam questions in a massive dataset and obtain the academic ability stratification label for each respondent. Based on the academic ability stratification model, the academic ability-related characteristics of all respondents to all historical exam questions in the ultra-large-scale dataset are analyzed, and the academic ability base of each respondent is calculated. Based on the academic ability base of all respondents with the same academic ability stratification label among all respondents of all historical exam questions in the ultra-large-scale dataset, the horizontal alienation index of academic ability for each respondent is calculated. Based on the academic ability base of all respondents to all historical exam questions in the ultra-large-scale dataset, the maximum and minimum academic ability base in each academic ability stratification label, the longitudinal alienation index of academic ability for each respondent is calculated; The horizontal and vertical academic ability differentiation indices of all respondents were used as the base for academic ability differentiation among respondents. Combined with the answer records and score distribution of each historical test question in the ultra-large-scale dataset, the difficulty coefficient, discrimination index, reliability value and validity value of each historical test question were obtained as the test question rating result.
6. The method for evaluating exam questions based on deep data mining according to claim 5, characterized in that, The horizontal and vertical academic ability differentiation indices of all respondents were used as the base for academic ability differentiation among respondents. Combined with the answer records and score distribution of each historical exam question in the massive dataset, the difficulty coefficient, discrimination index, reliability value, and validity value of each historical exam question were obtained as the exam question rating results, including: Based on the answer records of each historical exam question in the ultra-large-scale dataset, the ratio of the number of correct answers to the total number of people in each academic ability level label is determined as the original difficulty of the corresponding historical exam question under each academic ability level label. Based on the horizontal and vertical academic ability differentiation indices of all answerers under each academic ability level label, the original difficulty of the corresponding historical exam question under the corresponding academic ability level label is differentiated and corrected to obtain the difficulty coefficient of each historical exam question. The discrimination index of each historical exam question is calculated based on the score distribution of each historical exam question in the ultra-large-scale dataset and the horizontal alienation index of academic ability of all respondents. The reliability value of each historical exam question was calculated based on the score distribution of each historical exam question in the ultra-large-scale dataset and the cross-academic alienation index of all respondents. Based on the score distribution of each historical test question in the ultra-large-scale dataset, a knowledge point-differentiation basis matrix is constructed for each knowledge point. Based on the knowledge point-differentiation basis matrix of each knowledge point, the mastery of each knowledge point in different differentiation intervals is calculated. Based on the mastery of each knowledge point in different differentiation intervals, the validity value of each historical test question is calculated.
7. The method for evaluating exam questions based on deep data mining according to claim 6, characterized in that, Based on the score distribution of each historical exam question in the ultra-large-scale dataset and the horizontal differentiation index of academic ability of all respondents, the discrimination index of each historical exam question is calculated, including: Based on the score distribution analysis of each historical test question in the ultra-large-scale dataset, the correlation coefficient between the horizontal differentiation index of academic ability for each academic ability stratification label and the score of the corresponding historical test question is used as the intra-stratification discrimination of each historical test question. Based on the score distribution of each historical test question in the ultra-large-scale dataset, the average score and standard deviation of all respondents with the highest academic ability stratum label and the average score and standard deviation of all respondents with the lowest academic ability stratum label are determined. Based on the average score and standard deviation of all respondents with the highest academic ability stratum label and the average score and standard deviation of all respondents with the lowest academic ability stratum label, the cross-stratum discrimination of each historical test question is calculated. The discrimination index of each historical test question is calculated based on the intra-layer discrimination and cross-layer discrimination of each historical test question in the ultra-large-scale dataset.
8. The method for evaluating exam questions based on deep data mining according to claim 6, characterized in that, The reliability value of each historical exam question was calculated based on the score distribution of each historical exam question in the ultra-large-scale dataset and the cross-academic alienation index of all respondents, including: Based on the variance of the scores of all respondents to the corresponding historical exam questions for each academic ability stratification label in the ultra-large-scale dataset and the variance of the total score of the corresponding historical exam questions to which all respondents belong for the corresponding academic ability stratification labels, the stratified Cronbach's coefficient for each academic ability stratification label is calculated. The stratified Cronbach's coefficient of each academic ability stratification label is modified by using the standard deviation of the horizontal alienation index of all respondents for each academic ability stratification label. The minimum value among the modified stratified Cronbach's coefficients of all academic ability stratification labels is used as the reliability value for each historical test question.
9. A test question rating system based on deep data mining, characterized in that, include: The data integration module is used to integrate historical exam data, detailed student answer data, and teaching resource data, and to clean and normalize them to form a massive dataset containing exam question text, answer records, score distribution, and knowledge point tags. The question feature analysis module is used to generate a standard problem-solving logic tree based on the question text and knowledge point tags of each historical question in the ultra-large-scale dataset. Based on the structural features and structural relevance of all knowledge points in the standard problem-solving logic tree, it generates the application difficulty matrix and the correlation application degree matrix of each question as the question features of each historical question. The historical exam question rating module uses big data mining algorithms to analyze the answer records and score distribution of each historical exam question in a massive dataset, and takes into account the academic ability differentiation base among the respondents to obtain the difficulty coefficient, discrimination index, reliability value and validity value of each historical exam question as the exam question rating result. The rating model building module is used to learn the question-setting features and rating features of each historical exam question in a large-scale dataset using a neural network model, and obtain the exam question rating model. The model output rating module is used to obtain the rating result of the test questions based on the standard problem-solving logic tree and the test question rating model. Specifically, a standard problem-solving logic tree is generated based on the question text and knowledge point tags of each historical question in the ultra-large-scale dataset. The application difficulty matrix and relevance matrix of each question are generated based on the structural features and structural relevance of all knowledge points in the standard problem-solving logic tree, serving as the question-setting features for each historical question. These features include: Based on the question text and standard solution approach for each historical question in the ultra-large-scale dataset, at least one standard solution approach for each historical question is determined. A standard problem-solving logic tree is generated based on the standard problem-solving approach for each historical exam question and the knowledge point tags of all knowledge points involved in the standard problem-solving approach; Based on the level depth of each knowledge point involved in each standard solution approach for each historical exam question and the structural complexity of the corresponding standard solution logic tree, the application difficulty of the corresponding knowledge point is assessed. A difficulty matrix for each historical exam question is generated based on the application difficulty of all knowledge points involved in each standard problem-solving approach. Based on the structural correlation analysis of the different knowledge points involved in each standard solution approach of each historical exam question in the corresponding standard solution logic tree, the degree of correlation and application between the different knowledge points involved in each standard solution approach of each historical exam question is analyzed. A correlation application degree matrix for each historical exam question is generated based on the correlation application degree between different knowledge points involved in each standard problem-solving approach. The application difficulty matrix and the correlation application degree matrix of each test question are used as the question-setting features of each historical test question.
Citation Information
Patent Citations
Method for question bank quality evaluation
CN104732352A
Test question difficulty estimation method and device, electronic equipment and storage medium
CN111310463A