Proposition examination question rating method and system based on big data deep mining
By integrating and analyzing educational data to generate a standard problem-solving logic tree and an application difficulty matrix, combined with big data mining and neural network models, the shortcomings of existing test question grading technologies are addressed, scientific and accurate grading of test questions is achieved, and the quality of education is improved.
Patent Information
- Application Number
- CN202510576416.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-05-06
AI Technical Summary
The existing test question rating technology has problems such as insufficient data integration and processing capabilities, shallow test question feature mining, and inappropriate rating methods. It is difficult to generate accurate test question rating indicators, which affects the scientific nature and accuracy of education evaluation.
By integrating historical examination data, detailed student answer data and teaching resource data, a standard problem-solving logic tree and application difficulty matrix are generated. By combining big data mining algorithms and neural network models, the base of academic ability alienation is analyzed, and a test question rating model is constructed to achieve scientific rating of test questions.
It improves the accuracy and efficiency of test question rating, and can more accurately reflect the difficulty, discrimination and validity of test questions, provide valuable reference for education and teaching, and optimize teaching and examination arrangements.
Smart Images

Figure CN120632083A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of proposition and assessment technology, and in particular to a proposition and examination question grading method and system based on deep mining of big data. Background Art
[0002] In education, the quality of test questions directly impacts the accuracy and effectiveness of teaching evaluation. Accurately grading test questions helps teachers understand student learning, optimize teaching content and methods, and provide a basis for the rational allocation of educational resources. Traditional test grading methods rely heavily on teachers' experience and judgment, which is highly subjective and difficult to comprehensively and objectively assess all test indicators. With the rapid development of information technology, the application of big data in education is becoming increasingly widespread. Test grading methods and systems based on deep mining of big data have emerged. Leveraging massive amounts of historical test data, detailed student response data, and teaching resource data, these methods, through data processing, feature extraction, and algorithmic analysis, enable a more scientific and comprehensive assessment of key indicators such as test difficulty, discrimination, reliability, and validity. This not only improves the accuracy and objectivity of test grading, but also provides more targeted decision-making support for education and teaching, promoting the development of a scientific and intelligent education evaluation system, and possesses broad application prospects.
[0003] However, existing exam question rating technology suffers from numerous issues. First, data integration and processing capabilities are insufficient. This inadequate integration and processing of historical exam data, student responses, and teaching resources makes it difficult to effectively integrate and process data from past exams, student responses, and teaching resources, hindering the formation of ultra-large datasets and limiting the reliable data foundation for subsequent analysis. Second, the exploration of exam question features is shallow, failing to generate a standard problem-solving logic tree to accurately capture question characteristics. Third, inappropriate rating methods fail to utilize big data algorithms for comprehensive analysis and consider differences in academic ability, making it difficult to generate accurate rating indicators. This, in turn, impacts the rating capabilities of the constructed exam question rating model, making it impossible to scientifically assess new exam questions.
[0004] Therefore, the present invention proposes a method and system for grading test questions based on deep mining of big data. Summary of the Invention
[0005] The present invention provides a method and system for grading test questions based on deep mining of big data, the method comprising: in the data processing link, by integrating historical test data, detailed student answer data and teaching resource data, and performing cleaning and normalization processing, an ultra-large-scale data set is formed. This not only ensures the quality and standardization of the data, but also comprehensively covers all kinds of key information related to the test questions, provides a solid data foundation for subsequent in-depth analysis, and makes the rating results more reliable and comprehensive. In the process of generating propositional features, a standard problem-solving logic tree is generated based on the test question text and knowledge point labels, and the application difficulty matrix and the associated application degree matrix are derived based on this. This deeply explores the knowledge structure and application characteristics within the test questions, depicts the essential characteristics of the test questions from the perspective of propositions, helps to understand and evaluate the test questions more accurately, and provides an important internal basis for subsequent ratings. When obtaining the test question rating results, a big data mining algorithm is used to analyze the answer records and score distribution, and the base number of academic alienation is taken into account. This fully incorporates students' actual responses and comprehensively considers the differences in student abilities. This results in ratings such as difficulty coefficient, discrimination index, reliability, and validity that are more relevant to actual exam scenarios, truly reflecting the effectiveness of the exam questions on students and providing a more valuable reference for teaching and exam evaluation. A neural network model is used to learn question-setting and question-rating characteristics to develop a question-rating model. This model construction utilizes advanced machine learning techniques, which automatically learn underlying patterns from large amounts of data. This model possesses strong generalization capabilities and can effectively rate new exam questions, improving both efficiency and accuracy. Rating results are derived based on the standard solution logic tree for the question to be rated and the question-rating model, enabling rapid and scientific rating of new exam questions. This provides educators with powerful tools for setting and screening questions, as well as evaluating teaching effectiveness. This helps optimize teaching and exam schedules and improve educational quality.
[0006] The present invention provides a method for grading test questions based on deep mining of big data, comprising: S1: Integrate historical exam data, detailed student answer data, and teaching resource data, and perform cleaning and normalization to form a large-scale dataset containing exam text, answer records, score distribution, and knowledge point labels; S2: Generate a standard problem-solving logic tree based on the test text and knowledge point labels of each history test question in the ultra-large-scale dataset. Based on the structural features and structural relevance of all knowledge points in the standard problem-solving logic tree, generate an application difficulty matrix and an associated application degree matrix for each test question as the proposition features of each history test question. S3: Use big data mining algorithms to analyze the answer records and score distribution of each history test question in a large-scale data set, and consider the academic ability difference between the test takers to obtain the difficulty coefficient, discrimination index, reliability value and validity value of each history test question as the test question rating result; S4: Use a neural network model to learn the proposition characteristics and test question rating characteristics of each history test question in a large-scale dataset to obtain a test question rating model; S5: Obtain the test question rating result of the test question to be rated based on the standard problem-solving logic tree of the test question to be rated and the test question rating model.
[0007] Optionally, S2: Generate a standard problem-solving logic tree based on the test text and knowledge point labels of each history test question in the ultra-large-scale dataset, and generate an application difficulty matrix and a correlation application degree matrix for each test question as proposition features of each history test question based on the structural features and structural relevance of all knowledge points in the standard problem-solving logic tree, including: Based on the test text and standard solution ideas of each history test question in the ultra-large-scale data set, at least one standard solution idea for each history test question is determined; Generate a standard solution logic tree based on each standard solution idea for each history test question and the knowledge point labels of all knowledge points involved in the standard solution idea; Based on the hierarchical depth of each knowledge point involved in each standard solution of each history test question in the corresponding standard solution logic tree and the structural complexity of the corresponding standard solution logic tree, the application difficulty of the corresponding knowledge point is evaluated; Generate an application difficulty matrix for each history test question based on the application difficulty of all knowledge points involved in each standard solution approach; Analyze the degree of correlation and application between different knowledge points involved in each standard solution of each history test question based on the structural correlation of different knowledge points involved in each standard solution logic tree; Generate a correlation application matrix for each history test question based on the correlation application degree between different knowledge points involved in each standard solution idea of each history test question; The application difficulty matrix and associated application degree matrix of each test question are used as the proposition characteristics of each history test question.
[0008] Optionally, based on the hierarchical depth of each knowledge point involved in each standard solution idea for each history exam question in the corresponding standard solution logic tree and the structural complexity of the corresponding standard solution logic tree, the application difficulty of the corresponding knowledge point is evaluated, including: Calculate the structural complexity of the standard problem-solving logic tree based on the tree depth, molecular factor, node density, and loop degree of the standard logic tree; Determine the weight of each knowledge point involved in each standard solution of each history test question based on its hierarchical level in the knowledge point mastery expansion tree of the minimum scope of knowledge points covered by the corresponding history test question and its hierarchical level in the corresponding standard logic tree; Based on the hierarchical depth and corresponding weight value of each knowledge point involved in each standard solution idea of each history test question in the corresponding standard solution logic tree, as well as the structural complexity of the corresponding standard solution logic tree, the application difficulty of the corresponding knowledge point is evaluated.
[0009] Optionally, the degree of correlation and application between the different knowledge points involved in each standard solution of each history test question is analyzed based on the structural correlation of the different knowledge points involved in each standard solution of each history test question in the corresponding standard solution logic tree, including: Based on the minimum number of edges between each two knowledge points involved in each standard solution of each history test question in the corresponding standard solution logic tree, the minimum number of edges between the nearest common ancestor node and the root node, and the overlap between the corresponding subtree structures of the two knowledge points in the corresponding standard solution logic tree, a correlation vector of the solution logic structure of the two knowledge points is generated; Based on the minimum number of edges between each two knowledge points involved in each standard solution of each history test question in the knowledge point mastery expansion tree of the minimum range of knowledge points covered by the corresponding history test question, the minimum number of edges between the most recent common ancestor node and the root node, and the overlap between the corresponding subtree structures of the two knowledge points in the corresponding knowledge point mastery expansion tree, a knowledge point mastery expansion structure correlation vector corresponding to the two knowledge points is generated; Treat the nearest common ancestor node of each two knowledge points involved in each standard problem-solving idea of each history test question in the corresponding standard problem-solving logic tree and the nearest common ancestor node in the knowledge point mastery extension tree of the minimum range of knowledge point blocks involved in the corresponding history test question as the structural superordinate mapping knowledge point combination corresponding to the two nodes, and generate the knowledge point mastery extension structural correlation vector of the corresponding structural superordinate mapping knowledge point combination as the superordinate knowledge point mastery extension structural correlation vector corresponding to the two knowledge points; Based on the problem-solving logic structure correlation vector, knowledge point mastery and extension structure correlation vector, and superordinate knowledge point mastery and extension structure correlation vector of every two knowledge points involved in each standard problem-solving approach for each history test question, the correlation application degree between different knowledge points involved in each standard problem-solving approach for each history test question is analyzed.
[0010] Optionally, based on the problem-solving logic structure correlation vector, the knowledge point mastery and extension structure correlation vector, and the superordinate knowledge point mastery and extension structure correlation vector of each two knowledge points involved in each standard problem-solving approach for each history test question, the degree of association and application between different knowledge points involved in each standard problem-solving approach for each history test question is analyzed, including: Based on the hierarchical level of each two knowledge points involved in each standard problem-solving approach for each history test question in the corresponding standard problem-solving logic tree, the hierarchical level of each two knowledge points involved in the knowledge point mastery expansion tree of the minimum scope of knowledge points involved in the corresponding history test question, and the hierarchical level of the corresponding structural superordinate mapping knowledge point combination in the knowledge point mastery expansion tree of the minimum scope of knowledge points involved in the corresponding history test question, determine the weights of the problem-solving logic structure correlation vectors of the corresponding two knowledge points, the weights of the knowledge point mastery expansion structure correlation vectors, and the weights of the superordinate knowledge point mastery expansion structure correlation vectors; Based on the weight of the problem-solving logic structure correlation vector of each two knowledge points involved in each standard problem-solving idea of each history test question, the weight of the knowledge point mastery extension structure correlation vector, and the weight of the superordinate knowledge point mastery extension structure correlation vector, the problem-solving logic structure correlation vector, the knowledge point mastery extension structure correlation vector, and the superordinate knowledge point mastery extension structure correlation vector of the corresponding two knowledge points are weightedly summed to obtain the degree of correlation application between different knowledge points involved in each standard problem-solving idea of each history test question.
[0011] Optionally, S3: Use a big data mining algorithm to analyze the answer records and score distribution of each history test question in a large-scale data set, and consider the base of academic ability heterogeneity among the answerers to obtain the difficulty coefficient, discrimination index, reliability value, and validity value of each history test question as the test question rating result of each history test question, including: Using the academic ability stratification model, we analyze the academic ability-related features of all the historical exam takers in a large-scale dataset and obtain the academic ability stratification label for each respondent. Based on the academic ability stratification model, the academic ability-related characteristics of all respondents in all history exams in the ultra-large-scale dataset are analyzed to calculate the academic ability base of each respondent; Based on the academic ability base of all the answerers of all history test questions in the ultra-large-scale data set and the same academic ability stratification label, the academic ability horizontal alienation index of each answerer is calculated; Based on the academic ability bases of all the respondents for all the history exam questions in the ultra-large-scale dataset, as well as the maximum and minimum academic ability bases in each academic ability stratification label, the vertical alienation index of academic ability of each respondent is calculated. The horizontal alienation index and vertical alienation index of academic ability of all answering subjects are used as the base of academic ability alienation between answering subjects, and combined with the answer records and score distribution of each history test question in the ultra-large-scale data set, the difficulty coefficient, discrimination index, reliability value and validity value of each history test question are obtained as the test question rating result of each history test question.
[0012] Optionally, the horizontal and vertical alienation indices of academic ability of all answering subjects are used as the base of academic ability alienation between answering subjects, and combined with the answer records and score distribution of each history test question in the ultra-large-scale data set, the difficulty coefficient, discrimination index, reliability value, and validity value of each history test question are obtained as the test question rating result of each history test question, including: Based on the answer records of each history test question in the ultra-large-scale dataset, the ratio of the number of correct answers under each academic stratification label to the total number of people in the corresponding layer is determined as the original difficulty of the corresponding history test question under each academic stratification label. Based on the academic stratification horizontal and vertical alienation indices of all answer subjects under each academic stratification label, the original difficulty of the corresponding history test question under the corresponding academic stratification label is subjected to alienation correction to obtain the difficulty coefficient of each history test question; The discrimination index of each history test question is calculated based on the score distribution of each history test question in the ultra-large-scale data set and the horizontal alienation index of the academic ability of all test takers; The reliability value of each history test question is calculated based on the score distribution of each history test question in the ultra-large-scale data set and the horizontal alienation index of the academic ability of all test takers; Based on the score distribution of each history test question in the ultra-large-scale data set, a knowledge point-alienation basis matrix for each knowledge point is constructed. Based on the knowledge point-alienation basis matrix for each knowledge point, the mastery degree of each knowledge point in different alienation intervals is calculated. Based on the mastery degree of each knowledge point in different alienation intervals, the validity value of each history test question is calculated.
[0013] Optionally, the discrimination index of each history test question is calculated based on the score distribution of each history test question in the ultra-large-scale data set and the horizontal alienation index of the academic ability of all answering subjects, including: Based on the score distribution of each history test question in the ultra-large-scale data set, the correlation coefficient between the academic ability horizontal alienation index of each academic ability stratification label and the corresponding history test score is used as the intra-stratum discrimination of each history test question; Based on the score distribution of each history test question in the ultra-large-scale dataset, the average score and score standard deviation of all respondents with the highest academic ability stratification label and the average score and score standard deviation of all respondents with the lowest academic ability stratification label are determined. The cross-level discrimination of each history test question is calculated based on the average score and score standard deviation of all respondents with the highest academic ability stratification label and the average score and score standard deviation of all respondents with the lowest academic ability stratification label. The discrimination index of each history test question is calculated based on the intra-layer discrimination and cross-layer discrimination of each history test question in the ultra-large-scale data set.
[0014] Optionally, the reliability value of each history test question is calculated based on the score distribution of each history test question in the ultra-large-scale data set and the horizontal alienation index of the academic ability of all test takers, including: Based on the variance of the scores of all respondents for each academic ability stratification label for each history test question in the ultra-large-scale dataset and the variance of the total scores of the corresponding history test question for all respondents for the corresponding academic ability stratification label, the hierarchical Cronbach's coefficient for each academic ability stratification label is calculated. The standard deviation of the horizontal alienation index of the academic ability of all the respondents of each academic ability stratification label is used to perform alienation correction on the stratified Cronbach's coefficient of each academic ability stratification label to obtain the modified stratified Cronbach's coefficient of each academic ability stratification label; The minimum value of the modified stratified Cronbach's coefficient of all academic ability stratification labels is used as the reliability value of each history test question.
[0015] The present invention provides a question-setting and examination question rating system based on deep mining of big data, comprising: The data integration module is used to integrate historical exam data, detailed student answer data, and teaching resource data, and perform cleaning and normalization processing to form a large-scale data set containing exam text, answer records, score distribution, and knowledge point labels; The question feature analysis module is used to generate a standard solution logic tree based on the question text and knowledge point labels of each history question in the ultra-large-scale dataset. Based on the structural features and structural relevance of all knowledge points in the standard solution logic tree, the application difficulty matrix and associated application degree matrix of each question are generated as the question feature of each history question. The history exam question rating module uses big data mining algorithms to analyze the answer records and score distribution of each history exam question in a large-scale data set, and considers the base of academic ability alienation among the answerers to obtain the difficulty coefficient, discrimination index, reliability value, and validity value of each history exam question as the exam question rating result; The rating model building module is used to use a neural network model to learn the proposition characteristics and rating characteristics of each history exam question in a large-scale dataset to obtain an exam question rating model; The model output rating module is used to obtain the test question rating results of the test question to be rated based on the standard problem-solving logic tree of the test question to be rated and the test question rating model.
[0016] The beneficial effects of the present invention compared to the prior art are as follows: in the data processing link, by integrating historical test data, students' detailed answer data and teaching resource data, and performing cleaning and normalization processing, an ultra-large-scale data set is formed. This not only ensures the quality and standardization of the data, but also comprehensively covers all kinds of key information related to the test questions, providing a solid data foundation for subsequent in-depth analysis, and making the rating results more reliable and comprehensive. In the process of generating proposition features, a standard problem-solving logic tree is generated based on the test question text and knowledge point labels, and the application difficulty matrix and the associated application degree matrix are derived based on this. This deeply explores the knowledge structure and application characteristics within the test questions, and depicts the essential characteristics of the test questions from the perspective of the question proposition, which helps to understand and evaluate the test questions more accurately, and provides an important internal basis for subsequent ratings. When obtaining the test question rating results, a big data mining algorithm is used to analyze the answer records and score distribution, and the cardinality of academic alienation is taken into account. This fully incorporates students' actual responses and comprehensively considers the differences in student abilities. This results in ratings such as difficulty coefficient, discrimination index, reliability, and validity that are more relevant to actual exam scenarios, truly reflecting the effectiveness of the exam questions on students and providing a more valuable reference for teaching and exam evaluation. A neural network model is used to learn question-setting and question-rating characteristics to develop a question-rating model. This model construction utilizes advanced machine learning techniques, which automatically learn underlying patterns from large amounts of data. This model possesses strong generalization capabilities and can effectively rate new exam questions, improving both efficiency and accuracy. Rating results are derived based on the standard solution logic tree for the question to be rated and the question-rating model, enabling rapid and scientific rating of new exam questions. This provides educators with powerful tools for setting and screening questions, as well as evaluating teaching effectiveness. This helps optimize teaching and exam schedules and improve educational quality.
[0017] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present invention. The purpose and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in this application document.
[0018] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings: Figure 1 This is a flow chart of a method for grading test questions based on deep mining of big data in an embodiment of the present invention; Figure 2This is a design diagram of the distributed computing architecture of the proposition and examination question grading method based on deep mining of big data in an embodiment of the present invention. DETAILED DESCRIPTION
[0020] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0021] refer to Figure 1 The present invention provides an implementation of a method for grading test questions based on deep mining of big data, including: S1: Integrate historical exam data, detailed student response data, and teaching resource data, and perform cleaning and normalization to create a large-scale dataset containing exam text, response records, score distribution, and knowledge point labels. This integration of historical exam data, detailed student response data, and teaching resource data comes from a wide range of sources, covering various information about past exams, students' specific responses, and relevant teaching materials. Cleaning removes noise, errors, or duplication from the data, such as deleting obviously incorrect answers from response records and correcting typos in exam text. Normalization converts data of varying magnitudes or units to a unified scale, such as standardizing scores across different exams to a range of 0-100 for subsequent analysis. These operations create a large-scale dataset containing exam text (i.e., question content), response records (student responses to each question), score distribution (the distribution of students in different score ranges), and knowledge point labels (labeling the knowledge points covered by each question). For example, a piece of mathematics history test data includes the text of a geometry question, the students' answers to the question, the distribution of scores for the question among all students, and the labels of knowledge points such as triangles and similarity involved in the question.
[0022] S2: Generate a standard problem-solving logic tree based on the question text and knowledge point labels of each history test question in the ultra-large-scale dataset. Generate an application difficulty matrix and an associated application degree matrix for each test question based on the structural characteristics and structural correlations of all knowledge points in the standard problem-solving logic tree as the propositional characteristics of each history test question. Generate a standard problem-solving logic tree based on the question text and knowledge point labels of each history test question in the ultra-large-scale dataset. This is like building a logical framework for solving a problem, showing the various steps required from the question conditions to the answer and the logical relationship between the knowledge points involved. For example, for a physics mechanics problem, the standard problem-solving logic tree may start with analyzing the forces acting on the object, gradually deriving to the application of Newton's laws for calculation, clearly showing the position and role of each knowledge point in the problem-solving process.
[0023] Then, based on the structural characteristics of all knowledge points in the standard problem-solving logic tree (such as the level at which the knowledge point is located, how it connects to other knowledge points, etc.) and structural relevance (the degree of closeness between different knowledge points and the mutual influence relationship), an application difficulty matrix and an associated application degree matrix are generated for each test question, which serve as the proposition characteristics of each history test question. The application difficulty matrix reflects the difficulty of applying each knowledge point in the problem-solving process. For example, if a knowledge point is at a deeper level in the logic tree, it may mean that its application is more difficult, and the corresponding value in the matrix will reflect this feature. The associated application degree matrix reflects the degree to which different knowledge points are associated with each other when solving problems. For example, if two knowledge points are closely connected in the logic tree and are often used together in the problem-solving process, then their corresponding values in the associated application degree matrix will be higher.
[0024] S3: Using big data mining algorithms to analyze the answer records and score distributions for each history test question in a large-scale dataset, and taking into account the variation in academic ability among test takers, we determine the difficulty coefficient, discrimination index, reliability value, and validity value for each history test question, which serve as the test rating for each history test question. The variation in academic ability reflects differences in learning ability and knowledge mastery among students. For example, some students may have a solid foundation and strong learning ability, while others may be relatively weak. This difference is reflected in the variation in academic ability. Through comprehensive analysis, we determine the difficulty coefficient (measurement of the test question's difficulty for students), discrimination index (ability to distinguish between students of different academic ability levels), reliability value (reliability of the test question's measurement results), and validity value (whether the test question accurately tests the intended knowledge and abilities). These values serve as the test rating for each history test question.
[0025] S4: Use a neural network model to learn the propositional characteristics and rating characteristics of each history exam question in the ultra-large dataset to obtain a question rating model. Use a neural network model to learn the propositional characteristics (i.e., the characteristics represented by the application difficulty matrix and the associated application degree matrix generated previously) and the rating characteristics (characteristics reflected by the difficulty coefficient, discrimination index, reliability value, and validity value) of each history exam question in the ultra-large dataset. Neural network models have powerful learning capabilities and can automatically extract regularities and patterns from large amounts of data. By learning these characteristics, a question rating model is constructed. This model acts as an intelligent tool, capable of assigning corresponding ratings based on the input information related to the exam question.
[0026] S5: Obtain the test rating results of the test questions to be rated based on the standard solution logic tree of the test questions to be rated and the test question rating model. For the test questions to be rated, first generate its standard solution logic tree, and then combine this logic tree with the test question rating model obtained previously. The test question rating model analyzes the test questions to be rated based on the information contained in the standard solution logic tree of the test questions to be rated, and applies the previously learned rules to finally obtain the test rating results of the test questions to be rated, thereby achieving rapid and scientific rating of new test questions. For example, for a new chemistry test question, by generating its standard solution logic tree and inputting it into the model, the model can give rating results such as the difficulty coefficient and discrimination index of the question, helping educators understand the quality and scope of application of the question.
[0027] In an alternative embodiment, S2: generating a standard problem-solving logic tree based on the test text and knowledge point labels of each history test question in the ultra-large-scale dataset, and generating an application difficulty matrix and a correlation application degree matrix for each test question as proposition features for each history test question based on the structural features and structural relevance of all knowledge points in the standard problem-solving logic tree, including: Based on the text and standard solutions for each history question in the ultra-large dataset, we identified at least one standard solution for each question. For example, a math proof question might have both a conventional solution and a clever, simplified solution, resulting in at least two standard solutions. This step forms the foundation for the subsequent construction of a standard solution logic tree, ensuring a clear understanding of the solutions for each question.
[0028] A standard problem-solving logic tree is generated based on the knowledge point labels of each standard solution for each history exam question and all the knowledge points involved in the standard solution. For example, when solving a physics circuit problem, the solution might be to first analyze the circuit connections, then apply knowledge points such as Ohm's Law to calculate the current and voltage of each component. Based on this idea and the knowledge point labels such as "Circuit Connection" and "Ohm's Law", the constructed logic tree will display the order and interrelationships of these knowledge points in the problem-solving process in a hierarchical structure, such as "Circuit Connection" as the top node, and then the calculation steps using lower nodes such as "Ohm's Law" are derived from it.
[0029] The difficulty of applying each knowledge point involved in each standard solution to each history exam question is assessed based on the hierarchical depth of each knowledge point in the corresponding standard solution logic tree and the structural complexity of the corresponding standard solution logic tree. Hierarchical depth refers to the number of levels a knowledge point is from the root node in the logic tree. More levels may mean greater application difficulty. For example, a knowledge point at the bottom of a complex logic tree may require mastering the knowledge points in the previous layers before it can be applied. Structural complexity comprehensively considers various factors of the logic tree, such as tree depth (the maximum number of levels in the tree), numerator (which may be related to the degree of refinement of the knowledge point or the number of branches), node density (the proportional relationship between the number of nodes and the size of the tree), and circularity (whether there are circular references to knowledge points during the problem-solving process, etc.). For example, if a standard solution logic tree has a complex structure, a large tree depth, and high node density, and a knowledge point is at a deeper level, then the application difficulty of this knowledge point may be higher.
[0030] Based on the application difficulty of all knowledge points involved in each standard solution for each history exam question, an application difficulty matrix is generated for each question. This matrix presents the distribution of application difficulty of each knowledge point in a structured manner, facilitating analysis of the overall difficulty characteristics of the exam question. For example, the rows of the matrix can represent different knowledge points, the columns represent different standard solution strategies, and the matrix elements represent the application difficulty values of the corresponding knowledge points in the corresponding solution strategies.
[0031] Based on the structural correlation of the different knowledge points involved in each standard solution of each history test question in the corresponding standard solution logic tree, the degree of correlation and application between the different knowledge points involved in each standard solution of each history test question is analyzed; in this way, the positional relationship and structural characteristics of the knowledge points in the logic tree are comprehensively considered to determine the degree of their mutual correlation and application when solving the problem.
[0032] Based on the degree of interdependence between the different knowledge points involved in each standard solution for each history exam question, a matrix of interdependence is generated for each question. This matrix visually demonstrates the degree of interdependence between the knowledge points within each question, helping to further understand the interrelationships between knowledge points within the question. For example, the higher the value of the element corresponding to two knowledge points in the matrix, the closer their interdependence is in the problem-solving process.
[0033] Each history exam question is characterized by its application difficulty matrix and relevance matrix. These two matrices capture the inherent characteristics of the exam question from two key perspectives: difficulty and relevance. This provides a key basis for comprehensive and in-depth analysis and rating of the questions.
[0034] In an alternative embodiment, based on the hierarchical depth of each knowledge point involved in each standard solution of each history exam question in the corresponding standard solution logic tree and the structural complexity of the corresponding standard solution logic tree, the application difficulty of the corresponding knowledge point is evaluated, including: The structural complexity of a standard problem-solving logic tree is calculated based on its tree depth, numerator, node density, and cyclicity. This is the length of the path from the root node to the furthest leaf node (i.e., the longest chain of reasoning). For example, for the case of root → A → B → C, the tree depth is H = 3. A deeper tree generally indicates a more complex problem-solving process. The numerator is the average number of child nodes per node (B = total number of child nodes / total number of nodes). Node density is the ratio of total number of nodes to tree depth; denser nodes indicate more complex relationships between knowledge points. Circularity is the number of loops in the tree (C = 0 for a directed acyclic tree, but if knowledge points are allowed to be applied multiple times, repeated paths must be calculated). This indicates an increase in the complexity of the problem-solving logic. By combining these factors, a quantitative structural complexity value for the standard problem-solving logic tree can be derived. For example, for a standard problem-solving logic tree for a complex mathematical proof, a high tree depth, a high numerator, and a high node density indicate detailed knowledge points with many branches, a high node density, and a small number of circular references. These factors combined result in a higher structural complexity value.
[0035] The weight of each knowledge point involved in each standard solution for each history exam question is determined based on its hierarchical level in the knowledge point mastery expansion tree and its hierarchical level in the corresponding standard logic tree for the minimum scope of the knowledge point section covered by the corresponding history exam question. The hierarchical level of a knowledge point in these two structures reflects its relative importance within the overall knowledge system and the specific problem-solving logic. For example, if a knowledge point is at a higher level in the knowledge point mastery expansion tree, it indicates that it is fundamental or critical to the entire knowledge section; at the same time, it is at a higher level in the standard logic tree, indicating that it plays an important guiding role in the specific problem-solving process. By combining these two hierarchical levels, the weight of the knowledge point can be determined. If a knowledge point is at a higher level in both structures, its weight is likely to be higher. For example, the weight is calculated as weight = (hierarchical level in the knowledge point mastery expansion tree + hierarchical level in the standard logic tree) ÷ 2.
[0036] The difficulty of applying each knowledge point involved in each standard solution for each history exam question is assessed based on its hierarchical depth, corresponding weight, and the structural complexity of the corresponding standard solution logic tree. Hierarchical depth reflects the depth of the knowledge point within the problem-solving logic. A deeper level may mean that it requires mastering knowledge points from multiple levels before it can be applied. The weight reflects the relative importance of the knowledge point. The higher the importance, the greater the impact its application difficulty may have on the overall problem-solving difficulty. Structural complexity reflects the overall complexity of the problem-solving logic. The more complex the structure, the greater the difficulty of applying each knowledge point. For example, if a knowledge point is at a deeper level and has a larger weight in a standard solution logic tree with high structural complexity, its application difficulty will be higher. This comprehensive assessment method can more accurately measure the difficulty of applying each knowledge point in the problem-solving process. For example, the calculation formula is: Application Difficulty = Hierarchical Depth × Weight + (1 - Weight) × Structural Complexity.
[0037] In an alternative embodiment, the degree of correlation and application between the different knowledge points involved in each standard solution of each history test question is analyzed based on the structural correlation of the different knowledge points involved in each standard solution of each history test question in the corresponding standard solution logic tree, including: Based on the minimum number of edges between each two knowledge points involved in each standard solution of each history test question in the corresponding standard solution logic tree, the minimum number of edges between the nearest common ancestor node and the root node, and the overlap between the corresponding subtree structures of the two knowledge points in the corresponding standard solution logic tree, a correlation vector of the solution logic structure of the two knowledge points is generated; Based on the minimum number of edges between each two knowledge points involved in each standard solution of each history test question in the knowledge point mastery expansion tree of the minimum range of knowledge points covered by the corresponding history test question, the minimum number of edges between the most recent common ancestor node and the root node, and the overlap between the corresponding subtree structures of the two knowledge points in the corresponding knowledge point mastery expansion tree, a knowledge point mastery expansion structure correlation vector corresponding to the two knowledge points is generated; Treat the nearest common ancestor node of each two knowledge points involved in each standard problem-solving idea of each history test question in the corresponding standard problem-solving logic tree and the nearest common ancestor node in the knowledge point mastery extension tree of the minimum range of knowledge point blocks involved in the corresponding history test question as the structural superordinate mapping knowledge point combination corresponding to the two nodes, and generate the knowledge point mastery extension structural correlation vector of the corresponding structural superordinate mapping knowledge point combination as the superordinate knowledge point mastery extension structural correlation vector corresponding to the two knowledge points; Based on the problem-solving logic structure correlation vector, knowledge point mastery and extension structure correlation vector, and superordinate knowledge point mastery and extension structure correlation vector of every two knowledge points involved in each standard problem-solving approach for each history test question, the correlation application degree between different knowledge points involved in each standard problem-solving approach for each history test question is analyzed.
[0038] In an alternative embodiment, based on the problem-solving logic structure correlation vector, the knowledge point mastery and extension structure correlation vector, and the superordinate knowledge point mastery and extension structure correlation vector of each two knowledge points involved in each standard problem-solving approach for each history test question, the degree of association and application between different knowledge points involved in each standard problem-solving approach for each history test question is analyzed, including: Based on the hierarchical level of each two knowledge points involved in each standard problem-solving approach for each history test question in the corresponding standard problem-solving logic tree, the hierarchical level of each two knowledge points involved in the knowledge point mastery expansion tree of the minimum scope of knowledge points involved in the corresponding history test question, and the hierarchical level of the corresponding structural superordinate mapping knowledge point combination in the knowledge point mastery expansion tree of the minimum scope of knowledge points involved in the corresponding history test question, determine the weights of the problem-solving logic structure correlation vectors of the corresponding two knowledge points, the weights of the knowledge point mastery expansion structure correlation vectors, and the weights of the superordinate knowledge point mastery expansion structure correlation vectors; Based on the weight of the problem-solving logic structure correlation vector of each two knowledge points involved in each standard problem-solving idea of each history test question, the weight of the knowledge point mastery extension structure correlation vector, and the weight of the superordinate knowledge point mastery extension structure correlation vector, the problem-solving logic structure correlation vector, the knowledge point mastery extension structure correlation vector, and the superordinate knowledge point mastery extension structure correlation vector of the corresponding two knowledge points are weightedly summed to obtain the degree of correlation application between different knowledge points involved in each standard problem-solving idea of each history test question.
[0039] In this embodiment, a problem-solving logic structure correlation vector is generated: Taking a physics circuit analysis problem as an example, suppose that its standard solution logic tree includes two knowledge points: "Ohm's law" and "Characteristics of series circuits".
[0040] Minimum number of edges between nodes: In a standard problem-solving logic tree, a path from the "Ohm's Law" node to the "Characteristics of Series Circuits" node passes through at least three edges, so the minimum number of edges between nodes is three. Edges represent the logical deduction relationship between knowledge points. The fewer edges there are, the closer the two knowledge points are in terms of problem-solving logic.
[0041] Minimum number of edges between the LCA node and the root node: Assuming their LCA node is "Basic Circuit Concepts," there are at least two edges from "Basic Circuit Concepts" to the root node (e.g., "Known Conditions in the Problem"). This value reflects the relative depth of these two knowledge points in the entire problem-solving logic hierarchy.
[0042] The degree of overlap between the subtree structures corresponding to two knowledge points in their respective standard problem-solving logic trees: The subtree structure corresponding to "Ohm's Law" contains some derivation steps related to current, voltage, and resistance calculations, while the subtree structure corresponding to "Characteristics of Series Circuits" covers the current and voltage patterns in series circuits. Analysis revealed some overlap in calculating total resistance and current distribution. Assume that the degree of overlap is 0.6, as determined by some quantitative method (for example, by comparing the number of nodes in the overlapping part with the total number of nodes in the two subtrees).
[0043] To generate a problem-solving logic structure relevance vector, we can combine these three values into a vector, such as [3, 2, 0.6]. This vector describes the structural relevance of the two knowledge points in the standard problem-solving logic tree from different perspectives, providing a basis for subsequent analysis of the degree of associated application.
[0044] Similarly, for the two knowledge points of "Ohm's Law" and "Characteristics of Series Circuits", they are analyzed in the knowledge point mastery and expansion tree of the smallest scope of knowledge point sections (such as "Circuit Knowledge Section") covered by the corresponding historical test questions.
[0045] Minimum number of edges between two knowledge points: In the knowledge point mastery and expansion tree, there are at least four edges from "Ohm's Law" to "Characteristics of Series Circuits." This tree displays the relationship between knowledge points from the perspective of knowledge system expansion, and the number of edges reflects the distance between two knowledge points in the knowledge expansion context.
[0046] Minimum number of edges between the most recent common ancestor and the root node: For example, if their most recent common ancestor in the knowledge point mastery tree is "Electrical Basics," there are at least three edges from "Electrical Basics" to the root node (e.g., "Basic Concepts of Physics"). This reflects the depth of the hierarchy of these two knowledge points within the broader knowledge system.
[0047] The degree of overlap between the corresponding subtree structures of the two knowledge points in the corresponding knowledge point mastery expansion tree: the subtree of "Ohm's Law" in the knowledge point mastery expansion tree may include its application expansion in different circuit scenarios, and the subtree of "Characteristics of Series Circuits" includes expansion content such as the comparison between series circuits and other circuit types. After analysis, the overlap is 0.5.
[0048] Then, a knowledge point mastery and extension structure correlation vector is generated, such as [4, 3, 0.5]. This vector describes the correlation between two knowledge points from the perspective of knowledge extension, which helps to fully understand the relationship between knowledge points.
[0049] For "Ohm's Law" and "Characteristics of Series Circuits", the nearest common ancestor node in the standard problem-solving logic tree is "Basic Concepts of Circuits", and the nearest common ancestor node in the knowledge point mastery and expansion tree is "Electrical Foundations". These two nearest common ancestor nodes are combined into a structural superordinate mapping knowledge point combination, namely the Basic Concepts of Circuits and the Electrical Foundations.
[0050] Generate a knowledge point mastery expansion structure correlation vector for this combination. For example, by analyzing some structural features of "Basic Circuit Concepts" and "Electrical Fundamentals" in the knowledge point mastery expansion tree (such as the overlap of their subtree structures and their positional relationships within the tree), we hypothetically derive a vector [2, 0.7] (the specific meaning and generation method of the vector elements here are determined by specific analysis rules). This vector is the superordinate knowledge point mastery expansion structure correlation vector for the two knowledge points, reflecting the connection between the two knowledge points from a more macroscopic knowledge structure perspective.
[0051] Analyze the correlation and application between different knowledge points: The three vectors generated previously are used to analyze the correlation and application between the two knowledge points of "Ohm's law" and "Characteristics of series circuits".
[0052] Assume that we calculate the degree of associated application through a simple weighted summation method, and set the weight of the problem-solving logical structure correlation vector to w1=0.4, the weight of the knowledge point mastery and extension structure correlation vector to w2=0.3, and the weight of the upper-level knowledge point mastery and extension structure correlation vector to w3=0.3.
[0053] First, the problem-solving logic structure correlation vector [3,2,0.6] is normalized (assuming it is [0.3,0.2,0.6] after normalization, and the normalization method is determined according to specific needs). The knowledge point mastery and extension structure correlation vector [4,3,0.5] is normalized to [0.4,0.3,0.5], and the upper-level knowledge point mastery and extension structure correlation vector [2,0.7] is normalized to [0.2,0.7] (this is only an example, the actual normalization method must ensure that each element of the vector is within a reasonable range and is comparable).
[0054] Association application degree = w1×(0.3×a+0.2×b+0.6×c)+w2×(0.4×a+0.3×b+0.5×c)+w3×(0.2×a+0.7×b) (where a, b, and c are coefficients set according to the influence of vector elements on the association application degree, assuming a=0.3, b=0.4, and c=0.3).
[0055] Substitute into the calculation: The correlation vector of the logical structure of the problem-solving = 0.4 × (0.3 × 0.3 + 0.2 × 0.4 + 0.6 × 0.3) = 0.4 × (0.09 + 0.08 + 0.18) = 0.4 × 0.35 = 0.14; Knowledge point mastery expansion structure correlation vector part = 0.3 × (0.4 × 0.3 + 0.3 × 0.4 + 0.5 × 0.3) = 0.3 × (0.12 + 0.12 + 0.15) = 0.3 × 0.39 = 0.117; The upper knowledge point mastery of the extended structure correlation vector part = 0.3 × (0.2 × 0.3 + 0.7 × 0.4) = 0.3 × (0.06 + 0.28) = 0.3 × 0.34 = 0.102; Correlation application degree = 0.14 + 0.117 + 0.102 = 0.359; In this way, by comprehensively considering the structural correlations represented by different vectors, the degree of association between two knowledge points can be obtained, thereby gaining a more comprehensive understanding of the degree of their association in problem solving and knowledge systems.
[0056] In an alternative embodiment, reference Figure 2 S3: Use big data mining algorithms to analyze the answer records and score distribution of each history test question in the ultra-large-scale data set, and consider the academic ability alienation base between the answer subjects to obtain the difficulty coefficient, discrimination index, reliability value and validity value of each history test question as the test question rating result of each history test question, including: The academic ability stratification model is used to analyze the academic ability-related characteristics of all answerers of all history test questions in a super-large-scale dataset to obtain the academic ability stratification label of each answerer; it is assumed that the super-large-scale dataset contains a large number of students' answer information on multiple history test questions, where the academic ability-related characteristics include students' past test scores, homework completion status, classroom performance and other data indicators.
[0057] The academic ability stratification model analyzes this information to categorize students into different levels. For example, based on the average scores of past exams, the model might classify students with scores of 90 or above as "high academic ability," those with scores of 60-89 as "medium academic ability," and those with scores below 60 as "low academic ability." This gives each respondent a specific academic ability label, which roughly reflects their learning ability.
[0058] The academic stratification model analyzes the academic stratification characteristics of all history exam takers in a large-scale dataset to calculate each respondent's academic stratification base. Still based on these academic stratification characteristics, the model further calculates each respondent's academic stratification base. For example, the academic stratification base can be calculated using a comprehensive formula: Academic stratification base = Average past exam score × 0.6 + Assignment quality score × 0.3 + Classroom activity score × 0.1 (the weights used here are only examples; actual calculations may follow more complex and rational rules).
[0059] For example, let's say a student has an average score of 85 on past exams, a score of 90 for homework quality, and an 80 for class activity. Then, their academic ability base = 85 × 0.6 + 90 × 0.3 + 80 × 0.1 = 51 + 27 + 8 = 86. The academic ability base is a quantitative value that more accurately reflects a student's learning ability.
[0060] Based on the academic ability bases of all respondents with the same academic ability stratification label across all history exam questions in a large-scale dataset, we calculated the horizontal academic ability dissimilarity index for each respondent. For all respondents across all history exams, we analyzed the data within the student population with the same academic ability stratification label. For example, in the "high academic ability stratification" group, there are students A, B, and C, with academic ability bases of 92, 90, and 88, respectively.
[0061] Calculate each student's horizontal dissimilarity index. Assume that the calculation method used is to use the average academic base of the group as a benchmark and calculate the degree of difference between each student's academic base and the average. The average academic base of "high academic ability" students is (92 + 90 + 88) ÷ 3 = 90.
[0062] Student A's horizontal dissimilarity index is ∣92-90∣=2, Student B's horizontal dissimilarity index is ∣90-90∣=0, and Student C's horizontal dissimilarity index is ∣88-90∣=2. The horizontal dissimilarity index reflects the differences in learning ability among students at the same level of academic ability.
[0063] Based on the academic ability bases of all respondents across all history exams in the ultra-large dataset, as well as the maximum and minimum academic ability bases within each academic ability stratification label, we calculated the vertical dissimilarity index for each respondent's academic ability. For example, the "high academic ability stratum" has a maximum academic ability base of 95 and a minimum academic ability base of 85; the "medium academic ability stratum" has a maximum academic ability base of 84 and a minimum academic ability base of 65; and the "low academic ability stratum" has a maximum academic ability base of 64 and a minimum academic ability base of 50.
[0064] Taking Student A (who belongs to the "high academic ability tier" and has an academic ability base score of 92) as an example, we calculate their academic ability vertical alienation index. We assume the following calculation method: Academic ability vertical alienation index = (|92−95|+|92−85|) / (95−50) (the denominator here is the difference between the maximum and minimum academic ability base scores across all academic ability tiers, used for normalization).
[0065] The vertical alienation index of student A's academic ability is 453 + 7, which is ≈ 0.22. The vertical alienation index of academic ability reflects the differences in learning ability between students at different academic levels.
[0066] The horizontal alienation index and vertical alienation index of academic ability of all answering subjects are used as the base of academic ability alienation among answering subjects. Combined with the answer records of each history test question in the ultra-large-scale data set (such as how many students answered correctly or incorrectly) and the score distribution (the number of students in each score range), the difficulty coefficient, discrimination index, reliability value and validity value of each history test question are obtained as the test question rating result of each history test question.
[0067] In this embodiment, the difficulty coefficient is: For example, for a history question, 80% of students in the "high academic ability" group answered correctly, 50% of students in the "medium academic ability" group answered correctly, and 20% of students in the "low academic ability" group answered correctly. Taking into account the varying bases of academic ability among students in different academic ability groups, these correct answer percentages are adjusted (assuming the adjustment method is weighted based on the horizontal and vertical academic ability variation indices) to ultimately determine the difficulty coefficient for the question. For example, after a complex calculation (simplified here), the difficulty coefficient is 0.6, indicating that the question is of moderate difficulty.
[0068] Discrimination Index: This analyzes the relationship between the scores of students at each academic level and the Horizontal Differentiation Index of Academic Ability. For example, students in the "Higher Ability Level" with a high Horizontal Differentiation Index of Academic Ability will score higher on this question, indicating that the question has good discrimination for students at higher academic levels. A discrimination index is calculated based on the performance of each academic level. A value of 0.7 indicates that the question can effectively distinguish students of different academic levels.
[0069] Reliability: This value is calculated by analyzing the stability of scores on this question among students of the same academic level and applying corrections based on the Horizontal Variation Index (for example, if students within a certain academic level have a high Horizontal Variation Index, their scores will fluctuate significantly, affecting the reliability calculation). Assuming a final reliability value of 0.8, this indicates that the measurement result for this question is relatively reliable.
[0070] In an alternative implementation, the horizontal and vertical dissimilarities of academic ability indexes of all answering subjects are used as the base of academic dissimilarities among answering subjects. In combination with the answer records and score distribution of each history test question in the ultra-large-scale dataset, the difficulty coefficient, discrimination index, reliability value, and validity value of each history test question are obtained as the test question rating result of each history test question, including: Based on the answer records for each history question in the ultra-large-scale dataset, the ratio of the number of students who answered correctly for each academic ability stratification label to the total number of students in that stratification label is used as the raw difficulty of the corresponding history question for each academic ability stratification label. For example, suppose the answer records for a math history question in the ultra-large-scale dataset show that 200 students in the "high academic ability" stratification label answered the question correctly, of which 160 answered correctly; 300 students in the "medium academic ability" stratification label answered the question correctly, of which 150 answered correctly; and 250 students in the "low academic ability" stratification label answered the question correctly, of which 50 answered correctly. Therefore, the raw difficulty for the "high academic ability" stratification label is 160 ÷ 200 = 0.8; the raw difficulty for the "medium academic ability" stratification label is 150 ÷ 300 = 0.5; and the raw difficulty for the "low academic ability" stratification label is 50 ÷ 250 = 0.2. The raw difficulty reflects the proportion of students in each academic ability stratification label who answered the question correctly.
[0071] Based on the horizontal and vertical dissimilarity indices of all participants in each academic ability stratification label, the original difficulty of the corresponding history exam question under the corresponding academic ability stratification label is adjusted to obtain the difficulty coefficient of each history exam question. Assume that in the "high academic ability stratification," the average horizontal dissimilarity index of all participants is 3, and the average vertical dissimilarity index is 0.2. Assume that the dissimilarity correction formula is: correction coefficient = 1 + average horizontal dissimilarity index × 0.1 + average vertical dissimilarity index × 0.5. Then, the correction coefficient for the "high academic ability stratification" is 1 + 3 × 0.1 + 0.2 × 0.5 = 1 + 0.3 + 0.1 = 1.4. The adjusted difficulty coefficient for this stratification is 0.8 × 1.4 = 1.12. (In actual applications, the result may be normalized to a reasonable range, for example, by converting 1.12 to a value between 0 and 1 using a certain rule. This is just an example calculation process.) The same method is used to adjust the difficulty for both the "middle-level" and "low-level" levels, comprehensively factoring in the final difficulty coefficient for each history question. This adjustment takes into account differences in learning ability within the same level (horizontal dissimilarity) and between different levels (vertical dissimilarity), ensuring that the difficulty coefficient more accurately reflects the actual difficulty of the test for different students.
[0072] The discrimination index of each history test question is calculated based on the score distribution of each history test question in the ultra-large-scale data set and the horizontal alienation index of the academic ability of all test takers; The reliability value of each history test question is calculated based on the score distribution of each history test question in the ultra-large-scale data set and the horizontal alienation index of the academic ability of all test takers; Based on the score distribution of each history test question in the ultra-large-scale data set, a knowledge point-alienation basis matrix for each knowledge point is constructed. Based on the knowledge point-alienation basis matrix for each knowledge point, the mastery degree of each knowledge point in different alienation intervals is calculated. Based on the mastery degree of each knowledge point in different alienation intervals, the validity value of each history test question is calculated.
[0073] For example, in the matrix of the knowledge point "Ancient Political System", there are 30 people in the low alienation range, 50 people in the medium alienation range, and 20 people in the high alienation range for the "high academic ability layer"; there are 20 people in the low alienation range, 40 people in the medium alienation range, and 40 people in the high alienation range for the "medium academic ability layer"; there are 10 people in the low alienation range, 30 people in the medium alienation range, and 60 people in the high alienation range for the "low academic ability layer".
[0074] Calculate the mastery of each knowledge point at different alienation intervals. For example, for the "Ancient Political System" knowledge point in the "High Ability" low alienation interval, the mastery is calculated by dividing the number of correct answers in that interval by the total number of people in that interval (assuming 25 people in the low alienation interval correctly answered). The mastery is 25 ÷ 30 = 0.83. Similarly, calculate the mastery for other intervals and different ability levels.
[0075] Suppose there's a formula for calculating validity: The validity value is the weighted sum of the mastery of all alienated intervals across all academic ability strata, where i represents the academic ability stratum, j represents the alienated interval, and wij represents the weight. For example, assuming w11=0.2, w12=0.3, w13=0.1, w21=0.1, w22=0.2, w23=0.1, w31=0.05, w32=0.05, and w33=0.1). Substituting these into the formula yields the validity value. In this way, validity is calculated based on the mastery of knowledge points across different alienated intervals to measure whether the test accurately assesses the intended knowledge and abilities.
[0076] In an alternative embodiment, the discrimination index of each history test question is calculated based on the score distribution of each history test question in the ultra-large-scale dataset and the horizontal alienation index of the academic ability of all test takers, including: Based on the score distribution of each history test question in the ultra-large-scale dataset, we analyzed the correlation coefficient between the academic ability horizontal alienation index of each academic ability stratification label and the corresponding history test score as the intra-stratum discrimination of each history test question; assuming that in the ultra-large-scale dataset, for a Chinese history test question, we stratify students according to their academic ability into "high academic ability layer", "middle academic ability layer" and "low academic ability layer".
[0077] For the "high academic ability" group, we analyze the lateral heterogeneity index of students in this group and their score distribution on this question, using statistical methods (such as the Pearson correlation coefficient) to calculate the correlation coefficient between the two. For example, the calculation found that students with a higher lateral heterogeneity index in the "high academic ability" group also scored relatively higher on this question, resulting in a correlation coefficient of r1=0.7. This r1 represents the intra-level discrimination of this question within the "high academic ability" group. It reflects the degree to which students within the same academic ability level differ in their scores on this question due to differences in learning ability (reflected by the lateral heterogeneity index). A higher value indicates that the question is more effective in discriminating between students with different learning abilities within that group.
[0078] Using the same method, we can calculate the intra-stratum discrimination r2 for the middle school ability group and r3 for the low school ability group. Assume that r2 = 0.5 for the middle school ability group and r3 = 0.4 for the low school ability group.
[0079] Based on the score distribution of each history test question in the ultra-large-scale data set, the average score and score standard deviation of all answering objects with the highest academic ability stratification label and the average score and score standard deviation of all answering objects with the lowest academic ability stratification label are determined, and the cross-level discrimination of each history test question is calculated based on the average score and score standard deviation of all answering objects with the highest academic ability stratification label and the average score and score standard deviation of all answering objects with the lowest academic ability stratification label; taking this Chinese test question as an example, the average score xhigh and score standard deviation shigh of all answering objects with the highest academic ability stratification label (assuming it is "high academic ability layer"), as well as the average score xlow and score standard deviation slow of all answering objects with the lowest academic ability stratification label (assuming it is "low academic ability layer") are determined from the score distribution.
[0080] Assume that the average score of the "high academic ability class" is xhigh = 85 points, and the standard deviation of the score is shigh = 5; the average score of the "low academic ability class" is xlow = 45 points, and the standard deviation of the score is slow = 8.
[0081] Some common formulas can be used to calculate cross-layer discrimination, such as: cross-layer discrimination D = |xhigh−xlow| / (shigh 2 +slow 2 ) 1 / 2 .
[0082] Substituting the numerical value into the calculation: D ≈ 4.25 (in actual applications, the result may be further normalized or adjusted to fall within an appropriate numerical range). Cross-level discrimination measures the difference in scores between students in the highest and lowest academic ability levels on the test question, reflecting the test's ability to distinguish between students of different academic ability levels.
[0083] The discrimination index of each history test question is calculated based on the intra-layer discrimination and cross-layer discrimination of each history test question in the ultra-large-scale data set.
[0084] After obtaining the intra-layer discrimination (r1, r2, r3) and cross-layer discrimination D of each history test question, they are combined in a certain way to calculate the discrimination index.
[0085] Assuming the weighted average method is used, the weight of the intra-layer discrimination is set to w1 = 0.6, and the weight of the cross-layer discrimination is set to w2 = 0.4. First, calculate the average intra-layer discrimination r = (r1 + r2 + r3) / 3 = 30.7 + 0.5 + 0.4 = 0.53.
[0086] The discrimination index (DI) is then calculated as: w1 × mean value r + w2 × D = 0.6 × 0.53 + 0.4 × 4.25 = 0.318 + 1.7 = 2.018. (In practice, this result may be further processed to be within a reasonable range, such as by normalizing it to between 0 and 1 to better represent the degree of discrimination.) The discrimination index comprehensively considers the differences in test scores between students within the same and different academic ability levels, comprehensively reflecting the ability of the test to distinguish students of different academic ability levels.
[0087] In an alternative embodiment, the reliability value of each history test question is calculated based on the score distribution of each history test question in the ultra-large-scale dataset and the horizontal alienation index of the academic ability of all test takers, including: Based on the variance of the scores of all respondents for each academic ability stratification label for each history test question in the ultra-large-scale dataset and the variance of the total scores of the corresponding history test questions for all respondents for the corresponding academic ability stratification label, the stratified Cronbach's coefficient for each academic ability stratification label is calculated; taking a mathematics history test question as an example, assuming that in the ultra-large-scale dataset, students are divided into "high academic ability layer", "middle academic ability layer" and "low academic ability layer" according to their academic ability.
[0088] For the "high academic ability level", let the scores of all the respondents in this level for this history test question be x11, x12, ⋯, x1n, and the total scores of the corresponding history test paper be y11, y12, ⋯, y1n (n is the number of students in this level).
[0089] First, calculate the variance sx12 of the score value of each history test question and the variance sy12 of the total score of the test paper.
[0090] According to the Cronbach's coefficient formula α=1-(the quotient of the sum of the variance sx12 of the score values of all history test questions and the variance sy12 of the total score of the test paper).
[0091] Using the same method, we calculated the stratified Cronbach's coefficients α2 and α3 for the "middle school ability" and "low school ability" groups, respectively. Assuming that the calculated α2 for the "middle school ability" group is 0.78, the calculated α3 for the "low school ability" group is 0.75.
[0092] The standard deviation of the horizontal alienation index of the academic ability of all the respondents of each academic ability stratification label is used to perform alienation correction on the stratified Cronbach's coefficient of each academic ability stratification label to obtain the modified stratified Cronbach's coefficient of each academic ability stratification label; Continuing with the example of the “high academic ability layer”, let the academic ability horizontal alienation index of all respondents in this layer be z11, z12, ⋯, z1n.
[0093] First, calculate the standard deviation of the horizontal alienation index of academic ability; The formula for alienation correction is: Corrected stratified Cronbach coefficient α1′=α1×(1−σ1 / 10) (divided by 10 here to make the correction amplitude reasonable, this is only an example and can be adjusted in practice).
[0094] In the same way, the stratified Cronbach's coefficients of the "middle and low academic ability layers" were corrected.
[0095] The minimum value of the modified stratified Cronbach's coefficient of all academic ability stratification labels is used as the reliability value of each history test question.
[0096] The present invention also provides an implementation of a question-setting and test-question grading system based on deep mining of big data, including: The data integration module is used to integrate historical exam data, detailed student answer data, and teaching resource data, and perform cleaning and normalization processing to form a large-scale data set containing exam text, answer records, score distribution, and knowledge point labels; The question feature analysis module is used to generate a standard solution logic tree based on the question text and knowledge point labels of each history question in the ultra-large-scale dataset. Based on the structural features and structural relevance of all knowledge points in the standard solution logic tree, the application difficulty matrix and associated application degree matrix of each question are generated as the question feature of each history question. The history exam question rating module uses big data mining algorithms to analyze the answer records and score distribution of each history exam question in a large-scale data set, and considers the base of academic ability alienation among the answerers to obtain the difficulty coefficient, discrimination index, reliability value, and validity value of each history exam question as the exam question rating result; The rating model building module is used to use a neural network model to learn the proposition characteristics and rating characteristics of each history exam question in a large-scale dataset to obtain an exam question rating model; The model output rating module is used to obtain the test question rating results of the test question to be rated based on the standard problem-solving logic tree of the test question to be rated and the test question rating model.
[0097] This implementation of a question-based and exam-rating system based on deep big data mining has numerous beneficial effects. First, the data integration module integrates, cleans, and normalizes historical exam data, detailed student response data, and teaching resource data to form a comprehensive and standardized, ultra-large-scale dataset. This ensures more accurate and rich data for subsequent analysis, covering a wide range of information, from the exam questions themselves to student responses, laying a solid foundation for accurate question rating. Second, the question feature analysis module deeply analyzes the characteristics of each history exam question by generating a standard problem-solving logic tree and deriving an application difficulty matrix and a correlation application degree matrix based on this tree. This analysis, based on the knowledge point structure and relevance, allows educators to clearly understand the inherent logic and difficulty structure of the exam questions, helping to optimize question-setting strategies. Third, the history exam question rating module utilizes big data mining algorithms, comprehensively considering answer records, score distribution, and the base of academic ability alienation to derive question ratings, including difficulty coefficient, discrimination index, reliability value, and validity value. This process is closely linked to the actual situation of the students, so that the rating results can better reflect the test effects of the test questions in the actual test, and provide a targeted reference for teaching improvements. Fourthly, the rating model building module uses the neural network model to learn the characteristics of the questions and the rating characteristics of the test questions to construct a test question rating model. The powerful learning ability of the neural network model enables the model to extract key information from massive data, and has good generalization ability, which can adapt to the rating needs of different types of test questions. Fifth, the model output rating module derives the rating results based on the standard problem-solving logic tree of the test questions to be rated and the test question rating model, realizing a fast and scientific evaluation of the rated test questions. This provides educators with an efficient and reliable tool in daily question setting, test arrangement and other work, which helps to improve the quality of education and teaching and optimize the test evaluation system.
[0098] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the present invention and its equivalents, the present invention is intended to include these modifications and variations.
Claims
1. A method for grading test questions based on deep mining of big data, characterized in that: include: S1: Integrate historical exam data, detailed student answer data, and teaching resource data, and perform cleaning and normalization to form a large-scale dataset containing exam text, answer records, score distribution, and knowledge point labels; S2: Generate a standard problem-solving logic tree based on the test text and knowledge point labels of each history test question in the ultra-large-scale dataset. Based on the structural features and structural relevance of all knowledge points in the standard problem-solving logic tree, generate an application difficulty matrix and an associated application degree matrix for each test question as the proposition features of each history test question. S3: Use big data mining algorithms to analyze the answer records and score distribution of each history test question in a large-scale data set, and consider the academic ability difference between the test takers to obtain the difficulty coefficient, discrimination index, reliability value and validity value of each history test question as the test question rating result; S4: Use a neural network model to learn the proposition characteristics and test question rating characteristics of each history test question in a large-scale dataset to obtain a test question rating model; S5: Obtain the test question rating result of the test question to be rated based on the standard problem-solving logic tree of the test question to be rated and the test question rating model.
2. The method for grading test questions based on deep mining of big data according to claim 1 is characterized in that: S2: Generate a standard problem-solving logic tree based on the test text and knowledge point labels of each history test question in the ultra-large-scale dataset. Based on the structural features and structural relevance of all knowledge points in the standard problem-solving logic tree, generate an application difficulty matrix and a correlation application degree matrix for each test question as the proposition features of each history test question, including: Based on the test text and standard solution ideas of each history test question in the ultra-large-scale data set, at least one standard solution idea for each history test question is determined; Generate a standard solution logic tree based on each standard solution idea for each history test question and the knowledge point labels of all knowledge points involved in the standard solution idea; Based on the hierarchical depth of each knowledge point involved in each standard solution of each history test question in the corresponding standard solution logic tree and the structural complexity of the corresponding standard solution logic tree, the application difficulty of the corresponding knowledge point is evaluated; Generate an application difficulty matrix for each history test question based on the application difficulty of all knowledge points involved in each standard solution approach; Analyze the degree of correlation and application between different knowledge points involved in each standard solution of each history test question based on the structural correlation of different knowledge points involved in each standard solution logic tree; Generate a correlation application matrix for each history test question based on the correlation application degree between different knowledge points involved in each standard solution idea of each history test question; The application difficulty matrix and associated application degree matrix of each test question are used as the proposition characteristics of each history test question.
3. The method for grading test questions based on deep mining of big data according to claim 2 is characterized in that: Based on the hierarchical depth of each knowledge point involved in each standard solution of each history exam question in the corresponding standard solution logic tree and the structural complexity of the corresponding standard solution logic tree, the application difficulty of the corresponding knowledge point is evaluated, including: Calculate the structural complexity of the standard problem-solving logic tree based on the tree depth, molecular factor, node density, and loop degree of the standard logic tree; Determine the weight of each knowledge point involved in each standard solution of each history test question based on its hierarchical level in the knowledge point mastery expansion tree of the minimum scope of knowledge points covered by the corresponding history test question and its hierarchical level in the corresponding standard logic tree; Based on the hierarchical depth and corresponding weight value of each knowledge point involved in each standard solution idea of each history test question in the corresponding standard solution logic tree, as well as the structural complexity of the corresponding standard solution logic tree, the application difficulty of the corresponding knowledge point is evaluated.
4. The method for grading test questions based on deep mining of big data according to claim 2 is characterized in that: Based on the structural correlation of different knowledge points involved in each standard solution of each history test question in the corresponding standard solution logic tree, the correlation and application degree between different knowledge points involved in each standard solution of each history test question is analyzed, including: Based on the minimum number of edges between each two knowledge points involved in each standard solution of each history test question in the corresponding standard solution logic tree, the minimum number of edges between the nearest common ancestor node and the root node, and the overlap between the corresponding subtree structures of the two knowledge points in the corresponding standard solution logic tree, a correlation vector of the solution logic structure of the two knowledge points is generated; Based on the minimum number of edges between each two knowledge points involved in each standard solution of each history test question in the knowledge point mastery expansion tree of the minimum range of knowledge points covered by the corresponding history test question, the minimum number of edges between the most recent common ancestor node and the root node, and the overlap between the corresponding subtree structures of the two knowledge points in the corresponding knowledge point mastery expansion tree, a knowledge point mastery expansion structure correlation vector corresponding to the two knowledge points is generated; Treat the nearest common ancestor node of each two knowledge points involved in each standard problem-solving idea of each history test question in the corresponding standard problem-solving logic tree and the nearest common ancestor node in the knowledge point mastery extension tree of the minimum range of knowledge point blocks involved in the corresponding history test question as the structural superordinate mapping knowledge point combination corresponding to the two nodes, and generate the knowledge point mastery extension structural correlation vector of the corresponding structural superordinate mapping knowledge point combination as the superordinate knowledge point mastery extension structural correlation vector corresponding to the two knowledge points; Based on the problem-solving logic structure correlation vector, knowledge point mastery and extension structure correlation vector, and superordinate knowledge point mastery and extension structure correlation vector of every two knowledge points involved in each standard problem-solving approach for each history test question, the correlation application degree between different knowledge points involved in each standard problem-solving approach for each history test question is analyzed.
5. The method for grading test questions based on deep mining of big data according to claim 4 is characterized in that: Based on the problem-solving logic structure correlation vector, knowledge point mastery and extension structure correlation vector, and superordinate knowledge point mastery and extension structure correlation vector of each two knowledge points involved in each standard problem-solving approach for each history exam question, we analyze the degree of correlation and application between different knowledge points involved in each standard problem-solving approach for each history exam question, including: Based on the hierarchical level of each two knowledge points involved in each standard problem-solving approach for each history test question in the corresponding standard problem-solving logic tree, the hierarchical level of each two knowledge points involved in the knowledge point mastery expansion tree of the minimum scope of knowledge points involved in the corresponding history test question, and the hierarchical level of the corresponding structural superordinate mapping knowledge point combination in the knowledge point mastery expansion tree of the minimum scope of knowledge points involved in the corresponding history test question, determine the weights of the problem-solving logic structure correlation vectors of the corresponding two knowledge points, the weights of the knowledge point mastery expansion structure correlation vectors, and the weights of the superordinate knowledge point mastery expansion structure correlation vectors; Based on the weight of the problem-solving logic structure correlation vector of each two knowledge points involved in each standard problem-solving idea of each history test question, the weight of the knowledge point mastery extension structure correlation vector, and the weight of the superordinate knowledge point mastery extension structure correlation vector, the problem-solving logic structure correlation vector, the knowledge point mastery extension structure correlation vector, and the superordinate knowledge point mastery extension structure correlation vector of the corresponding two knowledge points are weightedly summed to obtain the degree of correlation application between different knowledge points involved in each standard problem-solving idea of each history test question.
6. The method for grading test questions based on deep mining of big data according to claim 1 is characterized in that: S3: Use big data mining algorithms to analyze the answer records and score distribution of each history test question in a large-scale data set, and consider the base of academic differentiation among the test takers to obtain the difficulty coefficient, discrimination index, reliability value, and validity value of each history test question as the test question rating result, including: Using the academic ability stratification model, we analyze the academic ability-related features of all the historical exam takers in a large-scale dataset and obtain the academic ability stratification label for each respondent. Based on the academic ability stratification model, the academic ability-related characteristics of all respondents in all history exams in the ultra-large-scale dataset are analyzed to calculate the academic ability base of each respondent; Based on the academic ability base of all the answerers of all history test questions in the ultra-large-scale data set and the same academic ability stratification label, the academic ability horizontal alienation index of each answerer is calculated; Based on the academic ability bases of all the respondents for all the history exam questions in the ultra-large-scale dataset, as well as the maximum and minimum academic ability bases in each academic ability stratification label, the vertical alienation index of academic ability of each respondent is calculated. The horizontal alienation index and vertical alienation index of academic ability of all answering subjects are used as the base of academic ability alienation between answering subjects, and combined with the answer records and score distribution of each history test question in the ultra-large-scale data set, the difficulty coefficient, discrimination index, reliability value and validity value of each history test question are obtained as the test question rating result of each history test question.
7. The method for grading test questions based on deep mining of big data according to claim 6 is characterized in that: The horizontal and vertical dissimilarities of academic ability of all test takers are used as the base of academic dissimilarities among test takers. Combined with the answer records and score distribution of each history test question in the ultra-large-scale dataset, the difficulty coefficient, discrimination index, reliability value, and validity value of each history test question are obtained as the test question rating results of each history test question, including: Based on the answer records of each history test question in the ultra-large-scale dataset, the ratio of the number of correct answers under each academic stratification label to the total number of people in the corresponding layer is determined as the original difficulty of the corresponding history test question under each academic stratification label. Based on the academic stratification horizontal and vertical alienation indices of all answer subjects under each academic stratification label, the original difficulty of the corresponding history test question under the corresponding academic stratification label is subjected to alienation correction to obtain the difficulty coefficient of each history test question; The discrimination index of each history test question is calculated based on the score distribution of each history test question in the ultra-large-scale data set and the horizontal alienation index of the academic ability of all test takers; The reliability value of each history test question is calculated based on the score distribution of each history test question in the ultra-large-scale data set and the horizontal alienation index of the academic ability of all test takers; Based on the score distribution of each history test question in the ultra-large-scale data set, a knowledge point-alienation basis matrix for each knowledge point is constructed. Based on the knowledge point-alienation basis matrix for each knowledge point, the mastery degree of each knowledge point in different alienation intervals is calculated. Based on the mastery degree of each knowledge point in different alienation intervals, the validity value of each history test question is calculated.
8. The method for grading test questions based on deep mining of big data according to claim 7 is characterized in that: Based on the score distribution of each history test question in the ultra-large-scale dataset and the horizontal alienation index of the academic ability of all test takers, the discrimination index of each history test question is calculated, including: Based on the score distribution of each history test question in the ultra-large-scale data set, the correlation coefficient between the academic ability horizontal alienation index of each academic ability stratification label and the corresponding history test score is used as the intra-stratum discrimination of each history test question; Based on the score distribution of each history test question in the ultra-large-scale dataset, the average score and score standard deviation of all respondents with the highest academic ability stratification label and the average score and score standard deviation of all respondents with the lowest academic ability stratification label are determined. The cross-level discrimination of each history test question is calculated based on the average score and score standard deviation of all respondents with the highest academic ability stratification label and the average score and score standard deviation of all respondents with the lowest academic ability stratification label. The discrimination index of each history test question is calculated based on the intra-layer discrimination and cross-layer discrimination of each history test question in the ultra-large-scale data set.
9. The method for grading test questions based on deep mining of big data according to claim 7, characterized in that: The reliability value of each history question is calculated based on the score distribution of each history question in the ultra-large-scale dataset and the horizontal alienation index of the academic ability of all respondents, including: Based on the variance of the scores of all respondents for each academic ability stratification label for each history test question in the ultra-large-scale dataset and the variance of the total scores of the corresponding history test question for all respondents for the corresponding academic ability stratification label, the hierarchical Cronbach's coefficient for each academic ability stratification label is calculated. The standard deviation of the horizontal alienation index of the academic ability of all the respondents of each academic ability stratification label is used to perform alienation correction on the stratified Cronbach's coefficient of each academic ability stratification label to obtain the modified stratified Cronbach's coefficient of each academic ability stratification label; The minimum value of the modified stratified Cronbach's coefficient of all academic ability stratification labels is used as the reliability value of each history test question.
10. A question-setting and test-question rating system based on deep mining of big data, characterized in that: include: The data integration module is used to integrate historical exam data, detailed student answer data, and teaching resource data, and perform cleaning and normalization processing to form a large-scale data set containing exam text, answer records, score distribution, and knowledge point labels; The question feature analysis module is used to generate a standard solution logic tree based on the question text and knowledge point labels of each history question in the ultra-large-scale dataset. Based on the structural features and structural relevance of all knowledge points in the standard solution logic tree, the application difficulty matrix and associated application degree matrix of each question are generated as the question feature of each history question. The history exam question rating module uses big data mining algorithms to analyze the answer records and score distribution of each history exam question in a large-scale data set, and considers the base of academic ability alienation among the answerers to obtain the difficulty coefficient, discrimination index, reliability value, and validity value of each history exam question as the exam question rating result; The rating model building module is used to use a neural network model to learn the proposition characteristics and rating characteristics of each history exam question in a large-scale dataset to obtain an exam question rating model; The model output rating module is used to obtain the test question rating results of the test question to be rated based on the standard problem-solving logic tree of the test question to be rated and the test question rating model.
Citation Information
Patent Citations
Method for question bank quality evaluation
CN104732352A
Test question difficulty estimation method and device, electronic equipment and storage medium
CN111310463A
Proposition examination question rating method and system based on big data
CN118885584A