Question bank generation method, device, storage medium and electronic device

By collecting the subjective and objective factors of students' wrong questions, using hierarchical clustering and wrong questions correlation evaluation models, a dynamically adjusted question bank was generated, and the problem of low accuracy in screening of knowledge points in the existing technology was solved, and the effect of adjusting the difficulty of questions based on students' real level was achieved.

CN119903243BActive Publication Date: 2025-07-04浙江海亮科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510399650.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-04
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

In the existing technology, the cluster of important knowledge points has not eliminated the knowledge points affected by subjective reasons during students' questions, resulting in low accuracy in screening of knowledge points and the difficulty of questions cannot be dynamically adjusted according to students' real learning level.

Method used

Through the smart education terminal, the subjective and objective factor feature data of students' wrong questions were collected, and the hierarchical clustering algorithm and the preset wrong question correlation evaluation model were used to filter out the most influential objective factor features, and generate question banks of different difficulty levels and map them on the learning map.

Benefits of technology

It improves the accuracy of the knowledge point cluster, can dynamically adjust the difficulty of the question based on students' real learning level, and improves the accuracy and learning effect of the question bank.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119903243B_ABST
    Figure CN119903243B_ABST
Patent Text Reader

Abstract

The present disclosure provides a question bank generation method, apparatus, storage medium, and electronic device. The method includes: collecting, by using an intelligent education terminal, a first feature data set including factors affecting students' wrong answers; screening, by using a hierarchical clustering algorithm, features in the first feature data set that are caused by objective factors for students' wrong answers; evaluating, by using a preset wrong-question correlation evaluation model, the correlation scores between the features that are caused by objective factors for students' wrong answers, sorting the features that are caused by objective factors for students' wrong answers in descending order of the correlation scores, and determining the knowledge points corresponding to the features with the highest correlation scores as the first target knowledge point clusters; generating, based on a basic knowledge point cluster and the first target knowledge point cluster, question banks corresponding to questions of different difficulty levels and mapping them on a learning map, where the learning map includes one or more key learning grids, and the key learning grids are the grids corresponding to the first target knowledge point clusters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of intelligent education, and particularly to a question bank generation method, device, storage medium, and electronic device. Background Art

[0002] In recent years, with the continuous advancement and development of computer technology and educational informatization, computer and artificial intelligence technologies have gradually been applied to various daily educational activities. Improving the accuracy of knowledge points in a knowledge point cluster can effectively help teachers improve the hit rate of important knowledge points, and at the same time, can recommend more accurate knowledge points to students to improve their learning effects.

[0003] In the current related technologies, the knowledge point cluster includes a basic knowledge point cluster and an important knowledge point cluster. However, in the important knowledge point cluster, the knowledge points affected by subjective reasons such as guessing, not doing, not knowing, having a half - understanding, misreading questions, and being unskilled during the process of students doing questions are not excluded. It is impossible to effectively obtain the students' mastery of knowledge points, which affects the accuracy of screening key knowledge points, resulting in a low accuracy of knowledge points in the question bank generated based on the basic knowledge point cluster and the important knowledge point cluster. Furthermore, it is impossible to dynamically adjust the difficulty of the questions recommended for students according to their true learning levels. Summary of the Invention

[0004] The present disclosure provides a question bank generation method, device, storage medium, and electronic device.

[0005] According to a first aspect of the present disclosure, there is provided a question bank generation method, the method including:

[0006] Collecting a first feature dataset containing the influencing factors of students' wrong - answering questions by using an intelligent education terminal, where the influencing factors of wrong - answering questions include subjective factors and objective factors;

[0007] Using a hierarchical clustering algorithm to screen out the features of students' wrong - answering questions caused by objective factors from the first feature dataset of the influencing factors of wrong - answering questions;

[0008] Using a preset wrong - question correlation evaluation model to evaluate the correlation scores between the features of students' wrong - answering questions caused by objective factors, sorting the features of students' wrong - answering questions caused by objective factors in descending order of the correlation scores, and determining the knowledge points corresponding to the feature with the highest correlation score as the first target knowledge point cluster;

[0009] Based on the basic knowledge point cluster and the first target knowledge point cluster, generating question banks corresponding to questions of different difficulty levels and mapping them on a learning map, where the learning map includes one or more key learning grids, the key learning grids are the grids corresponding to the first target knowledge point cluster, and the basic knowledge point cluster is obtained by extracting knowledge points from textbooks by using an intelligent education terminal.

[0010] In some embodiments of the present disclosure, a first feature dataset including factors affecting students' wrong answers is collected by using an intelligent education terminal, including:

[0011] Collect the students' test-taking video images and the students' test-taking situations based on the intelligent education terminal;

[0012] Use the facial expression recognition algorithm and the eye movement tracking algorithm to identify the subjective factor feature dataset that causes students to get wrong answers in the students' test-taking video images, and obtain the objective factor feature dataset of students' wrong answers by using the students' test-taking situations;

[0013] Based on the subjective factor feature dataset and the objective factor feature dataset, determine the first feature dataset including factors affecting students' wrong answers.

[0014] In some embodiments of the present disclosure, the subjective factor feature dataset and the objective factor feature dataset are used to determine the first feature dataset including factors affecting students' wrong answers, including:

[0015] Merge the subjective factor feature dataset and the objective factor feature dataset to obtain a second feature dataset including factors affecting students' wrong answers;

[0016] Perform preprocessing and normalization processing on all feature data in the second feature dataset to obtain the first feature dataset including factors affecting students' wrong answers.

[0017] In some embodiments of the present disclosure, use a preset wrong-question correlation evaluation model to evaluate the correlation score between the features that cause students to get wrong answers due to objective factors, including:

[0018] Use a preset correlation algorithm to evaluate the correlation of each factor affecting wrong answers in the first feature dataset, obtain the first correlation analysis result between the features in the first feature dataset, and obtain the correlation strength between the features of the factors affecting wrong answers according to the first correlation analysis result;

[0019] Based on the first correlation analysis result, generate a dendrogram of the first feature dataset of the factors affecting wrong answers according to the hierarchical clustering algorithm, determine the number of groups of the features in the first feature dataset according to the number of clusters in the dendrogram, and allocate the feature data in the first feature dataset to the corresponding groups to obtain at least two third feature datasets;

[0020] Use the pre-trained correlation evaluation model to evaluate the correlation of the third feature dataset, obtain the second correlation analysis result between the third feature datasets, and obtain the correlation score based on the first correlation analysis result and the second correlation analysis result.

[0021] In some embodiments of the present disclosure, a preset correlation algorithm is used to evaluate the correlation of each wrong-question influencing factor in the first feature dataset, obtaining a first correlation analysis result among the features in the first feature dataset, and obtaining the correlation strength among the wrong-question influencing factor features according to the first correlation analysis result, including:

[0022] The first correlation analysis result is calculated by a preset correlation algorithm:

[0023]

[0024] Wherein, represents the average value of the feature and the feature in the first feature dataset, represents the feature in the first feature dataset, represents the feature and the feature The first correlation analysis result therebetween;

[0025] The correlation strength among the wrong-question influencing factor features is obtained according to the first correlation analysis result, wherein, The value range is [0, 1].

[0026] In some embodiments of the present disclosure, a pre-trained correlation evaluation model is used to evaluate the correlation of the third feature dataset, obtaining a second correlation analysis result among the third feature datasets, and obtaining a correlation score based on the first correlation analysis result and the second correlation analysis result, including:

[0027] The second correlation analysis result is calculated by a pre-trained correlation evaluation model:

[0028]

[0029] Wherein, represents the second correlation analysis result, represents the constant term, represents the partial regression coefficient of the pre-set correlation evaluation model, represents the feature in the third feature dataset;

[0030] The correlation score is obtained based on the first correlation analysis result and the second correlation analysis result.

[0031] In some embodiments of the present disclosure, after using a preset wrong-question correlation evaluation model to evaluate the correlation scores between the characteristics of students' wrong answers caused by objective factors, sorting the characteristics of students' wrong answers caused by objective factors in descending order of the correlation scores, and determining the knowledge points corresponding to the characteristics with the highest correlation score as the first target knowledge point cluster, the method further includes:

[0032] Determining a basic knowledge point cluster based on the extraction of knowledge points from the textbook text, determining a first important knowledge point cluster based on the importance degree of the knowledge points in the basic knowledge point cluster, and determining a second important knowledge point cluster based on the extraction of knowledge points from the historical question bank and according to the difficulty level of the knowledge points in the historical question bank;

[0033] Determining the knowledge points that are the same among the first target knowledge point cluster, the first important knowledge point cluster, and the second important knowledge point cluster as the second target knowledge point cluster;

[0034] Generating question banks corresponding to questions of different difficulty levels based on the basic knowledge point cluster and the first target knowledge point cluster and mapping them on the learning map, including:

[0035] Generating question banks corresponding to questions of different difficulty levels based on the basic knowledge point cluster and the second target knowledge point cluster and mapping them on the learning map.

[0036] According to a second aspect of the present disclosure, there is provided a question bank generation device, which includes:

[0037] An acquisition unit, configured to use an intelligent education terminal to acquire a first feature data set including factors affecting students' wrong answers, and the factors affecting students' wrong answers include subjective factors and objective factors;

[0038] A screening unit, configured to use a hierarchical clustering algorithm to screen out the characteristics of students' wrong answers caused by objective factors from the first feature data set of factors affecting students' wrong answers;

[0039] A determination unit, configured to use a preset wrong-question correlation evaluation model to evaluate the correlation scores between the characteristics of students' wrong answers caused by objective factors, sort the characteristics of students' wrong answers caused by objective factors in descending order of the correlation scores, and determine the knowledge points corresponding to the characteristics with the highest correlation score as the first target knowledge point cluster;

[0040] A generation unit, configured to generate question banks corresponding to questions of different difficulty levels based on the basic knowledge point cluster and the first target knowledge point cluster and map them on the learning map, and the learning map includes one or more key learning grids, and the key learning grids are the grids corresponding to the first target knowledge point cluster, and the basic knowledge point cluster is obtained by extracting knowledge points from the textbook by using the intelligent education terminal.

[0041] According to a third aspect of the present disclosure, there is provided a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the method described in the foregoing first aspect is implemented.

[0042] According to a fourth aspect of the present disclosure, there is provided an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, and when the processor executes the computer program, the method described in the foregoing first aspect is implemented.

[0043] The question bank generation method, device, storage medium, and electronic device provided by the present disclosure collect a first feature dataset including factors affecting students' wrong answers through a smart education terminal. The factors affecting students' wrong answers include subjective factors and objective factors. The hierarchical clustering algorithm is used to screen out the features of wrong answers caused by objective factors from the first feature dataset of factors affecting students' wrong answers. The preset wrong-question correlation evaluation model is used to evaluate the correlation scores between the features of wrong answers caused by objective factors, and the features of wrong answers caused by objective factors are sorted in descending order according to the correlation scores. The knowledge point corresponding to the feature with the highest correlation score is determined as the first target knowledge point cluster. Based on the basic knowledge point cluster and the first target knowledge point cluster, question banks corresponding to questions of different difficulty levels are generated and mapped on the learning map. The learning map includes one or more key learning grids, and the key learning grids are the grids corresponding to the first target knowledge point cluster. The basic knowledge point cluster is obtained by extracting knowledge points from textbooks through a smart education terminal. By using the knowledge points in the objective factors that have the greatest impact on students' wrong answers as the knowledge points in the first target knowledge point cluster, the accuracy of the first target knowledge point cluster can be improved, and further the accuracy of generating question banks corresponding to questions of different difficulty levels based on the basic knowledge point cluster and the first target knowledge point cluster can be improved. The difficulty level of the questions recommended for students can be dynamically adjusted according to the students' actual learning levels.

[0044] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0046] Figure 1 is a flowchart of a question bank generation method provided by an embodiment of the present disclosure;

[0047] Figure 2 is a flowchart of a method for obtaining a first feature dataset provided by an embodiment of the present disclosure;

[0048] Figure 3 Schematic diagram of the knowledge point network provided by the embodiments of the present disclosure;

[0049] Figure 4 Schematic diagram of the screening process provided by the embodiments of the present disclosure;

[0050] Figure 5 Schematic diagram of a question bank marking provided by the embodiments of the present disclosure;

[0051] Figure 6 Schematic structural diagram of a question bank generation device provided by the embodiments of the present disclosure;

[0052] Figure 7 Schematic hardware structure diagram of an electronic device provided by the embodiments of the present disclosure. Detailed implementation manners

[0053] The following makes an explanation of the exemplary embodiments of the present disclosure with reference to the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the descriptions of well-known functions and structures are omitted below.

[0054] In recent years, with the continuous advancement and development of computer technology and educational informatization, computer and artificial intelligence technologies have gradually been applied to various daily educational teaching activities. Improving the accuracy of knowledge points in the knowledge point cluster can effectively help teachers improve the hit rate of important knowledge points, and at the same time can recommend more accurate knowledge points to students to improve the learning effect of students.

[0055] In the current related technologies, the knowledge point cluster includes a basic knowledge point cluster and an important knowledge point cluster, but the knowledge points affected by subjective reasons such as guessing, not doing, not knowing how to do, having a half - understanding, misreading questions, and being unskilled during the process of students doing questions are not excluded from the important knowledge point cluster, and the mastery of knowledge points by students cannot be effectively obtained, which affects the accuracy of screening key knowledge points, resulting in a low accuracy of knowledge points in the question bank generated based on the basic knowledge point cluster and the important knowledge point cluster, and further resulting in the inability to dynamically adjust the difficulty of the questions recommended to students according to the real learning level of students.

[0056] To solve the problems in the related art, the question bank generation method provided by the present disclosure collects a first feature data set including the influencing factors for students' wrong answers through a smart education terminal. The influencing factors for wrong answers include subjective factors and objective factors. The hierarchical clustering algorithm is used to screen out the features that objectively cause students to answer questions wrong from the first feature data set of the influencing factors for wrong answers. The preset wrong-question correlation evaluation model is used to evaluate the correlation scores between the features that objectively cause students to answer questions wrong, and the features that objectively cause students to answer questions wrong are sorted in descending order according to the correlation scores. The knowledge point corresponding to the feature with the highest correlation score is determined as the first target knowledge point cluster. Based on the basic knowledge point cluster and the first target knowledge point cluster, question banks corresponding to questions of different difficulty levels are generated and mapped on the learning map. The learning map includes one or more key learning grids, and the key learning grids are the grids corresponding to the first target knowledge point cluster. The basic knowledge point cluster is obtained by extracting the knowledge points in the textbook through the smart education terminal. By taking the knowledge points that have the greatest impact on students' wrong answers among the objective factors as the knowledge points in the first target knowledge point cluster, the accuracy of the first target knowledge point cluster can be improved, and further the accuracy of generating question banks corresponding to questions of different difficulty levels based on the basic knowledge point cluster and the first target knowledge point cluster can be improved. The difficulty level of the questions recommended for students can be dynamically adjusted according to the students' actual learning levels.

[0057] The following describes the question bank generation method, device, electronic device, and storage medium according to the embodiments of the present disclosure with reference to the accompanying drawings.

[0058] Figure 1 It is a flowchart of a question bank generation method provided by an embodiment of the present disclosure. As Figure 1 shown, the method includes:

[0059] Step 101, collect a first feature data set including the influencing factors for students' wrong answers through a smart education terminal. The influencing factors for wrong answers include subjective factors and objective factors;

[0060] In some embodiments, the smart education terminal can be implemented based on a hardware device (such as a tablet computer, an intelligent whiteboard), or can be a software application installed on a general computer or a mobile device. Usually, these terminals can be connected to a cloud service to realize functions such as accessing rich educational resources and data synchronization.

[0061] In some embodiments, the first feature data set including the influencing factors for students' wrong answers can be obtained from the log data of the learning management system.

[0062] In some embodiments, since there are many factors that affect students' wrong answers, including misreading questions, writing wrong answers, and truly not understanding, etc., it is impossible to effectively obtain students' mastery of knowledge points. Therefore, it is necessary to screen out the knowledge points that students truly cannot do from the feature dataset containing the influencing factors of students' wrong answers, so as to facilitate teachers to understand students' mastery of knowledge points.

[0063] In some embodiments, subjective factors may include emotional characteristics such as positive, neutral, negative, etc.; concentration characteristics such as concentrated, distracted, easily affected by external factors, etc.; question-solving style characteristics such as careful, careless, impatient, etc.; objective factors may include knowledge point mastery characteristics such as proficient, skilled, unskilled, etc.; question-solving efficiency characteristics such as fast, medium, slow, etc.; question difficulty level characteristics: easy, ordinary, difficult, etc.

[0064] Step 102: Use the hierarchical clustering algorithm to screen out the characteristics of wrong answers caused by objective factors from the first feature dataset of influencing factors of wrong answers.

[0065] In some embodiments, a bottom-up or top-down hierarchical clustering method can be selected. Taking the bottom-up hierarchical clustering as an example, each influencing factor of wrong answers can be regarded as a separate class first, and then the most similar classes can be merged in turn.

[0066] In some embodiments, before performing hierarchical clustering, a distance metric can be selected to measure the similarity or difference between different features. Commonly used distance metrics include Euclidean distance, Manhattan distance, etc.

[0067] In some embodiments, improving the screening of the characteristics of wrong answers caused by objective factors can exclude the influence of subjective factors on wrong answers, reduce data interference, and accurately judge whether students master the corresponding knowledge points.

[0068] Step 103: Use a preset wrong-question correlation evaluation model to evaluate the correlation scores between the characteristics of wrong answers caused by objective factors, and sort the characteristics of wrong answers caused by objective factors in descending order of correlation scores. Determine the knowledge points corresponding to the characteristics with the highest correlation score as the first target knowledge point cluster.

[0069] In some embodiments, the preset wrong-question correlation evaluation model can be a pre-trained correlation evaluation model.

[0070] In some embodiments, the correlation score can be directly obtained by a similarity algorithm, or obtained by the average value of the correlations obtained by multiple similarity algorithms, or obtained by weighted summation of the correlations obtained by multiple similarity algorithms.

[0071] In some embodiments, by sorting the characteristics of wrong answers made by students due to objective factors in descending order of relevance scores, the screening speed of the characteristics with the highest subsequent relevance scores can be improved.

[0072] In some embodiments, the characteristic with the highest relevance score is the one that has the greatest impact on wrong answers among objective factors. By determining the knowledge points corresponding to the objective factors that have the greatest impact on students' wrong answers as the first target knowledge point cluster, the knowledge points caused by subjective factors for students' wrong answers can be excluded, and the knowledge points corresponding to the objective factors that have the greatest impact on students' wrong answers can be obtained. The accuracy of the knowledge points in the obtained first target knowledge point cluster is higher.

[0073] Step 104: Based on the basic knowledge point cluster and the first target knowledge point cluster, generate question banks corresponding to questions of different difficulty levels and map them on the learning map. The learning map includes one or more key learning grids, and the key learning grids are the grids corresponding to the first target knowledge point cluster. The basic knowledge point cluster is obtained by extracting knowledge points from textbooks based on the intelligent education terminal.

[0074] In some embodiments, the learning map is a visualization tool that can be presented in the form of a network diagram or a flowchart. The learning map is used to help students understand and navigate complex knowledge systems, where each knowledge point cluster can be represented as a grid in the learning map. For example, the basic knowledge point cluster and the first target knowledge point cluster respectively correspond to a learning grid.

[0075] In some embodiments, the learning map is used to reflect the learning grid corresponding to the current user. According to the order of students' answering questions, the learning grids can be linked in sequence to obtain the learning path of the current user. The learning path is generated based on the learning map according to the students' learning goals, the students' current knowledge levels, and the students' learning progress. By combining the knowledge graph and the learning path, key knowledge points are identified and key monitoring is carried out when students enter the learning scope of these knowledge points. When students encounter difficulties in learning progress, the system will match users with similar learning paths to promote interaction and mutual assistance among students, thereby enhancing the learning experience.

[0076] In some embodiments, by classifying the knowledge points in the basic knowledge point cluster and the first target knowledge point cluster in detail and determining their difficulty levels according to the refinement degree of the knowledge points, the systematicness and hierarchy of the knowledge points are ensured, providing a basis for generating questions in the subsequent question bank. The questions in the question bank correspond to important knowledge points in the textbook content. This application does not limit the types of questions in the question bank, which can be one or more of multiple-choice questions, fill-in-the-blank questions, short-answer questions, essay questions, etc.

[0077] In some embodiments, important knowledge point clusters marked as core concepts or principles with wide applications in the knowledge graph can be marked with different colors or symbols to distinguish ordinary knowledge points from the first target knowledge points when marking. Design a learning map based on the knowledge graph to ensure that the relationships between knowledge points are clearly visible and highlight the first target knowledge point cluster.

[0078] In some embodiments, knowledge points can be distributed on a two-dimensional grid, and each knowledge point cluster occupies a cell. The size of the cell can be adjusted according to the complexity and importance of the knowledge point. For the cell where the first target knowledge point cluster is located, a special identifier can be used for marking to form a key learning grid.

[0079] In some embodiments, a navigation path can be added to the learning map to indicate the learning order from basic knowledge points to the first target knowledge points, which can help students better plan their learning paths.

[0080] In some embodiments, by determining the key learning grid on the learning map, key knowledge points can be highlighted while providing students with a clear learning path, thereby improving students' learning effects.

[0081] As a possible implementation, as Figure 2 shown in the flowchart of a question bank generation method, based on the above embodiments, the specific process of using the intelligent education terminal to collect the first feature data set including the influencing factors of students' wrong answers is as follows:

[0082] Step 201: Collect the video images of students' problem-solving and the students' problem-solving situations based on the intelligent education terminal;

[0083] In some embodiments, through the image acquisition device installed in the classroom, the intelligent education terminal can record the whole process of students' problem-solving in real time, including students' writing actions, thinking time, facial expressions, etc.

[0084] In some embodiments, the intelligent education terminal can automatically record the specific questions and answers completed by each student, the problem-solving time for each question, whether the answer is modified and the number of times the answer is modified, whether to view the hint, and other information to obtain the students' problem-solving situations.

[0085] Step 202: Use the facial expression recognition algorithm and eye movement tracking algorithm to identify the subjective factor feature data set that causes students to answer questions wrongly in the video images of students' problem-solving, and use the students' problem-solving situations to obtain the objective factor feature data set of students' wrong answers.

[0086] In some embodiments, the facial expression recognition algorithm (Single Shot Multibox Detector, SSD) can be a deep learning-based method for detecting and classifying the emotional states shown by students during the problem-solving process, such as positive, neutral, negative, confused, frustrated, confident, etc. Specifically, the emotional changes occurring within a specific time period can be marked, especially the facial expressions of students when facing difficult problems.

[0087] In some embodiments, the eye movement tracking algorithm is used to monitor the gaze movement patterns of students, including the positions of the students' fixation points, the paths of saccading through the questions, etc. By analyzing these movement patterns and the question content, such as whether the student spends too much time on irrelevant information or skips key parts, it can be analyzed whether the student is concentrated and careful when doing the questions.

[0088] In some embodiments, a dedicated eye tracker, a high-resolution camera combined with an image processing algorithm can be used for eye movement tracking.

[0089] Step 203: Based on the subjective factor feature dataset and the objective factor feature dataset, determine the first feature dataset containing the influencing factors for students to get questions wrong.

[0090] In some embodiments, subjective factors (such as emotional characteristics, concentration characteristics, problem-solving style characteristics) and objective factors (such as knowledge point mastery characteristics, problem-solving efficiency characteristics) are combined to form the first feature set. By determining the first feature dataset containing the influencing factors for students to get questions wrong, the influences of both subjective and objective factors on students getting questions wrong can be considered simultaneously.

[0091] As a possible implementation, based on the subjective factor feature dataset and the objective factor feature dataset on the basis of the above embodiments, the specific process of determining the first feature dataset containing the influencing factors for students to get questions wrong includes:

[0092] Merge the subjective factor feature dataset and the objective factor feature dataset to obtain a second feature dataset containing the influencing factors for students to get questions wrong;

[0093] Perform preprocessing and normalization on all the feature data in the second feature dataset to obtain the first feature dataset containing the influencing factors for students to get questions wrong.

[0094] In some embodiments, since the data in the subjective factor feature dataset and the objective factor feature dataset are collected, there may be outliers and missing values. Therefore, before performing wrong-question analysis, preprocessing is required to ensure the quality and consistency of the data. Specifically, the Isolation Forest can be used to detect outliers and decide whether to delete or correct these outliers. Records containing missing values can be selected for deletion or the interpolation method can be used to fill in the missing values.

[0095] In some embodiments, the second feature dataset includes all the feature data in the subjective factor feature dataset and the objective factor feature dataset, including outliers and missing values.

[0096] In some embodiments, normalization processing is used to map all the feature data in the subjective factor feature dataset and the objective factor feature dataset to the same range, so that data of different scales are comparable. The method of normalization processing in this application is not limited.

[0097] As a possible implementation, on the basis of the above embodiments, a preset wrong-question correlation evaluation model is used to evaluate the correlation score between the features that cause students to get wrong answers due to objective factors, including:

[0098] Use a preset correlation algorithm to evaluate the correlation of each wrong-answer influencing factor in the first feature dataset, obtain the first correlation analysis result between the features in the first feature dataset, and obtain the correlation strength between the wrong-answer influencing factor features according to the first correlation analysis result;

[0099] In some embodiments, the preset correlation algorithm can be the Pearson correlation coefficient distance algorithm or the Spearman rank correlation coefficient. As long as it is an algorithm that can obtain the correlation between two features, it can be applied here. In this application, the preset correlation algorithm takes the Pearson correlation coefficient distance algorithm as an example. Using the Pearson correlation coefficient distance algorithm, the correlation between two-by-two features in the first feature dataset can be obtained, and the size of the correlation is used to measure the correlation strength between two features.

[0100] In some embodiments, because there is a normal phenomenon that the question difficulty leads to slow problem-solving efficiency, it is necessary to first find the correlation between the features in the first feature dataset.

[0101] Based on the first correlation analysis result, generate a dendrogram for the first feature dataset of wrong-answer influencing factors according to the hierarchical clustering algorithm, so as to determine the number of groups of features in the first feature dataset according to the number of clusters in the dendrogram, and allocate the feature data in the first feature dataset to the corresponding groups to obtain at least two third feature datasets;

[0102] In some embodiments, the first correlation analysis result is embodied in the form of a correlation matrix. The correlation matrix can be used as the input of the hierarchical clustering algorithm to obtain the corresponding dendrogram. One cluster in the dendrogram represents a group of strongly correlated features. According to the dendrogram, the elbow method or the silhouette coefficient method can be used to automatically split the nodes in the dendrogram to obtain the number of clusters.

[0103] In some embodiments, each third feature dataset corresponds to a cluster.

[0104] In some embodiments, based on the first correlation analysis result, a dendrogram is generated from the first feature dataset of the influencing factors of wrong answers according to the hierarchical clustering algorithm, and the correlation distribution among the evaluation index features can be obtained. The shortest distance linkage method can be selected: d(Ccluster,Dcluster) = min{d(c,d): c ∈ C, d ∈ D}, where d(Ccluster,Dcluster) represents the distance between cluster C and cluster D, and c, d represent a point c in cluster C and a point d in cluster D.

[0105] The third feature dataset is evaluated for correlation using a pre-trained correlation evaluation model to obtain a second correlation analysis result among the third feature datasets, and a correlation score is obtained based on the first correlation analysis result and the second correlation analysis result.

[0106] In some embodiments, the pre-trained correlation evaluation model is used to obtain the correlation among the third datasets.

[0107] In some embodiments, based on the first correlation analysis result, the correlation distribution among the third feature sets is generated to determine whether there is a correlation among the feature sets, which solves the problems of low accuracy and lack of a global perspective in statistical calculation and averaging based on the correlation analysis result.

[0108] In some embodiments, after obtaining the first correlation analysis result and the second correlation analysis result, the two correlation analysis results can be averaged to obtain a correlation score, or the two correlation analysis results can be weighted and summed to obtain a correlation score. The method for calculating the correlation score in this application is not limited.

[0109] As a possible implementation, based on the above embodiments, a preset correlation algorithm is used to evaluate the correlation of each influencing factor of wrong answers in the first feature dataset to obtain a first correlation analysis result among the features in the first feature dataset, and the correlation strength among the influencing factors of wrong answers is obtained according to the first correlation analysis result, including:

[0110] The first correlation analysis result is calculated by the preset correlation algorithm:

[0111] Formula 1

[0112] In Formula 1, represents the average value of feature and feature in the first feature dataset, represents the feature in the first feature dataset, represents feature and features The first correlation analysis result between;

[0113] According to the first correlation analysis result, obtain the correlation strength between the influencing factors of wrong questions, where The value range is [0, 1].

[0114] In some embodiments, The closer the value is to 1, the weaker the correlation between the influencing factors of wrong questions and the characteristics of wrong questions.

[0115] In some embodiments, by using a preset correlation algorithm to evaluate the correlation of each influencing factor of wrong questions in the first feature dataset, the first correlation analysis result between the features in the first feature dataset is obtained, which is convenient to obtain the correlation between the influencing factors of wrong questions and the characteristics of wrong questions. The larger the first correlation analysis result, the weaker the correlation between the influencing factors of wrong questions and the characteristics of wrong questions.

[0116] As a possible implementation, on the basis of the above embodiments, use a pre-trained correlation evaluation model to evaluate the correlation of the third feature dataset, obtain the second correlation analysis result between the third feature datasets, and obtain a correlation score based on the first correlation analysis result and the second correlation analysis result, including:

[0117] The second correlation analysis result is calculated by a pre-trained correlation evaluation model:

[0118] Formula 2

[0119] In Formula 2, represents the second correlation analysis result, represents the constant term, represents the partial regression coefficient of the preset correlation evaluation model, represents the feature in the third feature dataset;

[0120] Obtain a correlation score based on the first correlation analysis result and the second correlation analysis result.

[0121] In some embodiments, by obtaining a correlation score based on the first correlation analysis result and the second correlation analysis result, it is possible to more finely screen out truly strongly correlated features.

[0122] As a possible implementation, based on the above embodiments, the method for generating a question bank further includes: evaluating the correlation score between the characteristics of students' wrong answers caused by objective factors by using a preset wrong-question correlation evaluation model, sorting the characteristics of students' wrong answers caused by objective factors in descending order of the correlation score, and determining the knowledge point corresponding to the characteristic with the highest correlation score as the first target knowledge point cluster.

[0123] Determining a basic knowledge point cluster based on the extraction of knowledge points from the textbook text, determining a first important knowledge point cluster based on the importance degree of the knowledge points in the basic knowledge point cluster, and determining a second important knowledge point cluster based on the extraction of knowledge points from the historical question bank and according to the difficulty level of the knowledge points in the historical question bank.

[0124] In some embodiments, the content in the textbook text can be uploaded to the first knowledge point analysis module, and the knowledge points corresponding to the content in the textbook text can be extracted by using text analysis. Specifically, the knowledge points corresponding to the content in the textbook text can be extracted by using keywords. Before performing text analysis, preprocessing of the content in the textbook text can also be included, such as word segmentation and removing stop words.

[0125] In some embodiments, the knowledge points are classified hierarchically according to the association relationship between the knowledge points to obtain a classification identifier corresponding to the knowledge points. Specifically, the classification identifier refers to assigning a specific mark to the knowledge points at each level in a multi-level classification structure for the purpose of distinction and management. Through the classification identifier of the knowledge points, the hierarchical relationship of the knowledge points can be clearly displayed, which helps to construct the knowledge point network. The knowledge points are used as nodes in the knowledge network, the knowledge points classified by chapter in the textbook are used as the primary classification nodes, corresponding to the main nodes in the knowledge network; the knowledge points classified by section in the textbook are used as the secondary classification nodes, corresponding to the subordinate nodes of the primary classification nodes in the knowledge network, and are used to represent each subsection in each chapter; the knowledge points classified by unit result can be used as the tertiary classification nodes, corresponding to the subordinate nodes of the secondary classification nodes in the knowledge network, and are used to further subdivide the specific knowledge points under each unit. The association relationship between the knowledge points refers to the above-mentioned chapters, sections, and units.

[0126] In some embodiments, the classification identifier corresponding to the knowledge points can be English letters. For example, A represents the primary classification node, B represents the secondary classification node, and C represents the tertiary classification node.

[0127] In some embodiments, the hierarchical identifiers corresponding to the knowledge points are encoded to obtain the identification code values of the knowledge points at each level. Specifically, each knowledge point corresponds to a unique identification code value, which can be one or a combination of multiple identification forms such as numbers, letters, and special symbols. When the hierarchical identifier is an English letter, for different knowledge points on the same hierarchical node, numbers can be used for encoding. For example, if there are 3 first-level classification nodes, they can be encoded as A1, A2, and A3 respectively. Another example is that if there are two second-level classification nodes, they can be encoded as B1 and B2. At the same time, the first-level classification nodes associated with B1 and B2 also need to be encoded.

[0128] In some embodiments, based on the hierarchical identifiers corresponding to the knowledge points and the identification code values of the knowledge points at each level, an importance level identification vector set corresponding to the knowledge points is determined, and a basic knowledge point cluster is generated according to the identification vector set. Specifically, the first-level classification nodes appear in the form of a single-unit identification vector, and the multi-level classification nodes appear in the form of a combination of multiple-unit identification vectors; the knowledge points in the textbook can be divided into n levels, such as Figure 3 shown Figure 3 is a schematic diagram of the knowledge point network provided by the embodiment of the present disclosure. When n = 3, the corresponding knowledge point cluster U = [An, Bm, Ck]. Specifically, the first-level classification An (A in the figure) in the knowledge point network is used as the main node and classified by chapter to obtain An = [A1, A2... An]; the second-level classification Bm is used as the subordinate node of An and classified by section to obtain Bm = [A1b1, A1b2, A2b1... Anbm]; the third-level classification Ck is used as the subordinate node of Bm and classified by unit result to obtain Ck = [A1b1c1, A1b2c2, A1b2c3... Anbmck]. The corresponding identification vector is generated according to the identification code, and the knowledge point network is constructed based on the identification vector. The knowledge point network is used to represent the relationship between knowledge points, and then multiple knowledge point networks are combined into a comprehensive identification vector set, that is, the basic knowledge point cluster Uori.

[0129] In some embodiments, a pre-trained evaluation model can be used to evaluate the importance level of each knowledge point in the basic knowledge point cluster, and then the first important knowledge point cluster is determined according to the importance level, such as Figure 4 shown Figure 4 is a schematic diagram of the screening process provided by the embodiment of the present disclosure. Specifically, the knowledge points with relatively high importance levels are screened out from the knowledge points in the basic knowledge point cluster Uori to obtain the first important knowledge point cluster Unew.

[0130] In some embodiments, based on all the knowledge points involved in the historical question bank and the pre-trained important knowledge point evaluation model, the difficulty level score of each knowledge point among all the knowledge points involved in the historical question bank is determined. Specifically, based on the historical question bank, according to the identification vectors (Wn, Rm, Hk) corresponding to the knowledge points in the question bank, the basic knowledge point cluster Zori = [Wn, Rm, Hk] corresponding to the question bank can be generated. The difficulty level score of each knowledge point among all the knowledge points involved in the historical question bank can be calculated by the pre-trained important knowledge point evaluation model, and its mathematical expression is as follows:

[0131] Formula 3

[0132] In Formula 3, Gj is the number of occurrences of the test point knowledge point, Gq is the number of error frequency occurrences of the test point knowledge point, Gr is the number of branches of the test point knowledge point, is the weight score corresponding to each index, is the difficulty level score of the knowledge point.

[0133] Based on the difficulty level score of each knowledge point among all the knowledge points involved in the historical question bank and the second preset screening threshold, screening is carried out among all the knowledge points involved in the historical question bank, and the set of screened knowledge points is determined as the second important knowledge point cluster. The second preset screening threshold is the initial value set for the difficulty level corresponding to the knowledge point. Specifically, the second preset screening threshold can be obtained by the same method as the first preset screening threshold, or can be obtained by a method different from the first preset screening threshold, that is, one of the methods such as self-definition, average method, median method, etc., or can be obtained by multiple of the methods such as self-definition, average method, median method, etc., and then taking the average of the multiple values obtained by the multiple methods. This embodiment does not make any limitations on this. The second preset screening threshold T2 can be determined. When the foregoing yG > T2, important knowledge points are screened out to obtain the second important knowledge point cluster; the selection of the second preset screening threshold is related to the importance of the knowledge points in the second important knowledge point cluster. If the second preset screening threshold is too small, it will result in a large number of knowledge points in the second important knowledge point cluster and low importance. If the second preset threshold is too large, it will result in too few knowledge points in the second important knowledge point cluster. The knowledge points in the second important knowledge point cluster are highly important, but some relatively important knowledge points may be excluded. Further, in the scenario of teacher lesson preparation, if the second preset threshold is too large, it will lead to a smaller range of teacher lesson preparation, which is not sufficient to cover most important knowledge points. Also, in the review scenario, if the second preset threshold is too large, some important test points will not be covered.

[0134] The knowledge points that are the same among the first target knowledge point cluster, the first important knowledge point cluster, and the second important knowledge point cluster are determined as the second target knowledge point cluster;

[0135] In some embodiments, one of algorithms such as Jaccard similarity algorithm, cosine similarity algorithm, Euclidean distance algorithm, and Manhattan distance algorithm can be used to determine the same knowledge points in the first target knowledge point cluster, the first important knowledge point cluster, and the second important knowledge point cluster, and then determine the second target knowledge point cluster.

[0136] Based on the basic knowledge point cluster and the first target knowledge point cluster, question banks corresponding to questions of different difficulty levels are generated and mapped on the learning map, including:

[0137] Based on the basic knowledge point cluster and the second target knowledge point cluster, question banks corresponding to questions of different difficulty levels are generated and mapped on the learning map.

[0138] In some embodiments, as Figure 5 shown, Figure 5 is a schematic diagram of a question bank marking method provided by an embodiment of the present disclosure. In the question bank, the basic knowledge point cluster is marked in white as the first information identifier, and the second target knowledge point cluster is marked in black as the second information identifier. The question bank includes questions corresponding to level 1 (easy difficulty), level 2 (medium difficulty), and level 3 (difficulty).

[0139] In some embodiments, since the second target knowledge point cluster is determined by identifying the same knowledge points in the first target knowledge point cluster, the first important knowledge point cluster, and the second important knowledge point cluster, using the basic knowledge point cluster and the second target knowledge point cluster to generate questions of different difficulty levels is more accurate. Therefore, the question bank composed of questions of different difficulty levels is also more accurate. Finally, mapping the more accurate question bank on the learning map can improve the learning effect of students.

[0140] Corresponding to the above question bank generation method, the present invention also proposes a question bank generation device. Since the device embodiment of the present invention corresponds to the above method embodiment, details not disclosed in the device embodiment can be referred to the above method embodiment, and will not be elaborated in the present invention.

[0141] Figure 6 is a structural schematic diagram of a question bank generation device provided by an embodiment of the present disclosure. As Figure 6 shown, the question bank generation device 600 includes:

[0142] An acquisition unit 601, configured to use a smart education terminal to acquire a first feature data set including factors affecting students' wrong answers. The factors affecting wrong answers include subjective factors and objective factors;

[0143] A screening unit 602, configured to use a hierarchical clustering algorithm to screen out the features of objective factors that cause students to answer questions wrongly from the first feature data set of factors affecting wrong answers;

[0144] A determination unit 603, configured to use a preset wrong-question correlation evaluation model to evaluate the correlation score between the characteristics of wrong questions caused by objective factors for students, sort the characteristics of wrong questions caused by objective factors for students in descending order of the correlation score, and determine the knowledge point corresponding to the characteristic with the largest correlation score as the first target knowledge point cluster;

[0145] A generation unit 604, configured to generate question banks corresponding to questions of different difficulty levels based on the basic knowledge point cluster and the first target knowledge point cluster and map them on a learning map, where the learning map includes one or more key learning grids, the key learning grids are the grids corresponding to the first target knowledge point cluster, and the basic knowledge point cluster is obtained by extracting knowledge points from textbooks based on an intelligent education terminal.

[0146] In some embodiments of the present disclosure, the acquisition unit 601 is configured to:

[0147] Collect the video images of students' question-solving and the students' question-solving situations based on an intelligent education terminal;

[0148] Use a facial expression recognition algorithm and an eye movement tracking algorithm to identify the subjective factor feature dataset that causes students to make wrong questions in the video images of students' question-solving, and obtain the objective factor feature dataset of students' wrong questions using the students' question-solving situations;

[0149] Based on the subjective factor feature dataset and the objective factor feature dataset, determine a first feature dataset including the influencing factors of students' wrong questions.

[0150] In some embodiments of the present disclosure, the acquisition unit 601 is configured to:

[0151] Merge the subjective factor feature dataset and the objective factor feature dataset to obtain a second feature dataset including the influencing factors of students' wrong questions;

[0152] Perform preprocessing and normalization processing on all feature data in the second feature dataset to obtain a first feature dataset including the influencing factors of students' wrong questions.

[0153] In some embodiments of the present disclosure, the determination unit 603 is configured to:

[0154] Use a preset correlation algorithm to evaluate the correlation of each influencing factor of wrong questions in the first feature dataset, obtain a first correlation analysis result between the features in the first feature dataset, and obtain the correlation strength between the feature datasets of influencing factors of wrong questions according to the first correlation analysis result;

[0155] Based on the first correlation analysis result, a dendrogram is generated from the first feature dataset of the influencing factors of wrong answers according to the hierarchical clustering algorithm, so as to determine the number of groups of features in the first feature dataset according to the number of clusters in the dendrogram, and the feature data in the first feature dataset are assigned to the corresponding groups to obtain at least two third feature datasets;

[0156] The third feature datasets are evaluated for correlation using a pre-trained correlation evaluation model to obtain a second correlation analysis result between the third feature datasets, and a correlation score is obtained based on the first correlation analysis result and the second correlation analysis result.

[0157] In some embodiments of the present disclosure, the determining unit 603 is configured to:

[0158] The first correlation analysis result is calculated by a preset correlation algorithm:

[0159]

[0160] Wherein, represents the average value of the feature and the feature in the first feature dataset, represents the feature in the first feature dataset, represents the feature and the feature the first correlation analysis result between;

[0161] The correlation strength between the influencing factors of wrong answers is obtained according to the first correlation analysis result, wherein, The value range is [0, 1].

[0162] In some embodiments of the present disclosure, the determining unit 603 is configured to:

[0163] The second correlation analysis result is calculated by a pre-trained correlation evaluation model:

[0164]

[0165] Wherein, represents the second correlation analysis result, represents the constant term, represents the partial regression coefficient of the preset correlation evaluation model, represents the feature in the third feature dataset;

[0166] A correlation score is obtained based on the first correlation analysis result and the second correlation analysis result.

[0167] In some embodiments of the present disclosure, the question bank generation device 600 further includes a second target knowledge point cluster determination unit, and the second target knowledge point cluster determination unit is configured to:

[0168] Determine a basic knowledge point cluster based on the extraction of knowledge points from the teaching material text, determine a first important knowledge point cluster based on the importance level of the knowledge points in the basic knowledge point cluster, and determine a second important knowledge point cluster based on the extraction of knowledge points from the historical question bank and according to the difficulty level of the knowledge points in the historical question bank;

[0169] Determine the same knowledge points in the first target knowledge point cluster, the first important knowledge point cluster, and the second important knowledge point cluster as the second target knowledge point cluster;

[0170] In some embodiments of the present disclosure, the generation unit 604 is configured to:

[0171] Generate question banks corresponding to questions of different difficulty levels based on the basic knowledge point cluster and the second target knowledge point cluster and map them on the learning map.

[0172] It should be noted that the foregoing explanations of the method embodiments also apply to the devices in this embodiment. The principles are the same and will not be limited in this embodiment.

[0173] Based on the above method as Figures 1 to 5 shown, correspondingly, this embodiment also provides a computer program product, including a computer program, and the computer program realizes the above method as Figures 1 to 5 shown when executed by a processor.

[0174] Based on the above method as Figures 1 to 5 shown, correspondingly, this embodiment also provides a computer-readable storage medium, on which a computer program is stored, and the computer program realizes the above method as Figures 1 to 5 shown when executed by a processor.

[0175] Based on such an understanding, the technical solution of the present application can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods in various implementation scenarios of the present application.

[0176] As Figure 7 shown is a schematic hardware structure diagram of an electronic device according to the present invention, including:

[0177] At least one processor 701; and,

[0178] A memory 702 communicatively connected to at least one of the processors 701; wherein,

[0179] The memory 702 stores instructions executable by at least one of the processors. The instructions are executed by at least one of the processors to enable at least one of the processors to execute the question bank generation method as described above.

[0180] Figure 7 Taking one processor 701 as an example.

[0181] The electronic device may further include: an input device 703 and a display device 704.

[0182] The processor 701, the memory 702, the input device 703, and the display device 704 may be connected through a bus or other means. In the figure, connection through a bus is taken as an example.

[0183] As a non-volatile computer-readable storage medium, the memory 702 can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the review content generation method in the embodiments of the present application. For example, Figures 1 to 5 the method flow shown. By running the non-volatile software programs, instructions, and modules stored in the memory 702, the processor 701 executes various functional applications and data processing, that is, implements the question bank generation method in the above embodiments.

[0184] The memory 702 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the review content generation method, etc. In addition, the memory 702 may include a high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 702 may optionally include a memory remotely set relative to the processor 701, and these remote memories can be connected to the device for executing the review content generation method through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0185] The input device 703 can receive input user clicks and generate signal inputs related to user settings and function controls of the review content generation method. The display device 704 may include a display screen and other display devices.

[0186] When the one or more modules are stored in the memory 702 and run by the one or more processors 701, the question bank generation method in any of the above method embodiments is executed.

[0187] Optionally, the above-mentioned physical device may further include a user interface, a network interface, a camera, a Radio Frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, etc. The user interface may include a display, an input unit such as a keyboard, etc. Optionally, the user interface may further include a USB interface, a card reader interface, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface), etc.

[0188] Those skilled in the art can understand that the above-mentioned physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or different component arrangements.

[0189] The storage medium may further include an operating system and a network communication module. The operating system is a program for managing the hardware and software resources of the above-mentioned physical device, and supports the operation of information processing programs and other software and / or programs. The network communication module is used to implement communication between components inside the storage medium, and communication between other hardware and software in the information processing physical device.

[0190] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus a necessary general hardware platform, or can also be implemented by hardware. By applying the solution of this embodiment, compared with the current existing technologies, in this embodiment, a first historical data set of multiple historical students is obtained, and based on the first historical data set, a learning performance prediction model is constructed, where the first historical data set includes the learning performances of multiple historical students; a second historical data set of multiple historical students is obtained, and based on the second historical data set, a learning performance prediction model is constructed, where the second historical data set includes the learning performances of multiple historical students; a question bank generation model is constructed based on the learning performance prediction model and the learning performance prediction model; based on the question bank generation model, the learning ability of the target student is determined, and by constructing the learning performance prediction model and the learning performance prediction model, and constructing the question bank generation model, the question bank generation model can comprehensively evaluate the learning performance of the target student in the objective aspect and the learning performance in the subjective aspect, improving the accuracy and diversity of the question bank generation model, helping schools or parents more intuitively and clearly understand the all-round development of students' learning ability, timely discover the reasons for students' performance decline, and intervene as early as possible to guide targeted learning tutoring programs.

[0191] It should be noted that, in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the said element.

[0192] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments described herein, but rather will conform to the broadest scope consistent with the principles and novel features claimed herein.

Claims

1. A question bank generation method, characterized in that, Including: Collecting a first feature dataset containing factors influencing students' wrong answers using an intelligent education terminal, where the factors influencing wrong answers include subjective factors and objective factors; Using a hierarchical clustering algorithm to screen out the features of wrong answers caused by objective factors from the first feature dataset of factors influencing wrong answers; Using a preset correlation algorithm to evaluate the correlation of each factor influencing wrong answers in the first feature dataset, obtaining a first correlation analysis result among the features in the first feature dataset, and obtaining the correlation strength among the features of factors influencing wrong answers according to the first correlation analysis result; Based on the first correlation analysis result, generating a dendrogram for the first feature dataset of factors influencing wrong answers according to the hierarchical clustering algorithm, determining the number of groups of features in the first feature dataset according to the number of clusters in the dendrogram, and allocating the feature data in the first feature dataset to the corresponding groups to obtain at least two third feature datasets; Using a pre-trained correlation evaluation model to evaluate the correlation of the third feature dataset, obtaining a second correlation analysis result among the third feature datasets, and obtaining the correlation score based on the first correlation analysis result and the second correlation analysis result, and sorting the features of wrong answers caused by objective factors in descending order according to the correlation score, determining the knowledge points corresponding to the features with the highest correlation score as the first target knowledge point cluster, and the feature with the highest correlation score is the feature with the greatest influence on wrong answers among objective factors; Based on the basic knowledge point cluster and the first target knowledge point cluster, generating question banks corresponding to questions of different difficulty levels and mapping them on a learning map, where the learning map includes one or more key learning grids, the key learning grids are the grids corresponding to the first target knowledge point cluster, and the basic knowledge point cluster is obtained by extracting knowledge points from textbooks using an intelligent education terminal.

2. The method according to claim 1, characterized in that The collecting a first feature dataset containing factors influencing students' wrong answers using an intelligent education terminal includes: Collecting students' question-solving video images and students' question-solving situations based on an intelligent education terminal; Using a facial expression recognition algorithm and an eye movement tracking algorithm to identify the subjective factor feature dataset that causes students' wrong answers in the students' question-solving video images, and obtaining the objective factor feature dataset of students' wrong answers using the students' question-solving situations; Based on the subjective factor feature dataset and the objective factor feature dataset, determining the first feature dataset containing factors influencing students' wrong answers.

3. The method according to claim 2, characterized in that, The determining the first feature dataset containing factors influencing students' wrong answers based on the subjective factor feature dataset and the objective factor feature dataset includes: Merging the subjective factor feature dataset and the objective factor feature dataset to obtain a second feature dataset containing factors influencing students' wrong answers; Performing preprocessing and normalization processing on all feature data in the second feature dataset to obtain the first feature dataset containing factors influencing students' wrong answers.

4. The method according to claim 1, wherein Performing a correlation evaluation on each wrong-question influencing factor in the first feature dataset using a preset correlation algorithm to obtain a first correlation analysis result among the features in the first feature dataset, and obtaining the correlation strength among the wrong-question influencing factor features according to the first correlation analysis result, including: The first correlation analysis result is calculated by a preset correlation algorithm: Among them, represents the average value of the features in the first feature dataset and the feature ; represents the feature in the first feature dataset, represents the feature and the feature ; the first correlation analysis result therebetween. Obtain the correlation strength between the influencing factor features of wrong answers based on the first correlation analysis result, where The value range is [0, 1].

5. The method according to claim 1, wherein Performing a correlation evaluation on the third feature dataset using a pre-trained correlation evaluation model to obtain a second correlation analysis result among the third feature datasets, and obtaining the correlation score based on the first correlation analysis result and the second correlation analysis result, including: The second correlation analysis result is calculated by a pre-trained correlation evaluation model: Among them, represents the second correlation analysis result, represents the constant term, represents the partial regression coefficient of the preset correlation evaluation model, represents the feature in the third feature dataset; Obtaining the correlation score based on the first correlation analysis result and the second correlation analysis result.

6. The method according to claim 1, characterized in that, After using a preset wrong-question correlation evaluation model to evaluate the correlation score among the features of wrong questions caused by the objective factors for students, sorting the features of wrong questions caused by the objective factors for students in descending order according to the correlation score, and determining the knowledge point corresponding to the feature with the highest correlation score as the first target knowledge point cluster, the method further includes: Determining a basic knowledge point cluster based on the extraction of knowledge points in the teaching material text, determining a first important knowledge point cluster based on the importance level of the knowledge points in the basic knowledge point cluster, and determining a second important knowledge point cluster based on the extraction of knowledge points in the historical question bank and according to the difficulty level of the knowledge points in the historical question bank; Determining the same knowledge points among the first target knowledge point cluster, the first important knowledge point cluster, and the second important knowledge point cluster as the second target knowledge point cluster; Generating question banks corresponding to questions of different difficulty levels based on the basic knowledge point cluster and the first target knowledge point cluster and mapping them on the learning map, including: Generating question banks corresponding to questions of different difficulty levels based on the basic knowledge point cluster and the second target knowledge point cluster and mapping them on the learning map.

7. A question bank generation device, characterized in that, Including: A collection unit, configured to collect a first feature dataset containing wrong-question influencing factors for students by using an intelligent education terminal, where the wrong-question influencing factors include subjective factors and objective factors; A screening unit, configured to screen out the features of wrong questions caused by objective factors for students from the first feature dataset of the wrong-question influencing factors by using a hierarchical clustering algorithm; A determination unit, configured to perform a correlation evaluation on each wrong-question influencing factor in the first feature dataset using a preset correlation algorithm to obtain a first correlation analysis result among the features in the first feature dataset, and obtaining the correlation strength among the wrong-question influencing factor features according to the first correlation analysis result; An allocation unit, configured to generate a dendrogram for the first feature dataset of the factors affecting students' wrong answers according to the hierarchical clustering algorithm based on the first correlation analysis result, so as to determine the number of groups of features in the first feature dataset according to the number of clusters in the dendrogram, and allocate the feature data in the first feature dataset to the corresponding groups to obtain at least two third feature datasets; A sorting unit, configured to use a pre-trained correlation evaluation model to evaluate the correlation of the third feature datasets, obtain a second correlation analysis result between the third feature datasets, and obtain the correlation score based on the first correlation analysis result and the second correlation analysis result, and sort the features of the objective factors causing students to answer questions wrongly from largest to smallest according to the correlation score, and determine the knowledge point corresponding to the feature with the largest correlation score as the first target knowledge point cluster, and the feature with the largest correlation score is the feature that has the greatest impact on wrong answers among the objective factors; A generation unit, configured to generate question banks corresponding to questions of different difficulty levels based on the basic knowledge point cluster and the first target knowledge point cluster and map them on a learning map, where the learning map includes one or more key learning grids, and the key learning grids are grids corresponding to the first target knowledge point cluster, and the basic knowledge point cluster is obtained by extracting knowledge points in the textbook based on an intelligent education terminal.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 6.

9. An electronic device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein, When the processor executes the computer program, it implements the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Question setting method and system

    CN107292785A

  • Power distribution network operation efficiency main influence factor mining method

    CN111144682A