A question bank quality analysis method and system based on answer behavior data
By analyzing user response records, a response state space for the question bank is constructed, which solves the problem of inaccurate question bank quality analysis in existing technologies and achieves more accurate question bank quality assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGXI ZHIJIE PRECISION TEST QUESTION BANK INFORMATION TECH SERVICE CO LTD
- Filing Date
- 2026-07-01
- Publication Date
- 2026-07-31
AI Technical Summary
Existing question bank quality analysis methods rely on manually set evaluation rules, which make it difficult to reflect changes in user behavior during continuous answering and to determine the applicability of the question bank to different user groups, resulting in inaccurate analysis results.
By collecting response records from the target user group, performing dual-window fragmentation based on segmentation points, generating response observation fingerprints and response association tags, constructing a response state space, and analyzing the quality of the question bank.
It improves the correlation between the question bank quality analysis results and actual use scenarios, and provides a more accurate basis for question bank maintenance and management.
Smart Images

Figure CN122490125A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of question bank quality analysis technology, specifically to a question bank quality analysis method and system based on answer behavior data. Background Technology
[0002] With the development of online education platforms, intelligent assessment systems, and exam training systems, question banks have been widely used in daily practice, periodic assessments, and qualification exam training. Users generate a large amount of answer behavior data during their use of question banks. How to analyze this answer behavior data to determine whether a question bank is suitable for a target user group or specific users is a crucial issue that needs to be addressed in question bank maintenance, question selection, and question delivery.
[0003] Some question bank quality analysis methods rely primarily on the statistical results of individual questions, such as accuracy rate, error rate, and average response time, and generate a question bank quality score by setting multiple evaluation indicators and assigning weights. While this approach reflects the basic usage of questions to some extent, it is heavily dependent on manually set evaluation rules, easily affected by indicator selection and weighting. Furthermore, its analysis focuses mainly on single questions or static statistical results, making it difficult to reflect changes in user behavior during continuous answering and to determine the actual applicability of a question at different answering stages or among different user groups. Consequently, it fails to fully reflect the dynamic performance of the question bank in real-world usage.
[0004] Therefore, it is necessary to provide a question bank quality analysis method and system that can analyze the usage adaptation of the question bank among the target user group based on the answer behavior data naturally generated by the question bank system. Summary of the Invention
[0005] This invention proposes a question bank quality analysis method and system based on answer behavior data, aiming to solve at least one of the technical problems mentioned in the background art above.
[0006] To achieve the above objectives, the first aspect of the present invention provides a method for analyzing the quality of a question bank based on answer behavior data, comprising: Collect answer records of multiple users in the target user group regarding the target question bank, and extract one or more answer result sequences for each user from the answer record data. The answer result sequence consists of multiple answer units arranged in the order of answering. Each response sequence is processed by a dual-window fragmentation based on the segmentation point, and the preceding behavior fragments and subsequent response fragments corresponding to multiple segmentation points are extracted. The preceding behavior fragments and subsequent response fragments constitute the behavior response samples corresponding to the segmentation points. For each behavioral response sample, a follow-up response analysis is performed to generate a response observation fingerprint for each preceding behavioral segment. Based on the response observation fingerprint, response association detection is performed on multiple preceding behavioral segments to determine the response association label between any two preceding behavioral segments. Based on response association tags, state merging is performed on multiple preceding behavior fragments to construct multiple response state nodes, where each response state node includes several preceding behavior fragments. The state transition relationship between response state nodes is extracted based on multiple behavior response samples. Based on the state transition relationship, a response state space corresponding to multiple response state nodes is constructed. Based on the response state space, a group response state observation and analysis of the target question bank is performed to generate the quality analysis results of the target question bank.
[0007] Preferably, a follow-up response analysis is performed on each behavioral response sample to generate a response observation fingerprint for each preceding behavioral segment, including: Based on the multiple exercises included in the preceding action segment, extract the intersection and union of exercises between any two preceding action segments, calculate the ratio between the number of exercises in the intersection and the number of exercises in the union, and obtain the item overlap rate between the two preceding action segments. Based on the answer results for each exercise, calculate the answer matching degree between two preceding action segments with respect to the intersection of the exercises, and determine the action matching weight between any two preceding action segments based on the question overlap rate and the answer matching degree. Extract the correct answer probability features for each question in the continuation response segment of the behavioral response sample, and construct the continuation response fingerprint of the preceding behavioral segment based on the correct answer probability features of the multiple questions. For the extraction of correct answer probability features, multiple behavior matching samples are identified for each behavior response sample based on the behavior matching weight. Based on the multiple exercises included in the continuation response fragment in each behavior matching sample, a reference response observation sample set for each exercise in the continuation response fragment of the behavior response sample is constructed. The correct answer probability features for the corresponding exercise are obtained statistically based on the reference response observation sample set.
[0008] Preferably, constructing the response state space corresponding to multiple response state nodes based on state transition relationships includes: For the extraction of state transition relations, the answer state node corresponding to the preceding behavior segment in each behavior response sample is determined. Based on the answer state nodes corresponding to multiple preceding behavior segments, the state transition edge corresponding to any two adjacent split points in the answer result sequence is constructed. Based on the answering units contained between two split points in the state transition edge, extract the exercises and corresponding answering results contained in the answering units, construct the state transition conditions of the state transition edge, and obtain the state transition relationship between any two answering state nodes. By combining the state transition relationships between multiple response state nodes, a response state space corresponding to multiple response state nodes is constructed.
[0009] Preferably, response association detection is performed on multiple preceding behavior segments based on response observation fingerprints to determine the response association label between any two preceding behavior segments, including: Extract the intersection and union of the problems between the response observation fingerprints corresponding to any two preceding behavior segments, and calculate the item overlap rate between the two response observation fingerprints based on the intersection and union of the problems; Construct a local correlation vector for each response observation fingerprint based on the intersection of exercises, and calculate the local correlation parameters between two response observation fingerprints based on the local correlation vector; The response association parameters between two preceding action segments are generated by integrating the overlap rate of the items and the local association parameters, and the response association labels between the two preceding action segments are determined based on the response association parameters.
[0010] Preferably, the group response state observation and analysis of the target question bank is performed based on the response state space, and the resulting quality analysis results of the target question bank include: The percentage of preceding behavior segments contained in each response state node is counted to obtain state distribution data for multiple response state nodes; Based on the answer state space, construct a set of candidate exercises for each answer state node, perform state transition effect analysis on multiple candidate exercises in the set, calculate the state transition effect parameters for each candidate exercise, and obtain the state transition effect data for multiple exercises; Based on state distribution data, high-frequency state nodes in the answer state space are identified. According to the candidate question set of answer state nodes and the state transition data of multiple questions, the observation coverage index of the target question bank on high-frequency state nodes is analyzed to generate the quality analysis results of the target question bank.
[0011] Preferably, the state transition action parameter and the observation coverage index also include: For a target candidate question in the candidate question set of a response status node, the state transition probability of multiple users transitioning to any other response status node after answering the target candidate question under the current response status node is calculated, and the state difference parameter between the current response status node and any other response status node is calculated. Based on the state transition probability and the state difference parameter, the state transition effect parameter of the target candidate question on the current response status node is calculated. For the observation coverage index, multiple valid observation problems in the candidate problem set of the answering state node are identified based on the state transition action parameter. The proportion of valid observation problems in the candidate problem set is calculated to obtain the observation coverage index of the high-frequency state node.
[0012] A second aspect of the present invention provides a question bank quality analysis system based on answer behavior data, comprising: The answer sequence construction module is used to collect answer record data of multiple users in the target user group about the target question bank, and extract one or more answer result sequences corresponding to each user from the answer record data. The answer result sequence consists of multiple answer units arranged in the answer order. The behavior response sample construction module is used to perform dual-window fragmentation processing based on segmentation points on each response result sequence, extract the preceding behavior fragments and subsequent response fragments corresponding to multiple segmentation points, and the preceding behavior fragments and subsequent response fragments constitute the behavior response samples corresponding to the segmentation points. The response association analysis module is used to perform continuation response analysis on each behavior response sample, generate response observation fingerprints for each preceding behavior segment, perform response association detection on multiple preceding behavior segments based on the response observation fingerprints, and determine the response association label between any two preceding behavior segments. The state transition analysis module is used to perform state merging processing on multiple preceding behavior fragments based on response association labels to construct multiple response state nodes, where each response state node includes several preceding behavior fragments, and extracts the state transition relationship between response state nodes based on multiple behavior response samples. The question bank quality analysis module is used to construct a question state space corresponding to multiple question state nodes based on the state transition relationship, and to perform group question state observation and analysis on the target question bank based on the question state space, thereby generating the quality analysis results of the target question bank.
[0013] The present invention has the following beneficial effects: This invention collects continuous answer records from a target user group within a target question bank, converts scattered answer records into answer result sequences, and further constructs behavioral response samples corresponding to preceding behavioral segments and subsequent response segments. By generating response observation fingerprints and response association tags, preceding behavioral segments with similar subsequent response behaviors are grouped into answer state nodes, and an answer state space is constructed based on the state transition relationships between answer state nodes. On this basis, question bank quality analysis results are generated according to high-frequency answer state nodes and their effective state transition coverage. This allows for the analysis of the target question bank's coverage and differentiation of major answer states during the target user group's actual continuous answering process, thereby improving the correlation between question bank quality analysis results and actual usage scenarios, and providing reliable data for management needs such as question bank maintenance, exercise supplementation, and question selection. Attached Figure Description
[0014] Figure 1 A flowchart illustrating a question bank quality analysis method based on answer behavior data, provided as an embodiment of the present invention.
[0015] Figure 2 This is a structural block diagram of a question bank quality analysis system based on answer behavior data, provided as an embodiment of the present invention. Detailed Implementation
[0016] To enable those skilled in the art to better understand the technical solutions of this invention, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for explaining the invention and are not intended to limit the invention.
[0017] The first aspect of this invention provides a question bank quality analysis method based on answer behavior data. This method can be applied to online question bank systems, intelligent assessment systems, online practice systems, or examination training platforms. It is used to analyze the usage of the target question bank by the target user group based on the actual answer behavior of the target user group in the target question bank, thereby generating question bank quality analysis results.
[0018] Please see Figure 1 The method includes the following steps: Step S1: Collect the answer records of multiple users in the target user group regarding the target question bank, and extract one or more answer result sequences corresponding to each user from the answer record data.
[0019] In this embodiment, the target user group can be a set of users who practice, test, or train using the same target question bank. For example, users learning the same course chapter, users preparing for the same exam subject, or users answering questions under the same online training task. The target question bank can include multiple question items that can generate standardized answer results, such as true / false questions, multiple choice questions, fill-in-the-blank questions, and standard answer calculation questions. Answer behavior data originates from the user's natural answering process in the question bank system and can include one or more of the following: user identifier, question number, answering order, answering time, submitted answer, answer result, and answering time.
[0020] The answer sequences of a single user answering multiple questions in a continuous answering process are arranged according to the answering order to obtain an answer result sequence. The answer result sequence consists of answer units, each of which includes at least a question number and an answer result. The question number identifies the specific question answered by the user, and the answer result represents the user's standardized response to that question. In one implementation, the answer result can be represented in binary form, for example, 1 for a correct answer and 0 for an incorrect answer. In another implementation, the answer unit may also include information such as answering time, answering time consumed, or answer submission, so that the answering process can be further analyzed as needed.
[0021] Step S2: Perform dual-window fragmentation processing on each response sequence based on the segmentation site, extract the preceding behavior fragments and subsequent response fragments corresponding to multiple segmentation sites, and construct the behavior response samples corresponding to the segmentation sites from the preceding behavior fragments and subsequent response fragments.
[0022] In this embodiment, a response sequence typically includes multiple consecutive response units. To analyze the relationship between the preceding response behavior formed by the user at a certain response stage and its subsequent response, multiple segmentation points are set in the response sequence. For example, a segmentation point can be set between any two adjacent response units in the response sequence. For any segmentation point, several response units before the segmentation point are extracted as preceding behavior segments, and several response units after the segmentation point are extracted as continuation response segments.
[0023] The preceding behavioral fragment represents the user's response behavior state before the segmentation point, while the subsequent response fragment represents the user's short-term response behavior after the segmentation point. By combining the preceding behavioral fragment and the subsequent response fragment for each segmentation point, a behavioral response sample can be generated for each segmentation point, thereby constructing several behavioral response samples corresponding to each response result sequence.
[0024] It should be noted that the dual-window fragmentation process can constrain the lengths of the preceding action fragment and the subsequent response fragment to avoid fragments that are too short, resulting in insufficient state representation, or fragments that are too long, introducing too much irrelevant response information. For example, minimum and maximum lengths of the preceding action fragment and the subsequent response fragment can be set. Segmentation points that satisfy the length constraints are determined as valid segmentation points, and their corresponding preceding action fragments and subsequent response fragments are used to construct action response samples.
[0025] Step S3: Perform follow-up response analysis on each behavior response sample to generate a response observation fingerprint for each preceding behavior segment. Based on the response observation fingerprint, perform response association detection on multiple preceding behavior segments to determine the response association label between any two preceding behavior segments.
[0026] In this embodiment, the response observation fingerprint is used to characterize the subsequent response performance corresponding to a certain preceding behavior segment. Specifically, for a certain preceding behavior segment, based on other behavior response samples with similar preceding answer characteristics, the correct answer probabilities of multiple subsequent questions are statistically analyzed to obtain the response observation fingerprint corresponding to the preceding behavior segment. The response observation fingerprint can be understood as the response distribution that a user may exhibit when answering several subsequent questions under the answering state corresponding to the preceding behavior segment.
[0027] After obtaining the response observation fingerprints of multiple preceding action segments, the response observation fingerprints of different preceding action segments are further compared to determine whether the effects of different preceding action segments on the subsequent response are consistent or similar, thereby determining whether there is a response correlation between the two and generating corresponding response correlation labels.
[0028] In one implementation, performing follow-up response analysis on each behavioral response sample to generate a response observation fingerprint for each preceding behavioral fragment specifically includes: Based on the multiple exercises included in the preceding action segment, extract the intersection and union of exercises between any two preceding action segments, calculate the ratio between the number of exercises in the intersection and the number of exercises in the union, and obtain the overlap rate of items between the two preceding action segments.
[0029] In this embodiment, based on the multiple exercises contained in each preceding action segment, the intersection and union of exercises between any two preceding action segments can be determined, namely, the set of exercises that appear in both preceding action segments and the set of all exercises that appear in both preceding action segments. The ratio between the number of exercises in the intersection and the number of exercises in the union is used as the question overlap rate, thereby representing the degree of overlap between the two preceding action segments in the range of answered exercises.
[0030] Based on the answer results for each exercise, calculate the answer matching degree between two preceding action segments with respect to the intersection of exercises, and determine the action matching weight between any two preceding action segments based on the question overlap rate and the answer matching degree.
[0031] In this embodiment, for each exercise in the exercise intersection set, the consistency of the answers to that exercise in the two preceding action segments is compared. If the answers are consistent, it is recorded as a match; otherwise, it is recorded as a mismatch. The ratio between the number of matched exercises in the exercise intersection set and the total number of exercises in the exercise intersection set is used as the answer matching degree between the two preceding action segments, representing the degree of consistency in the answers to the common exercises. Based on this, the product of the item overlap rate and the answer matching degree is used as the action matching weight, which characterizes the reference level of one preceding action segment in the continuation response analysis of another preceding action segment.
[0032] Extract the correct answer probability features for each question in the subsequent response segment of the behavioral response sample, and construct the subsequent response fingerprint of the preceding behavioral segment based on the correct answer probability features of the multiple questions.
[0033] In this embodiment, for a target behavior response sample among multiple behavior response samples, based on the behavior matching weight between the preceding behavior fragment of the target behavior response sample and the preceding behavior fragment of other behavior response samples, multiple behavior matching samples corresponding to the target behavior response sample are identified, that is, behavior response samples whose preceding behavior fragment has a valid behavior matching weight with the preceding behavior fragment of the target behavior response sample.
[0034] To avoid invalid samples from being included in the statistics, a behavior matching weight threshold can be set. When the behavior matching weight between two preceding behavior segments is greater than or equal to the behavior matching weight threshold, the corresponding behavior response sample is determined as the behavior matching sample of the target behavior response sample.
[0035] For each exercise in the follow-up response segment of the target behavior response sample, a reference response observation sample set corresponding to that exercise is constructed. Specifically, among multiple behavior matching samples corresponding to the target behavior response sample, behavior matching samples containing that exercise in the follow-up response segment are selected, and the selected behavior matching samples are used as the reference response observation sample set corresponding to that exercise. The reference response observation sample set is used to provide observation data on the user's answer results when continuing to answer the exercise under similar preceding answering behavior conditions.
[0036] After obtaining the reference response observation sample set for each exercise in the subsequent response segment of the target behavior response sample, the correct answer probability feature for that exercise is statistically obtained based on the answer results of each behavior matching sample in the reference response observation sample set regarding that exercise in the corresponding subsequent response segment. Specifically, the answer results of each behavior matching sample in the reference response observation sample set for that exercise are taken as observation values, where a correct answer is recorded as 1 and an incorrect answer as 0. Statistical analysis is performed on multiple observation values to calculate the proportion of exercises with correct answers, thereby obtaining the correct answer probability feature for that exercise.
[0037] In another implementation, the correct answer probability feature can be calculated by combining behavior matching weights. Specifically, based on the behavior matching weights between each behavior matching sample and the target behavior response sample, the observations of multiple behavior matching samples are weighted statistically, and the result of the weighted statistics is used as the correct answer probability feature corresponding to the exercise.
[0038] It is worth noting that, for the target behavior response sample, the probability feature of answering a certain question in the continuation response segment is not determined solely by the continuation response segment of the target behavior response sample itself, but is jointly provided by multiple behavior matching samples with similar preceding behavior segments. This reduces the impact of the randomness of a single sample on the continuation response analysis results, and makes the obtained probability feature more reflective of the user's overall response to the question under similar preceding answering behavior conditions.
[0039] Based on the above method, the above-described correct answer probability feature extraction process is performed on multiple questions included in the continuation response segment in the target behavior response sample to obtain the correct answer probability features corresponding to each question. The correct answer probability features of multiple questions are combined according to the question number to construct the continuation response fingerprint corresponding to the target preceding behavior segment.
[0040] Furthermore, if the number of reference response observation samples corresponding to a certain exercise is less than a preset support threshold, it can be considered that the exercise lacks reliable observation data under the current preceding behavior segment, and the exercise will be removed from the continuation response fingerprint to avoid distortion of the continuation response fingerprint due to insufficient sample size. Through the above processing, the continuation response fingerprint corresponding to each preceding behavior segment can be obtained. This continuation response fingerprint is not a simple similarity description of the preceding behavior segment itself, but is used to represent the probability distribution of correct answers that multiple users may produce for multiple continuation exercises under the preceding answering behavior represented by the preceding behavior segment.
[0041] In one implementation, determining the response association label between any two preceding behavior segments by performing response association detection on multiple preceding behavior segments based on response observation fingerprints specifically includes: Extract the intersection and union of the problem sets between the response observation fingerprints corresponding to any two preceding behavior segments, and calculate the item overlap rate between the two response observation fingerprints based on the intersection and union of the problem sets.
[0042] In this embodiment, since the successive response segments corresponding to different preceding behavior segments may be different, the exercise numbers and number of exercises contained in different response observation fingerprints may also be different. For any two response observation fingerprints, the exercise sets contained in each are extracted, and the intersection and union of the two exercise sets are determined. Similar to the aforementioned method, the item overlap rate between the two response observation fingerprints is calculated based on the exercise intersection and exercise union, which is used to represent the degree of overlap between the two response observation fingerprints in the coverage of successive exercises.
[0043] Construct a local correlation vector for each response observation fingerprint based on the intersection of the exercises, and calculate the local correlation parameters between two response observation fingerprints based on the local correlation vector.
[0044] In this embodiment, for multiple exercises included in the exercise intersection, the correct answer probability features of these exercises are extracted from the two response observation fingerprints respectively. These correct answer probability features are combined according to the order of the multiple exercises in the exercise intersection to construct a local association vector for each response observation fingerprint. This transforms two response observation fingerprints with potentially different dimensions into two equal-length local association vectors on a common exercise set. For calculating the local association parameter, the difference between the correct answer probability features of the same exercise can be calculated. The absolute values of multiple differences are taken, and the mean is calculated. The parameter obtained by subtracting the mean difference from the result is used as the local association parameter between the two response observation fingerprints.
[0045] In another implementation, those skilled in the art can also calculate the local association parameters based on parameters such as cosine similarity and Euclidean distance between local association vectors. The local association parameters are used to characterize the similarity of the correct answer probability features of two response observation fingerprints on a common problem; the larger the local association parameter, the closer the correct answer probability features of the two response observation fingerprints on a common problem.
[0046] After obtaining the item overlap rate and local correlation parameters, the two are fused to generate response correlation parameters between the two preceding behavioral segments. These response correlation parameters simultaneously characterize the degree of overlap between the two response observation fingerprints in the coverage of the subsequent exercises, as well as the similarity in the correct answer probabilities on the shared subsequent exercises. In one implementation, the product of the item overlap rate and the local correlation parameters can be used as the response correlation parameters.
[0047] Using the above fusion method, a high response association parameter will only be obtained between two response observation fingerprints if they simultaneously satisfy the conditions of having a large number of common exercises and similar correct answer probability characteristics on the common exercises. If the two response observation fingerprints have similar correct answer probability characteristics on the common exercises but a small number of common exercises, the item overlap rate will be low, and the final response association parameter will not be too high. If the two response observation fingerprints contain many common exercises but have large differences in the correct answer probability characteristics on the common exercises, the local association parameter will be low, and the final response association parameter will also be low.
[0048] For response association tags, the response association parameter can be compared with a preset association threshold. When the response association parameter is greater than or equal to the preset association threshold, the response association tag between the two preceding behavior segments is determined to be associated; when the response association parameter is less than the preset association threshold, the response association tag between the two preceding behavior segments is determined to be not associated.
[0049] It should be noted that the response association label is not a subjective classification of user ability level or question quality type, but rather a representation of the association between two preceding behavioral segments in the performance of the subsequent response. Through step S3, the relationship between different preceding behavioral segments can be elevated from whether the content of the preceding answers is similar to whether their effects on the subsequent response are similar.
[0050] Step S4: Based on the response association tags, perform state merging processing on multiple preceding behavior fragments to construct multiple response state nodes, and extract the state transition relationship between the response state nodes based on multiple behavior response samples.
[0051] In this embodiment, based on the response association tags obtained in step S3, multiple preceding behavioral segments with the same or similar follow-up response performance are grouped into the same answer state node. Each answer state node represents a type of answer state formed by the target user group in the target question bank. This answer state is not a pre-defined ability level, but is formed by the association relationship between multiple preceding behavioral segments in their follow-up response performance.
[0052] In one implementation, the process of merging multiple preceding behavior fragments based on response association tags to construct multiple response state nodes specifically includes: A prerequisite behavior fragment association graph is constructed based on multiple prerequisite behavior fragments and response association labels between any two prerequisite behavior fragments. Each prerequisite behavior fragment is treated as a fragment node. When the response association labels between two prerequisite behavior fragments indicate that they are associated, an association edge is established between the corresponding two fragment nodes; when the response association labels between two prerequisite behavior fragments indicate that they are not associated, no association edge is established.
[0053] Based on the preceding action fragment association graph, state merging is performed on multiple preceding action fragments. Specifically, multiple fragment nodes that are directly or indirectly related in the preceding action fragment association graph are divided into the same set of related fragments, and each set of related fragments is determined as a response state node. Thus, each response state node includes several preceding action fragments, and multiple preceding action fragments within the same response state node have the same or similar response association characteristics in their continuation response behavior.
[0054] For example, a connected component partitioning method is used for state merging. That is, the preceding behavior fragment association graph is traversed, interconnected fragment nodes are divided into the same connected component, and the set of preceding behavior fragments corresponding to each connected component is determined as a response state node.
[0055] In one implementation, extracting the state transition relationship between response state nodes based on multiple behavioral response samples specifically includes: Determine the response state node corresponding to the preceding behavior segment in each behavior response sample. Based on the response state nodes corresponding to multiple preceding behavior segments, construct the state transition edge corresponding to any two adjacent split points in the response result sequence.
[0056] In this embodiment, the aforementioned steps have already performed state merging processing on multiple preceding behavior fragments based on response association tags. For any behavior response sample, the answer state node corresponding to the preceding behavior fragment can be determined according to the state merging result of the preceding behavior fragment in the behavior response sample.
[0057] For any two adjacent valid segmentation points in the same response result sequence, determine the response state node corresponding to the preceding segmentation point and the response state node corresponding to the following segmentation point, respectively. If the preceding behavior segment corresponding to the preceding segmentation point belongs to the response state node Na, and the preceding behavior segment corresponding to the following segmentation point belongs to the response state node Nb, then a state transition edge can be constructed between the response state nodes Na and Nb to represent the change of the user's response state from the preceding segmentation point to the following segmentation point in the same response result sequence.
[0058] Furthermore, based on the answering units contained between two split points in the state transition edge, the exercises and corresponding answering results included in the answering unit are extracted, and the state transition conditions of the state transition edge are constructed. For example, for an answering unit e=(q,y), where q represents the exercise number and y represents the answering result, (q,y) can be used as the state transition condition C of the state transition edge. The final constructed state transition relationship consists of answering state node Na, answering state node Nb, and state transition condition C. This indicates that when a user is in answering state node Na, after completing state transition condition C, i.e., after completing exercise q and answering with result y, the updated preceding behavior segment corresponds to answering state node Nb. By performing the above processing on multiple answering result sequences from multiple users, the state transition relationships between multiple answering state nodes can be obtained.
[0059] It is worth noting that for any two adjacent valid split points, the response state node corresponding to the first split point and the response state node corresponding to the second split point may belong to the same response state node. Therefore, the constructed state transition relationship may include transitions between the same response state nodes.
[0060] Step S5: Construct a response state space corresponding to multiple response state nodes based on the state transition relationship, and perform group response state observation and analysis on the target question bank according to the response state space to generate the quality analysis results of the target question bank.
[0061] In this embodiment, the answer state space includes multiple answer state nodes and state transition relationships between answer state nodes. Answer state nodes represent different answer states of the target user group in the target question bank, and state transition relationships represent the possible changes in a user's answer state after completing a question and generating a corresponding answer.
[0062] After obtaining the state transition relationships between multiple answer state nodes, these relationships are combined to construct an answer state space corresponding to each answer state node. This answer state space can represent the distribution of the target user group's answer states in the target question bank and how these states change with the answer results.
[0063] The aforementioned answer state space can serve as the foundation for question bank quality analysis. For example, based on the number of preceding behavior fragments corresponding to each answer state node in the answer state space, the distribution of the target user group in each answer state node can be determined. Furthermore, based on the state transition edges in the answer state space, the impact of different exercises and their answer results on answer state changes can be analyzed, thus providing a data foundation for generating quality analysis results for the target question bank. This includes information such as the target question bank's coverage of the target user group's main answer states, the state transition effects of different question items in the answer state space, answer state areas that need to be supplemented or optimized, and suggestions for retaining, replacing, adjusting push priorities, or supplementing question items.
[0064] Through the above scheme, the present invention can conduct group-level observation and analysis of the answer state of the target question bank based on the answer state space formed during the user's actual answering process. This avoids relying solely on the accuracy rate of individual questions, average time taken, or manual weighted scoring to generate question bank quality analysis results, and improves the correlation between question bank quality analysis results and the actual usage process.
[0065] In one implementation, the process of observing and analyzing the collective response states of the target question bank based on the response state space to generate quality analysis results for the target question bank specifically includes: The number of preceding action fragments contained in each response state node is counted, and based on the number of preceding action fragments corresponding to each response state node, the fragment ratio of each response state node with respect to the total number of preceding action fragments contained in the response state space is calculated to obtain the state distribution data for multiple response state nodes.
[0066] Based on the answer state space, construct a set of candidate exercises for each answer state node, perform state transition effect analysis on multiple candidate exercises in the set, calculate the state transition effect parameters for each candidate exercise, and obtain the state transition effect data for multiple exercises.
[0067] In this embodiment, for any answer state node in the answer state space, the state transition edge starting from the answer state node is extracted, and according to the state transition conditions corresponding to these state transition edges, the exercise numbers that can trigger the state transition of the answer state node are extracted, and the exercises corresponding to these exercise numbers are determined as candidate exercises for the answer state node, so as to obtain the candidate exercise set corresponding to the answer state node.
[0068] The purpose of state transition analysis is to analyze the degree of state change that occurs after a user completes a candidate question at a certain answering state node. If a user can relatively stably transition to other answering states after completing a candidate question, it indicates that the candidate question has a strong state transition effect at that answering state node.
[0069] As an optional implementation process, the calculation of the state transition action parameters can be based on the difference between the response state node after the transition and the initial response state node.
[0070] For example, for a target candidate question in the candidate question set of a response status node, the state transition probability of multiple users moving to any other response status node after answering the target candidate question in the current response status node is calculated. Specifically, some users in the current response status node may move to other response status nodes or remain in the current response status node after answering the target candidate question. By calculating the proportion of users moving to any other response status node during this process, multiple state transition probabilities of the target candidate question moving from the current response status node to various response status nodes are obtained.
[0071] Building upon this, the state difference parameters between the current answering state node and any other answering state node are calculated. Specifically, the response observation fingerprints of multiple preceding behavioral segments included in each answering state node are first fused to construct the state node structure fingerprint of each answering state node. During this process, one or more correct answer probability features corresponding to each exercise are extracted and their averages are calculated to obtain the group correct answer probability features of multiple exercises. These group correct answer probability features are combined as the state node structure fingerprint of the answering state node. Then, the state difference parameters between any two answering state nodes are calculated based on the state node structure fingerprint.
[0072] To calculate the state difference parameter, first calculate the state coverage difference parameter between the structural fingerprints of state nodes, where the state coverage difference parameter = 1 - item overlap rate. Simultaneously, construct a local association vector for each structural fingerprint of state nodes, and calculate the local association parameter between two structural fingerprints of state nodes based on the local association vector. The calculation methods for item overlap rate and local association parameter are the same as described above and will not be repeated here. Finally, fuse the state coverage difference parameter and the local association parameter to obtain the state difference parameter between two answering state nodes, where the state difference parameter = state coverage difference parameter × (1 - local association parameter).
[0073] Based on the state transition probability and state difference parameter, multiple state transition probabilities are used to weight and fuse multiple state difference parameters. For any two answering state nodes, the state transition probability is used as the weight to weight and fuse the state difference parameter to calculate the state transition effect parameter of the target candidate exercise on the current answering state node.
[0074] It's worth noting that the state transition probability represents the likelihood that multiple users, after answering a target candidate question at a certain answering state node, will enter another answering state node. It reflects the direction and probability distribution of the target candidate question actually triggering a change in the user's answering state in the current answering state. If a user transitions relatively stably to another state after answering the question, it indicates that the question can update the current answering state to some extent. The state difference parameter represents the degree of difference in the subsequent response performance between two answering state nodes; that is, whether the user's subsequent answering response performance changes significantly after transitioning from the current answering state node to another answering state node. If the state response representations of the two state nodes differ significantly, the state difference parameter will be large. The state difference parameter avoids assuming the question has a strong effect simply because a state node change has occurred, but rather further determines whether the transitioned state is substantially different from the original state.
[0075] The calculated state transition effect parameter represents the overall state change that the target candidate question can cause under the current answering state node. The effect of the target candidate question under the current answering state node depends not only on which state nodes the user will transition to after answering the question, but also on the degree of difference between these target state nodes and the current state node. If the target candidate question causes the user to transition to an answering state node with a high probability of being significantly different from the current state, the corresponding state transition effect parameter is large; if the user mainly remains in the current state after answering the question, or transitions to a state node with a small difference from the current state, the corresponding state transition effect parameter is small.
[0076] Therefore, the state transition effect parameter is not used for subjective evaluation of the quality of a particular exercise, but rather to characterize the observation and updating effect of the exercise on the user's answer state at a specific answer state node. By calculating the state transition effect parameter of each candidate exercise at high-frequency answer state nodes, it is possible to determine which exercises can effectively observe the main answer states of the target user group, and then analyze the observation coverage of the target question bank on the main answer states of the target user group.
[0077] Based on the above, high-frequency state nodes in the answer state space are identified based on state distribution data. According to the candidate question set of answer state nodes and the state transition data of multiple questions, the observation coverage index of the target question bank on high-frequency state nodes is analyzed to generate the quality analysis results of the target question bank.
[0078] In this embodiment, answer status nodes with a segment percentage greater than or equal to a preset percentage threshold can be identified as high-frequency status nodes. Alternatively, answer status nodes can be selected from the top of the segment percentage ranking as high-frequency status nodes. These high-frequency status nodes are used to represent the main answer statuses that are frequently observed by the target user group when using the target question bank.
[0079] After identifying multiple high-frequency state nodes, the process of analyzing the observation coverage index of the target question bank for these high-frequency state nodes can be achieved by determining the effective observation question set corresponding to each high-frequency state node based on the state transition action parameters of each candidate question in the candidate question set corresponding to the high-frequency state node. Specifically, if the state transition action parameter of a candidate question is greater than or equal to the state transition action parameter threshold, it is determined as an effective observation question for the answering state node, thereby constructing the effective observation question sets corresponding to multiple high-frequency state nodes.
[0080] The observation coverage index of a high-frequency state node is determined based on the relationship between the set of effective observations and the set of candidate observations. For example, the proportion of effective observations in the candidate observation set is calculated. In this embodiment, the ratio between the number of effective observations and the number of candidate observations is used as the observation coverage index of the high-frequency state node.
[0081] Based on the observation coverage index corresponding to multiple high-frequency status nodes, the quality analysis results of the target question bank can be generated. This can include information such as the high-frequency status nodes corresponding to the target user group, the candidate question set corresponding to the high-frequency status nodes, the effective observation question set corresponding to the high-frequency status nodes, and the observation coverage index of the high-frequency status nodes. This quality analysis result can be used to indicate the status areas in the target question bank that need to be supplemented with questions, replaced with questions, or have their push priority adjusted, thereby providing data basis for question bank maintenance and question item optimization.
[0082] It is worth noting that this invention converts the continuous answering behavior of multiple users into an answer state space, and analyzes the observation coverage of the target question bank on the main answering states of the target user group based on this answer state space. Compared with the method of analyzing the quality of the question bank based solely on the accuracy, difficulty, or discrimination of individual questions, this invention can reflect the effect of questions in different answering states on the changes in users' subsequent answering states, thereby identifying areas in the question bank that are insufficiently covered for high-frequency answering states. Therefore, it avoids judging the quality of the question bank solely based on the static statistical results of individual questions, and improves the correlation between the question bank quality analysis results and the actual continuous answering process.
[0083] Please see Figure 2 The second aspect of the present invention provides a question bank quality analysis system based on answer behavior data, specifically used to implement the above-mentioned question bank quality analysis method based on answer behavior data, the system comprising: The answer sequence construction module is used to collect answer record data of multiple users in the target user group about the target question bank, and extract one or more answer result sequences corresponding to each user from the answer record data. The answer result sequence consists of multiple answer units arranged in the answer order. The behavior response sample construction module is used to perform dual-window fragmentation processing based on segmentation points on each response result sequence, extract the preceding behavior fragments and subsequent response fragments corresponding to multiple segmentation points, and the preceding behavior fragments and subsequent response fragments constitute the behavior response samples corresponding to the segmentation points. The response association analysis module is used to perform continuation response analysis on each behavior response sample, generate response observation fingerprints for each preceding behavior segment, perform response association detection on multiple preceding behavior segments based on the response observation fingerprints, and determine the response association label between any two preceding behavior segments. The state transition analysis module is used to perform state merging processing on multiple preceding behavior fragments based on response association labels to construct multiple response state nodes, where each response state node includes several preceding behavior fragments, and extracts the state transition relationship between response state nodes based on multiple behavior response samples. The question bank quality analysis module is used to construct a question state space corresponding to multiple question state nodes based on the state transition relationship, and to perform group question state observation and analysis on the target question bank based on the question state space, thereby generating the quality analysis results of the target question bank.
[0084] The above are merely specific embodiments of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art. Parts not described in detail in this specification are prior art known to those skilled in the art.
Claims
1. A method for analyzing the quality of a question bank based on answer behavior data, characterized in that, include: Collect answer records of multiple users in the target user group regarding the target question bank, and extract one or more answer result sequences for each user from the answer record data. The answer result sequence consists of multiple answer units arranged in the order of answering. Each response sequence is processed by a dual-window fragmentation based on the segmentation point, and the preceding behavior fragments and subsequent response fragments corresponding to multiple segmentation points are extracted. The preceding behavior fragments and subsequent response fragments constitute the behavior response samples corresponding to the segmentation points. For each behavioral response sample, a follow-up response analysis is performed to generate a response observation fingerprint for each preceding behavioral segment. Based on the response observation fingerprint, response association detection is performed on multiple preceding behavioral segments to determine the response association label between any two preceding behavioral segments. Based on response association tags, state merging is performed on multiple preceding behavior fragments to construct multiple response state nodes, where each response state node includes several preceding behavior fragments. The state transition relationship between response state nodes is extracted based on multiple behavior response samples. Based on the state transition relationship, a response state space corresponding to multiple response state nodes is constructed. Based on the response state space, a group response state observation and analysis of the target question bank is performed to generate the quality analysis results of the target question bank.
2. The question bank quality analysis method based on answer behavior data according to claim 1, characterized in that, For each behavioral response sample, a follow-up response analysis is performed to generate a response observation fingerprint for each preceding behavioral segment, including: Based on the multiple exercises included in the preceding action segment, extract the intersection and union of exercises between any two preceding action segments, calculate the ratio between the number of exercises in the intersection and the number of exercises in the union, and obtain the item overlap rate between the two preceding action segments. Based on the answer results for each exercise, calculate the answer matching degree between two preceding action segments with respect to the intersection of the exercises, and determine the action matching weight between any two preceding action segments based on the question overlap rate and the answer matching degree. Extract the correct answer probability features for each question in the continuation response segment of the behavioral response sample, and construct the continuation response fingerprint of the preceding behavioral segment based on the correct answer probability features of the multiple questions. For the extraction of correct answer probability features, multiple behavior matching samples are identified for each behavior response sample based on the behavior matching weight. Based on the multiple exercises included in the continuation response fragment in each behavior matching sample, a reference response observation sample set for each exercise in the continuation response fragment of the behavior response sample is constructed. The correct answer probability features for the corresponding exercise are obtained statistically based on the reference response observation sample set.
3. The question bank quality analysis method based on answer behavior data according to claim 2, characterized in that, Based on the state transition relationship, the response state space corresponding to multiple response state nodes is constructed, including: For the extraction of state transition relations, the answer state node corresponding to the preceding behavior segment in each behavior response sample is determined. Based on the answer state nodes corresponding to multiple preceding behavior segments, the state transition edge corresponding to any two adjacent split points in the answer result sequence is constructed. Based on the answering units contained between two split points in the state transition edge, extract the exercises and corresponding answering results contained in the answering units, construct the state transition conditions of the state transition edge, and obtain the state transition relationship between any two answering state nodes. By combining the state transition relationships between multiple response state nodes, a response state space corresponding to multiple response state nodes is constructed.
4. The question bank quality analysis method based on answer behavior data according to claim 2, characterized in that, Based on response observation fingerprints, response association detection is performed on multiple preceding behavior fragments to determine the response association labels between any two preceding behavior fragments, including: Extract the intersection and union of the problems between the response observation fingerprints corresponding to any two preceding behavior segments, and calculate the item overlap rate between the two response observation fingerprints based on the intersection and union of the problems; Construct a local correlation vector for each response observation fingerprint based on the intersection of exercises, and calculate the local correlation parameters between two response observation fingerprints based on the local correlation vector; The response association parameters between two preceding action segments are generated by integrating the overlap rate of the items and the local association parameters, and the response association labels between the two preceding action segments are determined based on the response association parameters.
5. The question bank quality analysis method based on answer behavior data according to claim 3, characterized in that, Based on the response state space, a group response state observation and analysis of the target question bank is performed, generating quality analysis results for the target question bank, including: The percentage of preceding behavior segments contained in each response state node is counted to obtain state distribution data for multiple response state nodes; Based on the answer state space, construct a set of candidate exercises for each answer state node, perform state transition effect analysis on multiple candidate exercises in the set, calculate the state transition effect parameters for each candidate exercise, and obtain the state transition effect data for multiple exercises; Based on state distribution data, high-frequency state nodes in the answer state space are identified. According to the candidate question set of answer state nodes and the state transition data of multiple questions, the observation coverage index of the target question bank on high-frequency state nodes is analyzed to generate the quality analysis results of the target question bank.
6. The question bank quality analysis method based on answer behavior data according to claim 5, characterized in that, The state transition action parameters and observation coverage index also include: For a target candidate question in the candidate question set of a response status node, the state transition probability of multiple users transitioning to any other response status node after answering the target candidate question under the current response status node is calculated, and the state difference parameter between the current response status node and any other response status node is calculated. Based on the state transition probability and the state difference parameter, the state transition effect parameter of the target candidate question on the current response status node is calculated. For the observation coverage index, multiple valid observation problems in the candidate problem set of the answering state node are identified based on the state transition action parameter. The proportion of valid observation problems in the candidate problem set is calculated to obtain the observation coverage index of the high-frequency state node.
7. A question bank quality analysis system based on answer behavior data, characterized in that, The system is used to implement the question bank quality analysis method based on answer behavior data as described in any one of claims 1-6, including: The answer sequence construction module is used to collect answer record data of multiple users in the target user group about the target question bank, and extract one or more answer result sequences corresponding to each user from the answer record data. The answer result sequence consists of multiple answer units arranged in the answer order. The behavior response sample construction module is used to perform dual-window fragmentation processing based on segmentation points on each response result sequence, extract the preceding behavior fragments and subsequent response fragments corresponding to multiple segmentation points, and the preceding behavior fragments and subsequent response fragments constitute the behavior response samples corresponding to the segmentation points. The response association analysis module is used to perform continuation response analysis on each behavior response sample, generate response observation fingerprints for each preceding behavior segment, perform response association detection on multiple preceding behavior segments based on the response observation fingerprints, and determine the response association label between any two preceding behavior segments. The state transition analysis module is used to perform state merging processing on multiple preceding behavior fragments based on response association labels to construct multiple response state nodes, where each response state node includes several preceding behavior fragments, and extracts the state transition relationship between response state nodes based on multiple behavior response samples. The question bank quality analysis module is used to construct a question state space corresponding to multiple question state nodes based on the state transition relationship, and to perform group question state observation and analysis on the target question bank based on the question state space, thereby generating the quality analysis results of the target question bank.