Education evaluation method and system based on visual identification data analysis

By constructing explanatory mind maps and using multimodal data analysis, combined with students' facial expressions and body movements, the problem of insufficient objectivity and accuracy in existing educational assessments has been solved, enabling real-time, personalized teaching assessments and improvement guidance.

CN120997008APending Publication Date: 2025-11-21SHENZHEN POLYTECHNIC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511116635.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-08-13
Filing Date
2025-08-11
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing educational assessment methods lack objectivity and comprehensiveness, making it difficult to capture students' dynamic responses in the teaching process in real time, failing to adapt to personalized teaching needs, and lacking accuracy in visual recognition during multimodal data fusion, resulting in unreal-time adjustments to the difficulty of knowledge points.

Method used

We construct mind maps for explanation, preset the level of understanding, combine semantic analysis and multimodal data analysis to obtain real-time classroom information, assess students' status through facial expressions and body movements, and make adjustments based on the level of understanding of the mind map nodes to ultimately form a teaching assessment effect.

Benefits of technology

It enables real-time and accurate teaching assessment, provides actionable improvement guidance, and enhances the comprehensiveness and accuracy of classroom feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997008A_ABST
    Figure CN120997008A_ABST
Patent Text Reader

Abstract

The invention discloses an education evaluation method and system based on visual identification data analysis, and relates to the technical field of education analysis, and the method comprises the steps: presetting the understanding difficulty for each mapping node, and taking the understanding difficulty as a teaching path reference; performing semantic analysis on the explanation scheme, extracting knowledge points and secondary and supplementary key phrases, and associating the knowledge points and the secondary and supplementary key phrases to corresponding mapping nodes to realize content structuring; acquiring real-time classroom voice information, extracting and analyzing characters to determine real-time keywords, positioning corresponding nodes in the mind map, sequentially connecting the nodes to form a real-time map node sequence line, and capturing an actual explanation path; calculating student initial state evaluation, combining node difficulty correction to obtain final evaluation, and accumulating all student evaluation as a teaching effect; taking a preset explanation order line teaching effect corresponding to the real-time order line as a standard, determining the relative positioning of actual evaluation, and achieving deviation quantification; according to the invention, through multi-modal fusion and a dynamic order line mechanism, the real-time performance and accuracy of evaluation are improved, and teaching improvement guidance is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of educational analytics, and in particular to educational assessment methods and systems based on visual recognition data analysis. Background Technology

[0002] In the current field of educational assessment, evaluating classroom teaching effectiveness has become a crucial link in improving teaching quality. With the rapid development of artificial intelligence and big data technologies, traditional assessment methods have revealed many shortcomings. Traditional classroom assessment mainly relies on teachers' subjective observations, student feedback, or final exam scores. These methods are often limited by human bias and sample size, failing to capture students' dynamic responses in real time during the teaching process, resulting in assessment results lacking objectivity and comprehensiveness. For example, early assessment systems only provided feedback through simple questionnaires or grade statistics, ignoring the real-time interaction of students' emotions, behaviors, and knowledge absorption, making it difficult to adapt to personalized teaching needs.

[0003] In recent years, multimodal data analytics has been increasingly used in educational assessment, integrating multi-source data such as voice, video, and text to achieve a more comprehensive analysis of classroom behavior. For example, some smart classroom systems utilize multimodal learning analytics models to extract features from students' facial expressions, postures, and voices, constructing a teaching behavior analysis framework to optimize teacher-student interaction and learning paths. Knowledge graphs, as a core tool, are widely used in educational informatization, representing the sequential relationships of knowledge points through entity-relationship modeling, and supporting intelligent search and recommendation systems.

[0004] However, existing multimodal analysis and knowledge graph technologies still have significant limitations. First, when dealing with complex relationships, knowledge graphs lack sufficient modeling capabilities to fully capture dynamic interactions, leading to real-time biases in assessments that overlook the teaching path. Second, while multimodal data fusion can extract facial expression and posture features, existing methods struggle to integrate the complex relationships between text and visual information. Particularly in visual recognition, the accuracy of edge detection and matching models is insufficient, increasing errors in student status assessment. Furthermore, there is a lack of subjective correction mechanisms for knowledge point difficulty; current systems mostly rely on static thresholds for assessment, failing to adjust response weights in real time. These limitations make it difficult for existing technologies to provide accurate and real-time feedback in complex classroom scenarios. Summary of the Invention

[0005] The purpose of this invention is to provide a more accurate and effective educational assessment method and system.

[0006] This invention discloses an educational assessment method based on visual recognition data analysis, comprising:

[0007] Step S100: Construct a mind map for the classroom explanation plan, preset the understanding difficulty of each mind map node, and determine the explanation order of the mind map nodes. Based on the constructed explanation order, connect the mind map nodes sequentially to obtain the explanation order line.

[0008] Step S200: Perform semantic analysis on the classroom explanation plan, identify several explanation keywords, and associate the explanation keywords with the corresponding mind map nodes;

[0009] Step S300: Obtain real-time classroom audio information, extract real-time explanation text information from the real-time classroom audio information, perform semantic analysis on the extracted real-time explanation text information, determine real-time explanation keywords, and determine the corresponding mind map nodes based on several real-time explanation keywords that appear in a preset time period, and connect the determined mind map nodes in sequence to obtain the real-time mind map node sequence line.

[0010] Step S400: Obtain real-time classroom video information, analyze students' facial expressions and body movements in the real-time classroom information, determine the initial state assessment of a single student, and make the first correction to the initial state assessment based on the understanding difficulty of different mind map nodes to obtain the final state assessment of the student. The sum of the final state assessments of all students is recognized as the teaching assessment effect.

[0011] Step S500: Using the preset teaching evaluation effect corresponding to the explanation sequence line corresponding to the real-time map node sequence line as the standard, determine the relative positioning of the teaching evaluation effect;

[0012] The method for analyzing students' facial expressions includes using a facial mapping model to determine a real-time facial mapping template for facial expressions, and matching the real-time facial mapping template with a standard facial mapping template to determine the expression type.

[0013] Methods for analyzing students' body movements include using body movement analysis models to determine the movement types of body movements.

[0014] In some embodiments disclosed in this invention, the method for semantic analysis of classroom teaching schemes includes:

[0015] Pre-define the key words of the knowledge points involved in the classroom teaching plan;

[0016] Based on the position of the knowledge point keywords in the text of the classroom explanation plan, the secondary keywords that have semantic relationships before and after the knowledge point keywords are identified, and based on the semantic relationships, the knowledge point keywords and secondary keywords are combined to obtain the explanation keyword group.

[0017] The scope of the text corresponding to the explanation keyword group is determined, and other supplementary keywords with a frequency greater than or equal to the preset value in the text information are identified. The supplementary keywords are added to the explanation keyword group, and the explanation keyword group is associated with the corresponding mind map node.

[0018] In some embodiments disclosed in this invention, a method for generating corresponding mind map nodes based on several explanatory keywords appearing within a preset time period includes:

[0019] Identify the mind map nodes that have already been explained, and use the mind map nodes that have just been explained as the current positioning nodes. Then, select the mind map nodes after the current positioning nodes as the comparison nodes.

[0020] The real-time explanation keywords appearing within the preset time period are compared with the explanation keyword groups corresponding to the comparison map nodes to determine the mapping number of real-time explanation keywords and different explanation keyword groups, and the ratio of the mapping number to the number of real-time explanation keywords within the preset time period is calculated and recorded as the comprehensive mapping ratio.

[0021] Based on the comprehensive mapping ratio and the ratio of various keywords mapped in the explanation keyword group, the mapping matching parameter between the real-time explanation keywords and the explanation keyword group appearing within the preset time period is determined. Based on the magnitude of the mapping matching parameter, the explanation keyword group corresponding to the real-time keyword is determined, and the mind map node corresponding to the explanation keyword group is identified as the corresponding mind map node of the real-time keyword.

[0022] In some embodiments disclosed in this invention, the method for determining the mapping match parameter between real-time explanation keywords and explanation keyword groups appearing within a preset time period includes:

[0023] Determine the number of first mappings of the knowledge point keywords mapped in the explanation keyword group, the number of second mappings of the mapped secondary keywords, and the number of third mappings of the mapped supplementary keywords;

[0024] Calculate the first mapping ratio between the first mapping quantity and the real-time keyword quantity within a preset time period; calculate the second mapping ratio between the second mapping quantity and the real-time keyword quantity within a preset time period; calculate the third mapping ratio between the third mapping quantity and the real-time keyword quantity within a preset time period.

[0025] Based on the first mapping ratio, the second mapping ratio, the third mapping ratio, and the comprehensive mapping ratio, the mapping matching parameters are determined;

[0026] The expression for calculating the mapping matching parameters is as follows:

[0027] Y = R × U zong ×exp{[K1×U1+K2×U2+K3×U3]+b};

[0028] Where Y is the mapping matching parameter, R is the mapping matching parameter transformation coefficient, and U zong The comprehensive mapping ratio is defined as follows: U1 is the first mapping ratio, U2 is the second mapping ratio, U3 is the third mapping ratio, K1 is the knowledge point keyword weight adjustment coefficient, K2 is the secondary keyword weight adjustment coefficient, K3 is the supplementary keyword weight adjustment coefficient, and b is the mapping ratio adjustment constant.

[0029] In some embodiments disclosed in this invention, the method for determining a real-time facial mapping template for facial expressions using a facial mapping model includes:

[0030] The system uses a pre-defined facial recognition model to identify the student's facial region and, based on the identified facial region, determines the real-time facial feature blocks within that region, including eye feature blocks, nose feature blocks, eyebrow feature blocks, and mouth feature blocks.

[0031] A planar facial block representation output model is constructed. Using the planar facial block representation output model, facial feature blocks are analyzed to determine the corresponding planar facial block representation. The planar facial block representation is then configured in a preset facial mapping template to obtain a real-time facial mapping template.

[0032] The methods for constructing a planar facial region representation output model include:

[0033] Edge detection technology is used to detect edges in historical facial feature blocks to form historical facial edge feature blocks. Important edges in the historical facial edge feature blocks are delineated, and undelineated edges are removed.

[0034] Construct several planar facial block representations and establish the correlation between the planar facial block representations and historical facial edge feature blocks;

[0035] Edge detection technology is used to detect edges in real-time facial feature blocks to form real-time facial edge feature blocks. The real-time facial edge feature blocks are compared with different historical facial edge feature blocks to determine the edge matching degree of each historical facial edge feature block participating in the comparison. Based on the edge matching degree, the matching historical facial edge feature blocks are selected, and the corresponding planar facial block representation is called for output.

[0036] The expression for calculating the degree of edge matching is as follows:

[0037]

[0038] Where W represents the degree of edge fit, μ xLet μ be the matching parameter for the x-th important edge. If the x-th important edge is matched by the edge mapping in the real-time facial edge feature block, then μ... x Output 1 otherwise output 0, where L is the edge mapping matching effect adjustment coefficient and D is the edge mapping matching effect adjustment constant.

[0039] The method for determining whether an important edge matches the edge mapping in the real-time facial edge feature block includes: pushing the important edge to coincide with the edge in the real-time facial edge feature block, calculating the intersection area between them, and calculating the ratio of the intersection area to the edge length of the important edge. If the ratio is less than or equal to a preset value, it is determined that the important edge matches the edge mapping in the real-time facial edge feature block.

[0040] In some embodiments disclosed in this invention, the method for analyzing students' facial expressions and body movements in real-time classroom video information to determine the initial state assessment of a single student includes:

[0041] Facial expressions include focused learning expressions, confused expressions, and unfocused learning expressions; body movements include static upright state, static unupright state, and active unupright state. Each facial expression and each body movement is equipped with a specific unit evaluation parameter.

[0042] We conducted continuous time-point analysis on each student's facial expressions and body movements, and based on the analysis results, we determined the initial state assessment for each student.

[0043] In some embodiments disclosed in this invention, the method for making the first correction to the initial state evaluation, taking into account the difficulty of understanding different map nodes, includes:

[0044] To address the teaching process in class, a time progress reference line is established. Combinations of facial expression factors and body movement factors determined at different time points are recorded as state factor groups, and these state factor groups are mapped onto the time progress reference line according to their corresponding time points.

[0045] Based on the time nodes corresponding to the mind map nodes explained in class, time segments are marked on the time progress reference line, and the difficulty of understanding the mind map nodes is associated with the corresponding time segments.

[0046] For each state factor, an initial sub-state evaluation is performed, and based on the time segment to which the corresponding time node of the state factor belongs, the correction features for the initial sub-state evaluation are determined.

[0047] In some embodiments disclosed in this invention, the expression for calculating the initial state assessment of a single student is as follows:

[0048]

[0049] Where P is the initial state assessment of a single student, and α t β is the unit evaluation parameter for facial expression at the i-th time point. t Let δ[t] be the unit evaluation parameter for the limb movement at the i-th time node, and let δ[t] be the judgment function for continuous time nodes. If there exists α for each continuous time node... t +β t If the difficulty of the t-th time node is greater than or equal to the preset value, the number of consecutive time nodes is determined, and the largest number of consecutive time nodes is selected. δ[t] is the preset continuity correction parameter based on the largest number of consecutive time nodes. H[...] is the difficulty judgment function. If the difficulty of the t-th time node is greater than or equal to the preset value, and α... t +β t If the value is negative, H[...] will output 0. If the difficulty at the t-th time node is less than or equal to the preset value, H[...] will output α. t +β t If the difficulty at time point t is greater than or equal to the preset value, then H[...] will output [α]. t +β t ]×h, where h is the preset difficulty multiplier amplification coefficient.

[0050] In some embodiments disclosed in this invention, an educational assessment system based on visual recognition data analysis is also disclosed, including:

[0051] The first module is used to construct a mind map for classroom explanation plans, preset the understanding difficulty of each mind map node, and define the explanation order of the mind map nodes. Based on the constructed explanation order, the mind map nodes are connected sequentially to obtain the explanation order line.

[0052] The second module is used to perform semantic analysis on the classroom explanation plan, identify several explanation keywords, and associate the explanation keywords with the corresponding mind map nodes.

[0053] The third module is used to acquire real-time classroom audio information, extract real-time explanation text information from the real-time classroom audio information, perform semantic analysis on the extracted real-time explanation text information, determine real-time explanation keywords, and determine the corresponding mind map nodes based on several real-time explanation keywords that appear in a preset time period, and connect the determined mind map nodes in sequence to obtain the real-time mind map node sequence line.

[0054] The fourth module is used to acquire real-time classroom video information, analyze students' facial expressions and body movements in the real-time classroom information, determine the initial state assessment of a single student, and make the first correction to the initial state assessment based on the understanding difficulty of different mind map nodes, so as to obtain the final state assessment of the student. The sum of the final state assessments of all students is recognized as the teaching assessment effect.

[0055] The fifth module is used to determine the relative positioning of the teaching evaluation effect by using the preset teaching evaluation effect corresponding to the explanation sequence line corresponding to the real-time map node sequence line as the standard.

[0056] The method for analyzing students' facial expressions includes using a facial mapping model to determine a real-time facial mapping template for facial expressions, and matching the real-time facial mapping template with a standard facial mapping template to determine the expression type.

[0057] Methods for analyzing students' body movements include using body movement analysis models to determine the movement types of body movements.

[0058] This invention discloses an educational assessment method and system based on visual recognition data analysis, belonging to the field of educational analysis technology. It presets the comprehension difficulty for each mind map node as a benchmark for the teaching path; performs semantic analysis on the explanation plan, extracting knowledge points, secondary and supplementary keyword groups, and associating them with corresponding mind map nodes to achieve content structuring; acquires real-time classroom audio information, extracts and analyzes the text to determine real-time keywords, locates corresponding nodes in the mind map, and sequentially connects them to form a real-time mind map node sequence line, capturing the actual explanation path; calculates the initial student state assessment, combines it with node difficulty correction to obtain the final assessment, and accumulates all student assessments as the teaching effect; uses the teaching effect of the preset explanation sequence line corresponding to the real-time sequence line as the standard to determine the relative positioning of the actual assessment, achieving deviation quantification; this invention improves the real-time performance and accuracy of assessment through multimodal fusion and dynamic sequence line mechanisms, providing guidance for teaching improvement.

[0059] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0060] Figure 1 This is a flowchart illustrating the steps of an educational assessment method based on visual recognition data analysis disclosed in some embodiments of the present invention. Detailed Implementation

[0061] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0062] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings and specific embodiments. It should be understood that the preferred embodiments described herein are only for illustration and explanation of the present invention and should not be construed as limiting the scope of protection of the present invention. Those skilled in the art can make some non-essential improvements and adjustments based on the following content of the present invention. In the present invention, unless otherwise expressly specified and limited, the technical terms used in the present invention should have the ordinary meaning understood by those skilled in the art.

[0063] Example:

[0064] This invention discloses an educational assessment method based on visual recognition data analysis. (See attached document.) Figure 1 ,include:

[0065] Step S100: Construct a mind map for the classroom explanation plan, preset the understanding difficulty of each mind map node, and determine the explanation order of the mind map nodes. Based on the constructed explanation order, connect the mind map nodes sequentially to obtain the explanation order line.

[0066] By constructing mind maps for classroom instruction, abstract teaching content is transformed into a radial hierarchical structure. Branch nodes expand from a central theme, with each node pre-defined to quantify the cognitive challenge of the knowledge point. The order of explanation for each node in the mind map is also defined to ensure the sequential and coherent teaching logic. Nodes are then connected sequentially based on this structured order to form an explanation sequence line. This design not only provides a visual benchmark for the teaching path but also facilitates subsequent dynamic comparison and difficulty adjustments, enabling structured and personalized optimization of the teaching plan.

[0067] Step S200: Perform semantic analysis on the classroom explanation plan, identify several explanation keywords, and associate the explanation keywords with the corresponding mind map nodes.

[0068] Semantic analysis of classroom teaching plans is performed by extracting knowledge points, secondary and supplementary keyword phrases and associating them with corresponding mind map nodes. This ensures the seamless integration of in-depth semantic analysis of teaching content with mind maps, avoids the superficial limitations of simple text processing, achieves refined representation of knowledge points, and lays a reliable foundation for real-time voice and video analysis.

[0069] Step S300: Obtain real-time classroom audio information, extract real-time explanation text information from the real-time classroom audio information, perform semantic analysis on the extracted real-time explanation text information, determine real-time explanation keywords, and determine the corresponding mind map nodes based on several real-time explanation keywords appearing in a preset time period, and connect the determined mind map nodes in sequence to obtain the real-time mind map node sequence line.

[0070] After acquiring real-time classroom audio information, real-time keywords are identified through text extraction and semantic analysis. Corresponding nodes are then located in a mind map based on preset time periods, and these nodes are sequentially connected to form a real-time mind map node sequence line. This mechanism uses mapping matching parameters to quantify the matching degree, captures the dynamic evolution of the actual explanation path, provides real-time visualization of progress tracking, and avoids the lag problem of static analysis.

[0071] Step S400: Obtain real-time classroom video information, analyze students' facial expressions and body movements in the real-time classroom information, determine the initial state assessment of a single student, and make the first correction to the initial state assessment based on the understanding difficulty of different mind map nodes to obtain the final state assessment of the student. The sum of the final state assessments of all students is recognized as the teaching assessment effect.

[0072] After acquiring real-time classroom video information, facial mapping models and body analysis models are used to determine expression types (e.g., focused learning, confusion) and action types (e.g., still and upright). An initial state assessment for each student is calculated, and this is corrected by considering the difficulty of understanding mind map nodes, resulting in a final assessment. All student assessments are then accumulated as the overall teaching effectiveness. This multimodal visual fusion mechanism achieves objective assessment of student responses through quantitative correction, improving the accuracy and comprehensiveness of classroom feedback.

[0073] Step S500: Using the preset teaching evaluation effect corresponding to the explanation sequence line corresponding to the real-time mind map node sequence line as the standard, determine the relative positioning of the teaching evaluation effect.

[0074] Using the explanation sequence line corresponding to the real-time mind map node sequence line as a benchmark, and introducing preset teaching evaluation effectiveness standards, the relative positioning of the actual teaching evaluation effectiveness is determined by comparison. This relative positioning mechanism, based on path deviation analysis, quantifies the degree of conformity between the explanation and the preset plan, provides actionable guidance for teaching improvement, and avoids the limitations of absolute evaluation.

[0075] The method for analyzing students' facial expressions includes using a facial mapping model to determine a real-time facial mapping template for facial expressions, and matching the real-time facial mapping template with a standard facial mapping template to determine the expression type.

[0076] Methods for analyzing students' body movements include using body movement analysis models to determine the movement types of body movements.

[0077] In some embodiments disclosed in this invention, the method for semantic analysis of classroom teaching schemes includes:

[0078] Key words for knowledge points involved in the classroom teaching plan are pre-defined.

[0079] By pre-identifying core knowledge keywords manually or through algorithms, a semantic foundation for teaching content is established, ensuring that subsequent analysis focuses on key concepts, avoids interference from irrelevant information, and achieves efficient knowledge extraction and structuring.

[0080] Based on the position of the knowledge point keywords in the text of the classroom explanation plan, the secondary keywords that have semantic relationships before and after the knowledge point keywords are identified. Based on the semantic relationships, the knowledge point keywords and secondary keywords are combined to obtain the explanation keyword group.

[0081] By leveraging location context analysis to capture semantic dependencies, such as dependency parsing or co-occurrence statistics, secondary keywords can be selected and combined to form extended groups, thereby enhancing the semantic richness of keywords and ensuring that the groups represent complete knowledge units.

[0082] The scope of the text corresponding to the explanation keyword group is determined, and other supplementary keywords with a frequency greater than or equal to the preset value in the text information are identified. The supplementary keywords are added to the explanation keyword group, and the explanation keyword group is associated with the corresponding mind map node.

[0083] By filtering highly relevant supplementary word extension groups through frequency thresholds and associating them with mind map nodes, the semantic groups are mapped to the structured graph, ensuring the integrity and hierarchy of knowledge representation and supporting dynamic evaluation.

[0084] In some embodiments disclosed in this invention, a method for generating corresponding mind map nodes based on several explanatory keywords appearing within a preset time period includes:

[0085] Identify the mind map nodes that have already been explained, and use the nodes that have just been explained as the current positioning nodes. Then, select several other mind map nodes after the current positioning nodes as comparison nodes.

[0086] By tracking the progress of the explanation to locate the current node, and filtering subsequent nodes forward as the comparison range, the matching is ensured to focus on logically adjacent parts, avoiding the inefficiency of global search and improving the targeting and accuracy of real-time positioning.

[0087] By tracking the progress of the explanation to locate the current node, and filtering subsequent nodes forward as the comparison range, the matching is ensured to focus on logically adjacent parts, avoiding the inefficiency of global search and improving the targeting and accuracy of real-time positioning.

[0088] The real-time explanation keywords appearing within the preset time period are compared with the explanation keyword groups corresponding to the comparison map nodes to determine the mapping number of real-time explanation keywords and different explanation keyword groups, and the ratio of the mapping number to the number of real-time explanation keywords within the preset time period is calculated and recorded as the comprehensive mapping ratio.

[0089] By comparing keyword groups, the mapping coverage is calculated, the overall overlap between real-time content and preset groups is quantified, preliminary matching indicators are provided, and subsequent fine-grained evaluation is supported.

[0090] Based on the comprehensive mapping ratio and the ratio of various keywords mapped in the explanation keyword group, the mapping matching parameter between the real-time explanation keywords and the explanation keyword group appearing within the preset time period is determined. Based on the magnitude of the mapping matching parameter, the explanation keyword group corresponding to the real-time keyword is determined, and the mind map node corresponding to the explanation keyword group is identified as the corresponding mind map node of the real-time keyword.

[0091] By integrating the overall and category ratios to calculate the matching parameters, the highest value is selected to correspond to the group and node, achieving accurate mapping, avoiding bias from a single indicator, and ensuring the reliability of node identification.

[0092] In some embodiments disclosed in this invention, the method for determining the mapping match parameter between real-time explanation keywords and explanation keyword groups appearing within a preset time period includes:

[0093] Determine the number of first mappings of the knowledge point keywords mapped in the explanation keyword group, the number of second mappings of the mapped secondary keywords, and the number of third mappings of the mapped supplementary keywords.

[0094] Calculate the first mapping ratio between the first mapping quantity and the real-time keyword quantity within a preset time period, calculate the second mapping ratio between the second mapping quantity and the real-time keyword quantity within a preset time period, and calculate the third mapping ratio between the third mapping quantity and the real-time keyword quantity within a preset time period.

[0095] Based on the first mapping ratio, the second mapping ratio, the third mapping ratio, and the comprehensive mapping ratio, the mapping matching parameters are determined.

[0096] The expression for calculating the mapping matching parameters is as follows:

[0097] Y = R × U zong ×exp{[K1×U1+K2×U2+K3×U3]+b}.

[0098] Where Y is the mapping matching parameter, R is the mapping matching parameter transformation coefficient, and U zong The comprehensive mapping ratio is defined as follows: U1 is the first mapping ratio, U2 is the second mapping ratio, U3 is the third mapping ratio, K1 is the knowledge point keyword weight adjustment coefficient, K2 is the secondary keyword weight adjustment coefficient, K3 is the supplementary keyword weight adjustment coefficient, and b is the mapping ratio adjustment constant.

[0099] This expression quantifies the degree of matching between real-time explanation keywords and explanation keyword groups within a preset time period. It aims to assess the semantic consistency and completeness of classroom explanation content, supporting dynamic positioning and correction in educational assessment. The formula design is based on an exponential weighting model, combining linear weighting and exponential amplification mechanisms to ensure that the matching parameter Y reflects both the overall matching degree and amplifies key differences, achieving a non-linear sensitive response.

[0100] In some embodiments disclosed in this invention, the method for determining a real-time facial mapping template for facial expressions using a facial mapping model includes:

[0101] The system uses a pre-defined facial recognition model to identify the student's facial region and, based on the identified facial region, determines real-time facial feature blocks within that region, including eye feature blocks, nose feature blocks, eyebrow feature blocks, and mouth feature blocks.

[0102] The student's face bounding box is detected and located from video frames by a pre-trained facial recognition model (such as a deep learning-based CNN or Haar cascade). Then, key feature blocks are further segmented within this region. These blocks correspond to facial anatomical standards (such as eyes, nose, eyebrows, and mouth). Real-time features are extracted using edge detection or key point localization algorithms to ensure that subsequent expression analysis focuses on expression-sensitive areas, avoids interference from global image noise, and achieves efficient local feature extraction.

[0103] A planar facial block representation output model is constructed. Using the planar facial block representation output model, facial feature blocks are analyzed to determine the corresponding planar facial block representation. The planar facial block representation is then configured in a preset facial mapping template to obtain a real-time facial mapping template.

[0104] First, an output model (such as a machine learning-based classifier or generative model) is built to simplify historical feature blocks into a planar representation (such as a geometric shape or vector representation). Then, the model is applied to real-time feature blocks for analysis, outputting a simplified representation and mapping it onto a preset template to form a standardized real-time template. This mechanism simplifies complex facial data through dimensionality reduction and standardization, supports subsequent matching and expression classification, and improves computational efficiency and robustness.

[0105] The methods for constructing a planar facial region representation output model include:

[0106] Edge detection technology is used to detect edges in historical facial feature blocks to form historical facial edge feature blocks. Important edges in the historical facial edge feature blocks are delineated, and undelineated edges are removed.

[0107] By extracting pixel gradient changes from historical facial feature block images using classic edge detection algorithms (such as the Canny or Sobel operators), historical edge blocks representing edge contours are formed. Then, important edges (such as key lines of the eye contour) are identified based on preset criteria. This preprocessing mechanism ensures that historical data is simplified to core features, improving the efficiency and accuracy of subsequent model training and comparison, and avoiding the computational burden caused by redundant information.

[0108] Construct several planar facial block representations and establish the correlation between the planar facial block representations and historical facial edge feature blocks.

[0109] By using simplified geometric models (such as polygons or vector representations) to create planar block representations that represent abstract expressions of facial features, and then using association algorithms (such as hash mapping or clustering) to connect these planar representations with the processed historical edge blockchain, a one-to-many or many-to-one correspondence is formed. This construction mechanism aims to build an efficient feature library that supports fast retrieval and output, ensuring that standardized representations can be called to simulate complex expressions during real-time analysis, avoiding the complexity of directly processing high-dimensional image data.

[0110] Edge detection technology is used to detect edges in real-time facial feature blocks to form real-time facial edge feature blocks. The real-time facial edge feature blocks are compared with different historical facial edge feature blocks to determine the edge matching degree of each historical facial edge feature block participating in the comparison. Based on the edge matching degree, the matching historical facial edge feature blocks are selected, and the corresponding planar facial block representation is called for output.

[0111] The expression for calculating the degree of edge matching is as follows:

[0112]

[0113] Where W represents the degree of edge fit, μ x Let μ be the matching parameter for the x-th important edge. If the x-th important edge is matched by the edge mapping in the real-time facial edge feature block, then μ... x Output 1 otherwise output 0. L is the edge mapping matching effect adjustment coefficient, and D is the edge mapping matching effect adjustment constant.

[0114] This expression is used to quantify the degree of similarity between historical facial edge feature blocks and real-time facial edge feature blocks, supporting template matching in facial expression recognition. It aims to assess the similarity of edge overlap to achieve accurate analysis of student status in educational assessment.

[0115] The method for determining whether an important edge matches the edge mapping in the real-time facial edge feature block includes: pushing the important edge to coincide with the edge in the real-time facial edge feature block, calculating the intersection area between them, and calculating the ratio of the intersection area to the edge length of the important edge. If the ratio is less than or equal to a preset value, it is determined that the important edge matches the edge mapping in the real-time facial edge feature block.

[0116] In some embodiments disclosed in this invention, the method for analyzing students' facial expressions and body movements in real-time classroom video information to determine the initial state assessment of a single student includes:

[0117] Facial expressions include focused learning expressions, confused expressions, and unfocused learning expressions; body movements include static upright state, static unupright state, and active unupright state. Each facial expression and each body movement is assigned a specific unit of evaluation parameter.

[0118] We conducted continuous time-point analysis on each student's facial expressions and body movements, and based on the analysis results, we determined the initial state assessment for each student.

[0119] In some embodiments disclosed in this invention, the method for making the first correction to the initial state evaluation, taking into account the difficulty of understanding different map nodes, includes:

[0120] To guide the classroom explanation process, a time progress reference line is established. The combination of facial expression factors and body movement factors determined at different time points is recorded as a state factor group, and the state factor group is mapped onto the time progress reference line according to the corresponding time point.

[0121] By constructing a timeline as a reference benchmark, facial expression factors (such as quantitative values ​​of focus or confusion) and body movement factors (such as quantitative values ​​of uprightness or activity) are integrated at each node to form a composite state factor group. These groups are then mapped onto the reference line in chronological order. This mechanism enables the temporal visualization of student behavior data, ensuring that the assessment captures dynamic changes, avoids the limitations of static analysis, and supports subsequent segmentation and correction.

[0122] Based on the time nodes corresponding to the mind map nodes explained in class, time segments are marked on the time progress reference line, and the difficulty of understanding the mind map nodes is associated with the corresponding time segments.

[0123] By using the explanation timestamps of the mind map nodes to divide the time intervals on the reference line (such as the time interval from node A to B), and binding the preset understanding difficulty of each node (such as quantified score) to the corresponding interval, this association mechanism aligns the teaching content structure with time behavior data, ensuring that difficulty factors are integrated into the time sequence analysis, avoiding blind assessment that is detached from the knowledge difficulty, and achieving contextual relevance and accuracy of the assessment.

[0124] For each state factor, an initial sub-state evaluation is performed, and based on the time segment to which the corresponding time node of the state factor belongs, the correction features for the initial sub-state evaluation are determined.

[0125] First, a preliminary evaluation value (such as the weighted sum of facial expressions and actions) is calculated for each state factor group. Then, based on the segment in which the time node falls, the associated difficulty is extracted as a correction factor (such as amplifying the negative impact when the difficulty is high). This step-by-step mechanism achieves dynamic optimization of sub-states through difficulty-driven adjustments, avoids subjective bias, ensures that the final evaluation reflects the real student response under knowledge challenges, and provides targeted guidance for teaching improvement.

[0126] In some embodiments disclosed in this invention, the expression for calculating the initial state assessment of a single student is as follows:

[0127] Where P is the initial state assessment of a single student, and α t β is the unit evaluation parameter for facial expression at the i-th time point. t Let δ[t] be the unit evaluation parameter for the limb movement at the i-th time node, and let δ[t] be the judgment function for continuous time nodes. If there exists α for each continuous time node... t +β t If the difficulty of the t-th time node is greater than or equal to the preset value, the number of consecutive time nodes is determined, and the largest number of consecutive time nodes is selected. δ[t] is the preset continuity correction parameter based on the largest number of consecutive time nodes. H[...] is the difficulty judgment function. If the difficulty of the t-th time node is greater than or equal to the preset value, and α... t +β t If the value is negative, H[...] will output 0. If the difficulty at the t-th time node is less than or equal to the preset value, H[...] will output α. t +β t If the difficulty at time point t is greater than or equal to the preset value, then H[...] will output [α]. t +β t ]×h, where h is the preset difficulty multiplier amplification coefficient.

[0128] This expression is used to quantify the learning status assessment of a single student over a classroom time series. It aims to integrate factors such as facial expressions, body language, continuity, and difficulty to achieve dynamic adjustments in educational assessment. The expression design is based on an additive model combined with multiplicative corrections to ensure that P reflects the overall strength of the student's response, while amplifying the impact of continuous positive behavior and challenging situations, providing a non-linearly sensitive assessment.

[0129] In some embodiments disclosed in this invention, an educational assessment system based on visual recognition data analysis is also disclosed, including:

[0130] The first module is used to construct a mind map for classroom explanation plans, preset the understanding difficulty of each mind map node, and define the explanation order of the mind map nodes. Based on the constructed explanation order, the mind map nodes are connected sequentially to obtain the explanation order line.

[0131] The second module is used to perform semantic analysis on the classroom explanation plan, identify several explanation keywords, and associate the explanation keywords with the corresponding mind map nodes.

[0132] The third module is used to acquire real-time classroom audio information, extract real-time explanation text information from the real-time classroom audio information, perform semantic analysis on the extracted real-time explanation text information, determine real-time explanation keywords, and determine the corresponding mind map nodes based on several real-time explanation keywords that appear in a preset time period. The determined mind map nodes are then connected in sequence to obtain the real-time mind map node sequence line.

[0133] The fourth module is used to acquire real-time classroom video information, analyze students' facial expressions and body movements in the real-time classroom information, determine the initial state assessment of a single student, and make the first correction to the initial state assessment based on the understanding difficulty of different mind map nodes, so as to obtain the final state assessment of the student. The sum of the final state assessments of all students is recognized as the teaching assessment effect.

[0134] The fifth module is used to determine the relative positioning of the teaching evaluation effect by using the preset teaching evaluation effect corresponding to the explanation sequence line corresponding to the real-time map node sequence line as the standard.

[0135] The method for analyzing students' facial expressions includes using a facial mapping model to determine a real-time facial mapping template for facial expressions, and matching the real-time facial mapping template with a standard facial mapping template to determine the expression type.

[0136] Methods for analyzing students' body movements include using body movement analysis models to determine the movement types of body movements.

[0137] This invention discloses an educational assessment method and system based on visual recognition data analysis, belonging to the field of educational analysis technology. It presets the comprehension difficulty for each mind map node as a benchmark for the teaching path; performs semantic analysis on the explanation plan, extracting knowledge points, secondary and supplementary keyword groups, and associating them with corresponding mind map nodes to achieve content structuring; acquires real-time classroom audio information, extracts and analyzes the text to determine real-time keywords, locates corresponding nodes in the mind map, and sequentially connects them to form a real-time mind map node sequence line, capturing the actual explanation path; calculates the initial student state assessment, combines it with node difficulty correction to obtain the final assessment, and accumulates all student assessments as the teaching effect; uses the teaching effect of the preset explanation sequence line corresponding to the real-time sequence line as the standard to determine the relative positioning of the actual assessment, achieving deviation quantification; this invention improves the real-time performance and accuracy of assessment through multimodal fusion and dynamic sequence line mechanisms, providing guidance for teaching improvement.

[0138] Through the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented in hardware or by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) and includes several instructions to cause a computer device (such as a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0139] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for educational assessment based on visual recognition data analysis, characterized in that, The method comprises the following steps: Step S100, constructing a lecture mind map for a classroom lecture plan, presetting understanding difficulty for each mind map node, and determining a lecture sequence of the mind map nodes of the lecture mind map; sequentially connecting the mind map nodes based on the constructed lecture sequence to obtain a lecture sequence line; Step S200, performing semantic analysis on the classroom lecture plan to determine a plurality of lecture keywords, and associating the lecture keywords with corresponding mind map nodes respectively; Step S300, obtaining real-time classroom voice information, performing real-time lecture text information extraction on the real-time classroom voice information, performing semantic analysis on the extracted real-time lecture text information to determine real-time lecture keywords, and based on a plurality of real-time lecture keywords appearing in a preset time period, determining corresponding mind map nodes in the mind map and sequentially connecting the determined mind map nodes to obtain a real-time mind map node sequence line; Step S400, obtaining real-time classroom video information, analyzing the facial expressions and body movements of students in the real-time classroom information to determine the initial state evaluation of a single student, and combining the understanding difficulty of different mind map nodes to perform a first correction on the initial state evaluation to obtain the final state evaluation of the student, and cumulatively determining the final state evaluation of all students as the teaching evaluation effect; Step S500, taking the preset teaching evaluation effect corresponding to the lecture sequence line corresponding to the real-time mind map node sequence line as a standard to determine the relative positioning of the teaching evaluation effect; The method for analyzing the facial expressions of students comprises: determining a real-time facial mapping template of the facial expressions by using a facial mapping model, and matching the real-time facial mapping template with a standard facial mapping template to determine the expression type; The method for analyzing the body movements of students comprises: determining the movement type of the body movements by using a body analysis model.

2. The education evaluation method based on visual recognition data analysis according to claim 1, characterized in that, The method for performing semantic analysis on the classroom lecture plan comprises: previously determining the knowledge point keywords related to the knowledge points in the classroom lecture plan; based on the positions of the knowledge point keywords in the text segments of the classroom lecture plan, determining secondary keywords having a semantic association relationship before and after the knowledge point keywords, and combining the knowledge point keywords and the secondary keywords based on the semantic association relationship to obtain a lecture keyword group; determining the text segment range corresponding to the lecture keyword group, determining other supplementary keywords in the text information with an occurrence frequency greater than or equal to a preset value, supplementing the supplementary keywords into the lecture keyword group, and associating the lecture keyword group with the corresponding mind map nodes.

3. The education evaluation method based on visual recognition data analysis according to claim 2, characterized in that, The method for determining corresponding mind map nodes in the mind map based on a plurality of lecture keywords appearing in a preset time period comprises: determining the mind map nodes that have been lectured, taking the just-lectured mind map nodes as the current positioning nodes, and screening a plurality of mind map nodes as comparison nodes after the current positioning nodes. The real-time explanation keywords appearing in the preset time period are compared with the explanation keyword groups corresponding to the guide node for comparison, the mapping quantity of the real-time explanation keywords and different explanation keyword groups is determined, and the ratio of the mapping quantity to the keyword quantity of the real-time explanation keywords in the preset time period is calculated, which is recorded as a comprehensive mapping ratio; Based on the comprehensive mapping ratio and the ratio of each type of mapped keywords in the explanation keyword group, the mapping coincidence parameter of the real-time explanation keywords appearing in the preset time period and the explanation keyword group is determined, and based on the size of the mapping coincidence parameter, the explanation keyword group corresponding to the real-time keyword is determined, and the guide node corresponding to the explanation keyword group is identified as the corresponding guide node.

4. The education evaluation method based on visual recognition data analysis according to claim 3, characterized in that, The method for determining the mapping coincidence parameter of the real-time explanation keywords appearing in the preset time period and the explanation keyword group includes: The first mapping quantity of the knowledge point keywords mapped in the explanation keyword group, the second mapping quantity of the secondary keywords mapped, and the third mapping quantity of the supplementary keywords mapped are determined; The first mapping ratio of the first mapping quantity to the real-time keyword quantity in the preset time period is calculated, the second mapping ratio of the second mapping quantity to the real-time keyword quantity in the preset time period is calculated, and the third mapping ratio of the third mapping quantity to the real-time keyword quantity in the preset time period is calculated; Based on the first mapping ratio, the second mapping ratio, the third mapping ratio, and the comprehensive mapping ratio, the mapping coincidence parameter is determined; The expression for calculating the mapping coincidence parameter is: Y = R x U zong x exp{[K1 x U1 + K2 x U2 + K3 x U3] + b}; Wherein, Y is a mapping matching parameter, R is a mapping matching parameter conversion coefficient, U zong is a comprehensive mapping ratio, U1 is a first mapping ratio, U2 is a second mapping ratio, U3 is a third mapping ratio, K1 is a knowledge point keyword weight adjustment coefficient, K2 is a secondary keyword weight adjustment coefficient, K3 is a supplementary keyword weight adjustment coefficient, and b is a mapping ratio adjustment constant. 5.The education evaluation method based on visual recognition data analysis of claim 1, wherein, The method for determining the real-time face mapping template of the facial expression by using the face mapping model includes: The student face area is identified by using the preset face recognition model, and based on the determined face area, the real-time face feature block in the face area is determined, including the eye feature block, the nose feature block, the eyebrow feature block, and the mouth feature block; A plane face block performance output model is constructed, the face feature block is analyzed by using the plane face block performance output model, the corresponding plane face block performance is determined, and the plane face block performance is configured in the preset face mapping template to obtain the real-time face mapping template; The method for constructing the plane face block performance output model includes: The edges in the historical face feature block are detected by using the edge detection technology to form the historical face edge feature block, and the important edges in the historical face edge feature block are marked, and the unmarked edges are removed; A plurality of plane face block performances are constructed, and the association between the plane face block performance and the historical face edge feature block is established; The edges in the real-time face feature block are detected by using the edge detection technology to form the real-time face edge feature block, the real-time face edge feature block and different historical face edge feature blocks are compared, the edge coincidence degree of each historical face edge feature block participating in the comparison is determined, and based on the edge coincidence degree, the historical face edge feature block that matches is selected, and the corresponding plane face block performance is called for output; The expression for calculating the edge coincidence degree is: Wherein, W is the edge matching degree, μ x is the matching parameter of the xth important edge, if the xth important edge is matched by the edge mapping in the real-time face edge feature block, then μ x Output 1, otherwise output 0, L is the edge mapping matching influence adjustment coefficient, and D is the edge mapping matching influence adjustment constant; The method for judging whether the important edge is matched by the edge in the real-time face edge feature block comprises: overlapping the important edge and the edge in the real-time face edge feature block, calculating the intersection area between the two, and calculating the ratio of the intersection area to the edge length of the important edge; if the ratio is less than or equal to a preset value, it is determined that the important edge is matched by the edge in the real-time face edge feature block. 6.The education evaluation method based on visual recognition data analysis of claim 1, wherein, The method for determining the initial state evaluation of a single student by analyzing the facial expressions and body movements of students in real-time classroom video information comprises: The facial expressions include focused learning expressions, doubt expressions and unfocused learning expressions, and the body movements include stationary and correct states, stationary and incorrect states and active and incorrect states, wherein each facial expression and each body movement is configured with a specific unit evaluation parameter; The initial state evaluation of a single student is determined by continuously analyzing the time nodes of the facial expressions and body movements of each student and based on the analysis results.

7. The education evaluation method based on visual recognition data analysis according to claim 1, characterized in that, The method for making a first correction to the initial state evaluation in combination with the understanding difficulty of different guide map nodes comprises: For the explanation process of the classroom, a time process reference line is established, the combination of the facial expression factors and the body movement factors determined at different time nodes is recorded as a state factor group, and the state factor group is mapped on the time process reference line according to the corresponding time nodes; Based on the time nodes corresponding to the guide map nodes explained in the classroom explanation process, time sections are marked on the time process reference line, and the understanding difficulty of the guide map nodes is associated with the corresponding time sections; An initial sub-state evaluation is made for each state factor, and a correction feature of the initial sub-state evaluation is determined based on the time section to which the time node corresponding to the state factor belongs.

8. The education evaluation method based on visual recognition data analysis according to claim 7, characterized in that, The expression of the initial state evaluation of a single student is: Wherein, P is the initial state evaluation of a single student, a t is the unit evaluation parameter of facial expression of the i th time node, β t is the unit evaluation parameter of body movement of the i th time node, δ[t] is a continuous time node judgment function, if there is a continuous time node respectively corresponding to a t +β t is greater than or equal to a preset value, the corresponding continuous time node number is determined, and the maximum continuous time node number is selected, δ[t] determines the preset continuity correction parameter according to the maximum continuous time node number, H[...] is a difficulty judgment function, if the difficulty of the t th time node is greater than or equal to a preset value, and a t +β t is negative, 0 is output by H[...], if the difficulty of the t th time node is less than or equal to a preset value, a t +β t is output by H[...], if the difficulty of the t th time node is greater than or equal to a preset value, [a t +β t ]×h is output by H[...], wherein h is a preset difficulty amplification coefficient.

9. An educational assessment system based on visual recognition data analysis, characterized in that, The educational evaluation method for executing any one of claims 1-8 comprises: A first module is configured to construct an explanation mind map for a classroom explanation scheme, preset an understanding difficulty for each guide map node, and demarcate the explanation sequence of the guide map nodes of the explanation mind map, sequentially connect the guide map nodes based on the constructed explanation sequence, and obtain an explanation sequence line; A second module is configured to perform semantic analysis on the classroom explanation scheme, determine a plurality of explanation keywords, and associate the explanation keywords with the corresponding guide map nodes respectively; A third module is configured to obtain real-time classroom audio information, perform real-time explanation text information extraction on the real-time classroom audio information, perform semantic analysis on the extracted real-time explanation text information, determine real-time explanation keywords, and based on a plurality of real-time explanation keywords appearing in a preset time period, determine the corresponding guide map nodes in the mind map, and sequentially connect the determined guide map nodes to obtain a real-time guide map node sequence line. The fourth module is configured to acquire real-time classroom video information, analyze facial expressions and body movements of students in the real-time classroom information, determine initial state evaluation of a single student, make a first correction to the initial state evaluation in combination with understanding difficulties of different guide map nodes, obtain final state evaluation of the student, and determine a cumulative sum of final state evaluations of all students as a teaching evaluation effect; The fifth module is configured to determine a relative positioning of the teaching evaluation effect by taking a preset teaching evaluation effect corresponding to an explanation sequence line corresponding to a real-time guide map node sequence line as a standard; The facial expression analysis method for the student includes determining a real-time facial mapping template of the facial expression by using a facial mapping model, matching the real-time facial mapping template with a standard facial mapping template, and determining an expression type; The body movement analysis method for the student includes determining a movement type of the body movement by using a body analysis model.

Citation Information

Patent Citations

  • Face beauty assessment method based on video

    CN101305913A

  • Driving danger identification method and system

    CN106127155A

  • Classroom teaching quality assessment method

    CN107895244A

  • Teaching behavior analysis system for teaching feature fusion and modeling based on knowledge base

    CN115239527A

  • Classroom effect evaluation method and system based on multi-modal attention data analysis

    CN119312809A