An intelligent education teaching evaluation method, system and device
By establishing the mapping relationship between teaching data and test data, and establishing a teaching evaluation model in combination with monitoring data slices, the problem of overfitting the image recognition teaching evaluation model in the existing technology is solved, and a more accurate and objective teaching evaluation is achieved.
Patent Information
- Application Number
- CN202510217057.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-02-26
AI Technical Summary
The existing schemes that use image recognition for teaching evaluation rely on posture recognition and expression recognition, which is prone to overfitting, resulting in a biased evaluation of AI model output.
By obtaining teaching data and test data, analyzing and establishing mapping relationships, combining monitoring data slices to establish and train teaching evaluation models, inputting the monitoring data to be evaluated to obtain objective and accurate evaluation results.
It improves the accuracy and interpretability of teaching evaluation, reduces the risk of overfitting the model, and ensures the objectivity and fairness of the evaluation.
Smart Images

Figure CN119692823B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to an intelligent education teaching evaluation method, system and device. Background Art
[0002] With the rapid development of information technology, the application of artificial intelligence (AI for short) has gradually deepened in various fields. Especially in the field of education, the wide application of artificial intelligence technology has provided new opportunities for the innovation of education models and the improvement of education quality. In recent years, teaching evaluation, as an important link in the education process, has received more and more attention. Traditional teaching evaluation mainly relies on manual scoring and subjective judgment. The evaluation process is inefficient and it is difficult to comprehensively and accurately reflect the learning status of students and the teaching quality of teachers. Especially in a large-scale education environment, the limitations of manual evaluation are more prominent.
[0003] Traditional teaching evaluation methods usually include indicators such as written tests, assignments, and classroom participation. The evaluation subjects are mostly teachers, and the evaluation criteria and methods often have certain subjectivity and inconsistency. With the continuous expansion of the education scale, the diversification of evaluation objects and the complexity of evaluation content, the traditional evaluation methods are facing huge challenges. Currently, in the field of classroom teaching quality evaluation, the application of artificial intelligence has been gradually penetrated, including using computer vision technology to conduct real-time analysis on the teacher's teaching process through video monitoring, and automatically identifying indicators such as the teacher's teaching behavior, the transmission of teaching content, and the student's classroom participation to evaluate the teaching quality.
[0004] Currently, the solutions for teaching evaluation using image recognition generally rely on pose recognition and facial expression recognition to evaluate the state of teachers and / or students in class, and evaluate education and teaching based on the state of teachers and / or students in class. However, the models used in the above solutions are prone to overfitting, and the teaching style of teachers is also likely to affect the final results. For example, teachers with a serious classroom atmosphere generally receive higher evaluations than those with an active classroom atmosphere, but in fact, it may be that teachers with an active classroom atmosphere have better teaching effects. The evaluations based on the output of AI models are prone to be biased. Summary of the Invention
[0005] The present invention provides an intelligent education teaching evaluation method, system and device, and provides a teaching effect evaluation solution based on artificial intelligence, which at least solves the problem that the solutions for teaching evaluation using image recognition generally rely on pose recognition and facial expression recognition to evaluate the state of teachers and / or students in class, and evaluate education and teaching based on the state of teachers and / or students in class. However, the models used in the above solutions are prone to overfitting, and the evaluations based on the output of AI models are prone to be biased.
[0006] This application provides an intelligent education teaching evaluation method, including:
[0007] Obtain first teaching data and first test data, parse the first teaching data and the first test data, and obtain a first mapping of the first test data and the first teaching data, where the first mapping is configured as a mapping of each question in the first test data to each teaching period in the first teaching data;
[0008] Obtain first monitoring data corresponding to the first teaching data according to the first teaching data;
[0009] Obtain a second mapping of the first test data and the first monitoring data according to the first teaching data, the first monitoring data, and the first mapping, where the second mapping is configured as a mapping of each question in the first test data to each monitoring data slice in the first monitoring data;
[0010] Conduct a test on the target to be tested according to the first test data, and obtain a first test result;
[0011] Obtain the teaching evaluation of each monitoring data slice in the first monitoring data according to the score of each question in the first test result;
[0012] Establish a teaching evaluation model with the teaching evaluations of each monitoring data slice in the first monitoring data as a data set, where the teaching evaluation model is configured to take the monitoring data slice as input and the teaching evaluation corresponding to the monitoring data slice as output;
[0013] Input the monitoring data to be evaluated according to the teaching evaluation model, and obtain the evaluation result corresponding to the monitoring data to be evaluated.
[0014] Optionally, the obtaining of the first teaching data and the first test data, parsing the first teaching data and the first test data, and obtaining the first mapping of the first test data and the first teaching data includes:
[0015] Obtain the first test data, parse the first test data to obtain a first question, and the first question belongs to the first test data;
[0016] Obtain a first parsing corresponding to the first question according to the first question;
[0017] Obtain at least one first embedding vector of the first parsing according to the first parsing;
[0018] Obtain the first teaching data, where the first teaching data includes at least one of text data, voice data, and image data, parse the first teaching data, and obtain a first teaching period, and the first teaching period belongs to the first teaching data;
[0019] Extract the features of the data corresponding to the first teaching period and perform feature fusion to obtain a second embedding vector according to the first teaching period;
[0020] Obtain the mapping relationship between the first question and the first teaching period according to the at least one first embedding vector and the second embedding vector;
[0021] Obtain the first mapping between the first test data and the first teaching data according to the mapping relationship between the first question and the first teaching period.
[0022] Optionally, the first teaching data includes at least voice data;
[0023] The parsing of the first teaching data to obtain the first teaching period includes;
[0024] Obtain the ASR data corresponding to the voice data according to the voice data;
[0025] Perform semantic analysis on the ASR data according to the ASR data, slice the ASR data according to the semantic analysis result, and obtain ASR slice data;
[0026] Obtain the slice data of the first teaching data according to the ASR slice data;
[0027] Obtain the first teaching period according to the slice data of the first teaching data.
[0028] Optionally, the first teaching data includes text data, voice data, and image data;
[0029] The extracting the features of the data corresponding to the first teaching period and performing feature fusion to obtain a second embedding vector according to the first teaching period includes:
[0030] Extract the context features of the text data using a pre-trained BERT model according to the text data to obtain text features;
[0031] Obtain the ASR data corresponding to the voice data according to the voice data, and extract the context features of the ASR data using a pre-trained BERT model to obtain voice features;
[0032] Extract image features using an image recognition model according to the image data to obtain first image features;
[0033] Extract the text in the image using OCR according to the image data, and extract the context features of the text in the image using a pre-trained BERT model to obtain second image features;
[0034] According to the text feature, the speech feature, the first image feature, and the second image feature, use a multimodal fusion model to map the text feature, the speech feature, the first image feature, and the second image feature to a unified embedding space, and obtain a second embedding vector.
[0035] Optionally, obtaining the teaching evaluation of each monitoring data slice in the first monitoring data according to the scoring situation of each question in the first test result includes:
[0036] Obtain the questions in the first test data corresponding to each monitoring data slice according to the second mapping;
[0037] According to the scoring situation of each question in the first test data corresponding to each monitoring data slice and the proportion of the content involved in each monitoring data slice in each question, obtain the scoring result of each monitoring data slice in the first monitoring data;
[0038] Perform normalization processing on the scoring result to obtain the teaching evaluation of each monitoring data slice in the first monitoring data.
[0039] Optionally, establishing a teaching evaluation model based on the teaching evaluation of each monitoring data slice in the first monitoring data as a data set includes:
[0040] Establish an initial fusion model, which is configured to include a pose recognition sub-model, an expression analysis sub-model, and a temporal fusion sub-model. The initial fusion model is configured to take video data as input and evaluation data as output;
[0041] Obtain a training set and a test set according to the data set;
[0042] Train the initial fusion model according to the training set and the test set to obtain a teaching evaluation model.
[0043] Optionally, according to the teaching evaluation model, input the monitoring data to be evaluated, and obtain the evaluation result corresponding to the monitoring data to be evaluated, including:
[0044] Obtain the slice data to be evaluated according to the monitoring data to be evaluated, and the slice data to be evaluated belongs to the monitoring data to be evaluated;
[0045] According to the teaching evaluation model, input the slice data to be evaluated, and obtain the evaluation of each slice data to be evaluated;
[0046] Obtain the evaluation result corresponding to the monitored data to be evaluated according to the evaluation of each slice data to be evaluated.
[0047] On the other hand, an intelligent education teaching evaluation system includes a model management platform and an evaluation platform;
[0048] The model management platform is configured to:
[0049] Obtain first teaching data and first test data, parse the first teaching data and the first test data, and obtain a first mapping of the first test data and the first teaching data, where the first mapping is configured as the mapping of each question in the first test data to each teaching period in the first teaching data;
[0050] According to the first teaching data, obtain first monitoring data corresponding to the first teaching data;
[0051] According to the first teaching data, the first monitoring data, and the first mapping, obtain a second mapping of the first test data and the first monitoring data, where the second mapping is configured as the mapping of each question in the first test data to each monitoring data slice in the first monitoring data;
[0052] According to the first test data, conduct a test on the target to be tested and obtain a first test result;
[0053] According to the score situation of each question in the first test result, obtain the teaching evaluation of each monitoring data slice in the first monitoring data;
[0054] Using the teaching evaluation of each monitoring data slice in the first monitoring data as a data set, establish a teaching evaluation model, where the teaching evaluation model is configured to take the monitoring data slice as the input and the teaching evaluation corresponding to the monitoring data slice as the output;
[0055] The model evaluation platform is configured to:
[0056] According to the teaching evaluation model, input the monitoring data to be evaluated and obtain an evaluation result corresponding to the monitoring data to be evaluated.
[0057] Optionally, it further includes a data acquisition platform;
[0058] The data acquisition platform is configured to collect and store at least one of the first teaching data, the first test data, the first monitoring data, and the monitoring data to be evaluated.
[0059] On the other hand, an embodiment of the present application further provides a device, which includes a memory and a processor. A computer program is stored in the memory, and the processor executes the computer program to implement the above method.
[0060] In another aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and a processor executes the computer program to implement the above method.
[0061] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0062] An intelligent education and teaching evaluation method, system and device of the present invention include obtaining first teaching data and first test data, parsing the first teaching data and the first test data to obtain a first mapping of the first test data and the first teaching data, and the first mapping is configured as a mapping of each question in the first test data to each teaching period in the first teaching data; obtaining first monitoring data corresponding to the first teaching data according to the first teaching data; obtaining a second mapping of the first test data and the first monitoring data according to the first teaching data, the first monitoring data and the first mapping, and the second mapping is configured as a mapping of each question in the first test data to each monitoring data slice in the first monitoring data; performing a test on a target to be tested according to the first test data to obtain a first test result; obtaining a teaching evaluation of each monitoring data slice in the first monitoring data according to the score of each question in the first test result; establishing a teaching evaluation model with the teaching evaluations of each monitoring data slice in the first monitoring data as a data set, and the teaching evaluation model is configured to take a monitoring data slice as an input and the teaching evaluation corresponding to the monitoring data slice as an output; inputting the monitoring data to be evaluated according to the teaching evaluation model to obtain an evaluation result corresponding to the monitoring data to be evaluated. At least it solves the problem that the current solutions for teaching evaluation using image recognition generally rely on pose recognition and expression recognition to evaluate the states of teachers and / or students in class, and evaluate education and teaching based on the states of teachers and / or students in class, but the models used in the above solutions are prone to overfitting, and the evaluations based on the output of the AI model are prone to be one-sided. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale.
[0064] Figure 1 It is a flowchart of an intelligent education and teaching evaluation method in the present application;
[0065] Figure 2 It is a structural diagram of a device in the present application;
[0066] Markings in the figure: 101 - Processor, 102 - Communication bus, 103 - Network interface, 104 - User interface, 105 - Memory.
[0067] The realization, functional features and advantages of the present application will be further described in conjunction with the embodiments with reference to the accompanying drawings. Detailed implementation manners
[0068] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0069] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above accompanying drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0070] Embodiment 1
[0071] As Figure 1 shown, an intelligent education and teaching evaluation method includes:
[0072] S1. Obtain first teaching data and first test data, analyze the first teaching data and the first test data, and obtain a first mapping of the first test data and the first teaching data.
[0073] Specifically, the first mapping is configured as the mapping of each question in the first test data to each teaching period in the first teaching data.
[0074] Optionally, the first teaching data is historical classroom data, including one or more of the lesson plans, audio data, and video data corresponding to each class, where the audio data at least includes the audio data corresponding to the teacher, and the video data at least includes the video data corresponding to the blackboard writing.
[0075] Optionally, the first test data may include exam questions, the analysis of exam questions, or test questions specifically used to evaluate teaching quality, and the analysis of test questions.
[0076] The main purpose of this step is to associate the test results with the teaching data, so as to quantitatively evaluate the teaching quality with the test results and make the final education and teaching evaluation results more in line with the objective results.
[0077] S2. Obtain first monitoring data corresponding to the first teaching data according to the first teaching data.
[0078] Optionally, the first monitoring data may be video data obtained by a camera, or may include physiological data such as the user's heart rate obtained by wearable devices such as children's watches.
[0079] S3. Obtain a second mapping between the first test data and the first monitoring data according to the first teaching data, the first monitoring data, and the first mapping.
[0080] Specifically, the second mapping is configured as the mapping between each question in the first test data and each monitoring data slice in the first monitoring data.
[0081] Optionally, the monitoring data slice is data obtained by slicing the first monitoring data according to time, and each monitoring data slice corresponds to a time period in the first monitoring data.
[0082] By adopting the above steps, the relationship between the knowledge points of each question and the monitoring data slices can be established. The purpose is to make the corresponding relationship between the characteristics of the monitoring data and the evaluation results clearer and improve the interpretability of the model.
[0083] S4. Conduct a test on the target to be tested according to the first test data to obtain a first test result.
[0084] S5. Obtain the teaching evaluation of each monitoring data slice in the first monitoring data according to the score of each question in the first test result.
[0085] Optionally, according to the score of the question corresponding to each monitoring data slice, obtain the teaching evaluation of the corresponding monitoring data slice. The higher the score rate of the corresponding question, the higher the teaching evaluation of the monitoring data slice.
[0086] S6. Establish a teaching evaluation model with the teaching evaluation of each monitoring data slice in the first monitoring data as the data set.
[0087] Specifically, the teaching evaluation model is configured to take the monitoring data slice as the input and the teaching evaluation corresponding to the monitoring data slice as the output.
[0088] Optionally, the base of the teaching evaluation model can be selected to include the base of a posture recognition sub-model, an expression analysis sub-model, and a temporal fusion sub-model.
[0089] S7. According to the teaching evaluation model, input the monitoring data to be evaluated, and obtain the evaluation result corresponding to the monitoring data to be evaluated.
[0090] By adopting the above solution, establishing the association between the exam or test and the monitoring data slices can more accurately locate the behaviors or other data that affect education and teaching. The evaluation result is more objective and has stronger interpretability compared to the existing solutions that establish evaluation models relying on indicator systems or expert evaluation results, and can accurately evaluate the effect of education and teaching. At least it solves the problem that the current solutions for teaching evaluation using image recognition generally rely on posture recognition and expression recognition to evaluate the states of teachers and / or students in class, and evaluate education and teaching based on the states of teachers and / or students in class, but the models used in the above solutions are prone to overfitting, and the evaluations based on the outputs of AI models are prone to be biased.
[0091] Embodiment 2
[0092] Based on Embodiment 1, an intelligent education and teaching evaluation method includes:
[0093] S1. Obtain the first teaching data and the first test data, parse the first teaching data and the first test data, and obtain the first mapping of the first test data and the first teaching data.
[0094] Specifically, the first mapping is configured as the mapping between each question in the first test data and each teaching period in the first teaching data.
[0095] Optionally, obtaining the first teaching data and the first test data, parsing the first teaching data and the first test data, and obtaining the first mapping of the first test data and the first teaching data includes:
[0096] Obtain the first test data, parse the first test data to obtain the first question, and the first question belongs to the first test data;
[0097] According to the first question, obtain the first parsing corresponding to the first question;
[0098] According to the first parsing, obtain at least one first embedding vector of the first parsing;
[0099] Obtain the first teaching data, where the first teaching data includes at least one of text data, voice data, and image data, parse the first teaching data, and obtain the first teaching period, and the first teaching period belongs to the first teaching data;
[0100] Extract the features of the data corresponding to the first teaching period and perform feature fusion to obtain a second embedding vector;
[0101] Obtain the mapping relationship between the first question and the first teaching period according to at least one first embedding vector and the second embedding vector;
[0102] Obtain the first mapping of the first test data and the first teaching data according to the mapping relationship between the first question and the first teaching period.
[0103] Optionally, according to the first analysis, based on methods such as named entity recognition and keyword extraction, obtain at least one first embedding vector of the first analysis, that is, each knowledge point corresponds to a first embedding vector. Similarly, according to the first teaching period, based on methods such as named entity recognition and keyword extraction, extract the features of the data corresponding to the first teaching period and perform feature fusion to obtain at least one second embedding vector;
[0104] The method for obtaining the mapping relationship between the first question and the first teaching period according to at least one first embedding vector and the second embedding vector includes:
[0105] Determine whether there is an embedding vector group with a similarity greater than the threshold in at least one first embedding vector and at least one second embedding vector. If so, the first question is associated with the first teaching period.
[0106] Optionally, the first teaching data at least includes voice data;
[0107] Parse the first teaching data to obtain the first teaching period, including:
[0108] Obtain the ASR data corresponding to the voice data according to the voice data;
[0109] Perform semantic analysis on the ASR data according to the ASR data, slice the ASR data according to the semantic analysis result, and obtain the ASR slice data;
[0110] Obtain the slice data of the first teaching data according to the ASR slice data;
[0111] Obtain the first teaching period according to the slice data of the first teaching data.
[0112] Optionally, the above scheme can be implemented by the following pseudocode:
[0113] S101. Install the necessary libraries, such as installing SpeechRecognition for speech recognition, installing pyaudio for audio input, and installing transformers for the BERT model;
[0114] S102. Import necessary libraries, such as importing the speech_recognition library for speech recognition, importing BertTokenizer and BertModel from the transformers library for text feature extraction, and importing torch for tensor operations;
[0115] S103. Initialize the speech recognizer, such as creating a Recognizer object to perform the speech recognition task;
[0116] S104. Load the audio file and convert it to text, including defining the path of the audio file, opening the audio file using AudioFile, reading the audio data through the speech recognizer, and using Google's ASR engine to convert the audio to text;
[0117] S105. Use BERT to extract text features, including loading the pre-trained model and vocabulary of BERT (BertTokenizer and BertModel), tokenizing and encoding the text obtained from the audio conversion, using the BERT model to extract the features of the text, and using the output of the first token of the last hidden layer of the model as the embedding vector of the text;
[0118] S106. Output the text embedding features, including the extracted text embedding features and printing the shape of the embedding vector.
[0119] Optionally, the first teaching data includes text data, speech data, and image data;
[0120] According to the first teaching period, extract the features of the data corresponding to the first teaching period and perform feature fusion to obtain the second embedding vector, including:
[0121] According to the text data, use the pre-trained BERT model to extract the context features of the text data to obtain text features;
[0122] According to the speech data, obtain the ASR data corresponding to the speech data, and use the pre-trained BERT model to extract the context features of the ASR data to obtain speech features;
[0123] According to the image data, use the image recognition model to extract the image features to obtain the first image features;
[0124] According to the image data, use OCR to extract the text in the image, and use the pre-trained BERT model to extract the context features of the text in the image to obtain the second image features;
[0125] According to the text features, speech features, first image features, and second image features, use a multimodal fusion model to map the text features, speech features, first image features, and second image features to a unified embedding space to obtain a second embedding vector.
[0126] Optionally, the multimodal fusion model adopts a CLIP model. After mapping the features from text, audio, and images to a unified embedding space, splicing fusion or weighted fusion can be used for fusion. In this embodiment, weighted fusion is preferably used for fusion, where the weight of the language features is greater than or equal to the weight of the text features, which is greater than or equal to the weight of the first image features, which is greater than or equal to the weight of the second image features.
[0127] S2. Obtain first monitoring data corresponding to the first teaching data according to the first teaching data.
[0128] S3. Obtain a second mapping of the first test data and the first monitoring data according to the first teaching data, the first monitoring data, and the first mapping.
[0129] Specifically, the second mapping is configured as the mapping of each question in the first test data to each monitoring data slice in the first monitoring data.
[0130] S4. Conduct a test on the target to be tested according to the first test data to obtain a first test result.
[0131] S5. Obtain the teaching evaluation of each monitoring data slice in the first monitoring data according to the score of each question in the first test result.
[0132] Optionally, obtaining the teaching evaluation of each monitoring data slice in the first monitoring data according to the score of each question in the first test result includes:
[0133] Obtain the questions in the first test data corresponding to each monitoring data slice according to the second mapping;
[0134] Obtain the scoring result of each monitoring data slice in the first monitoring data according to the score of each question in the first test data corresponding to each monitoring data slice and the proportion of the content involved in each monitoring data slice in each question;
[0135] Perform a normalization process on the scoring result to obtain the teaching evaluation of each monitoring data slice in the first monitoring data.
[0136] Optionally, obtaining the scoring result of each monitoring data slice in the first monitoring data according to the score of each question in the first test data corresponding to each monitoring data slice and the proportion of the content involved in each monitoring data slice in each question includes:
[0137] Obtain at least one first slice topic corresponding to the first monitoring data slice. The first monitoring data slice belongs to the monitoring data slices. According to the first slice topic, obtain the knowledge points corresponding to each first slice topic;
[0138] According to the total number of knowledge points corresponding to the first slice topic and the number of knowledge points corresponding to the first monitoring data slice, obtain the importance of the first monitoring data slice relative to the first slice topic;
[0139] According to the importance of the first monitoring data slice relative to all first slice topics, obtain the weights of each first slice topic; the scoring rate of the topic, and obtain the scoring result of the first monitoring data slice;
[0140] According to the scoring result of the first monitoring data slice, obtain the scoring results of each monitoring data slice in the first monitoring data.
[0141] Optionally, the scoring result of the first monitoring data slice is calculated by the following formula:
[0142]
[0143] Among them, P 1 represents the scoring result of the first monitoring data slice, N i represents the i weight of the D i represents the i scoring rate of the n th first slice topic,
[0144] is the total number of first slice topics. i Optionally, the weight of the i th first slice topic is equal to the number of knowledge points associated with the i th first slice topic and the first monitoring data slice divided by the total number of knowledge points of the
[0145] th first slice topic.
[0146] Using the above solution, a solution for improving the teaching evaluation accuracy of each monitoring data slice is provided. The core of this education and teaching evaluation method lies in more accurate and objective evaluation of each monitoring data slice. Using the above method, the accuracy of this solution can be greatly improved, and the problems of poor model interpretability and robustness caused by directly training the model through the exam results and a whole piece of monitoring data can be solved.
[0147] Specifically, the teaching evaluation model is configured to take the monitored data slices as input and the corresponding teaching evaluations of the monitored data slices as output.
[0148] Optionally, a teaching evaluation model is established based on the teaching evaluations of each monitored data slice in the first monitored data as a data set, including:
[0149] An initial fusion model is established. The initial fusion model is configured to include a pose recognition sub-model, a facial expression analysis sub-model, and a temporal fusion sub-model. The initial fusion model is configured to take video data as input and evaluation data as output;
[0150] According to the data set, a training set and a test set are obtained;
[0151] The initial fusion model is trained according to the training set and the test set to obtain the teaching evaluation model.
[0152] Optionally, the overall architecture of the initial fusion model is configured to include a video frame extraction and preprocessing module, a pose recognition sub-model, a facial expression analysis sub-model, and a temporal fusion sub-model;
[0153] The video frame extraction and preprocessing module is configured to: decompose the video into each frame image and perform necessary preprocessing on each frame, such as size adjustment, normalization, etc.;
[0154] The pose recognition sub-model is configured to: use OpenPose or HRNet for pose estimation, identify the body key points of the person in the video frame, and extract the pose information of the person from each frame.
[0155] The facial expression analysis sub-model is configured to: use VGG-Face or FACS for facial expression analysis, extract the facial expressions of the person from each frame, and infer their emotional states.
[0156] The temporal fusion sub-model is configured to: adopt LSTM or a convolutional neural network for time series modeling to fuse the spatial features (pose, expression) and temporal features (behavior and emotional states changing over time) of the video frames, and then fuse the outputs of each sub-model through weighted average, a fusion network, or a multi-layer perceptron (MLP), combine the results of pose estimation and facial expression analysis with the output of the time series model, and generate the final score.
[0157] Optionally, the initial fusion model is trained according to the training set and the test set to obtain the teaching evaluation model, including:
[0158] S601, data preprocessing and data augmentation;
[0159] S602. Pose estimation sub - model training: Use video data with labeled key - point positions to train OpenPose or HRNet, perform regression training on the key - points in each frame, and optimize the model to minimize the key - point prediction error.
[0160] S603. Facial expression analysis sub - model training: Use a dataset containing facial expression labels to train a facial expression recognition model, extract facial features for each frame, and analyze the emotion category through CNN.
[0161] S604. Temporal fusion sub - model training: Input the features extracted from pose estimation and facial expression analysis into LSTM or CNN, train the network to capture the temporal dynamic changes in the video, perform time - series training on each video segment, predict the emotional state at each moment, fuse the output results of each sub - model, and train a simple fully - connected network to generate the final score.
[0162] S605. Loss function and optimization: Use the MSE loss function or cross - entropy loss function, use the Adam optimizer to optimize the model training, and dynamically adjust the learning rate.
[0163] S7. According to the teaching evaluation model, input the monitoring data to be evaluated, and obtain the evaluation result corresponding to the monitoring data to be evaluated.
[0164] Optionally, according to the teaching evaluation model, input the monitoring data to be evaluated, and obtain the evaluation result corresponding to the monitoring data to be evaluated, including:
[0165] Obtain the sliced data to be evaluated according to the monitoring data to be evaluated, and the sliced data to be evaluated belongs to the monitoring data to be evaluated.
[0166] According to the teaching evaluation model, input the sliced data to be evaluated, and obtain the evaluation of each sliced data to be evaluated.
[0167] Obtain the evaluation result corresponding to the monitored data to be evaluated according to the evaluation of each sliced data to be evaluated.
[0168] Adopting the above - mentioned scheme, on the basis of the scheme of Embodiment 1, each step is optimized, further improving the accuracy and interpretability of the evaluation model. Further solving the problem that the current scheme for teaching evaluation using image recognition generally relies on pose recognition and expression recognition to evaluate the state of teachers and / or students in class, and evaluates education and teaching based on the state of teachers and / or students in class, but the models used in the above - mentioned scheme are prone to overfitting, and the evaluation based on the output of the AI model is prone to being biased.
[0169] Embodiment 3
[0170] An intelligent education and teaching evaluation system, including a model management platform and an evaluation platform;
[0171] The model management platform is configured to:
[0172] Obtain first teaching data and first test data, parse the first teaching data and the first test data, and obtain a first mapping of the first test data and the first teaching data. The first mapping is configured as a mapping of each question in the first test data to each teaching period in the first teaching data;
[0173] According to the first teaching data, obtain first monitoring data corresponding to the first teaching data;
[0174] According to the first teaching data, the first monitoring data, and the first mapping, obtain a second mapping of the first test data and the first monitoring data. The second mapping is configured as a mapping of each question in the first test data to each monitoring data slice in the first monitoring data;
[0175] According to the first test data, conduct a test on the target to be tested and obtain a first test result;
[0176] According to the score of each question in the first test result, obtain the teaching evaluation of each monitoring data slice in the first monitoring data;
[0177] Using the teaching evaluation of each monitoring data slice in the first monitoring data as a data set, establish a teaching evaluation model. The teaching evaluation model is configured to take the monitoring data slice as the input and the teaching evaluation corresponding to the monitoring data slice as the output;
[0178] The model evaluation platform is configured to:
[0179] According to the teaching evaluation model, input the monitoring data to be evaluated and obtain an evaluation result corresponding to the monitoring data to be evaluated.
[0180] Optionally, it further includes a data acquisition platform;
[0181] The data acquisition platform is configured to collect and store at least one of the first teaching data, the first test data, the first monitoring data, and the monitoring data to be evaluated.
[0182] Optionally, obtaining the first teaching data and the first test data, parsing the first teaching data and the first test data, and obtaining the first mapping of the first test data and the first teaching data includes:
[0183] Obtain the first test data, parse the first test data to obtain the first question, and the first question belongs to the first test data;
[0184] According to the first question, obtain a first analysis corresponding to the first question;
[0185] According to the first parsing, at least one first embedding vector of the first parsing is obtained;
[0186] Obtain the first teaching data, the first teaching data includes at least one of text data, voice data, and image data, parse the first teaching data, and obtain the first teaching period, and the first teaching period belongs to the first teaching data;
[0187] According to the first teaching period, extract the features of the data corresponding to the first teaching period and perform feature fusion to obtain a second embedding vector;
[0188] According to at least one first embedding vector and the second embedding vector, obtain the mapping relationship between the first question and the first teaching period;
[0189] According to the mapping relationship between the first question and the first teaching period, obtain the first mapping of the first test data and the first teaching data.
[0190] Optionally, the first teaching data includes at least voice data;
[0191] Parse the first teaching data to obtain the first teaching period, including;
[0192] According to the voice data, obtain the ASR data corresponding to the voice data;
[0193] According to the ASR data, perform semantic analysis on the ASR data, and slice the ASR data according to the semantic analysis result to obtain ASR slice data;
[0194] According to the ASR slice data, obtain the slice data of the first teaching data;
[0195] According to the slice data of the first teaching data, obtain the first teaching period.
[0196] Optionally, the first teaching data includes text data, voice data, and image data;
[0197] According to the first teaching period, extract the features of the data corresponding to the first teaching period and perform feature fusion to obtain a second embedding vector, including:
[0198] According to the text data, use the pre-trained BERT model to extract the context features of the text data to obtain text features;
[0199] According to the voice data, obtain the ASR data corresponding to the voice data, and use the pre-trained BERT model to extract the context features of the ASR data to obtain voice features;
[0200] According to the image data, use the image recognition model to extract image features to obtain the first image features;
[0201] According to the image data, use OCR to extract the text in the image, and use the pre-trained BERT model to extract the context features of the text in the image to obtain the second image features;
[0202] According to the text features, speech features, first image features, and second image features, use a multi-modal fusion model to map the text features, speech features, first image features, and second image features to a unified embedding space to obtain the second embedding vector.
[0203] Optionally, according to the scores of each question in the first test result, obtain the teaching evaluation of each monitoring data slice in the first monitoring data, including:
[0204] According to the second mapping, obtain the questions in the first test data corresponding to each monitoring data slice;
[0205] According to the scores of each question in the first test data corresponding to each monitoring data slice and the proportion of the content involved in each monitoring data slice in each question, obtain the scoring results of each monitoring data slice in the first monitoring data;
[0206] Perform standardization processing on the scoring results to obtain the teaching evaluation of each monitoring data slice in the first monitoring data.
[0207] Optionally, use the teaching evaluation of each monitoring data slice in the first monitoring data as a data set to establish a teaching evaluation model, including:
[0208] Establish an initial fusion model, which is configured to include a pose recognition sub-model, an expression analysis sub-model, and a temporal fusion sub-model. The initial fusion model is configured to take video data as input and evaluation data as output;
[0209] According to the data set, obtain the training set and the test set;
[0210] Train the initial fusion model according to the training set and the test set to obtain the teaching evaluation model.
[0211] Optionally, according to the teaching evaluation model, input the monitoring data to be evaluated to obtain the evaluation results corresponding to the monitoring data to be evaluated, including:
[0212] According to the monitoring data to be evaluated, obtain the slice data to be evaluated, and the slice data to be evaluated belongs to the monitoring data to be evaluated;
[0213] According to the teaching evaluation model, input the slice data to be evaluated to obtain the evaluation of each slice data to be evaluated;
[0214] According to the evaluation of each slice data to be evaluated, obtain the evaluation results corresponding to the monitored data to be evaluated.
[0215] Example 4
[0216] This embodiment provides a device, which includes a memory and a processor. A computer program is stored in the memory, and the processor executes the computer program to implement any of the above methods.
[0217] Specifically, as Figure 2 shown, Figure 2 is a schematic structural diagram of a device for the hardware operating environment involved in the solution of this embodiment of the application. This device is an electronic device and may include: a processor 101, such as a central processing unit (CPU), a communication bus 102, a user interface 104, a network interface 103, and a memory 105. Among them, the communication bus 102 is used to realize the connection and communication between these components. The user interface 104 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 104 may further include a standard wired interface and a wireless interface. The network interface 103 may optionally include a standard wired interface and a wireless interface (such as a wireless fidelity (WI-FI) interface). The memory 105 may optionally be a storage device independent of the aforementioned processor 101. The memory 105 may be a high-speed random access memory (Random Access Memory, RAM) or a stable non-volatile memory (Non-Volatile Memory, NVM), such as at least one disk memory; the processor 101 may be a general-purpose processor, including a central processor, a network processor, etc., or may also be a digital signal processor, an application-specific integrated circuit, a field programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0218] Those skilled in the art can understand that the structure shown in the appendix Figure 2 does not constitute a limitation on the electronic device, and may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0219] As Figure 2 shown, the memory 105 as a storage medium may include an operating system, a network communication module, a user interface module, and an application program for implementing an intelligent education and teaching evaluation method.
[0220] In Figure 2In the electronic device shown, the network interface 103 is mainly used for data communication with a network server; the user interface 104 is mainly used for data interaction with a user; in the present application, the processor 101 and the memory 105 can be arranged in the electronic device, and the electronic device calls, through the processor 101, an application program stored in the memory 105 for implementing an intelligent education and teaching evaluation method to implement the above method.
[0221] Embodiment 5
[0222] This embodiment provides a computer-readable storage medium, on which a computer program is stored, and the processor executes the computer program to implement any of the above methods.
[0223] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or may be various devices including one or any combination of the above memories. The computer may be various computing devices including intelligent terminals and servers.
[0224] In the above embodiments of the present disclosure, the descriptions of the respective embodiments each have their own emphasis. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0225] In the several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of units can be a logical function division, and in actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of units or modules can be in an electrical or other form.
[0226] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0227] In addition, in each of the various embodiments of the present disclosure, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0228] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable non-volatile storage medium. Based on this understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a non-volatile storage medium and includes several instructions for causing a device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present disclosure. The aforementioned non-volatile storage medium includes: USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), external hard drives, magnetic disks, or optical discs and other various media that can store program codes.
[0229] The above are only the preferred embodiments of the present disclosure. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present disclosure, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present disclosure.
Claims
1. An intelligent education and teaching evaluation method, characterized in that: include: Acquire first teaching data and first test data, parse the first teaching data and the first test data, and obtain a first mapping between the first test data and the first teaching data, wherein the first mapping is configured as a mapping between each question in the first test data and each teaching period in the first teaching data; According to the first teaching data, obtaining first monitoring data corresponding to the first teaching data; According to the first teaching data, the first monitoring data and the first mapping, obtaining a second mapping between the first test data and the first monitoring data, wherein the second mapping is configured as a mapping between each question in the first test data and each monitoring data slice in the first monitoring data; Conducting a test on the target to be tested according to the first test data to obtain a first test result; According to the score of each question in the first test result, obtaining the teaching evaluation of each monitoring data slice in the first monitoring data; Establishing a teaching evaluation model based on the teaching evaluation of each monitoring data slice in the first monitoring data as a data set, wherein the teaching evaluation model is configured to take the monitoring data slice as input and take the teaching evaluation corresponding to the monitoring data slice as output; According to the teaching evaluation model, the monitoring data to be evaluated is input to obtain the evaluation result corresponding to the monitoring data to be evaluated; The step of establishing a teaching evaluation model based on the teaching evaluation of each monitoring data slice in the first monitoring data as a data set includes: Establishing an initial fusion model, and obtaining a training set and a test set according to the data set; Training the initial fusion model according to the training set and the test set to obtain a teaching evaluation model; The overall architecture of the initial fusion model is configured to include a gesture recognition sub-model, an expression analysis sub-model and a temporal fusion sub-model; The posture recognition sub-model is configured to: identify the body key points of the person in the video frame and extract the person's posture information from each frame; The expression analysis sub-model is configured to: extract the facial expression of the character from each frame and infer its emotional state; The temporal fusion sub-model is configured to: perform time series modeling to fuse the spatial and temporal features of video frames, fuse the outputs of each sub-model, combine the results of posture estimation and expression analysis with the output of the time series model, and generate the final score.
2. An intelligent education and teaching evaluation method according to claim 1, characterized in that: The acquiring of the first teaching data and the first test data, parsing the first teaching data and the first test data, and obtaining a first mapping between the first test data and the first teaching data includes: Acquire first test data, parse the first test data to obtain a first question, the first question belongs to the first test data; According to the first question, obtaining a first analysis corresponding to the first question; According to the first analysis, obtaining at least one first embedding vector of the first analysis; Acquire first teaching data, the first teaching data including at least one of text data, voice data and image data, parse the first teaching data, obtain a first teaching period, the first teaching period belongs to the first teaching data; According to the first teaching period, extracting features of data corresponding to the first teaching period and performing feature fusion to obtain a second embedding vector; Obtaining a mapping relationship between the first topic and the first teaching period according to the at least one first embedding vector and the second embedding vector; According to the mapping relationship between the first question and the first teaching period, a first mapping between the first test data and the first teaching data is obtained.
3. An intelligent education and teaching evaluation method according to claim 2, characterized in that: The first teaching data at least includes voice data; The step of parsing the first teaching data to obtain a first teaching period includes: According to the voice data, obtaining ASR data corresponding to the voice data; According to the ASR data, performing semantic analysis on the ASR data, and slicing the ASR data according to the semantic analysis result to obtain ASR slice data; According to the ASR slice data, obtaining slice data of the first teaching data; A first teaching period is obtained according to the slice data of the first teaching data.
4. The intelligent education and teaching evaluation method according to claim 2 is characterized in that: The first teaching data includes text data, voice data and image data; The extracting, according to the first teaching period, features of the data corresponding to the first teaching period and performing feature fusion to obtain a second embedding vector includes: According to the text data, extract context features of the text data using a pre-trained BERT model to obtain text features; According to the speech data, obtaining ASR data corresponding to the speech data, extracting context features of the ASR data using a pre-trained BERT model, and obtaining speech features; Extracting image features using an image recognition model according to the image data to obtain a first image feature; According to the image data, extract text in the image using OCR, and extract context features of the text in the image using a pre-trained BERT model to obtain a second image feature; According to the text features, the voice features, the first image features and the second image features, a multimodal fusion model is used to map the text features, the voice features, the first image features and the second image features to a unified embedding space to obtain a second embedding vector.
5. The intelligent education and teaching evaluation method according to claim 1 is characterized in that: The step of obtaining the teaching evaluation of each monitoring data slice in the first monitoring data according to the score of each question in the first test result includes: According to the second mapping, obtaining questions in the first test data corresponding to each monitoring data slice; Obtaining scoring results for each monitoring data slice in the first monitoring data according to the scores of each question in the first test data corresponding to each monitoring data slice and the proportion of the content involved in each monitoring data slice in each question; The scoring results are standardized to obtain teaching evaluations of each monitoring data slice in the first monitoring data.
6. The intelligent education and teaching evaluation method according to claim 5, characterized in that: According to the teaching evaluation model, inputting the monitoring data to be evaluated, and obtaining the evaluation results corresponding to the monitoring data to be evaluated, comprises: According to the monitoring data to be evaluated, the slice data to be evaluated is obtained, wherein the slice data to be evaluated belongs to the monitoring data to be evaluated; According to the teaching evaluation model, the slice data to be evaluated is input to obtain the evaluation of each slice data to be evaluated; According to the evaluation of each slice data to be evaluated, an evaluation result corresponding to the evaluated monitoring data is obtained.
7. An intelligent education and teaching evaluation system, characterized in that: Includes model management platform and evaluation platform; The model management platform is configured to: Acquire first teaching data and first test data, parse the first teaching data and the first test data, and obtain a first mapping between the first test data and the first teaching data, wherein the first mapping is configured as a mapping between each question in the first test data and each teaching period in the first teaching data; According to the first teaching data, obtaining first monitoring data corresponding to the first teaching data; According to the first teaching data, the first monitoring data and the first mapping, a second mapping between the first test data and the first monitoring data is obtained, wherein the second mapping is configured as a mapping between each question in the first test data and each monitoring data slice in the first monitoring data; Conducting a test on the target to be tested according to the first test data to obtain a first test result; According to the score of each question in the first test result, obtaining the teaching evaluation of each monitoring data slice in the first monitoring data; Establishing a teaching evaluation model based on the teaching evaluation of each monitoring data slice in the first monitoring data as a data set, wherein the teaching evaluation model is configured to take the monitoring data slice as input and take the teaching evaluation corresponding to the monitoring data slice as output; According to the teaching evaluation model, the monitoring data to be evaluated is input to obtain the evaluation result corresponding to the monitoring data to be evaluated; The step of establishing a teaching evaluation model based on the teaching evaluation of each monitoring data slice in the first monitoring data as a data set includes: Establishing an initial fusion model, and obtaining a training set and a test set according to the data set; Training the initial fusion model according to the training set and the test set to obtain a teaching evaluation model; The overall architecture of the initial fusion model is configured to include a gesture recognition sub-model, an expression analysis sub-model and a temporal fusion sub-model; The posture recognition sub-model is configured to: identify the body key points of the person in the video frame and extract the person's posture information from each frame; The expression analysis sub-model is configured to: extract the facial expression of the character from each frame and infer its emotional state; The temporal fusion sub-model is configured to: perform time series modeling to fuse the spatial and temporal features of the video frames, fuse the outputs of each sub-model, combine the results of posture estimation and expression analysis with the output of the time series model, and generate the final score; The model evaluation platform is configured to: According to the teaching evaluation model, the monitoring data to be evaluated is input to obtain the evaluation results corresponding to the monitoring data to be evaluated.
8. An intelligent education and teaching evaluation system according to claim 7, characterized in that: It also includes a data collection platform; The data collection platform is configured to collect and store at least one of the first teaching data, the first test data, the first monitoring data and the monitoring data to be evaluated.
9. A device, characterized in that: The device comprises a memory and a processor, wherein a computer program is stored in the memory, and the processor executes the computer program to implement the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
System and method for evaluating students' mastery degree in class based on multiple sensors
CN106878677A
Vocational education teaching evaluation method based on big data
CN113487213A