Multimedia recording and broadcasting teaching method and device based on artificial intelligence
By analyzing video and audio data from teachers and students, as well as data from the smart blackboard, and employing artificial intelligence methods, the problem of insufficient teaching quality assessment in multimedia recording and broadcasting teaching systems has been solved. This has enabled real-time assessment and report generation of the status of teachers and students, thereby improving teaching quality.
Patent Information
- Application Number
- CN202511775072.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2026-03-06
AI Technical Summary
Existing multimedia recording and broadcasting teaching systems lack in-depth analysis and intelligent evaluation of the teaching process, and cannot generate comprehensive teaching quality reports in real time and automatically. In particular, they lack sufficient monitoring and feedback on teachers' focus, explanation quality, and students' learning status.
The teaching method adopts an AI-based multimedia recording and broadcasting approach. By analyzing the video and audio data of teachers and students and combining it with the blackboard writing data of the smart blackboard, it calculates various scoring indicators for teachers and students, including status scores, language scores, activity levels and emotion scores, and generates short-term teaching quality scores and student status reports.
It enables real-time, automated evaluation of teaching quality for both teachers and students, improves the quality and efficiency of multimedia recording and broadcasting teaching, and provides comprehensive teaching quality reports.
Smart Images

Figure CN121617293A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of educational informatization technology, and in particular to a multimedia recording and broadcasting teaching method and device based on artificial intelligence. Background Technology
[0002] With the development of educational informatization, recorded courses have become an important part of the current education field. Through recorded teaching, teachers can share their lectures with students, and students can review the lectures at any time, improving the flexibility and efficiency of learning.
[0003] Existing multimedia recording and broadcasting teaching systems generally adopt a method based on video and audio acquisition. Some systems have also begun to introduce automation technology, using speech recognition technology to convert the content explained by the teacher into text, making it convenient for students to search and review during playback.
[0004] While the methods described above can enable multimedia recording and playback teaching, existing systems only focus on recording and playing back course content, lacking in-depth analysis and intelligent evaluation of the teaching process. The teacher's focus, the quality of their explanation, and the students' learning status are not effectively monitored and monitored. Current systems generally rely on manual or simple rules to evaluate teaching quality, failing to generate comprehensive teaching quality reports in real time and automatically. Therefore, there is an urgent need for a multimedia recording and playback teaching method and device that integrates artificial intelligence technology to improve the quality of multimedia recording and playback teaching. Summary of the Invention
[0005] This invention provides an artificial intelligence-based multimedia recording and broadcasting teaching method and a computer-readable storage medium, the main purpose of which is to improve the teaching quality of multimedia recording and broadcasting teaching.
[0006] To achieve the above objectives, this invention provides a multimedia recording and teaching method based on artificial intelligence, comprising: The recording and teaching environment and total course duration have been confirmed. The recording and teaching environment includes: teacher, multiple students, teacher's camera, student cameras, and smart blackboard. Number multiple students to obtain multiple numbered students; The total course duration is divided into multiple consecutive time periods according to preset time intervals; Perform the following operation for each of the multiple consecutive time periods: Based on the teacher's camera in a continuous time period and the teacher's confirmation of the teaching video in the recorded teaching environment, the teaching video includes: multiple video frames and audio stream; Teacher status scores were determined based on multiple video frames in the instructional video. Teacher language scores were identified based on the audio stream in the instructional videos; The smart blackboard is used to monitor teachers' handwriting and obtain handwriting data, which includes: average writing pressure, number of erasures, number of strokes, and writing speed sequence. Short-term teaching quality scores were determined based on the average writing pressure, number of erasures, number of strokes, writing speed sequence, teacher status score, and teacher language score in the blackboard writing data. Based on the student cameras in the recorded teaching environment and multiple numbered students, a student status score set is identified. The student status score set includes multiple student status scores, and each student status score corresponds one-to-one with a numbered student. By summarizing the short-term teaching quality scores, multiple short-term teaching quality scores are obtained. By summing up the student status rating sets, multiple student status rating sets are obtained; Based on multiple short-term teaching quality scores and multiple student status score sets, the teaching quality scores and multiple student status reports were confirmed, and multimedia recording and broadcasting teaching was completed.
[0007] Optionally, the step of determining the teacher's status score based on multiple video frames in the teaching video includes: Perform the following operation on each of the multiple video frames: The video frame resolution is determined, which includes: pixel length and pixel width; Target detection is performed on video frames to obtain teacher target data, which includes: teacher target image, vertex x coordinate, vertex y coordinate, bounding box length and bounding box width; Calculate the proportion of the teacher's image based on pixel length, pixel width, bounding box length, and bounding box width; By summarizing the teacher target data, multiple sets of teacher target data are obtained. By summing the percentage of teachers in the screen, we can obtain multiple percentages of teachers in the screen. The average percentage was determined based on the percentage of multiple teachers' screen shots, where the average percentage is the average of the percentages of multiple teachers' screen shots; Teacher activity and emotion scores were identified based on multiple teacher target data. Teacher status scores are calculated based on average percentage, teacher activity level, and emotional score.
[0008] Optionally, the step of determining teacher activity and emotion scores based on multiple teacher target data includes: For each teacher goal graph in multiple teacher goal datasets, perform the following operation: Keypoint detection is performed on the teacher's target image to obtain keypoint data. The keypoint data includes multiple keypoints, and each keypoint includes its x-coordinate and y-coordinate. Summarize the key point data to obtain multiple key point data; The amplitude of the movement is calculated based on the x-coordinates and y-coordinates of multiple key points in the data. Calculate the total displacement of the teacher based on the x-coordinates and y-coordinates of multiple vertices in multiple teacher target data; Teacher activity level is calculated based on the range of motion and the teacher's total displacement. Emotional scores were identified based on multiple teacher goal maps from multiple teacher goal data.
[0009] Optionally, the step of determining the emotion score based on multiple teacher goal maps from multiple teacher goal data includes: For each teacher goal graph in multiple teacher goal datasets, perform the following operation: Face extraction is performed on the teacher target image to obtain the teacher's face image; A pre-built facial expression recognition model is used to analyze the facial expressions of teachers in images, and the teacher's facial expression results are: happy, natural, or tired. If the teacher's expression is happy or natural, the preset single value will be used as the expression score. Otherwise, the preset zero value will be used as the expression score; Summarize the facial expression scores to obtain multiple facial expression scores; An emotion score is calculated based on multiple facial expression scores.
[0010] Optionally, the step of determining the teacher's language score based on the audio stream in the instructional video includes: The audio stream is divided into multiple audio segments using a preset segmentation interval; Number multiple audio segments to obtain multiple numbered audio segments; Perform the following operation on each of the multiple numbered audio segments: The numbered audio segments are divided according to the preset frame length to obtain multiple audio frames, wherein each audio frame includes multiple audio sample values. For each of the multiple audio frames, perform the following operation: Calculate the energy of a single frame based on multiple audio sample values in the audio frame; Calculate the root mean square value of a single frame based on the energy of that single frame. Sum the root mean square values of a single frame to obtain multiple root mean square values of a single frame. Summarize the energy of a single frame to obtain the energy of multiple single frames; The average energy is determined based on the energy of multiple single frames, where the average energy is the average value of the energy of multiple single frames; For each of the multiple single-frame energies, perform the following operation: Compare the single-frame energy with the average energy. If the single-frame energy is greater than or equal to the average energy, then the numbered audio segment corresponding to the single-frame energy is taken as the numbered speech segment. Otherwise, the numbered audio segment corresponding to the energy of a single frame will be used as the numbered silent segment; The numbered speech segments and numbered silence segments are summarized separately to obtain multiple numbered speech segments and multiple numbered silence segments; Based on multiple numbered speech segments and multiple numbered silent segments, the class rhythm score and speech-text set were determined. Obtain teaching content, perform keyword analysis on the teaching content, and obtain a keyword set; Keyword coverage was determined based on the speech-text set and keyword set; The volume fluctuation was determined based on multiple single-frame root mean square values, where the volume fluctuation was the standard deviation of the multiple single-frame root mean square values. The teacher's language score is calculated based on the lesson pacing, keyword coverage, and volume fluctuation, using the formula shown below: in, Indicates teacher language score, The score indicates the pace of the lesson. Indicates keyword coverage. Indicates volume fluctuation. It represents the natural logarithm.
[0011] Optionally, the step of determining the class rhythm score and speech-text set based on multiple numbered speech segments and multiple numbered silence segments includes: Multiple numbered speech segments are merged to obtain multiple target speech segments; For each of the multiple target speech segments, perform the following operation: Once the duration of the target speech is determined, the target speech segment is converted into text to obtain the speech-text. The total number of words is obtained by counting the words in the spoken text. The class speaking speed is calculated based on the total number of words and the duration of the target audio. The class speaking speed is the ratio of the total number of words to the duration of the target audio. Summarize the speech and text to obtain a speech and text set; By summarizing the speaking speed during class, multiple speaking speeds can be obtained. The average speaking speed was determined based on multiple speaking speeds in class, where the average speaking speed is the average of multiple speaking speeds in class. Multiple numbered silent segments are merged to obtain multiple target silent segments; For each of the multiple target silence segments, perform the following operation: The target silence duration is determined, and the target silence duration is compared with the preset duration threshold. If the target silence segment duration is greater than or equal to the duration threshold, then the target silence segment corresponding to the target silence segment duration is taken as the real silence segment; By summarizing the actual silent segments, multiple actual silent segments are obtained; The pause frequency was determined based on multiple target speech segments, multiple target silence segments, and multiple real silence segments; The lesson pace score is calculated based on the average speaking speed and the frequency of pauses.
[0012] Optionally, the short-term teaching quality score is determined based on the average writing pressure, erasure frequency, stroke count, writing speed sequence, teacher status score, and teacher language score from the blackboard data, including: The average writing speed was determined based on the writing speed sequence; The short-term teaching quality score is calculated based on average writing pressure, number of erasures, number of strokes, average writing speed, teacher status score, and teacher language score. The calculation formula is as follows: in, Indicates short-term teaching quality score. Indicates average writing pressure. Indicates the number of strokes. Indicates average writing speed. Indicates the number of erase / write cycles. The preset pressure influence coefficient, The preset writing influence coefficient, This is the preset teacher influence coefficient.
[0013] Optionally, the student status rating set confirmed based on student cameras and multiple numbered students in the recorded teaching environment includes: Perform the following operation on each of the multiple numbered students: The status of numbered students is monitored using student cameras in the recorded teaching environment to obtain student status videos, which include multiple status video frames. For each of the multiple state video frames, perform the following operation: The status video frame is converted to grayscale to obtain a grayscale video frame; Face detection is performed on grayscale video frames to obtain student facial images; The student's facial image is processed by illumination normalization to obtain the target image; The facial expression recognition model is used to identify facial expressions in the target image, and the recognition results are: smiling, frowning, sleepy or focused. The gaze direction vector is determined based on a pre-constructed pupil localization algorithm and the target image; Key point detection is performed on the target image to obtain student key point data; The recognition results are summarized to obtain multiple recognition results; By summing the line-of-sight vectors, multiple line-of-sight vectors are obtained; Summarize the key point data of students to obtain multiple key point data of students; Multiple recognition results are statistically merged to obtain a merged result, which includes: the total number of results, the number of smiles, and the number of times attention is focused. The student's focus level is determined based on the total number of results, the number of smiles, and the number of times they focus. The student's focus level is the ratio of the sum of the number of smiles and the number of times they focus to the total number of results. Obtain the space vector of the smart blackboard; The average line-of-sight vector is determined based on multiple line-of-sight direction vectors; Calculate student gaze level based on the smart blackboard spatial vector and average gaze vector; The range of student movements was determined based on multiple key data points of the students. Student status scores are calculated based on student concentration, student gaze, and student movement range. The student status scores are compiled to obtain a set of student status scores.
[0014] Optionally, the step of determining the teaching quality score and multiple student status reports based on multiple short-term teaching quality scores and multiple student status score sets includes: The teaching quality score is determined based on multiple short-term teaching quality scores, where the teaching quality score is the average of multiple short-term teaching quality scores; Statistical analysis was performed on multiple student status rating sets to obtain comprehensive student ratings. Multiple student status reports were identified based on the comprehensive scores of multiple students.
[0015] To achieve the above objectives, the present invention also provides a multimedia recording and teaching device based on artificial intelligence, comprising: The basic environment confirmation module is used to confirm the recording and teaching environment and the total course duration. The recording and teaching environment includes: teacher, multiple students, teacher's camera, student cameras and smart blackboard. Multiple students are numbered to obtain multiple numbered students. The total course duration is divided according to preset time intervals to obtain multiple consecutive time periods. The teacher status monitoring module is used to perform the following operations for each of multiple consecutive time periods: based on the consecutive time period, the teacher's camera in the recorded teaching environment, and the teacher's confirmed teaching video, wherein the teaching video includes: multiple video frames and audio stream; based on the multiple video frames in the teaching video, the teacher's status score is confirmed; and based on the audio stream in the teaching video, the teacher's language score is confirmed. The student status monitoring module is used to monitor the teacher's handwriting on the smart blackboard and obtain handwriting data. The handwriting data includes: average writing pressure, number of erasures, number of strokes, and writing speed sequence. Based on the average writing pressure, number of erasures, number of strokes, writing speed sequence, teacher status score, and teacher language score in the handwriting data, a short-term teaching quality score is determined. Based on the student cameras in the recorded teaching environment and multiple numbered students, a student status score set is determined. The student status score set includes: multiple student status scores, and each student status score corresponds one-to-one with a numbered student. The teaching quality assessment module is used to summarize short-term teaching quality scores, obtain multiple short-term teaching quality scores, summarize student status score sets, obtain multiple student status score sets, and confirm the teaching quality score and multiple student status reports based on multiple short-term teaching quality scores and multiple student status score sets, thus completing multimedia recording and broadcasting teaching.
[0016] To address the above problems, the present invention also provides an electronic device, the electronic device comprising: Memory, storing at least one instruction; and The processor executes the instructions stored in the memory to implement the aforementioned AI-based multimedia recording and teaching method.
[0017] To address the aforementioned problems, the present invention also provides a computer-readable storage medium storing at least one instruction, which is executed by a processor in an electronic device to implement the aforementioned artificial intelligence-based multimedia recording and teaching method.
[0018] To address the problems described in the background art, this invention identifies the recording and broadcasting teaching environment and the total course duration. The recording and broadcasting teaching environment includes: a teacher, multiple students, a teacher's camera, student cameras, and a smart blackboard. This invention provides a material basis for subsequent video capture of the teacher and students by identifying the recording and broadcasting teaching environment. Identifying the total course duration allows for subsequent division of the total course duration, and multiple students are then numbered, resulting in multiple numbered students. This invention simplifies the subsequent analysis and processing of each student by numbering them. Dividing the total course duration according to preset time intervals yields multiple consecutive time periods. This invention effectively manages the duration of a lesson... The recording is divided into multiple time periods to facilitate timely analysis of teacher teaching quality and student learning quality, thereby improving the teaching quality of multimedia recording and broadcasting. For each consecutive time period, the following operations are performed: based on the consecutive time period, the teacher's camera in the recording and broadcasting environment, and the teacher's confirmed teaching video, the teaching video includes multiple video frames and an audio stream. This embodiment of the invention obtains the teacher's teaching video by filming the teacher, enabling subsequent frame-by-frame analysis and improving the teaching quality of multimedia recording and broadcasting. Teacher status scores are determined based on multiple video frames in the teaching video, and teacher language scores are determined based on the audio stream in the teaching video. This embodiment of the invention improves the teaching quality of multimedia recording and broadcasting by analyzing the video frame by frame. The system analyzes video frames to calculate teacher status scores and analyzes audio streams to calculate teacher language scores. This provides a prerequisite for calculating short-term teaching quality scores based on teacher status and language scores, thus improving the teaching quality of multimedia recorded teaching. The system also uses a smart blackboard to monitor teacher handwriting, obtaining handwriting data including average writing pressure, erasing frequency, stroke count, and writing speed sequence. This embodiment of the invention utilizes a smart blackboard to acquire teacher handwriting data during class, providing a foundation for subsequent processing and calculation. Based on the average writing pressure, erasing frequency, stroke count, writing speed sequence, teacher status score, and teacher language score in the handwriting data, short-term teaching quality scores are determined. The invention, through analysis of blackboard writing data, calculates a blackboard writing quality score reflecting the teacher's blackboard writing quality. This score, combined with teacher status and language scores, is used to calculate a comprehensive short-term teaching quality score, thus improving the teaching quality of multimedia recorded teaching. Based on student cameras and multiple numbered students in the recorded teaching environment, a student status score set is identified. This set includes multiple student status scores, with each score corresponding to a unique student number. The invention utilizes student cameras to capture video of students and then analyzes the captured video frame-by-frame to calculate the student status score set, providing the basis for generating subsequent student status reports and summarizing the short-term teaching quality score.This invention obtains multiple short-term teaching quality scores, summarizes student status score sets, and generates multiple student status score sets. Based on these scores, a teaching quality score and multiple student status reports are determined, thus completing multimedia recorded teaching. Therefore, this invention improves the teaching quality of multimedia recorded teaching by calculating a teaching quality score from multiple short-term scores and generating multiple student status reports from multiple student status score sets. Attached Figure Description
[0019] Figure 1 A flowchart illustrating an artificial intelligence-based multimedia recording and teaching method according to an embodiment of the present invention; Figure 2 A functional block diagram of an artificial intelligence-based multimedia recording and teaching device provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of an electronic device for implementing the AI-based multimedia recording and teaching method, as provided in an embodiment of the present invention.
[0020] Explanation of reference numerals in the attached figures: 10. Electronic device; 11. Processor; 12. Memory; 13. Bus.
[0021] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0022] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0023] This application provides an AI-based multimedia recording and teaching method. The executing entity of this AI-based multimedia recording and teaching method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the AI-based multimedia recording and teaching method can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster.
[0024] Reference Figure 1 The diagram shown is a flowchart illustrating an artificial intelligence-based multimedia recording and teaching method according to an embodiment of the present invention. In this embodiment, the artificial intelligence-based multimedia recording and teaching method includes: S1. Confirm the recording and teaching environment and the total course duration. The recording and teaching environment includes: teacher, multiple students, teacher's camera, student cameras, and smart blackboard.
[0025] For example, Zhang is the school's director of teaching. He needs to evaluate the teaching status of teachers and students in the classroom and implement multimedia online recording and broadcasting teaching. Therefore, Zhang determines the recording and broadcasting filtering teaching environment and the total course duration to facilitate subsequent multimedia online recording and broadcasting teaching and teaching quality evaluation.
[0026] It should be explained that the recorded teaching environment refers to the necessary environment for conducting recorded teaching activities. This environment includes: a teacher, multiple students, a teacher's camera, student cameras, and a smart blackboard. The teacher is the one who lectures and imparts knowledge in the classroom, while the students are the learners. Both the teacher's and student cameras are single-type cameras; optionally, the FOTRIC 600C series can be used. The smart blackboard is an interactive blackboard for smart classrooms; optionally, the Oudi fifth-generation smart classroom blackboard can be used. The total course duration refers to the time of one class period.
[0027] S2. Number the multiple students to obtain multiple numbered students. Divide the total course duration into multiple consecutive time periods according to the preset time intervals.
[0028] For example, if the multiple students are: Zhang, Qian, Sun, Li, and Wang, then the multiple students are numbered, resulting in multiple numbered students as St1, St2, St3, St4, and St5. If the total course duration is 45 minutes and the time interval is 1 minute, then the total course duration is divided according to the preset time interval, resulting in multiple consecutive time intervals as: 0-1, 1-2, 2-3, ..., 43-44, 44-45.
[0029] S3. Perform the following operation for each of the multiple consecutive time periods: Based on the consecutive time period, the teacher's camera in the recording and broadcasting teaching environment, and the teacher's confirmed teaching video, wherein the teaching video includes: multiple video frames and audio stream.
[0030] It should be explained that the "based on a continuous time period, a teacher's camera in a recorded teaching environment, and the teacher's confirmation of the teaching video" refers to recording the teacher's class using a teacher's camera in a recorded teaching environment within a continuous time period. This video of the teacher teaching is the teaching video. Furthermore, the method of recording the teacher's class using a teacher's camera in a recorded teaching environment within a continuous time period is existing technology and will not be elaborated upon here. A video frame is the basic unit that constitutes a video; it can be understood as a static image. An audio stream is a sound signal continuously recorded over time, usually stored in waveform data format, and has a certain duration.
[0031] S4. Teacher status score is determined based on multiple video frames in the teaching video.
[0032] Specifically, the method of determining the teacher's status score based on multiple video frames in the teaching video includes: Perform the following operation on each of the multiple video frames: The video frame resolution is determined, which includes: pixel length and pixel width; Target detection is performed on video frames to obtain teacher target data, which includes: teacher target image, vertex x coordinate, vertex y coordinate, bounding box length and bounding box width; The teacher's screen proportion is calculated based on pixel length, pixel width, bounding box length, and bounding box width. The calculation formula is as follows: in, This indicates the proportion of the teacher's image in the picture. Indicates the length of the bounding box. Indicates the width of the bounding box. Indicates pixel length. Indicates pixel width; By summarizing the teacher target data, multiple sets of teacher target data are obtained. By summing the percentage of teachers in the screen, we can obtain multiple percentages of teachers in the screen. The average percentage was determined based on the percentage of multiple teachers' screen shots, where the average percentage is the average of the percentages of multiple teachers' screen shots; Teacher activity and emotion scores were identified based on multiple teacher target data. Teacher status scores are calculated based on average percentage, teacher activity level, and emotional rating, using the following formula: in, Indicates teacher status rating. Indicates the average percentage. Indicates teacher activity level, Indicates mood rating. The preset percentage influence parameter, The preset active influence parameters, These are preset parameters for the impact of emotions.
[0033] It should be explained that video frame resolution refers to the number of pixels in a single video frame in the horizontal and vertical directions, usually expressed as: pixel width. Pixel length. The number of pixels in a video frame in the horizontal direction is the pixel length, and the number of pixels in a video frame in the vertical direction is the pixel width. The video frame resolution can be obtained from the teacher's camera. For example, if the resolution of the teacher's camera is set to 1280×720 when shooting, then the video frame resolution is 1280×720. Target detection of the video frame to obtain teacher target data means: using a target detection model to identify the image of the teacher in the video frame. The method for using a target detection model to identify the image of the teacher in the video frame is existing technology and will not be elaborated here. Optionally, the YOLOv3 algorithm can be used as the target detection model. The teacher target image refers to the image of the teacher in the video frame identified by the target detection model.
[0034] Understandably, the vertex x-coordinate refers to the x-coordinate of the top-left vertex of the teacher's target image in the video frame, and the vertex y-coordinate refers to the y-coordinate of the top-left vertex of the teacher's target image in the video frame. The bounding box length refers to the length of the teacher's target image, and the bounding box width refers to the width of the teacher's target image. The teacher's screen share refers to the ratio of the area occupied by the teacher in a video frame to the total area of the entire video frame, used to measure the teacher's visibility in the video frame. The larger the teacher's screen share, the larger the ratio of the area occupied by the teacher in a video frame to the total area of the entire video frame, and the greater the teacher's visibility in the video frame. The teacher's status score reflects the teacher's activity level; the higher the teacher's status score, the greater the teacher's activity level. The percentage influence parameter, activity influence parameter, and emotion influence parameter are all values manually set by the school's academic director, and are optional. The percentage influence parameter is 0.3, the activity influence parameter is 0.3, and the emotion influence parameter is 0.4.
[0035] Specifically, the teacher activity and emotion scores identified based on multiple teacher target data include: For each teacher goal graph in multiple teacher goal datasets, perform the following operation: Keypoint detection is performed on the teacher's target image to obtain keypoint data. The keypoint data includes multiple keypoints, and each keypoint includes its x-coordinate and y-coordinate. Summarize the key point data to obtain multiple key point data; The motion amplitude is calculated based on the x-coordinates and y-coordinates of multiple key points from multiple key point data. The calculation formula is as follows: in, Indicates the range of motion. This represents the first of multiple key data points. The first key data point The x-coordinates of the key points. This represents the first of multiple key data points. The first key data point The x-coordinates of the key points. This represents the first of multiple key data points. The first key data point The y-coordinate of each key point This represents the first of multiple key data points. The first key data point The y-coordinate of each key point The number of teacher objective data points in a set of multiple teacher objective data points; The total displacement of the teacher is calculated based on the x-coordinates and y-coordinates of multiple vertices in multiple teacher target data. The calculation formula is as follows: in, Indicates the total displacement of teachers. This represents the x-coordinate of multiple vertices in multiple vertex coordinate data. The x-coordinates of the vertices This represents the y-coordinate of multiple vertices in multiple vertex coordinate data. The y-coordinates of each vertex This represents the x-coordinate of multiple vertices in multiple vertex coordinate data. The x-coordinates of the vertices This represents the y-coordinate of multiple vertices in multiple vertex coordinate data. The y-coordinates of each vertex; Teacher activity level is calculated based on the range of motion and total teacher displacement, using the following formula: in, Represents the natural constant. To achieve the preset optimal range of motion, The preset optimal total displacement; Emotional scores were identified based on multiple teacher goal maps from multiple teacher goal data.
[0036] It should be explained that the keypoint detection in the teacher target image refers to using a keypoint detection model to detect key parts of the teacher in the teacher target image. The method for using a keypoint detection model to detect key parts of the teacher in the teacher target image is existing technology and will not be elaborated here. Optionally, the OpenPose model can be used as the keypoint detection model. Keypoint data refers to the data of key parts of the teacher in the teacher target image detected by the keypoint detection model, and the keypoint data includes: multiple keypoints, keypoint x-coordinates, and keypoint y-coordinates. Here, a keypoint refers to a key part of the teacher, the keypoint x-coordinate refers to the x-coordinate of the keypoint in the teacher target image, and the keypoint y-coordinate refers to the y-coordinate of the keypoint in the teacher target image.
[0037] Understandably, the range of motion reflects the teacher's level of activity; the larger the range of motion, the more active the teacher. The total displacement reflects the distance the teacher moves; the greater the total displacement, the greater the distance the teacher moves. Teacher activity reflects the richness of the teacher's movements during class; the greater the activity, the richer the teacher's movements. The optimal range of motion and the optimal total displacement are both manually set by the school's academic director based on the average range of motion and displacement of teachers when their teaching quality is high. For example, if the average range of motion and displacement of teachers when their teaching quality is high is 50 centimeters and 10 meters respectively, then the optimal range of motion is 50 centimeters and the optimal total displacement is 10 meters.
[0038] Specifically, the process of identifying emotion scores based on multiple teacher goal maps from multiple teacher goal data includes: For each teacher goal graph in multiple teacher goal datasets, perform the following operation: Face extraction is performed on the teacher target image to obtain the teacher's face image; A pre-built facial expression recognition model is used to analyze the facial expressions of teachers in images, and the teacher's facial expression results are: happy, natural, or tired. If the teacher's expression is happy or natural, the preset single value will be used as the expression score. Otherwise, the preset zero value will be used as the expression score; Summarize the facial expression scores to obtain multiple facial expression scores; An emotion score is calculated based on multiple facial expression scores, using the following formula: in, This indicates the score of multiple expressions. Each expression scores a point.
[0039] It should be explained that the face extraction of the teacher target image refers to: using a face detection model to extract the image of the teacher's face in the teacher target image. The method of using a face detection model to extract the image of the teacher's face in the teacher target image is existing technology and will not be described in detail here. The image of the teacher's face in the teacher target image is the teacher's face picture. Optionally, the Viola-Jones algorithm can be used as the face detection model.
[0040] The operating principle of the facial expression recognition model is as follows: First, a picture of a teacher's face or a target image of a student's face is input into the model. A normalized face image is obtained through face detection and alignment. Next, a deep learning network automatically extracts local and overall facial features to obtain facial feature information. Finally, this facial feature information is mapped to an expression category space to obtain the corresponding expression recognition result. If the model input is a picture of a teacher's face, the corresponding expression recognition result is happy, natural, or tired. If the model input is a target image of a student's face, the corresponding expression recognition result is smiling, frowning, sleepy, or focused. The teacher's expression result refers to the expression recognition result output by the model after the teacher's face image is input. The emotion score reflects the positiveness of the teacher's expression; the higher the emotion score, the more positive the teacher's expression. Optionally, a single value is "1", and a zero value is "0".
[0041] S5. Teacher language scores are determined based on the audio stream in the instructional video.
[0042] Specifically, the method of determining teacher language scores based on audio streams from instructional videos includes: The audio stream is divided into multiple audio segments using a preset segmentation interval; Number multiple audio segments to obtain multiple numbered audio segments; Perform the following operation on each of the multiple numbered audio segments: The numbered audio segments are divided according to the preset frame length to obtain multiple audio frames, wherein each audio frame includes multiple audio sample values. For each of the multiple audio frames, perform the following operation: The energy of a single frame is calculated based on multiple audio sample values in the audio frame, using the following formula: in, Indicates the energy of a single frame. Represents the first of multiple audio sample values Each audio sample value, The number of audio sample values in a set of multiple audio sample values; The root mean square value of a single frame is calculated based on the energy of that single frame, using the following formula: in, This represents the root mean square value of a single frame. Indicates frame length; Sum the root mean square values of a single frame to obtain multiple root mean square values of a single frame. Summarize the energy of a single frame to obtain the energy of multiple single frames; The average energy is determined based on the energy of multiple single frames, where the average energy is the average value of the energy of multiple single frames; For each of the multiple single-frame energies, perform the following operation: Compare the single-frame energy with the average energy. If the single-frame energy is greater than or equal to the average energy, then the numbered audio segment corresponding to the single-frame energy is taken as the numbered speech segment. Otherwise, the numbered audio segment corresponding to the energy of a single frame will be used as the numbered silent segment; The numbered speech segments and numbered silence segments are summarized separately to obtain multiple numbered speech segments and multiple numbered silence segments; Based on multiple numbered speech segments and multiple numbered silent segments, the class rhythm score and speech-text set were determined. Obtain teaching content, perform keyword analysis on the teaching content, and obtain a keyword set; Keyword coverage was determined based on the speech-text set and keyword set; The volume fluctuation was determined based on multiple single-frame root mean square values, where the volume fluctuation was the standard deviation of the multiple single-frame root mean square values. The teacher's language score is calculated based on the lesson pacing, keyword coverage, and volume fluctuation, using the formula shown below: in, Indicates teacher language score, The score indicates the pace of the lesson. Indicates keyword coverage. Indicates volume fluctuation. It represents the natural logarithm.
[0043] For example, if the audio stream is 45 minutes long and the segmentation interval is 1 minute, the audio stream is segmented using the preset segmentation interval, resulting in multiple audio segments: 0-1, 1-2, 2-3, ..., 43-44, 44-45. After numbering these multiple audio segments, the resulting numbered audio segments are: So1, So2, So3, ..., So44, So45. If an audio segment is 1 minute long and the preset frame length is 25 ms, the numbered audio segments are segmented according to the preset frame length, resulting in multiple audio frames: F1, F2, F3, ..., F2399, F2400.
[0044] It should be explained that audio sampling value refers to the instantaneous amplitude of the sound wave, and single-frame energy reflects the strength of the sound signal in the audio frame; the greater the single-frame energy, the stronger the sound signal in the audio frame. The root mean square (RMS) value of a single frame is used to measure the average amplitude of the audio signal in a single frame; the greater the RMS value, the greater the average amplitude of the audio signal in the single frame. Numbered speech segments refer to the numbered audio segments corresponding to single-frame energies greater than or equal to the average energy, and numbered silence segments refer to the numbered audio segments corresponding to single-frame energies less than the average energy. Teaching content refers to the content taught by the teacher in class, which can be obtained from the teacher's lesson preparation notes. The keyword analysis of the teaching content refers to: segmenting the text data of the teaching content into words and performing word frequency statistics, then combining this with keyword extraction algorithms (such as TF-IDF) to filter out words that can represent the teaching theme and core information, and compiling these words into a set, i.e., a keyword set.
[0045] Understandably, determining the keyword coverage rate based on the speech-text set and keyword set means comparing the speech-text set with the keyword set, counting how many keywords from the keyword set appear in the speech-text set, and expressing the keyword coverage rate as the ratio of the number of appearing keywords to the total number of keywords. The formula for calculating the keyword coverage rate is as follows: ,in, Indicates keyword coverage. Indicates the number of keywords. This indicates the total number of keywords. Keyword coverage rate measures the degree to which the teacher's content covers the key teaching points; the higher the keyword coverage rate, the better the teacher's content covers the key teaching points. Teacher language score reflects the teacher's level of engagement in class; the higher the teacher language score, the more engaged the teacher is in using language during class.
[0046] In detail, the process of determining the lesson rhythm score and speech-text set based on multiple numbered speech segments and multiple numbered silence segments includes: Multiple numbered speech segments are merged to obtain multiple target speech segments; For each of the multiple target speech segments, perform the following operation: Once the duration of the target speech is determined, the target speech segment is converted into text to obtain the speech-text. The total number of words is obtained by counting the words in the spoken text. The class speaking speed is calculated based on the total number of words and the duration of the target audio. The class speaking speed is the ratio of the total number of words to the duration of the target audio. Summarize the speech and text to obtain a speech and text set; By summarizing the speaking speed during class, multiple speaking speeds can be obtained. The average speaking speed was determined based on multiple speaking speeds in class, where the average speaking speed is the average of multiple speaking speeds in class. Multiple numbered silent segments are merged to obtain multiple target silent segments; For each of the multiple target silence segments, perform the following operation: The target silence duration is determined, and the target silence duration is compared with the preset duration threshold. If the target silence segment duration is greater than or equal to the duration threshold, then the target silence segment corresponding to the target silence segment duration is taken as the real silence segment; By summarizing the actual silent segments, multiple actual silent segments are obtained; The pause frequency was determined based on multiple target speech segments, multiple target silence segments, and multiple real silence segments; The lesson pacing score is calculated based on average speaking speed and pause frequency, using the following formula: in, Indicates average speaking speed. Indicates the frequency of pauses. The preset ideal speaking speed.
[0047] It should be understood that the process of merging multiple numbered speech segments to obtain multiple target speech segments involves: first, arranging all numbered speech segments in numerical order; then, grouping the numbered speech segments according to the continuity of their numbers: when the numbers of several speech segments are numerically consecutive, these segments are merged into one target speech segment; when the number of a speech segment is numerically discontinuous with either the preceding or following segment, that segment independently becomes a target speech segment. Thus, consecutively numbered speech segments are merged into a single target speech segment, while discontinuous segments are treated as separate target speech segments. Finally, the target speech segments are aggregated to obtain multiple target speech segments. For example, if multiple numbered speech segments are: So1, So2, So3, So5, So6, So8, So10, So11, So12, then the multiple numbered speech segments are merged into multiple target speech segments as: (So1, So2, So3), (So5, So6), (So8), (So10, So11, So12).
[0048] It should be explained that the target speech duration refers to the length of the target speech segment. The text-to-text processing of the target speech segment refers to performing text recognition on the target speech segment to obtain its text information. The method for performing text recognition on the target speech segment is existing technology and will not be elaborated here. The text information of the target speech segment is the speech-to-text format. The formula for calculating the speaking speed in class is as follows: ,in, Indicates the speaking speed during class. Indicates the total number of words. This indicates the duration of the target speech. The word count for the speech-text refers to counting the number of words in the text information; this number of words is the total word count. A speech-text set is a collection of multiple speech-texts.
[0049] It should be understood that the method of merging multiple numbered silence segments to obtain multiple target silence segments is the same as the method of merging multiple numbered speech segments to obtain multiple target speech segments, and will not be described again here.
[0050] Understandably, the target silence duration refers to the time of the target silence segment. The duration threshold is a value set manually by the school's academic director; optionally, the duration threshold is 2 seconds. A true silence segment refers to a target silence segment whose duration is greater than or equal to the duration threshold. Determining the pause frequency based on multiple target speech segments, multiple target silence segments, and multiple true silence segments involves counting the number of target speech segments, the number of target silence segments, and the number of true silence segments. The number of target speech segments represents the number of target speech occurrences, the number of target silence segments represents the number of target silence occurrences, and the number of true silence segments represents the number of true silence occurrences. The pause frequency is calculated using the ratio of the number of true silence occurrences to the sum of the number of target silence occurrences and the number of target speech occurrences. The formula for calculating the pause frequency is as follows: ,in, Indicates the frequency of pauses. Indicates the actual number of times the device was muted. Indicates the number of times the target is silenced. This indicates the number of times the target speech is heard. The lesson rhythm score reflects the naturalness of the teacher's pacing during class; the higher the score, the more natural the teacher's pacing. The ideal speaking speed is a value set by the school's academic director based on the average speaking speed of teachers assuming excellent teaching quality. For example, if the average speaking speed of a teacher assuming excellent teaching quality is 3 words per second, then this is 3 words per second.
[0051] S6. Use the smart blackboard to monitor the teacher's handwriting and obtain handwriting data, including: average writing pressure, number of erasures, number of strokes, and writing speed sequence. Based on the average writing pressure, number of erasures, number of strokes, writing speed sequence, teacher status score, and teacher language score in the handwriting data, a short-term teaching quality score is determined.
[0052] It should be explained that the use of smart blackboards to monitor teachers' handwriting refers to using smart blackboards to monitor the data of teachers' handwriting during class. The method of using smart blackboards to monitor the data of teachers' handwriting during class is existing technology and will not be elaborated here. The data of teachers' handwriting during class is the blackboard data, which includes: average writing pressure, number of erasures, number of strokes, and writing speed sequence. Among them, average writing pressure refers to the average contact pressure of teachers when writing on the smart blackboard with writing tools (such as smart pens), number of erasures refers to the total number of times teachers use the erase function to clear the blackboard during the teaching process, number of strokes refers to the number of effective writing strokes formed during the blackboard writing process, and writing speed sequence refers to the writing speed records at different times during the entire blackboard writing process.
[0053] In detail, the short-term teaching quality score is determined based on the average writing pressure, erasure frequency, stroke count, writing speed sequence, teacher status score, and teacher language score from the blackboard data, including: The average writing speed was determined based on the writing speed sequence; The short-term teaching quality score is calculated based on average writing pressure, number of erasures, number of strokes, average writing speed, teacher status score, and teacher language score. The calculation formula is as follows: in, Indicates short-term teaching quality score. Indicates average writing pressure. Indicates the number of strokes. Indicates average writing speed. Indicates the number of erase / write cycles. The preset pressure influence coefficient, The preset writing influence coefficient, This is the preset teacher influence coefficient.
[0054] It should be explained that determining the average writing speed based on the writing speed sequence means: statistically processing multiple writing speed values recorded by the teacher over time during the blackboard writing process, and taking the arithmetic mean of this speed sequence to obtain the teacher's average writing speed throughout the entire blackboard writing process. Short-term teaching quality scores reflect the quality of a teacher's lessons over a continuous period; the higher the short-term teaching quality score, the higher the quality of the teacher's lessons over that continuous period. The stress influence coefficient, writing influence coefficient, and teacher influence coefficient are all values arbitrarily set by the school's academic director. Optionally, the stress influence coefficient is 0.3, the writing influence coefficient is 0.3, and the teacher influence coefficient is 0.4.
[0055] S7. Based on the student cameras in the recorded teaching environment and multiple numbered students, a student status score set is identified. The student status score set includes multiple student status scores, and each student status score corresponds one-to-one with a numbered student. Short-term teaching quality scores are summarized to obtain multiple short-term teaching quality scores. The student status score set is then summarized to obtain multiple student status score sets.
[0056] Specifically, the student status rating set confirmed based on student cameras and multiple numbered students in the recorded teaching environment includes: Perform the following operation on each of the multiple numbered students: The status of numbered students is monitored using student cameras in the recorded teaching environment to obtain student status videos, which include multiple status video frames. For each of the multiple state video frames, perform the following operation: The status video frame is converted to grayscale to obtain a grayscale video frame; Face detection is performed on grayscale video frames to obtain student facial images; The student's facial image is processed by illumination normalization to obtain the target image; The facial expression recognition model is used to identify facial expressions in the target image, and the recognition results are: smiling, frowning, sleepy or focused. The gaze direction vector is determined based on a pre-constructed pupil localization algorithm and the target image; Key point detection is performed on the target image to obtain student key point data; The recognition results are summarized to obtain multiple recognition results; By summing the line-of-sight vectors, multiple line-of-sight vectors are obtained; Summarize the key point data of students to obtain multiple key point data of students; Multiple recognition results are statistically merged to obtain a merged result, which includes: the total number of results, the number of smiles, and the number of times attention is focused. The student's focus level is determined based on the total number of results, the number of smiles, and the number of times they focus. The student's focus level is the ratio of the sum of the number of smiles and the number of times they focus to the total number of results. Obtain the space vector of the smart blackboard; The average line-of-sight vector is determined based on multiple line-of-sight direction vectors; Student gaze intensity is calculated based on the smart blackboard's spatial vector and average gaze vector, using the following formula: in, Indicates student attention level. Represents the average line-of-sight vector. Represents the space vector of the smart blackboard. The preset line-of-sight offset parameters, Represents the inverse cosine function. represents the natural exponential function, and represents the magnitude of the vector; The range of student movements was determined based on multiple key data points of the students. Student status scores are calculated based on student focus, student gaze intensity, and student movement range, using the following formula: in, Indicates student status rating. Indicates student concentration. Indicates the range of motion of the student. These are preset parameters affecting the action. The student status scores are compiled to obtain a set of student status scores.
[0057] It should be explained that the phrase "using student cameras in the recording and teaching environment to monitor the status of numbered students and obtain student status videos" refers to: using student cameras in the recording and teaching environment to capture videos of numbered students. These videos of numbered students constitute the student status videos. Furthermore, the method of capturing videos of numbered students using student cameras in the recording and teaching environment is existing technology and will not be elaborated upon here. A status video frame is the basic unit constituting a status video, and can be understood as a static image reflecting the student's status. The phrase "converting the status video frame to grayscale" refers to: converting each pixel in the status video frame to grayscale. The method of converting each pixel in the status video frame to grayscale is existing technology and will not be elaborated upon here. A grayscale video frame refers to a video frame after grayscale processing.
[0058] It should be understood that the method for performing face detection on grayscale video frames to obtain student facial images is the same as the method for extracting faces from teacher target images to obtain teacher facial images, and will not be repeated here. The student facial image refers to an image of a student's face. The method for performing illumination normalization processing on the student facial image refers to eliminating brightness differences in the student facial image. This method for eliminating brightness differences in student facial images is existing technology and will not be repeated here. The target image refers to the student facial image after illumination normalization processing. The method for performing expression recognition on the target image using an expression recognition model to obtain recognition results is the same as the method for performing expression analysis on teacher facial images using a pre-built expression recognition model to obtain teacher expression results, and will not be repeated here. The recognition result refers to the expression recognition result output after inputting the target image into the expression recognition model.
[0059] It should be explained that the determination of the gaze direction vector based on the pre-constructed pupil localization algorithm and the target image refers to: using the pupil localization algorithm to accurately detect the pupil position in the student's facial image, obtaining the pupil position, and then calculating the student's gaze direction vector based on the positional relationship of the pupil in the student's facial image. The gaze direction vector represents a three-dimensional vector of the student's gaze direction. The method for detecting key points in the target image to obtain student key point data is the same as the method for detecting key points in the teacher target image to obtain key point data, and will not be repeated here.
[0060] It should be understood that the formula for calculating student concentration is as follows: ,in, Indicates student concentration. Indicates the number of smiles. Indicates the number of times focused. This indicates the total number of results.
[0061] For example, if multiple recognition results are: smiling, smiling, smiling, frowning, drowsy, focused, then the multiple recognition results are statistically merged, and the merged result is (6, 3, 1), where 6 represents the total number of results, 3 represents the number of smilings, and 1 represents the number of focused moments.
[0062] Understandably, the smart blackboard spatial vector refers to the unit vector of the smart blackboard in the classroom space, used for comparison with the student's gaze direction vector. The determination of the average gaze vector based on multiple gaze direction vectors involves statistically processing multiple gaze direction vectors and obtaining a single vector reflecting the overall gaze direction of the students—the average gaze vector—through vector arithmetic averaging. Student gaze intensity reflects the student's level of focus in class; the greater the gaze intensity, the greater the student's level of focus. The gaze offset parameter is a value set manually by the school's academic director; optionally, the gaze offset parameter is 0.2 radians. The method for determining the student's movement amplitude based on multiple student keypoint data is the same as the method for calculating the movement amplitude based on the x-coordinates and y-coordinates of multiple keypoints in the keypoint data, and will not be elaborated here. The student status score reflects the student's level of attentiveness in class; the higher the student status score, the higher the student's level of attentiveness. The action influence parameter is a value set manually by the school's academic director; optionally, the action influence parameter is 1. The student status score set is a collection of multiple student status scores.
[0063] S8. Based on multiple short-term teaching quality scores and multiple student status score sets, confirm the teaching quality score and multiple student status reports, and complete the multimedia recording and broadcasting teaching.
[0064] In detail, the process of determining the teaching quality score and multiple student status reports based on multiple short-term teaching quality scores and multiple student status score sets includes: The teaching quality score is determined based on multiple short-term teaching quality scores, where the teaching quality score is the average of multiple short-term teaching quality scores; Statistical analysis was performed on multiple student status rating sets to obtain comprehensive student ratings. Multiple student status reports were identified based on the comprehensive scores of multiple students.
[0065] It should be explained that the statistical analysis of multiple student status score sets to obtain multiple student comprehensive scores refers to: statistically processing multiple student status score sets (such as averaging or weighted averaging) to generate a comprehensive score for each student, thus obtaining multiple student comprehensive scores. The confirmation of multiple student status reports based on multiple student comprehensive scores means: using each student's comprehensive score in class as input, evaluating and classifying each student's performance according to the scoring criteria, and generating a corresponding student status report. The scoring criteria can be obtained from the school's student assessment standards.
[0066] For example, after obtaining the teaching quality score and individual student status report, Xiao Zhang completed the multimedia online recording and teaching quality assessment.
[0067] To address the problems described in the background art, this invention identifies the recording and broadcasting teaching environment and the total course duration. The recording and broadcasting teaching environment includes: a teacher, multiple students, a teacher's camera, student cameras, and a smart blackboard. This invention provides a material basis for subsequent video capture of the teacher and students by identifying the recording and broadcasting teaching environment. Identifying the total course duration allows for subsequent division of the total course duration, and multiple students are then numbered, resulting in multiple numbered students. This invention simplifies the subsequent analysis and processing of each student by numbering them. Dividing the total course duration according to preset time intervals yields multiple consecutive time periods. This invention effectively manages the duration of a lesson... The recording is divided into multiple time periods to facilitate timely analysis of teacher teaching quality and student learning quality, thereby improving the teaching quality of multimedia recording and broadcasting. For each consecutive time period, the following operations are performed: based on the consecutive time period, the teacher's camera in the recording and broadcasting environment, and the teacher's confirmed teaching video, the teaching video includes multiple video frames and an audio stream. This embodiment of the invention obtains the teacher's teaching video by filming the teacher, enabling subsequent frame-by-frame analysis and improving the teaching quality of multimedia recording and broadcasting. Teacher status scores are determined based on multiple video frames in the teaching video, and teacher language scores are determined based on the audio stream in the teaching video. This embodiment of the invention improves the teaching quality of multimedia recording and broadcasting by analyzing the video frame by frame. The system analyzes video frames to calculate teacher status scores and analyzes audio streams to calculate teacher language scores. This provides a prerequisite for calculating short-term teaching quality scores based on teacher status and language scores, thus improving the teaching quality of multimedia recorded teaching. The system also uses a smart blackboard to monitor teacher handwriting, obtaining handwriting data including average writing pressure, erasing frequency, stroke count, and writing speed sequence. This embodiment of the invention utilizes a smart blackboard to acquire teacher handwriting data during class, providing a foundation for subsequent processing and calculation. Based on the average writing pressure, erasing frequency, stroke count, writing speed sequence, teacher status score, and teacher language score in the handwriting data, short-term teaching quality scores are determined. The invention, through analysis of blackboard writing data, calculates a blackboard writing quality score reflecting the teacher's blackboard writing quality. This score, combined with teacher status and language scores, is used to calculate a comprehensive short-term teaching quality score, thus improving the teaching quality of multimedia recorded teaching. Based on student cameras and multiple numbered students in the recorded teaching environment, a student status score set is identified. This set includes multiple student status scores, with each score corresponding to a unique student number. The invention utilizes student cameras to capture video of students and then analyzes the captured video frame-by-frame to calculate the student status score set, providing the basis for generating subsequent student status reports and summarizing the short-term teaching quality score.This invention obtains multiple short-term teaching quality scores, summarizes student status score sets, and generates multiple student status score sets. Based on these scores, a teaching quality score and multiple student status reports are determined, thus completing multimedia recorded teaching. Therefore, this invention improves the teaching quality of multimedia recorded teaching by calculating a teaching quality score from multiple short-term scores and generating multiple student status reports from multiple student status score sets.
[0068] like Figure 2 The diagram shown is a functional block diagram of a multimedia recording and teaching device based on artificial intelligence provided in an embodiment of the present invention.
[0069] The AI-based multimedia recording and teaching device 100 of this invention can be installed in an electronic device. Depending on the functions implemented, the AI-based multimedia recording and teaching device 100 may include a basic environment verification module 101, a teacher status monitoring module 102, a student status monitoring module 103, and a teaching quality evaluation module 104. The module described in this invention can also be called a unit, which refers to a series of computer program segments that can be executed by the processor of an electronic device and can perform a fixed function, and are stored in the memory of the electronic device.
[0070] The basic environment confirmation module 101 is used to confirm the recording and broadcasting teaching environment and the total course duration. The recording and broadcasting teaching environment includes: teacher, multiple students, teacher's camera, student's camera and smart blackboard. Multiple students are numbered to obtain multiple numbered students. The total course duration is divided according to a preset time interval to obtain multiple consecutive time periods. The teacher status monitoring module 102 is used to perform the following operations for each of the multiple consecutive time periods: based on the consecutive time period, the teacher's camera in the recording teaching environment, and the teacher's confirmation of the teaching video, wherein the teaching video includes: multiple video frames and an audio stream; based on the multiple video frames in the teaching video, the teacher's status score is confirmed; and based on the audio stream in the teaching video, the teacher's language score is confirmed. The student status monitoring module 103 is used to monitor the teacher's handwriting using a smart blackboard and obtain handwriting data. The handwriting data includes: average writing pressure, number of erasures, number of strokes, and writing speed sequence. Based on the average writing pressure, number of erasures, number of strokes, writing speed sequence, teacher status score, and teacher language score in the handwriting data, a short-term teaching quality score is determined. Based on the student cameras in the recorded teaching environment and multiple numbered students, a student status score set is determined. The student status score set includes: multiple student status scores, and each student status score corresponds one-to-one with a numbered student. The teaching quality assessment module 104 is used to summarize short-term teaching quality scores to obtain multiple short-term teaching quality scores, summarize student status score sets to obtain multiple student status score sets, and confirm the teaching quality score and multiple student status reports based on multiple short-term teaching quality scores and multiple student status score sets, thereby completing multimedia recording and broadcasting teaching.
[0071] In detail, the modules in the AI-based multimedia recording and teaching device 100 described in this embodiment of the invention employ the same methods as described above during use. Figure 1 The method uses the same technical means as the AI-based multimedia recording and teaching method described above, and can produce the same technical effects, so it will not be repeated here.
[0072] like Figure 3 The diagram shown is a structural schematic of an electronic device for implementing a multimedia recording and teaching method based on artificial intelligence, according to an embodiment of the present invention.
[0073] The electronic device 1 may include a processor 10, a memory 11 and a bus 12, and may also include a computer program stored in the memory 11 and capable of running on the processor 10, such as a multimedia recording and teaching method program based on artificial intelligence.
[0074] The memory 11 includes at least one type of readable storage medium, such as flash memory, portable hard drive, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 11 can be an internal storage unit of the electronic device 1, such as the portable hard drive of the electronic device 1. In other embodiments, the memory 11 can be an external storage device of the electronic device 1, such as a plug-in portable hard drive, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the electronic device 1. Furthermore, the memory 11 includes both internal storage units and external storage devices of the electronic device 1. The memory 11 can be used not only to store application software and various types of data installed on the electronic device 1, such as the code of a multimedia recording and teaching method program based on artificial intelligence, but also to temporarily store data that has been output or will be output.
[0075] In some embodiments, the processor 10 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control unit of the electronic device, connecting various components of the entire electronic device through various interfaces and lines. It executes programs or modules stored in the memory 11 (e.g., multimedia recording and teaching methods based on artificial intelligence) and calls data stored in the memory 11 to perform various functions of the electronic device 1 and process data.
[0076] The bus 12 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus 12 can be divided into an address bus, a data bus, a control bus, etc. The bus 12 is configured to realize the connection and communication between the memory 11 and at least one processor 10, etc.
[0077] Figure 3 Only electronic devices with components are shown; it will be understood by those skilled in the art that... Figure 3 The structure shown does not constitute a limitation on the electronic device 1, and may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0078] For example, although not shown, the electronic device 1 may also include a power supply (such as a battery) to power the various components. Preferably, the power supply can be logically connected to the at least one processor 10 through a power management device, thereby enabling functions such as charging management, discharging management, and power consumption management. The power supply may also include one or more DC or AC power supplies, recharging devices, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The electronic device 1 may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be described in detail here.
[0079] Furthermore, the electronic device 1 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a Wi-Fi interface, a Bluetooth interface, etc.), which is typically used to establish communication connections between the electronic device 1 and other electronic devices.
[0080] Optionally, the electronic device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), and optionally, a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen, etc. The display may also be appropriately referred to as a screen or display unit, used to display information processed in the electronic device 1 and to display a visual user interface.
[0081] The multimedia recording and teaching method program based on artificial intelligence stored in the memory 11 of the electronic device 1 is a combination of multiple instructions. When run in the processor 10, it can achieve the following: The recording and teaching environment and total course duration have been confirmed. The recording and teaching environment includes: teacher, multiple students, teacher's camera, student cameras, and smart blackboard. Number multiple students to obtain multiple numbered students; The total course duration is divided into multiple consecutive time periods according to preset time intervals; Perform the following operation for each of the multiple consecutive time periods: Based on the teacher's camera in a continuous time period and the teacher's confirmation of the teaching video in the recorded teaching environment, the teaching video includes: multiple video frames and audio stream; Teacher status scores were determined based on multiple video frames in the instructional video. Teacher language scores were identified based on the audio stream in the instructional videos; The smart blackboard is used to monitor teachers' handwriting and obtain handwriting data, which includes: average writing pressure, number of erasures, number of strokes, and writing speed sequence. Short-term teaching quality scores were determined based on the average writing pressure, number of erasures, number of strokes, writing speed sequence, teacher status score, and teacher language score in the blackboard writing data. Based on the student cameras in the recorded teaching environment and multiple numbered students, a student status score set is identified. The student status score set includes multiple student status scores, and each student status score corresponds one-to-one with a numbered student. By summarizing the short-term teaching quality scores, multiple short-term teaching quality scores are obtained. By summing up the student status rating sets, multiple student status rating sets are obtained; Based on multiple short-term teaching quality scores and multiple student status score sets, the teaching quality scores and multiple student status reports were confirmed, and multimedia recording and broadcasting teaching was completed.
[0082] Specifically, the processor 10's implementation method for the above instructions can be found in [reference needed]. Figures 1 to 3 The descriptions of the relevant steps in the corresponding embodiments are not repeated here.
[0083] Furthermore, if the modules / units integrated in the electronic device 1 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, or a read-only memory (ROM).
[0084] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor of an electronic device, can perform the following: The recording and teaching environment and total course duration have been confirmed. The recording and teaching environment includes: teacher, multiple students, teacher's camera, student cameras, and smart blackboard. Number multiple students to obtain multiple numbered students; The total course duration is divided into multiple consecutive time periods according to preset time intervals; Perform the following operation for each of the multiple consecutive time periods: Based on the teacher's camera in a continuous time period and the teacher's confirmation of the teaching video in the recorded teaching environment, the teaching video includes: multiple video frames and audio stream; Teacher status scores were determined based on multiple video frames in the instructional video. Teacher language scores were identified based on the audio stream in the instructional videos; The smart blackboard is used to monitor teachers' handwriting and obtain handwriting data, which includes: average writing pressure, number of erasures, number of strokes, and writing speed sequence. Short-term teaching quality scores were determined based on the average writing pressure, number of erasures, number of strokes, writing speed sequence, teacher status score, and teacher language score in the blackboard writing data. Based on the student cameras in the recorded teaching environment and multiple numbered students, a student status score set is identified. The student status score set includes multiple student status scores, and each student status score corresponds one-to-one with a numbered student. By summarizing the short-term teaching quality scores, multiple short-term teaching quality scores are obtained. By summing up the student status rating sets, multiple student status rating sets are obtained; Based on multiple short-term teaching quality scores and multiple student status score sets, the teaching quality scores and multiple student status reports were confirmed, and multimedia recording and broadcasting teaching was completed.
[0085] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative, and actual implementations may have other classification methods.
[0086] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0087] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0088] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0089] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. An artificial intelligence-based multimedia recording and teaching method, characterized in that, The method comprises: Confirming a recording and broadcasting teaching environment and a total course length, wherein the recording and broadcasting teaching environment comprises a teacher, a plurality of students, a teacher camera, a student camera and an intelligent blackboard; Numbering the plurality of students to obtain a plurality of numbered students; Dividing the total course length according to a preset time interval to obtain a plurality of continuous time periods; For each of the plurality of continuous time periods, the following operations are performed: Based on the continuous time period, the teacher camera in the recording and broadcasting teaching environment and the teacher, a teaching video is confirmed, wherein the teaching video comprises a plurality of video frames and an audio stream; Based on the plurality of video frames in the teaching video, a teacher state score is confirmed; Based on the audio stream in the teaching video, a teacher language score is confirmed; Using the intelligent blackboard to monitor the board writing of the teacher to obtain board writing data, wherein the board writing data comprises an average writing pressure, an erasing and writing frequency, a stroke number and a writing speed sequence; Based on the average writing pressure, the erasing and writing frequency, the stroke number, the writing speed sequence, the teacher state score and the teacher language score in the board writing data, a short-term teaching quality score is confirmed; Based on the student camera in the recording and broadcasting teaching environment and the plurality of numbered students, a student state score set is confirmed, wherein the student state score set comprises a plurality of student state scores, and the student state score corresponds to the numbered student in a one-to-one manner; The short-term teaching quality scores are summarized to obtain a plurality of short-term teaching quality scores; The student state score sets are summarized to obtain a plurality of student state score sets; Based on the plurality of short-term teaching quality scores and the plurality of student state score sets, a teaching quality score and a plurality of student state reports are confirmed, and multimedia recording and broadcasting teaching is completed.
2. The artificial intelligence-based multimedia recording and teaching method of claim 1, wherein, The method comprises: For each of the plurality of video frames, the following operations are performed: Confirming a video frame resolution, wherein the video frame resolution comprises a pixel length and a pixel width; Performing target detection on the video frame to obtain teacher target data, wherein the teacher target data comprises a teacher target graph, a vertex x coordinate, a vertex y coordinate, a bounding box length and a bounding box width; Calculating a teacher picture proportion according to the pixel length, the pixel width, the bounding box length and the bounding box width; Summarizing the teacher target data to obtain a plurality of teacher target data; Summarizing the teacher picture proportion to obtain a plurality of teacher picture proportions; Based on the plurality of teacher picture proportions, an average proportion is confirmed, wherein the average proportion is the average value of the plurality of teacher picture proportions; Based on the plurality of teacher target data, a teacher activity level and an emotion score are confirmed; According to the average proportion, the teacher activity level and the emotion score, a teacher state score is calculated.
3. The artificial intelligence-based multimedia recording and broadcasting teaching method of claim 2, wherein The method comprises: For each of the plurality of teacher target graphs in the plurality of teacher target data, the following operations are performed: Performing key point detection on the teacher target graph to obtain key point data, wherein the key point data comprises a plurality of key points, wherein the key point comprises a key point x coordinate and a key point y coordinate; Summarizing the key point data to obtain a plurality of key point data; According to the plurality of key point x coordinates and the plurality of key point y coordinates in the plurality of key point data, an action amplitude is calculated; Calculate a teacher total displacement according to a plurality of vertex x coordinates and a plurality of vertex y coordinates in the plurality of teacher target data; Calculate a teacher activity level according to the action amplitude and the teacher total displacement; Confirm an emotion score based on a plurality of teacher target pictures in the plurality of teacher target data.
4. The artificial intelligence-based multimedia recording and broadcasting teaching method of claim 3, wherein, The confirming the emotion score based on the plurality of teacher target pictures in the plurality of teacher target data comprises: The following operations are performed on each of the plurality of teacher target pictures in the plurality of teacher target data: Perform face extraction on the teacher target picture to obtain a teacher face picture; Perform expression analysis on the teacher face picture by using a pre-constructed expression recognition model to obtain a teacher expression result, wherein the teacher expression result is happy, natural or tired; If the teacher expression result is happy or natural, a preset single value is taken as an expression score; Otherwise, a preset zero value is taken as the expression score; Sum up the expression scores to obtain a plurality of expression scores; Calculate the emotion score according to the plurality of expression scores.
5. The artificial intelligence-based multimedia recording and broadcasting teaching method of claim 4, wherein The confirming the teacher language score based on the audio stream in the teaching video comprises: Segment the audio stream by using a preset segmentation interval to obtain a plurality of audio segments; Number the plurality of audio segments to obtain a plurality of numbered audio segments; The following operations are performed on each of the plurality of numbered audio segments: Cut the numbered audio segment according to a preset frame length to obtain a plurality of audio frames, wherein the audio frame comprises a plurality of audio sample values; The following operations are performed on each of the plurality of audio frames: Calculate a single-frame energy according to the plurality of audio sample values in the audio frame; Calculate a single-frame root mean square value according to the single-frame energy; Sum up the single-frame root mean square values to obtain a plurality of single-frame root mean square values; Sum up the single-frame energies to obtain a plurality of single-frame energies; Confirm an average energy based on the plurality of single-frame energies, wherein the average energy is an average value of the plurality of single-frame energies; The following operations are performed on each of the plurality of single-frame energies: Compare the single-frame energy with the average energy, if the single-frame energy is greater than or equal to the average energy, the numbered audio segment corresponding to the single-frame energy is taken as a numbered speech segment; Otherwise, the numbered audio segment corresponding to the single-frame energy is taken as a numbered silence segment; Respectively sum up the numbered speech segments and the numbered silence segments to obtain a plurality of numbered speech segments and a plurality of numbered silence segments; Confirm a class rhythm score and a speech text set based on the plurality of numbered speech segments and the plurality of numbered silence segments; Obtain teaching content, and perform keyword analysis on the teaching content to obtain a keyword set; Confirm a keyword coverage rate based on the speech text set and the keyword set; Confirm a volume fluctuation degree based on the plurality of single-frame root mean square values, wherein the volume fluctuation degree is a standard deviation of the plurality of single-frame root mean square values; Calculate the teacher language score according to the class rhythm score, the keyword coverage rate and the volume fluctuation degree, and the calculation formula is as follows: wherein, represents a teacher language score, represents a lesson pace score, represents a keyword coverage, represents a volume fluctuation, represents a natural logarithm.
6. The artificial intelligence-based multimedia recording and broadcasting teaching method of claim 5, wherein, The confirming the class rhythm score and the speech text set based on the plurality of numbered speech segments and the plurality of numbered silence segments comprises: Multi-element combine the plurality of numbered speech segments to obtain a plurality of target speech segments; The following operations are performed on each of the plurality of target speech segments: Confirm a target speech duration, and perform speech-to-text processing on the target speech segment to obtain a speech text; Count the number of words in the speech text to obtain a total number of words; Calculate a class speech speed according to the total number of words and a target speech time length, wherein the class speech speed is a ratio of the total number of words to the target speech time length; Aggregate the speech text to obtain a speech text set; Aggregate the class speech speeds to obtain a plurality of class speech speeds; Confirm an average speech speed based on the plurality of class speech speeds, wherein the average speech speed is an average of the plurality of class speech speeds; Perform multi-element merging on the plurality of numbered silent segments to obtain a plurality of target silent segments; For each target silent segment in the plurality of target silent segments, perform the following operations: Confirm a target silent segment time length, and compare the target silent segment time length with a preset time length threshold; If the target silent segment time length is greater than or equal to the time length threshold, take the target silent segment corresponding to the target silent segment time length as a real silent segment; Aggregate the real silent segments to obtain a plurality of real silent segments; Confirm a pause frequency based on the plurality of target speech segments, the plurality of target silent segments, and the plurality of real silent segments; Calculate a class rhythm score according to the average speech speed and the pause frequency.
7. The artificial intelligence-based multimedia recording and broadcasting teaching method of claim 6, wherein The confirming the short-term teaching quality score based on the average writing pressure, the erasing times, the number of strokes, the writing speed sequence, the teacher state score and the teacher language score in the board writing data comprises: Confirming an average writing speed based on the writing speed sequence; Calculating the short-term teaching quality score according to the average writing pressure, the erasing times, the number of strokes, the average writing speed, the teacher state score and the teacher language score, and the calculation formula is as follows: wherein, represents a short-term teaching quality score, represents an average writing pressure, represents a number of strokes, represents an average writing speed, represents a number of erasures, is a preset pressure influence coefficient, is a preset writing influence coefficient, is a preset teacher influence coefficient.
8. The artificial intelligence-based multimedia recording and broadcasting teaching method of claim 7, wherein, The confirming the student state score set based on the student camera and the plurality of numbered students in the recording and broadcasting teaching environment comprises: For each numbered student in the plurality of numbered students, perform the following operations: Monitoring the state of the numbered student by using the student camera in the recording and broadcasting teaching environment to obtain a student state video, wherein the student state video comprises a plurality of state video frames; For each state video frame in the plurality of state video frames, perform the following operations: Performing grayscale processing on the state video frame to obtain a grayscale video frame; Performing face detection on the grayscale video frame to obtain a student face image; Performing illumination normalization processing on the student face image to obtain a target image; Performing expression recognition on the target image by using an expression recognition model to obtain a recognition result, wherein the recognition result is a smile, a frown, a drowsiness or a concentration; Confirming a gaze direction vector based on a pre-constructed pupil positioning algorithm and the target image; Performing key point detection on the target image to obtain student key point data; Aggregating the recognition results to obtain a plurality of recognition results; Aggregating the gaze direction vectors to obtain a plurality of gaze direction vectors; Aggregating the student key point data to obtain a plurality of student key point data; Statistically merging the plurality of recognition results to obtain a merged result, wherein the merged result comprises a total result number, a smile frequency and a concentration frequency; Confirming a student concentration degree based on the total result number, the smile frequency and the concentration frequency, wherein the student concentration degree is a ratio of a sum of the smile frequency and the concentration frequency to the total result number; Obtaining an intelligent blackboard space vector; Confirming an average gaze vector based on the plurality of gaze direction vectors; Calculating a student gaze degree according to the intelligent blackboard space vector and the average gaze vector; Confirming a student action amplitude based on a plurality of student key point data; Calculating a student state score according to a student concentration, a student gaze and the student action amplitude; Summarizing the student state scores to obtain a student state score set.
9. The artificial intelligence-based multimedia recording and broadcasting teaching method of claim 8, wherein, The confirming of the teaching quality score and the student state reports based on the plurality of short-term teaching quality scores and the student state score sets comprises: Confirming the teaching quality score based on the plurality of short-term teaching quality scores, wherein the teaching quality score is an average of the plurality of short-term teaching quality scores; Statistically analyzing the plurality of student state score sets to obtain a plurality of student comprehensive scores; Confirming the plurality of student state reports based on the plurality of student comprehensive scores.
10. An artificial intelligence-based multimedia recording and teaching device, characterized in that, The device comprises: A basic environment confirming module configured to confirm a recording and broadcasting teaching environment and a total course duration, wherein the recording and broadcasting teaching environment comprises a teacher, a plurality of students, a teacher camera, a student camera and an intelligent blackboard, the plurality of students are numbered to obtain a plurality of numbered students, and the total course duration is divided according to a preset time interval to obtain a plurality of continuous time periods; A teacher state monitoring module configured to confirm a teaching video based on a continuous time period, the teacher camera and the teacher in the recording and broadcasting teaching environment for each of the plurality of continuous time periods, wherein the teaching video comprises a plurality of video frames and an audio stream, confirm a teacher state score based on the plurality of video frames in the teaching video, and confirm a teacher language score based on the audio stream in the teaching video; A student state monitoring module configured to monitor a blackboard writing of the teacher by using the intelligent blackboard to obtain blackboard writing data, wherein the blackboard writing data comprises an average writing pressure, an erasing and writing times, a stroke number and a writing speed sequence, confirm a short-term teaching quality score based on the average writing pressure, the erasing and writing times, the stroke number, the writing speed sequence, the teacher state score and the teacher language score in the blackboard writing data, and confirm a student state score set based on the student camera and the plurality of numbered students in the recording and broadcasting teaching environment, wherein the student state score set comprises a plurality of student state scores, and the student state scores correspond to the numbered students one by one; A teaching quality evaluation module configured to summarize the short-term teaching quality scores to obtain a plurality of short-term teaching quality scores, summarize the student state score sets to obtain a plurality of student state score sets, confirm a teaching quality score and a plurality of student state reports based on the plurality of short-term teaching quality scores and the student state score sets, and complete the multimedia recording and broadcasting teaching.