Method and terminal device for detecting learning concentration in online courses
By detecting the audio data and video data of online teaching and matching teaching instructions and action categories, the problem of low accuracy in learning concentration detection in the existing technology is solved, and more efficient learning concentration monitoring and teaching quality improvement is achieved.
Patent Information
- Application Number
- CN202210155293.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-21
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-02-21
AI Technical Summary
The learning concentration detection method used for online teaching in the prior art is relatively low in accuracy and cannot effectively monitor students' learning concentration, resulting in a decline in teaching quality.
By detecting teaching instructions in the audio data of online teaching, and obtaining the action categories to be matched during the validity period of the teaching instructions, combining the action detection results of the target object in the video data, a matching operation is performed to determine the learning concentration.
This improves the accuracy of the learning concentration test results, allowing teachers to more accurately monitor students' learning concentration, thereby improving the quality of teaching.
Smart Images

Figure CN114529872B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent image processing technology, and in particular to a method and terminal device for detecting learning concentration in online teaching. Background Art
[0002] Online teaching is a particularly important teaching method under the new situation. During online teaching, students usually open application software through terminal devices such as tablets, connect to the teacher's computer, and receive education through video. However, many problems are exposed in current online teaching, such as teachers cannot monitor students' concentration like offline classes, resulting in a decline in teaching quality. Therefore, consumers, especially parents of students, very much hope that when students use terminal devices to study online courses, they can improve their concentration through effective methods to ensure that students can focus on learning.
[0003] In the related art, although some terminal devices have concentration detection technology, the method used by these terminal devices for classroom concentration detection is to detect the line of sight of the human eye or the direction of the human head to determine whether the terminal user is focused on learning. These detection methods are too rough and the detection results of learning concentration are less accurate. Summary of the invention
[0004] The purpose of this application is to provide a method and terminal device for detecting learning concentration in online teaching, so as to solve the problem of low accuracy of the detection results of learning concentration in related technologies.
[0005] In a first aspect, the present application provides a method for detecting learning concentration in online teaching, the method comprising:
[0006] Detecting teaching instructions in audio data of online teaching, and obtaining a to-be-matched action category within the validity period of the teaching instructions; the to-be-matched action category is obtained by performing action detection on a target object in video data;
[0007] The action category to be matched is matched with the teaching instruction to obtain the learning concentration of the target object.
[0008] In a possible implementation, the performing action detection on the target object in the video data includes:
[0009] Taking every n consecutive frames of images in the video data as an image sequence, where n is a positive integer;
[0010] For each image sequence, execute:
[0011] Acquire key point data of the target object in each frame of the image sequence to obtain a key point sequence formed by the key point data of each frame of the image; the key point data includes face key points and body key points;
[0012] The key point sequence is input into the action recognition deep learning model to obtain the action category of the target object.
[0013] In a possible implementation manner, matching the action category to be matched with the teaching instruction to obtain the learning concentration of the target object includes:
[0014] Searching in a pre-stored correspondence table whether the action category to be matched and the teaching instruction have a correspondence;
[0015] If there is a corresponding relationship, determining the learning concentration of the target object to be a first value;
[0016] If there is no corresponding relationship, the learning concentration of the target object is determined to be a second value.
[0017] In a possible implementation manner, obtaining the action category to be matched within the validity period of the teaching instruction includes:
[0018] On the premise that the teaching instruction is the latest teaching instruction, the detected action category of the target object is determined as the action category to be matched; or,
[0019] Record the first detection time of the teaching instruction and the second detection time of each action category. If the second detection time of the action category is within the specified time interval, determine that the action category is the action category to be matched, the start time of the specified time interval is the first detection time of the teaching instruction, and the end time of the specified time interval is the first detection time of the next teaching instruction.
[0020] In a possible implementation, the method further includes:
[0021] determining a second detection time of a first action category matching the teaching instruction;
[0022] determining a time difference between a second detection time of the first action category and a first detection time of the teaching instruction;
[0023] The preset action delay time is divided by the time difference and then multiplied by the first value to obtain the updated first value, wherein the preset action delay time is used to represent the time required from issuing the teaching instruction to executing the teaching instruction.
[0024] In a possible implementation, the method further includes:
[0025] The duration of the validity period of the teaching instruction is used as the effective duration, and the duration of the action category to be matched that matches the teaching instruction is determined;
[0026] The continuation duration is divided by the effective duration and then multiplied by the updated first value to obtain a final value of the first value of the target object.
[0027] In a possible implementation, the detecting teaching instructions in the audio data of the online teaching includes:
[0028] Performing audio recognition on the audio data to obtain a text sequence;
[0029] The text sequence is detected based on preset teaching instruction keywords to obtain teaching instructions.
[0030] In a possible implementation, the acquiring key point data of the target object in each frame of the image sequence includes:
[0031] Based on a human key point detection model, obtaining human key points of each frame of the image sequence;
[0032] Based on the facial key point detection model, the facial key points of each frame image in the image sequence are obtained.
[0033] In a possible implementation, the method further includes:
[0034] For multiple frames of similar images in the image sequence whose similarity is higher than a similarity threshold, the human body key point detection model and the facial key point detection model are used to process one frame of the multiple frames of similar images to obtain the human body key points and facial key points corresponding to the multiple frames of similar images respectively.
[0035] In a second aspect, the present application provides a terminal device, including:
[0036] Display, processor and memory;
[0037] The display is used to display the screen display area;
[0038] The memory is used to store the processor executable instructions;
[0039] The processor is configured to execute the instructions to implement a method for detecting learning concentration in online courses as described in any one of the first aspects above.
[0040] In a third aspect, the present application provides a computer-readable storage medium. When the instructions in the computer-readable storage medium are executed by a terminal device, the terminal device can execute the method for detecting learning concentration in online courses as described in any one of the first aspects above.
[0041] In a fourth aspect, the present application provides a computer program product, including a computer program:
[0042] When the computer program is executed by a processor, the method for detecting learning concentration in online courses as described in any one of the first aspects above is implemented.
[0043] The technical solution provided by the embodiments of the present application brings at least the following beneficial effects:
[0044] The embodiment of the present application detects the teaching instructions in the audio data of the online teaching, and obtains the action category to be matched within the validity period of the teaching instructions; the action category to be matched is obtained by detecting the action of the target object in the video data; the action category to be matched is matched with the teaching instructions to obtain the learning concentration of the target object. Thus, by comparing the degree of conformity between the teacher's teaching instructions and the action category of the target object within the validity period of the teaching instructions, the learning concentration of the target object is determined according to the degree of conformity, which can improve the accuracy of the detection result of the learning concentration of the target object, so that the teacher can more accurately monitor the learning concentration of the target object and improve the teaching quality.
[0045] Other features and advantages of the present application will be described in the following description, and partly become apparent from the description, or be understood by practicing the present application. The purpose and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. Obviously, the drawings introduced below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0047] Figure 1 A schematic diagram of the structure of a terminal device provided in an embodiment of the present application;
[0048] Figure 2 A software structure diagram of a terminal device provided in an embodiment of the present application;
[0049] Figure 3A flow chart of a method for detecting learning concentration in online courses provided in an embodiment of the present application;
[0050] Figure 4 A flowchart of a method for detecting teaching instructions in audio data of online teaching provided in an embodiment of the present application;
[0051] Figure 5 A schematic diagram of a flow chart of a method for detecting an action category of a target object in an image sequence provided in an embodiment of the present application;
[0052] Figure 6 A schematic diagram of a process for obtaining a key point sequence of a target object in an image sequence provided by an embodiment of the present application;
[0053] Figure 7 A schematic diagram of key points in a human key point detection model provided in an embodiment of the present application;
[0054] Figure 8 A schematic diagram of key points in a facial key point detection model provided in an embodiment of the present application;
[0055] Fig. 9 A schematic diagram of a process of identifying an action category of a target object in an image sequence provided by an embodiment of the present application;
[0056] Fig.10 A schematic diagram of determining an action category and a teaching instruction that match the detection time provided by an embodiment of the present application;
[0057] Fig.11 A schematic diagram of another method for determining action categories and teaching instructions that match detection time provided in an embodiment of the present application;
[0058] Fig.12 A flowchart of another method for detecting learning concentration in online courses provided in an embodiment of the present application. DETAILED DESCRIPTION
[0059] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Among them, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0060] Furthermore, in the description of the embodiments of the present application, unless otherwise specified, “ / ” means or. For example, A / B can mean A or B. The “and / or” in the text is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, “multiple” refers to two or more than two.
[0061] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood as suggesting or implying relative importance or implicitly indicating the number of technical features indicated. Thus, features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more.
[0062] Online teaching is a particularly important teaching method under the new situation. During online teaching, students usually open application software through terminal devices such as tablets, connect to the teacher's computer, and receive education through video. However, many problems are exposed in current online teaching, such as teachers cannot monitor students' concentration like offline classes, resulting in a decline in teaching quality. Therefore, consumers, especially parents of students, very much hope that when students use terminal devices to study online courses, they can improve their concentration through effective methods to ensure that students can focus on learning.
[0063] In the related technologies, some terminal devices do not have concentration detection technology, while some terminal devices have concentration detection technology, but the method used by these terminal devices for classroom concentration detection is to detect the line of sight of the human eye or the direction of the human head to determine whether the terminal user is focused on learning. These detection methods are too rough and ignore the fact that the terminal user is not simply watching the video during the learning process, but also needs to perform dictation, follow-up reading, action imitation and other operations, resulting in low accuracy in the detection results of user learning concentration in dictation, follow-up reading, action imitation and other operations.
[0064] In view of this, the present application provides a method and terminal device for detecting learning concentration in online courses, so as to solve the problem of low accuracy of detection results of user learning concentration in related technologies.
[0065] The inventive concept of the present application can be summarized as follows: the present application embodiment detects the teaching instructions in the audio data of the online teaching, and obtains the action category to be matched within the validity period of the teaching instructions; the action category to be matched is obtained by detecting the action of the target object in the video data; the action category to be matched is matched with the teaching instructions to obtain the learning concentration of the target object. Therefore, by comparing the degree of conformity between the teacher's teaching instructions and the action category of the target object within the validity period of the teaching instructions, the learning concentration of the target object is determined according to the degree of conformity, which can improve the accuracy of the detection result of the learning concentration of the target object, so that the teacher can more accurately monitor the learning concentration of the target object and improve the teaching quality.
[0066] After introducing the inventive concept of the present application, the terminal device provided by the present application will be described below. Figure 1 FIG. 1 shows a schematic diagram of the structure of a terminal device 100. It should be understood that Figure 1 The terminal device 100 shown is only an example, and the terminal device 100 may have more Figure 1 The more or less components shown in the figure can be combined with two or more components, or can have different component configurations. The various components shown in the figure can be implemented in hardware, software, or a combination of hardware and software including one or more signal processing and / or application specific integrated circuits.
[0067] Figure 1 FIG. 1 is a block diagram showing a hardware configuration of a terminal device 100 according to an exemplary embodiment. Figure 1 As shown, the terminal device 100 includes: a radio frequency (RF) circuit 110, a memory 120, a display unit 130, a camera 140, a sensor 150, an audio circuit 160, a wireless fidelity (Wi-Fi) module 170, a processor 180, a Bluetooth module 181, and a power supply 190 and other components.
[0068] The RF circuit 110 can be used to receive and send signals during the process of sending and receiving information or making calls. It can receive downlink data from the base station and hand it over to the processor 180 for processing; it can send uplink data to the base station. Generally, the RF circuit includes but is not limited to antennas, at least one amplifier, a transceiver, a coupler, a low noise amplifier, a duplexer and other devices.
[0069] The memory 120 can be used to store software programs and data. The processor 180 executes various functions and data processing of the terminal device 100 by running the software programs or data stored in the memory 120. The memory 120 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. The memory 120 stores an operating system that enables the terminal device 100 to run. In the present application, the memory 120 can store an operating system and various application programs, and can also store program codes for executing the method for detecting the learning concentration of online teaching described in the embodiment of the present application.
[0070] The display unit 130 may be used to receive input digital or character information and generate signal input related to user settings and function control of the terminal device 100. Specifically, the display unit 130 may include a touch screen 131 disposed on the front of the terminal device 100, which may collect user touch operations thereon or near the touch screen, such as clicking a button.
[0071] The display unit 130 can also be used to display information input by the user or information provided to the user and a graphical user interface (GUI) of various menus of the terminal device 100. Specifically, the display unit 130 may include a display screen 132 disposed on the front of the terminal device 100. The display screen 132 may be configured in the form of a liquid crystal display, a light emitting diode, etc. The display unit 130 may be used to display the screen display area of the terminal device in the present application.
[0072] The touch screen 131 may be covered on the display screen 132, or the touch screen 131 and the display screen 132 may be integrated to realize the input and output functions of the terminal device 100, and the integrated touch screen may be referred to as a touch display screen. In the present application, the display unit 130 may display the application and the corresponding operation steps.
[0073] The camera 140 can be used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, which is then passed to the processor 180 for conversion into a digital image signal.
[0074] The terminal device 100 may further include at least one sensor 150, such as an acceleration sensor 151, a distance sensor 152, a fingerprint sensor 153, and a temperature sensor 154. The terminal device 100 may also be configured with other sensors such as a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, a light sensor, and a motion sensor.
[0075] The audio circuit 160, the speaker 161, and the microphone 162 can provide an audio interface between the user and the terminal device 100. The audio circuit 160 can transmit the electrical signal converted from the received audio data to the speaker 161, which is converted into a sound signal for output. The terminal device 100 can also be configured with a volume button for adjusting the volume of the sound signal, and can also be used to combine other buttons to adjust the closed area. On the other hand, the microphone 162 converts the collected sound signal into an electrical signal, which is received by the audio circuit 160 and converted into audio data, and then the audio data is output to the RF circuit 110 to be sent to, for example, another terminal device, or the audio data is output to the memory 120 for further processing.
[0076] Wi-Fi is a short-range wireless transmission technology. The terminal device 100 can help users send and receive emails, browse web pages, and access streaming media through the Wi-Fi module 170, which provides users with wireless broadband Internet access.
[0077] The processor 180 is the control center of the terminal device 100. It uses various interfaces and lines to connect various parts of the entire terminal device. It executes various functions of the terminal device 100 and processes data by running or executing software programs stored in the memory 120 and calling data stored in the memory 120. In some embodiments, the processor 180 may include one or more processing units; the processor 180 may also integrate an application processor and a baseband processor, wherein the application processor mainly processes the operating system, the user interface, and the application program, and the baseband processor mainly processes wireless communication. It is understandable that the above-mentioned baseband processor may not be integrated into the processor 180. In the present application, the processor 180 can run the operating system, the application program, the user interface display and the touch response, as well as the detection method for the learning concentration of the online teaching described in the embodiment of the present application. In addition, the processor 180 is coupled to the display unit 130.
[0078] The Bluetooth module 181 is used to exchange information with other Bluetooth devices having Bluetooth modules through the Bluetooth protocol. For example, the terminal device 100 can establish a Bluetooth connection with a wearable electronic device (such as a smart watch) that also has a Bluetooth module through the Bluetooth module 181 to exchange data.
[0079] The terminal device 100 also includes a power supply 190 (such as a battery) for supplying power to various components. The power supply can be logically connected to the processor 180 through a power management system, so that the power management system can manage functions such as charging, discharging, and power consumption. The terminal device 100 can also be configured with a power button for powering on and off the terminal device, as well as locking the screen and other functions.
[0080] Figure 2 It is a software structure block diagram of the terminal device 100 according to an embodiment of the present application.
[0081] The layered architecture divides the software into several layers, each with clear roles and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system can be divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system library, and the kernel layer.
[0082] The application layer can include a series of application packages.
[0083] like Figure 2 As shown, the application package may include Tencent Classroom, DingTalk, Xueersi, Gaotu, Camera, Video, Gallery, Calendar, WLAN, Bluetooth and other applications.
[0084] The application framework layer provides an application programming interface (API) and a programming framework for the applications in the application layer. The application framework layer includes some predefined functions.
[0085] like Figure 2 As shown, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, and the like.
[0086] The window manager is used to manage window programs. The window manager can obtain the display screen size, determine whether there is a status bar, lock the screen, capture the screen, etc.
[0087] Content providers are used to store and retrieve data and make it accessible to applications. The data may include videos, images, audio, calls made and received, browsing history and bookmarks, phone books, short messages, etc.
[0088] The view system includes visual controls, such as controls for displaying text, controls for displaying images, etc. The view system can be used to build applications. A display interface can be composed of one or more views. For example, a display interface including a short message notification icon can include a view for displaying text and a view for displaying images.
[0089] The phone manager is used to provide communication functions of the terminal device 100, such as management of call status (including connection, disconnection, etc.).
[0090] The resource manager provides various resources for applications, such as localized strings, icons, images, layout files, video files, etc.
[0091] The notification manager enables applications to display notification information (such as the content of short messages) in the status bar. It can be used to convey notification-type messages and can automatically disappear after a short stay without user interaction. For example, the notification manager is used to notify download completion, message reminders, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or scroll bar text, such as notifications of applications running in the background, or a notification that appears on the screen in the form of a dialog window. For example, a text message is prompted in the status bar, a prompt sound is emitted, the terminal device vibrates, the indicator light flashes, etc.
[0092] Android Runtime includes core libraries and virtual machines. Android runtime is responsible for scheduling and management of the Android system.
[0093] The core library consists of two parts: one part is the function that needs to be called by the Java language, and the other part is the Android core library.
[0094] The application layer and the application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as object life cycle management, stack management, thread management, security and exception management, and garbage collection.
[0095] The system library may include multiple functional modules, such as surface manager, media library, 3D graphics processing library (such as OpenGL ES), 2D graphics engine (such as SGL), etc.
[0096] The surface manager is used to manage the display subsystem and provide the fusion of 2D and 3D layers for multiple applications.
[0097] The media library supports playback and recording of a variety of commonly used audio and video formats, as well as static image files, etc. The media library can support a variety of audio and video encoding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.
[0098] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0099] A 2D (animation mode) graphics engine is a drawing engine for 2D drawing.
[0100] The kernel layer is the layer between hardware and software. The kernel layer contains at least display driver, camera driver, audio driver, and sensor driver.
[0101] The terminal 100 in the embodiment of the present application may be an electronic device including but not limited to a smart phone, a tablet computer, a wearable electronic device (such as a smart watch), a laptop computer, etc.
[0102] In order to facilitate understanding of the method for detecting learning concentration in online courses provided in the embodiment of the present application, it is further explained below with reference to the accompanying drawings.
[0103] Figure 3 The flowchart of the method for detecting the learning concentration of online teaching provided in the embodiment of the present application is as follows. Figure 3 As shown, the method comprises the following steps:
[0104] In step 301, the teaching instructions in the audio data of the online teaching are detected, and the action categories to be matched within the validity period of the teaching instructions are obtained; wherein the action categories to be matched are obtained by performing action detection on the target object in the video data.
[0105] In a possible implementation, detecting teaching instructions in audio data of online teaching in the embodiment of the present application can be implemented as follows: performing audio recognition on the audio data to obtain a text sequence; and detecting the text sequence based on preset teaching instruction keywords to obtain the teaching instructions.
[0106] Exemplarily, a flow chart of a method for detecting teaching instructions in audio data of online teaching is as follows: Figure 4 As shown, the following steps are included:
[0107] In step 401, audio data from the video stream of the online course is extracted and input into a speech recognition module.
[0108] In step 402, feature extraction is performed on the audio data, and the feature-extracted audio data is converted into a text sequence.
[0109] The speech recognition module can use the open source Kaldi HMM-GMM (Kaldi-Hidden MarkovModel-Gaussian Mixture Model) model to extract features from the audio data, such as extracting speech feature parameters mfcc (Mel Frequency Cepstrum Coefficient). The speech recognition module can use the GMM model (Gaussian Mixture Model) to convert the feature-extracted audio data into a text sequence.
[0110] In step 403, the teaching instruction is found based on the preset teaching instruction keywords.
[0111] First, the text sequence output in step 402 is input into the teaching instruction analysis module, and then the teaching instruction analysis module finds the teaching instruction from the text sequence output by the speech recognition module based on the preset teaching instruction keywords.
[0112] Among them, the preset teaching instructions include vocabulary commonly used by teachers in teaching, such as: "Please read with me", "Please recite aloud", "Please look", "Please look at the screen", "Dictation test", "Do eye exercises", etc.
[0113] If any preset teaching instruction keywords are not detected in the text sequence, it is determined that the default teaching instruction is detected. For example, "Please read" is set as the default teaching instruction. When any preset teaching instruction keywords are not detected, it is considered that the teacher's teaching instruction in the audio data of the online teaching is "Please read". In this way, the teacher's teaching instruction in the audio data of the online teaching can be accurately obtained according to the listening state.
[0114] In a possible implementation, in order to perform action detection on a target object in the video data based on the video data of the target object, in the embodiment of the present application, each consecutive n-frame image in the video data can be regarded as an image sequence, where n is a positive integer, and the following steps are performed for each image sequence respectively. Figure 5 Steps shown:
[0115] In step 501, key point data of a target object in each frame of an image sequence is obtained to obtain a key point sequence of the target object formed by the key point data of each frame of the image; the key point data includes face key points and body key points.
[0116] In a possible implementation, key point data of the target object in each frame of the image sequence is obtained, which can be specifically implemented as follows: based on a human key point detection model, the human key points of each frame in the image sequence are obtained; based on a face key point detection model, the face key points of each frame in the image sequence are obtained.
[0117] Exemplarily, every 10 consecutive frames of images in a segment of video data can be regarded as an image sequence. For example, the 1st to 10th frames of images are first regarded as the first image sequence, the 2nd to 11th frames of images are regarded as the second image sequence, the 3rd to 12th frames of images are regarded as the third image sequence, and so on to obtain multiple image sequences.
[0118] In a possible implementation, obtaining a key point sequence of a target object in an image sequence can be implemented as follows: Figure 6 Steps shown:
[0119] In step 601, each frame of an image in an image sequence is acquired;
[0120] In step 602, each frame of the image in the image sequence is input into a human key point detection model and a face key point detection model;
[0121] In step 603, the human body key point detection model and the face key point detection model output human body key points and face key points respectively, and the human body key points and face key points of all images in the image sequence constitute a key point sequence.
[0122] like Figure 7 As shown in the figure, there are 25 human key points output by the human key point detection model, and the corresponding human positions are: (0, "nose"), (1, "neck"), (2, "right shoulder"), (3, "right elbow"), (4, "right wrist"), (5, "left shoulder"), (6, "left elbow"), (7, "left wrist"), (8, "mid-hip"), (9, "right hip"), (10, "right knee"), (11, "right ankle"), (12, "left hip"), (13, "left knee"), (14, "left ankle"), (15, "right eye"), (16, "left eye"), (17, "right ear"), (18, "left ear"), (19, "left big toe"), (20, "left little toe"), (21, "left heel"), (22, "right big toe"), (23, "right little toe"), (24, "right heel"). Figure 8 As shown in Figure 1, there are 98 facial key points output by the facial key point detection model, corresponding to various parts of the face. Figure 7 The key points of the human body shown in Figure 8After the facial key points shown in the figure, the human body key points and the facial key points can be combined to obtain the key point data of a frame of image, such as {human body key points 0, 1...24, facial key points 0, 1...97}, and then the key point data of all images in the image sequence are combined to obtain a key point sequence, such as {human body key points 0, 1...24 of the first frame image, facial key points 0, 1...97 of the first frame image, human body key points 0, 1...24 of the second frame image, facial key points 0, 1...97, ...} as a key point sequence.
[0123] Of course, the number of human key points detected by the human key point detection model in this application can also be more or less, for example, only 18 human key points are detected or only the human key points at and above the waist are detected, as long as the action category of the target object can be identified. The number of facial key points detected by the face key point detection model in this application can also be respectively from Figure 8 A part of the facial key points of the 98 facial key points shown are extracted from the eyebrows, eyes, nose, and mouth to form a facial key point set. More detailed facial key points can also be added to achieve more facial key points. It can be set according to actual conditions, and the embodiment of the present application is not limited.
[0124] Among them, the human key point detection models applicable to the embodiments of the present application include but are not limited to CPM (Convolutional Pose Machines), OpenPose (human posture recognition model), HourGlass (human posture estimation model), etc., and the facial key point detection models applicable to the embodiments of the present application include but are not limited to DAN (Deep Alignment Network, face alignment model), PFLD (Practical Facial Landmark Detector, face detection model) and DLIB (face detection) open source library.
[0125] Among them, the embodiments of the present application do not limit the order of obtaining the human body key points and obtaining the facial key points of each frame in the image sequence. They can be performed simultaneously, or the human body key points of each frame in the image sequence can be obtained first and then the facial key points of each frame in the image sequence. The facial key points of each frame in the image sequence can also be obtained first and then the human body key points of each frame in the image sequence.
[0126] Therefore, based on the human key point detection model and the face key point detection model, the human key points and face key points of the target object in each frame of the image sequence can be obtained, and then the human key points and face key points of each frame are combined to obtain the key point sequence of the target object. Then, in step 502, the key point sequence is input into the action recognition deep learning model to obtain the action category of the target object.
[0127] Exemplarily, the action category recognition process is as follows Fig. 9 As shown, after acquiring the key point data of the target object in each frame of the image sequence, a key point sequence of the target object formed by the key point data of each frame is obtained, and then the key point sequence of the target object is input into the action recognition deep learning model. After calculation by the action recognition deep learning model, the action category of the target object in the image sequence is finally obtained.
[0128] Among them, in the embodiment of the present application, the training action recognition deep learning model can collect learning-related videos of multiple target objects in advance, and then manually label and classify the action categories in the video clips according to the action classification, such as labeling the video clips into corresponding action categories such as "reading aloud" and "writing", and then converting the labeled video files into key point sequences. The key point sequence is used as a training sample for training the action recognition deep learning model, input into the action recognition deep learning model, and the action category of the output target object is obtained. It is compared with the labeled action category to obtain the loss function, and then the model parameters are adjusted based on the loss function. After multiple trainings, a stable action recognition deep learning model is finally obtained.
[0129] Among them, the action recognition deep learning models applicable to the embodiments of the present application include but are not limited to st-GCN (SpatialTemporal-Graph Convolution Networks), AS-gcn (Actional-Structural Graph Convolutional Networks), 2s-AGCN (Two-stream-adaptive graph convolutional networks), etc.
[0130] Therefore, based on the video data of the target object, the action category of the target object can be obtained by performing action detection on the target object in the video data.
[0131] Among them, in the embodiment of the present application, the execution order of detecting the teacher's teaching instructions in the audio data of the online teaching and performing action detection on the target object in the video data based on the video data of the target object is not limited.
[0132] In a possible implementation, after performing action detection on the target object in the video data based on the video data of the target object to obtain the action category of the target object, it is necessary to obtain the action category to be matched within the validity period of the teaching instruction. Therefore, in an embodiment of the present application, on the premise that the teaching instruction is the latest teaching instruction, the action category of the detected target object can be determined as the action category to be matched; or the first detection time of the teaching instruction and the second detection time of each action category can be recorded. If the second detection time of the action category is within a specified time interval, the action category is determined to be the action category to be matched, wherein the start time of the specified time interval is the first detection time of the teaching instruction, and the end time of the specified time interval is the first detection time of the next teaching instruction.
[0133] For example, Fig.10 As shown, teaching instruction 1 is detected at time A, teaching instruction 2 is detected at time B, and action category 1 is detected at time C. Time A is earlier than time B, and time B is earlier than time C. Therefore, teaching instruction 2 is determined to be the latest teaching instruction, and action category 1 is determined to be the action category to be matched for teaching instruction 2.
[0134] For example, Fig.11 As shown, the first detection time of teaching instruction 1 is moment A, the second detection time of action category 1 is moment B, and the first detection time of teaching instruction 2 is moment C. Moment A is earlier than moment B, and moment B is earlier than moment C. If the second detection time of action category 1 is detected between the first detection time of teaching instruction 1 and the first detection time of teaching instruction 2, then action category 1 is determined to be the action category to be matched for teaching instruction 1.
[0135] In step 302, the action category to be matched is matched with the teaching instruction to obtain the learning concentration of the target object.
[0136] In a possible implementation, in the embodiment of the present application, the action category to be matched is matched with the teaching instruction to obtain the learning concentration of the target object. Specifically, it can be implemented as follows: searching in a pre-stored correspondence table whether the action category to be matched and the teaching instruction have a correspondence; if there is a correspondence, determining the learning concentration of the target object to be a first value; if there is no correspondence, determining the learning concentration of the target object to be a second value.
[0137] The pre-stored corresponding relationship table is shown in Table 1:
[0138] Table 1
[0139]
[0140] Exemplarily, for example, check in Table 1 whether there is a corresponding relationship between the action category to be matched and the teaching instruction. If there is a corresponding relationship, the learning concentration of the target object is set to 1; if there is no corresponding relationship, the learning concentration of the target object is set to 0.
[0141] If the action category to be matched is detected as "gaze" and the teaching instruction is "please read with me", and the action category to be matched and the teaching instruction are not found to have a corresponding relationship in Table 1, then the learning concentration of the target object is determined to be 0; if the action category to be matched is detected as "gaze" and the teaching instruction is "please look", and the action category to be matched and the teaching instruction are found to have a corresponding relationship in Table 1, then the learning concentration of the target object is determined to be 1.
[0142] In a possible implementation, in order to further improve the accuracy of the detection results of the target object's learning concentration, in addition to considering whether the teaching instructions and the action categories to be matched have a corresponding relationship, the embodiment of the present application also considers the timeliness of the action categories to be matched that match the teaching instructions. Specifically, it can be implemented as follows: determining the second detection time of the first action category to be matched that matches the teaching instructions; and determining the time difference between the second detection time of the first action category to be matched and the first detection time of the teaching instructions; dividing the preset action delay time by the time difference and then multiplying it by the first value to obtain an updated first value.
[0143] The preset action delay time is used to indicate the time required from issuing a teaching instruction to executing the teaching instruction, which is an empirical value. For example, if the teaching instruction of "write" is detected at 9:15:00, and the preset action delay time is 30 seconds, the latest time allowed to execute the teaching instruction is 9:15:30, that is, the latest time allowed to detect the action category of "write" of the target object is 9:15:30.
[0144] For example, if the learning concentration of the target object has been detected to be the first value, and the first detection time of the teaching instruction is Tx, and the second detection time of the first action category to be matched that matches the teaching instruction is Tx', then the time difference between the second detection time of the first action category to be matched and the first detection time of the teaching instruction is ΔTx=Tx'-Tx, and the preset action delay time is T delay , then the updated learning concentration of the target object is (T delay / ΔTx)×first value. For example, if the learning concentration of the target object has been detected to be 1, the first detection time of the teaching instruction is the 8th second, and the second detection time of the first action category to be matched that matches the teaching instruction is the 2nd second, then the time difference between the second detection time of the first action category to be matched and the first detection time of the teaching instruction is 6 seconds, and the preset action delay time is 5 seconds, then the updated learning concentration of the target object is (5 / 6)×1, and the updated learning concentration of the target object is 5 / 6.
[0145] In one possible implementation, in order to further improve the accuracy of the detection results of the target object's learning concentration, in addition to considering whether the teaching instructions and the action categories to be matched have a corresponding relationship and the timeliness of the action categories to be matched that match the teaching instructions, the embodiments of the present application may also consider the proportion of the continuation duration of the action categories to be matched that match the teaching instructions in the effective duration of the teaching instructions. Specifically, it can be implemented as follows: taking the duration of the validity period of the teaching instructions as the effective duration, and determining the continuation duration of the action categories to be matched that match the teaching instructions; and dividing the continuation duration by the effective duration and then multiplying it by the updated first value to obtain the final value of the first value of the target object.
[0146] The effective duration of a teaching instruction refers to the time between issuing a teaching instruction and issuing the next teaching instruction. For example, if the teacher issues the teaching instruction "write" at 9:15 and then issues the next teaching instruction "read" at 9:30, the effective duration of the teaching instruction "write" is 15 minutes.
[0147] Exemplarily, if the learning concentration of the target object has been detected, and the effective duration of the teaching instruction is T1, and the duration of the action category to be matched that matches the teaching instruction is T1', then the duration of the action category to be matched that matches the teaching instruction accounts for T1' / T1 in the effective duration of the teaching instruction, and the updated learning concentration of the target object is (T1' / T1)×learning concentration. For example, after considering the timeliness of the action category to be matched that matches the teaching instruction, the learning concentration of the target object has been detected to be 5 / 6, and the effective duration of the teaching instruction is 5 minutes, and the duration of the action category to be matched that matches the teaching instruction is 3 minutes, then the final learning concentration of the target object is (3 / 5)×(5 / 6), and the final learning concentration of the target object is 0.5.
[0148] Therefore, by considering various factors such as whether there is a corresponding relationship between the teaching instructions and the action categories to be matched, the timeliness of the action categories to be matched that match the teaching instructions, and the proportion of the duration of the action categories to be matched that match the teaching instructions in the effective duration of the teaching instructions, the learning concentration of the target object can be detected more accurately.
[0149] In one possible implementation, while improving the accuracy of the detection results of the target object's learning concentration, it is also necessary to save computing resources and reduce the complexity of the calculation. Therefore, in an embodiment of the present application, for multiple frames of similar images in an image sequence whose similarity is higher than a similarity threshold, a human key point detection model and a face key point detection model can be used to process one frame of the multiple frames of similar images to obtain the human key points and face key points corresponding to the multiple frames of similar images.
[0150] For example, if the human body key point detection model and the facial key point detection model have been used to process the first frame of 10 frames of images to obtain the human body key points and facial key points of the first frame, then if it is detected that the similarity between the remaining 9 frames of images and the first frame is higher than the similarity threshold, there is no need to use the human body key point detection model and the facial key point detection model to process the remaining 9 frames of images. It is only necessary to directly use the human body key points and facial key points of the first frame as the human body key points and facial key points of the remaining 9 frames.
[0151] In order to determine whether the similarity of multiple frames in an image sequence is higher than the similarity threshold, it is necessary to detect the similarity between the current frame and the previous frame in the image sequence in advance. After adding the similarity between the current frame and the previous frame, the flow chart of the detection method for learning concentration in online teaching is as follows: Fig.12 As shown, the specific steps include:
[0152] In step 121, a current frame image is input.
[0153] In step 122, the similarity between the current frame image and the previous frame image is calculated.
[0154] In step 123, determine whether the similarity is higher than the similarity threshold. If the similarity is higher than the similarity threshold, then in step 124, the key point sequence of the previous frame image is used as the key point sequence of the current frame image, and then execute step 126; if the similarity is less than or equal to the similarity threshold, then in step 125, based on the human key point detection model, obtain the human key points of the current frame image and based on the face key point detection model, obtain the face key points of the current frame image to obtain the key point sequence.
[0155] The similarity of multiple similar images in the image sequence can be determined by using a brightness mean comparison algorithm or a hash algorithm such as LSB (Least Significant Bit).
[0156] In step 126, the key point sequence is input into the action recognition deep learning model to obtain the corresponding action category.
[0157] In step 127, the teaching instructions in the audio data of the online teaching are detected.
[0158] The process of executing step 121 to step 125 is performed synchronously with the process of executing step 127.
[0159] In step 128, the action category with the detected time coincident with the teaching instruction is matched.
[0160] In step 129, the learning concentration of the target object is obtained.
[0161] Therefore, while improving the accuracy of the detection results of the target object's learning concentration, it also greatly saves the computing resources of the terminal device, reduces the complexity of the calculation, and simplifies the calculation process.
[0162] In a possible implementation, if it is necessary to simultaneously obtain the learning concentration of the target object in a shorter period of time and the learning concentration of the target object in a longer period of time, it is necessary to detect the learning concentration of the target object in the shorter period of time and the learning concentration of the target object in the longer period of time respectively, and it is necessary to repeat many calculation processes, wasting computing resources. Therefore, in order to save computing resources and simplify the calculation process, in the embodiment of the present application, the mean of the learning concentration of the target object in the first cycle can also be determined as the average concentration of the target object; then the learning concentration of the first cycle is averaged or weighted summed to obtain the learning concentration of the second cycle; the second cycle includes multiple first cycles.
[0163] For example, if it is necessary to obtain the target object's learning concentration in each class and each day at the same time, the first cycle is one class and the second cycle is one day. At this time, it is only necessary to detect the target object's learning concentration in each class, and then calculate the average or weighted sum of the target object's learning concentration in all classes of the day, so as to obtain the target object's learning concentration for that day, without having to detect the target object's learning concentration for that day again.
[0164] Among them, different weights can be set for the learning concentration of different classes according to the different requirements for learning concentration in each class, and then the learning concentration of all classes is weighted and summed to obtain the learning concentration of the target object every day. For example, if there is only one physical education class and one mathematics class in a day, the learning concentration required for the physical education class may be lower than that required for the mathematics class. When calculating the learning concentration of the target object on this day, the weight of the learning concentration of the physical education class can be set to 0.3, and the weight of the learning concentration of the mathematics class can be set to 0.7, and then the learning concentration of the physical education class can be multiplied by 0.3 and the learning concentration of the mathematics class can be multiplied by 0.7 to obtain the learning concentration of the target object on this day.
[0165] Therefore, when detecting the learning concentration of a target object over a longer period of time, the computing resources of the terminal device can be saved and the computing process is simplified.
[0166] Based on the foregoing description, the embodiment of the present application detects the teaching instructions in the audio data of the online teaching, and obtains the action category to be matched within the validity period of the teaching instructions; the action category to be matched is obtained by performing action detection on the target object in the video data; the action category to be matched is matched with the teaching instructions to obtain the learning concentration of the target object. Thus, by comparing the degree of conformity between the teacher's teaching instructions and the action category of the target object within the validity period of the teaching instructions, the learning concentration of the target object is determined according to the degree of conformity. Compared with the concentration detection method based only on line of sight detection and head movement detection, the present application can improve the accuracy of the detection results of the learning concentration of the target object, so that the teacher can more accurately monitor the learning concentration of the target object and improve the quality of teaching.
[0167] In an exemplary embodiment, the present application further provides a computer-readable storage medium including instructions, such as a memory 120 including instructions, and the above instructions can be executed by the processor 180 of the terminal device 100 to complete the above-mentioned method for detecting learning concentration for online teaching. Optionally, the computer-readable storage medium can be a non-temporary computer-readable storage medium, for example, the non-temporary computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0168] In an exemplary embodiment, a computer program product is also provided, including a computer program, which, when executed by the processor 180, implements the method for detecting learning concentration in online courses as provided in the present application.
[0169] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0170] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0171] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0172] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0173] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. A method for detecting learning concentration in online teaching, characterized in that: The method comprises: Detect teaching instructions in audio data of online teaching, and obtain the action category to be matched within the validity period of the teaching instructions; the action category to be matched is obtained by performing action detection on the target object in the video data; the obtaining of the action category to be matched within the validity period of the teaching instructions includes: on the premise that the teaching instructions are the latest teaching instructions, determining the action category of the detected target object as the action category to be matched; or, recording the first detection time of the teaching instructions and the second detection time of each action category, if the second detection time of the action category is within a specified time interval, determining the action category as the action category to be matched, the start time of the specified time interval is the first detection time of the teaching instructions, and the end time of the specified time interval is the first detection time of the next teaching instruction; The action category to be matched is matched with the teaching instruction to obtain the learning concentration of the target object; the method also includes: determining the second detection time of the first action category to be matched that matches the teaching instruction; determining the time difference between the second detection time of the first action category to be matched and the first detection time of the teaching instruction; dividing the preset action delay time by the time difference and then multiplying by a first value to obtain an updated first value, wherein the preset action delay time is used to represent the time required from issuing the teaching instruction to executing the teaching instruction, and the first value is the learning concentration of the target object determined when the first action category to be matched and the teaching instruction have a corresponding relationship.
2. The method according to claim 1, characterized in that The performing action detection on the target object in the video data includes: Taking every n consecutive frames of images in the video data as an image sequence, where n is a positive integer; For each image sequence, execute: Acquire key point data of the target object in each frame of the image sequence to obtain a key point sequence formed by the key point data of each frame of the image; the key point data includes face key points and body key points; The key point sequence is input into the action recognition deep learning model to obtain the action category of the target object.
3. The method according to claim 1, characterized in that The matching operation of the action category to be matched with the teaching instruction to obtain the learning concentration of the target object includes: Searching in a pre-stored correspondence table whether the action category to be matched and the teaching instruction have a correspondence; If there is a corresponding relationship, determining the learning concentration of the target object to be a first value; If there is no corresponding relationship, the learning concentration of the target object is determined to be a second value.
4. The method according to claim 1, characterized in that: The method further comprises: The duration of the validity period of the teaching instruction is used as the effective duration, and the duration of the action category to be matched that matches the teaching instruction is determined; The continuation duration is divided by the effective duration and then multiplied by the updated first value to obtain a final value of the first value of the target object.
5. The method according to claim 1, characterized in that The detecting of teaching instructions in the audio data of the online teaching comprises: Performing audio recognition on the audio data to obtain a text sequence; The text sequence is detected based on preset teaching instruction keywords to obtain teaching instructions.
6. The method according to claim 2, characterized in that The step of acquiring key point data of the target object in each frame of the image sequence includes: Based on a human key point detection model, obtaining human key points of each frame of the image sequence; Based on the facial key point detection model, the facial key points of each frame image in the image sequence are obtained.
7. The method according to claim 6, characterized in that The method further comprises: For multiple frames of similar images in the image sequence whose similarity is higher than a similarity threshold, the human body key point detection model and the facial key point detection model are used to process one frame of the multiple frames of similar images to obtain the human body key points and facial key points corresponding to the multiple frames of similar images respectively.
8. A terminal device, characterized in that: include: Display, processor and memory; The display is used to display the screen display area; The memory is used to store the processor executable instructions; The processor is configured to execute the instructions to implement the method for detecting learning concentration in online courses as described in any one of claims 1-7.
Citation Information
Patent Citations
Feedback type fatigue detecting system
CN101968918A