Multi-mode English teaching device for English teaching
By integrating detection and control modules in the multimodal English teaching device, the brightness of the video player and the speaker sound are automatically adjusted, which solves the problem of students being idle and distracted and improves the teaching quality.
Patent Information
- Application Number
- CN202510132835.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing multimodal English teaching device can easily cause students to be idle and distracted when students watch videos of the same brightness or lectures of the same sound intensity for a long time. Teachers need to frequently adjust the brightness of the video player or the speaker sound, which is troublesome and affects the teaching quality.
A multimodal English teaching device is designed, including a video player, a speaker, a camera, a detection module, a control module and a processing module. The detection module recognizes the student's daze rate, video player brightness and speaker volume, and uses the processing module to build a control coefficient, automatically adjusts the video player brightness and speaker sound to maintain students' attention.
It realizes automatic and timely adjustment of the brightness of the video player and the speaker sound, attracts students' attention, improves teaching quality, and reduces the trouble of the teachers' operation.
Smart Images

Figure CN119992894A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of teaching devices, and in particular to a multimodal English teaching device for English teaching. Background Art
[0002] With the continuous development of multimedia technology, network technology and scientific means, traditional teaching methods are gradually being eliminated and replaced by multimedia teaching. Multimodal English teaching is a multimodal interaction between subjects and objects of teaching, including interaction between students and problem situations, interaction between teachers and network environment, etc. This interaction is manifested in communication between them, the ability to give evaluation and feedback, and the ability to jointly promote the development of the teaching process, realizing the combination of multimodal and multimedia classroom teaching.
[0003] The multimodal English teaching device in the prior art mainly displays the teaching content through a display screen and uses a player and a microphone for lectures. In this multimodal English teaching device, students are prone to daze and distraction when watching videos with the same brightness or lectures with the same sound intensity for a long time. At this time, experienced teachers will adjust the brightness of the video player or the sound of the speaker in time to attract students' attention, but this operation is more troublesome and affects the teacher's energy. For this reason, we propose a multimodal English teaching device for English teaching. Summary of the invention
[0004] The present invention provides a multimodal English teaching device for English teaching to solve the problems raised in the above background technology.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions: A multimodal English teaching device for English teaching, comprising a video player and a speaker, a camera, a detection module, a control module, and a processing module; The detection module includes a daze rate detection unit for identifying the daze rate of students in the class, a brightness detection unit for detecting the brightness of the video player, and a volume detection unit for detecting the sound of the speaker; The control module is connected to the video player and the speaker by electrical signals; The processing module includes an evaluation unit and a judgment unit.
[0006] As a further improvement of the technical solution: the camera is used to take real-time facial pictures of all students in the class.
[0007] As a further improvement of the technical solution: the detection method of the daze rate detection unit is: sending the real-time facial pictures of all students in the class taken by the camera to the daze state recognition model to identify the number of students who are distracted and dazed; According to the formula: the rate of daydreaming during class = the number of students who are distracted / the number of students in the class, we can get the rate of daydreaming during class.
[0008] As a further improvement of this technical solution: the method for constructing the daze state recognition model is: S1: Obtain a large number of facial specimen images of dazed students; S2: marking the tissue region of interest in the dazed student face specimen image to obtain a marked dazed student face specimen image; S3: A large number of labeled facial specimen images of dazed students and annotation information are used as a data set, and the data set is used to train a machine learning model, wherein the machine learning model uses the standardized facial specimen images of dazed students and annotation information as input, and whether the student is dazed as output, and a dazed state recognition model is obtained through supervised learning training.
[0009] As a further improvement of the present technical solution: in S3, the selected machine learning model is at least one of logistic regression, linear discriminant analysis, K nearest neighbor, naive Bayes, support vector machine, random forest, and neural network.
[0010] As a further improvement of the present technical solution: during training, the data set is divided into a prediction set and a training set. After the training, the validity of the model is verified in the test set. If the verification fails, the process goes to S2 to add more labeled data; if the verification passes, the training ends.
[0011] As a further improvement of this technical solution: if the average AP value of the test set is <0.95, the verification fails; if the average AP value is ≥0.95, the verification passes.
[0012] As a further improvement of this technical solution: the teaching method of the multimodal English teaching device is: In the first step, when teaching English, the teacher uses a video player to display the teaching content and uses a loudspeaker to give lectures; The second step is to detect the daze rate of students in class, the brightness of the video player, and the volume of the speaker through the daze rate detection unit, the brightness detection unit, and the volume detection unit in the detection module, respectively, to form the daze rate M, the brightness N, and the volume B; The third step is to summarize the acquired daze rate M, brightness N, and volume B to form the detection condition information, build a data model and optimize it, and perform regression analysis on the acquired daze rate M, brightness N, and volume B through the analysis software SPSS to obtain the influence of daze rate M, brightness N, and volume B on teaching quality, and output the daze rate influencing factor Am, brightness influencing factor An, and volume influencing factor Ab; In the fourth step, the evaluation unit obtains the daze rate M, brightness N, and volume B, and performs dimensionless processing on them, and associates them to form the control coefficient JX. The association model is:
[0013] Among them, Am represents the influencing factor of daze rate, 0.17≤Am≤0.53; An represents the influencing factor of brightness, 0.44≤An≤0.79; Ab represents the influencing factor of volume, 0.27≤Ab≤0.85; The fifth step is to transmit the obtained control coefficient JX to the judgment unit for comparison with the preset threshold. If the control coefficient JX is not within the range of the preset threshold, a corresponding control instruction is generated to adjust the brightness of the video player and the volume of the speaker. At the same time, the control coefficient JX continues to be evaluated until the control coefficient JX is within the preset threshold, thereby ensuring the quality of teaching.
[0014] Compared with the prior art, the present invention has the following beneficial effects: The present invention can automatically and timely regulate the brightness of the video player and the volume of the speaker by constructing the regulation coefficient JX, so as to attract students' attention and ensure the quality of teaching.
[0015] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention and implement it according to the contents of the specification, the following is a detailed description of the preferred embodiments of the present invention in conjunction with the accompanying drawings. The specific implementation of the present invention is given in detail by the following embodiments and their accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings: Figure 1 A system schematic diagram of a multimodal English teaching device for English teaching proposed by the present invention; Figure 2 This is a schematic diagram of the method for constructing the daze state recognition model proposed in the present invention. DETAILED DESCRIPTION
[0017] The principles and features of the present invention are described below in conjunction with the accompanying drawings. The examples given are only used to explain the present invention and are not used to limit the scope of the present invention. The present invention is described in more detail by way of example with reference to the accompanying drawings in the following paragraphs. It should be noted that the accompanying drawings are all in a very simplified form and are not in precise proportions, and are only used to facilitate and clearly assist in explaining the purpose of the embodiments of the present invention.
[0018] See also Figures 1-2,In an embodiment of the present invention, a multimodal English teaching device for teaching English includes a video player and a speaker, a camera, a detection module, a control module, and a processing module, where the camera is used to take real-time facial pictures of all students in the class; The control module is connected with the video player and the speaker by electrical signals; A processing module, including an evaluation unit and a judgment unit; The detection module includes a daze rate detection unit for identifying the daze rate of students in the class, a brightness detection unit for detecting the brightness of the video player, and a volume detection unit for detecting the sound of the speaker; The detection method of the daze rate detection unit is as follows: the real-time facial pictures of all students in the class taken by the camera are sent to the daze state recognition model to identify the number of students who are distracted and dazed; According to the formula: the rate of daydreaming during class = the number of students who are distracted / the number of students in the class, we can get the rate of daydreaming during class.
[0019] The construction method of the daze state recognition model is: S1: Obtain a large number of facial specimen images of dazed students; S2: marking the tissue region of interest in the dazed student face specimen image to obtain a marked dazed student face specimen image; S3: A large number of labeled dazed student facial specimen images and annotation information are used as a data set, and the data set is used to train a machine learning model, wherein the machine learning model uses the standardized dazed student facial specimen images and annotation information as input, and whether the student is dazed as output, and a dazed state recognition model is obtained through supervised learning training.
[0020] In S3, the selected machine learning model is at least one of logistic regression, linear discriminant analysis, K nearest neighbor, naive Bayes, support vector machine, random forest, and neural network.
[0021] During training, the data set is divided into a prediction set and a training set. After the training, the effectiveness of the model is verified in the test set. If the verification fails, it is transferred to S2 to add more labeled data; if the verification passes, the training ends.
[0022] If the average AP value of the test set is <0.95, the verification fails; if the average AP value is ≥0.95, the verification passes The teaching method of the multimodal English teaching device is: In the first step, when teaching English, the teacher uses a video player to display the teaching content and uses a loudspeaker to give lectures; The second step is to detect the daze rate of students in class, the brightness of the video player, and the volume of the speaker through the daze rate detection unit, the brightness detection unit, and the volume detection unit in the detection module, respectively, to form the daze rate M, the brightness N, and the volume B; The third step is to summarize the acquired daze rate M, brightness N, and volume B to form the detection condition information, build a data model and optimize it, and perform regression analysis on the acquired daze rate M, brightness N, and volume B through the analysis software SPSS to obtain the influence of daze rate M, brightness N, and volume B on teaching quality, and output the daze rate influencing factor Am, brightness influencing factor An, and volume influencing factor Ab; In the fourth step, the evaluation unit obtains the daze rate M, brightness N, and volume B, and performs dimensionless processing on them, and associates them to form the control coefficient JX. The association model is:
[0023] Among them, Am represents the influencing factor of daze rate, 0.17≤Am≤0.53; An represents the influencing factor of brightness, 0.44≤An≤0.79; Ab represents the influencing factor of volume, 0.27≤Ab≤0.85; The fifth step is to transmit the obtained control coefficient JX to the judgment unit for comparison with the preset threshold. If the control coefficient JX is not within the range of the preset threshold, a corresponding control instruction is generated to adjust the brightness of the video player and the volume of the speaker. At the same time, the control coefficient JX continues to be evaluated until the control coefficient JX is within the preset threshold, thereby ensuring the quality of teaching.
[0024] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Any ordinary technician in the industry can smoothly implement the present invention as shown in the drawings and described above. However, any equivalent changes, modifications and evolutions made by technicians familiar with the profession without departing from the scope of the technical solution of the present invention using the technical content disclosed above are all equivalent embodiments of the present invention. At the same time, any equivalent changes, modifications and evolutions made to the above embodiments based on the essential technology of the present invention are still within the protection scope of the technical solution of the present invention.
Claims
1. A multimodal English teaching device for English teaching, characterized in that: Including video player and speaker, camera, detection module, control module, processing module; The detection module includes a daze rate detection unit for identifying the daze rate of students in the class, a brightness detection unit for detecting the brightness of the video player, and a volume detection unit for detecting the sound of the speaker; The control module is connected to the video player and the speaker by electrical signals; The processing module includes an evaluation unit and a judgment unit.
2. The multimodal English teaching device for English teaching according to claim 1, characterized in that: The camera is used to take real-time facial images of all students in the class.
3. The multimodal English teaching device for English teaching according to claim 2, characterized in that: The detection method of the daze rate detection unit is as follows: the real-time facial pictures of all students in the class taken by the camera are sent to the daze state recognition model to identify the number of students who are distracted and dazed; According to the formula: the rate of daydreaming during class = the number of students who are distracted / the number of students in the class, we can get the rate of daydreaming during class.
4. The multimodal English teaching device for English teaching according to claim 3, characterized in that: The construction method of the daze state recognition model is: S1: Obtain a large number of facial specimen images of dazed students; S2: marking the tissue region of interest in the dazed student face specimen image to obtain a marked dazed student face specimen image; S3: A large number of labeled facial specimen images of dazed students and annotation information are used as a data set, and the data set is used to train a machine learning model, wherein the machine learning model uses the standardized facial specimen images of dazed students and annotation information as input, and whether the student is dazed as output, and a dazed state recognition model is obtained through supervised learning training.
5. The multimodal English teaching device for English teaching according to claim 4, characterized in that: In S3, the selected machine learning model is at least one of logistic regression, linear discriminant analysis, K nearest neighbor, naive Bayes, support vector machine, random forest, and neural network.
6. The multimodal English teaching device for English teaching according to claim 5, characterized in that: During training, the data set is divided into a prediction set and a training set. After the training is completed, the validity of the model is verified in the test set. If the verification fails, it is transferred to S2 to add more labeled data; If the verification passes, the training ends.
7. The multimodal English teaching device for English teaching according to claim 6, characterized in that: If the average AP value of the test set is <0.95, the verification fails; if the average AP value is ≥0.95, the verification passes.
8. The multimodal English teaching device for English teaching according to claim 7, characterized in that: The teaching method of the multimodal English teaching device is: In the first step, when teaching English, the teacher uses a video player to display the teaching content and uses a loudspeaker to give lectures; The second step is to detect the daze rate of students in class, the brightness of the video player, and the volume of the speaker through the daze rate detection unit, the brightness detection unit, and the volume detection unit in the detection module, respectively, to form the daze rate M, the brightness N, and the volume B; The third step is to summarize the acquired daze rate M, brightness N, and volume B to form the detection condition information, build a data model and optimize it, and perform regression analysis on the acquired daze rate M, brightness N, and volume B through the analysis software SPSS to obtain the influence of daze rate M, brightness N, and volume B on teaching quality, and output the daze rate influencing factor Am, brightness influencing factor An, and volume influencing factor Ab; In the fourth step, the evaluation unit obtains the daze rate M, brightness N, and volume B, and performs dimensionless processing on them, and associates them to form the control coefficient JX. The association model is:
9. Among them, Am represents the influencing factor of daze rate, 0.17≤Am≤0.53; An represents the influencing factor of brightness, 0.44≤An≤0.79; Ab represents the influencing factor of volume, 0.27≤Ab≤0.85; The fifth step is to transmit the obtained control coefficient JX to the judgment unit for comparison with the preset threshold. If the control coefficient JX is not within the range of the preset threshold, a corresponding control instruction is generated to adjust the brightness of the video player and the volume of the speaker. At the same time, the control coefficient JX continues to be evaluated until the control coefficient JX is within the preset threshold, thereby ensuring the quality of teaching.