Multimodal extraction methods, systems, and storage media for teachers' implicit experiences
By acquiring teaching data from outstanding teachers, extracting verbal, physical, and emotional features, constructing an implicit experience expression model, and forming an experience database, the problem of the difficulty in extracting implicit experiences from ordinary teachers is solved, thereby improving teachers' teaching level.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUIZHOU JINGSHI CHENGTOU INTELLIGENT EDUCATION IND CO LTD
- Filing Date
- 2022-11-30
- Publication Date
- 2026-05-26
AI Technical Summary
Existing technologies make it difficult to systematically extract and organize teachers' implicit experiences, resulting in the inability to effectively improve the teaching abilities of ordinary teachers.
By acquiring raw teaching data from outstanding teachers, we extract behavioral features across multiple dimensions, including verbal, physical, and emotional characteristics, to construct an implicit experience expression model and form an experience database. This helps ordinary teachers identify their own implicit experience deficiencies.
It enables implicit experience to be made explicit into a learnable experience base, helping ordinary teachers to discover their own strengths and weaknesses and improve their teaching abilities.
Smart Images

Figure CN115712764B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of educational information technology, specifically to a multimodal extraction method, system, and storage medium for teachers' implicit experience. Background Technology
[0002] Basic education is crucial for cultivating talent, and teacher education, as the key to the education system, bears the responsibility of training teachers for schools at all levels and of all types. Although senior or outstanding teachers are the best in the basic education teaching force, their numbers are extremely small and their distribution is uneven. Their teaching level can only reflect the room for improvement, but cannot represent the average level. Therefore, the teaching ability of the vast majority of ordinary teachers is at a mid-level. How to improve the ability of ordinary teachers has always been a core issue of teacher education.
[0003] Generally speaking, teacher experience is divided into two types: explicit and implicit. Explicit experience refers to knowledge that can be expressed verbally, while implicit experience refers to knowledge that cannot be expressed in language. In fact, the current teacher education system focuses on training in explicit experience, such as through common methods like videos, on-site learning, textbooks, and discussion-based learning. However, the shortcomings in the transmission of knowledge and skills based on explicit experience are quite obvious. When there is an unequal experience gap between excellent and ordinary teachers, relying solely on the exchange of experiences that can be expressed verbally is far from sufficient, which means "knowing but not understanding."
[0004] The key to learning from outstanding teachers lies not in explicit knowledge, but in implicit knowledge. Essentially, the teaching activities of excellent teachers are a combination of explicit and implicit experience. Implicit knowledge is not only unclear and difficult to express, but also lacks integration, refinement, and holistic processing. It cannot be automatically extracted by excellent teachers, resulting in a vast amount of "silent knowledge" existing only within them. This "silent knowledge," as an objective reality, guides teachers' professional practice, thinking, and actions in a state of "using without knowing." Therefore, how to systematically refine and organize teachers' implicit experience is crucial to improving the level of the teaching force and is an urgent problem to be solved. Summary of the Invention
[0005] The present invention aims to provide a multimodal extraction method, system and storage medium for teachers' implicit experience, which can extract the implicit experience of excellent teachers and help other teachers learn from the implicit experience of excellent teachers.
[0006] The basic solution provided by this invention: a multimodal extraction method for teachers' implicit experience, comprising the following steps:
[0007] S100: Obtain raw teaching data from excellent teachers;
[0008] S200: Extract behavioral features from multiple dimensions related to the implicit experience of excellent teachers from the original teaching data, and generate behavioral feature sets under multiple dimensions respectively;
[0009] S300: Construct an implicit experience representation model for behavioral features, and incorporate a set of behavioral features from multiple dimensions into the implicit experience representation model;
[0010] S400: Repeat S100-S300, based on the implicit experience expression model, extract the implicit experience representations of multiple excellent teachers and store them to form an experience base.
[0011] The principle of this invention lies in acquiring raw teaching data from outstanding teachers, including video and audio recordings of their lectures. From this raw data, feature sets related to the teachers' implicit experiences across multiple dimensions are extracted. These feature sets are then fed into an implicit experience expression model, where each dimension represents a different aspect of the teacher's behavior. The implicit experience expression model is used to analyze the implicit experiences of outstanding teachers across these dimensions, quantifying these implicit experiences in each dimension. By analyzing the feature sets from multiple outstanding teachers across various dimensions, commonalities in their implicit experiences are extracted, forming an experience database.
[0012] Compared to existing technologies, it has the following advantages:
[0013] 1. It can extract and represent implicit experiences that are difficult to express. By modeling implicit experience extraction, it can extract and refine the implicit experiences of excellent teachers, and find commonalities among the implicit experiences of excellent teachers. It forms an experience database by summarizing the commonalities of a large number of excellent teachers' implicit experiences under the implicit experience expression model. This makes it easy for other teachers to compare with the data in the experience database and find the shortcomings of their own experience.
[0014] 2. The analysis process helps outstanding teachers discover their own undiscovered strengths. Implicit experience may exist without teachers' awareness, perhaps as a behavioral habit or subconscious tone of voice or gesture. Even outstanding teachers may not know which of their most prominent implicit experiences they possess, nor which areas they are weak in. Analyzing a set of characteristics across multiple dimensions helps outstanding teachers further clarify their strengths and weaknesses, enabling them to make improvements and address these shortcomings.
[0015] Furthermore, S200 includes the following steps:
[0016] S210: Extracting the verbal characteristics of excellent teachers from raw teaching data;
[0017] S211: Classify speech features separately through semantic recognition, and establish a speech feature set based on the speech feature classification results.
[0018] The speech features of outstanding teachers are extracted from the raw data. Speech features refer to the content, expression, tone and other features related to the way teachers speak when they teach. Through semantic recognition, the speech of outstanding teachers is classified, and various types of speech content and speaking styles are divided into multiple speech feature sets, thereby analyzing the speech features of outstanding teachers.
[0019] Furthermore, S211 includes the following steps:
[0020] S2111: Extract independent features, statistical features, and combined features from speech features according to preset rules. The independent features include question features and statement features.
[0021] S2112: Establish a set of speech features ;
[0022] in - For question-type features, Questions that focus on the subject matter Questions that stimulate metacognition Questions designed to encourage student participation. Questions that focus on evaluation - Represents statement-type features, It indicates the expression of content, tone, and speaking speed. This indicates a response to the student's question;
[0023] - For statistical features, Statistics on question types that focus on subject-specific content. The first question type focuses on actual answers;
[0024] This is a combination feature, representing a combination of question and answer.
[0025] Verbal features are categorized into independent features, statistical features, and combined features. Independent features include question-based features and declarative features. Teachers primarily deliver lessons through verbal expression; therefore, their speaking style plays a crucial role in teaching quality. This includes considerations such as when to ask questions, the content to focus on, and the target audience and timing of questions. By extracting and analyzing the verbal features of outstanding teachers, we can identify their strengths and facilitate learning from their methods.
[0026] Furthermore, step S200 also includes the following steps:
[0027] S220: Extracting nonverbal features of outstanding teachers from raw teaching data;
[0028] S221: Extract body posture and emotional features from nonverbal features through visual recognition analysis, and generate body posture feature sets and emotional feature sets respectively.
[0029] The nonverbal characteristics of outstanding teachers were extracted from the raw data and divided into body language characteristics and emotional characteristics. Body language characteristics refer to the teacher's actions when teaching, while emotional characteristics refer to the facial expressions and emotions expressed when teaching.
[0030] Furthermore, S221 includes the following steps:
[0031] S2211: Based on the recognition results, body posture features are divided into non-mobile body postures and mobile body postures;
[0032] S2212: Establish a set of body posture features
[0033] in Non-mobile features, including standing and sitting postures. Motion features include head pose, gestures, and gait.
[0034] A teacher's physical characteristics are also an important part of the teaching process. For example, a teacher's standing, sitting, and walking postures can convey a lot of information and express meanings that cannot be expressed in words. When students encounter knowledge points that are difficult to understand or distinguish, teachers can use gestures to illustrate them. For example, when teaching students about consonants and vowels, teachers can use gestures to imitate the position of the tongue inside the cavity when pronouncing a certain syllable.
[0035] Furthermore, S221 also includes the following steps:
[0036] S2213: Classify the emotion characteristics based on the recognition results;
[0037] S2214: Establish a set of emotional features
[0038] in To express happiness, To express sadness, To express anger, Indicates fear. To express disgust, Expressing surprise, It indicates a neutral emotion.
[0039] Meanwhile, facial expressions enhance the expressive effect of language. When students are listening to a lecture, they not only "listen to what the teacher says" but also "observe what the teacher does." Therefore, facial expressions should change with the teaching content, showing joy when appropriate and sorrow when appropriate, so that students can understand the knowledge through these changes in expression. However, the changes in facial expressions should not be excessive. Therefore, we extract and analyze the characteristics of teachers' emotional expressions.
[0040] Furthermore, S300 includes the following steps:
[0041] S310: Construct implicit experiential representation models of verbal features, body language features, and emotional features.
[0042]
[0043] S311: The sets of verbal features, body features, and emotional features extracted from the original teaching data are incorporated into the implicit experience expression model to generate implicit experience norms about excellent teachers in three dimensions.
[0044] An implicit experience expression model is established based on three dimensions: verbal features, body language features, and emotional features. Sets of verbal features, body language features, and emotional features extracted from raw data of excellent teachers' lectures are input into the implicit experience expression model to generate implicit experience norms for excellent teachers across these three dimensions. Other ordinary teachers can compare their own data with the implicit experience norms of excellent teachers, thereby identifying gaps in their own implicit experience by comparing discrepancies.
[0045] The present invention also discloses a multimodal extraction system for teachers' implicit experience, which uses the above-mentioned multimodal extraction method for teachers' implicit experience.
[0046] The present invention also discloses a storage medium for storing computer-executable instructions, which, when executed, enable the aforementioned multimodal extraction method for teachers' implicit experience. Attached Figure Description
[0047] Figure 1This is a flowchart illustrating an embodiment of the multimodal extraction method for teachers' implicit experience according to the present invention. Detailed Implementation
[0048] The following detailed description illustrates the specific implementation method:
[0049] The basic implementation examples are as follows: Figure 1 As shown:
[0050] A multimodal extraction method for teachers' implicit experiences includes the following steps:
[0051] S100: Obtain the original teaching data of outstanding teachers. Specifically, in this embodiment, the original teaching data is the video footage of outstanding teachers giving lectures. The data is entered into the system by means of obtaining it from the Internet or manually uploading it. The video data of outstanding teachers giving lectures in various subjects and fields are collected and stored in the database.
[0052] S200: Extract behavioral features from the original teaching data, focusing on multiple dimensions related to the implicit experiences of excellent teachers, and generate behavioral feature sets for each dimension. In this embodiment, the extracted behavioral features include verbal and nonverbal features, with nonverbal features including body language and emotional features. Verbal features specifically refer to the teacher's language organization and expression during teaching; body language features specifically refer to the teacher's body movements during teaching; and emotional features specifically refer to the teacher's facial expressions during teaching. In S200, verbal features, body language features, and emotional features are extracted separately, and sets of verbal features, body language features, and emotional features are generated respectively.
[0053] Specifically, S200 includes the following steps:
[0054] S210: Extract speech features of excellent teachers from the original teaching data. Specifically, this involves extracting audio from the video data, filtering out noise, and obtaining the audio data of the excellent teachers' speech as speech features.
[0055] By recognizing video footage, the audio of excellent teachers' speech during lectures is extracted. The recognition algorithm distinguishes the teachers' speaking audio data based on timbre, using this data as the speech features of the dataset. Simultaneously, audio data of students asking and answering questions is also identified, providing a basis and support for subsequent classification of the speech features of excellent teachers.
[0056] S211: Speech features are classified separately through semantic recognition, and a speech feature set is established based on the classification results. Using existing semantic recognition algorithms, the content of teachers' speech is identified, and classified according to preset conditions. At the same time, by recognizing the tone and intonation of excellent teachers, their expression style is determined. The audio data of excellent teachers' lectures are classified according to the content and style of speech, thereby establishing a speech feature set of excellent teachers.
[0057] Specifically, S211 includes the following steps:
[0058] S2111: Extract independent features, statistical features, and combined features from the speech features according to preset rules. Independent features include question features and statement features. In this embodiment, speech features are subdivided into three dimensions, including 6 independent features, 2 statistical features, and 1 combined feature. The 6 independent features include 4 question features and 2 statement features.
[0059] By analyzing the tone and intonation of excellent teachers' speech, we can distinguish between interrogative and declarative sentences, thus identifying their respective characteristics. Semantic recognition is then performed on the content of interrogative sentences, combining this with the course content and the content and time intervals of student responses to determine whether the question focuses on subject content, stimulates metacognition, encourages student participation, or focuses on evaluation.
[0060] At the same time, the frequency of question types asked by excellent teachers is statistically analyzed to obtain statistical characteristics. In this embodiment, the statistics include question types that focus on subject content and the first question type that focuses on the actual answer.
[0061] Furthermore, the system identifies the combination features of question and answer combinations in the audio data of excellent teachers based on the content.
[0062] Then, a speech feature set S is constructed based on the classification results, which includes the following steps:
[0063] S2112: Establish a set of speech features .in - For question-type features, Questions that focus on the subject matter Questions that stimulate metacognition Questions designed to encourage student participation. This indicates a question that focuses on evaluation.
[0064] in - For question-type features, Questions that focus on the subject matter Questions that stimulate metacognition Questions designed to encourage student participation. Questions that focus on evaluation - Represents statement-type features, It indicates the expression of content, tone, and speaking speed. This indicates a response to the student's question;
[0065] - For statistical features, Statistics on question types that focus on subject-specific content. The first question type focuses on actual answers;
[0066] This is a combination feature, representing a combination of question and answer.
[0067] The above steps complete the classification of the speech feature set. By examining question-based features, we can focus on the timing, frequency, and content of questions asked by excellent teachers. This analysis reveals when, where, and what kind of questions are likely to stimulate student thinking, metacognition, and participation. By examining declarative features, we can focus on how excellent teachers express the course content, their tone and intonation when explaining it, and how they respond to student questions.
[0068] The above describes the process of establishing a set of verbal features from excellent teachers' lectures. For the extraction of nonverbal features and the establishment of the nonverbal feature set, step S200 also includes the following steps:
[0069] S220: Extracting nonverbal features of outstanding teachers from raw teaching data;
[0070] S221: Extract body posture features and emotional features from nonverbal features through visual recognition analysis, and generate body posture feature sets and emotional feature sets respectively.
[0071] In this embodiment, verbal features specifically include body language features and emotional features. Existing video analysis algorithms are used to identify the nonverbal features of excellent teachers. Specifically, this embodiment employs Action Recognition, using the UCF-101 and HMDB-51 datasets to identify and analyze the body movements, facial movements, facial expressions, and interactive actions of excellent teachers in video footage. Based on the identified body movements and interactive actions combined with the spoken content of the audio data, the correlation between the body movements or interactive actions and the spoken content is determined to identify whether the body language features are related to teaching. For example, when encountering knowledge points that students find difficult to understand or distinguish, teachers may use body language to explain them; or when teaching students about initials and finals, teachers may use gestures to imitate the position of the tongue inside the cavity when pronouncing a syllable. This embodiment also divides body language features into non-movement features and movement features. Non-movement features include standing and sitting postures, while movement features include head posture, gestures, and gait.
[0072] After extracting non-moving and moving features, a body posture feature set is established. ,in Non-mobile features, including standing and sitting postures. These are movement features, including head posture, gestures, and gait. This completes the establishment of a set of body posture features.
[0073] Based on the identified facial movements and expressions, combined with the spoken content, tone, and intonation in the audio data, it is determined whether the emotions are related to the lesson, such as whether the teacher's expressions change with the teaching content, showing joy when appropriate and sorrow when appropriate. The emotional characteristics identified in this embodiment include happiness, sadness, anger, fear, disgust, surprise, and neutral emotions.
[0074] After extracting the emotional features, establish the emotional feature set S2214: Establish the emotional feature set
[0075] in To express happiness, To express sadness, To express anger, Indicates fear. To express disgust, Expressing surprise, It indicates a neutral emotion.
[0076] Through the above steps, the sets of verbal features, body language features, and emotional features are established, and the implicit experiences of excellent teachers are evaluated under these three dimensions.
[0077] S300: Construct an implicit experience expression model for behavioral characteristics, incorporating sets of verbal features, body language features, and emotional features into the implicit experience expression model.
[0078] Specifically, S300 includes the following steps:
[0079] S310: Construct implicit experiential representation models of verbal features, body language features, and emotional features.
[0080]
[0081] S311: The sets of verbal features, body features, and emotional features extracted from the original teaching data are incorporated into the implicit experience expression model to generate implicit experience norms about excellent teachers in three dimensions.
[0082] Specifically, the implicit experience representation model is a multi-dimensional vector composed of verbal features, body language features, and emotional features, forming multiple values in the three dimensions. The results of analyzing the three dimensions of verbal features, body language features, and emotional features in the teacher's teaching process are representations of implicit experience. The algorithm is a disjunctive method, which obtains the surface of excellent teachers. It compares the deviation from the norm surface and gives feedback based on the dimension with the largest deviation, providing personalized recommendations.
[0083] S400: Repeat S100-S300, based on the implicit experience expression model, extract the implicit experience representations of multiple excellent teachers and store them to form an experience base.
[0084] Prior to extracting the implicit experience of outstanding teachers, the implicit experience representations of multiple outstanding teachers are repeatedly extracted, and an experience database is formed on this basis. Ordinary teachers can upload their own videos and compare scores under multiple indicators, which can be used as a basis to analyze the shortcomings of ordinary teachers in the classroom teaching process.
[0085] This embodiment also discloses a multimodal extraction system for teachers' implicit experience, which uses the above-described multimodal extraction method for teachers' implicit experience.
[0086] This embodiment also discloses a storage medium for storing computer-executable instructions, which, when executed, can implement the above-described multimodal extraction method for teachers' implicit experience.
[0087] The above are merely embodiments of the present invention. Commonly known structures and characteristics are not described in detail here. Those skilled in the art are aware of all common technical knowledge in the field prior to the application date or priority date, are aware of all existing technologies in that field, and have the ability to apply conventional experimental methods prior to that date. Those skilled in the art can, under the guidance of this application, improve and implement this solution in combination with their own capabilities. Some typical known structures or methods should not be obstacles for those skilled in the art to implement this application. It should be noted that those skilled in the art can make several modifications and improvements without departing from the structure of the present invention. These should also be considered within the scope of protection of the present invention, and will not affect the effectiveness of the implementation of the present invention or the practicality of the patent. The scope of protection claimed in this application should be determined by the content of its claims, and the specific embodiments described in the specification can be used to interpret the content of the claims.
Claims
1. A multi-modal extraction method of teacher's tacit experience, characterized in that: Includes the following steps: S100: Obtain raw teaching data from excellent teachers; S200: Extract behavioral features from multiple dimensions related to the implicit experience of excellent teachers from the original teaching data, and generate behavioral feature sets under multiple dimensions respectively; S300: Construct an implicit experience representation model for behavioral features, and incorporate a set of behavioral features from multiple dimensions into the implicit experience representation model; S400: Repeat S100-S300, based on the implicit experience expression model, extract the implicit experience representations of multiple excellent teachers and store them to form an experience base. S200 includes the following steps: S210: Extracting the verbal characteristics of excellent teachers from raw teaching data; S211: Classify speech features separately through semantic recognition, and establish a speech feature set based on the speech feature classification results; S211 includes the following steps: S2111: Extract independent features, statistical features, and combined features from speech features according to preset rules. The independent features include question features and statement features. S2112: Establish a set of speech features ; in - For question-type features, Questions that focus on the subject matter Questions that stimulate metacognition Questions designed to encourage student participation. Questions that focus on evaluation - Represents statement-type features, It indicates the expression of content, tone, and speaking speed. This indicates a response to the student's question; - For statistical features, Statistics on question types that focus on subject-specific content. The first question type focuses on actual answers; This is a composite feature, representing a combination of questions and answers; S200 further includes the following steps: S220: Extracting nonverbal features of outstanding teachers from raw teaching data; S221: Extract body posture and emotional features from nonverbal features through visual recognition analysis, and generate body posture feature set and emotional feature set respectively; S221 includes the following steps: S2211: Based on the recognition results, body posture features are divided into non-mobile body postures and mobile body postures; S2212: Establish a set of body posture features in Non-mobile features include standing and sitting postures. Motion features include head pose, gestures, and gait. S221 further includes the following steps: S2213: Classify the emotion characteristics based on the recognition results; S2214: Establish a set of emotional features in To express happiness, To express sadness, To express anger, Indicates fear. To express disgust, Expressing surprise, Indicates a neutral mood; S300 includes the following steps: S310: Construct implicit experiential representation models of verbal features, body language features, and emotional features. S311: The sets of verbal features, body features, and emotional features extracted from the original teaching data are incorporated into the implicit experience expression model to generate implicit experience norms about excellent teachers in three dimensions.
2. A multimodal extraction system for teachers' implicit experience, characterized in that: The multimodal extraction method for teachers' implicit experience as described in claim 1 was used.
3. A storage medium for storing computer-executable instructions, characterized in that, When the computer executes the instructions, it can implement the multimodal extraction method of teachers' implicit experience as described in claim 1.