Teaching case recommendation method, system, equipment and medium

By performing voice and role recognition of classroom transcripts, combining sequence labeling and text classification models, we extract the knowledge points and cognitive levels of questions asked by teachers, and recommend personalized teaching cases, solving the problem of insufficient analysis of teachers' questions in the existing technology, and improving the quality of classroom teaching.

CN120067434APending Publication Date: 2025-05-30SHANGHAI QIAOCHUANG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411965950.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing technology lacks in-depth analysis of the semantics, structures and cognitive levels of questions asked by teachers, making it difficult to provide teachers with personalized improvement suggestions and affects the quality of classroom teaching.

Method used

By performing speech recognition and role recognition on classroom transcripts, character labeling text is generated, and the sequence labeling model and text classification model are used to extract the knowledge points summary and cognitive levels of questions asked by teachers, and personalized teaching cases are recommended.

Benefits of technology

It realizes multi-dimensional analysis of teachers' questions, provides personalized teaching cases, greatly improves the quality of classroom teaching, and helps teachers optimize teaching strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067434A_ABST
    Figure CN120067434A_ABST
Patent Text Reader

Abstract

The invention relates to a teaching case recommendation method, system and device and a medium. The method comprises the steps of obtaining a classroom record; performing voice recognition and role recognition on the classroom record to correspondingly obtain a transcriptional text and a role identifier of a speaker, and mapping the role identifier to the transcriptional text to generate a role labeling text; extracting a teacher utterance text from the role labeling text, and inputting the teacher utterance text into an abstract generation model to generate a corresponding knowledge point abstract; inputting the role labeling text into a question recognition model, and extracting a classroom question proposed by the teacher role; inputting the classroom question into the question hierarchy classification model to obtain a cognitive hierarchy corresponding to the classroom question; and selecting a corresponding teaching case for recommendation based on each cognitive level and knowledge point abstract of the classroom record. The problem that personalized improvement suggestions are difficult to provide for teachers due to the lack of analysis on semantics, structures and cognitive levels of questions asked by the teachers in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent recommendation, and particularly to a method, system, device and medium for recommending teaching cases. Background Art

[0002] In order to improve the classroom teaching effect, it is crucial to comprehensively evaluate teachers' teaching activities. Classroom teaching evaluation usually starts from multiple dimensions such as teaching content, teaching design, teaching process and teaching effect. Among them, the quality of classroom questions, as a key indicator, directly affects the teaching effect. Therefore, classroom questions play multiple roles in teaching. It can not only help students consolidate knowledge, cultivate logical thinking, but also stimulate students' learning interest and improve classroom participation. At the same time, teachers can also understand students' learning situations through students' feedback, so as to optimize teaching strategies. However, there are some problems in actual teaching, such as insufficient teacher-student interaction in asking questions and overly simple question design, which to a certain extent limit the effectiveness of classroom questions. In addition, the current classroom teaching evaluation usually stays at the observation of teachers' behaviors, lacking in-depth analysis of the quality of teachers' questions. In order to effectively improve the quality of classroom questions and optimize the teaching effect, it is necessary to provide a method, system, device and medium for recommending teaching cases. Summary of the Invention

[0003] In view of the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a method, system, device and medium for recommending teaching cases, which improves the problem that the prior art lacks analysis of the semantics, structure and cognitive levels of teachers' questions, resulting in difficulty in providing personalized improvement suggestions for teachers.

[0004] To achieve the above object and other related objects, the present invention provides a method for recommending teaching cases, including: obtaining a classroom record; performing speech recognition and role recognition on the classroom record respectively, obtaining a transcribed text and a role identifier of the speaker correspondingly, and mapping the role identifier to the transcribed text to generate a role-annotated text; extracting teacher's speech text from the role-annotated text, and inputting the teacher's speech text into a summary generation model to generate a corresponding knowledge point summary; inputting the role-annotated text into a question recognition model to extract classroom questions raised by the teacher role; wherein, the question recognition model is a sequence annotation model; inputting the classroom questions into a question level classification model to obtain the cognitive level corresponding to the classroom questions; wherein, the question level classification model is a text classification model; based on the respective cognitive levels of the classroom record and the knowledge point summary, selecting corresponding teaching cases for recommendation.

[0005] In an embodiment of the present invention, the obtaining of the classroom record includes: obtaining an initial classroom record; performing audio noise reduction and echo cancellation processing on the initial classroom record to obtain a final classroom record.

[0006] In an embodiment of the present invention, the speech recognition and role recognition are respectively performed on the classroom record, and the corresponding transcription text and the role identifier of the speaker are obtained, and the role identifier is mapped to the transcription text to generate a role-annotated text, including: performing speech recognition on the classroom record to obtain the transcription text corresponding to the classroom record; performing speech activity detection on the classroom record, and based on the pause intervals in the classroom record, segmenting the classroom record into a plurality of discourse units; wherein each discourse unit belongs to a speech statement of the classroom record; for each discourse unit: inputting the discourse unit into a speaker embedding recognition model to generate a speech embedding vector; based on a clustering algorithm, aggregating the generated speech embedding vectors into a plurality of role clusters; counting the number of speech embedding vectors in each role cluster, and identifying the role cluster with the largest number of speech embedding vectors as the teacher cluster, and identifying the remaining at least one role cluster as the student cluster; based on the speech statement to which the discourse unit belongs, mapping the role identifiers corresponding to the teacher cluster and the student cluster to the text statements corresponding to the transcription text to generate a role-annotated text.

[0007] In an embodiment of the present invention, the abstract generation model is a sequence annotation model. The teacher's speech text is extracted from the role-annotated text, and the teacher's speech text is input into the abstract generation model to generate a corresponding knowledge point abstract, including: parsing the role identifier from the role-annotated text, identifying and extracting the teacher's speech text; performing word segmentation processing on the teacher's speech text to obtain a word segmentation sequence corresponding to the teacher's speech text; inputting the word segmentation sequence into the abstract generation model to obtain the predicted category of each word; selecting the corresponding words with the predicted category of knowledge points and arranging them according to the order of the words in the teacher's speech text to generate the knowledge point abstract corresponding to the teacher's speech text.

[0008] In an embodiment of the present invention, the abstract generation model is a sequence-to-sequence model. The teacher's speech text is extracted from the role-annotated text, and the teacher's speech text is input into the abstract generation model to generate a corresponding knowledge point abstract, including: parsing the role identifier from the role-annotated text, identifying and extracting the teacher's speech text; performing word segmentation processing on the teacher's speech text to obtain a word segmentation sequence corresponding to the teacher's speech text; inputting the word segmentation sequence into the encoder of the abstract generation model to extract the semantic features of the teacher's speech text and generate a semantic vector sequence corresponding to the teacher's speech text; inputting the semantic vector sequence into the decoder of the abstract generation model, and performing semantic recombination on each semantic vector based on the attention mechanism to obtain the knowledge point abstract corresponding to the teacher's speech text.

[0009] In an embodiment of the present invention, the cognitive levels include a first cognitive level, a second cognitive level, and a third cognitive level with gradually increasing levels. Selecting corresponding teaching cases for recommendation based on each cognitive level of the classroom record and the knowledge point summary includes: calculating the similarity between the knowledge point summary and each knowledge point in the case library, and selecting the teaching case set pre-associated with the corresponding knowledge point based on the similarity; wherein each teaching case is associated with a cognitive level; determining whether the proportion of the first cognitive level in all the questions in the classroom record is greater than or equal to a preset first threshold: if so, selecting teaching cases of the second cognitive level and / or the third cognitive level from the teaching case set for recommendation; if not, selecting corresponding teaching cases for recommendation based on the proportion of the second cognitive level in all the questions in the classroom record.

[0010] In an embodiment of the present invention, selecting corresponding teaching cases for recommendation based on the proportion of the second cognitive level in all the questions in the classroom record includes: determining whether the proportion of the second cognitive level in all the questions is greater than or equal to a preset second threshold: if so, selecting the teaching cases associated with the third cognitive level from the teaching case set for recommendation; if not, recommending the teaching cases associated with the second cognitive level and the third cognitive level from the teaching case set for recommendation.

[0011] In an embodiment of the present invention, a teaching case recommendation system is further provided. The system includes: a classroom record acquisition module for acquiring a classroom record; a role annotation module for performing speech recognition and role recognition on the classroom record respectively, obtaining a transcription text and a role identifier of the speaker correspondingly, and mapping the role identifier to the transcription text to generate a role-annotated text; a knowledge point recognition module for extracting the teacher's speech text from the role-annotated text and inputting the teacher's speech text into a summary generation model to generate a corresponding knowledge point summary; a question recognition module for inputting the role-annotated text into a question recognition model to extract the classroom questions proposed by the teacher role; wherein the question recognition model is a sequence annotation model; a cognitive level recognition module for inputting the classroom questions into a question level classification model to obtain the corresponding cognitive level; wherein the question level classification model is a text classification model; a case recommendation module for selecting corresponding teaching cases for recommendation based on each cognitive level of the classroom record and the knowledge point summary.

[0012] In an embodiment of the present invention, an electronic device is further provided, including: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, enabling the electronic device to implement the teaching case recommendation method described in any one of the above.

[0013] In an embodiment of the present invention, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed by a processor of a computer, the computer is enabled to execute the recommended method for teaching cases described in any one of the above.

[0014] As described above, a recommended method, system, device, and medium for teaching cases of the present invention have the following beneficial effects: By performing speech recognition and role recognition on classroom recordings, the speech content of teachers and students can be accurately distinguished, and this speech content is mapped to role-annotated text. Through a sequence annotation model and a text classification model, teacher questions are accurately extracted from the role-annotated text and classified according to their cognitive levels, and a quantitative analysis of the questions is performed from the structural and semantic perspectives. Through the sequence annotation model, the content of the teacher's questions can be semantically understood, and the questions raised by the teacher can be accurately identified. In addition, a corresponding knowledge point summary is extracted according to the teacher's speech content through a summary generation model. According to the knowledge point summary and the cognitive levels corresponding to each question raised by the teacher in the classroom recording, appropriate teaching cases are recommended. The present invention analyzes the teacher's lecture content in multiple dimensions through multi-dimensional methods such as speech recognition, role recognition, question recognition, and question level classification, so as to provide personalized teaching cases for teachers, greatly improving the quality of classroom teaching. It improves the problem in the prior art that it is difficult to provide personalized improvement suggestions for teachers due to the lack of analysis of the semantics, structure, and cognitive levels of teachers' questions. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a flowchart showing a recommended method for teaching cases provided by an embodiment of the present invention;

[0016] Figure 2 It is a block diagram showing the structure of a recommended system for teaching cases provided by an embodiment of the present invention;

[0017] Figure 3 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] The following specific examples illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0019] It should be noted that the illustrations provided in the following embodiments only schematically illustrate the basic concept of the present invention. Therefore, only the components related to the present invention are shown in the drawings, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and proportions of the components in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0020] In the following description, numerous details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid making the embodiments of the present invention difficult to understand.

[0021] The inventors have found that most of the existing classroom teaching quality assessment systems adopt the method of separately assessing teachers and students. These systems usually use video analysis technology based on neural networks to locate teachers and students, and classify classroom behaviors through face, expression, and behavior recognition algorithms. For example, students' behaviors can be classified into categories such as interaction, listening, raising hands, answering, reading and writing, etc.; teachers' behaviors include writing on the blackboard, interacting, patrolling, lecturing, etc. The system samples the classroom video periodically, records the behaviors of teachers and students, and generates analysis results based on this. The existing classroom teaching quality assessment systems have certain limitations. These systems often only evaluate based on the behavioral performances of teachers and students, and lack in-depth analysis of more detailed aspects in the teaching process, such as the quality of teachers' questions and students' answers. In addition, these systems usually fail to provide targeted improvement suggestions for teachers.

[0022] The present invention provides a method for recommending teaching cases. By performing speech recognition and role recognition on classroom recordings, the speech content of teachers and students can be accurately distinguished, and this speech content is mapped to role-annotated text. Through a sequence annotation model and a text classification model, teacher questions are accurately extracted from the role-annotated text and classified according to their cognitive levels. A quantitative analysis of the questions is carried out from the structural and semantic perspectives. Through the sequence annotation model, the content of the teacher's questions can be semantically understood, and the questions raised by the teacher can be accurately identified. In addition, a knowledge point summary is extracted according to the teacher's speech content through a summary generation model. According to the knowledge point summary and the cognitive levels corresponding to each question raised by the teacher in the classroom recording, appropriate teaching cases are recommended. The present invention analyzes the teacher's lecture content in multiple dimensions through multi-dimensional methods such as speech recognition, role recognition, question recognition, and question classification, so as to provide personalized teaching cases for teachers, greatly improving the quality of classroom teaching. The present invention conducts intelligent analysis on the teacher's classroom questions and provides targeted feedback suggestions to help the teacher improve the quality of classroom questions, thereby achieving the purpose of optimizing teaching effects. It improves the problem in the prior art that the lack of analysis of the semantics, structure, and cognitive levels of the teacher's questions makes it difficult to provide personalized improvement suggestions for teachers.

[0023] Please refer to Figure 1 , the method for recommending teaching cases includes the following steps:

[0024] S1. Obtain a classroom recording.

[0025] Audio and video recording devices (such as microphones and high-definition cameras, etc.) are installed in the classroom to capture the speech of teachers and students and classroom interactions, obtaining a classroom recording. The types of classroom recordings include, but are not limited to, audio recordings, video recordings, and comprehensive recordings formed by combining multi-channel audio and video. It can be understood that the classroom recording described in the present invention can be either a complete record of an entire class or a partial excerpt of the class. Those skilled in the art can adaptively select based on the actual teaching scenario requirements and are not limited herein. To ensure the accuracy of subsequent data processing, in an embodiment of the present invention, obtaining a classroom recording includes:

[0026] Obtain an initial classroom recording;

[0027] Perform audio noise reduction and echo cancellation processing on the initial classroom recording to obtain a final classroom recording.

[0028] Through the audio and video recording equipment installed in the classroom, information such as the speeches and interactions of teachers and students and the classroom environment is recorded, thereby obtaining an initial classroom record. However, considering that the initial classroom record may be affected by environmental noise and echoes, resulting in a decrease in the accuracy of subsequent audio processing. To improve this situation, audio noise reduction and echo cancellation need to be performed on the initial classroom record. Specifically, audio noise reduction can be performed through methods such as spectral subtraction, Wiener filtering, and noise estimation, so that the noise-reduced classroom record not only retains the speech clarity of the speaker but also reduces noise interference. Echo cancellation is performed on the noise-reduced classroom record through an echo model to make the classroom record clearer, thereby obtaining the finally processed classroom record.

[0029] S2. Perform speech recognition and role recognition on the classroom record respectively, obtain a transcription text and the role identifier of the speaker correspondingly, and map the role identifier to the transcription text to generate a role-annotated text.

[0030] After obtaining the classroom record, perform speech recognition on the classroom record to convert the audio into text information and obtain a transcription text. And perform role recognition on the classroom record, and determine the role of the speaker by analyzing information such as the speech frequency in the audio signal, where the role of the speaker includes but is not limited to teachers and students. Map the recognized role identifier to the corresponding statement of the transcription text to obtain a role-annotated text carrying the role identifier.

[0031] In an embodiment of the present invention, the performing speech recognition and role recognition on the classroom record respectively, obtaining a transcription text and the role identifier of the speaker correspondingly, and mapping the role identifier to the transcription text to generate a role-annotated text includes:

[0032] Perform speech recognition on the classroom record to obtain the transcription text corresponding to the classroom record;

[0033] Perform voice activity detection on the classroom record, and based on the pause intervals in the classroom record, segment the classroom record into multiple discourse units; wherein, each discourse unit belongs to a speech statement of the classroom record;

[0034] For each discourse unit: input the discourse unit into a speaker embedding recognition model to generate a voice embedding vector;

[0035] Based on a clustering algorithm, aggregate the generated voice embedding vectors into multiple role clusters;

[0036] Count the number of voice embedding vectors in each role cluster, and identify the role cluster with the largest number of voice embedding vectors as the teacher cluster, and identify the remaining at least one role cluster as the student cluster;

[0037] Based on the speech sentences to which the discourse units belong, map the role identifiers corresponding to the teacher cluster and the student cluster to the text sentences corresponding to the transcription text to generate a role-annotated text.

[0038] Input the classroom record into a pre-trained speech recognition model, convert the audio content of the classroom record into text information, and generate a transcription text corresponding to the entire classroom record. Among them, the speech recognition model includes but is not limited to models with an end-to-end Transformer architecture, DeepSpeech models, etc., as long as it can achieve speech-to-text recognition, which is not limited here. Perform voice activity detection (VAD) on the classroom record. By analyzing the pause intervals in the audio information, the complete audio stream of the entire classroom record is segmented into multiple discourse units, and each discourse unit corresponds to a speech sentence in the classroom record. For each discourse unit: Input the discourse unit into a pre-trained speaker embedding recognition model. The model extracts the speech features of the input discourse unit to generate a high-dimensional vector representation, which is used to characterize the voice features such as the timbre of the speaker, and use it as the speech embedding vector corresponding to the discourse unit for distinguishing different speakers in the subsequent clustering process. Among them, the speaker embedding recognition model includes but is not limited to models with a convolutional neural network, a recurrent neural network, or a Transformer architecture, as long as it can effectively extract speech features and perform speaker identification, which is not limited here.

[0039] After generating corresponding speech embedding vectors for all discourse units in the entire classroom record, through a clustering algorithm, grouping is performed according to the similarity of the speech embedding vectors, aggregating all speech embedding vectors into multiple role clusters, and the number of role clusters is the number of speakers in the classroom record. There is at least one speech embedding vector in each role cluster, and the speech embedding vectors in each role cluster do not overlap. Among them, the clustering algorithm includes but is not limited to K-means, DBSCAN (Density-Based Spatial Clustering of Applications with Noise), or hierarchical clustering, etc., as long as it can achieve aggregating speech embedding vectors into different role clusters, which is not limited here. For the speech embedding vectors in the same role cluster, since these speech embedding vectors have similar timbres, it means that the same speaker is speaking. Since the teacher's speaking frequency is relatively high in a class, and the number of times is greater than the number of times the students speak, the role cluster with the largest number of speech embedding vectors is marked as the teacher cluster, and the remaining at least one role cluster is marked as the student cluster. It can be understood that if there is only dialogue interaction between the teacher and students in the classroom, the role clusters can be divided into two types: student clusters and teacher clusters. If there are other roles (such as teaching assistants or listeners, etc.), in addition to student clusters and teacher clusters, there are other role clusters. According to the mapping relationship between the discourse unit and the relevant statements in the transcribed text, the corresponding role cluster labels are matched with the statements in the transcribed text, so as to form text with role identifiers.

[0040] S3. Extract the teacher's speech text from the role-annotated text, and input the teacher's speech text into the summary generation model to generate the corresponding knowledge point summary.

[0041] Analyze the role-annotated text, extract the teacher's speech content separately to obtain the teacher's speech text. Input the extracted teacher's speech text into a pre-trained summary generation model, process the input text through natural language processing methods, extract the key information and perform semantic analysis, so as to obtain the knowledge point summary of this class. Among them, the knowledge point summary refers to a general summary of the core knowledge points of this class. It can be understood that the teacher's speech text can be either all the speech content of the teacher in the entire classroom record or the preset number of speech contents of the teacher at the beginning stage, which is not limited here.

[0042] In an embodiment of the present invention, the summary generation model is a sequence annotation model, and the extracting the teacher's speech text from the role-annotated text and inputting the teacher's speech text into the summary generation model to generate the corresponding knowledge point summary includes:

[0043] Parse the role identifier from the role-annotated text, identify and extract the teacher's speech text;

[0044] Perform word segmentation on the teacher's speech text to obtain a word segmentation sequence corresponding to the teacher's speech text;

[0045] Input the word segmentation sequence into the abstract generation model to obtain the predicted category of each word segment;

[0046] Select the word segments corresponding to the knowledge points with the predicted category, and arrange them according to the order of the word segments in the teacher's speech text to generate a knowledge point abstract corresponding to the teacher's speech text.

[0047] Parse the role annotation text, identify and separate the teacher's speech content according to the role labels in the text, obtain all the text sentences with the role identifier "teacher", and arrange these text sentences in the order in which they appear in the role annotation text to form the teacher's speech text. Perform word segmentation on the teacher's speech text to divide the continuous text into multiple independent words, obtaining a word segmentation sequence corresponding to the teacher's speech text. Input the word segmentation sequence into a pre-trained abstract generation model. Since the abstract generation model is a sequence annotation model, perform binary classification for each word segment through the sequence annotation method to predict a category, where the predicted categories include knowledge points and non-knowledge points. Arrange the relevant word segments with the category of knowledge points according to their original order in the teacher's speech text to obtain the corresponding knowledge point abstract. Therefore, when the abstract generation model is a sequence annotation model, the generated knowledge point abstract is composed of directly extracting keywords from the original teacher's speech text. Exemplarily, the teacher's speech text is "T: Hello, everyone. Today we are going to learn about the history of the ancient Mesopotamia. The ancient Mesopotamia, also known as the Fertile Crescent, is a magical land that gave birth to a glorious civilization and left a magnificent chapter in the long river of human history...". After performing word segmentation on it and inputting it into the abstract generation model, through summarizing the knowledge points of the input text, the obtained knowledge point abstract is "the history of Mesopotamia". Through this method, the knowledge points involved in the classroom teaching content are obtained. It can be understood that if the abstract generation model is a sequence annotation model, the abstract generation model includes, but is not limited to, a fine-tuned BERT model, as long as it can effectively identify the key knowledge point abstract from the long text information.

[0048] In an embodiment of the present invention, the abstract generation model is a sequence-to-sequence model. The extracting the teacher's speech text from the role annotation text and inputting the teacher's speech text into the abstract generation model to generate a corresponding knowledge point abstract includes:

[0049] Parse the role identifier from the role annotation text, and identify and extract the teacher's speech text;

[0050] Perform word segmentation on the teacher's speech text to obtain a word segmentation sequence corresponding to the teacher's speech text;

[0051] Input the tokenized sequence into the encoder of the abstract generation model to extract the semantic features of the teacher's utterance text and generate a semantic vector sequence corresponding to the teacher's utterance text;

[0052] Input the semantic vector sequence into the decoder of the abstract generation model, and perform semantic recombination on each semantic vector based on the attention mechanism to obtain the knowledge point abstract corresponding to the teacher's utterance text.

[0053] Parse the role annotation text through natural language processing to extract the teacher's speech content and form the teacher's utterance text. Use a tokenizer to tokenize the teacher's utterance text, splitting the entire teacher's utterance text into independent words to obtain a tokenized sequence. Since the abstract generation model consists of an encoder and a decoder, input the tokenized sequence into the encoder of the abstract generation model. Adopt methods such as bidirectional LSTM or Transformer encoder to obtain a semantic vector sequence corresponding to the tokenized sequence by capturing the context information and deep semantic features in the tokenized sequence. Input the semantic vector sequence into the decoder. Based on a multi-layer LSTM or Transformer decoder, etc., use the attention mechanism to dynamically adjust the weight of each word output by the decoder, and select the words with higher weights and arrange them in sequence to form the knowledge point abstract. Therefore, when the abstract generation model is a sequence-to-sequence model, the generated knowledge point abstract is not just a repetition or simple extraction of the original text, but information reorganized and refined from the original content. It can be understood that if the abstract generation model is a sequence-to-sequence model, the abstract generation model includes but is not limited to fine-tuned BART models or T5 models, etc., as long as it can effectively identify the key knowledge point abstract from the long text information.

[0054] S4. Input the role annotation text into the question recognition model to extract the classroom questions raised by the teacher role; wherein, the question recognition model is a sequence annotation model.

[0055] Based on the natural language processing method, the role annotation text is segmented, and the role annotation text is divided into multiple independent tokens to obtain a structured segmentation sequence. As an example, the role annotation text can be segmented by BERT's WordPiece segmenter, and the entire text can be segmented by identifying spaces, punctuation marks, and special characters. For example, for the sentence "Is this a corner?", the segmenter can be split into 8 tokens ["this", "a", "is", "no", "is", "corner", "what", "?"]. Each token in the structured segmentation sequence is processed by word2vec, sentence-transformers and other pre-trained models to generate a word embedding vector corresponding to the token, and each word embedding vector is arranged in order to obtain a word embedding vector sequence corresponding to the entire role annotation text. The word embedding vector sequence is input into the pre-trained question recognition model, and a category label is assigned to each token using context information, wherein the category label represents the role of the current token in the entire sentence. The category label can be used to identify the specific questions raised by the teacher and the corresponding answers of the students. It should be noted that the question recognition model can be any model that can extract the question structure from text information, including but not limited to BERT (Bidirectional Encoder Representations from Transformers), recurrent neural network models, etc.

[0056] Exemplarily, the input of the question recognition model is a string, which is the aforementioned role-labeled text, such as "T: Let's take a look at this student's work. Is this a corner? S: No. T: Then why is this not a corner?" Among them, the sentence starting with T represents the teacher's words, and the sentence starting with S represents the student's words. After the input string is segmented into tokens one by one, it is input into a fine-tuned pre-trained language model, such as BERT. The output label corresponds one-to-one to the input token. The labels include: B-question, I-question, B-answer, I-answer, 0. They respectively represent: question starting point, question subordination, answer starting point, answer subordination, and others. For the above example, the expected output label of the model is "T(0): (0) Let's (0) take a look (0) at (0) this (0) classmate's (0) work (0) below (0). (0) Is this (B-question) (I-question) the (I-question) corner (I-question)? (I-question) S(B-answer): (I-answer) is (I-answer) not (I-answer). (I-answer) T(I-question): (I-question) Then (I-question) why is this (I-question) not (I-question) the (I-question) corner (I-question) (I-question)". From the above output, we can extract the following two classroom questions and corresponding answers: (1) Question: Is this an angle? Answer: No; (2) Why is this not an angle?

[0057] S5. Input the classroom questions into the problem hierarchy classification model to obtain the cognitive level corresponding to the classroom questions; wherein the problem hierarchy classification model is a text classification model.

[0058] Use a tokenization tool to split all the classroom questions raised by the teacher, obtaining that each classroom question is split into several independent tokens, and getting the classroom token sequence corresponding to the classroom question. Use a pre-trained language model (such as BERT) to perform word embedding processing on each classroom token sequence respectively, generating the word embedding vector sequence corresponding to each classroom question. Input the generated word embedding vector sequences into a pre-trained question recognition model, and through the forward propagation algorithm, obtain the probability that the word embedding vector sequence belongs to each preset cognitive level, and select the category label with the highest probability as the cognitive level label of the classroom question. Among them, the setting of the cognitive level can be adaptively set according to actual needs. In one embodiment, the cognitive level includes three cognitive level labels of fact and memory, application and analysis, and evaluation and creation with increasing cognitive levels in sequence. Exemplarily, input the classroom question "Is this an angle?" into the question level classification model. The final output of the model has three categories (such as: fact and memory, application and analysis, evaluation and creation) and the corresponding probability values for each category. The expected output of the model is: (0.99, 0.01, 0). Take the label with the largest probability as the category to which the question belongs, that is, "fact and memory". Further, to ensure the accuracy of the recognition result of the question level classification model, before inputting the classroom question into the question level classification model, it also includes data cleaning and standardization processing of the classroom question to improve the quality of the processed classroom question text. Among them, data cleaning includes but is not limited to removing stop words, removing irrelevant characters, etc., and standardization processing includes but is not limited to converting the text to lowercase uniformly, converting the text to simplified Chinese uniformly, etc.

[0059] S6. Based on each cognitive level of the classroom record and the knowledge point summary, select corresponding teaching cases for recommendation.

[0060] Classify all the classroom questions in the classroom record according to the cognitive level, and by calculating the similarity between the knowledge point summary and each knowledge point in the case library, select several knowledge points with higher similarity. Through the cognitive level to which the classroom question belongs, find the corresponding knowledge points from the selected knowledge points and recommend cases.

[0061] Specifically, in one embodiment of the present invention, the cognitive level includes a first cognitive level, a second cognitive level, and a third cognitive level with increasing levels in sequence. The selecting corresponding teaching cases for recommendation based on each cognitive level of the classroom record and the knowledge point summary includes:

[0062] Calculate the similarity between the knowledge point summary and each knowledge point in the case library, and based on the similarity, select the teaching case set pre-associated with the corresponding knowledge points; wherein, each teaching case is associated with a cognitive level;

[0063] Determine whether the proportion of the first cognitive level in all the questions in the classroom record is greater than or equal to a preset first threshold:

[0064] If so, select teaching cases associated with the second cognitive level and / or the third cognitive level from the teaching case set for recommendation;

[0065] If not, select corresponding teaching cases for recommendation based on the proportion of the second cognitive level in all the questions in the classroom record.

[0066] There are several knowledge points pre-stored in the case library, and at least one teaching case is associated with each knowledge point. This teaching case is used to represent the teaching method of the corresponding knowledge point in actual teaching. Among them, the teaching case is pre-associated with a specific cognitive level label, and each teaching case corresponds to a specific level (such as the first cognitive level, the second cognitive level, or the third cognitive level) to ensure that the recommended cases can meet the problem requirements of different levels in the classroom. Each teaching case includes, but is not limited to, relevant teaching content, teaching methods, teaching objectives, and related classroom questions, classroom interactions, student feedback, and other information. These teaching cases can provide teachers with specific teaching scenarios and strategies, helping teachers understand how to better explain a specific knowledge point and improve their classroom teaching effect. The knowledge points in the case library can be set in layers. By calculating the similarity between the knowledge point summary and each knowledge point at the bottom layer of the case library, the relevant cases are initially screened to obtain several knowledge points with a relatively high degree of association with this knowledge point, and the teaching cases associated with these knowledge points form a teaching case set.

[0067] Further, the knowledge point summary is converted into a summary vector representation through a pre-trained language model (such as BERT), each knowledge point in the case base is respectively converted into a knowledge point vector representation, and the similarity between the summary vector representation and each knowledge point vector representation is calculated. The calculation methods of similarity include, but are not limited to, cosine similarity and Euclidean distance, etc. The similarity is used to represent the semantic similarity between the knowledge point summary and the corresponding knowledge point. Select the relevant knowledge points with similarity greater than a preset threshold or the preset number of knowledge points with the highest similarity, and extract the teaching cases associated with these knowledge points to form a teaching case set. Count the number of various types of questions in the classroom record, and calculate the proportion of the first cognitive level (such as facts and memory) in all questions. When the proportion of the first cognitive level is greater than or equal to the first threshold (such as 60%), it indicates that the classroom questions mainly focus on the lower level and the interaction depth needs to be improved. Therefore, in order to improve the teaching quality, it is necessary to select teaching cases of the second cognitive level (such as application and analysis) and / or the third cognitive level (such as evaluation and creation) from the teaching case set obtained above, so as to encourage teachers to introduce higher-level questions in the classroom and promote students to think and analyze at a deeper level. Exemplarily, for the "History of the Two Rivers Basin" output above, through retrieval, the two most similar knowledge points "Ancient Two Rivers Basin" and "Development History of the Two Rivers Basin" existing in the teaching design database are found, and it can be known from the question level that the case corresponding to the "Ancient Two Rivers Basin" meets the requirements, and then the excellent teaching design cases under the knowledge point of "Ancient Two Rivers Basin" are returned for recommendation. On the contrary, if the proportion of the first cognitive level is less than the first threshold, it is necessary to further recommend corresponding teaching cases based on the proportion of the second cognitive level in all questions.

[0068] In an embodiment of the present invention, the selecting corresponding teaching cases for recommendation based on the proportion of the second cognitive level in the classroom record includes:

[0069] Determine whether the proportion of the second cognitive level in all questions is greater than or equal to a preset second threshold:

[0070] If so, select the teaching cases associated with the third cognitive level from the teaching case set for recommendation;

[0071] If not, recommend the teaching cases associated with the second cognitive level and the third cognitive level from the teaching case set for recommendation.

[0072] If the proportion of the second cognitive level is higher than or equal to a preset second threshold (e.g., 40%), teaching cases of the third cognitive level will be recommended to further improve the level of classroom questioning and promote higher-order evaluation and creative thinking. If the proportion of the second cognitive level is still lower than the second threshold, teaching cases of the second and third cognitive levels will be comprehensively recommended, aiming to comprehensively improve the quality of classroom questioning, ensure that teachers can find a balance in question design at different levels, and optimize the classroom interaction effect.

[0073] It should be noted that the case library of the present invention is built by collecting and sorting textbooks and establishing a knowledge point list covering various disciplines and educational stages in primary and secondary schools. Specifically, a large number of excellent teaching design cases in various disciplines and educational stages are collected, and the following field information is extracted from the unstructured case document data: discipline, grade, knowledge point, teaching objective, teaching process, etc., and is associated with the existing knowledge point list. The model for extracting the above information from the teaching design case is similar to the model used in "question extraction" and belongs to a sequence labeling model. The input text is the teaching design case text, and the output labels include: B-discipline, I-discipline, B-grade, I-grade, B-title, I-title, B-objective, I-objective, B-process, I-process, O, which represent the following labels respectively: discipline, grade, knowledge point, teaching objective, teaching process. Thus, the case library is constructed. Further, it should be noted that the present invention can also perform statistical analysis on classroom interaction data. This includes the number of questions asked by the teacher, the type of questions asked each time, the students' answering situations, etc., and the above information is formed into a feedback report and fed back to the teacher, and at the same time, the retrieved excellent teaching cases are displayed.

[0074] Please refer to Figure 2, the recommendation system 100 for teaching cases includes: a classroom record acquisition module 110, a role annotation module 120, a knowledge point identification module 130, a question identification module 140, a cognitive level identification module 150, and a case recommendation module 160. The above-mentioned classroom record acquisition module 110 is used to acquire classroom records. The role annotation module 120 is used to perform speech recognition and role recognition on the classroom record respectively, obtain the corresponding transcribed text and the role identifier of the speaker, and map the role identifier to the transcribed text to generate a role-annotated text. The knowledge point identification module 130 is used to extract the teacher's speech text from the role-annotated text, and input the teacher's speech text into a summary generation model to generate a corresponding knowledge point summary. The question identification module 140 is used to input the role-annotated text into a question identification model to extract the classroom questions raised by the teacher role; wherein, the question identification model is a sequence annotation model. The cognitive level identification module 150 is used to input the classroom questions into a question level classification model to obtain the corresponding cognitive level; wherein, the question level classification model is a text classification model. The case recommendation module 160 is used to select corresponding teaching cases for recommendation based on the respective cognitive levels of the classroom record and the knowledge point summary.

[0075] For the specific limitations of the recommendation system for teaching cases, reference can be made to the limitations of the recommendation method for teaching cases in the above text, which will not be elaborated here. Each module in the above-mentioned recommendation system for teaching cases can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in or independent of the processor in the computer device in a hardware format, or stored in the memory in the computer device in a software format, so as to facilitate the processor to call the corresponding operations of each of the above modules.

[0076] It should be noted that, in order to highlight the innovative part of the present invention, modules that are not closely related to solving the technical problems proposed by the present invention are not introduced in this embodiment, but this does not mean that there are no other modules in this embodiment.

[0077] Please refer to Figure 3 , the electronic device 1 may include a memory 12, a processor 13, and a bus, and may also include a computer program stored in the memory 12 and executable on the processor 13, such as a recommendation program for teaching cases.

[0078] Among them, the memory 12 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disk, etc. The memory 12 can be an internal storage unit of the electronic device 1 in some embodiments, such as the mobile hard disk of the electronic device 1. The memory 12 can also be an external storage device of the electronic device 1 in other embodiments, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 1. Further, the memory 12 can also include both the internal storage unit and the external storage device of the electronic device 1. The memory 12 can be used not only to store application software installed on the electronic device 1 and various types of data, such as the code recommended for teaching cases, etc., but also to temporarily store data that has been output or will be output.

[0079] The processor 13 can be composed of integrated circuits in some embodiments. For example, it can be composed of a single packaged integrated circuit, or can be composed of multiple integrated circuits with the same or different functions packaged, including the combination of one or more Central Processing Units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips, etc. The processor 13 is the control core (Control Unit) of the electronic device 1, connecting various components of the entire electronic device 1 through various interfaces and lines, and by running or executing programs or modules stored in the memory 12 (such as the recommended program for teaching cases, etc.), and calling the data stored in the memory 12, to execute various functions of the electronic device 1 and process data.

[0080] The processor 13 executes the operating system of the electronic device 1 and various installed application programs. The processor 13 executes the application program to implement the steps in the above-mentioned recommended method for teaching cases.

[0081] Exemplarily, the computer program can be divided into one or more modules, and the one or more modules are stored in the memory 12 and executed by the processor 13 to complete this application. The one or more modules can be a series of computer program instruction segments capable of completing specific functions, and this instruction segment is used to describe the execution process of the computer program in the electronic device 1. For example, the computer program can be divided into a classroom record acquisition module 110, a role annotation module 120, a knowledge point identification module 130, a question identification module 140, a cognitive level identification module 150, and a case recommendation module 160.

[0082] The integrated unit implemented in the form of software functional modules can be stored in a computer-readable storage medium, which can be non-volatile or volatile. The above software functional modules are stored in a storage medium and include several instructions for causing a computer device (which can be a personal computer, a computer device, or a network device, etc.) or a processor to execute some functions of the teaching case recommendation method described in each embodiment of the present application.

[0083] In summary, for a teaching case recommendation method, system, device, and medium disclosed by the present invention, speech recognition and role recognition are performed on the classroom record to accurately distinguish the speech content of teachers and students, and map this speech content to role-annotated text. Through a sequence annotation model and a text classification model, teacher questions are accurately extracted from the role-annotated text and classified according to their cognitive levels, and a quantitative analysis of the questions is performed from the structural and semantic perspectives. Through the sequence annotation model, the content of the teacher's questions can be semantically understood, and the questions raised by the teacher can be accurately identified. In addition, a knowledge point summary is extracted according to the teacher's speech content through a summary generation model. According to the knowledge point summary and the cognitive levels corresponding to each question raised by the teacher in the classroom record, appropriate teaching cases are recommended. The present invention performs multi-dimensional analysis on the teacher's lecture content through multi-dimensional methods such as speech recognition, role recognition, question recognition, and question level classification, so as to provide personalized teaching cases for teachers, greatly improving the quality of classroom teaching. It improves the problem that the prior art lacks the analysis of the semantics, structure, and cognitive levels of teachers' questions, making it difficult to provide personalized improvement suggestions for teachers. Therefore, the present invention effectively overcomes various disadvantages in the prior art and has high industrial utilization value.

[0084] The above embodiments are only illustrative of the principles and effects of the present invention, and are not used to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.

Claims

1. A method for recommending teaching cases, characterized in that: The method comprises: Get classroom transcripts; Performing speech recognition and role recognition on the classroom transcript respectively, obtaining a transcription text and a speaker's role identification correspondingly, and mapping the role identification to the transcription text to generate a role-labeled text; Extracting the teacher's speech text from the role annotation text, and inputting the teacher's speech text into the summary generation model to generate a corresponding knowledge point summary; Input the role-annotated text into a question recognition model to extract classroom questions raised by the teacher role; wherein the question recognition model is a sequence labeling model; Input the classroom questions into the problem hierarchy classification model to obtain the cognitive level corresponding to the classroom questions; wherein the problem hierarchy classification model is a text classification model; Based on the various cognitive levels of the classroom transcript and the summary of the knowledge points, corresponding teaching cases are selected for recommendation.

2. The method for recommending teaching cases according to claim 1, characterized in that: The acquisition of classroom records includes: Obtain initial classroom transcripts; The initial classroom recording is subjected to audio noise reduction and echo cancellation processing to obtain a final classroom recording.

3. The method for recommending teaching cases according to claim 1, characterized in that: The method of performing speech recognition and role recognition on the classroom transcript to obtain a transcription text and a speaker's role identifier, and mapping the role identifier to the transcription text to generate a role-labeled text includes: Performing speech recognition on the classroom record to obtain a transcribed text corresponding to the classroom record; Performing voice activity detection on the classroom record, and dividing the classroom record into a plurality of speech units based on pause intervals in the classroom record; wherein each speech unit belongs to a speech sentence in the classroom record; For each speech unit: input the speech unit into the speaker embedding recognition model to generate a speech embedding vector; Based on the clustering algorithm, the generated speech embedding vectors are aggregated into multiple role clusters; Counting the number of speech embedding vectors in each role cluster, identifying the role cluster with the largest number of speech embedding vectors as a teacher cluster, and identifying at least one remaining role cluster as a student cluster; Based on the speech sentences to which the discourse units belong, the role identifiers corresponding to the teacher cluster and the student cluster are mapped to the text sentences corresponding to the transcription text to generate role-labeled text.

4. The method for recommending teaching cases according to claim 1, characterized in that: The summary generation model is a sequence labeling model, and the teacher's speech text is extracted from the role labeling text, and the teacher's speech text is input into the summary generation model to generate a corresponding knowledge point summary, including: Parsing the role identifier from the role annotation text, identifying and extracting the teacher's speech text; Performing word segmentation processing on the teacher's speech text to obtain a word segmentation sequence corresponding to the teacher's speech text; Inputting the word segmentation sequence into the summary generation model to obtain a predicted category of each word segmentation; The corresponding participles whose predicted categories are knowledge points are selected, and the participles are arranged in order in the teacher's speech text to generate a knowledge point summary corresponding to the teacher's speech text.

5. The method for recommending teaching cases according to claim 1, characterized in that: The summary generation model is a sequence-to-sequence model, and the teacher's speech text is extracted from the role annotation text, and the teacher's speech text is input into the summary generation model to generate a corresponding knowledge point summary, including: Parsing the role identifier from the role annotation text, identifying and extracting the teacher's speech text; Performing word segmentation processing on the teacher's speech text to obtain a word segmentation sequence corresponding to the teacher's speech text; Inputting the word segmentation sequence into the encoder of the summary generation model, extracting the semantic features of the teacher's speech text, and generating a semantic vector sequence corresponding to the teacher's speech text; The semantic vector sequence is input into the decoder of the summary generation model, and each semantic vector is semantically reorganized based on the attention mechanism to obtain the knowledge point summary corresponding to the teacher's speech text.

6. The method for recommending teaching cases according to claim 1, characterized in that: The cognitive levels include a first cognitive level, a second cognitive level, and a third cognitive level in ascending order. The corresponding teaching cases are selected for recommendation based on the cognitive levels of the classroom records and the knowledge point summary, including: Calculating the similarity between the knowledge point summary and each knowledge point in the case library, and selecting a teaching case set pre-associated with the corresponding knowledge point based on the similarity; wherein each teaching case is associated with a cognitive level; Determine whether the proportion of the first cognitive level in the classroom record to all questions in the classroom record is greater than or equal to a preset first threshold: If so, selecting teaching cases at the second cognitive level and / or the third cognitive level from the teaching case collection for recommendation; If not, then based on the proportion of the second cognitive level in all questions in the classroom transcript, select corresponding teaching cases for recommendation.

7. The method for recommending teaching cases according to claim 6, characterized in that: Based on the proportion of the second cognitive level in all questions in the classroom transcript, corresponding teaching cases are selected for recommendation, including: Determine whether the proportion of the second cognitive level to all questions is greater than or equal to the preset second threshold: If so, selecting teaching cases related to the third cognitive level from the teaching case collection for recommendation; If not, then recommend teaching cases related to the second cognitive level and the third cognitive level from the teaching case collection.

8. A teaching case recommendation system, characterized in that: The system comprises: Classroom record acquisition module, used to obtain classroom records; A role labeling module is used to perform speech recognition and role recognition on the classroom transcript, obtain a transcription text and a speaker's role identification, and map the role identification to the transcription text to generate a role labeling text; A knowledge point recognition module is used to extract the teacher's speech text from the role annotation text, and input the teacher's speech text into the summary generation model to generate a corresponding knowledge point summary; A question recognition module is used to input the role-annotated text into a question recognition model to extract classroom questions raised by the teacher role; wherein the question recognition model is a sequence labeling model; A cognitive level identification module is used to input classroom questions into a problem level classification model to obtain a corresponding cognitive level; wherein the problem level classification model is a text classification model; The case recommendation module is used to select corresponding teaching cases for recommendation based on the various cognitive levels of the classroom records and the knowledge point summary.

9. An electronic device, characterized in that: The electronic device comprises: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, enables the electronic device to implement the recommendation method of the teaching case as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor of a computer, the computer is caused to execute the method for recommending teaching cases as described in any one of claims 1 to 7.