Teacher sensitive speech detection method in classroom teaching scene

By constructing a teacher's teaching speech behavior classification network and combining it with a graph convolutional neural network and a multi-head attention reconstruction graph structure neural network, the accuracy and intelligence problems of teacher sensitive speech detection in the classroom are solved, and efficient automatic recognition and classification of teacher speech behavior are achieved, thereby improving teaching quality.

CN120744112APending Publication Date: 2025-10-03SHAANXI NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510849720.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing technologies make it difficult to accurately identify and monitor teachers' sensitive verbal behaviors in the classroom. Traditional detection methods have low intelligence levels and high false alarm rates, and cannot adapt to real classroom contexts.

Method used

A teacher teaching speech behavior classification network is constructed, and a graph convolutional neural network and a multi-head attention reconstruction graph structure neural network are used. The output of the graph convolutional neural network is combined with the output of the multi-head attention reconstruction graph structure neural network. The text graph structure is constructed through point-by-point mutual information and word frequency-inverse document frequency algorithm. The teacher teaching speech behavior classification network is trained, and sensitive speech is judged using mathematical statistical methods.

Benefits of technology

The classification accuracy of teacher-sensitive speech detection and the generalization ability of the network model have been improved, and the automatic recognition and sub-categorization of teachers' classroom speech behaviors have been realized, reducing manual intervention and optimizing teaching quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744112A_ABST
    Figure CN120744112A_ABST
Patent Text Reader

Abstract

The invention discloses a teacher sensitive speech detection method in a classroom teaching scene. The method comprises the following steps: constructing a teacher classroom speech behavior data set; constructing a teacher teaching speech behavior classification network; constructing a text graph structure; training a teacher teaching speech behavior classification network; detecting a teacher teaching speech behavior classification network by using the test set; and performing speech sensitivity judgment on the classroom of the teacher according to a classification result. According to the classification network for detecting the teaching speech behaviors of the teacher, through a graph reconstruction layer and a multi-head attention mechanism, the graph structure can be dynamically adjusted, the complex relation between nodes can be captured, and the semantic understanding of the classroom speech behaviors of the teacher can be enhanced, so that the speech behaviors of the teacher can be identified and classified more accurately; an objective basis is provided for quantitative evaluation of the teaching ability of the teacher, a good classroom atmosphere can be constructed, the teaching method is optimized, and the teaching quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the intersection of artificial intelligence and pedagogy, and specifically relates to a method for detecting teacher-sensitive speech in classroom teaching scenarios. Background Art

[0002] The classroom is the primary venue for teachers to carry out their teaching and educating responsibilities. Teachers' classroom teaching behaviors have a profound impact on students' knowledge acquisition, ability development, and mental health. Good teacher verbal behavior not only helps stimulate students' interest in learning and participation, but also optimizes teaching structure and improves teaching effectiveness. However, sensitive language such as sarcasm, derogatory language, discrimination, and threats may still appear in the classroom, causing damage to students' self-esteem, emotional depression, and even long-term psychological trauma. Currently, there is a lack of effective means to identify and regulate sensitive language. Manual class observation and evaluation have limited coverage and are difficult to capture individual teachers' abnormal verbal behavior in a timely manner. At the same time, teachers themselves lack sensitivity to non-standard language, making it difficult to detect and improve it in a timely manner. Traditional keyword detection methods have low intelligence levels and high false positive rates. They have difficulty understanding complex semantics and emotional expressions and are unable to adapt to real classroom contexts.

[0003] Against the backdrop of the continuous development of information technology and artificial intelligence, the field of education is accelerating towards a new stage of intelligence and personalization. Using computer technologies such as natural language processing to automatically analyze teacher speech texts in the classroom and implement behavior classification and feedback has become a viable path to improving teaching quality and efficiency. However, there is currently a lack of standardized teacher behavior classification systems and high-quality annotated datasets. Research on the automated identification of sensitive speech behaviors in teacher texts is still in its infancy, making it difficult to support large-scale teaching behavior modeling and application. Research and practical exploration of relevant technical means are urgently needed. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a method for detecting teacher-sensitive speech in classroom teaching scenarios with high classification accuracy and strong network model generalization ability.

[0005] The technical solution adopted to solve the above technical problems is: a method for detecting teacher sensitive speech in a classroom teaching scenario, comprising the following steps:

[0006] Step 1. Construct a dataset of teacher classroom speech behaviors

[0007] In a real classroom environment, teachers' lecture audio is collected and transcribed into text. The text is segmented by sentence and annotated with category labels to form a dataset of teachers' classroom speech behavior. The dataset is divided into training set, test set, and validation set according to a certain ratio.

[0008] Step 2. Construct a network to classify teachers’ teaching speech behaviors

[0009] The teacher teaching speech behavior classification network includes a graph convolutional neural network, a multi-head attention-based graph structure reconstruction neural network, and an adder; the graph convolutional neural network is used to extract node features from graph structure data and capture the relationship between nodes; the multi-head attention-based graph structure reconstruction neural network is used to dynamically reconstruct the graph structure through a multi-head attention mechanism to enhance the network's ability to capture complex relationships between nodes; the adder is used to fuse the output of the graph convolutional neural network with the output of the multi-head attention-based graph structure reconstruction neural network;

[0010] Step 3. Build the text graph structure

[0011] The training set, test set, and validation set are each converted into a text graph structure as the input of the teacher's teaching speech behavior classification network. In the text graph structure, the document node represents the entire sentence or document, the word node represents each word in the sentence, and the edge represents the semantic association of the word and the attribution relationship between the word node and the document node. All word nodes are traversed and the edge weights between word nodes are calculated using the point-by-point mutual information algorithm. The edge weights between word nodes and document nodes are calculated using the word frequency-inverse document frequency algorithm.

[0012] Step 4. Train the teacher’s teaching speech act classification network

[0013] Initialize the parameters of the teacher teaching speech behavior classification network using the Xavier method and set the super parameters. Input the training set into the teacher teaching speech behavior classification network, perform forward propagation, and determine the loss function as the cross entropy loss L. Use the Adam optimizer to minimize the loss, and update all parameters until the loss function converges. During each round of training, calculate the confidence value of each word node in the text graph structure, mark the word nodes with confidence values ​​higher than the threshold as keywords, add them to the text graph structure, update the edge weights between word nodes and the edge weights between word nodes and document nodes, and obtain the updated text graph structure as the input for the next round of training. After the training is completed, the trained teacher teaching speech behavior classification network is obtained.

[0014] Step 5. Use the test set to test the teacher's teaching speech behavior classification network

[0015] The validation set is input into the trained teacher teaching speech behavior classification network, and the parameters of the teacher teaching speech behavior classification network are adjusted to obtain the optimal teacher teaching speech behavior classification network. The test set is then input into the teacher teaching speech behavior classification network and the classification results are output.

[0016] Step 6. Determine the teacher's classroom speech sensitivity based on the classification results

[0017] The classification results were analyzed using mathematical statistical methods, and teachers' classroom speech behaviors were classified into sensitive speech classrooms and non-sensitive speech classrooms according to whether they contained sensitive speech categories; sensitive speech classrooms were divided into low-sensitivity speech classrooms and high-sensitivity speech classrooms according to whether they contained verbal violence categories.

[0018] Preferably, in step 2, the multi-head attention-based graph structure reconstruction neural network is composed of a parallel graph reconstruction layer and a graph embedding reconstruction module and a Concat connection layer, a fully connected layer, a Dropout layer, and a graph embedding attention module connected in series in sequence; the graph reconstruction layer is used to reconstruct the adjacency matrix of the graph and output the reconstructed adjacency matrix and node feature matrix; the graph embedding reconstruction module is used to further extract the feature representation of the node, enhance the semantic information of the node, and output the reconstructed node feature representation; the Concat connection layer is used to fuse the output of the graph reconstruction layer and the output of the graph embedding reconstruction module; the fully connected layer is used for feature dimensionality reduction; the Dropout layer is used to randomly discard some features to prevent model overfitting; the graph embedding attention module is used to further extract and optimize the feature representation of the node through the attention mechanism, and enhance the semantic information and structural information of the node.

[0019] Preferably, the confidence value of the word node is:

[0020]

[0021] Where, is the word node w i The confidence value, num is the category label, p num To mark the document containing the word node w i The number of documents in the numth category, where n is the number of categories.

[0022] Preferably, the threshold is 0.9.

[0023] Preferably, the super parameters include: batch size of 64, learning rate of 0.01, point-by-point mutual information window of 10, dropout rate of 0.5 in two-layer graph convolutional neural network, dropout rate of 0.55 in graph reconstruction neural network, K of KNN algorithm of 5, and minimum threshold of 0.95.

[0024] The beneficial effects of the present invention are as follows:

[0025] The classification network for detecting teachers' teaching speech behaviors constructed by the present invention can dynamically adjust the graph structure through the graph reconstruction layer and multi-head attention mechanism, capture the complex relationships between nodes, enhance the semantic understanding of teachers' classroom speech behaviors, and thus more accurately identify and classify teachers' speech behaviors.

[0026] The classification network for detecting teacher teaching speech behavior constructed by the present invention fuses the output of the graph convolutional neural network with the output of the neural network based on multi-head attention reconstruction graph structure, making full use of the advantages of the two networks, further improving the feature expression ability, and thus improving the accuracy of classification.

[0027] This invention can automatically identify and categorize teachers' classroom speech behaviors, enabling automatic detection of sensitive speech behaviors and their subcategories, improving classification accuracy and analysis efficiency while reducing manual intervention. It provides an objective basis for the quantitative evaluation of teachers' teaching abilities, helping to foster a positive classroom atmosphere, optimize teaching methods, and enhance teaching quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 It is a flow chart of the teacher's sensitive speech detection method in the classroom teaching scenario of the present invention.

[0029] Figure 2 This is a diagram of the teacher’s classroom speech behavior dataset.

[0030] Figure 3 It is a schematic diagram of the network structure for classifying teachers’ teaching speech behaviors.

[0031] Figure 4 This is a schematic diagram of the graph embedding attention module structure. DETAILED DESCRIPTION

[0032] The present invention will be further described in detail below with reference to the accompanying drawings and examples, but the present invention is not limited to the following embodiments.

[0033] exist Figure 1 In this embodiment, a method for detecting teacher sensitive speech in a classroom teaching scenario includes the following steps:

[0034] Step 1. Construct a dataset of teacher classroom speech behaviors

[0035] In a real classroom environment, teachers' lecture audio is collected and transcribed into text. The text is segmented by sentence and annotated with category labels to form a dataset of teachers' classroom speech behavior. The dataset is divided into training set, test set, and validation set according to a certain ratio.

[0036] The category labels include: classroom teaching, guiding questions, feedback evaluation, instruction management, and sensitive language. The classroom teaching category includes the teacher's behavior of systematically imparting knowledge to students in the classroom. The content imparted includes information such as the teacher's personal understanding, authoritative opinions, and objective facts.

[0037] Guided questioning involves teachers asking questions to guide students to participate in classroom interactions and stimulate their thinking. This includes asking questions to individual students, asking questions to the whole class, and teachers participating in student discussions.

[0038] The feedback evaluation category includes teachers’ behaviors of evaluating students’ performance during the teaching process, including positive or negative feedback on students’ answers, as well as motivational evaluations such as praise for students’ behaviors;

[0039] The directive management category includes the teacher's management behaviors in the classroom that are not related to the teaching content, mainly including maintaining classroom order, organizing students into groups, and issuing disciplinary instructions;

[0040] Sensitive language is divided into non-normative language and violent language. Non-normative language includes all speech that is inappropriate for the classroom, specifically the use of inappropriate humor and internet memes to liven up the atmosphere.

[0041] The verbal violence category in the sensitive language category includes all discriminatory, insulting and mocking teaching language behaviors used by teachers against students in the classroom, such as Figure 2 .

[0042] Step 2. Construct a network to classify teachers’ teaching speech behaviors

[0043] like Figure 3 The teacher teaching speech behavior classification network includes a graph convolutional neural network, a graph structure neural network based on multi-head attention reconstruction, and an adder; the graph convolutional neural network is used to extract node features from graph structure data and capture the relationship between nodes; the graph structure neural network based on multi-head attention reconstruction is used to dynamically reconstruct the graph structure through a multi-head attention mechanism to enhance the network's ability to capture complex relationships between nodes; the adder is used to fuse the output of the graph convolutional neural network with the output of the graph structure neural network based on multi-head attention reconstruction;

[0044] The multi-head attention-based graph structure reconstruction neural network consists of a parallel graph reconstruction layer and a graph embedding reconstruction module, connected in series with a concat layer, a fully connected layer, a dropout layer, and a graph embedding attention module. The graph reconstruction layer reconstructs the graph's adjacency matrix and outputs a reconstructed adjacency matrix and a node feature matrix. The graph embedding reconstruction module further extracts node feature representations, enhances node semantic information, and outputs reconstructed node feature representations. The concat layer fuses the output of the graph reconstruction layer with the output of the graph embedding reconstruction module. The fully connected layer performs feature dimensionality reduction. The dropout layer randomly discards some features to prevent model overfitting. The graph embedding attention module further extracts and optimizes node feature representations through an attention mechanism, enhancing the semantic and structural information of nodes.

[0045] like Figure 4The graph embedding attention module is composed of a multi-head attention layer, a first Dropout layer, a first residual connection layer, a second Dropout layer, and a second residual connection layer connected in sequence. The multi-head attention layer is used to calculate the attention scores between nodes and capture the complex relationships between nodes. The first Dropout layer is used to randomly discard some features to prevent overfitting; the first residual connection layer is used to add the input features to the processed features to enhance the feature transfer and avoid gradient disappearance; the second Dropout layer is used to randomly discard some features again to further prevent overfitting; the second residual connection layer adds the input features to the processed features again to further enhance the feature transfer.

[0046] Step 3. Build the text graph structure

[0047] The training set, test set, and validation set are each converted into a text graph structure as the input of the teacher's teaching speech behavior classification network. In the text graph structure, the document node represents the entire sentence or document, the word node represents each word in the sentence, and the edge represents the word semantic association and the attribution relationship between the word node and the document node; all word nodes are traversed, and the point-by-point mutual information algorithm is used to calculate the edge weights between word nodes, and the word frequency-inverse document frequency algorithm is used to calculate the edge weights between word nodes and document nodes.

[0048] Step 4. Train the teacher’s teaching speech act classification network

[0049] The Xavier method was used to initialize the parameters of the teacher's teaching speech behavior classification network, and the batch size was set to 64, the learning rate was 0.01, the point-by-point mutual information window was 10, the dropout rate in the two-layer graph convolutional neural network was 0.5, the dropout rate in the graph reconstruction neural network was 0.55, the K in the KNN algorithm was 5, and the minimum threshold was 0.95.

[0050] The training set is input into the teacher teaching speech behavior classification network, and the forward propagation is performed to determine the loss function as the cross entropy loss L. The Adam optimizer is used to minimize the loss, and all parameters are updated until the loss function converges. During each round of training, the confidence value of each word node in the text graph structure is calculated. Word nodes with a confidence value higher than a threshold of 0.9 are marked as keywords and added to the text graph structure. The edge weights between word nodes and the edge weights between word nodes and document nodes are updated to obtain the updated text graph structure as the input for the next round of training. At the end of training, the trained teacher teaching speech behavior classification network is obtained.

[0051] Among them, the confidence value of the word node is:

[0052]

[0053] Where, is the word node w i The confidence value, num is the category label, p num To mark the document containing the word node w i The number of documents in the numth category, where n is the number of categories.

[0054] Step 5. Use the test set to test the teacher's teaching speech behavior classification network

[0055] The validation set is input into the trained teacher teaching speech behavior classification network, and the parameters of the teacher teaching speech behavior classification network are adjusted to obtain the optimal teacher teaching speech behavior classification network. The test set is then input into the teacher teaching speech behavior classification network, and the teacher teaching behavior classification results are output.

[0056] Step 6. Determine the teacher's classroom speech sensitivity based on the classification results

[0057] The classification results were analyzed using mathematical statistical methods, and teachers' classroom speech behaviors were classified into sensitive speech classrooms and non-sensitive speech classrooms according to whether they contained sensitive speech categories; sensitive speech classrooms were divided into low-sensitivity speech classrooms and high-sensitivity speech classrooms according to whether they contained verbal violence categories.

[0058] In order to verify the beneficial effects of the teacher teaching speech behavior classification network of the present invention, the inventors conducted the following comparative experiments:

[0059] We conducted comparative experiments with “Kim Y.Convolutional neural networks for sentence classification[J].arXiv preprint arXiv:1408.5882,2014.” (referred to as Comparison 1), “Yao, Liang, Chengsheng Mao, and Yuan Luo."Graph convolutional networks for text classification."Proceedings of the AAAI conference on artificial intelligence.Vol.33.No.01.2019." (referred to as Comparison 2), and “Jing, Rongrong, et al."Self-Training based semi-Supervised and semi-Paired hashing cross-modal retrieval."2022International Joint Conference on Neural Networks(IJCNN).IEEE,2022." (referred to as Comparison 3). The classification results were evaluated using test accuracy and mean macro-F1. The experimental results are shown in Table 1.

[0060] Table 1 Experimental results of the present invention and comparative experiments

[0061] Experimental group Test accuracy Macro-F1 Comparison plan 1 0.6714 0.6023 Comparison Plan 2 0.6882 0.6525 Comparison Plan 3 0.7 0.6689 The present invention 0.7184 0.6948

[0062] As shown in Table 1, compared with Comparative Experiments 1-3, the network for classifying teacher teaching speech behaviors of the present invention has the highest test accuracy and Macro-F1 value, and the scores of various indicators are significantly improved. The test accuracy and Macro-F1 value of the network for classifying teacher teaching speech behaviors of Example 1 are 4.7% and 9.25% higher than those of Comparative Experiment 1, 3.02% and 4.23% higher than those of Comparative Experiment 2, and 1.84% and 2.59% higher than those of Comparative Experiment 3.

[0063] The above experiments show that the scores of various indicators of the teacher teaching speech behavior classification network proposed in the present invention are better than those of the comparison scheme, and the best effect is achieved on the self-built teacher teaching behavior dataset, which can accurately classify teacher teaching behaviors.

Claims

1. A method for detecting teacher sensitive speech in classroom teaching scenarios, characterized in that: The following steps are involved: Step 1. Construct a dataset of teacher classroom speech behaviors In a real classroom environment, teachers' lecture audio is collected and transcribed into text. The text is segmented by sentence and annotated with category labels to form a dataset of teachers' classroom speech behavior. The dataset is divided into training set, test set, and validation set according to a certain ratio. Step 2. Construct a network to classify teachers’ teaching speech behaviors The teacher teaching speech behavior classification network includes a graph convolutional neural network, a multi-head attention-based graph structure reconstruction neural network, and an adder; the graph convolutional neural network is used to extract node features from graph structure data and capture the relationship between nodes; the multi-head attention-based graph structure reconstruction neural network is used to dynamically reconstruct the graph structure through a multi-head attention mechanism to enhance the network's ability to capture complex relationships between nodes; the adder is used to fuse the output of the graph convolutional neural network with the output of the multi-head attention-based graph structure reconstruction neural network; Step 3. Build the text graph structure The training set, test set, and validation set are each converted into a text graph structure as the input of the teacher's teaching speech behavior classification network. In the text graph structure, the document node represents the entire sentence or document, the word node represents each word in the sentence, and the edge represents the semantic association of the word and the attribution relationship between the word node and the document node. All word nodes are traversed and the edge weights between word nodes are calculated using the point-by-point mutual information algorithm. The edge weights between word nodes and document nodes are calculated using the word frequency-inverse document frequency algorithm. Step 4. Train the teacher’s teaching speech act classification network Initialize the parameters of the teacher teaching speech behavior classification network using the Xavier method and set hyperparameters. Input the training set into the teacher teaching speech behavior classification network. Forward propagation is performed and the loss function is determined to be the cross-entropy loss L. The Adam optimizer is used to minimize the loss, and all parameters are updated until the loss function converges. During each round of training, the confidence value of each word node in the text graph structure is calculated. Word nodes with confidence values ​​above the threshold are marked as keywords and added to the text graph structure. The edge weights between word nodes and the edge weights between word nodes and document nodes are updated to obtain the updated text graph structure as input for the next round of training. After the training is completed, a trained teacher's teaching speech behavior classification network is obtained; Step 5. Use the test set to test the teacher's teaching speech behavior classification network The validation set is input into the trained teacher teaching speech behavior classification network, and the parameters of the teacher teaching speech behavior classification network are adjusted to obtain the optimal teacher teaching speech behavior classification network. The test set is then input into the teacher teaching speech behavior classification network and the classification results are output. Step 6. Determine the teacher's classroom speech sensitivity based on the classification results Mathematical statistical methods were used to analyze the classification results, and teachers' classroom speech behaviors were classified into sensitive speech classrooms and non-sensitive speech classrooms according to whether they contained sensitive speech categories. Sensitive language classes are divided into low-sensitivity language classes and high-sensitivity language classes based on whether they contain verbal violence.

2. The method for detecting teacher sensitive speech in classroom teaching scenarios according to claim 1 is characterized in that: In the step 2, the multi-head attention-based graph structure reconstruction neural network is composed of a parallel graph reconstruction layer and a graph embedding reconstruction module, and a Concat connection layer, a fully connected layer, a Dropout layer, and a graph embedding attention module connected in series in sequence; the graph reconstruction layer is used to reconstruct the adjacency matrix of the graph and output the reconstructed adjacency matrix and node feature matrix; the graph embedding reconstruction module is used to further extract the feature representation of the node, enhance the semantic information of the node, and output the reconstructed node feature representation; the Concat connection layer is used to fuse the output of the graph reconstruction layer and the output of the graph embedding reconstruction module; the fully connected layer is used for feature dimensionality reduction; the Dropout layer is used to randomly discard some features to prevent model overfitting; the graph embedding attention module is used to further extract and optimize the feature representation of the node through the attention mechanism, and enhance the semantic information and structural information of the node.

3. The method for detecting teacher sensitive speech in classroom teaching scenarios according to claim 1 is characterized in that: The confidence value of the word node is: Where, is the word node w i The confidence value, num is the category label, p num To mark the document containing the word node w i The number of documents in the numth category, where n is the number of categories.

4. The method for detecting teacher sensitive speech in classroom teaching scenarios according to claim 1, characterized in that: The threshold is 0.

9.

5. The method for detecting teacher sensitive speech in classroom teaching scenarios according to claim 1 is characterized in that: The hyper parameters include: batch size of 64, learning rate of 0.01, point-by-point mutual information window of 10, dropout rate of 0.5 in the two-layer graph convolutional neural network, dropout rate of 0.55 in the graph reconstruction neural network, K of 5 in the KNN algorithm, and minimum threshold of 0.95.