Classroom interaction evaluation method and system based on voice data

By building a classroom interaction network and using social network analysis technology to deeply evaluate classroom interaction behavior, the problem of lack of in-depth analysis in the existing technology is solved, and efficient and accurate classroom interaction analysis and improvement suggestions are achieved.

CN120260607APending Publication Date: 2025-07-04CHONGQING COLLEGE OF ELECTRONICS ENG +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510389522.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing classroom interaction analysis technology lacks in-depth analysis and structured representation, and cannot provide teachers with targeted teaching improvement suggestions.

Method used

By collecting classroom voice data, converting it into digital signals, pre-processing and separating the voice data of different speakers, building a classroom interactive network, and using social network analysis technology for in-depth evaluation.

Benefits of technology

It improves the accuracy and efficiency of classroom interaction analysis and provides teachers with targeted suggestions for improving classroom interaction organization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260607A_ABST
    Figure CN120260607A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of classroom interaction, and discloses a classroom interaction evaluation method and system based on voice data, and the method comprises the steps: collecting classroom voice data, and converting an audio signal in the voice data into a digital signal; the collected audio signals are preprocessed, the audio is converted into characters, and voice data of different speakers are separated; analyzing the voice data of different speakers to obtain a speaker interaction sequence; carrying out statistics on the speaker interaction sequence to obtain interaction frequency characteristics among different speakers; constructing a classroom interaction network based on the speakers based on the obtained interaction durations and interaction frequencies among the different speakers; a social network analysis technology is utilized to perform deep analysis on a classroom interaction network, and classroom interaction behaviors are evaluated through quantitative evaluation parameters. According to the invention, the accuracy and efficiency of classroom interaction analysis are improved, and targeted classroom interaction organization and improvement suggestions are provided for teachers through deep analysis of interaction behaviors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of classroom interaction, and specifically relates to a method and device for classroom interaction evaluation based on voice data. Background Art

[0002] In the current educational environment, classroom interaction is one of the important indicators for evaluating teaching quality and students' learning effects. Traditional classroom interaction analysis mainly relies on manual observation or questionnaires. These methods are not only time-consuming and laborious, but also difficult to comprehensively and objectively reflect the actual interaction situation in the classroom.

[0003] With the development of speech recognition technology, it has become possible to automatically and real-time capture and analyze interaction behaviors in the classroom using language analysis technology. However, most of the existing language analysis technologies focus on single speech recognition or simple speech duration statistics, lacking in-depth analysis and structured representation of classroom interaction behaviors, and unable to provide targeted teaching improvement suggestions for teachers. Summary of the Invention

[0004] Aiming at the deficiencies of the existing technology, the present invention proposes a method and system for classroom interaction evaluation based on voice data to solve the above technical problems.

[0005] In a first aspect, a method for classroom interaction evaluation based on voice data is provided, including:

[0006] Collect classroom voice data and convert the audio signal in the voice data into a digital signal;

[0007] Preprocess the collected audio signal, convert the audio into text using a large model, and separate the voice data of different speakers;

[0008] Analyze the voice data of different speakers after segmentation to obtain the speaker interaction sequence and mark the interaction duration;

[0009] Statistically analyze the speaker interaction sequence to obtain the interaction frequency characteristics between different speakers;

[0010] Based on the obtained interaction duration and interaction frequency between different speakers, construct a classroom interaction network based on speakers;

[0011] Use social network analysis technology to deeply analyze the classroom interaction network and evaluate classroom interaction behaviors through quantitative evaluation parameters.

[0012] Further, the audio signal preprocessing further includes:

[0013] Convert the audio signal into a frequency domain representation through short-time Fourier transform;

[0014] Generate a compact feature representation perceived by the human ear through Mel Frequency Cepstral Coefficients.

[0015] Furthermore, the classroom interaction network is a directed multi - edge network structure, where nodes represent speakers and edges represent interaction relationships.

[0016] Furthermore, the optimization suggestions include:

[0017] Based on the speech duration ratio, interaction frequency, and network density between teachers and students, provide suggestions for adjusting the classroom interaction mode.

[0018] Furthermore, the calculation formula for the network density includes:

[0019] density = 2k / n(n - 1)

[0020] where k represents the actual number of edges and n represents the number of nodes.

[0021] Furthermore, the quantitative evaluation parameters also include the comprehensive index of the interaction network.

[0022] Furthermore, the calculation formula for the comprehensive index of the interaction network includes:

[0023] IT = ET × density

[0024] where ET is the total weighted interaction feature of all edges.

[0025] In a second aspect, a classroom interaction evaluation system based on speech data is provided. According to the classroom interaction evaluation method based on speech data described in any one of the foregoing, it includes:

[0026] A conversion module, configured to collect classroom speech data and convert the audio signal in the speech data into a digital signal;

[0027] A pre - processing module, configured to pre - process the collected audio signal, convert the audio into text using a large model, and separate the speech data of different speakers;

[0028] An analysis module, configured to analyze the speech data of different speakers after segmentation, obtain the speaker interaction sequence, and mark the interaction duration;

[0029] A statistics module, configured to statistically analyze the speaker interaction sequence to obtain the interaction frequency characteristics between different speakers;

[0030] A construction module, configured to construct a classroom interaction network based on speakers based on the obtained interaction duration and interaction frequency between different speakers;

[0031] An evaluation module, configured to deeply analyze the classroom interaction network by using social network analysis technology, and evaluate classroom interaction behaviors through quantifying evaluation parameters.

[0032] In a third aspect, it includes a processor and a memory storing program instructions. The processor is configured to execute the classroom interaction evaluation method based on voice data as described in any one of the foregoing when running the program instructions.

[0033] The invention adopting the above technical solution has the following advantages:

[0034] Through voice recognition and speaker identification technologies, the present invention collects voice data in the classroom, performs preprocessing, and then analyzes the voice data to obtain the interaction frequency characteristics between different speakers, and further constructs the classroom interaction network of the speakers. This method not only improves the accuracy and efficiency of classroom interaction analysis, but also provides targeted suggestions for teachers to improve classroom interaction organization by deeply analyzing interaction behaviors. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the specific embodiments of the present invention, the drawings required for use in the specific embodiments will be briefly introduced below. In all the drawings, the elements or parts are not necessarily drawn to actual scale.

[0036] Figure 1 It is a flowchart of the classroom interaction evaluation method based on voice data of the present invention;

[0037] Figure 2 It is a framework diagram of the Whisper-SNA adaptive classroom voice interaction analysis in the classroom interaction evaluation method based on voice data of the present invention;

[0038] Figure 3 It is a block diagram of the whisper model structure in the classroom interaction evaluation method based on voice data of the present invention;

[0039] Figure 4 It is a schematic diagram of network relationships in the classroom interaction evaluation method based on voice data of the present invention Figure 1 ;

[0040] Figure 5 It is a schematic diagram of network relationships in the classroom interaction evaluation method based on voice data of the present invention Figure 2 ;

[0041] Figure 6 It is a flowchart of the classroom interaction evaluation system based on voice data of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] The embodiments of the technical solution of the present invention will be described in detail below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, so they are only examples and cannot be used to limit the protection scope of the present invention.

[0043] It should be noted that unless otherwise specified, the technical terms or scientific terms used in this application should have the ordinary meaning understood by those skilled in the art to which the present invention belongs. The terms "first", "second", etc. in the description of the embodiments of the present disclosure, the claims and the above drawings are used to distinguish similar objects and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so as to implement the embodiments of the present disclosure described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. Unless otherwise specified, the term "plurality" means two or more. In the embodiments of the present disclosure, the character " / " indicates that the objects before and after are in an "or" relationship. For example, A / B means: A or B. The term "and / or" is a description of the relationship between objects and indicates that three relationships can exist. For example, A and / or B means: A or B, or, the three relationships of A and B. The term "corresponding" may refer to an association relationship or a binding relationship. A corresponding to B means that there is an association relationship or a binding relationship between A and B.

[0044] As Figures 1 to 6 shown, the classroom interaction evaluation method based on voice data of the present invention includes:

[0045] Step S01, collect classroom voice data and convert the audio signal in the voice data into a digital signal;

[0046] Step S02, preprocess the collected audio signal, use a large model to convert the audio into text, and separate the voice data of different speakers (converting the audio into text is to verify the application effect of the large model and confirm its accuracy);

[0047] Step S03, analyze the voice data of different speakers after segmentation to obtain the speaker interaction sequence and mark the interaction duration;

[0048] Step S04, count the speaker interaction sequence to obtain the interaction frequency characteristics between different speakers;

[0049] Step S05, based on the obtained interaction duration and interaction frequency between different speakers, construct a speaker-based classroom interaction network;

[0050] Step S06, use social network analysis technology to deeply analyze the classroom interaction network and evaluate the classroom interaction behavior through quantitative evaluation parameters.

[0051] Specifically, the present invention collects voice data in the classroom through voice recognition and speaker identification technologies, performs preprocessing, and then analyzes the voice data to obtain the interaction frequency characteristics between different speakers, and further constructs the classroom interaction network of the speakers. This method not only improves the accuracy and efficiency of classroom interaction analysis, but also provides targeted suggestions for improving classroom interaction organization for teachers by deeply analyzing interaction behaviors.

[0052] In some embodiments, the audio signal preprocessing further includes:

[0053] Converting the audio signal into a frequency-domain representation through short-time Fourier transform;

[0054] Generating a compact feature representation perceptible to the human ear through Mel-frequency cepstral coefficients.

[0055] In some embodiments, the speech recognition model is the Whisper model, which converts the audio signal into text data through the structure of an encoder and a decoder.

[0056] Specifically, the audio signal is represented as a discrete time function x(t), and its formula is:

[0057]

[0058] where x(t) is the audio signal at time, A i , f i and φ i are the amplitude, frequency, and phase, respectively.

[0059] With this configuration, the system can capture the voice signals of different speakers in the classroom, providing high-quality audio input for subsequent processing steps.

[0060] The steps of using the Whisper model to perform speech-to-text conversion on the audio are as follows:

[0061] (1) Audio feature extraction:

[0062] Before the audio signal enters the Whisper model, feature extraction is first required. This process extracts the time-frequency features of the audio signal through short-time Fourier transform (STFT) and Mel-frequency cepstral coefficients (MFCC). The specific steps are as follows:

[0063] ① Performing short-time Fourier transform on the audio signal to convert the time-domain signal into a frequency-domain representation:

[0064]

[0065] where:

[0066] X(t,f) is the result of the Fourier transform of time t and frequency f;

[0067] x[n] is the sampled value of the audio signal;

[0068] w[n] is the window function used to prevent boundary effects;

[0069] N is the length of the window.

[0070] ② Convert the spectrum to Mel Frequency Cepstral Coefficients (MFCC) through a Mel filter bank:

[0071]

[0072] Where:

[0073] H k is the response of the Mel filter;

[0074] K is the number of filters;

[0075] X(n,k) is the spectral coefficient.

[0076] This feature extraction step generates a compact feature representation by capturing the frequency characteristics perceived by the human ear in the audio signal, facilitating subsequent model processing.

[0077] In some embodiments, the encoder uses a multi-head self-attention mechanism to capture dependencies in the audio signal, and the decoder uses a multi-layer self-attention mechanism to generate the predicted text sequence.

[0078] Specifically, (1) Encoder

[0079] The encoder adopts a multi-head self-attention mechanism (Multi-head Self-Attention Mechanism), which can effectively capture the dependencies between different time steps in the audio signal. Through the self-attention mechanism, the model can simultaneously focus on multiple positions in the input sequence, making the modeling of long-term dependencies more accurate. The calculation process of the encoder is as follows:

[0080] ① Linear transformation: Map the input features to query, key, and value representations through a linear transformation of the input vector:

[0081] Q = XW Q , K = XW K , V = XW V

[0082] Where, W Q , W K , W V are the learned weight matrices.

[0083] ② Calculate the attention scores:

[0084]

[0085] Among them, d k is the dimension of the key vector, which is used to ensure numerical stability.

[0086] ③ Output representation:

[0087] Attention(X) = AV

[0088] Through the above calculations, the Whisper encoder generates the context representation for each frame of the input. These representations capture the correlations between different time steps in the input audio.

[0089] (2) Decoder

[0090] The decoder of Whisper is responsible for converting the context representation output by the encoder into the corresponding text. It uses a multi-layer self-attention mechanism combined with an encoder-decoder cross-attention mechanism to generate the predicted text sequence. The decoder starts generating the predicted text sequence from the start token. Each time the input to the decoder is the text token generated in the previous step. Let the decoder input be Y, then the calculation process is as follows:

[0091] ① Self-attention:

[0092]

[0093] ② Encoder-decoder attention

[0094]

[0095] ③ Finally, the output probability distribution generated by the decoder represents the predicted text sequence:

[0096] Y output = softmax(W o (A self + A cross ))

[0097] Among them, W o represents the weight matrix of the output layer.

[0098] The decoder will repeatedly execute this process, gradually generating the complete text sequence until the end token is generated.

[0099] The Whisper model, through the encoder-decoder structure and the multi-layer self-attention mechanism, can efficiently convert classroom audio into text content and shows strong robustness when dealing with noisy or complex audio environments.

[0100] In some embodiments, structured data is generated in the following manner:

[0101] Extract the speaker identity, the corresponding speech time, and the text content according to the speaker identifier and timestamp in the speech recognition result.

[0102] In some embodiments, the classroom interaction network is a directed multi - edge network structure, where nodes represent speakers and edges represent interaction relationships.

[0103] In some embodiments, the optimization suggestions include:

[0104] Provide suggestions for adjusting the classroom interaction mode based on the speech duration ratio, interaction frequency, and network density between teachers and students.

[0105] Specifically, classroom interaction is a dynamically spreading process, which needs to be represented by a directed network graph. Nodes in the graph represent roles, and directed edges represent the frequency and total duration of interactions, pointing from the previous role to the next role in the speech file.

[0106] First, it is not known how many people there are in the class. Only the number of speakers is known through the audio file. In this process, teachers and students gradually activate nodes, and node activation indicates that they are having a language interaction.

[0107] One audio file corresponds to one graph. The roles having speech interactions in the audio file are represented by nodes, and its edges represent the occurrence and sequence of interactions. The edge value records the total number and total duration of interactions.

[0108] Suppose there are two roles a and b in the speech. Then there are nodes a and b in the graph representing the two roles. The directed edge (a, b) ∈ E indicates the occurrence of the interaction a → b, and the parameter pair (times, duration) on its edge annotates the total number and total duration of the interaction a → b in this audio file, where (a, b) ≠ (b, a). If a graph has k edges, that is, |E| = k, then the probability value p(a, b) of each edge occurrence ∈ [0, 1], where:

[0109]

[0110] The duration of communication from role a to role b is:

[0111]

[0112] The speech interaction feature of roles a and b can be represented by f(a, b) = p(a, b)t(a, b). Obtain the adjacency matrix of the network.

[0113] Among them, p(a, b) is the interaction probability from role a to role b, E a→b (times) represents the number of interactions from a to b, k is the total number of edges, and t(a, b) represents the communication duration between a and b.

[0114] If there is communication between a and b, there is an edge between a and b.

[0115] The value of t(a, b) is the communication duration between a and b; if there is no communication between a and b, then

[0116] the edge between a and b does not exist.

[0117] the value of t(a, b) is 0.d E a→b (duration) represents the communication duration between a and b, and E represents the edge.

[0118] Regarding the characteristics of the network:

[0119]

[0120] Among them, ET represents the sum of weighted interaction characteristics of all edges, and f(a, b) = p(a, b)t(a, b), representing the interaction characteristic value between a and b.

[0121] f(i, j) represents the interaction characteristic value between i and j, i represents a node, j represents a node, and n represents a natural number.

[0122] Network topology structure characteristics:

[0123] The density of the topological structure characteristic graph of the network.

[0124] The density of a complex network graph is an index to measure the density of edges in the graph, representing the ratio of the actual number of edges to the possible maximum number of edges.

[0125] For a network with n nodes and k actual connected edges, the calculation formula for network density includes:

[0126] density = 2k / n(n - 1)

[0127] Among them, k represents the actual number of edges, and n represents the number of nodes.

[0128] The greater the network density, the more complex the network.

[0129] Evaluation of the interaction network integrating audio and network topology features

[0130] The calculation formula for obtaining the interaction network evaluation by integrating the audio evaluation and the network topology structure includes:

[0131] IT = ET × density

[0132] Among them, ET is the sum of weighted interaction characteristics of all edges.

[0133] In some embodiments, data collection and preprocessing:

[0134] The acquisition of voice data is completed through the real-time recording module of the system. Specifically, the microphone captures the audio signals in the classroom.

[0135] The system uses the Python library PyAudio, which is widely used in audio processing, as a tool for processing audio streams.

[0136] PyAudio provides a simple interface to PyAudio, enabling Python programs to perform efficient audio input and output operations.

[0137] This library is suitable for a variety of audio processing applications, including speech recognition, music generation, and signal processing.

[0138] To ensure the acquisition quality of audio signals, this system is configured with a sampling rate (16000Hz) and an appropriate buffer size, enabling it to efficiently and clearly capture the voice interactions in the classroom.

[0139] Data format: In this embodiment, the design of the output result format aims to clearly and structurally present the results of classroom voice analysis, including speech recognition text, speaker identity information, and timestamps.

[0140] The specific output result format is as follows: [Time range]:[Speaker]:[Recognized text]

[0141] For example: Time range: 10:00:00~10:00:15

[0142] Speaker: Teacher

[0143] Recognized text: Classmates, today we are going to learn new knowledge points.

[0144] Result:

[0145] [10:00:00~10:00:15]:[Teacher]:[Classmates, today we are going to learn new knowledge points.]

[0146] 10:00:00~10:00:15, Teacher: Classmates, today we are going to learn new knowledge points.

[0147] 10:00:16~10:00:30, Student A: Teacher, can you explain this knowledge point again?

[0148] Whisper selection:

[0149] Model comparison diagram

[0150] model CER WER Whisper-medium 0.23883 1.18577 deepspeech-0.9.3-models.pbmm 1.13295 4.0583 Whisper-large-v3 0.22691 1.11528

[0151] Application to different data result comparisons:

[0152] Comparison Chart of Speech Durations between Teachers and Students

[0153]

[0154] In some other embodiments, a classroom interaction evaluation system based on voice data is provided. According to the classroom interaction evaluation method based on voice data in any of the foregoing items, it includes:

[0155] A conversion module, configured to collect classroom voice data and convert the audio signal in the voice data into a digital signal;

[0156] A preprocessing module, configured to preprocess the collected audio signal, convert the audio into text using a large model, and separate the voice data of different speakers;

[0157] An analysis module, configured to analyze the voice data of different speakers after segmentation, obtain the speaker interaction sequence, and mark the interaction duration;

[0158] A statistics module, configured to statistically analyze the speaker interaction sequence to obtain the interaction frequency characteristics between different speakers;

[0159] A construction module, configured to construct a classroom interaction network based on speakers based on the obtained interaction duration and interaction frequency between different speakers;

[0160] An evaluation module, configured to deeply analyze the classroom interaction network using social network analysis technology and evaluate the classroom interaction behavior by quantifying evaluation parameters.

[0161] In some other embodiments, it includes a processor and a memory storing program instructions. The processor is configured to execute the classroom interaction evaluation method based on voice data in any of the foregoing items when running the program instructions.

[0162] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered by the scope of the claims and the description of the present invention.

Claims

1. A classroom interaction evaluation method based on voice data, characterized in that, including: Collecting classroom voice data and converting the audio signal in the voice data into a digital signal; Preprocessing the collected audio signal, converting the audio into text using a large model, and separating the voice data of different speakers; Analyzing the voice data of different speakers after segmentation to obtain the speaker interaction sequence and annotating the interaction duration; Statistically analyzing the speaker interaction sequence to obtain the interaction frequency characteristics between different speakers; Constructing a speaker-based classroom interaction network based on the obtained interaction duration and interaction frequency between different speakers; Using social network analysis technology to deeply analyze the classroom interaction network and evaluating the classroom interaction behavior through quantitative evaluation parameters.

2. The classroom interaction evaluation method based on voice data according to claim 1, wherein The audio signal preprocessing further includes: Converting the audio signal into a frequency domain representation through short-time Fourier transform; Generating a compact feature representation perceptible to the human ear through mel-frequency cepstral coefficients.

3. The classroom interaction evaluation method based on voice data according to claim 1, characterized in that The classroom interaction network is a directed multi-edge network structure, where nodes represent speakers and edges represent interaction relationships.

4. The classroom interaction evaluation method based on voice data according to claim 1, characterized in that, The optimization suggestions include: Providing suggestions for adjusting the classroom interaction mode based on the speech duration ratio, interaction frequency, and network density between teachers and students.

5. The classroom interaction evaluation method based on voice data according to claim 4, wherein The calculation formula for the network density includes: density = 2k / n(n - 1) where k represents the actual number of edges and n represents the number of nodes.

6. The classroom interaction evaluation method based on voice data according to claim 1, characterized in that The quantitative evaluation parameter further includes an interaction network comprehensive index.

7. The method for classroom interaction evaluation based on voice data according to claim 6, wherein The calculation formula for the interaction network comprehensive index includes: IT = ET × density where ET is the total weighted interaction feature of all edges.

8. A classroom interaction evaluation system based on voice data, characterized in that, The classroom interaction evaluation method based on voice data according to any one of claims 1 to 7 includes: A conversion module configured to collect classroom voice data and convert the audio signal in the voice data into a digital signal; A preprocessing module configured to preprocess the collected audio signal, convert the audio into text using a large model, and separate the voice data of different speakers; An analysis module configured to analyze the voice data of different speakers after segmentation to obtain the speaker interaction sequence and annotate the interaction duration; A statistics module configured to statistically analyze the speaker interaction sequence to obtain the interaction frequency characteristics between different speakers; A construction module configured to construct a speaker-based classroom interaction network based on the obtained interaction duration and interaction frequency between different speakers; An evaluation module configured to use social network analysis technology to deeply analyze the classroom interaction network and evaluate the classroom interaction behavior through quantitative evaluation parameters.

9. A classroom interaction evaluation system based on voice data, comprising a processor and a memory storing program instructions, characterized in that, The processor is configured to execute the classroom interaction evaluation method based on voice data according to any one of claims 1 to 7 when running the program instructions.

Citation Information

Cited By

  • Network course teaching effect analysis and optimization method and system based on big data

    CN121031992A

  • Methods and Systems for Analyzing and Optimizing the Teaching Effectiveness of Online Courses Based on Big Data

    CN121031992B