Audio-Based Instructor Discourse Analysis With Local-Global Context

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems for evaluating classroom instructor discourse suffer from halo effects, unreliable human observations, limited rater supply, and privacy concerns with video data, while automated systems lack fine-grained analysis capabilities and raise privacy issues.

Innovation Solution

A method and system using automatic speech recognition to convert audio signals into transcripts, extract features, filter student talk, and analyze local and global context predictions to provide fine-grained feedback on instructor discourse, incorporating machine learning for robust classification and reliability estimation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If human-based methods are used to observe and rate instructor discourse, then detailed qualitative feedback can be obtained, but halo effects and reliability issues reduce measurement accuracy

Engineering Contradiction:
Improvediscourse evaluation accuracyVSAvoidrater consistency
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent replaces the human mechanical observation system with an automated computational system that uses machine learning models to analyze instructor discourse. This substitution eliminates human rater biases and halo effects while maintaining detailed qualitative feedback capabilities through natural language processing of transcript data.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces an automated analysis system as an intermediary between the instructor discourse and the evaluation process. This intermediary uses speech recognition to convert audio to text, then applies trained machine learning models to objectively measure discourse quality without human intervention, thereby improving reliability.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If automated systems use wearable recording to capture instructor discourse, then productivity increases, but analysis granularity is too coarse to provide meaningful feedback

Engineering Contradiction:
Improveevaluation throughputVSAvoiddiscourse analysis granularity
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent segments the discourse analysis into multiple hierarchical levels: utterance-level features, turn-level interactions, and session-level patterns. This segmentation allows the system to process large volumes of data efficiently while simultaneously capturing fine-grained instructional details that provide meaningful feedback.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds temporal and contextual dimensions to the analysis by examining discourse patterns across multiple time scales (individual utterances, conversational turns, and entire sessions). This multi-dimensional approach enables both high productivity through automated processing and fine-grained precision through contextual analysis.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Measurement precision

If video recordings are used to enable automated fine-grained analysis, then measurement precision improves, but privacy concerns increase

Engineering Contradiction:
Improvediscourse analysis capabilityVSAvoidprivacy concerns
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent extracts only the necessary audio data from the classroom environment and converts it to text transcripts, discarding all visual information. This extraction approach maintains the ability to perform fine-grained discourse analysis while eliminating privacy concerns associated with video recording of students and instructors.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent substitutes video-based analysis with audio-based speech recognition and natural language processing. This substitution achieves comparable or superior measurement precision for discourse quality while avoiding the privacy intrusions of video recording, as audio data can be processed without capturing visual images of individuals.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Loss of information

If human raters are used to provide detailed feedback, then analysis depth increases, but the limited supply of raters reduces productivity

Engineering Contradiction:
Improvefeedback detail qualityVSAvoidevaluation capacity
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The patent enables the discourse evaluation system to serve itself by using automated machine learning models that can analyze unlimited amounts of discourse data without requiring human rater involvement. This self-service capability maintains detailed feedback quality while dramatically increasing evaluation capacity to serve multiple instructors simultaneously.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent changes the operational parameters of the evaluation system from human-dependent to computationally-driven, allowing parallel processing of multiple discourse samples. This parameter change enables both deep analytical feedback and high productivity by removing the bottleneck of limited human rater availability.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12400658B2System and method for automated observation and analysis of instructional discourse
Publication Date: 2025.08.26 UNIV OF PITTSBURGH OF THE COMMONWEALTH SYST OF HIGHER EDUCATION
  • US12400658B2 patent drawing
  • US12400658B2 patent drawing
  • US12400658B2 patent drawing

AI summary

A method of analyzing instructor discourse includes recording an audio signal representing speech of the instructor during a class session, converting the audio signal to a session transcript comprising speech data for the session using an automatic speech recognition tool and segmenting the transcript into utterances, extracting a set of features from the session transcript, filtering student talk out from the utterances, analyzing a first subset of the features to produce a number of local context predictions for each utterance of the session transcript, analyzing a second subset of the features to produce a number of global context predictions for the session transcript, and combining a subset of the number of local context predictions and the number of global context predictions into a classification that attends to differential reliability.