Lecturer Speech Signal Processing With Latent-Space Voice Alignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing automated speech recognition tools in lecture rooms suffer from bias, leading to poor performance for lecturers with voices dissimilar to those used in training, exacerbated by acoustic properties and ambient noise, resulting in inaccurate transcriptions.

Innovation Solution

A method that modifies lecturer speech signals to align with voices that perform well on automated speech recognition tools by encoding and decoding spectrograms using an autoencoder, adjusting the signal in a multi-dimensional space to improve accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If automated speech recognition tools are used in lecture rooms, then transcription capability is provided, but performance deteriorates for lecturers with voices dissimilar to training data due to bias

Engineering Contradiction:
Improvetranscription capabilityVSAvoidrecognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent introduces an intermediary voice modification system that translates the lecturer's voice into a synthetic voice that resembles voices from the training data. This intermediary layer (the voice modification model) bridges the gap between the lecturer's unique voice and the biased recognition tool's training data distribution, allowing accurate transcription without requiring the lecturer's voice to match the training data characteristics.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the acoustic parameters of the lecturer's voice by generating a modified audio signal that has different voice characteristics. The system transforms the input audio signal through voice conversion to change parameters such as pitch, timbre, and other acoustic properties, making the modified voice more similar to voices in the training data and thereby improving recognition accuracy.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If voice modification is applied to align lecturer speech with known performers, then word accuracy rate improves, but processing complexity increases

Engineering Contradiction:
Improveword accuracy rateVSAvoidsignal processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary voice modification of the lecturer's speech signal before it is input to the automated speech recognition tool. By pre-processing the audio signal to transform it into a voice that resembles training data voices, the system prepares the signal in advance to be more compatible with the recognition tool's expectations, thereby improving accuracy without adding complexity during the recognition process itself.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a synthetic copy of the lecturer's voice that mimics the characteristics of voices from the training data. Instead of modifying the original voice directly, the system generates a synthetic counterpart that replicates the acoustic properties of training data voices, allowing the recognition tool to process this copied signal more effectively and achieve higher accuracy.

Inventive Principle:
Principle #26Copying

3Productivity

If existing automated speech recognition tools are used, then transcription is generated, but performance deteriorates due to acoustic properties and ambient noise

Engineering Contradiction:
Improvetranscription generationVSAvoidtranscription accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent extracts and removes the acoustic characteristics of the lecturer's voice from the audio signal through voice conversion. By separating and transforming the voice components, the system isolates the speech content from the lecturer's unique vocal characteristics and acoustic interference, making the remaining signal more suitable for accurate recognition despite acoustic properties and ambient noise.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent converts the harmful effect of the lecturer's unique voice characteristics into a benefit by using voice conversion to transform these characteristics into desirable acoustic properties. Instead of dealing with the challenges posed by the lecturer's voice, the system transforms it into a synthetic voice that has the acoustic properties needed for high-accuracy recognition, turning a potential liability into an asset.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enhances the word accuracy rate of automated speech recognition transcriptions by aligning lecturer voices with known performers, reducing bias and noise interference.

Implementation Method 1

A Fourier transform is applied to the sound signal to produce a spectrogram

Methodology Applied
Scientific EffectFourier transform:

Data Source

PatentUS12424225B2Lecturer speech signal processing
Publication Date: 2025.09.23 HABITAT LEARN LTD
  • US12424225B2 patent drawing
  • US12424225B2 patent drawing
  • US12424225B2 patent drawing

AI summary

A sound signal is received from a room of people, the sound signal including speech of a lecturer. A Fourier transform is applied to the sound signal to produce a spectrogram. An encoding of the spectrogram in a multi-dimensional space is computed using an encoder. A seed is found in the multi-dimensional space, where the seed is a position in the multi-dimensional space which encodes a spectrogram from another lecturer known to have performance on an automated speech recognition tool above a threshold. The encoding is modified by moving the location of the encoding in the multi-dimensional space towards the seed. The modified encoding is decoded into a decoded signal. A reverse Fourier transform is applied to the decoded signal to produce an output sound signal. The output sound signal is sent to the automated speech recognition tool to generate a transcript.