Contextualizing Unstructured Spoken Text via Multi-Stage Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems for indexing and contextualizing unstructured text from spoken words, particularly in telephone calls, are inefficient due to the time-consuming nature of human transcription and the limitations of machine transcription in providing contextual understanding, especially when dealing with chopped and broken conversations and coded language.

Innovation Solution

A system comprising a grammar processor, sentence processor, frequency processor, emotion processor, and data significance processor that indexes and time-records unstructured text, segments sentences, compares phrases with a real phrase database, evaluates sentiment, and determines contextual significance, while also utilizing an acoustic-to-electric transducer and signal processor to optimize sound energy conversion and filter frequencies for enhanced emotion analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If human transcription is used to transcribe audio, then accuracy is improved, but time consumption increases significantly

Engineering Contradiction:
Improvetranscription accuracyVSAvoidtranscription time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system merges automated speech recognition with multiple layers of computational analysis (grammar processing, sentence segmentation, phrase comparison, frequency analysis, emotion detection) to create a hybrid system that achieves both speed and contextual accuracy, eliminating the need for purely manual transcription while maintaining high accuracy through multi-stage verification

Inventive Principle:
Principle #5Merging (Combining)

2Loss of time

If machine transcription is used to transcribe audio, then time consumption is reduced, but contextual understanding deteriorates

Engineering Contradiction:
Improvetranscription timeVSAvoidcontextual understanding
Core Design Contradiction:
Loss of timeVSLoss of information

Solution Approach 1:

The system segments the transcription process into distinct functional modules: automated speech recognition for rapid text generation, grammar processing for structural analysis, sentence segmentation for phrase identification, phrase comparison for contextual matching, frequency analysis for pattern recognition, and emotion detection for sentiment analysis. Each module processes specific aspects of the data independently before integrating results, enabling comprehensive contextual understanding without sacrificing speed

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces multiple intermediary processing layers between the raw audio and the final contextualized output. These intermediaries include the grammar processor that analyzes sentence structure, the phrase comparator that matches against databases, the frequency processor that identifies patterns, and the emotion processor that detects sentiment. These intermediaries transform raw machine transcription into contextually enriched data without requiring manual intervention

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If automated speech recognition is used to process spoken word, then processing speed is improved, but contextual significance deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidcontextual significance
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The system performs preliminary processing of the audio signal before text generation, including acoustic-to-electric conversion with optimized frequency filtering (300Hz-3400Hz for speech, additional frequencies for emotion analysis). This preliminary action prepares the data in advance for rapid processing while preserving emotional and contextual information that would otherwise be lost in standard ASR

Inventive Principle:
Principle #10Preliminary action

4Measurement precision

If telephone conversations are monitored for coded language, then detection accuracy is improved, but system complexity increases

Engineering Contradiction:
Improvecoded language detection accuracyVSAvoidmonitoring system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system creates multiple representations or 'copies' of the same data at different levels of analysis: the raw audio signal is copied and processed separately for emotion analysis, the automated transcription is copied and processed through grammar and phrase comparison modules, and the results are integrated. This multi-copy approach enables detection of coded language patterns without requiring a single complex system, distributing complexity across independent processing stages

Inventive Principle:
Principle #26Copying

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enables efficient indexing and contextual monitoring of spoken words, reducing the time required to identify significant phrases and emotions, and supplements human listeners by providing actionable insights from emotional intensity, such as alerting response teams for urgent situations.

Implementation Method 1

optimise the transducing of sound energy, such as an audio information signal representative of the spoken word, into a digitised signal

Methodology Applied
Scientific EffectAcoustic-to-electric transduction:

Implementation Method 2

signal processor to optimize sound energy conversion and filter frequencies for enhanced emotion analysis

Methodology Applied
Scientific EffectFrequency filtering: Filter (electronic)

Data Source

PatentUS10546064B2System and method for contextualising a stream of unstructured text representative of spoken word
Publication Date: 2020.01.28 VERINT SYST UK LTD
  • US10546064B2 patent drawing
  • US10546064B2 patent drawing
  • US10546064B2 patent drawing

AI summary

A system for contextualising an unstructured stream of text, representative of spoken word, including a grammar processor, a sentence processor, a frequency processor, a summer and an emotion processor. The unstructured stream of text is processed and outputs an audio file total for each matched phrase, word and proper noun, determined from the unstructured text, to a data significance processor. The data significance processor receives and audio file total for each name, proper noun and matched real phrase, determined from the unstructured text, and outputs a list including the names, proper nouns and matched real phrases in order of contextual significance.