Audio Event Analytics via Acoustic Segmentation and Context Correlation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Automatic speech recognition systems are not sufficiently accurate for conversational speech recorded at events like business meetings due to challenges such as multiple speakers, overlapping speech, and interruptions, and existing audio analytics systems fail to effectively utilize supplementary contextual information like meeting agendas and participant roles.

Innovation Solution

A method for directly analyzing audio signals and correlating them with additional sources of information, such as meeting agendas and participant roles, to generate insights and summaries of events without relying solely on automatic speech recognition, using techniques like acoustic segmentation, speaker identification, prosody analysis, and correlation-based speech analysis to identify key characteristics and interactions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If automatic speech recognition is used to analyze audio recordings, then speech content can be identified, but accuracy is insufficient for conversational speech with multiple speakers, overlapping speech, and interruptions

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidhandling of conversational speech scenarios
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The audio recording is segmented into multiple audio segments based on acoustic features, speaker identification, and prosody analysis. This segmentation divides the complex conversational speech into manageable units that can be analyzed more accurately, addressing the limitation of traditional speech recognition systems that treat audio as a continuous stream.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces supplementary contextual information (meeting agendas, participant roles, event metadata) as an intermediary layer between the audio signal and the analysis process. This intermediary provides additional context that helps disambiguate speakers, identify key phrases, and understand the conversational flow, thereby improving accuracy in multi-speaker scenarios.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If traditional audio analytics are used, then basic speech content can be extracted, but supplementary contextual information like meeting agendas and participant roles is not utilized

Engineering Contradiction:
Improvecontextual information utilizationVSAvoidsystem architecture
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent merges audio signal processing with contextual information processing in a unified analysis framework. Audio segments are correlated with meeting agendas, participant roles, and event metadata to produce comprehensive insights. This combining approach ensures that supplementary contextual information is fully utilized without creating separate siloed systems.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The analysis system is designed to handle multiple types of input (audio signals, text agendas, metadata) and produce multiple types of output (summaries, key phrases, emotional characteristics, communication efficiency metrics). This multi-functional approach allows the same system architecture to process diverse event data uniformly.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If comprehensive audio analysis is performed including segmentation, speaker identification, and prosody analysis, then accurate insights can be generated, but processing complexity increases

Engineering Contradiction:
Improveanalysis accuracyVSAvoidprocessing system
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

Audio recordings are pre-processed with acoustic segmentation, speaker identification, and prosody analysis before the main correlation step. These preliminary actions prepare the audio data in a structured format that facilitates more efficient correlation with contextual information, reducing the computational burden of the subsequent analysis.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent analyzes audio not just in the time domain but also in the contextual dimension by correlating audio segments with meeting agendas, participant roles, and event metadata. This dimensional expansion allows the system to derive deeper insights without proportionally increasing processing complexity, as the additional dimension leverages existing contextual data.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS10043517B2Audio-based event interaction analytics
Publication Date: 2018.08.07 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10043517B2 patent drawing
  • US10043517B2 patent drawing
  • US10043517B2 patent drawing

AI summary

In an approach for audio based event analytics, a processor receives a recording of audio from an event. A processor collects information about the event and a list of participants. A processor segments the recording into, at least, a plurality of utterances. A processor analyzes the segmented recording. A processor summarizes the recording based on the segmentation and the analysis. A processor generates insights about interactions patterns of the event based on the segmentation and the analysis.