Unsupervised Audio Stream Thematic Analysis Using Dynamic Document Sizing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for analyzing audio streams struggle with precise segmentation and thematic structuring due to undefined document sizes and reliance on text-based analysis models, which are not suited for audio streams, leading to inefficiencies in detecting recurring sound patterns and handling background noise variability.

Innovation Solution

An unsupervised system and method that extracts acoustic parameters, performs low-level segmentation, and uses fuzzy classification and dynamic programming to determine optimal document sizes and segmentations, adapting probabilistic latent semantic analysis models to identify recurring themes and patterns in audio streams without requiring predefined document sizes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If text-based analysis models (LSA, PLSA, LDA) are used for audio stream analysis, then semantic analysis capability is improved, but the models are not suited for audio streams and require arbitrary fixed document sizes

Engineering Contradiction:
Improvethematic analysis precisionVSAvoidadaptability to audio streams
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent transforms the fixed document size parameter into a variable parameter that can be optimized. It introduces a document size optimization step that searches for the best document size within a range [tmin, tmax] by evaluating likelihood scores, allowing the system to adapt to different audio stream characteristics rather than using arbitrary fixed sizes.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent makes the document size dynamic rather than static. Instead of using a single fixed document size, the system evaluates multiple document sizes and selects the optimal one based on likelihood metrics. This dynamic approach allows the analysis to adapt to the specific characteristics of each audio stream being analyzed.

Inventive Principle:
Principle #15Dynamics

2Ease of operation

If a single fixed document size is used, then processing simplicity is improved, but segmentation precision deteriorates due to inability to capture varying theme durations

Engineering Contradiction:
Improveprocessing simplicityVSAvoidsegmentation precision
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent applies segmentation at multiple levels: first segmenting the audio stream into frames, then into segments, and finally into documents of variable sizes. This multi-level segmentation allows the system to capture themes at different granularities and durations, improving segmentation precision while maintaining manageable processing complexity through hierarchical organization.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces dynamic document sizing where the document size is not fixed but optimized for each analysis task. The system evaluates multiple document sizes and selects the optimal one based on likelihood scores, making the segmentation adaptive to the specific audio stream characteristics rather than using a static size for all cases.

Inventive Principle:
Principle #15Dynamics

3Reliability

If conventional systems rely on image and video processing technologies, then environmental monitoring capability is improved, but audio stream thematic analysis capability deteriorates

Engineering Contradiction:
Improveenvironmental monitoring capabilityVSAvoidaudio thematic analysis precision
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent substitutes image and video processing methodologies with audio-specific processing techniques. It adapts the probabilistic latent semantic analysis framework originally designed for text to work with audio streams by using acoustic parameter extraction and audio segment classification, creating an audio-optimized analysis system rather than forcing visual processing methods onto auditory data.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Measurement precision

If exhaustive search of temporal supports is performed to find optimal document sizes, then segmentation accuracy is improved, but computational complexity increases

Engineering Contradiction:
Improvesegmentation accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements a balanced search approach where it evaluates document sizes within a practical range [tmin, tmax] rather than searching all possible sizes. This partial action approach finds sufficiently accurate segmentations without requiring exhaustive computation of every possible document size, achieving a practical compromise between accuracy and computational complexity.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentEP2766825B1Non-supervised system and method for multiresolution thematic analysis and structuring of audio streams
Publication Date: 2016.01.06 THALES SA
  • EP2766825B1 patent drawingFigure 1
  • EP2766825B1 patent drawingFigure 2
  • EP2766825B1 patent drawingFigure 3

AI summary

Non-supervised method and system for thematic analysis and structuring of an audio stream, said audio stream comprising a set of documents each consisting of several segments, said audio stream comprising one or more patterns or themes repeated over time, said system being adapted to implement the method according to the invention characterized in that it comprises at least the following elements: . a database (11) comprising a pre-established dictionary of audio words, and statistical models, . one or more acoustic sensors (2, 20) linked to an analysis module comprising a first module (22) for extracting acoustic parameters from said audio stream, said first module being linked to a buffer memory (23), the output of the buffer memory being linked to a module for fuzzy classification (24) of said segments emanating from the buffer memory (23), a module for constructing co-occurrence matrices and for adapting the models (25), a module (26) adapted for determining the likelihood of the models and a module for optimal segmentation searching executing an algorithm of dynamic programming type applied to the likelihood measurements.