VoIP Conversing Pair Identification via Binary Voice Activity Patterns

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current VoIP systems face challenges in identifying conversing pairs while maintaining user anonymity, as existing methods lack efficient mechanisms to detect coordinated speech patterns between users.

Innovation Solution

The method involves converting voice activity streams into binary streams, calculating complementary similarity metrics, and using a progressive clustering technique to pair voice streams based on coordinated speech patterns, without revealing user identities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If overlay networks are used to preserve user anonymity in VoIP, then user anonymity is improved, but the ability to identify conversing pairs deteriorates

Engineering Contradiction:
Improveuser anonymityVSAvoidconversing pair identification
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent introduces voice activity detection (VAD) patterns as an intermediary that indirectly identifies conversing pairs without exposing user identities. The VAD pattern captures speech activity characteristics that serve as a mediator between anonymous voice streams and conversing pair identification, allowing identification while maintaining the anonymity provided by overlay networks

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transforms voice streams into binary voice activity patterns, effectively changing the 'color' or representation format of the data. This transformation converts continuous voice signals into discrete binary patterns that reveal conversational structure without revealing underlying user identities, enabling identification through pattern recognition rather than direct identity exposure

Inventive Principle:
Principle #32Color changes

2Device complexity

If traditional pairing methods are used in VoIP networks, then implementation simplicity is improved, but pairing accuracy deteriorates due to inability to detect coordinated speech patterns

Engineering Contradiction:
Improvepairing mechanism complexityVSAvoidconversing pair detection accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent utilizes the rhythmic and periodic nature of conversational speech patterns, treating voice activity detection patterns as vibrations in the data domain. By analyzing these temporal patterns and their coordination between parties, the system detects conversing pairs through pattern synchronization rather than complex identity verification mechanisms

Inventive Principle:
Principle #18Mechanical vibration

Solution Approach 2:

The patent replaces traditional identity-based pairing mechanisms with pattern-based detection using voice activity analysis. Instead of relying on metadata or authentication information, the system substitutes a mechanical signal processing approach that analyzes temporal speech patterns to identify conversing pairs

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If voice streams are analyzed to identify conversing pairs, then identification capability is improved, but computational complexity deteriorates

Engineering Contradiction:
Improveconversing pair identificationVSAvoidanalysis process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts only the essential voice activity detection patterns from complete voice streams, separating the critical identification features from the full audio data. This extraction process removes unnecessary computational burden by focusing only on speech activity presence/absence patterns rather than analyzing complete voice content

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments voice streams into discrete binary voice activity patterns, dividing continuous audio data into manageable temporal segments. This segmentation simplifies analysis by breaking down complex voice data into discrete binary states that are easier to process and compare for identification purposes

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS7822604B2Method and apparatus for identifying conversing pairs over a two-way speech medium
Publication Date: 2010.10.26 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US7822604B2 patent drawing
  • US7822604B2 patent drawing
  • US7822604B2 patent drawing

AI summary

One embodiment of the present method and apparatus for identifying a conversing pair of users of a two-way speech medium includes receiving a plurality of binary voice activity streams, where the plurality of voice activity streams includes a first voice activity stream associated with a first user, and pairing the first voice activity stream with a second voice activity stream associated with a second user, in accordance with a complementary similarity between the first voice activity stream and the second voice activity stream.