Speech Section Extraction Using Transition Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional technologies struggle to determine and extract important speech sections in call data, as they fail to consider the transition of speech sections and rely solely on keyword-based methods or similarity analysis.

Innovation Solution

A speech section extraction device and method that identifies speech sections, determines their types, extracts speech types, and extracts important speech sections based on the combination and transition of speech section types and speech types.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If keyword-based methods or similarity analysis are used to determine important speech sections, then the extraction process can be automated, but the accuracy of determining important speech sections deteriorates because the transition of speech sections is not considered

Engineering Contradiction:
Improveautomation of speech section extractionVSAvoidaccuracy of important speech section determination
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The speech text is segmented into multiple speech sections based on speech transitions, and each section is further analyzed for speech types. This segmentation allows the system to consider both individual speech characteristics and section transitions, resolving the contradiction between automation and accuracy by enabling automated multi-level analysis that captures temporal dynamics.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The analysis transitions from a single-dimensional keyword matching approach to a multi-dimensional approach that incorporates both speech-level features (speech types) and section-level features (transition patterns). This dimensional expansion enables the system to automatically capture complex interaction patterns between consecutive speech sections, improving accuracy while maintaining automation.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Device complexity

If speech section information is determined for each section independently, then the analysis process is simplified, but the ability to determine important speech sections deteriorates because combination and transition information is lost

Engineering Contradiction:
Improvecomplexity of analysis processVSAvoidprecision of important speech section identification
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The system merges speech-level analysis (identifying speech types within each section) with section-level analysis (identifying transition patterns between consecutive sections). This combining of multiple analysis levels allows the system to maintain simplicity in individual processing steps while achieving high precision through the integration of comprehensive features that capture both local and global characteristics of speech sections.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20250029613A1Utterance section extraction device, utterance section extraction method and utterance section extraction program
Publication Date: 2025.01.23 NIPPON TELEGRAPH & TELEPHONE CORP
  • US20250029613A1 patent drawing
  • US20250029613A1 patent drawing
  • US20250029613A1 patent drawing

AI summary

A speech section extraction device includes: a speech section identification unit that identifies a speech section including at least one speech from speech text data including speeches of two or more people; a speech section type determination unit that determines a speech section type for each of the speech section that has been identified; a speech type extraction unit that extracts a speech type of each speech included in the speech text data from the speech text data; and a speech section extraction unit that extracts an important speech section among the speech section that has been identified, based on a combination and transition of the speech section type that has been determined, and a combination and transition of the speech type that has been extracted.