Speech Content Summary Using Focus Signal Weighting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech content summary systems cannot effectively indicate the relative importance of different segments of speech content, leading to unsatisfactory text summaries as they treat all segments as equally significant.

Innovation Solution

A system that uses a 'focus more' or 'focus less' button to associate multiple focus signals with time windows, determining the importance of speech content segments based on the number of signals, allowing for prioritization in generating a text content summary.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If the system uses a single indication button to mark speech segments, then the operation is simple, but the system cannot distinguish the relative importance of different segments

Engineering Contradiction:
Improveease of operationVSAvoidmeasurement precision
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent divides the indication mechanism into multiple levels by introducing different types of indication buttons (first indication button and second indication button) that can be pressed separately. This segmentation allows the system to distinguish between different levels of importance for speech segments, resolving the contradiction by maintaining operational simplicity while enabling precise measurement of segment importance through differentiated user interactions.

Inventive Principle:
Principle #1Segmentation

2Productivity

If the system treats all marked segments as equal importance, then the processing is simple, but the generated summary does not reflect user priorities

Engineering Contradiction:
ImproveproductivityVSAvoidloss of information
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent applies local quality by assigning different weights or priorities to different speech segments based on which indication button was pressed and how many times. Segments marked with the first indication button receive different processing weight compared to those marked with the second indication button, and repeated presses increase the weight further. This allows the summary generation to reflect user priorities while maintaining efficient processing by using weighted aggregation rather than complex analysis.

Inventive Principle:
Principle #3Local quality

3Measurement precision

If the user presses the indication button multiple times to indicate importance, then the system can prioritize segments, but the operation becomes more complex

Engineering Contradiction:
Improvemeasurement precisionVSAvoidease of operation
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent merges the functionality of multiple indication buttons into a unified interface where both the first indication button and the second indication button can be pressed independently or in combination. The system processes these merged inputs by accumulating indication counts and assigning weights accordingly, maintaining ease of operation while enabling precise measurement of segment importance through the combined effect of multiple presses.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS8868419B2Generalizing text content summary from speech content
Publication Date: 2014.10.21 CERENCE OPERATING CO
  • US8868419B2 patent drawing
  • US8868419B2 patent drawing
  • US8868419B2 patent drawing

AI summary

A text content summary is created from speech content. A focus more signal is issued by a user while receiving the speech content. The focus more signal is associated with a time window, and the time window is associated with a part of the speech content. It is determined whether to use the part of the speech content associated with the time window to generate a text content summary based on a number of the focus more signals that are associated with the time window. The user may express relative significance to different portions of speech content, so as to generate a personal text content summary.