Speaker Identification Using Operation History Feedback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech data processing systems face inaccuracies in assigning speaker names due to changes in speech features based on physical conditions, leading to user effort in correcting identification information and names.

Innovation Solution

An information processing apparatus that receives and divides speech data into utterance data, assigns speaker identification based on acoustic features, and generates a candidate list for user selection using operation history information to determine accurate speaker names.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If speaker identification is performed based on speech features, then speaker identification information can be automatically assigned, but accuracy deteriorates when speech features change due to physical conditions

Engineering Contradiction:
Improveautomatic speaker identificationVSAvoidspeaker identification accuracy
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The system stores operation history information including utterance identification information, speaker identification information, and speaker names in a database. When generating candidate lists, the system retrieves and utilizes this historical data to improve identification accuracy, creating a feedback loop that learns from past corrections and adjustments

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system pre-stores operation history information and speaker data in advance before actual speaker identification tasks. This preliminary preparation allows the system to quickly generate accurate candidate lists without real-time computation, improving both automation and accuracy

Inventive Principle:
Principle #10Preliminary action

2Ease of operation

If arbitrary speaker identification information is assigned to classified speech data, then user task support is improved, but correction effort increases when identification is inaccurate

Engineering Contradiction:
Improveuser task supportVSAvoidcorrection time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system uses stored operation history to generate candidate lists that reflect actual user corrections and preferences. This feedback mechanism ensures that previously corrected data is learned from, reducing the need for repeated corrections and minimizing user time investment

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system automatically generates candidate lists with speaker names based on operation history without requiring user intervention for each individual assignment. This self-service approach handles routine identification tasks autonomously, freeing users to focus only on exceptional cases

Inventive Principle:
Principle #25Self-service

3Ease of operation

If speaker names are assigned based on phonetic similarity comparison, then candidate suggestions can be provided, but accuracy deteriorates when speech features vary

Engineering Contradiction:
Improvecandidate list provisionVSAvoidspeaker name assignment accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The system pre-stores operation history information including accurate speaker names and their corresponding utterance data before actual identification tasks. This preliminary data preparation ensures that candidate lists are generated from verified historical data rather than raw phonetic comparison alone

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system combines multiple information sources including phonetic similarity data, operation history, speaker identification information, and utterance data to generate candidate lists. This composite approach integrates multiple data types to overcome the limitations of any single method

Inventive Principle:
Principle #40Composite materials

Data Source

PatentUS9196253B2Information processing apparatus for associating speaker identification information to speech data
Publication Date: 2015.11.24 TOSHIBA DIGITAL SOLUTIONS CORP
  • US9196253B2 patent drawing
  • US9196253B2 patent drawing
  • US9196253B2 patent drawing

AI summary

According to an embodiment, an information processing apparatus includes a dividing unit, an assigning unit, and a generating unit. The dividing unit is configured to divide speech data into pieces of utterance data. The assigning unit is configured to assign speaker identification information to each piece of utterance data based on an acoustic feature of the each piece of utterance data. The generating unit is configured to generate a candidate list that indicates candidate speaker names so as to enable a user to determine a speaker name to be given to the piece of utterance data identified by instruction information, based on operation history information in which at least pieces of utterance identification information, pieces of the speaker identification information, and speaker names given by the user to the respective pieces of utterance data are associated with one another.