Speaker Identification Using Operation History Feedback
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech data processing systems face inaccuracies in assigning speaker names due to changes in speech features based on physical conditions, leading to user effort in correcting identification information and names.
Innovation Solution
An information processing apparatus that receives and divides speech data into utterance data, assigns speaker identification based on acoustic features, and generates a candidate list for user selection using operation history information to determine accurate speaker names.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If speaker identification is performed based on speech features, then speaker identification information can be automatically assigned, but accuracy deteriorates when speech features change due to physical conditions
Solution Approach 1:
The system stores operation history information including utterance identification information, speaker identification information, and speaker names in a database. When generating candidate lists, the system retrieves and utilizes this historical data to improve identification accuracy, creating a feedback loop that learns from past corrections and adjustments
Solution Approach 2:
The system pre-stores operation history information and speaker data in advance before actual speaker identification tasks. This preliminary preparation allows the system to quickly generate accurate candidate lists without real-time computation, improving both automation and accuracy
2Ease of operation
If arbitrary speaker identification information is assigned to classified speech data, then user task support is improved, but correction effort increases when identification is inaccurate
Solution Approach 1:
The system uses stored operation history to generate candidate lists that reflect actual user corrections and preferences. This feedback mechanism ensures that previously corrected data is learned from, reducing the need for repeated corrections and minimizing user time investment
Solution Approach 2:
The system automatically generates candidate lists with speaker names based on operation history without requiring user intervention for each individual assignment. This self-service approach handles routine identification tasks autonomously, freeing users to focus only on exceptional cases
3Ease of operation
If speaker names are assigned based on phonetic similarity comparison, then candidate suggestions can be provided, but accuracy deteriorates when speech features vary
Solution Approach 1:
The system pre-stores operation history information including accurate speaker names and their corresponding utterance data before actual identification tasks. This preliminary data preparation ensures that candidate lists are generated from verified historical data rather than raw phonetic comparison alone
Solution Approach 2:
The system combines multiple information sources including phonetic similarity data, operation history, speaker identification information, and utterance data to generate candidate lists. This composite approach integrates multiple data types to overcome the limitations of any single method
Data Source
AI summary
According to an embodiment, an information processing apparatus includes a dividing unit, an assigning unit, and a generating unit. The dividing unit is configured to divide speech data into pieces of utterance data. The assigning unit is configured to assign speaker identification information to each piece of utterance data based on an acoustic feature of the each piece of utterance data. The generating unit is configured to generate a candidate list that indicates candidate speaker names so as to enable a user to determine a speaker name to be given to the piece of utterance data identified by instruction information, based on operation history information in which at least pieces of utterance identification information, pieces of the speaker identification information, and speaker names given by the user to the respective pieces of utterance data are associated with one another.


