Speaker Retrieval Using Score Vector Conversion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speaker retrieval systems require pre-registered speech samples of the desired speaker for searching, making it difficult to find speakers with similar voice qualities unless the desired speaker's speech is available in advance.
Innovation Solution
A speaker retrieval device that converts pre-registered acoustic models into score vectors using a learned conversion model, allowing for the search of speakers based on subjective voice quality features represented by score vectors, enabling the retrieval of speakers with high similarity in voice quality without the need for pre-registered speech samples.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If speech samples of the desired speaker are prepared in advance as a query, then speakers having high degrees of similarity in voice quality can be retrieved, but the system cannot function unless the speech of the desired speaker is available in advance
Solution Approach 1:
The patent introduces score vectors as an intermediary representation that bridges the gap between acoustic models and voice quality similarity search. Instead of requiring actual speech samples of the desired speaker, the system converts acoustic models into score vectors that capture voice quality characteristics, enabling retrieval without the desired speaker's speech being available in advance
Solution Approach 2:
The patent transforms the search parameter from requiring actual speech samples to using score vectors derived from acoustic models. This parameter change allows the system to represent voice quality characteristics in a compressed, processed format that can be searched without needing the original speech data of the desired speaker
2Measurement precision
If acoustic feature quantities are extracted from input speech and similarity is obtained with speakers in the database, then candidate speakers with similar voice qualities can be retrieved, but the system requires the speech of the desired speaker to be prepared in advance
Solution Approach 1:
The patent performs preliminary conversion of acoustic models into score vectors during the database construction phase. This preliminary action stores the transformed representations in advance, so that during retrieval operations, the system can directly compare score vectors without needing to extract acoustic features from the desired speaker's speech at query time
Solution Approach 2:
The patent creates a transformed copy of the acoustic model in the form of a score vector. This copy captures the essential voice quality characteristics in a simplified format that can be used for similarity search, replacing the need to use the original speech samples for comparison
Data Source
AI summary
A speaker retrieval device includes a first converting unit, a receiving unit, and a searching unit. The first converting unit converts, using an inverse transform model of a first conversion model for converting score vectors representing the features of voice quality into acoustic models, pre-registered acoustic models into score vectors; and registers the score vectors in a corresponding manner to a speaker identifier in score management information. The receiving unit receives input of a score vector. The searching unit searches the score management information for the speaker identifiers whose score vectors are similar to the received score vector.


