Speaker Retrieval Using Score Vector Conversion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speaker retrieval systems require pre-registered speech samples of the desired speaker for searching, making it difficult to find speakers with similar voice qualities unless the desired speaker's speech is available in advance.

Innovation Solution

A speaker retrieval device that converts pre-registered acoustic models into score vectors using a learned conversion model, allowing for the search of speakers based on subjective voice quality features represented by score vectors, enabling the retrieval of speakers with high similarity in voice quality without the need for pre-registered speech samples.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech samples of the desired speaker are prepared in advance as a query, then speakers having high degrees of similarity in voice quality can be retrieved, but the system cannot function unless the speech of the desired speaker is available in advance

Engineering Contradiction:
Improvevoice quality similarity retrieval accuracyVSAvoidsystem functionality without pre-registered speech
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent introduces score vectors as an intermediary representation that bridges the gap between acoustic models and voice quality similarity search. Instead of requiring actual speech samples of the desired speaker, the system converts acoustic models into score vectors that capture voice quality characteristics, enabling retrieval without the desired speaker's speech being available in advance

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transforms the search parameter from requiring actual speech samples to using score vectors derived from acoustic models. This parameter change allows the system to represent voice quality characteristics in a compressed, processed format that can be searched without needing the original speech data of the desired speaker

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If acoustic feature quantities are extracted from input speech and similarity is obtained with speakers in the database, then candidate speakers with similar voice qualities can be retrieved, but the system requires the speech of the desired speaker to be prepared in advance

Engineering Contradiction:
Improvevoice quality similarity measurementVSAvoidease of speaker retrieval without pre-registered speech
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent performs preliminary conversion of acoustic models into score vectors during the database construction phase. This preliminary action stores the transformed representations in advance, so that during retrieval operations, the system can directly compare score vectors without needing to extract acoustic features from the desired speaker's speech at query time

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a transformed copy of the acoustic model in the form of a score vector. This copy captures the essential voice quality characteristics in a simplified format that can be used for similarity search, replacing the need to use the original speech samples for comparison

Inventive Principle:
Principle #26Copying

Data Source

PatentUS10978076B2Speaker retrieval device, speaker retrieval method, and computer program product
Publication Date: 2021.04.13 KK TOSHIBA
  • US10978076B2 patent drawing
  • US10978076B2 patent drawing
  • US10978076B2 patent drawing

AI summary

A speaker retrieval device includes a first converting unit, a receiving unit, and a searching unit. The first converting unit converts, using an inverse transform model of a first conversion model for converting score vectors representing the features of voice quality into acoustic models, pre-registered acoustic models into score vectors; and registers the score vectors in a corresponding manner to a speaker identifier in score management information. The receiving unit receives input of a score vector. The searching unit searches the score management information for the speaker identifiers whose score vectors are similar to the received score vector.