Talker Identification Using Auxiliary Context Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice operation apparatuses face challenges in accurately identifying talkers, especially when the operation voice is atypical or short, and voice quality is affected by environmental factors, leading to low identification accuracy.
Innovation Solution
Incorporating a talker identification unit that uses voice operation information, position information, direction information, distance information, and time information as auxiliary data to improve talker identification accuracy, alongside a voice quality model registered in advance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If voice operation apparatus uses only voice quality model for talker identification, then system complexity is low, but identification accuracy is insufficient especially for atypical or short sentences
Solution Approach 1:
The patent combines multiple information sources (voice quality model, voice operation information, position information, direction information, distance information, and time information) into a unified talker identification system. The talker identification unit integrates these diverse data types to make comprehensive identification decisions, thereby improving accuracy without relying on a single data source.
Solution Approach 2:
The patent transitions from one-dimensional voice quality analysis to multi-dimensional identification by incorporating spatial information (position, direction, distance) and temporal information. This dimensional expansion allows the system to distinguish talkers more accurately by considering multiple attributes simultaneously rather than relying solely on voice characteristics.
2Measurement precision
If voice operation apparatus uses multiple auxiliary information types for talker identification, then identification accuracy improves, but processing complexity increases
Solution Approach 1:
The patent divides the talker identification process into distinct functional modules: voice quality model processing, voice operation information extraction, position information processing, direction information processing, distance information processing, and time information processing. Each module handles specific data types independently before the results are integrated, making the complex processing task more manageable and systematic.
Solution Approach 2:
The talker identification unit serves as an intermediary that receives and processes multiple types of information from different sources. It acts as a mediator that integrates voice quality data with auxiliary information (position, direction, distance, time) to produce accurate talker identification results, managing the complexity through coordinated data fusion.
Data Source
AI summary
A voice operation apparatus and a control method thereof that can further improve accuracy of talker identification are provided. Provided is a voice operation apparatus including a talker identification unit that identifies a user as a talker of a voice operation based on voice information and a voice quality model of a user registered in advance, and a voice operation recognition unit that performs voice recognition on the voice information and generates voice operation information, wherein the talker identification unit identifies a talker by using, as auxiliary information, at least one of the voice operation information, position information on a voice operation apparatus, direction information on a talker, distance information on a talker, and time information.


