Speaker Recognition Anchor Model Generation for Embedded Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speaker recognition systems require a large number of anchors to accurately describe a target speaker, leading to increased computation load and difficulty in implementation on embedded systems, as they can only reflect the distribution of the anchor space.
Innovation Solution
The proposed solution involves generating a reference anchor set using enrollment speeches based on an anchor space, which includes principal and associate anchors, allowing for the creation of smaller-sized anchor models through speaker adaptation techniques, thereby reducing computational load and memory requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a large number of anchors are used to describe the target speaker better, then the speaker recognition accuracy is improved, but the computation load increases and it becomes more difficult to implement on embedded systems
Solution Approach 1:
The patent extracts only the essential and representative anchors from the anchor space to form a compact reference anchor set. Instead of using all anchors or a large number of them, the system selectively extracts the most relevant anchors that capture the essential speaker characteristics, thereby reducing computation load while maintaining recognition accuracy.
Solution Approach 2:
The patent segments the large anchor space into multiple clusters and selects representative anchors from each cluster. This segmentation approach allows the system to cover the entire speaker characteristic space with a smaller number of strategically selected anchors, reducing the overall reference anchor set size while preserving recognition accuracy.
2Measurement precision
If a large number of anchors are used to describe the target speaker better, then the speaker recognition accuracy is improved, but the memory requirements increase
Solution Approach 1:
The patent extracts only the essential and representative anchors from the anchor space to form a compact reference anchor set. Instead of using all anchors or a large number of them, the system selectively extracts the most relevant anchors that capture the essential speaker characteristics, thereby reducing computation load while maintaining recognition accuracy.
Solution Approach 2:
The patent segments the large anchor space into multiple clusters and selects representative anchors from each cluster. This segmentation approach allows the system to cover the entire speaker characteristic space with a smaller number of strategically selected anchors, reducing the overall reference anchor set size while preserving recognition accuracy.
Data Source
AI summary
A modeling device comprises a front end which receives enrollment speech data from each target speaker, a reference anchor set generation unit which generates a reference anchor set using the enrollment speech data based on an anchor space, and a voice print generation unit which generates voice prints based on the reference anchor set and the enrollment speech data. By taking the enrollment speech and speaker adaptation technique into account, anchor models with a smaller size can be generated, so reliable and robust speaker recognition with a smaller size reference anchor set is possible.


