Speaker Identification Correction via Visual Integration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speaker identification methods in portable electronic devices, such as tablets and smartphones, are not 100% accurate due to various environmental influences, leading to incorrect identification of speech from multiple speakers or a single speaker.
Innovation Solution
An electronic device with a receiver and display controller that integrates speech periods from multiple speakers into a single speaker and allows users to correct speaker identification errors through a user interface, enabling integration or division of speech periods based on audio data analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If speaker identification is performed automatically using known methods, then speech can be visualized to support preparation of meeting minutes, but identification accuracy deteriorates under various environmental influences causing speech of the same person to be identified as speech of multiple persons or vice versa
Solution Approach 1:
The system performs preliminary speaker identification automatically to create an initial visualization, then allows users to perform corrective actions by integrating or dividing speech periods. This preliminary action enables efficient meeting minutes preparation while providing opportunity to improve accuracy through user correction of identification errors.
Solution Approach 2:
The system provides visual feedback by displaying speech periods with speaker identifiers, allowing users to detect identification errors. Users can then correct these errors through integration or division operations, creating a feedback loop that improves identification accuracy while maintaining efficient automated processing.
2Speed
If speech periods are automatically identified and displayed without user intervention, then processing speed is improved, but identification reliability deteriorates due to environmental influences causing incorrect speaker attribution
Solution Approach 1:
The system performs preliminary automatic speaker identification to quickly process and display speech periods, then allows users to perform corrective actions. This maintains high processing speed while providing opportunity to improve reliability through user correction of identification errors.
Solution Approach 2:
The system dynamically adjusts between automatic processing mode (for speed) and user correction mode (for reliability). Users can intervene to integrate or divide speech periods when identification errors are detected, creating a dynamic system that balances speed and reliability based on actual needs.
3Measurement precision
If the system provides detailed speaker identification results, then measurement precision is improved, but device complexity increases due to the need for integration and division functions
Solution Approach 1:
The system automatically performs speaker identification and provides detailed results without requiring complex manual intervention. The integration and division functions are triggered by simple user actions on the displayed results, allowing the system to serve itself while maintaining high precision with minimal added complexity.
Solution Approach 2:
The visual display of speech periods with speaker identifiers serves as an intermediary between automatic identification and user correction. This intermediary interface allows users to make precise corrections through simple actions, improving identification precision without requiring complex direct manipulation of the identification system.
Data Source
AI summary
According to one embodiment, an electronic device includes a display controller and circuitry. The display controller displays a first object indicative of a first speaker, a first object indicative of a second speaker different from the first speaker, a second object indicative of a first speech period identified as a speech of the first speaker, and a second object indicative of a second speech period identified as a speech of the second speaker. The circuitry integrates the first speech period and the second speech period into a speech period of a same speaker when a first operation of associating the first object indicative of the first speaker with the first object indicative of the second speaker is operated.


