Speaker Identification Correction via Visual Integration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speaker identification methods in portable electronic devices, such as tablets and smartphones, are not 100% accurate due to various environmental influences, leading to incorrect identification of speech from multiple speakers or a single speaker.

Innovation Solution

An electronic device with a receiver and display controller that integrates speech periods from multiple speakers into a single speaker and allows users to correct speaker identification errors through a user interface, enabling integration or division of speech periods based on audio data analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If speaker identification is performed automatically using known methods, then speech can be visualized to support preparation of meeting minutes, but identification accuracy deteriorates under various environmental influences causing speech of the same person to be identified as speech of multiple persons or vice versa

Engineering Contradiction:
Improveefficiency of meeting minutes preparationVSAvoidspeaker identification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system performs preliminary speaker identification automatically to create an initial visualization, then allows users to perform corrective actions by integrating or dividing speech periods. This preliminary action enables efficient meeting minutes preparation while providing opportunity to improve accuracy through user correction of identification errors.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system provides visual feedback by displaying speech periods with speaker identifiers, allowing users to detect identification errors. Users can then correct these errors through integration or division operations, creating a feedback loop that improves identification accuracy while maintaining efficient automated processing.

Inventive Principle:
Principle #23Feedback

2Speed

If speech periods are automatically identified and displayed without user intervention, then processing speed is improved, but identification reliability deteriorates due to environmental influences causing incorrect speaker attribution

Engineering Contradiction:
Improvespeech processing speedVSAvoidspeaker identification reliability
Core Design Contradiction:
SpeedVSReliability

Solution Approach 1:

The system performs preliminary automatic speaker identification to quickly process and display speech periods, then allows users to perform corrective actions. This maintains high processing speed while providing opportunity to improve reliability through user correction of identification errors.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system dynamically adjusts between automatic processing mode (for speed) and user correction mode (for reliability). Users can intervene to integrate or divide speech periods when identification errors are detected, creating a dynamic system that balances speed and reliability based on actual needs.

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If the system provides detailed speaker identification results, then measurement precision is improved, but device complexity increases due to the need for integration and division functions

Engineering Contradiction:
Improvespeaker identification precisionVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system automatically performs speaker identification and provides detailed results without requiring complex manual intervention. The integration and division functions are triggered by simple user actions on the displayed results, allowing the system to serve itself while maintaining high precision with minimal added complexity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The visual display of speech periods with speaker identifiers serves as an intermediary between automatic identification and user correction. This intermediary interface allows users to make precise corrections through simple actions, improving identification precision without requiring complex direct manipulation of the identification system.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS9536526B2Electronic device with speaker identification, method and storage medium
Publication Date: 2017.01.03 DYNABOOK INC
  • US9536526B2 patent drawing
  • US9536526B2 patent drawing
  • US9536526B2 patent drawing

AI summary

According to one embodiment, an electronic device includes a display controller and circuitry. The display controller displays a first object indicative of a first speaker, a first object indicative of a second speaker different from the first speaker, a second object indicative of a first speech period identified as a speech of the first speaker, and a second object indicative of a second speech period identified as a speech of the second speaker. The circuitry integrates the first speech period and the second speech period into a speech period of a same speaker when a first operation of associating the first object indicative of the first speaker with the first object indicative of the second speaker is operated.