Participant-Driven Speech Annotation for Conversation-State Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning models for estimating conversation states from speech data face challenges due to insufficient data annotation and subjective label determination by third parties, leading to inaccurate annotations.

Innovation Solution

An information processing device that determines whether speech data is an insufficient annotation candidate and requests annotation directly from the conversation's participants, using provisional labels to narrow down the data and reduce user burden.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If annotation is performed by third parties who do not participate in actual conversation, then annotation can be performed without involving conversation participants, but annotation accuracy deteriorates due to subjective view limitations

Engineering Contradiction:
Improveease of annotationVSAvoidannotation accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent enables conversation participants to annotate their own speech data. The notification device sends annotation requests to the actual speakers or listeners, who then provide annotations based on their direct experience of the conversation. This self-service approach ensures that annotations reflect the true context and meaning of the speech, resolving the contradiction between ease of operation and annotation accuracy.

Inventive Principle:
Principle #25Self-service

2Quantity of substance

If all speech data is annotated to improve model training, then annotation completeness improves, but user burden increases due to excessive annotation requests

Engineering Contradiction:
Improveannotation completenessVSAvoiduser burden
Core Design Contradiction:
Quantity of substanceVSEase of manufacture

Solution Approach 1:

The patent implements selective annotation by determining whether each speech data requires annotation before sending requests. The notification device evaluates speech data characteristics and only notifies users to annotate when necessary, rather than requesting annotation for all speech data. This partial action approach maintains annotation completeness for critical data while reducing overall user burden.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system performs preliminary determination of annotation necessity before sending annotation requests. By pre-assessing which speech data requires annotation based on predefined criteria, the system avoids unnecessary notification and annotation requests, thereby reducing user burden while maintaining necessary annotation completeness.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12417342B2Information processing device and non-transitory recording medium storing an information processing program
Publication Date: 2025.09.16 TOYOTA JIDOSHA KK
  • US12417342B2 patent drawing
  • US12417342B2 patent drawing
  • US12417342B2 patent drawing

AI summary

An information processing device that is configured to: in a case of annotation of a machine learning model that estimates a state of a party to a conversation from speech data of the conversation, determine whether or not the speech data is an insufficient predetermined annotation candidate, and in a case in which it is determined that the speech data is an insufficient predetermined annotation candidate, request annotation from the party to the conversation, who is at least one of a speaker or a listener of the conversation.