Voice Recognition Model Tag-Based Data Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Creating a voice recognition model with high accuracy is challenging due to the difficulty in preparing a data set with a sufficient amount of voice data collected under specific conditions, such as environment and speaker characteristics.

Innovation Solution

A voice recognition device that acquires and processes voice data using a first recognition model, extracts significant tags based on recognition results, and creates a second model by combining data sets associated with these tags to enhance recognition accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a recognition model is created based on a data set with sufficient data amount, then recognition accuracy is improved, but data preparation difficulty increases

Engineering Contradiction:
Improverecognition accuracyVSAvoiddata preparation difficulty
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The patent creates synthetic voice data by copying and transforming existing voice data characteristics. A learning model generates artificial voice data that mimics real voice patterns, allowing the system to expand the training data set without requiring additional real-world data collection. This copying approach enables sufficient data quantity for high recognition accuracy while avoiding the difficulty of collecting enough real voice samples.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs self-service by automatically generating its own training data. The learning model uses existing voice data to create synthetic samples, eliminating the need for external data collection efforts. This self-service mechanism allows the system to independently expand its data set to achieve high recognition accuracy without relying on difficult external data preparation.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If voice data is collected under specific conditions (environment, speaker characteristics), then recognition accuracy under those conditions is improved, but data collection complexity increases

Engineering Contradiction:
Improverecognition accuracy under specific conditionsVSAvoiddata collection complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies local quality by generating voice data with specific local characteristics (environmental conditions, speaker features) as needed. Rather than collecting comprehensive data under all possible conditions, the learning model synthesizes voice data with particular local qualities required for specific recognition scenarios, reducing data collection complexity while maintaining high accuracy for target conditions.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system uses parameter changes to transform existing voice data into synthetic samples with different environmental and speaker characteristics. By adjusting parameters such as background noise levels, speaker identity features, and recording conditions during synthesis, the system can generate data for specific conditions without physically collecting data under those conditions, thereby reducing collection complexity.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11600262B2Recognition device, method and storage medium
Publication Date: 2023.03.07 KK TOSHIBA
  • US11600262B2 patent drawing
  • US11600262B2 patent drawing
  • US11600262B2 patent drawing

AI summary

According to one embodiment, a recognition device includes storage and a processor. The storage is configured to store a first recognition model, a first data set, and tags, for each first recognition model. The processor is configured to acquire a second data set, execute recognition processing of the second recognition target data in the second data set by using the first recognition model, extract a significant tag of the tags stored in the storage in association with the first recognition model, based on the recognition processing result and the second correct data in the second data set, and create a second recognition model based on the acquired second data set and the first data set stored in the storage in association with the extracted tag.