Speech Recognition Noise-Based Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition technologies face inefficiencies in processing recognized strings from input speech due to errors caused by noise, leading to increased user effort and time in selecting and correcting processing units.

Innovation Solution

An information processing device and method that acquires processing units from a recognition string based on noise volume, allowing for efficient processing and correction of selected units, including the ability to adjust the length and number of processing units based on noise levels and user input.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If speech recognition is performed on input speech in noisy environments, then speech recognition functionality is provided, but recognition accuracy deteriorates due to noise

Engineering Contradiction:
Improvespeech recognition functionalityVSAvoidrecognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The recognition string is divided into multiple processing units based on noise volume. When noise is detected, the string is segmented into smaller units, allowing users to select and correct only specific portions rather than the entire string, thereby maintaining functionality while improving accuracy in noisy environments.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The number and size of processing units are dynamically adjusted based on the detected noise volume. In high-noise conditions, more and smaller processing units are created; in low-noise conditions, fewer and larger units are used. This dynamic adaptation optimizes recognition accuracy while preserving speech recognition functionality across varying environmental conditions.

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If multiple processing units are generated from recognition string for user selection, then correction capability is improved, but user operation complexity increases

Engineering Contradiction:
Improvecorrection capabilityVSAvoiduser operation complexity
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The recognition string is segmented into processing units that are presented to users for selection. This segmentation enables precise correction of only the problematic portions of the recognized text, improving correction capability while keeping the interface manageable through structured division rather than presenting the entire string at once.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The number and size of processing units are changed based on noise volume parameters. By adjusting these parameters dynamically, the system optimizes the balance between providing sufficient correction options and avoiding overwhelming users with too many small units, thereby improving correction capability while controlling operation complexity.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If processing units are divided into smaller segments for noise-based correction, then recognition accuracy improves, but processing time increases

Engineering Contradiction:
Improverecognition accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The recognition string is segmented into processing units based on noise volume. This segmentation improves recognition accuracy by enabling focused correction on affected portions. The processing time increase is mitigated by only requiring user interaction with the segmented units that need correction rather than processing the entire string uniformly.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The granularity of processing units is dynamically adjusted based on noise conditions. In high-noise environments, smaller units are created to improve accuracy where needed. The system balances processing time by adapting the level of segmentation to the actual noise conditions, avoiding unnecessary fine-grained processing in low-noise situations.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS10540968B2Information processing device and method of information processing
Publication Date: 2020.01.21 SONY GROUP CORP
  • US10540968B2 patent drawing
  • US10540968B2 patent drawing
  • US10540968B2 patent drawing

AI summary

Provided is an information processing device including a processing unit acquisition portion that acquires one or more processing units, on the basis of noise, from a first recognition string obtained by performing speech recognition on first input speech, and a processor that processes a processing target, when any one of the one or more processing units is selected as the processing target.