Speech Recognition Noise-Based Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition technologies face inefficiencies in processing recognized strings from input speech due to errors caused by noise, leading to increased user effort and time in selecting and correcting processing units.
Innovation Solution
An information processing device and method that acquires processing units from a recognition string based on noise volume, allowing for efficient processing and correction of selected units, including the ability to adjust the length and number of processing units based on noise levels and user input.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If speech recognition is performed on input speech in noisy environments, then speech recognition functionality is provided, but recognition accuracy deteriorates due to noise
Solution Approach 1:
The recognition string is divided into multiple processing units based on noise volume. When noise is detected, the string is segmented into smaller units, allowing users to select and correct only specific portions rather than the entire string, thereby maintaining functionality while improving accuracy in noisy environments.
Solution Approach 2:
The number and size of processing units are dynamically adjusted based on the detected noise volume. In high-noise conditions, more and smaller processing units are created; in low-noise conditions, fewer and larger units are used. This dynamic adaptation optimizes recognition accuracy while preserving speech recognition functionality across varying environmental conditions.
2Measurement precision
If multiple processing units are generated from recognition string for user selection, then correction capability is improved, but user operation complexity increases
Solution Approach 1:
The recognition string is segmented into processing units that are presented to users for selection. This segmentation enables precise correction of only the problematic portions of the recognized text, improving correction capability while keeping the interface manageable through structured division rather than presenting the entire string at once.
Solution Approach 2:
The number and size of processing units are changed based on noise volume parameters. By adjusting these parameters dynamically, the system optimizes the balance between providing sufficient correction options and avoiding overwhelming users with too many small units, thereby improving correction capability while controlling operation complexity.
3Measurement precision
If processing units are divided into smaller segments for noise-based correction, then recognition accuracy improves, but processing time increases
Solution Approach 1:
The recognition string is segmented into processing units based on noise volume. This segmentation improves recognition accuracy by enabling focused correction on affected portions. The processing time increase is mitigated by only requiring user interaction with the segmented units that need correction rather than processing the entire string uniformly.
Solution Approach 2:
The granularity of processing units is dynamically adjusted based on noise conditions. In high-noise environments, smaller units are created to improve accuracy where needed. The system balances processing time by adapting the level of segmentation to the actual noise conditions, avoiding unnecessary fine-grained processing in low-noise situations.
Data Source
AI summary
Provided is an information processing device including a processing unit acquisition portion that acquires one or more processing units, on the basis of noise, from a first recognition string obtained by performing speech recognition on first input speech, and a processor that processes a processing target, when any one of the one or more processing units is selected as the processing target.


