Voice Recognition Device Using Adjusted Waveform Group Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In manufacturing sites, slight shifts in voice attributes such as cutout position, background noise, or utterance speed can significantly disturb voice recognition results, leading to reduced accuracy and reproducibility, making it difficult to investigate recognition failures.
Innovation Solution
A voice recognition device that generates multiple adjusted voice signals by finely adjusting predetermined attributes of an input voice signal, such as utterance speed, and determines the most frequent recognition result across these adjustments as the correct outcome.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If voice recognition is performed on a single input voice signal, then the processing is simple and fast, but the recognition accuracy is reduced due to disturbances in voice attributes
Solution Approach 1:
The patent segments the voice recognition process by dividing it into multiple parallel recognition tasks. Instead of performing one recognition on the original signal, the system performs multiple recognitions on adjusted versions of the signal (with modified attributes like pitch, tempo, and timbre). This segmentation allows the system to handle attribute disturbances by distributing the recognition task across multiple variations, ultimately selecting the most reliable result.
Solution Approach 2:
The patent applies preliminary action by pre-adjusting the voice signal attributes before recognition. The system proactively modifies the voice signal to create multiple adjusted signals with different attribute values before performing recognition. This preliminary adjustment ensures that even if the original signal has attribute disturbances, the adjusted signals may compensate for these disturbances, improving the likelihood of accurate recognition.
2Reliability
If multiple adjusted voice signals are generated and recognized, then the recognition accuracy is improved, but the processing time increases
Solution Approach 1:
The patent applies partial action by performing recognition on a limited number of adjusted signals rather than exhaustively testing all possible attribute variations. The system generates multiple adjusted signals (e.g., 3-5 variations) with moderate attribute adjustments, performing recognition on these partial variations. This approach achieves sufficient accuracy improvement without the excessive time cost of comprehensive attribute space exploration.
3Reliability
If multiple adjusted voice signals are generated and recognized, then the reproducibility of recognition results is improved, but the device complexity increases
Solution Approach 1:
The patent systematically changes voice signal parameters (attributes) to create adjusted signals. Instead of complex structural modifications, the system modifies parameters such as pitch, tempo, and timbre using standard signal processing techniques. This parameter-based approach improves reproducibility by accounting for attribute variations while maintaining relatively simple system architecture, as it relies on conventional parameter manipulation rather than complex additional hardware or algorithms.
Data Source
AI summary
A voice recognition device according to the present disclosure performs voice recognition on a voice signal inputted on manufacturing premises and uses the result as a voice command, the voice recognition device comprising: an adjustment waveform group generation unit for performing a plurality of different adjustments on a prescribed attribute of an inputted voice signal and generating a plurality of adjusted voice signals corresponding to the same; and a voice recognition unit for performing voice recognition on the plurality of adjusted voice signals and the voice signal outputted by the adjustment waveform group generation unit. The adjustment performed by the adjustment waveform group generation unit includes, as an attribute to be adjusted, the speech speed.


