Speech Signal Processing System Using Environmental Sound Estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech signal processing systems fail to effectively incorporate characteristics such as environmental noise and volume of input speech during speech recognition, leading to inadequate speech conversion and recognition results.
Innovation Solution
A speech signal processing system that includes a speech input unit, storage unit, characteristic estimation unit, reference speech output unit, and characteristic adding unit, which estimates and adds environmental sound and volume characteristics to a reference speech signal, enabling accurate speech recognition by replicating the input speech environment.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If environmental sound is always superimposed at current time point, then speech conversion can be performed, but characteristics of input speech such as volume and signal blocking cannot be added
Solution Approach 1:
The system performs preliminary estimation of input speech characteristics (volume, environmental sound, signal blocking) before speech conversion. The characteristic estimation unit analyzes the input speech signal in advance to extract these characteristics, which are then stored and applied during the speech conversion process, ensuring that the converted speech reflects the original acoustic conditions.
Solution Approach 2:
The system creates a copy of the input speech signal and uses it to estimate characteristics without altering the original. The characteristic estimation unit refers to the input speech signal stored in the storage unit to extract environmental sound and volume information, then applies these characteristics to the converted speech through the characteristic adding unit, effectively copying the acoustic characteristics into the output.
2Measurement precision
If noise model is normalized to improve speech recognition accuracy, then recognition result accuracy improves, but characteristics of input speech such as environmental sound and volume cannot be utilized
Solution Approach 1:
The system segments the speech signal processing into distinct functional units: a speech recognition unit that processes the input speech signal, a characteristic estimation unit that extracts environmental sound and volume characteristics, and a characteristic adding unit that incorporates these characteristics into the converted speech. This segmentation allows each unit to specialize in its function while working together to achieve both accurate recognition and characteristic preservation.
Solution Approach 2:
The characteristic estimation unit acts as an intermediary between the input speech signal and the speech conversion process. It extracts relevant characteristics (environmental sound, volume) from the input speech and serves as a bridge to the characteristic adding unit, which then incorporates these characteristics into the converted speech, enabling the system to utilize input characteristics without compromising recognition accuracy.
3Device complexity
If speech signal is processed without considering volume and blocking characteristics, then processing is simplified, but accuracy of speech conversion deteriorates
Solution Approach 1:
The system performs preliminary estimation of speech characteristics (volume, environmental sound, signal blocking) before the main speech conversion process. The characteristic estimation unit analyzes the input speech signal in advance to extract these characteristics, which are then stored and applied during conversion, ensuring accurate speech conversion without adding complexity to the main processing flow.
Solution Approach 2:
The system changes parameters of the speech signal based on estimated characteristics. The characteristic adding unit modifies the converted speech by adjusting parameters such as volume and adding environmental sound characteristics, thereby improving speech conversion accuracy while maintaining a relatively simple processing architecture through parameter-based adjustments rather than complex structural changes.
Data Source
AI summary
A speech signal processing system that includes a speech input unit for inputting a speech signal; input speech storage unit for storing an input speech signal that is the speech signal inputted through the speech input unit; characteristic estimation unit for referring to the input speech signal stored in the input speech storage unit, and estimating characteristics of an input speech indicated by the input speech signal, the characteristics including an environmental sound included in the input speech signal; reference speech output unit for causing a predetermined speech signal that becomes a reference speech, to output; and characteristic adding unit for adding the characteristics of the input speech estimated by the characteristic estimation unit, in a reference speech signal that is the speech signal caused to output by the reference speech output unit.


