Speech Signal Processing System Using Environmental Sound Estimation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech signal processing systems fail to effectively incorporate characteristics such as environmental noise and volume of input speech during speech recognition, leading to inadequate speech conversion and recognition results.

Innovation Solution

A speech signal processing system that includes a speech input unit, storage unit, characteristic estimation unit, reference speech output unit, and characteristic adding unit, which estimates and adds environmental sound and volume characteristics to a reference speech signal, enabling accurate speech recognition by replicating the input speech environment.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If environmental sound is always superimposed at current time point, then speech conversion can be performed, but characteristics of input speech such as volume and signal blocking cannot be added

Engineering Contradiction:
Improvespeech conversion capabilityVSAvoidinput speech characteristics
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The system performs preliminary estimation of input speech characteristics (volume, environmental sound, signal blocking) before speech conversion. The characteristic estimation unit analyzes the input speech signal in advance to extract these characteristics, which are then stored and applied during the speech conversion process, ensuring that the converted speech reflects the original acoustic conditions.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates a copy of the input speech signal and uses it to estimate characteristics without altering the original. The characteristic estimation unit refers to the input speech signal stored in the storage unit to extract environmental sound and volume information, then applies these characteristics to the converted speech through the characteristic adding unit, effectively copying the acoustic characteristics into the output.

Inventive Principle:
Principle #26Copying

2Measurement precision

If noise model is normalized to improve speech recognition accuracy, then recognition result accuracy improves, but characteristics of input speech such as environmental sound and volume cannot be utilized

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidutilization of input speech characteristics
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system segments the speech signal processing into distinct functional units: a speech recognition unit that processes the input speech signal, a characteristic estimation unit that extracts environmental sound and volume characteristics, and a characteristic adding unit that incorporates these characteristics into the converted speech. This segmentation allows each unit to specialize in its function while working together to achieve both accurate recognition and characteristic preservation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The characteristic estimation unit acts as an intermediary between the input speech signal and the speech conversion process. It extracts relevant characteristics (environmental sound, volume) from the input speech and serves as a bridge to the characteristic adding unit, which then incorporates these characteristics into the converted speech, enabling the system to utilize input characteristics without compromising recognition accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Device complexity

If speech signal is processed without considering volume and blocking characteristics, then processing is simplified, but accuracy of speech conversion deteriorates

Engineering Contradiction:
Improveprocessing complexityVSAvoidspeech conversion accuracy
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The system performs preliminary estimation of speech characteristics (volume, environmental sound, signal blocking) before the main speech conversion process. The characteristic estimation unit analyzes the input speech signal in advance to extract these characteristics, which are then stored and applied during conversion, ensuring accurate speech conversion without adding complexity to the main processing flow.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system changes parameters of the speech signal based on estimated characteristics. The characteristic adding unit modifies the converted speech by adjusting parameters such as volume and adding environmental sound characteristics, thereby improving speech conversion accuracy while maintaining a relatively simple processing architecture through parameter-based adjustments rather than complex structural changes.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS8793128B2Speech signal processing system, speech signal processing method and speech signal processing method program using noise environment and volume of an input speech signal at a time point
Publication Date: 2014.07.29 NEC CORP
  • US8793128B2 patent drawing
  • US8793128B2 patent drawing
  • US8793128B2 patent drawing

AI summary

A speech signal processing system that includes a speech input unit for inputting a speech signal; input speech storage unit for storing an input speech signal that is the speech signal inputted through the speech input unit; characteristic estimation unit for referring to the input speech signal stored in the input speech storage unit, and estimating characteristics of an input speech indicated by the input speech signal, the characteristics including an environmental sound included in the input speech signal; reference speech output unit for causing a predetermined speech signal that becomes a reference speech, to output; and characteristic adding unit for adding the characteristics of the input speech estimated by the characteristic estimation unit, in a reference speech signal that is the speech signal caused to output by the reference speech output unit.