Voice Signal Processing Using Frequency-Domain Phase Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In environments like automobile interiors, it is challenging to distinguish between target and non-target sounds due to sound reflections, making it difficult to accurately identify and separate voice signals using phase difference techniques.

Innovation Solution

A voice signal processing method that converts voice signals to frequency signals, sets coefficients for sound existence and non-existence based on phase differences, and applies suppression coefficients to determine the likelihood of target sounds, allowing for effective identification and suppression of non-target sounds.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If phase difference technique is used to distinguish target sound from non-target sound, then voice signal separation is improved, but measurement precision deteriorates in environments with sound reflections

Engineering Contradiction:
Improvevoice signal separation accuracyVSAvoidsound reflection interference
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent segments the frequency spectrum into multiple frequency components and processes each frequency independently. By dividing the voice signal into frequency bands and calculating phase differences separately for each frequency, the system can identify target sound existence at different frequency levels, making the overall system more robust to reflections that affect specific frequency ranges.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a frequency dimension to the phase difference analysis. Instead of using only temporal phase differences in the time domain, the system transforms signals to the frequency domain and analyzes phase differences across multiple frequencies. This dimensional transformation provides additional information that helps distinguish target sounds from reflections more accurately.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If phase difference is calculated for all frequencies, then target sound detection is improved, but processing complexity increases

Engineering Contradiction:
Improvetarget sound detection accuracyVSAvoidprocessing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies different processing strategies to different frequency regions. By analyzing phase differences locally at each frequency component and comparing them against frequency-specific criteria, the system maintains high detection accuracy while avoiding the need to process all frequencies uniformly, thus reducing overall computational burden.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system calculates phase differences for multiple frequency components (excessive action) but only processes and compares those that fall within the target sound existence region defined by the microphone array geometry. This partial processing approach maintains reliability while reducing unnecessary computations at frequencies unlikely to contain target sound information.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10497380B2Medium for voice signal processing program, voice signal processing method, and voice signal processing device
Publication Date: 2019.12.03 FUJITSU LTD
  • US10497380B2 patent drawing
  • US10497380B2 patent drawing
  • US10497380B2 patent drawing

AI summary

A voice signal processing method includes: converting a first and a second voice signals to a first and a second frequency signals; setting a coefficient of existence representing degree of existence of a target sound and a coefficient of non-existence representing degree of existence of a non-target sound based on a phase difference for each of the predetermined frequencies between the first and the second frequency signals and a target sound existence region indicating an existence position of the target sound; and judging whether the first voice and/or the second voice include the target sound, based on the coefficient of existence, the coefficient of non-existence and a representative value corresponding to either one of the first and the second frequency signals.