Audio Processing Device Pitch Distribution Suppression

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio processing techniques fail to accurately estimate the impression of an utterance due to hesitation, as they incorrectly evaluate bright voices as dark when the speaker hesitates, due to an improper specification of the spread range of pitch frequency distribution.

Innovation Solution

An audio processing device that detects acoustic feature amounts, calculates a coefficient based on time change, and generates a statistical amount to accurately estimate the impression by suppressing the spread of the pitch frequency distribution during periods of hesitation, using weight coefficients to differentiate between first and second pitch frequencies.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If the spread range of pitch frequency distribution is calculated using all detected pitch frequencies, then the calculation is simple, but the estimation accuracy of utterance impression deteriorates when hesitation is present

Engineering Contradiction:
Improvecalculation simplicityVSAvoidutterance impression estimation accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent segments the pitch frequency distribution into two groups: first pitch frequencies (detected during periods with small time change amounts) and second pitch frequencies (detected during periods with large time change amounts). By separating these segments and applying different weight coefficients, the system achieves accurate impression estimation while maintaining computational feasibility through structured processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different weight coefficients to different segments of pitch frequency data based on local characteristics. Specifically, a first weight coefficient is applied to pitch frequencies detected during stable periods (small time change), while a second weight coefficient is applied to pitch frequencies detected during unstable periods (large time change). This local differentiation resolves the contradiction by improving accuracy without requiring complete recalculation of all data.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If weight coefficients are applied to differentiate between pitch frequencies detected during different time change periods, then the estimation accuracy improves, but the device complexity increases

Engineering Contradiction:
Improveutterance impression estimation accuracyVSAvoidprocessing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces dynamic weight coefficients that automatically adjust based on the time change amount of pitch frequency. The system dynamically determines which weight coefficient to apply by comparing the time change amount against a threshold, enabling adaptive processing that improves accuracy without requiring complex manual configuration or multiple separate processing systems.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system performs self-service by automatically determining the appropriate weight coefficient based on the inherent characteristics of the pitch frequency data itself. The time change amount of pitch frequency serves as the criterion for selecting weight coefficients, eliminating the need for external intervention or complex decision-making mechanisms, thus improving accuracy while maintaining relatively simple system architecture.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS10832687B2Audio processing device and audio processing method
Publication Date: 2020.11.10 FUJITSU LTD
  • US10832687B2 patent drawing
  • US10832687B2 patent drawing
  • US10832687B2 patent drawing

AI summary

There is provided an audio processing device including a memory, and a processor coupled to the memory and the processor configured to detect a first acoustic feature amount and a second acoustic feature amount of an input audio, calculate a coefficient for the second acoustic feature amount based on a time change amount by calculating the time change amount of the first acoustic feature amount, and calculate a statistical amount for the second acoustic feature amount based on the coefficient.