Voice Processing Apparatus Speaker Association Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice processing systems lack the ability to accurately determine the combination of voice data corresponding to multiple speakers in conversation from recorded voice data, making it difficult for evaluators to identify speaker associations, especially when the data size is large.

Innovation Solution

A voice processing apparatus that includes an acquisition unit to collect voice signals, a detecting unit to calculate signal intensities, and a determining unit to calculate correlation coefficients between signal intensities, determining whether voices are in a conversation state based on specified thresholds, thereby identifying speaker associations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If voice data is continuously recorded to learn communication patterns, then the quantity of voice data increases, but the difficulty of identifying speaker associations increases

Engineering Contradiction:
Improvequantity of voice dataVSAvoiddifficulty of identifying speaker associations
Core Design Contradiction:
Quantity of substanceVSDifficulty of detecting and measuring

Solution Approach 1:

The patent applies preliminary action by calculating correlation coefficients between signal intensities of different voice data in advance. The determining unit computes these correlations proactively before evaluation is needed, storing the results for later use. This allows evaluators to quickly identify speaker associations without manually analyzing large volumes of voice data, resolving the contradiction between data quantity and identification difficulty.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If manual evaluation is used to identify speaker associations, then accuracy can be maintained, but the workload and time consumption increase

Engineering Contradiction:
Improveaccuracy of speaker identificationVSAvoidtime consumption for evaluation
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements self-service by enabling the system to automatically determine speaker associations through correlation coefficient calculations. The determining unit autonomously analyzes signal intensity patterns and identifies which voice data corresponds to speakers in conversation, eliminating the need for manual evaluation. This maintains accuracy while dramatically reducing time consumption and workload.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If correlation coefficient calculation is performed on all voice data combinations, then identification accuracy improves, but computational complexity increases

Engineering Contradiction:
Improveidentification accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the voice data into individual data units, each with its own signal intensity time sequence. The determining unit then calculates correlation coefficients between pairs of these segmented units, comparing signal intensity patterns of each voice data against others. This segmented approach maintains identification accuracy while making the computational process more manageable compared to analyzing all data as a single complex unit.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS9916843B2Voice processing apparatus, voice processing method, and non-transitory computer-readable storage medium to determine whether voice signals are in a conversation state
Publication Date: 2018.03.13 FUJITSU LTD
  • US9916843B2 patent drawing
  • US9916843B2 patent drawing
  • US9916843B2 patent drawing

AI summary

A voice processing apparatus including a memory, and a processor coupled to the memory and the processor configured to acquire a first input signal containing a first voice, and a second input signal containing a second voice, obtain a first signal intensity of the first input signal, and a second signal intensity of the second input signal, specify a correlation coefficient between a time sequence of the first signal intensity and a time sequence of the second signal intensity, determine whether the first voice and the second voice are in the conversation state or not based on the specified correlation coefficient, and output information indicating an association between the first voice and the second voice when it is determined that the first voice and the second voice are in the conversation state.