Multi-Speaker Capture Testing via Dual-Phase Signal Assessment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for testing the performance of devices in multi-party conferencing scenarios, such as conference phones, face limitations in objectively measuring performance with artificial signals, which fail to reflect subjective quality and may not accurately represent real-world usage.

Innovation Solution

A method involving a reference phase where speech signals are applied at different angles, followed by a test phase where both signals are applied simultaneously, using a quality assessment model like POLQA to generate performance indices that simulate concurrent speech and account for device-specific processing characteristics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional artificial signals are used for testing device performance, then testing can be performed objectively and efficiently, but the measurement results fail to reflect subjective quality and do not accurately represent real-world usage

Engineering Contradiction:
Improvetesting efficiencyVSAvoidaccuracy of performance assessment
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent applies preliminary action by pre-processing speech signals to simulate various acoustic scenarios (single talker, multi-talker, different room acoustics) before testing the device. This allows the testing system to present realistic speech conditions in advance, making the performance assessment more accurate while maintaining testing efficiency. The pre-processed signals incorporate characteristics of real-world usage including speech variability and acoustic environment effects.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses copying by creating synthetic speech signals that replicate the statistical and spectral characteristics of real human speech. Instead of using actual human speakers during testing, the system generates artificial speech signals that copy the essential features of real speech (phonetic content, spectral shape, temporal structure), thereby achieving both objective measurability and realistic assessment of device performance in multi-talker scenarios.

Inventive Principle:
Principle #26Copying

2Measurement precision

If speech signals are applied separately at different angles in reference phase, then device response to individual speakers can be measured, but the interaction between simultaneous speakers cannot be captured

Engineering Contradiction:
Improveaccuracy of individual speaker measurementVSAvoidability to handle concurrent speech
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies segmentation by dividing the testing process into distinct phases: a reference phase where individual speech signals are applied separately from different angles to establish baseline device response, and a test phase where multiple speech signals are applied simultaneously to assess multi-talker performance. This segmented approach allows precise measurement of individual speaker handling while also evaluating the device's ability to manage concurrent speech interactions.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamics by transitioning from static testing conditions (separate signal application in reference phase) to dynamic testing conditions (simultaneous signal application in test phase). The system dynamically adjusts the testing scenario to match real-world conferencing situations where multiple speakers talk concurrently, thereby improving the device's adaptability assessment while maintaining measurement precision through the structured two-phase approach.

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If a quality assessment model is applied to combined signals, then perceptual quality metrics can be obtained, but the device-specific processing characteristics are not accounted for

Engineering Contradiction:
Improveperceptual quality metric accuracyVSAvoiddevice performance representation
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent applies feedback by using the quality assessment model results to iteratively refine the testing approach and device configuration. The perceptual quality metrics obtained from applying the model to combined signals during the test phase are fed back to adjust device processing parameters and capture properties. This feedback loop ensures that both perceptual quality accuracy and device-specific characteristics are properly accounted for in the final performance assessment.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10490206B2Testing device capture performance for multiple speakers
Publication Date: 2019.11.26 DOLBY LABORATORIES LICENSING CORP
  • US10490206B2 patent drawing
  • US10490206B2 patent drawing
  • US10490206B2 patent drawing

AI summary

Systems and methods are described for measuring capture performance of multiple voice signals. A first speech signal is applied to a device, and measured at far-end at a far-end of a testing environment. A second speech signal is separately applied to the device, and is also measured at the far end. The measured speech signals are added, and a quality assessment model is applied to the first far-end combined signal to obtain a first quality metric. The first speech signal and the second speech signal are then both applied at the same time to the device and measured at the far-end. The quality assessment model is applied to the second far-end combined signal to obtain a second quality metric. The quality metric for the second far-end combined signal is normalized, based on the first quality metric, to obtain a performance index for the device.