Multi-Speaker Capture Testing via Dual-Phase Signal Assessment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for testing the performance of devices in multi-party conferencing scenarios, such as conference phones, face limitations in objectively measuring performance with artificial signals, which fail to reflect subjective quality and may not accurately represent real-world usage.
Innovation Solution
A method involving a reference phase where speech signals are applied at different angles, followed by a test phase where both signals are applied simultaneously, using a quality assessment model like POLQA to generate performance indices that simulate concurrent speech and account for device-specific processing characteristics.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional artificial signals are used for testing device performance, then testing can be performed objectively and efficiently, but the measurement results fail to reflect subjective quality and do not accurately represent real-world usage
Solution Approach 1:
The patent applies preliminary action by pre-processing speech signals to simulate various acoustic scenarios (single talker, multi-talker, different room acoustics) before testing the device. This allows the testing system to present realistic speech conditions in advance, making the performance assessment more accurate while maintaining testing efficiency. The pre-processed signals incorporate characteristics of real-world usage including speech variability and acoustic environment effects.
Solution Approach 2:
The patent uses copying by creating synthetic speech signals that replicate the statistical and spectral characteristics of real human speech. Instead of using actual human speakers during testing, the system generates artificial speech signals that copy the essential features of real speech (phonetic content, spectral shape, temporal structure), thereby achieving both objective measurability and realistic assessment of device performance in multi-talker scenarios.
2Measurement precision
If speech signals are applied separately at different angles in reference phase, then device response to individual speakers can be measured, but the interaction between simultaneous speakers cannot be captured
Solution Approach 1:
The patent applies segmentation by dividing the testing process into distinct phases: a reference phase where individual speech signals are applied separately from different angles to establish baseline device response, and a test phase where multiple speech signals are applied simultaneously to assess multi-talker performance. This segmented approach allows precise measurement of individual speaker handling while also evaluating the device's ability to manage concurrent speech interactions.
Solution Approach 2:
The patent implements dynamics by transitioning from static testing conditions (separate signal application in reference phase) to dynamic testing conditions (simultaneous signal application in test phase). The system dynamically adjusts the testing scenario to match real-world conferencing situations where multiple speakers talk concurrently, thereby improving the device's adaptability assessment while maintaining measurement precision through the structured two-phase approach.
3Measurement precision
If a quality assessment model is applied to combined signals, then perceptual quality metrics can be obtained, but the device-specific processing characteristics are not accounted for
Solution Approach 1:
The patent applies feedback by using the quality assessment model results to iteratively refine the testing approach and device configuration. The perceptual quality metrics obtained from applying the model to combined signals during the test phase are fed back to adjust device processing parameters and capture properties. This feedback loop ensures that both perceptual quality accuracy and device-specific characteristics are properly accounted for in the final performance assessment.
Data Source
AI summary
Systems and methods are described for measuring capture performance of multiple voice signals. A first speech signal is applied to a device, and measured at far-end at a far-end of a testing environment. A second speech signal is separately applied to the device, and is also measured at the far end. The measured speech signals are added, and a quality assessment model is applied to the first far-end combined signal to obtain a first quality metric. The first speech signal and the second speech signal are then both applied at the same time to the device and measured at the far-end. The quality assessment model is applied to the second far-end combined signal to obtain a second quality metric. The quality metric for the second far-end combined signal is normalized, based on the first quality metric, to obtain a performance index for the device.


