Voice Processing Device Speaker Position Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice processing systems struggle to accurately identify and separate voice signals from multiple speakers in a shared space and assign authority levels to these signals, leading to inefficiencies in processing and control.

Innovation Solution

A voice processing device equipped with a voice data receiving circuit, wireless signal receiving circuit, memory, and processor that generates terminal position data from wireless signals, matches this data with terminal IDs, and separates voice signals based on speaker positions and authority levels.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If a microphone receives all voices from multiple speakers in a shared space, then the microphone can capture all speaker voices, but it becomes difficult to separate and identify which speaker each voice signal comes from

Engineering Contradiction:
Improvevoice signals capturedVSAvoidspeaker identification information
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent segments the mixed voice signals by spatial location, dividing the capture space into multiple zones with microphones positioned at different locations. Each microphone captures voices from its specific zone, allowing separation of speaker signals based on their spatial distribution and enabling identification of which speaker produced each voice signal.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If the system separates voice signals by speakers, then speaker identification is achieved, but additional processing complexity is required to determine speaker positions and match signals to speakers

Engineering Contradiction:
Improvespeaker position identificationVSAvoidsignal processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces terminal IDs as an intermediary element that links speaker terminals to their position information and voice signals. The terminal ID serves as a mediator that connects the physical speaker position with the captured voice signal, simplifying the matching process by providing a direct identifier rather than requiring complex analysis of signal characteristics to determine speaker identity.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If the system processes separated voice signals with authority levels, then control security is improved, but additional processing steps are required to determine authority levels for each speaker

Engineering Contradiction:
Improvecontrol authority verificationVSAvoidsignal processing efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent assigns authority levels to speaker terminals in advance, before voice signal processing occurs. By pre-configuring the authority levels associated with each terminal ID and speaker position, the system eliminates the need for real-time analysis of speaker credentials during voice processing, thereby maintaining security verification while improving processing efficiency.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20230260509A1Voice processing device for processing voices of speakers
Publication Date: 2023.08.17 AMOSENSE CO LTD
  • US20230260509A1 patent drawing
  • US20230260509A1 patent drawing
  • US20230260509A1 patent drawing

AI summary

Disclosed is a voice processing device. The voice processing device comprises: a voice data reception circuit configured to receive input voice data associated with the voice of a speaker; a wireless signal reception circuit configured to receive a wireless signal including a terminal ID from a speaker terminal of the speaker; a memory; and a processor configured to generate terminal location data indicating the location of the speaker terminal on the basis of the wireless signal, and match and store the generated terminal location data and the terminal ID in the memory, wherein the processor uses the input voice data to generate first speaker location data and first output voice data associated with a first voice spoken at the first location and matches a first terminal ID corresponding to the first speaker location data and the first output voice data.