Audio Command Processing With Speaker Position Verification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio processing systems fail to execute appropriate processing when the position of the speaker cannot be identified, leading to potential unintended operations.
Innovation Solution
An audio processing system that includes a processor and memory, capable of receiving a voice command, determining speaker position identification, and limiting output of speaker position commands when identification is not possible, thereby preventing unintended processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If the audio processing system processes voice commands based on speaker position, then the system can execute location-specific operations, but unintended processing may occur when speaker position cannot be identified
Solution Approach 1:
The system dynamically adjusts its operation mode based on speaker position identification results. When position can be identified, it executes location-specific operations; when position cannot be identified, it switches to a safe mode that processes only position-independent commands. This dynamic adaptation resolves the contradiction by making the system versatile when conditions permit while maintaining reliability when conditions are uncertain.
Solution Approach 2:
The system changes the operational parameters of voice command processing based on the identified speaker position. Different position parameters trigger different processing behaviors: when position is known, the system uses position-aware processing with higher adaptability; when position is unknown, it switches to position-agnostic processing with higher reliability. This parameter-based control resolves the technical contradiction.
2Reliability
If the system limits output of speaker position commands when position cannot be identified, then unintended operations are prevented, but system functionality is reduced
Solution Approach 1:
The system applies partial action by selectively limiting only the speaker position command output when position cannot be identified, while still allowing other types of voice commands to be processed normally. This resolves the contradiction by preventing unintended operations through selective restriction rather than complete system shutdown, thus maintaining overall system functionality while ensuring reliability.
Solution Approach 2:
The voice command processing is segmented into different types: position-dependent commands and position-independent commands. When speaker position cannot be identified, the system segments the processing by allowing only position-independent commands to execute while blocking position-dependent ones. This segmentation approach prevents unintended operations while preserving essential system functionality.
3Speed
If the system executes all voice commands without position verification, then processing speed is maintained, but unintended processing may occur
Solution Approach 1:
The system performs preliminary verification of speaker position before executing voice commands that depend on position information. This preliminary action occurs rapidly in the background, allowing the system to maintain high processing speed by pre-determining which commands can be executed immediately and which require position verification, thus resolving the contradiction between speed and reliability.
Data Source
AI summary
An audio processing system includes a memory, and a processor. The processor is coupled to the memory, and, when executing a program stored in the memory, performs: receiving a first voice that is a voice uttered by a first speaker; determining whether or not a position of the first speaker can be identified; and outputting, to a target device, a voice command that is specified by a voice and is a signal for controlling the target device, the processor limiting output of a speaker position command related to a position of a speaker among the voice command in a case where the processor determines that the position of the first speaker cannot be identified.


