Multi-Speaker Speech Recognition with Driver-Passenger Role Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional navigation systems fail to differentiate between driver and passenger speech commands, leading to unintended changes in driving environment and potential safety risks, as they treat all commands equally without considering the relationship between speakers.
Innovation Solution
A method and system that analyze conversation voices to determine the relationship between a driver and passenger, allowing personalized service execution and limiting passenger commands based on their role, using semantic analysis and voice features to differentiate and manage speech commands accordingly.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the navigation system processes all speech commands equally without differentiation, then the system is simple to operate, but safety risks increase due to unintended changes in driving environment
Solution Approach 1:
The patent segments speech command processing by speaker role. It divides the system into separate processing paths for driver speech commands and passenger speech commands. The processor identifies which speaker uttered each command and applies different processing rules: driver commands are processed with full authority, while passenger commands are processed with restricted authority and require driver confirmation for certain operations. This segmentation resolves the contradiction by enabling safety improvements through role-based differentiation without requiring complete system redesign.
Solution Approach 2:
The patent implements preliminary action by pre-establishing authority levels and permission sets for different speaker roles before speech command processing begins. The system pre-identifies speakers as driver or passenger and pre-configures which command types each role can execute independently versus which require confirmation. This preliminary role assignment and permission configuration enables safe command processing without adding complex real-time decision-making logic during command execution.
2Adaptability or versatility
If the navigation system treats all speech commands equally, then the system is easy to manufacture, but adaptability decreases as it cannot provide personalized service
Solution Approach 1:
The patent applies local quality by providing different processing qualities to different speech commands based on the speaker's role. Driver speech commands receive full processing authority with immediate execution capability, while passenger speech commands receive restricted processing with confirmation requirements. The system also provides personalized service by associating command responses with the appropriate speaker (e.g., playing music preferred by the passenger when they request it). This local differentiation achieves adaptability and personalization without requiring complete system reconfiguration.
3Reliability
If the system restricts passenger authority for certain commands, then safety improves, but ease of operation decreases
Solution Approach 1:
The patent implements partial action by applying restricted authority selectively to specific command types rather than universally to all passenger commands. Passenger speech commands are categorized: some commands (e.g., music playback, climate control) can be executed directly with partial authority, while other commands (e.g., navigation changes, communication functions) require driver confirmation. This partial restriction approach maintains ease of operation for routine tasks while improving safety for critical functions, resolving the contradiction between safety and operational ease.
Data Source
AI summary
A speech recognizing method in a multi-speaker environment is disclosed. The method may comprise determining a relation type between first and second speakers by analyzing a first conversation voice of the speakers inputted through a microphone of a computing system; receiving a first conversation voice input of the first or second speaker, and determining a speaker role in the relation type by semantic analysis of the first conversation voice; and extracting a voice feature of the first conversation voice and determining a voice feature of the speaker; receiving a second conversation voice input of the first or second speaker, and determining a speaker role of the second conversation voice in the relation type by using the voice feature extracted from the second conversation voice; and determining a personalized service corresponding to the second conversation voice by using the speaker role of the second conversation voice in the relation type.


