Speech Processing User Identification via Serial Number Association
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Speech recognition systems face challenges in accurately identifying and differentiating between users, particularly children, due to limited vocabulary and speech characteristics, which can lead to incorrect execution of commands or refusal to execute intended actions.
Innovation Solution
Associating a serial number with a user profile indicating a child user, allowing the system to confidently determine the user's identity even with insufficient speech characteristics, and using this association to manage command execution and access control.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If speech recognition systems rely on speech characteristics to identify users, then user identification can be performed without additional data, but children with limited vocabulary and speech characteristics cannot be accurately differentiated from adults
Solution Approach 1:
The patent combines multiple data sources (speech characteristics, vocabulary analysis, and serial number associations) to identify user identity. When speech characteristics are insufficient for accurate identification (as with children), the system merges in additional data from serial numbers associated with devices, thereby resolving the contradiction between maintaining identification accuracy and adapting to different user types.
Solution Approach 2:
The patent introduces serial numbers as an intermediary element that mediates between insufficient speech characteristics and accurate user identification. The serial number associated with a child's device serves as a bridge, allowing the system to confidently determine user identity even when direct speech analysis would be unreliable.
2Productivity
If the system uses speech characteristics alone for user identification, then the process is simple and quick, but it leads to incorrect execution of commands or refusal to execute intended actions for children
Solution Approach 1:
The system performs preliminary user identification using speech characteristics, then validates or corrects this identification by checking associated serial numbers before executing commands. This preliminary action maintains speed while ensuring accuracy, as the serial number verification happens in advance of command execution rather than delaying it.
Solution Approach 2:
The system uses feedback from serial number associations to correct potential misidentifications. When speech-based identification is ambiguous (common with children), the serial number feedback provides definitive user identity information, ensuring commands are executed by the correct user profile.
3Reliability
If the system lacks confidence in user identity determination, then it may refuse to execute commands, but this prevents intended actions from being performed
Solution Approach 1:
The serial number acts as an intermediary that provides definitive user identification when speech characteristics alone are insufficient. This intermediary mechanism increases confidence in user identity determination without creating barriers to command execution, as the serial number lookup is an automated process that does not require additional user input.
Data Source
AI summary
Techniques for implementing a “volatile” user ID are described. A system receives first input audio data and determines first speech processing results therefrom. The system also determines a first user that spoke an utterance represented in the first input audio data. The system establishes a multi-turn dialog session with a first content source and receives first output data from the first content source based on the first speech processing results and the first user. The system causes a device to present first output content associated with the first output data. The system then receives second input audio data and determines second speech processing results therefrom. The system also determines the second input audio data corresponds to the same multi-turn dialog session. The system determines a second user that spoke an utterance represented in the second input audio data and receives second output data from the first content source based on the second speech processing results and the second user. The system causes the device to present second output content associated with the second output data.


