Vocal-Characteristic Models for Multiuser Session Emotional State Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Multiuser virtual environment services face challenges in managing anonymous interactions, leading to online disinhibition and emotional distress due to aggressive or rude behavior, as participants may not be aware of the harm caused by their actions.
Innovation Solution
Deploying vocal-characteristic models to analyze audio streams and generate probability scores for emotional states like fear, sadness, anger, or disgust, enabling proactive remedial actions to improve user experience and reduce toxicity in multiuser sessions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If in-session voice chat functionality is enabled to enhance player immersion, then user engagement and immersion are improved, but online disinhibition and emotional distress increase due to anonymous aggressive behavior
Solution Approach 1:
The system implements real-time feedback by analyzing audio streams during voice chat to detect emotional states (fear, sadness, anger, disgust) and immediately responding with remedial actions such as muting aggressive participants or providing support notifications, creating a closed-loop system that continuously monitors and adjusts the social environment
Solution Approach 2:
The patent introduces an intermediary moderation system that acts between participants in voice chat, using vocal-characteristic models to detect emotional distress and automatically intervene through remedial actions, thereby mediating the harmful interactions without requiring direct human moderation
2Ease of operation
If anonymous participation is allowed to enable free expression, then ease of operation is improved, but harmful behavior increases due to lack of accountability
Solution Approach 1:
The system enables self-service moderation by automatically analyzing participant vocal characteristics and detecting emotional states without human intervention, allowing the system to self-regulate toxic behavior through automated remedial actions while maintaining anonymous participation
3Measurement precision
If real-time audio analysis is performed to detect emotional states, then user experience monitoring is improved, but system complexity and processing requirements increase
Solution Approach 1:
The patent replaces complex human moderation mechanics with automated vocal-characteristic models that analyze audio streams using machine learning techniques, substituting the mechanical system of human judgment with an automated computational system that detects emotional states through vocal patterns
4Object-affected harmful factors
If remedial actions are automatically performed based on detected emotional states, then emotional harm reduction is improved, but loss of user autonomy increases due to system intervention
Solution Approach 1:
The system applies partial intervention by performing remedial actions selectively based on detected emotional states rather than continuously monitoring and intervening in all interactions, using targeted actions such as muting only when specific emotional distress indicators are detected, thereby minimizing disruption to user autonomy while still providing protection
Data Source
AI summary
Techniques for adjusting user experiences for participants of a multiuser session by deploying vocal-characteristic models to analyze audio streams received in association with the participants are disclosed herein. The vocal-characteristic models are used to identify emotional state indicators corresponding to certain vocal properties being exhibited by individual participants. Based on the identified emotional state indicators, probability scores are generated indicating a likelihood that individual participants are experiencing a predefined emotional state. For example, a specific participant's voice may be continuously received and analyzed using a vocal-characteristic model designed to detect whether vocal properties are consistent with a predefined emotional state. Probability scores may be generated based on how strongly the detected vocal properties correlate with the vocal-characteristic model. Responsive to the probability score that results from the vocal-characteristic model exceeding a threshold score, some remedial action may be performed with respect to the specific participant that is experiencing the predefined emotional state.


