Vocal-Characteristic Models for Multiuser Session Emotional State Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Multiuser virtual environment services face challenges in managing anonymous interactions, leading to online disinhibition and emotional distress due to aggressive or rude behavior, as participants may not be aware of the harm caused by their actions.

Innovation Solution

Deploying vocal-characteristic models to analyze audio streams and generate probability scores for emotional states like fear, sadness, anger, or disgust, enabling proactive remedial actions to improve user experience and reduce toxicity in multiuser sessions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If in-session voice chat functionality is enabled to enhance player immersion, then user engagement and immersion are improved, but online disinhibition and emotional distress increase due to anonymous aggressive behavior

Engineering Contradiction:
Improveuser immersionVSAvoidemotional distress
Core Design Contradiction:
Adaptability or versatilityVSObject-affected harmful factors

Solution Approach 1:

The system implements real-time feedback by analyzing audio streams during voice chat to detect emotional states (fear, sadness, anger, disgust) and immediately responding with remedial actions such as muting aggressive participants or providing support notifications, creating a closed-loop system that continuously monitors and adjusts the social environment

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent introduces an intermediary moderation system that acts between participants in voice chat, using vocal-characteristic models to detect emotional distress and automatically intervene through remedial actions, thereby mediating the harmful interactions without requiring direct human moderation

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If anonymous participation is allowed to enable free expression, then ease of operation is improved, but harmful behavior increases due to lack of accountability

Engineering Contradiction:
Improveease of participationVSAvoidtoxic behavior
Core Design Contradiction:
Ease of operationVSObject-generated harmful factors

Solution Approach 1:

The system enables self-service moderation by automatically analyzing participant vocal characteristics and detecting emotional states without human intervention, allowing the system to self-regulate toxic behavior through automated remedial actions while maintaining anonymous participation

Inventive Principle:
Principle #25Self-service

3Measurement precision

If real-time audio analysis is performed to detect emotional states, then user experience monitoring is improved, but system complexity and processing requirements increase

Engineering Contradiction:
Improveemotional state detection accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces complex human moderation mechanics with automated vocal-characteristic models that analyze audio streams using machine learning techniques, substituting the mechanical system of human judgment with an automated computational system that detects emotional states through vocal patterns

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Object-affected harmful factors

If remedial actions are automatically performed based on detected emotional states, then emotional harm reduction is improved, but loss of user autonomy increases due to system intervention

Engineering Contradiction:
Improveemotional harmVSAvoiduser autonomy
Core Design Contradiction:
Object-affected harmful factorsVSLoss of information

Solution Approach 1:

The system applies partial intervention by performing remedial actions selectively based on detected emotional states rather than continuously monitoring and intervening in all interactions, using targeted actions such as muting only when specific emotional distress indicators are detected, thereby minimizing disruption to user autonomy while still providing protection

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11170800B2Adjusting user experience for multiuser sessions based on vocal-characteristic models
Publication Date: 2021.11.09 MICROSOFT TECHNOLOGY LICENSING LLC
  • US11170800B2 patent drawing
  • US11170800B2 patent drawing
  • US11170800B2 patent drawing

AI summary

Techniques for adjusting user experiences for participants of a multiuser session by deploying vocal-characteristic models to analyze audio streams received in association with the participants are disclosed herein. The vocal-characteristic models are used to identify emotional state indicators corresponding to certain vocal properties being exhibited by individual participants. Based on the identified emotional state indicators, probability scores are generated indicating a likelihood that individual participants are experiencing a predefined emotional state. For example, a specific participant's voice may be continuously received and analyzed using a vocal-characteristic model designed to detect whether vocal properties are consistent with a predefined emotional state. Probability scores may be generated based on how strongly the detected vocal properties correlate with the vocal-characteristic model. Responsive to the probability score that results from the vocal-characteristic model exceeding a threshold score, some remedial action may be performed with respect to the specific participant that is experiencing the predefined emotional state.