Automated Offensive Action Detection in Visual Sessions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing visual communication systems face challenges in consistently detecting and addressing offensive actions within mixed reality, augmented reality, and virtual reality sessions, as subjective determinations can lead to inconsistent enforcement and negative stigma against reporting objectionable behavior, resulting in unresolved offensive actions that affect other users.
Innovation Solution
A processing system that passively observes participants' activities to detect objectionable actions using machine learning algorithms and action detection models, which can identify patterns in visual features and wearable device inputs, allowing for automated filtering and modification of visual content to block or modify offensive actions without manual annotation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual annotation and human review are used to detect offensive actions, then measurement precision can be maintained, but productivity decreases and loss of time increases
Solution Approach 1:
The patent replaces manual human review processes with automated machine learning-based detection systems. The processing system uses action detection models that analyze visual content from communication sessions to automatically identify offensive actions, eliminating the need for manual annotation while maintaining detection accuracy and significantly improving processing speed.
Solution Approach 2:
The system performs self-service by automatically detecting, analyzing, and addressing offensive actions without requiring human intervention. The action detection models continuously monitor visual content, identify patterns associated with offensive behavior, and execute modifications autonomously based on predefined policies and user configurations.
2Productivity
If automated detection systems are implemented, then productivity increases, but device complexity increases
Solution Approach 1:
The patent segments the detection system into multiple independent action detection models, each specialized in identifying specific types of offensive actions. This modular architecture allows the system to handle complex detection tasks by combining simpler, specialized models rather than using a single complex monolithic system.
Solution Approach 2:
The processing system is designed as a universal platform that can detect multiple types of offensive actions through different action detection models. The same core system infrastructure supports various detection functions by loading different models, reducing overall system complexity compared to having separate specialized systems for each detection task.
3Reliability
If human reviewers investigate complaints, then reliability can be maintained through subjective judgment, but loss of time increases and productivity decreases
Solution Approach 1:
The system incorporates feedback mechanisms where users can report offensive actions, and the action detection models use this feedback to refine their detection accuracy. The processed visual content and detection results provide continuous feedback loops that improve the system's reliability over time while maintaining rapid processing speeds.
Solution Approach 2:
The action detection models perform preliminary analysis of visual content in real-time during communication sessions, identifying potential offensive actions before they escalate or affect other users. This proactive detection eliminates the need for post-event human investigation, reducing time loss while maintaining consistent determination through algorithmic analysis.
Data Source
AI summary
A processing system having at least one processor may establish a communication session between a first communication system of a first user and a second communication system of a second user, the communication session including first visual content, the first visual content including a first visual representation of the first user, and detecting a first action of the first visual representation in the first visual content in accordance with a first action detection model. The processing system may modify, in response to the detecting the first action, the first visual content in accordance with a first configuration setting of the first user for the communication session, which may include modifying the first action of the first visual representation of the first user in the first visual content. In addition, the processing system may transmit the first visual content that is modified to the second communication system of the second user.


