Dynamic Semantic Class Expansion for Toxic Content Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing social media networks face challenges in effectively detecting and categorizing toxic speech within vast volumes of user-generated content, relying on human moderators and basic automated systems.

Innovation Solution

The technology expands semantic classes through user feedback by starting with a predefined set of labels and dynamically refining them based on user-generated tags. When a threshold of similarly tagged content items contains similar terminology, the system consolidates this terminology into new semantic classes, providing a more granular representation of content item classification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If basic automated systems and human moderators are used to detect toxic speech, then the system can handle the volume of content, but the detection precision and categorization accuracy remain insufficient

Engineering Contradiction:
Improvedetection precisionVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments toxic speech detection into multiple hierarchical semantic classes (e.g., toxic, toxic-racist, toxic-gun-violence). Each class represents a specific category of toxic content, allowing the system to detect and categorize different types of toxic speech with increasing precision through progressive classification levels.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system implements feedback loops where user reports and moderator classifications are used to continuously refine and expand the semantic class hierarchy. User feedback on flagged content informs the creation of new sub-classes, enabling the system to adapt and improve detection precision over time based on real-world data.

Inventive Principle:
Principle #23Feedback

2Adaptability or versatility

If a fixed set of labels is used for content classification, then the system structure remains simple, but the system cannot detect nuanced forms of toxic speech

Engineering Contradiction:
Improveclassification versatilityVSAvoidclassification structure
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The semantic class hierarchy is designed to be dynamic rather than static. New sub-classes can be created and added to the hierarchy based on emerging patterns in user-generated content and feedback. This allows the classification system to adapt to new forms of toxic speech while maintaining the existing structure.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent implements a nested hierarchical structure where semantic classes are organized in levels of granularity. For example, 'toxic' contains sub-classes like 'toxic-racist' and 'toxic-gun-violence', which can themselves contain further sub-classes. This nested structure allows the system to maintain simplicity at higher levels while providing detailed classification capabilities at lower levels.

Inventive Principle:
Principle #7Nested doll (Nesting)

3Measurement precision

If manual review by human moderators is used for all content, then classification accuracy improves, but the productivity and scalability of the system decreases

Engineering Contradiction:
Improveclassification accuracyVSAvoidcontent processing throughput
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system applies automated detection and classification to the majority of content using the semantic class hierarchy, reserving human moderator review for cases that require nuanced judgment or represent new categories. This partial automation approach maintains high productivity while ensuring accuracy for critical cases.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system enables user self-service through automated reporting and classification features where users can flag and categorize toxic content themselves. This distributes the moderation workload to the user base, maintaining classification accuracy through community involvement while preserving system productivity by reducing the burden on professional moderators.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12229839B2Expanding semantic classes via user feedback
Publication Date: 2025.02.18 DISCORD INC
  • US12229839B2 patent drawing
  • US12229839B2 patent drawing
  • US12229839B2 patent drawing

AI summary

The present technology extends to methods, systems, and computer program products for expanding semantic classes via user feedback. Aspects of the technology learn how a set of labels can be expanded from user-generated tags. Text labels applied by human reviewers to digital content can be inspected and compared to one another. When a threshold of human-generated text tags contain similar terminology, the set of labels can be expanded to define a representation of the similar terminology. Similar terminology can include terms that originate from the same base term, are synonyms, are more specific terms related to a general term category, etc. Similar terminology can be consolidated into a defining term that is used to generate a new (more granular) label or a new top level label. Accordingly, new semantic classes can be discovered from user-generated feedback. New semantic classes can provide a more granular representation of content item classification.