Moral Image Classification via Text-Image Embedding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies for classifying immoral images are limited in their ability to generalize beyond specific areas, such as violence or sexuality, and lack suitable visual learning datasets with universal human morality, making it challenging to classify images at the level of general human intelligence.

Innovation Solution

An apparatus and method utilizing a text encoder and image encoder based on the CLIP model to create textual and visual embedding vectors, which are then processed by a morality classification unit to determine if input texts or images are moral or immoral, using a binary classification approach and a loss function to minimize errors in classification results.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional immoral content classification technologies are used, then classification can be performed for specific areas such as violence and sexuality, but the system cannot classify images at the level of general human intelligence without being limited to specific areas

Engineering Contradiction:
Improveclassification scopeVSAvoidclassification accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent applies universality by training a single deep learning model to perform multiple classification tasks across different moral domains (violence, sexuality, honesty, fairness, etc.) rather than requiring separate models for each specific area. The model learns universal moral judgment patterns that can generalize across diverse image types and moral categories, enabling classification at the level of general human intelligence.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If visual learning datasets are used for training, then the model can learn image classification, but there is no suitable visual learning dataset with universal human morality available

Engineering Contradiction:
Improvelearning capabilityVSAvoiddataset availability
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent uses text descriptions as an intermediary to bridge the gap between available data and moral classification goals. Since suitable visual datasets with moral labels are unavailable, the system employs text-based moral judgments and descriptions as a intermediate representation that can be learned from existing text data, which is then applied to image classification through the multi-modal model.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent performs preliminary training on text data to establish moral judgment capabilities before applying these capabilities to image classification. The model first learns moral concepts and patterns from text datasets, then transfers this learned knowledge to the visual domain, enabling moral image classification without requiring pre-existing labeled visual moral datasets.

Inventive Principle:
Principle #10Preliminary action

3Ease of manufacture

If the system is trained with limited specific area data, then training can be completed, but the model cannot generalize to classify images across all moral domains

Engineering Contradiction:
Improvetraining feasibilityVSAvoidgeneralization capability
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent changes the training parameters and data representation to enable broader generalization. By using text-based moral descriptions and multi-modal input representations, the model learns more abstract and transferable moral concepts that can be applied across different domains, rather than memorizing domain-specific patterns. This parameter change in the learning approach allows the model to generalize from limited training data to diverse moral classification tasks.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240046616A1Apparatus and method for classifying immoral images using deep learning technology
Publication Date: 2024.02.08 KOREA UNIV RES & BUSINESS FOUND
  • US20240046616A1 patent drawing
  • US20240046616A1 patent drawing
  • US20240046616A1 patent drawing

AI summary

Disclosed is an apparatus for classifying immoral images according to one embodiment of the present invention, comprising: a text encoder unit that receives a learning text as an input to create a textual embedding vector; an image encoder unit that receives an image as an input to create a visual embedding vector; and a morality classification unit that receives either the textual embedding vector or the visual embedding vector as an input to create and output a classification result indicating whether the input text or image is moral or immoral, wherein the morality classification unit learns the classification results of the input learning texts only from a learning dataset containing a plurality of learning texts in which whether each learning text is moral or immoral is mapped to a binary class.