Automated Text Classification via Cosine Similarity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manual annotation of large datasets for training deep learning models is time-consuming, expensive, and prone to errors, especially when experts are not available to label feedback effectively for text classification tasks.

Innovation Solution

A method using statistical methods to automatically label user feedback by generating machine-readable vectors from sample keywords and phrases, employing cosine similarity scores to classify feedback into categories without human intervention, allowing for efficient training of neural networks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual annotation is used to label training data, then classification accuracy can be maintained, but time consumption and cost increase significantly

Engineering Contradiction:
Improveclassification accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs self-service by automatically generating labels through statistical methods and cosine similarity calculations without requiring human annotators. The feedback is labeled by comparing it with sample feedback using vector similarity, enabling the system to annotate its own training data at scale.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical human annotation process with an automated computational system. Instead of manual reading and classification by experts, the system uses statistical methods, vector representations, and cosine similarity algorithms to automatically generate labels, eliminating the need for human labor in the annotation process.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If manual annotation is used to label training data, then label quality can be ensured, but cost increases significantly

Engineering Contradiction:
Improvelabel qualityVSAvoidcost
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The system performs self-service by automatically generating labels through statistical methods and cosine similarity calculations without requiring human annotators. The feedback is labeled by comparing it with sample feedback using vector similarity, enabling the system to annotate its own training data at scale.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical human annotation process with an automated computational system. Instead of manual reading and classification by experts, the system uses statistical methods, vector representations, and cosine similarity algorithms to automatically generate labels, eliminating the need for human labor in the annotation process.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Manufacturing precision

If extensive manual labeling is performed, then training data quality improves, but productivity decreases

Engineering Contradiction:
Improvetraining data qualityVSAvoidmodel training speed
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The system performs self-service by automatically generating labels through statistical methods and cosine similarity calculations without requiring human annotators. The feedback is labeled by comparing it with sample feedback using vector similarity, enabling the system to annotate its own training data at scale.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary action by pre-computing vector representations of sample feedback and establishing similarity thresholds before actual labeling occurs. This preparation enables rapid automated labeling of new feedback without requiring real-time human intervention, thus improving training data generation speed.

Inventive Principle:
Principle #10Preliminary action

4Productivity

If automated labeling is used, then productivity increases, but measurement precision may decrease

Engineering Contradiction:
Improvelabeling speedVSAvoidclassification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent replaces the mechanical human annotation process with an automated computational system. Instead of manual reading and classification by experts, the system uses statistical methods, vector representations, and cosine similarity algorithms to automatically generate labels, eliminating the need for human labor in the annotation process.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system uses parameter changes by adjusting the similarity threshold and controlling the number of iterations in the labeling process. By optimizing these parameters, the system achieves high labeling speed while maintaining sufficient accuracy for training purposes, balancing productivity and precision.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20230162021A1Text classification using one or more neural networks
Publication Date: 2023.05.25 NVIDIA CORP
  • US20230162021A1 patent drawing
  • US20230162021A1 patent drawing
  • US20230162021A1 patent drawing

AI summary

Apparatuses, systems, and techniques are presented to generate one or more images. In at least one embodiment, one or more neural networks are used to generate information about a computer program based, at least in part, on unannotated user feedback.