LLM-Labeled Neural Networks for Categorical Text Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems face challenges in categorizing large volumes of natural language text data due to the variability in human expression and the lack of high-quality labels for neural networks, leading to incomplete data analysis and missed insights in contexts like customer interactions and multimedia content.

Innovation Solution

Utilizing Large Language Models (LLMs) to identify concepts and generate labels, combined with clustering and neural networks trained on these labels to categorize concepts into categories and sub-categories, enabling comprehensive extraction and classification of intents, reasons, and actions in transcripts.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual data entry is used in CRM tools, then data can be recorded, but data completeness and accuracy deteriorate due to human error and omission

Engineering Contradiction:
Improvedata completenessVSAvoiddata extraction efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system enables self-service by allowing the LLM to automatically extract and categorize concepts from transcripts without manual intervention. The neural networks autonomously process the data, identifying intents, reasons, and actions while categorizing them into company-specific taxonomies, eliminating the need for manual data entry and improving both completeness and efficiency.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical system of manual data entry with an automated intelligent system. The LLM and neural networks substitute human agents in the data extraction process, using natural language processing to identify and categorize concepts from transcripts, thereby improving reliability while maintaining high productivity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If neural networks are trained without high-quality labels, then training can proceed, but classification accuracy deteriorates due to lack of supervised learning signals

Engineering Contradiction:
Improvetraining speedVSAvoidclassification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system applies preliminary action by first using the LLM to generate high-quality labeled training data from transcripts before training the neural networks. The LLM extracts concepts and categorizes them into company-specific taxonomies, creating a supervised learning dataset that enables accurate classification while maintaining efficient training processes.

Inventive Principle:
Principle #10Preliminary action

3Loss of information

If LLM is used to identify concepts from transcripts, then concept extraction completeness improves, but processing time increases due to analysis of entire transcript

Engineering Contradiction:
Improveconcept extraction completenessVSAvoidprocessing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system applies segmentation by dividing the transcript analysis into distinct conceptual categories (intents, reasons, actions) that are extracted and processed separately. The LLM identifies these specific concept types from the transcript, and the neural networks further categorize them into company-specific taxonomies, enabling complete extraction while optimizing processing efficiency through structured segmentation.

Inventive Principle:
Principle #1Segmentation

4Adaptability or versatility

If company-specific categories are used for classification, then relevance to organizational needs improves, but system complexity increases due to custom taxonomy mapping

Engineering Contradiction:
Improveorganizational relevanceVSAvoidtaxonomy mapping complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system uses an intermediary approach by employing the LLM as a bridge between the transcript data and company-specific categories. The LLM generates initial concept extractions and categorizations, which are then mapped to the organization's custom taxonomy through neural networks. This intermediary processing layer simplifies the complexity of direct mapping while maintaining high organizational relevance.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250217603A1Large language model and neural networks for categorical classification of natural language text
Publication Date: 2025.07.03 THE BANK OF NEW YORK MELLON
  • US20250217603A1 patent drawing
  • US20250217603A1 patent drawing
  • US20250217603A1 patent drawing

AI summary

The disclosure relates to systems and methods of identifying concepts in content having natural language text using a Large Language Model (LLM), training neural networks in a discovery phase to classify the identified concepts into categories, sub-categories, or other groupings of concepts, and executing the neural networks in an operational phase to classify identified concepts.