Subdomain Terminology Extraction via Language Model Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing language models struggle to automatically identify and differentiate business-specific terminology across various subdomains within a language domain, leading to inefficiencies in understanding and retrieval of domain-specific information.

Innovation Solution

A method that involves creating language models for specific domains, identifying common terminology, subtracting it to find subdomain-specific terms, calculating a priori appearance probabilities, and determining similar meanings across subdomains using word lattices and embeddings to automatically extract and rank subdomain-specific terminology.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a general language model is used for a language domain, then it can handle common terminology across all subdomains, but it cannot accurately identify or differentiate subdomain-specific terminology

Engineering Contradiction:
Improveterminology identification accuracyVSAvoidlanguage model structure
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the language model into multiple specialized models, each dedicated to a specific subdomain. This segmentation allows each model to focus on subdomain-specific terminology while maintaining high accuracy for that particular domain, resolving the contradiction between identification accuracy and model complexity by distributing the complexity across multiple specialized models rather than one complex general model

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts subdomain-specific terminology from the general language model by identifying and isolating terms unique to each subdomain. This extraction process creates specialized vocabulary sets that can be applied to enhance the general model's performance in specific subdomains without requiring complete model restructuring

Inventive Principle:
Principle #2Taking out (Extraction)

2Reliability

If separate glossaries are created for each subdomain, then subdomain-specific terminology can be accurately captured, but the system complexity and maintenance burden increase significantly

Engineering Contradiction:
Improvesubdomain terminology accuracyVSAvoidglossary management system
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges multiple subdomain-specific glossaries into a unified glossary structure that automatically integrates terminology from different subdomains. This unified approach maintains the accuracy of subdomain-specific terms while reducing system complexity by providing a single management interface and consistent structure across all subdomains

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a universal glossary framework that serves multiple subdomains simultaneously. This multi-functional glossary system can adapt to different subdomain requirements while maintaining a consistent structure, reducing the need for separate management systems for each subdomain and lowering overall maintenance burden

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If manual terminology extraction is performed for each subdomain, then high precision terminology can be obtained, but the time and resources required increase substantially

Engineering Contradiction:
Improveterminology extraction accuracyVSAvoidterminology extraction speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent implements self-service terminology extraction through automated algorithms that independently identify and extract subdomain-specific terminology from text corpora. The system uses statistical methods and pattern recognition to automatically distinguish subdomain terms from general terms without human intervention, maintaining high precision while dramatically improving extraction speed and reducing resource requirements

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent employs parameter changes in statistical algorithms to automatically adjust extraction thresholds and criteria based on subdomain characteristics. By dynamically modifying extraction parameters, the system adapts to different subdomains and maintains high precision terminology extraction across diverse domains without requiring manual tuning for each subdomain

Inventive Principle:
Principle #35Parameter changes

4Measurement precision

If domain-specific language models are created for each subdomain, then language processing accuracy improves, but the computational resources and development time required increase

Engineering Contradiction:
Improvelanguage processing accuracyVSAvoidmodel development time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary action by pre-processing and organizing text corpora from multiple subdomains before model creation. This preliminary organization includes identifying common terminology, structuring subdomain-specific content, and preparing training data in advance, which significantly reduces the time required for actual model development and deployment while maintaining high language processing accuracy

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11741310B2Automatic discovery of business-specific terminology
Publication Date: 2023.08.29 VERINT AMERICAS INC
  • US11741310B2 patent drawing
  • US11741310B2 patent drawing

AI summary

An IVR and chatbot, or other system, employing a language model, the language model resulting from a method and computer product encoding the method is available for preparing a domain or subdomain specific glossary. The method included using probabilities, word context, common terminology and different terminology to identify domain and subdomain specific language and a related glossary updated according to the method.