Subdomain Terminology Extraction via Language Model Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing language models struggle to automatically identify and differentiate business-specific terminology across various subdomains within a language domain, leading to inefficiencies in understanding and retrieval of domain-specific information.
Innovation Solution
A method that involves creating language models for specific domains, identifying common terminology, subtracting it to find subdomain-specific terms, calculating a priori appearance probabilities, and determining similar meanings across subdomains using word lattices and embeddings to automatically extract and rank subdomain-specific terminology.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a general language model is used for a language domain, then it can handle common terminology across all subdomains, but it cannot accurately identify or differentiate subdomain-specific terminology
Solution Approach 1:
The patent segments the language model into multiple specialized models, each dedicated to a specific subdomain. This segmentation allows each model to focus on subdomain-specific terminology while maintaining high accuracy for that particular domain, resolving the contradiction between identification accuracy and model complexity by distributing the complexity across multiple specialized models rather than one complex general model
Solution Approach 2:
The patent extracts subdomain-specific terminology from the general language model by identifying and isolating terms unique to each subdomain. This extraction process creates specialized vocabulary sets that can be applied to enhance the general model's performance in specific subdomains without requiring complete model restructuring
2Reliability
If separate glossaries are created for each subdomain, then subdomain-specific terminology can be accurately captured, but the system complexity and maintenance burden increase significantly
Solution Approach 1:
The patent merges multiple subdomain-specific glossaries into a unified glossary structure that automatically integrates terminology from different subdomains. This unified approach maintains the accuracy of subdomain-specific terms while reducing system complexity by providing a single management interface and consistent structure across all subdomains
Solution Approach 2:
The patent creates a universal glossary framework that serves multiple subdomains simultaneously. This multi-functional glossary system can adapt to different subdomain requirements while maintaining a consistent structure, reducing the need for separate management systems for each subdomain and lowering overall maintenance burden
3Measurement precision
If manual terminology extraction is performed for each subdomain, then high precision terminology can be obtained, but the time and resources required increase substantially
Solution Approach 1:
The patent implements self-service terminology extraction through automated algorithms that independently identify and extract subdomain-specific terminology from text corpora. The system uses statistical methods and pattern recognition to automatically distinguish subdomain terms from general terms without human intervention, maintaining high precision while dramatically improving extraction speed and reducing resource requirements
Solution Approach 2:
The patent employs parameter changes in statistical algorithms to automatically adjust extraction thresholds and criteria based on subdomain characteristics. By dynamically modifying extraction parameters, the system adapts to different subdomains and maintains high precision terminology extraction across diverse domains without requiring manual tuning for each subdomain
4Measurement precision
If domain-specific language models are created for each subdomain, then language processing accuracy improves, but the computational resources and development time required increase
Solution Approach 1:
The patent performs preliminary action by pre-processing and organizing text corpora from multiple subdomains before model creation. This preliminary organization includes identifying common terminology, structuring subdomain-specific content, and preparing training data in advance, which significantly reduces the time required for actual model development and deployment while maintaining high language processing accuracy
Data Source
AI summary
An IVR and chatbot, or other system, employing a language model, the language model resulting from a method and computer product encoding the method is available for preparing a domain or subdomain specific glossary. The method included using probabilities, word context, common terminology and different terminology to identify domain and subdomain specific language and a related glossary updated according to the method.

