Text Normalization via Statistical and Human-in-the-Loop Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems fail to effectively normalize text from various sources, leading to inconsistencies in language modeling, as they do not account for domain-specific vocabularies and evolving internet language, which hampers downstream analysis and machine learning performance.

Innovation Solution

A system and method combining statistical models, data models, and human-in-the-loop (HITL) normalization, where terms are extracted, searched, and prioritized based on context and probability, with human intervention to add new terms to a digitized data model, ensuring dynamic and consistent text representation across sources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If text is gathered from various sources to create a language model, then the vocabulary coverage is improved, but inconsistencies in terminology and syntax across domains worsen

Engineering Contradiction:
Improvevocabulary coverageVSAvoidterminology consistency
Core Design Contradiction:
Adaptability or versatilityVSStability of the object's composition

Solution Approach 1:

The patent introduces a normalization layer as an intermediary between text gathering and language model construction. This normalization process standardizes terminology and syntax from various sources before they are incorporated into the language model, thereby maintaining vocabulary diversity while ensuring consistency. The normalization module acts as a mediator that transforms heterogeneous text data into a unified format.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Stability of the object's composition

If manual normalization is performed to ensure terminology consistency, then terminology consistency is improved, but the time and effort required worsen

Engineering Contradiction:
Improveterminology consistencyVSAvoidnormalization time
Core Design Contradiction:
Stability of the object's compositionVSLoss of time

Solution Approach 1:

The system implements self-service through automated normalization algorithms that operate without continuous human intervention. The language model construction process automatically handles terminology standardization by gathering text from multiple sources and applying consistent normalization rules, eliminating the need for manual review of each term while maintaining high consistency standards.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent applies preliminary normalization actions during the text gathering phase rather than performing manual normalization later. By pre-processing text data to standardize terminology before model construction, the system eliminates the need for subsequent manual intervention, thereby maintaining consistency while reducing time investment.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If domain-specific vocabularies are incorporated to improve accuracy, then the domain accuracy is improved, but the complexity of managing multiple vocabularies worsens

Engineering Contradiction:
Improvedomain accuracyVSAvoidvocabulary management complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent creates a universal normalization framework that handles multiple domain-specific vocabularies through a single unified process. The language model construction system is designed to accommodate various domains (insurance, telecommunications, etc.) using the same normalization methodology, thereby maintaining domain accuracy while avoiding the complexity of separate management systems for each vocabulary.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11599723B2System and method of combining statistical models, data models, and human-in-the-loop for text normalization
Publication Date: 2023.03.07 VERINT AMERICAS INC
  • US11599723B2 patent drawing
  • US11599723B2 patent drawing

AI summary

According to principles described herein, unsupervised statistical models, semi-supervised data models, and HITL methods are combined to create a text normalization system that is both robust and trainable with a minimum of human intervention. This system can be applied to data from multiple sources to standardize text for insertion into knowledge bases, machine learning model training and evaluation corpora, and analysis tools and databases