Token Entropy Analysis for Vocabulary Consistency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Inconsistency of token units in text corpora for Asian languages impairs the quality of language models used in natural language processes like speech recognition and machine translation.

Innovation Solution

A method that splits or merges tokens in an initial vocabulary based on entropy calculations in a bootstrap language model, determining whether to delete or add tokens to improve vocabulary consistency, involving a processor or programmable circuitry to execute instructions for token processing, updating, and re-aligning text corpora.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If token units are kept as-is in the initial vocabulary, then the vocabulary structure is simple and processing is fast, but token inconsistency impairs language model quality

Engineering Contradiction:
Improvelanguage model qualityVSAvoidvocabulary processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments tokens by splitting inconsistent token units into multiple consistent sub-tokens based on entropy analysis. This segmentation resolves token inconsistency by breaking down ambiguous tokens into well-defined components, thereby improving language model quality while managing vocabulary complexity through structured segmentation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent merges multiple inconsistent token variants into a single unified token representation by analyzing their entropy and co-occurrence patterns. This merging reduces vocabulary complexity by consolidating redundant token forms while maintaining language model quality through the creation of a more consistent token space.

Inventive Principle:
Principle #5Merging (Combining)

2Measurement precision

If tokens are split into multiple sub-tokens to improve consistency, then token representation becomes more precise, but computational resources increase

Engineering Contradiction:
Improvetoken representation precisionVSAvoidcomputational resource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent changes the entropy parameter of token representations to determine splitting decisions. By calculating entropy values and comparing them against thresholds, the system dynamically adjusts token granularity, achieving precise token representation only where necessary and thereby controlling computational resource consumption.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent applies partial splitting actions by selectively dividing only those tokens that exhibit high entropy and inconsistency, rather than uniformly splitting all tokens. This partial action approach achieves sufficient token representation precision for problematic cases while avoiding unnecessary computational overhead from exhaustive splitting.

Inventive Principle:
Principle #16Partial or excessive action

3Stability of the object's composition

If entropy calculations are performed for all tokens, then optimal vocabulary consistency is achieved, but processing time increases

Engineering Contradiction:
Improvevocabulary consistencyVSAvoidprocessing time
Core Design Contradiction:
Stability of the object's compositionVSLoss of time

Solution Approach 1:

The patent performs preliminary entropy calculations on a bootstrap language model before final vocabulary optimization. This preliminary action identifies high-entropy tokens that require splitting or merging, allowing the system to focus subsequent processing only on problematic tokens rather than uniformly processing the entire vocabulary, thereby achieving vocabulary consistency with reduced processing time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent employs self-service mechanisms where the entropy calculation process automatically identifies and flags tokens requiring vocabulary optimization. The system uses its own entropy analysis results to guide subsequent splitting and merging operations, eliminating the need for external manual intervention or exhaustive processing of all tokens, thus balancing vocabulary consistency with processing efficiency.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11276394B2Method for re-aligning corpus and improving the consistency
Publication Date: 2022.03.15 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11276394B2 patent drawing
  • US11276394B2 patent drawing
  • US11276394B2 patent drawing

AI summary

Vocabulary consistency for a language model may be improved by splitting a target token in an initial vocabulary into a plurality of split tokens, calculating an entropy of the target token and an entropy of the plurality of split tokens in a bootstrap language model, and determining whether to delete the target token from the initial vocabulary based on at least the entropy of the target token and the entropy of the plurality of split tokens.