Token Entropy Analysis for Vocabulary Consistency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Inconsistency of token units in text corpora for Asian languages impairs the quality of language models used in natural language processes like speech recognition and machine translation.
Innovation Solution
A method that splits or merges tokens in an initial vocabulary based on entropy calculations in a bootstrap language model, determining whether to delete or add tokens to improve vocabulary consistency, involving a processor or programmable circuitry to execute instructions for token processing, updating, and re-aligning text corpora.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If token units are kept as-is in the initial vocabulary, then the vocabulary structure is simple and processing is fast, but token inconsistency impairs language model quality
Solution Approach 1:
The patent segments tokens by splitting inconsistent token units into multiple consistent sub-tokens based on entropy analysis. This segmentation resolves token inconsistency by breaking down ambiguous tokens into well-defined components, thereby improving language model quality while managing vocabulary complexity through structured segmentation.
Solution Approach 2:
The patent merges multiple inconsistent token variants into a single unified token representation by analyzing their entropy and co-occurrence patterns. This merging reduces vocabulary complexity by consolidating redundant token forms while maintaining language model quality through the creation of a more consistent token space.
2Measurement precision
If tokens are split into multiple sub-tokens to improve consistency, then token representation becomes more precise, but computational resources increase
Solution Approach 1:
The patent changes the entropy parameter of token representations to determine splitting decisions. By calculating entropy values and comparing them against thresholds, the system dynamically adjusts token granularity, achieving precise token representation only where necessary and thereby controlling computational resource consumption.
Solution Approach 2:
The patent applies partial splitting actions by selectively dividing only those tokens that exhibit high entropy and inconsistency, rather than uniformly splitting all tokens. This partial action approach achieves sufficient token representation precision for problematic cases while avoiding unnecessary computational overhead from exhaustive splitting.
3Stability of the object's composition
If entropy calculations are performed for all tokens, then optimal vocabulary consistency is achieved, but processing time increases
Solution Approach 1:
The patent performs preliminary entropy calculations on a bootstrap language model before final vocabulary optimization. This preliminary action identifies high-entropy tokens that require splitting or merging, allowing the system to focus subsequent processing only on problematic tokens rather than uniformly processing the entire vocabulary, thereby achieving vocabulary consistency with reduced processing time.
Solution Approach 2:
The patent employs self-service mechanisms where the entropy calculation process automatically identifies and flags tokens requiring vocabulary optimization. The system uses its own entropy analysis results to guide subsequent splitting and merging operations, eliminating the need for external manual intervention or exhaustive processing of all tokens, thus balancing vocabulary consistency with processing efficiency.
Data Source
AI summary
Vocabulary consistency for a language model may be improved by splitting a target token in an initial vocabulary into a plurality of split tokens, calculating an entropy of the target token and an entropy of the plurality of split tokens in a bootstrap language model, and determining whether to delete the target token from the initial vocabulary based on at least the entropy of the target token and the entropy of the plurality of split tokens.


