Variable-Length Word Embedding for Neural Network Storage Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Machine Learning systems face increased storage requirements and processing times due to the need for large word lengths to accurately represent objects across various contexts, which complicates the mapping of objects to numeric representations.

Innovation Solution

A data structure and neural network system that adjusts the word length of vector representations dynamically based on context, starting with a minimum parameter length and increasing only as necessary to achieve accurate representation, allowing for zero-value parameters to be clustered and stored efficiently.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If large word lengths are used to accurately represent objects across multiple contexts, then representation accuracy is improved, but storage requirements and processing time increase

Engineering Contradiction:
Improverepresentation accuracyVSAvoidstorage requirements
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent implements dynamic word length adjustment where the vector representation size is not fixed but adapts based on the specific object and context. The system starts with a minimum word length and increases it only when necessary to achieve accurate representation, rather than using a uniform large size for all objects. This dynamic approach resolves the contradiction by making representation accuracy context-dependent while minimizing overall storage requirements.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent applies local quality by allowing different objects to have different word lengths suited to their specific representation needs. Instead of uniformly increasing word length for all objects, the system selectively extends vector size only for objects that require it to capture their contextual meanings accurately. This localized adaptation reduces the total quantity of stored data while maintaining necessary representation accuracy for each object.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If large word lengths are used to accurately represent objects across multiple contexts, then representation accuracy is improved, but processing time increases

Engineering Contradiction:
Improverepresentation accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The dynamic word length adjustment mechanism reduces processing time by avoiding unnecessary computations with excessive vector dimensions. The system determines the minimum required word length for each object and context combination, performing only the necessary processing steps rather than uniformly processing all objects with maximum word length. This dynamic adaptation directly reduces computational overhead and processing time.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent applies partial action by using only the necessary portion of the vector representation needed for accurate object identification in each context. Instead of always processing full-dimensional vectors, the system uses partial vector lengths sufficient for the task at hand, reducing the computational workload and processing time while maintaining adequate representation accuracy.

Inventive Principle:
Principle #16Partial or excessive action

3Adaptability or versatility

If fixed large word lengths are used in embedded databases, then all objects can be represented uniformly, but storage efficiency decreases

Engineering Contradiction:
Improveuniform representation capabilityVSAvoidstorage efficiency
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent replaces fixed uniform word length with dynamic adjustment, where each object's vector representation size is determined by its specific needs. The system maintains adaptability by being able to represent any object with appropriate word length while improving storage efficiency through this variable-length approach, eliminating the waste inherent in fixed-size allocations.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameter of word length from a fixed constant to a variable parameter that adapts to each object's representation requirements. This parameter change enables the system to maintain versatility in representing diverse objects while significantly improving storage efficiency by allocating vector space only where necessary, rather than reserving space for all possible contexts uniformly.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11586652B2Variable-length word embedding
Publication Date: 2023.02.21 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11586652B2 patent drawing
  • US11586652B2 patent drawing
  • US11586652B2 patent drawing

AI summary

A data structure is used to configure and transform a computer machine learning system. The data structure has one or more records where each record is a (vector) representation of a selected object in a corpus. One or more non-zero parameters in the records define the selected object and the number of the non-zero parameters define a word length of the record. One or more zero-value parameters are in one or more of the records. The word length of the object representation varies, e.g. can increase, as necessary to accurately represent the object within one or more contexts provided during training of a neural network used to create the database, e.g. as more and more contexts are introduced during the training. A minimum number of non-zero parameters are needed and zero-value parameters can be clustered together and compressed to save large amounts of system storage and shorten execution times.