Language Data Management via Word Vector Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies face challenges in efficiently managing and utilizing language data for AI training, particularly in determining similarity and categorizing data effectively.
Innovation Solution
A method for managing language data involves generating word vectors based on the number of words in each piece of language data and using the dot product function to measure similarity scores. The system groups word vectors into clusters based on reference values, allowing for categorization and utilization of the data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If language data is managed without classification for AI training, then data processing speed is improved, but data quality and training effectiveness deteriorate
Solution Approach 1:
The patent applies preliminary action by pre-classifying language data into categories and levels before AI training begins. The management server performs classification of language data based on categories (e.g., topic, domain) and levels (e.g., difficulty, specificity) in advance, creating an organized data structure that can be efficiently retrieved during training without requiring real-time classification, thus maintaining both speed and quality
Solution Approach 2:
The patent applies segmentation by dividing language data into distinct categories and levels. Each piece of language data is tagged with category identifiers and level indicators, allowing the system to segment the overall dataset into manageable subsets that can be selectively applied to different AI training scenarios based on specific requirements
2Measurement precision
If similarity determination is performed on all language data without pre-classification, then comprehensive analysis is improved, but processing time increases
Solution Approach 1:
The patent applies preliminary action by pre-determining similarity relationships between language data items and storing these relationships in a data structure. The management server calculates similarity scores between language data items in advance and stores them, so that during AI training, the system can directly retrieve pre-calculated similarity information without performing time-consuming calculations
Solution Approach 2:
The patent applies segmentation by dividing the similarity determination process into manageable segments based on data categories and levels. Instead of calculating similarity across the entire dataset, the system segments the comparison process to only evaluate similarity between items within the same or related categories, reducing the overall computational scope while maintaining accuracy for relevant comparisons
3Adaptability or versatility
If language data is organized in a detailed tree structure with multiple nodes, then data categorization is improved, but system complexity increases
Solution Approach 1:
The patent applies universality by designing a standardized tree structure that serves multiple functions simultaneously. The same tree structure is used for both categorization (organizing data by topic/domain) and hierarchical organization (structuring data by levels such as general to specific). This universal structure can accommodate different types of language data without requiring separate organizational systems, reducing overall system complexity while maintaining high adaptability
Solution Approach 2:
The patent applies local quality by allowing different parts of the tree structure to have specialized properties appropriate to their function. Leaf nodes contain specific language data items with detailed attributes, while intermediate nodes provide categorical groupings. Each node level has optimized properties for its specific purpose, creating a structure that is complex only where necessary while remaining simple in other areas
Data Source
AI summary
In a state in which the language data in a tree structure includes at least one node and the at least one node includes at least one word, a plurality of word vectors including a first word vector and a second word vector is generated by a management server, based on the number of words included in each of a plurality of pieces of language data, and a dot product function of the first word vector and the second word vector is used by the management server to measure a score of similarity between first language data corresponding to the first word vector and second language data corresponding to the second word vector so that language data is managed for determining similarity by a management server.


