Word Embedding Model Training Using Semantic Windows
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing word embedding methods are inefficient and ineffective for training models on skill words without sequential relationships and fail to utilize correspondence information between different corpora, leading to suboptimal semantic expression and computation.
Innovation Solution
A method for training a word embedding model that generates reduced-dimensional word vectors by using a semantic window defined across multiple corpora, incorporating correspondence information to improve semantic expression and computation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional word embedding methods are used on skill words without sequential relationships, then the model can process skill words, but the training efficiency is low and the semantic expression ability is weak
Solution Approach 1:
The patent segments the training process into two distinct phases: pre-training on general corpora to learn basic language patterns, and then fine-tuning on skill word corpora to specialize in skill-related semantics. This segmentation allows the model to efficiently learn general language skills first, then focus computational resources on the specific skill word domain, resolving the contradiction between training efficiency and semantic expression ability.
Solution Approach 2:
The patent performs preliminary pre-training on general corpora before training on skill word corpora. This preliminary action equips the model with basic language understanding and processing capabilities, so that when the model subsequently trains on skill words, it can converge faster and achieve better semantic representation without wasting computational resources on learning basic language patterns during the skill-specific training phase.
2Reliability
If word embedding is performed without utilizing correspondence information between corpora, then the training process is simple, but the semantic computability across different corpora is poor
Solution Approach 1:
The patent makes the word embedding model universal across different corpora by training it on multiple corpora with the same architecture and objective function. The model learns corpus-independent skill word representations that can be applied across different domains and text types. This multi-functionality allows the same model to process skill words from various corpora while maintaining consistent semantic computability, without requiring separate models for each corpus.
Solution Approach 2:
The patent merges multiple corpora into a unified training framework, combining general corpora and skill word corpora in a multi-task learning setup. By merging the training objectives and sharing parameters across different corpus types, the model learns integrated representations that capture both general language patterns and domain-specific skill semantics, thereby achieving good semantic computability across corpora while managing complexity through a unified architecture.
3Loss of information
If high-dimensional word vectors are used, then complete word information is preserved, but the computational burden and storage requirements increase
Solution Approach 1:
The patent extracts only the essential and discriminative features from high-dimensional word vectors during the training process. Rather than preserving all dimensions equally, the model learns to identify and retain the most informative dimensions that capture skill word semantics, effectively extracting the core information while discarding redundant or noisy dimensions. This extraction process maintains word information completeness for skill-related semantics while reducing computational burden.
Solution Approach 2:
The patent changes the parameter representation from fixed high-dimensional one-hot encodings to learned low-dimensional dense vectors. By transforming the parameter space and learning optimal dimensionality during training, the model achieves a balance between preserving necessary word information and reducing computational requirements. The learned embeddings compress the information efficiently, maintaining semantic completeness for skill words while significantly reducing the computational burden compared to raw high-dimensional representations.
Data Source
AI summary
A method of training a model, a method of determining a word vector, a device, a medium, and a product are provided, which may be applied to fields of natural language processing, information processing, etc. The method includes: acquiring a first word vector set corresponding to a first word set; and generating a reduced-dimensional word vector for each word vector in the first word vector set based on a word embedding model, generating, for other word vector in the first word vector set, a first probability distribution in the first word vector set based on the reduced-dimensional word vector, and adjusting a parameter of the word embedding model so as to minimize a difference between the first probability distribution and a second probability distribution for the other word vector determined by a number of word vector in the first word vector set.


