Systems and methods for de-entangled word embeddings

CN115204177BActive Publication Date: 2026-09-01BAIDU USA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210344777.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-04-07
Filing Date
2022-03-31
Publication Date
2026-09-01
Estimated Expiration
2042-03-31

AI Technical Summary

Benefits of technology

[0014]本文描述的是改进词表征学习的系统和方法实施例。概率先验的实施例可以将统计解纠缠与词嵌入无缝集成。与以前的确定性方法不同,词嵌入可以被视为概率生成模型,它可以强加先验,所述先验可以识别生成词表征向量的独立因素。概率先验不仅增强了词嵌入的表示,还提高了模型的鲁棒性和稳定性。此外,公开的方法的实施例可以灵活地插入各种词嵌入模型中。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115204177B_ABST
    Figure CN115204177B_ABST
Patent Text Reader

Abstract

A system and method for de-entangled word embeddings are provided. The method includes: given a modified sequence comprising multiple context words surrounding at least one modified word, generating a prior distribution for each context word using a prior; generating word embeddings for each context word conditioned on the corresponding prior distributions using an encoder; and generating a decoder output based on the word embeddings of the multiple context words using a decoder, the decoder output comprising words recovered from at least one modified word.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to systems and methods for computer learning, which can provide improved computer performance, features, and uses. More specifically, this disclosure relates to systems and methods for de-entangled word embeddings. Background Technology

[0002] Deep neural networks have achieved great success in many fields, such as computer vision, natural language processing, and recommender systems.

[0003] Unsupervised word embeddings or item embeddings generate fundamental representations for downstream information retrieval systems or natural language processing models. Recent advances in pre-trained models have significantly improved language processing tasks. While learned word representations can reveal some simple semantic properties, they do not directly provide more information about word representations and are typically used as direct input to downstream black-box neural models.

[0004] Therefore, systems and methods for learning word embeddings are needed to improve performance. Summary of the Invention

[0005] Firstly, a method for untangling word embeddings is provided, including:

[0006] Given a modified sequence of multiple context words surrounding at least one modified word, generate a prior distribution for each context word using the prior.

[0007] Conditioned by the corresponding prior distribution, an encoder is used to generate word embeddings for each context word; and

[0008] The decoder output is generated based on word embeddings of multiple context words, and the decoder output includes a word recovery of at least one altered word.

[0009] In a second aspect, a non-transitory computer-readable medium is provided comprising one or more sequences of instructions, which, when executed by at least one processor, cause the execution of steps of the method for untangling word embeddings as described in the first aspect.

[0010] Thirdly, a system for untangling word embeddings is provided, comprising:

[0011] One or more processors; and

[0012] One or more non-transitory computer-readable media, including one or more sets of instructions, which, when executed by at least one of one or more processors, cause the steps of the method for unentangled word embeddings as described in the first aspect to be performed.

[0013] Fourthly, a computer program product is provided, comprising a computer program that, when executed by a processor, causes the processor to perform the method described in the first aspect.

[0014] This paper describes embodiments of systems and methods for improving word representation learning. Probabilistic prior embodiments seamlessly integrate statistical disentanglement with word embeddings. Unlike previous deterministic methods, word embeddings can be viewed as probabilistic generative models that can impose priors that identify independent factors generating word representation vectors. Probabilistic priors not only enhance the representation of word embeddings but also improve the robustness and stability of the model. Furthermore, embodiments of the disclosed methods can be flexibly inserted into various word embedding models. Attached Figure Description

[0015] Reference will be made to embodiments of this disclosure, examples of which may be illustrated in the accompanying drawings. These drawings are intended to be illustrative and not restrictive. Although this disclosure has been generally described in the context of these embodiments, it should be understood that it is not intended to limit the scope of this disclosure to these particular embodiments. Items in the drawings may not be to scale.

[0016] Figure 1 A network diagram of a model with a de-entanglement prior is depicted according to one or more embodiments of the present disclosure.

[0017] Figure 2 A process for untangling word embeddings according to one or more embodiments of the present disclosure is described.

[0018] Figure 3 An improved deentangled word embedding process according to embodiments of the present disclosure is described.

[0019] Figure 4 The average scores of three models according to embodiments of this disclosure at different times are depicted.

[0020] Figure 5 A simplified block diagram of a computing device / information processing system according to embodiments of the present disclosure is depicted. Detailed Implementation

[0021] In the following description, specific details are set forth for purposes of explanation in order to provide an understanding of this disclosure. However, it will be apparent to those skilled in the art that this disclosure may be practiced without these details. Furthermore, those skilled in the art will recognize that the embodiments of this disclosure described below can be implemented in various ways, such as as processes, apparatuses, systems, devices, or methods on tangible computer-readable media.

[0022] The components or modules shown in the figures are illustrative of exemplary embodiments of this disclosure and are intended to avoid obscuring this disclosure. It should also be understood that throughout the discussion, components can be described as individual functional units, which may include subunits; however, those skilled in the art will recognize that various components or portions thereof may be divided into individual components or may be integrated together, including, for example, in a single system or component. It should be noted that the functions or operations discussed herein can be implemented as components. Components can be implemented in software, hardware, or a combination thereof.

[0023] Furthermore, the connections between components or systems in the diagram are not intended to be limited to direct connections. Instead, data between these components can be modified, reformatted, or otherwise altered by intermediate components. Additionally, more or fewer connections can be used. It should also be noted that the terms “coupled,” “connection,” “communication coupling,” “interface,” “access,” or any derivative thereof should be understood to include direct connections, indirect connections via one or more intermediate devices, and wireless connections. It should also be noted that any communication, such as signals, responses, replies, acknowledgments, messages, queries, etc., can include one or more exchanges of information.

[0024] References to "one or more embodiments," "preferred embodiments," "embodiments," "some embodiments," etc., in the specification mean that a particular feature, structure, characteristic, or function described in connection with an embodiment is included in at least one embodiment of the invention and may be included in more than one embodiment. Furthermore, the above phrases appearing in multiple places in the specification do not necessarily refer to the same one or more embodiments.

[0025] The use of certain terms in different places in this specification is for illustrative purposes and should not be construed as limiting. A service, function, or resource is not limited to a single service, function, or resource; the use of these terms may refer to a group of related services, functions, or resources that may be distributed or integrated. The terms “including,” “contains,” and “have” should be understood as open-ended terms, and any list below is illustrative and does not imply limitation to the listed items. A “layer” may include one or more operations. The terms “optimal,” “optimized,” “optimization,” etc., refer to an improvement in a result or process and do not require that the specified result or process has reached an “optimal” or peak state. The use of memory, database, repository, data storage, table, hardware, cache, etc., in this document may refer to one or more system components that can input or otherwise record information.

[0026] In one or more embodiments, the stopping condition may include: (1) a set number of iterations have been performed; (2) a certain processing time has been reached; (3) convergence (e.g., the difference between successive iterations is less than a first threshold); (4) divergence (e.g., performance degradation); and (5) an acceptable result has been achieved.

[0027] Those skilled in the art should recognize that: (1) certain steps may be selectively performed; (2) the steps may not be limited to the specific order specified herein; (3) certain steps may be performed in different orders; and (4) certain steps may be performed simultaneously.

[0028] Any headings used herein are for organizational purposes only and should not be used to limit the scope of the specification or claims. Every reference / document mentioned in this patent document is incorporated herein by reference in its entirety.

[0029] It should be noted that any experiments and results provided herein are provided in an illustrative manner and were performed under specific conditions using one or more specific embodiments; therefore, these experiments and their results should not be used to limit the scope of disclosure of this patent document.

[0030] A. Overview

[0031] Unsupervised word embeddings or item embeddings provide fundamental representations for downstream information retrieval systems or natural language processing models. Recent advances in pre-trained models have significantly improved language processing tasks. To further improve word representation learning, this patent disclosure proposes an embodiment of neural probabilistic priors to integrate disentangled generative models and word representation learning.

[0032] The following section provides a brief overview of word embedding models and deentanglement.

[0033] 1. Word embedding

[0034] Classical shallow word embedding methods, such as skip-gram, continuous bag of words (CBOW), and GloVe (Global Vectors for Word Representation), learn word embeddings based on the occurrence of words in a sliding window. These self-evident word embedding methods learn word meaning based on co-occurrence and often ignore structural order information in the sequence. The learned word representations can exhibit some simple semantic properties, which are usually used as input to downstream black-box neural network models. Therefore, shallow word embedding models can capture simple semantic information, but they may lose the complex syntactic and semantic information contained in the structure of sentences and corpora.

[0035] Pre-trained word embeddings and language models can overcome the shortcomings of classical word embedding methods. Pre-trained word embedding models can effectively integrate learned prior knowledge with information from the specific task at hand. These models are often able to capture syntactic information from large sentences in a corpus by utilizing recurrent neural networks and / or attention mechanisms. However, pre-trained models are costly to train, due to large training corpora, long computation times, and high financial costs. These costs can also reduce the flexibility of the model in application scenarios when the training corpus or dataset is small.

[0036] To maintain the effectiveness and flexibility of representation learning, researchers have attempted to incorporate syntactic and semantic information into shallow or small models. This structural information can be flexibly encoded in the learned representation using graph convolutional neural networks (GCNs). Some researchers have improved word embeddings by leveraging syntactic and semantic structural information extracted from training corpora. In one or more embodiments disclosed in this patent, a CBOW model equipped with a GCN is used.

[0037] 2. Untangling

[0038] Variations of the variational autoencoder (VAE) are considered state-of-the-art techniques in unsupervised unentanglement learning. Some researchers have proposed β-VAEs, which introduce hyperparameters to the Kullback-Leibler (KL) regularizer of ordinary VAEs to maximize the following objective:

[0039]

[0040] Here, h is the latent representation of x. By setting β > 1, the encoder is forced to match the unified Gaussian prior of the decomposition. This process introduces an additional constraint on the capacity of the potential bottleneck, encouraging the encoder to learn a disentangled representation of the data. Variations of β-VAE have been used to obtain disentangled representations from diverse datasets including images and text.

[0041] Recently, nonlinear independent component analysis (ICA) theory has been used for unentanglement in latent spaces. The theoretical basis of this series of works is that the latent factors of data distribution can be approximated by utilizing weak data labels or data structure information.

[0042] This patent disclosure proposes an embodiment of a model for improving word representations by utilizing statistical disentanglement techniques. Unlike classical deterministic word embedding methods, in this patent disclosure, word embeddings are treated as probabilistic generative models, specifically, conditional variational autoencoders (CVAEs). This approach allows for probabilistic priors, enabling the model to learn independent latent factors that generate embedding vectors. The model can learn disentangled word representations by utilizing nonlinear independent component analysis (ICA). Furthermore, the method can be readily plugged into any word embedding model to improve performance. Experimental results on different evaluation tasks validate the advantages of the proposed method embodiment.

[0043] B. Methodological Examples

[0044] This section introduces word embedding expectations from CVAE. Implementations based on CVAE and nonlinear ICA priors are also disclosed.

[0045] 1. CVAE view of word embedding

[0046] Most classic word embedding methods can be categorized as variations of CVAE. For a text sequence s = [w1, ..., w...], ... n ], It is the altered sequence of s. The term "altered" should be understood to mean that one or more words in the sequence may be corrupted, masked, missing, altered, added to, adjusted, or any combination thereof. In one or more embodiments, the altered sequence It has one or more words masked by a corresponding binary mask m. In one or more embodiments of the invention, y = {s, m} is taken as a relation to... The label information and conditional distribution Modeling is done using a conditional variational autoencoder. When the distribution... When it is multimodal, it outperforms deterministic models. In one or more embodiments, the objective of the CVAE model can be expressed as:

[0047]

[0048] For traditional word embedding methods, h is the sequential concatenation of the embedding vectors of the unmasked words. The encoder (q) φ The indicator function (p) transforms the unmasked tokens into their embedding vectors. The decoder (p) θ Map from h to masked or missing word tags.

[0049] Figure 1 A network diagram of a model 100 for unentangled word embeddings according to one or more embodiments of the present disclosure is depicted. Figure 2The use of one or more embodiments of the present disclosure is described. Figure 1 The model shown illustrates the process of de-entangled word embeddings. For example... Figure 1 As shown, the model includes a prior network 110, an encoder 120, and a decoder 130. In one or more embodiments, the prior network 110 is implemented using one or more fully connected layers with rectified linear unit (ReLU) activation. In one or more embodiments, the prior network 110 is a fully connected network, such as a multilayer perceptron (MLP).

[0050] Given a modified sequence 105 having at least one modified word and including multiple context words (e.g., w1, w2, w4, and w5) surrounding the at least one modified word, a priori network 110 receives the modified sequence 105 to generate a (205) prior distribution 115 for each context word. The phrase “modified” should be understood to encompass embodiments of one or more words in a modified sequence where the corresponding word in the text sequence is disrupted, masked, or altered, one or more words are missing from the modified sequence, or any combination thereof. In one or more embodiments, the modified sequence 105 is obtained from a text sequence (e.g., w1, w2, w3, w4, and w5) having a binary mask (e.g., m) applied to at least one word (e.g., w3). In one or more embodiments, the word covered by the binary mask is not present in the modified sequence 105.

[0051] In one or more embodiments, the prior distribution 115 is a Gaussian distribution parameterized by mean μ and standard deviation σ. Encoder 120 generates (210) word embeddings 125 for each context word, conditioned on the corresponding prior distribution 115. Decoder 130 generates (215) decoder outputs based on the word embeddings of multiple context words, the decoder outputs including at least one recovered word corresponding to at least one modified word. In one or more embodiments, the multiple context words are words within a window of a predetermined size around each modified word. In one or more embodiments, context words can be defined using an analytic graph as described in detail in subsection C. At least one of the prior network, encoder, and decoder can be trained (220) using a target involving at least the decoder output and a ground truth label consisting of a text sequence and a binary mask.

[0052] In classic embedding methods, both the encoder and decoder are deterministic functions. The prior p of h... ψ It is a deterministic indicator function in conventional methods, and the KL term has also disappeared. In one or more embodiments of this disclosure, the prior (p) of h ψThe indicator is replaced with a real distribution function and regularization is applied to the embedding vector. Priors enable the model to achieve disentanglement across embedding entries. The decoder can be implemented using a neural network or a GCN. By leveraging syntactic and semantic structural information, GCNs can learn task-independent word representations. GCNs can flexibly aggregate structural knowledge between words and improve the representation of the learned word representations.

[0053] 2. Implementation Examples of Unentangled Priors

[0054] In this section, an example of the word embedding distribution prior (WEP) of the above CVAE model will be introduced, and the prior can be parameterized using a neural network.

[0055] 2.1 Nonlinear ICA for Unentanglement

[0056] One or more embodiments of this disclosure aim to achieve deentanglement of word embeddings by utilizing sequential label information y = {s, m}. Observing the conditional VAE framework of word embeddings, embodiments of this disclosure use nonlinear ICA to improve word embeddings. For word w t Regarding h t The distribution of can be a factorial member of an exponential family with sufficient statistics for v, labeled w. t As a condition. In one or more embodiments, h t Prior distribution p ψ The general form can be written as:

[0057]

[0058] Here Q i It is the basic metric, Z i It is the normalization constant, T i,j The components of a sufficient statistic, and λ i,j These are the corresponding parameters, depending on w. k In one or more embodiments, It is the output of an arbitrarily complex, unavoidable, and deterministic decoder that moves from the latent space to the data space. In other words, the decoder D maps from the latent space h to masked or missing word tokens. In one or more embodiments, using nonlinear ICA and with sufficient training samples, the conditional VAE can reveal the generated word embedding vector h. t One or more factors, namely,

[0059]

[0060] here and They are learned from the model, and they correspond to the truth values ​​T and h, respectively.t And Q. A is a full-rank matrix, and c is a constant vector. Therefore, as long as there is sufficient information about the conditional prior distribution p ψ (h t |w t Given word tags, the model can linearly recover information about the word embedding vector h. t The basic metric Q and the sufficient statistic T.

[0061] This disclosure presents one or more embodiments of applying a prior distribution to word embeddings via a nonlinear ICA generation framework. A Gaussian distribution can be used as the prior p. ψ (h t |w t ), mean μ and variance σ 2 It can be parameterized by a neural network, where the word tokens (w) t As input, the neural network p ψ It can accumulate knowledge of all words and linearly identify one or more independent latent factors that generate the vocabulary.

[0062] 2.2 Example Implementation Examples

[0063] In one or more embodiments of this disclosure, word embeddings are modeled using CVAE, and context words (tags) are used as labels for the masked sentence sequence. The concatenation of the embedding vectors (h) of the context words can be latent variables of the conditional VAE model. Figure 1 A schematic diagram of a CVAE model according to an embodiment of this disclosure is shown. An exponential family prior can be added to the embedding vector of each word, and the prior can depend on the word's label information. In one or more embodiments, a Gaussian distribution is used as the prior distribution of the embedding vectors. Embodiments of a sequence context model (sliding window) and a graph context case extended from the sequence context model are described below.

[0064] In one or more embodiments, the word w k The context can be defined as For the window size c, the number of positive integers. In one or more embodiments, for words... Regarding the embedding vector h i The prior distribution parameter can be defined as μ(w) t ) and σ(w t ), where μ(w t ) and σ(w t ) are the mean and standard deviation, and they are used to label w. t The neural network used as input is parameterized.

[0065] In one or more embodiments, the prior distribution can be integrated with word embeddings under the previously discussed conditional VAE framework. In one or more embodiments, regarding w k and The goal can be written as:

[0066]

[0067] In equation (5), the objective includes the reconstruction term and the regularization term. It is a deterministic indicator function. The regularization term is KL dispersion, which can be replaced by Gaussian likelihood. Therefore, the objective can be expressed as:

[0068]

[0069] Here, α is a non-negative value that controls the influence of the prior or the weights. Reconstruction term This can be achieved using the loss function of any word embedding model. The distributed parameter neural network μ and σ aggregates information on the structural distribution across all words and reveals subsequent factors in generating word representations. Table 1 gives the prior p... ψ One embodiment of the neural network architecture. In one or more embodiments, the word tag ID can be copied multiple times, for example, 32 times as shown in Table 1, and then used to form the input vector for the prior network. The output of the prior network can be divided into two parts, μ and logσ. In one or more embodiments, the target (6) can have two sets of parameters, θ and ψ, corresponding to the parameters of the decoder and the prior. In one or more embodiments, SynGCN and SemGCN are used for the reconstruction term in (6). In both models, the decoder uses GCN to combine syntactic and semantic structure knowledge.

[0070] Table 1. A priori network structure embodiment

[0071]

[0072] Variations of GCNs are used in different applications. In one or more embodiments, two GCNs, SynGCN and SemGCN (Vashishth et al., Incorporating Syntactic and Semantic Information in Word Embeddings using Graph Convolutional Networks, Proceedings of the 57th ACP Conference on Computational Linguistics, pp. 3308-3318, Florence, Italy, all of which are incorporated herein by reference) are used as the base model. SynGCN can utilize syntactic context for word representation learning, while SemGCN can incorporate semantic knowledge. In one or more embodiments of the invention, word embeddings can be improved by combining graph syntactic and semantic information, and the sliding window of contextual words can be replaced or combined with adjacent words in the contextual syntactic or semantic graph.

[0073] Figure 3 The process of improving disentangled word embeddings using GCN according to embodiments of the present disclosure is described. In one or more embodiments of the present disclosure, target and context embeddings are defined individually as parameters in the disentangled word embedding (DWE) model for each word in the vocabulary. For a given text sequence s = [w1, ..., w n First, an NLP parser is used to extract the (305) dependency resolution graph. An analytical graph consists of a set of vertices. and a set of labeled directed dependency edges ε s Each edge is of the form (w i ,w j ,l ij ), where l ij It is the source vertex (w) i ) to the destination vertex (w j The dependency relationship of words (w). In one or more embodiments, the word (w) k The context of ) is defined (310) as its in Adjacent words in the text, that is,

[0074] In one or more embodiments, word embeddings can be further improved by incorporating semantic knowledge in addition to syntactic information. The decoder learns a corpus-level labeled graph, where words are nodes and edges represent semantic relationships between them from different sources. In one or more embodiments, semantic relationships such as hyponyms, hypernyms, and synonyms can be represented together in a single graph. In one or more embodiments disclosed in this patent, WEPSyn and WEPSem are used to represent the disclosed WEP model, respectively, by incorporating syntactic and semantic information. In one or more embodiments, both WEPSyn and WEPSem can use the same objective (6) and the same prior network shown in Table 1.

[0075] C. Experimental Results

[0076] It should be noted that these experiments and results are provided for illustrative purposes and were performed under specific conditions using one or more specific embodiments; therefore, these experiments and their results should not be used to limit the scope of disclosure of this patent document.

[0077] Implementations of this method were compared with existing methods on various datasets. Table 1 presents an example of a prior network (μ and σ) for both WEPSyn and WEPSem. In one or more embodiments, the value of α in (6) can be varied to adjust the regularization of the prior distribution. In the experiments, α∈{0.5,0.1,1.0e-4,1.0e-6} was used.

[0078] 1. Baseline

[0079] The following baseline methods were considered in one or more experimental settings.

[0080] Word2vec: Continuous bag of words.

[0081] GloVe: A log-bilinear regression model that utilizes global co-occurrence statistics from a corpus.

[0082] Dependency-based syntactic contexts (Deps): A modification to the skip-gram model that uses dependency contexts instead of sequential contexts.

[0083] Extended Dependency Skip (EXT): An extension of Deps that utilizes second-order dependency context features.

[0084] SynGCN: A graph convolution-based method that uses graph convolutional neural networks and syntactic word relations to improve word embeddings.

[0085] SemGCN: A graph convolution-based method that uses graph convolutional neural networks and semantic word relationships to improve word embeddings.

[0086] In one or more evaluations, WEPSyn can be constructed as a SynGCN model combined with a prior network implementation. Similarly, WEPSem can be constructed as a SemGCN model enhanced with a prior network.

[0087] 2. Evaluation Methods

[0088] In one or more evaluations, embodiments of this method are compared to a baseline on the following intrinsic tasks:

[0089] Word similarity is a task that assesses the proximity between semantically similar words. Different methods were compared across different datasets.

[0090] The word analogy task involves predicting word b2 given three words a1, a2, and b1, such that the relation b1:b2 is the same as the relation a1:a2. Implementations of the proposed method are compared with baselines on various datasets to evaluate this task.

[0091] Concept classification involves grouping nominal concepts into natural categories. Evaluations were conducted on multiple datasets in one or more experiments.

[0092] 3. Results

[0093] This section first presents the comparison results between WEPSyn and the baseline, and then the comparison results between WEPSem and SemGCN.

[0094] 3.1WEPSyn

[0095] In one set of experiments, WEPSyn and SynGCN used the same syntactic knowledge extracted from the corpus. In this set of experiments, target and context embeddings were defined for each word in the vocabulary. After preprocessing, a corpus consisting of multiple sentences with multiple tags and multiple syntactic dependencies was used. The average sentence length in the corpus was approximately 20 words. Table 2 shows the performance of the different methods.

[0096] As shown in Table 2, the WEPSyn (Unentangled Word Embeddings) implementation outperforms existing methods on 8 out of 10 tasks. WEPSyn achieves the greatest improvement on the 4 concept classification tasks compared to all baseline methods. This implies that the presented prior significantly enhances the model's ability to capture concept structure. WEPSyn also outperforms the baselines on the three word similarity tasks. In particular, WEPSyn scores higher on all four word similarity tasks compared to SynGCN, which also utilizes word syntactic information. WEPSyn achieves absolute performance improvements of approximately 4.1% and 6.0% in word similarity and concept classification, respectively. For the word analogy task, WEPSyn achieves results comparable to other methods. As mentioned above, WEPSyn can be viewed as a SynGCN model with an inserted prior; therefore, the experimental results in Table 2 demonstrate that the prior helps the model identify potential independent factors in word embeddings. Thus, the WEPSyn model achieves better results than SynGCN on most settings.

[0097] Table 2: WEPSyn's evaluation of performance in word similarity (Spearman relevance), concept classification (cluster purity), and word analogy (Spearman relevance). WEPSyn outperforms existing methods in 8 out of 10 settings.

[0098]

[0099]

[0100] 3.2 WEPSem

[0101] In a set of experiments, WEPSem and SemGCN used the same semantic information. Both WEPSem and SemGCN used hypernyms, hyponyms, antonyms, and synonyms. Both models were trained using initial embeddings from WEPSyn or SynGCN. In Table 3, SG and WS are used to represent SynGCN and WEPSyn, respectively. WEPSem(SG) represents the WEPSem model trained using initial embeddings from SynGCN. Table 3 presents the results for SemGCN(SG), WEPSem(SG), and WEPSem(WS) on three tasks, representing three different evaluation metrics from Section C.2. Table 3 shows that WEPSem(SG) achieved similar performance compared to SemGCN(SG). However, WEPSem(WS) achieved improved performance on all three tasks compared to SemGCN(SG). WEPSem achieved an absolute performance increase of approximately 2.8% overall. WEPSem is equivalent to the SemGCN model using the proposed prior enhancements. It demonstrates that by leveraging syntactic and semantic information, prior implementations can enable models to capture independent latent factors to identify word similarity, concept classification, and word analogy.

[0102] Table 3: Evaluation of WEPSem using pre-trained embeddings

[0103]

[0104] The proposed WEP prior implementation not only enhances the unentanglement of learned word embeddings but also improves the stability of the model. Figure 4 The average scores for word similarity, concept classification, and word analogy for SemGCN(SG)405, WEPSem(SG)410, and WEPSem(WS)415 at different training epochs are shown. The performance of WEPSem(WS) continues to increase with the number of training epochs. Without prior embedding vectors, the performance of SemGCN(SG) decreases, likely due to model overfitting. Figure 4 The proposed WEP prior implementation demonstrates that it can normalize the dynamics of the embedding vectors and improve the training stability of the model.

[0105] D. Some observations

[0106] This disclosure presents one or more system and method embodiments for improving word representation learning. The proposed model embodiments can integrate probabilistic generative models and nonlinear ICA, and equip them with word embedding models. By utilizing nonlinear ICA, the embodiments can enhance word deentanglement representations. Experiments on various test settings validate the advantages of the proposed methods. In addition to GCN models, plug-and-play priors can be integrated with any word or item embedding method to achieve better performance.

[0107] E. Computing System Implementation

[0108] In one or more embodiments, aspects of this patent document may be directed to, may include, or may be implemented on one or more information processing systems (or computing systems). An information processing system / computing system may include any tool or set of tools operable for calculating, determining, classifying, processing, sending, receiving, retrieving, initiating, routing, switching, storing, displaying, communicating, detecting, recording, copying, processing, or utilizing information, intelligence, or data of any form. For example, a computing system may be or may include a personal computer (e.g., a laptop computer), a tablet computer, a mobile device (e.g., a personal digital assistant (PDA), a smartphone, a phablet, a tablet computer, etc.), a smartwatch, a server (e.g., a blade server or a rack server), a network storage device, a camera, or any other suitable device, and may vary in size, shape, performance, functionality, and price. A computing system may include random access memory (RAM), one or more processing resources such as a central processing unit (CPU) or hardware or software control logic, read-only memory (ROM), and / or other types of memory. Additional components of the computing system may include one or more disk drives, one or more network ports for communicating with external devices and various input and output (I / O) devices, such as a keyboard, mouse, stylus, touchscreen, and / or video display. The computing system may also include one or more buses operable to transmit communication between various hardware components.

[0109] Figure 5 A simplified block diagram of an information processing system (or computing system) according to embodiments of the present disclosure is depicted. It should be understood that the functionality shown in system 500 can operate to support various embodiments of the computing system—although it should be understood that the computing system can be configured differently and include different components, including those with fewer or more components, such as… Figure 5 As shown in the image.

[0110] like Figure 5As shown, the computing system 500 includes one or more central processing units (CPUs) 501 that provide computing resources and control the computer. The CPU 501 may be implemented using a microprocessor or the like, and may also include one or more graphics processing units (GPUs) 502 and / or floating-point coprocessors for mathematical calculations. In one or more embodiments, one or more GPUs 502 may be incorporated into a display controller 509, such as part of one or more graphics cards. The system 500 may also include system memory 519, which may include RAM, ROM, or both.

[0111] Multiple controllers and peripherals can also be provided, such as Figure 5 As shown. Input controller 503 represents an interface to various input devices 504, such as a keyboard, mouse, touchscreen, and / or stylus. Computing system 500 may also include a storage controller 507 for interfacing with one or more storage devices 508, each storage device 508 including a storage medium such as magnetic tape or disk, or an optical medium that can be used to record instruction programs for operating systems, utilities, and applications, which may include embodiments of programs implementing various aspects of this disclosure. Storage devices 508 may also be used to store processed data or data to be processed according to this disclosure. System 500 may also include a display controller 509 for providing an interface to a display device 511, which may be a cathode ray tube (CRT) display, a thin-film transistor (TFT) display, an organic light-emitting diode (OLED) display, an electroluminescent panel, a plasma panel, or any other type of display. Computing system 500 may also include one or more peripheral controllers or interfaces 505 for one or more peripheral devices 506. Examples of peripheral devices may include one or more printers, scanners, input devices, output devices, sensors, etc. The communication controller 514 can interface with one or more communication devices 515, enabling the system 500 to connect to remote devices via any of a variety of networks, including the Internet, cloud resources (e.g., Ethernet cloud, Ethernet Fibre Channel (FCoE) / Data Center Bridge (DCB) cloud, etc.), local area network (LAN), wide area network (WAN), storage area network (SAN), or via any suitable electromagnetic carrier signal, including infrared signals. As shown in the depicted embodiment, the computing system 500 includes one or more fans or fan trays 518 and one or more cooling subsystem controllers 517 that monitor the thermal temperature of the system 500 (or its components) and operate the fans / fan trays 518 to aid in temperature regulation.

[0112] In the illustrated system, all major system components can be connected to bus 516, which can represent more than one physical bus. However, the various system components may or may not be physically close to each other. For example, input data and / or output data can be remotely transmitted from one physical location to another. Furthermore, programs implementing various aspects of this disclosure can be accessed from a remote location (e.g., a server) via a network. Such data and / or programs can be transmitted through any of a variety of machine-readable media, including, for example: magnetic media such as hard disks, floppy disks, and magnetic tapes; optical media such as CD-ROMs and holographic devices; magneto-optical media; and hardware devices specifically configured to store or store and execute program code, such as application-specific integrated circuits (ASICs), programmable logic devices (PLDs), flash memory devices, other non-volatile memory (NVM) devices (such as 3D XPoint-based devices), and ROM and RAM devices.

[0113] Various aspects of this disclosure may be encoded on one or more non-transitory computer-readable media having instructions for one or more processors or processing units to result in execution steps. It should be noted that the one or more non-transitory computer-readable media should include volatile and / or non-volatile memory. It should be noted that alternative implementations are possible, including hardware implementations or software / hardware implementations. The functionality of a hardware implementation can be implemented using an ASIC, a programmable array, digital signal processing circuitry, etc. Therefore, the term "means" in any claim is intended to cover both software and hardware implementations. Similarly, the term "computer-readable medium" as used herein includes software and / or hardware, or a combination thereof, on which a program of instructions is contained. In consideration of these alternative implementations, it should be understood that the accompanying drawings and description provide functional information necessary for those skilled in the art to write program code (i.e., software) and / or assemble circuitry (i.e., hardware) to perform the desired processing.

[0114] It should be noted that embodiments of this disclosure may further relate to computer products having a non-transitory tangible computer-readable medium having computer code thereon for performing operations of various computer implementations. The medium and computer code may be those specifically designed and constructed for the purposes of this disclosure, or they may be of types known or available to those skilled in the art. Examples of tangible computer-readable media include, for example: magnetic media such as hard disks, floppy disks, and magnetic tapes; optical media such as CD-ROMs and holographic devices; magneto-optical media; and hardware devices specifically configured to store or store and execute program code, such as application-specific integrated circuits (ASICs), programmable logic devices (PLDs), flash memory devices, other non-volatile memory (NVM) devices (such as 3D XPoint-based devices), and ROM and RAM devices. Examples of computer code include machine code such as that generated by a compiler, and files containing higher-level code executed by a computer using an interpreter. Embodiments of this disclosure may be implemented, in whole or in part, as machine-executable instructions that may reside in program modules executed by a processing device. Examples of program modules include libraries, programs, routines, objects, components, and data structures. In a distributed computing environment, program modules can be physically located in local, remote, or both settings.

[0115] Those skilled in the art will recognize that no computing system or programming language is essential to the practice of this disclosure. They will also recognize that the aforementioned components can be physically and / or functionally separated into modules and / or submodules or combined together.

[0116] Those skilled in the art will understand that the foregoing examples and embodiments are exemplary and do not limit the scope of this disclosure. All arrangements, enhancements, equivalents, combinations, and modifications that will be apparent to those skilled in the art upon reading the specification and studying the accompanying drawings are intended to be included within the true spirit and scope of this disclosure. It should also be noted that the elements of any claim may be arranged differently, including having multiple dependencies, configurations, and combinations.

Claims

1. A method for untangling word embeddings, comprising: Given a modified sequence comprising multiple context words surrounding at least one modified word, the modified sequence is composed of a text sequence having a binary mask over one or more words in a text sequence; A prior distribution for each context word is generated using priors; wherein the prior distribution for each context word has parameters including a mean and a standard deviation, the mean and the standard deviation being parameterized by a neural network that has the ability to accumulate knowledge of all words and linearly identify one or more independent latent factors that generate the vocabulary. Conditioned by the corresponding prior distribution, an encoder is used to generate word embeddings for each context word; and The decoder output is generated based on the word embeddings of multiple context words, and the decoder output includes the word recovered from at least one modified word. The method further includes training at least one of a prior network, an encoder, and a decoder using an objective function that involves at least the decoder output and ground truth labels including text sequences and binary masks; wherein the objective function includes regularization terms and reconstruction terms.

2. The method according to claim 1, wherein, The plurality of context words are within a window of a predetermined size surrounding the at least one modified word.

3. The method according to claim 1, wherein, The multiple context words were selected using the following steps: The dependency parsing graph of the text sequence is extracted using a Natural Language Processing (NLP) parser; and In the dependency resolution graph, the adjacent words of a word are defined as the context words of the word.

4. The method according to claim 3, wherein, The decoder is a graph convolutional network (GCN).

5. The method according to claim 1, wherein, The objective function includes a regularization term and a reconstruction term, which are represented as a Kullback-Leibler (KL) dispersion between the encoder distribution and the prior distribution.

6. The method according to claim 1, wherein, The objective function includes a regularization term and a reconstruction term, which are represented as Gaussian likelihood values, and the Gaussian likelihood values ​​are a function of the mean and standard deviation of the prior distribution.

7. A non-transitory computer-readable medium comprising one or more sequences of instructions, which, when executed by at least one processor, cause the execution of the steps of the method for unentangled word embeddings as claimed in any one of claims 1 to 6.

8. A system for untangling word embeddings, comprising: One or more processors; as well as One or more non-transitory computer-readable media, comprising one or more sets of instructions, which, when executed by at least one of one or more processors, cause the steps of the method for unentangled word embeddings as described in any one of claims 1 to 6 to be performed.

9. A computer program product comprising a computer program that, when executed by a processor, causes the processor to perform the method as described in any one of claims 1 to 6.