Word embedding vector extraction method and system based on large language model

By embedding global context attention modules and dynamic semantic coding fusion mechanisms in each layer of self-attention network of large language model, the problem of difficulty in capturing dynamic context and long-range global semantics in the prior art is solved, and a more refined semantic expression of word embedding vectors is realized.

CN120218061APending Publication Date: 2025-06-27DATA TRANSMISSION GRP
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510237694.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Existing word embedding methods are difficult to capture dynamic context information and long-range global semantics, resulting in the inadequate expression of word embeddings.

Method used

Embed the global context attention module in each layer of the self-attention network of the large language model, combines the cross-position global attention mechanism and dynamic semantic coding fusion mechanism, and generates enhanced global semantic coding through multi-head attention joint coding and dynamic attention weight optimization.

Benefits of technology

It significantly improves the global semantic expression ability of word embedding vectors, enhances the long-range dependency modeling ability, and realizes the flexibility of cross-layer semantic fusion and the optimization of dynamic attention weights.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure QLYQS_3
    Figure QLYQS_3
  • Figure QLYQS_4
    Figure QLYQS_4
  • Figure QLYQS_7
    Figure QLYQS_7
Patent Text Reader

Abstract

The invention discloses a word embedding vector extraction method and system based on a large language model, and belongs to the field of natural language processing, and the word embedding vector extraction method based on the large language model comprises the following steps: S1, embedding a global context attention module, generating a Q / K / V vector through linear transformation, and calculating a global attention weight in a cross-position manner; s2, splicing coding features of adjacent layers, and performing gating fusion to generate a global context dynamic semantic state; s3, multi-head attention joint coding is carried out on the dynamic semantic state of the current layer and the global state; s4, performing average pooling to generate a global semantic vector, and calculating a dynamic attention weight; s5, performing weighted fusion on the dynamic weight and the global vector to generate an enhanced code; s6, outputting a final word vector by a GELU nonlinear transformation full-connection layer; the method has the beneficial effects that the global context semantic capture capability is enhanced, and the expression effect of the word embedding vector is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and more specifically, to a method and system for extracting word embedding vectors based on a large language model. Background Art

[0002] Traditional word embedding methods (such as Word2Vec, GloVe) generate word vectors through static mapping, but cannot capture dynamic context information. Large language models based on Transformer (such as GPT, BERT) extract word embeddings through self-attention mechanisms, but their self-attention layers usually focus on local dependencies and have limited ability to model long-range global semantics. In addition, existing methods have deficiencies in multi-level semantic fusion, dynamic weight allocation, etc., resulting in insufficiently fine semantic expressions of word embeddings. Summary of the Invention

[0003] The present invention proposes a method and system for extracting word embedding vectors based on a large language model, aiming to improve the global semantic expression ability of word embedding vectors by introducing technologies such as a global context attention module, a dynamic semantic coding fusion mechanism, and multi-head attention joint coding.

[0004] Technical Solution: A method for extracting word embedding vectors based on a large language model includes the following steps:

[0005] S1. Embed a global context attention module in each layer of the self-attention network layer of the GPT model, generate query, key, and value vectors through linear transformation, and calculate global attention weights through an intra-layer cross-position attention mechanism.

[0006] S2. Based on the generated global semantic coding, perform feature dimension splicing on the dynamic semantic coding states of adjacent self-attention network layers, and fuse them through a gating mechanism to generate a global context dynamic semantic coding state.

[0007] S3. Based on the global context dynamic semantic coding state, use 4-8 independent attention heads to perform multi-head attention joint coding on the current layer dynamic semantic coding state and the global dynamic semantic coding state.

[0008] S4. Perform average pooling on the dynamic semantic coding states of each token in the jointly encoded sequence to generate a global semantic coding vector, and calculate dynamic attention weights.

[0009] S5. Weightedly fuse the dynamic weights with the global semantic coding vector to generate an enhanced global semantic coding.

[0010] S6. Perform a non-linear transformation on the enhanced global semantic coding through a fully connected layer containing a Gaussian error linear unit activation to generate and extract the final word embedding vector.

[0011] Preferably, S1 specifically includes the following steps:

[0012] S1-1. Perform a linear transformation on the current layer's dynamic semantic encoding state to generate a query vector, a key vector, and a value vector, where the dimension of the key vector is 1 / 8 of the dimension of the dynamic semantic encoding layer;

[0013] Specifically, the dynamic semantic encoding state of the current layer is where n is the sequence length and d is the model dimension. Through three independent trainable weight matrices W Q 、W K 、W V perform a linear projection to generate a query vector Q, a key vector K, and a value vector V: Q = XW Q , K = XW K , V = XW V where the dimension of the key vector is compressed to while the query vector and the value vector maintain the original dimension d.

[0014] S1-2. Calculate the similarity matrix between the query vector and the key vector through the scaled dot-product attention mechanism and normalize it using the Softmax function;

[0015] The calculation formula of the similarity matrix is: where the scaling factor is used to control the magnitude of the dot product and prevent gradient vanishing. Then, apply the Softmax function to the similarity matrix row by row to generate the attention weight matrix: A = Softmax(A). This operation maps the similarity to a probability distribution, representing the global dependence strength between tokens.

[0016] S1-3. Weightedly sum the normalized attention weights and the value vector to generate the global semantic encoding, and input it into the next layer after residual connection with the original dynamic semantic encoding state. The residual connection is used to retain the feature representation ability of the original dynamic semantic encoding state.

[0017] The generation formula of the global semantic encoding is: where the encoding of each token aggregates the semantic information of the entire sequence. Then, perform residual superposition of the global context representation and the original input to retain the underlying semantic information: The output Y is used as the input state of the next layer.

[0018] Preferably, S2 specifically includes the following steps:

[0019] S2-1. Concatenate the global semantic encoding generated by S1, the current layer's dynamic semantic encoding state, and the previous layer's dynamic semantic encoding state along the feature dimension to form a fused state with doubled dimension;

[0020] The formula for three-way feature splicing is as follows: where Z is the global semantic encoding generated at the l-th layer, and X l and X l-1 are the dynamic semantic encodings of the current layer and the previous layer respectively. By splicing the global encoding, the current layer encoding, and the previous layer encoding, a fusion state containing cross-layer global context and local dynamic features is formed.

[0021] S2-2. Perform a non-linear mapping on the fusion state through a trainable gating parameter matrix to generate dynamic fusion weights, where the generation of the gating matrix includes cross-layer attention queries of the global semantic encoding;

[0022] The formula for generating the gating weights is as follows:

[0023] G = σ(FW F + ZW Z )

[0024] where W F and W Z are the gating parameter matrices for the fusion features and the global semantic encoding respectively, σ is the Sigmoid function, and the output range is [0,1]. Each row of the dynamic weight G contains 3 elements, corresponding to the fusion weights of the global encoding, the current layer, and the previous layer respectively

[0025] S2-3. Perform a weighted sum on the dynamic semantic encoding states of the previous layer and the current layer according to the dynamic fusion weights to generate the global dynamic semantic encoding state;

[0026] The formula for the weighted sum is as follows:

[0027] Y = LayerNorm(g1⊙Z + g2⊙X l + g3⊙X l-1 )

[0028] where g1, g2, and g3 are the dynamic fusion weights of the global encoding, the current layer, and the previous layer respectively, and LayerNorm is the layer normalization operation.

[0029] Preferably, S3 specifically includes the following steps:

[0030] S3-1. Use 4 - 8 independent attention heads to process the current layer and the global dynamic semantic encoding state respectively;

[0031] Each independent attention head i performs the following operations: - linearly project to generate query, key, and value vectors:

[0032] Cross-attention calculation:

[0033]

[0034] Among them, d h = d / h, where h is the number of attention heads

[0035] S3-2. The output dimension of each attention head is an equal division of the dynamic semantic encoding layer dimension, and the calculation results of each head are concatenated along the feature dimension;

[0036] The formula for concatenating the multi-head outputs is:

[0037]

[0038] Among them, W O is a trainable projection matrix to ensure alignment with the input dimension after dimensionality reduction.

[0039] S3-3. Reduce the dimension of the concatenated result through a trainable projection matrix, and perform residual connection and layer normalization with the dynamic semantic encoding state of the current layer.

[0040] The formulas for residual connection and layer normalization are:

[0041] Y = LayerNorm(O + X l )

[0042] Preferably, the said S4 specifically includes the following steps:

[0043] S4-1. Perform average pooling on the dynamic semantic encoding states of all tokens in the sequence to generate a global semantic encoding vector;

[0044] The formula for average pooling is:

[0045]

[0046] S4-2. Introduce a learnable temperature coefficient to adjust the similarity distribution between the dynamic semantic encoding state of the current token and the global context vector;

[0047] - The formula for similarity calculation is: Among them, τ is the temperature coefficient, with an initial value of 0.5, and is optimized through backpropagation.

[0048] S4-3. Generate a normalized dynamic attention weight through the Softmax function, where the initial value of the temperature coefficient is 0.5 and is automatically optimized during training.

[0049] The formula for generating the dynamic attention weight is:

[0050]

[0051] Preferably, the said S5 specifically includes the following steps:

[0052] S5-1. Multiply the dynamic attention weights and the global semantic encoding vector element-wise to generate a weighted context representation;

[0053] The formula for the weighted context representation is:

[0054] c i = α i ⊙ G;

[0055] S5-2. Concatenate the weighted context representation and the current token dynamic semantic encoding state along the feature dimension to form a fused input vector.

[0056] The formula for the fused input vector is:

[0057]

[0058] Preferably, S6 specifically includes the following steps:

[0059] S6-1. Input the fused input vector into a fully connected layer, and use the Gaussian Error Linear Unit (GeLU) as the activation function;

[0060] The non-linear transformation formula of the fully connected layer is:

[0061]

[0062] S6-2. Perform a residual connection between the output of the fully connected layer and the original token dynamic semantic encoding state to generate the final word embedding vector.

[0063] The formula for the residual connection and the output is:

[0064] e i = LayerNorm(h i + y i )

[0065] A word embedding vector extraction system based on a large language model, comprising:

[0066] An output module for text tokenization and generating initial word embedding vectors; a global context processing module including cascaded enhanced attention encoding network layers, each layer configured with cross-layer attention units and gated fusion units; a dynamic weight calculation module for calculating token-level dynamic attention weights; a word embedding generation module including a fully connected network activated by the Gaussian Error Linear Unit (GeLU) for fusing context information and dynamic semantic encoding states; an output module for outputting the final word embedding vectors as a dense matrix in sequence order;

[0067] Preferably, in the enhanced self-attention module:

[0068] The dimension of the key vector of the cross-layer attention unit is set to 1 / 8 of the model, and the dimension of the key vector is equal to the dimension of the value vector; the gated fusion unit generates dynamic weights through the Sigmoid function and combines them with the residual connection to retain the original semantic information.

[0069] Preferably, the word embedding generation module adopts a parameter-sharing architecture, where the weight matrix and bias parameters of all fully connected layers are shared among each token in the input sequence, and the same weighted fusion and non-linear transformation operations are performed on each token.

[0070] Compared with the prior art, the advantages of the present invention are as follows:

[0071] (1) The global context modeling ability is significantly enhanced: Traditional self-attention mechanisms (such as GPT, BERT) focus on local dependencies and have limited ability to capture long-range semantics. The present invention embeds a global context attention module in each layer of the self-attention network, combines the cross-position global attention mechanism with the key vector dimension compression technology, effectively reduces the computational complexity, and at the same time retains the original semantic information through the residual connection, significantly improving the model's ability to model long-range dependencies. Through the global attention module and the key vector compression technology, the long-range dependency modeling efficiency is significantly improved.

[0072] (2) The cross-layer semantic fusion is more flexible: Existing methods usually use fixed weights or simple splicing for multi-level semantic fusion, making it difficult to adaptively adjust the contributions of different-level features. The present invention introduces a dynamic gated fusion mechanism, which dynamically weights the global semantic encoding, the semantic states of the current layer and the previous layer through a learnable Sigmoid gating parameter matrix, realizing fine-grained fusion of cross-layer features. For example, when long-range dependencies are detected, the gating weight automatically biases towards the global encoding; when local features are more critical, the weights of the current layer or the previous layer are enhanced, improving the flexibility of semantic expression.

[0073] (3) The dynamic attention weight optimization focuses on key information: Existing methods rely on static attention weight allocation and are difficult to dynamically adjust the importance of tokens. The present invention adjusts the sharpness of the similarity distribution through a learnable temperature coefficient and combines it with Softmax to generate normalized dynamic weights, enabling the model to adaptively focus on important tokens. For example, the temperature coefficient is set to 0.5 at the beginning of training, and through backpropagation optimization, the weight distribution in the key semantic region becomes sharper, while the secondary region becomes smoother, effectively improving the pertinence of semantic enhancement. Detailed implementation manners

[0074] Embodiment, a method for extracting word embedding vectors based on a large language model, comprising the following steps:

[0075] S1. Embed a global context attention module in each self-attention network layer of the GPT model, generate query, key, and value vectors through linear transformation, and calculate global attention weights through the intra-layer cross-position attention mechanism.

[0076] S2. Based on the generated global semantic encoding, perform feature dimension splicing on the dynamic semantic encoding states of adjacent self-attention network layers, and fuse them through a gating mechanism to generate a global context dynamic semantic encoding state.

[0077] S3. Based on the global context dynamic semantic encoding state, use 4 - 8 independent attention heads to perform multi-head attention joint encoding on the current layer dynamic semantic encoding state and the global dynamic semantic encoding state.

[0078] S4. Perform average pooling on the dynamic semantic encoding states of each token in the jointly encoded sequence to generate a global semantic encoding vector, and calculate dynamic attention weights.

[0079] S5. Weightedly fuse the dynamic weights with the global semantic encoding vector to generate an enhanced global semantic encoding.

[0080] S6. Perform a non-linear transformation on the enhanced global semantic encoding through a fully connected layer containing Gaussian error linear unit activation to generate and extract the final word embedding vector.

[0081] The specific steps of S1 are as follows:

[0082] S1-1. Perform a linear transformation on the current layer dynamic semantic encoding state to generate a query vector, a key vector, and a value vector, where the dimension of the key vector is 1 / 8 of the dimension of the dynamic semantic encoding layer;

[0083] The current layer dynamic semantic encoding state is n is the sequence length, d model is the model dimension, and perform linear projection through three independent trainable weight matrices:

[0084] Q (l) =E (l) W Q ,K (l) =E (l) W K ,V (l) =E (l) W V

[0085] Among them:

[0086]

[0087] Key vector dimension compression: And

[0088] The output dimension satisfies The remaining vectors maintain their original dimensions

[0089] S1-2. Calculate the similarity matrix between the query vector and the key vector through the scaled dot-product attention mechanism, and normalize it using the Softmax function;

[0090] S1-2-1. Similarity matrix calculation: Calculate the scaled dot-product of the query vector and the compressed key vector:

[0091]

[0092] where the scaling factor is used to control the magnitude of the dot product and prevent the vanishing gradient.

[0093] S1-2-2. Probability distribution generation: Apply the Softmax function row-wise to the similarity matrix to generate the attention weight matrix:

[0094] A (l) = Softmax(S (l) ) satisfies

[0095] This operation maps the similarity to a probability distribution, representing the global dependency strength between tokens.

[0096] S1-3. Weightedly sum the normalized attention weights and the value vectors to generate the global semantic encoding, and connect it with the original dynamic semantic encoding state residually and input it into the next layer. The residual connection is used to retain the feature representation ability of the original dynamic semantic encoding state.

[0097] S1-3-1. Global semantic encoding generation: Weightedly sum the attention weights and the value vectors to generate the global context representation:

[0098]

[0099] where the encoding of each token aggregates the semantics of the entire sequence.

[0100] S1-3-2. Residual connection and output transfer: Residually add the global context representation to the original input to retain the underlying semantic information:

[0101]

[0102] Output E (l+1) as the input state of the next layer.

[0103] The said S2 specifically includes the following steps:

[0104] S2-1. Concatenate the global semantic encoding generated by S1, the dynamic semantic encoding state of the current layer, and the dynamic semantic encoding state of the previous layer along the feature dimension to form a fused state with doubled dimensions;

[0105] Three-way feature concatenation:

[0106]

[0107] Symbol definition:

[0108] The global semantic encoding generated by the l-th layer, from step S1;

[0109] The dynamic semantic encodings of the current layer and the previous layer;

[0110] By concatenating the global encoding (C (l) ), the current layer (E (l) ), and the previous layer (E (l-1) ) encodings, a fused state containing cross-layer global context and local dynamic features is formed;

[0111] Dimension expansion: If d c = d model , then the dimension after concatenation is 3d model , achieving multi-level semantic fusion.

[0112] S2-2. Perform a non-linear mapping on the fused state through a trainable gating parameter matrix to generate dynamic fusion weights, where the generation of the gating matrix includes cross-layer attention queries of the global semantic encoding;

[0113] Gating weight generation:

[0114]

[0115] Parameter description:

[0116] The gating parameter matrix of the fused features;

[0117] The gating parameter matrix of the global semantic encoding;

[0118] σ: Sigmoid function, output range [0,1]

[0119] Technical effect:

[0120] Each row of the dynamic weight G (l) contains three elements (g c , g l , g l-1 ), corresponding to the fusion weights of the global encoding, the current layer, and the previous layer respectively;

[0121] Introduce C (l) W c terms to make the gating decision guided by the global semantics. For example, when the global encoding detects long-range dependencies, increase the g c weight; when the local features are more critical, boost g l or g l-1 .

[0122] S2-3. Perform a weighted sum of the dynamic semantic encoding states of the previous layer and the current layer according to the dynamic fusion weights to generate the global dynamic semantic encoding state;

[0123] H (l) = LayerNorm(g c ⊙ C (l) + g l ⊙ E (l) + g l-1 ⊙ E (l-1) )

[0124] Definition:

[0125] ⊙: Element-wise multiplication (broadcasting mechanism), and the weight vectors [g c , g l , g l-1 are extended to the matrix dimension;

[0126] LayerNorm: Layer normalization, and the calculation formula is:

[0127] where μ, σ are the mean and standard deviation, and γ, β are learnable parameters;

[0128] Dynamic weight example:

[0129] For the i-th token, its fusion weight satisfies:

[0130]

[0131] If it means that 60% of the semantics of this token depends on the global context, 30% comes from the local features of the current layer, and 10% inherits the state of the previous layer.

[0132] The said S3 specifically includes the following steps:

[0133] S3-1. Use 4 - 8 independent attention heads to process the current layer and the global dynamic semantic encoding state respectively;

[0134] Independent attention head processing: For the current layer dynamic semantic encoding state and the global dynamic semantic encoding state Each independent attention head \(i\in\{1,\ldots,h\}\) (\(h\in[4,8]\)) performs the following operations:

[0135] - Generate query, key, and value vectors through linear projection:

[0136]

[0137] Dimension constraint: \(d\) k \(=d\) model / 8 (inheriting the key compression logic in claim 2), \(d\) v \(=d\) model / \(h\) (equal division of the output dimension);

[0138] Technical significance: Compressing the key vector (\(d\) k \(=d\) model / 8) reduces the computational complexity, and the dimension sharding of the value vector adapts to multi - head parallelism. Cross - attention calculation

[0139]

[0140] Core logic: Generate a query (\(Q\) i ) from the current layer state, generate keys and values (\(K\) i , \(V\) i ) from the global state, and capture cross - layer global dependencies.

[0141] S3 - 2. The output dimension of each attention head is an equal division of the dimension of the dynamic semantic encoding layer, and the calculation results of each head are concatenated along the feature dimension;

[0142] Multi - head output concatenation: Concatenate the outputs of \(h\) attention heads along the feature dimension and reduce the dimension through a projection matrix:

[0143]

[0144] Dimension consistency: \(h\cdot d\) v \(=h\cdot(d\) model / \(h\)) \(=d\) model , ensuring alignment with the input dimension after dimension reduction;

[0145] Technical effect: Integrate multi - head heterogeneous subspace features (such as syntax and semantics) and enhance the joint encoding ability.

[0146] S3 - 3. Reduce the dimension of the concatenated result through a trainable projection matrix, and perform residual connection and layer normalization with the current layer's dynamic semantic encoding state.

[0147] Residual connection and layer normalization: Perform residual connection between the multi - head output and the current layer state and perform layer normalization:

[0148]

[0149] Residual connection: Preserve the original dynamic semantic information and alleviate the vanishing gradient (inherit the residual design in claim 2); - Normalization operation: where μ and σ are the mean and standard deviation, and γ and β are learnable parameters.

[0150] Step S4 specifically includes the following steps:

[0151] S4-1. Perform average pooling on the dynamic semantic encoding states of all tokens in the sequence to generate a global semantic encoding vector;

[0152] Average pooling generates a global semantic encoding vector

[0153] Formula definition:

[0154] Parameter description:

[0155] H joint : The matrix of dynamic semantic encoding states after joint encoding, from the output of S3-3 in claim 4;

[0156] n: The length of the input sequence, the number of tokens;

[0157] d: The model dimension, consistent with d in claim 2 model ;

[0158] Logical association:

[0159] Input H joint Inherited from the LayerNorm output of S3-3 in claim 4: H joint = LayerNorm(H current + Proj(Concat(head1,..., head h )));

[0160] The average pooling operation compresses along the sequence dimension (n), preserves the feature dimension (d), and generates a global semantic vector G.

[0161] S4-2. Introduce a learnable temperature coefficient to adjust the similarity distribution between the current token's dynamic semantic encoding state and the global context vector;

[0162] Temperature coefficient adjusts the similarity distribution:

[0163] Formula definition: (Learnable parameter)

[0164] Parameter description:

[0165] W q , W k : Query / key projection matrix, inherit W in claim 2 Q , WK Compression logic, d k = d / 8;

[0166] τ: Temperature coefficient, with an initial value of 0.5, optimized through backpropagation;

[0167] s i : Similarity score of the i-th token with the global vector;

[0168] Logical association:

[0169] Projection consistency: Adopt the key vector compression method of S1-1 in claim 2:

[0170] Similarity calculation: Inherit the scaled dot product mechanism of S1-2 in claim 2, and add a temperature coefficient τ to control the distribution sharpness:

[0171] S4-3. Generate normalized dynamic attention weights through the Softmax function, where the initial value of the temperature coefficient is 0.5 and is automatically optimized during training.

[0172] Softmax generates dynamic attention weights:

[0173] Formula definition:

[0174] Parameter description:

[0175] α i : Dynamic attention weight of the i-th token;

[0176] The temperature coefficient τ participates in backpropagation, and the optimization goal is to make the weight distribution of important tokens sharper (τ tending to be small) or smoother (τ tending to be large);

[0177] Logical association:

[0178] Weight generation mechanism: Complementary to the gated weight generation of S2-2 in claim 3:

[0179] In claim 3: Static gating based on cross-layer feature concatenation (Sigmoid(W f [C fused +W g G));

[0180] This step: Dynamic attention based on global similarity (α i ), and the two jointly guide feature fusion;

[0181] Normalization constraint: Inherit the row-wise Softmax normalization of S1-2 in claim 2 to ensure ∑αi = 1.

[0182] S5 specifically includes the following steps:

[0183] S5-1. Multiply the dynamic attention weight and the global semantic encoding vector element-wise to generate a weighted context representation;

[0184] c i = α i ⊙ g avg

[0185] Symbol description:

[0186] α i The dynamic attention weight of token i, from S4-3 in claim 5; g avg The global semantic encoding vector, from S4-1 in claim 5;

[0187] Technical effect:

[0188] Adjust the contribution intensity of the global vector to the current token through dynamic weights to achieve fine-grained semantic enhancement.

[0189] S5-2. Concatenate the weighted context representation and the current token's dynamic semantic encoding state along the feature dimension to form a fused input vector.

[0190] Where

[0191] Proof of formula relevance:

[0192] Dimension alignment:

[0193] Is the dynamic semantic encoding of the current layer, inherited from S1-3 in claim 2;

[0194] c i The dimension is made consistent with through the broadcast mechanism, and d = 768 is a typical value; it conforms to the three-way concatenation logic of S2-1 in claim 3, i.e., feature dimension superposition;

[0195] Dynamic weight inheritance:

[0196]

[0197] Inherit the similarity calculation in S4-2 of claim 5

[0198] Gating mechanism extension:

[0199] Fused vector and τ in S2-3 of claim 3 -1 Form a two-path fusion.

[0200] S6 specifically includes the following steps:

[0201] S6-1. Input the fused input vector into the fully connected layer, and use the Gaussian Error Linear Unit (GeLU) as the activation function;

[0202] Nonlinear transformation of the fully connected layer:

[0203] GeLU(W fc ·E g +b fc )

[0204] (from S5-2 of claim 6)

[0205]

[0206] Formula explanation:

[0207] E g is the fused vector after splicing in claim 6, with a dimension of 2d( and splicing);

[0208] The fully connected layer compresses the 2d-dimensional input to d dimensions through the weight matrix W fc while retaining the same dimension as the original encoding; The expression of the GeLU activation function is: GeLU(x) = xΦ(x), where Φ(x) is the cumulative distribution function of the standard normal distribution, which can better retain negative value information compared to ReLU;

[0209] S6-2. Perform a residual connection between the output of the fully connected layer and the original token dynamic semantic encoding state to generate the final word embedding vector.

[0210] Residual connection and output:

[0211]

[0212] Formula explanation:

[0213] is the original dynamic semantic encoding of S1-1 in claim 2 (the i-th token of H l );

[0214] The projection matrix W proj ensures the alignment of the output of the fully connected layer with the original encoding dimension;

[0215] After the residual connection, layer normalization (LayerNorm) is performed, and the calculation formula is:

[0216]

[0217] where μ and σ are the mean and standard deviation, and γ and β are learnable parameters.

[0218] The foregoing has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and the above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.

Claims

1. A word embedding vector extraction method based on a large language model, characterized in that: The following steps are involved: S1. A global context attention module is embedded in each self-attention network layer of the GPT model, query, key, and value vectors are generated through linear transformation, and the global attention weight is calculated through the cross-position attention mechanism within the layer. S2. Based on the generated global semantic coding, the dynamic semantic coding states of adjacent self-attention network layers are spliced ​​in feature dimensions, and the global context dynamic semantic coding states are generated through fusion through a gating mechanism. S3. Based on the global context dynamic semantic coding state, use 4-8 independent attention heads to perform multi-head attention joint encoding on the current layer dynamic semantic coding state and the global dynamic semantic coding state. S4. Perform average pooling on the dynamic semantic coding states of each word in the jointly encoded sequence to generate a global semantic coding vector, and calculate the dynamic attention weight. S5. Weighted fusion of the dynamic weight and the global semantic coding vector to generate an enhanced global semantic coding. S6. The enhanced global semantic encoding is nonlinearly transformed through a fully connected layer containing Gaussian error linear unit activation to generate the final word embedding vector and extract it.

2. According to the method for extracting word embedding vectors based on a large language model according to claim 1, it is characterized in that: The S1 specifically includes the following steps: S1-1. Perform a linear transformation on the dynamic semantic coding state of the current layer to generate a query vector, a key vector and a value vector, where the dimension of the key vector is 1 / 8 of the dimension of the dynamic semantic coding layer; Specifically, the dynamic semantic encoding state of the current layer is Where n is the sequence length and d is the model dimension. Through three independent trainable weight matrices W Q , W K , W V Perform linear projection to generate query vector Q, key vector K and value vector V: Q = XW Q ,K=XW K ,V=XW V The dimension of the key vector is compressed to The query vector and value vector retain the original dimension d. S1-2. Calculate the similarity matrix between the query vector and the key vector through the scaled dot product attention mechanism and normalize it using the Softmax function; The calculation formula of the similarity matrix is: The scaling factor is It is used to control the amplitude of the dot product and prevent the gradient from disappearing. Next, the Softmax function is applied row by row to the similarity matrix to generate the attention weight matrix: A = Softmax(A) This operation maps the similarity into a probability distribution, which represents the global dependency strength between word units. S1-3. The normalized attention weights are weighted and summed with the value vector to generate a global semantic code, which is then residually connected to the original dynamic semantic code state and input into the next layer. The residual connection is used to retain the feature representation capability of the original dynamic semantic code state. The generation formula of global semantic coding is: The encoding of each word is Aggregates the semantic information of the entire sequence. Then, the global context representation is residually superimposed with the original input to retain the underlying semantic information: The output Y is used as the input state of the next layer.

3. The word embedding vector extraction method based on a large language model according to claim 1, characterized in that: The S2 specifically includes the following steps: S2-1. The global semantic coding generated by S1, the dynamic semantic coding state of the current layer and the dynamic semantic coding state of the predecessor layer are spliced ​​along the feature dimension to form a fusion state with doubled dimension; The formula for three-way feature splicing is: Among them, Z is the global semantic code generated by the lth layer, X l and X l-1 They are the dynamic semantic encodings of the current layer and the predecessor layer respectively. By concatenating the global encoding, the current layer and the predecessor layer encoding, a fusion state containing cross-layer global context and local dynamic features is formed. S2-2. Generate dynamic fusion weights by nonlinearly mapping the fusion state through a trainable gating parameter matrix, where the production of the gating matrix contains cross-layer attention queries of global semantic encoding; The formula for generating the gate weight is: G = σ (FW F +ZW Z )W F and W Z are the gating parameter matrices of fusion features and global semantic coding, σ is the Sigmoid function, and the output range is [0,1]. Each row of the dynamic weight G contains 3 elements, corresponding to the fusion weights of the global coding, the current layer, and the predecessor layer. S2-3. Perform weighted summation of the dynamic semantic coding state of the predecessor layer and the current layer according to the dynamic fusion weight to generate a global dynamic semantic coding state; The formula for weighted summation is: Y = LayerNorm (g1⊙Z + g2⊙X l +g3⊙X l-1 ) Among them, g1, g2, g3 are the dynamic fusion weights of the global encoding, current layer, and predecessor layer respectively, and LayerNorm is the layer normalization operation.

4. The word embedding vector extraction method based on a large language model according to claim 1, characterized in that: The S3 specifically includes the following steps: S3-1. Use 4-8 independent attention heads to process the current layer and global dynamic semantic encoding state respectively; Each independent attention head i performs the following operations: - Linear projection generates query, key, value vectors: Cross-attention calculation: Among them, d h =d / h, where h is the number of attention heads S3-2. The output dimension of each attention head is an equal division of the dimension of the dynamic semantic encoding layer, and the calculation results of each head are concatenated along the feature dimension; The formula for multi-head output splicing is: Among them, W O is a trainable projection matrix that ensures alignment with the input dimension after dimensionality reduction. S3-3. Reduce the dimension of the splicing result through a trainable projection matrix, and perform residual connection and layer normalization with the dynamic semantic encoding state of the current layer. The formula for residual connection and layer normalization is: Y = LayerNorm (O + X l ) 5. The word embedding vector extraction method based on a large language model according to claim 1, characterized in that: The S4 specifically comprises the following steps: S4-1. Average pooling of dynamic semantic encoding states of all word units in the sequence to generate a global semantic encoding vector; The formula for average pooling is: S4-2. Introduce a learnable temperature coefficient to adjust the similarity distribution between the current word unit dynamic semantic encoding state and the global context vector; The formula for calculating similarity is: Among them, τ is the temperature coefficient, the initial value is 0.5, and it is optimized by back propagation. S4-3. Generate normalized dynamic attention weights through the Softmax function, where the initial value of the temperature coefficient is 0.5 and is automatically optimized during the training process. The formula for generating dynamic attention weights is:

6. The word embedding vector extraction method based on a large language model according to claim 1, characterized in that: The S5 specifically includes the following steps: S5-1. Multiply the dynamic attention weights by the global semantic encoding vector element-wise to generate a weighted context representation; The formula for weighted context representation is: c i =a i ⊙G; S5-2. Concatenate the weighted context representation and the current word-unit dynamic semantic encoding state along the feature dimension to form a fused input vector. The formula for fusing the input vector is:

7. The word embedding vector extraction method based on a large language model according to claim 1, characterized in that: The S6 specifically comprises the following steps: S6-1. Input the fused input vector into the fully connected layer and use Gaussian error linear unit (GeLU) as the activation function; The nonlinear transformation formula of the fully connected layer is: S6-2. Perform a residual connection between the output of the fully connected layer and the dynamic semantic encoding state of the original word unit to generate the final word embedding vector. The formula for residual connection and output is: e i =LayerNorm(h i +y i ) 8. A word embedding vector extraction system based on a large language model, designed according to the word embedding vector extraction system based on a large language model according to claims 1-7, characterized in that: include: Output module, used for text segmentation and generating initial word embedding vectors; The global context processing module includes a cascade of enhanced attention encoding network layers, each of which is equipped with a cross-layer attention unit and a gated fusion unit; a dynamic weight calculation module is used to calculate the word-level dynamic attention weight; The word embedding generation module contains a fully connected network activated by Gaussian error linear unit (GeLU), which is used to fuse contextual information with dynamic semantic encoding state; the output module outputs the final word embedding vector as a dense matrix in sequence order.

9. A word embedding vector extraction system based on a large language model according to claim 8, characterized in that: In the enhanced self-attention module: The key vector dimension of the cross-layer attention unit is set to 1 / 8 of the model, and the key vector dimension is equal to the value vector dimension; the gated fusion unit generates dynamic weights through the Sigmoid function and is combined with the residual connection to retain the original semantic information.

Citation Information

Cited By

  • Text classification method, electronic equipment and storage medium

    CN120744127A

  • Semantic information generation method and device, storage medium and electronic device

    CN120975232A

  • Text enhancement time sequence prediction method and system based on large model word embedding space

    CN122087743A

  • A Text Augmentation Temporal Prediction Method and System Based on Large Model Word Embedding Space

    CN122087743B

  • Model processing method and device for semantic understanding and medium

    CN122088515A