Scene learning method about scene text replacement invariance
By improving the attention mechanism and positional encoding of the Transformer structure, the sensitivity of large language models to sample order in contextual learning is solved, achieving output consistency and performance improvement while reducing computational costs.
Patent Information
- Application Number
- CN202410457823.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-16
- Publication Date
- 2025-10-28
Smart Images

Figure CN120849607A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of machine learning technology in artificial intelligence, and relates to contextual learning technology for large models, specifically a contextual learning method that is invariant to contextual text substitution. Background Technology
[0002] Contextual learning, initially proposed by GPT-3, is an inference method that utilizes a pre-trained large model for few-shot learning. Previously, applying a language model to downstream text processing tasks typically required fine-tuning the language model using downstream text data. However, as the number of parameters in the language model increases, the computational cost of fine-tuning becomes very high. Therefore, contextual learning directly converts some sample-label pairs from the downstream task into natural language and concatenates them to form contextual text. This contextual text is then used as a prefix, combined with the test question statement, and input into the language model to obtain a natural language output. Compared to directly inputting the question into the language model, the output generated using this process is generally more accurate due to the use of task-related training data.
[0003] In contextual learning, the contextual text X is typically composed of n samples and test text, i.e., X = (x1, ..., x2). n x t ). where each x i It consists of text-label pairs from the dataset, while the test text x t This only contains text, and its labels are generated by the model. For a given set of samples {x1, ..., x...} n There are t! different permutations to compose a contextual text sequence. Intuitively, if the same data is input into a large model in different orders, the model's output should be the same or similar, since the knowledge acquired through contextual learning is essentially consistent. However, existing research has found that for the most commonly used large language models—autoregressive large language models (such as the GPT series)—the output of contextual learning is highly sensitive to the order of samples in the contextual text. Even using the same data, different orders of contextual samples result in very large variance in the model's output. For example, on the SST-2 sentiment classification dataset, the inference accuracy of contextual learning using different orders of contextual text can fluctuate between 51.6% and 88.7%. Furthermore, the sensitivity of contextual learning to the order of text samples leads to a decrease in its performance.
[0004] To address this issue, Sebastian Riedel et al. proposed a filtering module to select the best-performing order from a given set of samples. While this filtering method improves the performance of context learning, its filtering process relies on additional generation operations on the contextual text using a large model, which significantly increases the computational cost of context learning.
[0005] Since the samples in the contextual text are essentially independent and identically distributed, designing a contextual learning method that is invariant to sample permutations in the contextual text is a research approach to address the issue of sensitivity to the order of contextual learning. Existing research suggests that inductive bias that achieves sample symmetry in the model is beneficial for both learning and generalization. From a model structure perspective, the sensitivity of autoregressive large language models to sample order mainly stems from their autoregressive nature. This is achieved by using asymmetric causal masks and positional encodings in the model to determine the next word p(x) based on the existing sequence. i |x1, ...,x i-1 This prevents the first word from acquiring information about subsequent words, thus violating the permutation invariance of the original Transformer model. However, no work has yet addressed the order sensitivity issue in context-based learning, making it difficult to guarantee the accuracy of inference. Summary of the Invention
[0006] To overcome the shortcomings of the existing technologies, this invention proposes a context learning method that is permutation-invariant to contextual text. This method mainly improves the attention mechanism and positional encoding in the Transformer structure of large language models, thereby achieving permutation invariance in context learning from the model structure. This solves the problem that context learning in large language models is sensitive to the order of samples in contextual text, and improves the performance of context learning in large language models.
[0007] The technical solution of this invention:
[0008] A context-based learning method that is invariant to contextual text substitution. Specifically, for the Transformer language model, it includes the following steps:
[0009] A. Copy the context text and use it as input to the language model (x1, ..., x). i , ..., x n ,x′1,...,x′ i ,...,x′ n x t ). Where x i =x′ i x i This represents each contextual text sample; x tThis represents the test text; n is the number of scenario text samples; t is an abbreviation for the test text "test".
[0010] B. The language model's embedding layer is used to map the input to the feature space, obtaining the embedding representation of the input sequence in the feature space (z1, ..., z). i , ..., z n ,z′1,...,z′ i , ..., z′ n , z t ). Among them, z i For the embedded representation of contextual text samples; z t This is an embedded representation of the test text.
[0011] C. Change the original location encoding allocation method of the language model. Specifically, assign each contextual text sample x i All are considered to start from the first position and apply the same positional encoding. Furthermore, for the test text x... t Assign appropriate positional codes so that their corresponding positions are after all contextual text samples. Finally, combine the assigned positional codes with the features obtained in step B (i.e., the embedding representations (z1, ..., z2) according to the positional coding type of the language model. i , ..., z n ,z′1,...,z′ i , ..., z′ n , z t To integrate.
[0012] D. Text features are calculated using a permutation-invariant attention mechanism, including:
[0013] D1. Calculate (z1, ..., z) n Independent encoding of z i ←f(z i Where f is the autoregressive self-attention encoding, and A←B means assigning value A to B;
[0014] D2. Using (z1, ..., z) i-1 , z i+1 , ..., z n ) and z′ i Calculate z′ i Independent encoding z′ i ←f(z1, ..., z) i-1 , z i+1 , ..., z n , z′ i );
[0015] D3. Using (z′1,...,z′) n ) and zt z was calculated t test code z t ←f(z′1,...,z′ n , z t );
[0016] E. Use the language model's generator head module to compute the test code z. t Converted into generated text.
[0017] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0018] The permutation-invariant context learning method provided by this invention can fundamentally eliminate the negative impact of context learning's sensitivity to sample order, eliminating the need to design sample order when large models perform context learning. Simultaneously, this invention preserves the expressive power of large language models, significantly improving the generalization performance of context learning across multiple tasks. Attached Figure Description
[0019] Figure 1 The flowchart illustrates the permutation-invariant scenario learning method provided by this invention.
[0020] Figure 2 This is a schematic diagram of Transformer causal encoding.
[0021] Figure 3 A schematic diagram of a Transformer causal mask with permutation invariance designed for this invention. Detailed Implementation
[0022] The present invention will be further described below with reference to the accompanying drawings and embodiments, but the scope of the invention is not limited in any way.
[0023] This invention proposes a context learning method that is permutation-invariant to contextual text. It mainly improves the attention mechanism and positional encoding in the Transformer structure of large language models to achieve permutation invariance in context learning from the model structure, solves the problem that context learning of large language models is sensitive to the sample order in contextual text, and improves the performance of context learning of large language models.
[0024] The specific implementation of the present invention is as follows. The flow of the permutation-invariant scenario learning method provided by the present invention is as follows: Figure 1 As shown, we will use the commonly used GPT model structure and the context-based sentiment classification task as examples to further clarify and fully illustrate the present invention.
[0025] Suppose the classification task to be completed is to determine whether the sentiment of the text statement "I am angry" is positive or negative, and the dataset is provided as {("I am happy", positive), ("I am very frustrated", negative)}. Then the text dataset can be transformed into situational texts x1 = "Statement: I am happy. Sentiment: positive." and x2 = "Statement: I am very frustrated. Sentiment: negative.", and the task can be transformed into test text x. t = "Statement: I am angry". Assuming the language model's tokenizer converts each character into a token, then x1, x2, x... t They contain 14, 15, and 14 morphemes respectively.
[0026] A. Transform the dataset into contextual text (x1, ..., x) as described above. n ) and test text x t The context text is copied and then concatenated with the test text as input to the language model (x1, ..., x). n ,x′1,...,x′ n x t ), where x i =x′ i .
[0027] B. Map the language model's pre-trained embedding layer onto the feature space to obtain (z1, ..., z). n ,z′1,...,z′ n , z t ), where z i ∈r k , where k is a network hyperparameter.
[0028] C. Assign position codes to the word elements in x1 sequentially according to positions 0, 1, ..., 13; assign position codes to the word elements in x2 sequentially according to positions 0, 1, ..., 14; and assign position codes to x... t The lexical units are assigned positional codes sequentially at positions 15, 16, ..., 28. These positional codes are then fused with the features obtained in step B, according to the specific positional coding type used by the language model.
[0029] D. Features are computed using a permutation-invariant attention mechanism. Specifically, the causal mask in the original Transformer is as follows: Figure 2 As shown, the present invention adopts Figure 3 The causal mask shown is permutation-invariant. First, since the context text is copied once in step A, we also expand the causal mask accordingly. Second, for each iteration of the context text, we remove the cross-attention portion (z). i For z j Attention (i≠j), only self-attention (z) is retained.i For z i Attention). For two-pass context text encoding (z1, ..., z... n ), (z′1,...,z′ n Between z1, ..., z2, we retain (z1, ..., z2) i-1 , z i+1 , ..., z n ) and z′ i Cross-attention. Finally, we retain (z′1, ..., z′) n ) and z t Cross-attention. As shown in the figure, the designed causal mask allows the model to perform the following attention mechanism calculation process in parallel:
[0030] D1. Calculate (z1, ..., z) n Independent encoding of z i ←f(z i ), where f is the autoregressive self-attention encoding.
[0031] D2. Using (z1, ..., z) i-1 , z i+1 , ..., z n ) and z′ i Calculate z′ i The encoding z′ i ←f(z1, ..., z) i-1 , z i+1 , ..., z n , z′ i )
[0032] D3. Using (z′1,...,z′) n ) and z t Calculate z t The encoding z t ←f(z′1,...,z′ n , z t )
[0033] E. Use the language model's generator head module to compute the test code z. t Converted into generated text.
[0034] This enables a contextual learning method that maintains the integrity of contextual text substitution.
[0035] It should be noted that the purpose of disclosing the embodiments is to help further understand the present invention. However, those skilled in the art will understand that various substitutions and modifications are possible without departing from the scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the content disclosed in the embodiments, and the scope of protection of the present invention is defined by the scope of the claims.
Claims
1. A contextual learning method with invariant contextual text substitution, characterized in that, By improving the attention mechanism and positional encoding in the Transformer structure of large language models, permutation invariance for context-based learning of contextual text is achieved from the model structure; including the following steps: A. Convert the text dataset into contextual text (x1,…,x) n ) and test text x t The context text is copied and then concatenated with the test text as input to the language model, represented as (x1,…,x). i ,…,x n ,x′1,…,x′ i ,…,x′ n ,x t ); where x i =x′ i x i This represents each contextual text sample; x t This represents the test text; n is the number of scenario text samples. B. The language model's embedding layer maps the input to the feature space, obtaining the embedding representation of the input sequence in the feature space (z1,…,z). i ,…,z n ,z′1,…,z′ i ,…,z′ n ,z t ); where z i For the embedded representation of contextual text samples; z t This is an embedded representation of the test text; C. Change the original positional encoding allocation method of the language model; Specifically, each scenario text sample x i Starting with the first position, apply the same positional encoding; Then test text x t Assign a positional code so that its corresponding position is after all contextual text samples; Finally, the assigned positional codes are fused with the embedding representation features obtained in step B according to the positional code type of the language model; D. Using a causal mask with permutation invariance, text features are calculated based on a permutation-invariant attention mechanism, including: D1. Calculate (z1,…,z) n Independent encoding of z i ←f(z i ); where f is the autoregressive self-attention encoding, and A←B means assigning the value of A to B; D2. Using (z1,…,z) i-1 ,z i+1 ,…,z n ) and z′ i Calculate z′ i Independent encoding z′ i ←f(z1,…,z i-1 ,z i+1 ,…,z n ,z′ i ); D3. Using (z′1,…,z′) n ) and z t z was calculated t test code z t ←f(z′1,…,z′ n ,z t ); E. Using the language model's generation head module, the calculated test code z t Converted into generated text.
2. The contextual learning method with invariant contextual text substitution as described in claim 1, characterized in that, Step D, generating a causal mask with permutation invariance, specifically includes: First, the contextual text is copied in step A to expand the causal mask; Secondly, for each iteration of the context text, remove the cross-attention portion, i.e., z. i For z j For the attention of i ≠ j, only self-attention is retained, i.e., z. i For z i attention; For contextual text encoding (z1,…,z) n ),(z′1,…,z′ n Between z1, ..., z2, ..., z3, ..., z4, ... i-1 ,z i+1 ,…,z n ) and z′ i Cross attention; Finally, retain (z′1,…,z′) n ) and z t Cross attention.