Training Method, Device, Electronic Device and Readable Storage Medium for Transformer Model
By dividing the attention head of the Transformer model into self-attention head and cross-attention head, and introducing target domain feature vectors into the cross-attention head for feature fusion, the problems of difficulty in obtaining data, high computational cost and poor adaptability in domain migration of the pre-trained language model are solved, and efficient domain migration effect is achieved in low computing resource scenarios.
Patent Information
- Application Number
- CN202510280316.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-11
AI Technical Summary
The existing pre-trained language models for domain migration have challenges in problems such as difficulty in obtaining data, high computational costs and poor domain adaptability, especially in the low computing resource scenarios, where it is difficult to efficiently realize the domain migration task of large-scale pre-trained language models.
By dividing the attention heads in each layer of the original Transformer model into self-attention heads and cross-attention heads, introducing the target domain text feature vectors into the cross-attention heads for feature fusion, and training the original Transformer model using the target domain text to obtain the target Transformer model.
In the low computing resource scenario, a large-scale pre-trained language model is implemented with high-quality text migration tasks from the source field to the target field, which significantly reduces the demand for text training data in the target field, reduces the cost of computing resources, and improves the model's adaptability to text in different fields.
Smart Images

Figure CN119807714B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning, and particularly to a training method, device, electronic device and readable storage medium for a Transformer model. Background Art
[0002] In recent years, large-scale pre-trained language models based on the Transformer architecture (such as BERT, GPT, etc.) have made remarkable progress in the field of Natural Language Processing (NLP). These models can capture rich language features and knowledge by pre-training on large-scale general corpora. However, when these models need to be adapted to specific domains, such as the medical, legal, financial, etc. fields, traditional methods usually rely on a large amount of in-domain parallel corpora for fine-tuning, which brings the following problems:
[0003] Difficult data acquisition. Corpus data in many professional fields is scarce and difficult to obtain. For example, texts in the medical field may involve patient privacy, texts in the legal field may be protected by copyright, and texts in the financial field may contain sensitive business information; therefore, obtaining sufficient in-domain data for fine-tuning is a major challenge;
[0004] High computational cost. Fine-tuning large-scale pre-trained models requires a large amount of computational resources, especially when dealing with large-scale in-domain data;
[0005] Domain adaptability problem. Even with sufficient in-domain data, the model may overfit to the specific domain data during the fine-tuning process, resulting in a decline in performance on general tasks. This domain adaptability problem limits the generalization ability of the model between different tasks. Summary of the Invention
[0006] The present invention provides a training method, device, electronic device and readable storage medium for a Transformer model, to solve the problems of difficult data acquisition, high computational cost and poor domain adaptability of existing pre-trained language models for domain transfer, and to efficiently implement the domain transfer task of large-scale pre-trained language models in low-computational-resource scenarios.
[0007] The present invention provides a training method for a Transformer model, including: dividing the attention heads in each layer of the original Transformer model into self-attention heads and cross-attention heads; the original Transformer model is trained based on source domain texts; extracting features from target domain texts to obtain target domain text feature vectors; introducing the target domain text feature vectors into the cross-attention heads for feature fusion to obtain an intermediate Transformer model; and training the intermediate Transformer model using the target domain texts to obtain a target Transformer model.
[0008] Optionally, the training the intermediate Transformer model using the target domain texts to obtain a target Transformer model includes: freezing all other parameters except the LoRA parameters in the cross-attention heads of the intermediate Transformer model, and training the LoRA parameters using the target domain texts; the LoRA parameters are adjustment parameters introduced through low-rank decomposition; unfreezing the self-attention parameters of specified key layers of the intermediate Transformer model for fine-tuning to obtain the target Transformer model.
[0009] Optionally, the unfreezing the self-attention parameters of specified key layers of the intermediate Transformer model for fine-tuning to obtain the target Transformer model includes: unfreezing the self-attention parameters of the last 3 Transformer layers of the intermediate Transformer model for fine-tuning to obtain the target Transformer model.
[0010] Optionally, the extracting features from target domain texts to obtain target domain text feature vectors includes: vectorizing the target domain texts respectively using a vector model with the same structure as the original Transformer model to obtain multiple text vectors; calculating the arithmetic mean of the multiple text vectors to obtain a target domain average feature vector; and normalizing the target domain average feature vector to obtain target domain text feature vectors.
[0011] Optionally, the calculation formula for the cross-attention head is:
[0012]
[0013] where, is the cross-attention head, Q is the query matrix of the input sequence, K is the key matrix of the input sequence, V is the value matrix of the input sequence, is the average feature vector of the input sequence, is the dimension of the self-attention head, is the key matrix K 's transpose matrix, is the average feature vector after normalization calculation, is 's transpose matrix.
[0014] Optionally, the original Transformer model uses LLaMA-7B as the base model.
[0015] Optionally, dividing the attention heads in each layer of the original Transformer model into self-attention heads and cross-attention heads includes: dividing the 32 attention heads in each layer of the original Transformer model into 24 self-attention heads and 8 cross-attention heads.
[0016] Optionally, dividing the attention heads in each layer of the original Transformer model into self-attention heads and cross-attention heads includes: adjusting the quantity ratio of the self-attention heads to the cross-attention heads based on actual task requirements.
[0017] The present invention also provides a training device for a Transformer model, including the following modules:
[0018] A processing module, configured to divide the attention heads in each layer of the original Transformer model into self-attention heads and cross-attention heads; the original Transformer model is trained based on source domain texts;
[0019] A feature extraction module, configured to extract features from target domain texts to obtain target domain text feature vectors;
[0020] A feature fusion module, configured to introduce the target domain text feature vectors into the cross-attention heads for feature fusion to obtain an intermediate Transformer model;
[0021] A training module, configured to train the intermediate Transformer model using the target domain texts to obtain a target Transformer model.
[0022] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, it implements the training method of the Transformer model as described in any one of the above.
[0023] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the training method of the Transformer model as described in any one of the above is implemented.
[0024] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the training method of the Transformer model as described in any one of the above is implemented.
[0025] The training method, device, electronic device and readable storage medium of the Transformer model provided by the present invention divide the attention heads in each layer of the original Transformer model into self-attention heads and cross-attention heads, introduce the text feature vectors of the target domain in the cross-attention heads for feature fusion, and use the text of the target domain to train the original Transformer model to obtain the target Transformer model. It can achieve the text migration task of large-scale pre-trained language models from the source domain to the target domain with high quality in low-computing-resource scenarios, significantly reduce the demand for target-domain text training data, reduce the computing resource cost, and improve the adaptability of large-scale pre-trained language models to identify texts in different domains. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0027] Figure 1 It is a flowchart of a training method of a Transformer model provided by the present invention.
[0028] Figure 2 It is a schematic flow diagram of the sub-steps of step 102 provided by the present invention.
[0029] Figure 3 It is a schematic structural diagram of a training device of a Transformer model provided by the present invention.
[0030] Figure 4 It exemplifies a schematic diagram of the physical structure of an electronic device.
[0031] Reference Signs:
[0032] Training device 30 of the Transformer model; processing module 301; feature extraction module 302; feature fusion module 303; training module 304; processor 410; communication interface 420; memory 430; communication bus 440. Detailed implementation
[0033] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0034] The adaptation of large-scale pre-trained language models from the source domain to the target domain refers to adjusting or optimizing a language model pre-trained on a large amount of general data to improve its performance in a specific target domain, such as the medical, legal, financial and other fields. The adaptation methods of large-scale pre-trained language models from the source domain to the target domain mainly include the following categories:
[0035] Full fine-tuning method, directly using the text data in the target domain to perform complete fine-tuning on the pre-trained model. This method requires a large amount of text data and computing resources in the target domain and is prone to the problem of catastrophic forgetting; Prompt learning method, guiding the pre-trained model to adapt to the target domain by designing specific prompt templates. Although this method is parameter-efficient, the effect often depends on the quality of the manually designed prompts; Adapter method, inserting a small adapter module into the original pre-trained model. Although it reduces the number of parameters, it fails to fully utilize the feature extraction ability of the pre-trained model.
[0036] The above adaptation methods mainly have the following main problems:
[0037] Strong data dependence, requiring a large amount of data in the target text domain to achieve good adaptation effects; high consumption of computing resources, the full fine-tuning method needs to update all the parameters of the pre-trained language model, with high computing costs; insufficient feature fusion, existing methods often treat domain adaptation as an independent module and fail to deeply integrate with the core attention mechanism of the pre-trained language model; unstable transfer effect, in low-computing-resource scenarios, the domain transfer effects of existing methods are often not ideal enough.
[0038] These problems severely restrict the wide application of pre-trained language models in professional fields. Therefore, the present invention proposes a training method for the Transformer model that can efficiently achieve domain transfer, requires a small amount of data, has a stable transfer effect, and consumes low computing resources.
[0039] Among them, the Transformer model is a deep learning architecture based on the self-attention mechanism, which is widely used in natural language processing (NLP) tasks such as machine translation, text generation, and classification; the Transformer model completely relies on the self-attention mechanism (Self-Attention) to capture the relationships between different positions in the input sequence, including input embedding (InputEmbedding), positional encoding (Positional Encoding), encoder (Encoder), decoder (Decoder), and output layer (Output Layer), and also includes multi-head attention. By calculating multiple attention heads in parallel, the model can extract information from different subspaces and enhance the expressive ability.
[0040] Figure 1 The flowchart of a training method for the Transformer model provided by the present invention is as Figure 1 shown. This training method for the Transformer model is used in devices such as servers, desktops, and laptops, and includes the following steps.
[0041] In step 101, the attention heads in each layer of the original Transformer model are divided into self-attention heads and cross-attention heads.
[0042] Among them, the original Transformer model is trained based on source domain text, and the source domain text can be large-scale general corpus text; the corresponding is target domain text, which is text different from the initial domain. For example, if the source domain text is general corpus text, then the target domain text can be text in fields such as medicine, law, and finance.
[0043] One of the inventive points of the present invention is to improve the attention mechanism structure design of the original Transformer model, that is, in each layer of the original Transformer model, a total of H attention heads are divided into H1 self-attention heads and H2 cross-attention heads, where H = H1 + H2. The self-attention heads maintain the original in-sequence attention calculation mechanism and use the standard attention calculation formula. The calculation formula for the self-attention heads is:
[0044]
[0045] Among them, is the self-attention head,Q is the query matrix of the input sequence, K is the key matrix of the input sequence, V is the value matrix of the input sequence, is the dimension of the self-attention head, is the key matrix K 's transpose matrix, is the activation function; the self-attention head is responsible for maintaining the basic understanding ability of the Transformer model for the input sequence.
[0046] The calculation formula of the cross-attention head is:
[0047]
[0048] where is the cross-attention head, Q is the query matrix of the input sequence, K is the key matrix of the input sequence, V is the value matrix of the input sequence, is the average feature vector of the input sequence, is the dimension of the self-attention head, is the key matrix K 's transpose matrix, is the average feature vector after normalization calculation, is 's transpose matrix.
[0049] In one implementation, the original Transformer model uses LLaMA-7B as the base model; LLaMA-7B uses 32 Transformer layers, each layer contains 32 attention heads, the hidden dimension of LLaMA-7B is 4096, uses Rotary Position Embedding (RoPE) as the position encoding scheme of the text, and uses the Root Mean Square Layer Normalization (RMSNorm) method for normalization.
[0050] Among them, LLaMA-7B is an open-source large language model (LLM) and belongs to one of the LLaMA (Large Language Model Meta AI) series of models. The goal of the LLaMA series of models is to achieve or exceed the performance of existing large models through more efficient training methods and smaller model sizes. LLaMA-7B is a version with 7 billion parameters. Compared with larger models (such as LLaMA-13B, LLaMA-30B, and LLaMA-65B), LLaMA-7B has a lighter computational resource requirement while still maintaining strong language understanding and generation capabilities.
[0051] In the present invention, the original Transformer model using LLaMA-7B as the base model is taken as an example for illustration.
[0052] Exemplarily, the 32 attention heads in each layer of LLaMA-7B can be divided into 24 self-attention heads and 8 cross-attention heads.
[0053] In one implementation, the self-attention heads and cross-attention heads can be divided based on any one of the following two methods:
[0054] Task-driven allocation: According to specific task requirements, some heads are designed as self-attention heads and some as cross-attention heads; for example, in machine translation, the cross-attention heads are specifically used to handle the relationship between the source domain text and the target domain text.
[0055] Dynamic allocation: In more complex language models, the functions of the attention heads can be dynamically adjusted to automatically determine which heads are used for self-attention and which are used for cross-attention based on the input data.
[0056] Exemplarily, the ratio of the number of self-attention heads to cross-attention heads can be dynamically adjusted based on actual task requirements, and the present invention does not limit this.
[0057] By dividing the H attention heads into H1 self-attention heads and H2 cross-attention heads, the Transformer model can simultaneously capture the semantic relationships within the sequence and the interaction information between sequences; this design is particularly effective in domain transfer tasks because the model can better understand the relationship between the source domain and the target domain, thereby achieving a more natural domain transfer.
[0058] It should be noted that RoPE is a positional encoding method that encodes the positional information of tokens in the form of a rotation matrix, enabling LLaMA-7B to better understand the relative positional relationships of tokens in the input sequence. A token refers to a word or a part of a word in the input sequence text. The calculation formula of RoPE is as follows:
[0059]
[0060] where θ is the rotation angle related to the position, q is the query matrix of the input sequence text, k is the key matrix of the input sequence text, is the query matrix after calculation by the rotation matrix, is the key matrix after calculation by the rotation matrix.
[0061] RoPE can better handle sequence lengths beyond those processed during training, has better computational efficiency, and provides a flexible and efficient way to encode positional information, especially suitable for scenarios that require processing long sequences or have strict restrictions on computing resources.
[0062] In step 102, feature extraction is performed on the target domain text to obtain the target domain text feature vector.
[0063] The target domain text is the text corresponding to the source domain text, and the target domain text is text different from the initial domain. For example, if the source domain text is general corpus text, then the target domain text can be text in fields such as medicine, law, and finance.
[0064] It should be noted that step 102 also includes sub-step 1021, sub-step 1022, and sub-step 1023. The specific method for determining the target domain text feature vector will be described in detail in the sub-steps of step 102. Please refer to Figure 2 , Figure 2 which is the flow schematic diagram of the sub-steps of step 102 provided by the present invention.
[0065] In sub-step 1021, a vector model consistent with the original Transformer model structure is used to vectorize the target domain text respectively to obtain multiple text vectors.
[0066] Exemplarily, the vector model consistent with the original Transformer model structure can be the vector model built into LLaMA-7B. The vector model built into LLaMA-7B is used as the vectorization tool to vectorize the target domain text to ensure consistency during training.
[0067] For each target domain text in the target domain corpus , calculate its vector representation:
[0068]
[0069] Among them, E(x) is the vectorization function of the vector model, is the vector representation of the text in the target domain; to ensure the quality of the vector representation of the text in the target domain, the following preprocessing is performed before vectorization:
[0070] Perform text segmentation on the text in the target domain: Split the long text into natural paragraphs. For example, the boundaries of paragraphs can be identified based on punctuation marks; or blank lines can be used as the basis for paragraph separation.
[0071] Remove special formats from the text in the target domain: Clean up HyperText Markup Language (HTML) tags, special symbols, etc. For example, regular expressions or specialized HTML parsing libraries can be used to remove HTML tags from the text.
[0072] After vectorizing the text in the target domain separately through the vector model built into LLaMA-7B, multiple text vectors are obtained.
[0073] Sub-step 1022, calculate the arithmetic mean of multiple text vectors to obtain the average feature vector of the target domain.
[0074] In this step, the calculation formula for the arithmetic mean of multiple text vectors is:
[0075]
[0076] Among them, is the average feature vector of the target domain, is the vector representation of the text in the target domain, N is the number of.
[0077] Calculating the arithmetic mean of multiple text vectors in the target domain can effectively capture the global features of the target domain; the averaging operation can smooth out the noise or specificity in a single text and retain the common features of the target domain; the vector after arithmetic averaging can more stably represent the overall semantic distribution of the target domain and reduce the impact of deviations or outliers in a single text on the domain features.
[0078] Sub-step 1023, normalize the average feature vector of the target domain to obtain the text feature vector of the target domain.
[0079] To maintain the accuracy of numerical values during the calculation process and avoid calculation result deviations or failures caused by extreme values, L2 normalization is used to normalize the obtained average feature vector of the target domain, and its calculation formula is as follows:
[0080]
[0081] Among them, is the average feature vector of the target domain after normalization calculation, that is, the text feature vector of the target domain; is the average feature vector of the target domain.
[0082] The normalized text feature vector of the target domain can provide a clear feature representation of the target domain for the model; this representation can be used as a reference for the domain transfer task to help the model better adapt to the semantics and distribution of the target domain. In the domain transfer task, the distribution difference between the source domain and the target domain is a key challenge. By introducing the text feature vector of the target domain, the model can more effectively align the feature spaces of the source domain and the target domain, thereby reducing the negative impact brought by the domain difference.
[0083] In step 103, the text feature vector of the target domain is introduced into the cross-attention head for feature fusion to obtain an intermediate Transformer model.
[0084] Exemplarily, the 32 attention heads in each layer of LLaMA-7B can be divided into 24 self-attention heads and 8 cross-attention heads, and the following uses this example for illustration.
[0085] The self-attention head maintains the original rotary position encoding and attention calculation mechanism of LLaMA-7B. The original rotary position encoding of LLaMA-7B is RoPE, and the calculation formula of the self-attention head is as follows:
[0086]
[0087]
[0088] Among them, RoPE () is the position encoding function, is the query matrix encoded by the RoPE method, is the key matrix encoded by the RoPE method, V is the value matrix of the target domain text sequence; is the self-attention head, is the activation function, is the dimension of the self-attention head , is the transpose matrix of.
[0089] For the cross-attention head, while maintaining the RoPE positional encoding, the target domain text feature vector is introduced for interaction. The calculation formula of the cross-attention head is as follows:
[0090]
[0091] Among them, is the cross-attention head, is the query matrix encoded by the RoPE method, is the key matrix encoded by the RoPE method, V is the value matrix of the target domain text sequence, is the average feature vector of the target domain text sequence, is the activation function, is the average feature vector after normalization calculation, is the transposed matrix of.
[0092] By introducing the target domain text feature vector in the cross-attention head, the model can explicitly utilize the semantic information of the target domain to guide feature fusion. This guiding role can help the model better adapt to the distribution of the target domain, thereby improving the effect of domain transfer; the feature fusion in the cross-attention head can effectively align the feature spaces of the source domain and the target domain, reducing the negative impact of domain differences on the model performance; by fusing the target domain feature vector in the cross-attention head, the model can effectively transfer the knowledge of the source domain to the target domain, thereby improving the generalization ability of the model in the target domain.
[0093] In step 104, the intermediate Transformer model is trained using the target domain text to obtain the target Transformer model.
[0094] In this step, the intermediate Transformer model is trained using the target domain text to obtain the target Transformer model.
[0095] In terms of parameter optimization, exemplarily, the QLoRA (Quantized Low-Rank Adaptation) method can be used for efficient fine-tuning. First, the weights of the intermediate Transformer model obtained in the previous step are quantized to 4-bit precision, and then low-rank adapters are added on the basis of the quantized weights:
[0096]
[0097] Among them, is the weight matrix after adding low-rank adapters on the basis of the quantized weights, The original weights after 4-bit quantization, B and A are low-rank matrices with 16-bit precision, and the rank is set to 64. Fine-tuning is performed by the QLoRA method, which significantly reduces the computational cost and memory footprint while maintaining high performance.
[0098] It should be noted that B and A are low-rank adapters with complementary dimensions. A is a dimensionality reduction matrix that compresses the original weight dimensions into a low-rank space; B is a dimensionality increase matrix that maps the features in the low-rank space back to the original dimensions; through the collaborative action of the two low-rank matrices A and B, efficient parameter updates are achieved while maintaining the model performance.
[0099] In terms of the training strategy, exemplarily, a phased training scheme can be adopted:
[0100] The first stage: Freeze all other parameters except the low-rank adaptation (LoRA) parameters in the cross-attention heads, and use the target domain text to train the LoRA parameters; among them, the LoRA parameters are adjustment parameters introduced through low-rank decomposition.
[0101] The second stage: Unfreeze the self-attention parameters of the specified key layers of the intermediate Transformer model for fine-tuning to obtain the target Transformer model.
[0102] The AdaFactor optimizer can be used, with a base learning rate of 1e-4, and a cosine learning rate scheduling strategy is adopted.
[0103] In one implementation, the self-attention parameters of the last 3 Transformer layers of the intermediate Transformer model can be unfrozen for fine-tuning to obtain the target Transformer model. The last 3 Transformer layers are close to the output end and are responsible for high-level semantic feature integration. After unfreezing, they can more flexibly adapt to downstream recognition tasks.
[0104] The target Transformer model can be used to identify target domain text data, has a high recognition accuracy, good transfer stability, and the target Transformer model can achieve a good adaptation effect without a large amount of data in the target text domain, consumes less computing resources, and has a low computational cost; it can fully integrate the features of the target domain text and deeply integrate the features of the target domain text with the core attention mechanism of the pre-trained language model; the transfer effect is stable, and in low-computing resource scenarios, an efficient domain transfer effect can be achieved.
[0105] The training method, device, electronic device and readable storage medium of the Transformer model provided by the present invention can divide the attention heads in each layer of the original Transformer model into self-attention heads and cross-attention heads, introduce the text feature vectors of the target domain in the cross-attention heads for feature fusion, and use the text of the target domain to train the original Transformer model to obtain the target Transformer model. It can achieve the text transfer task of large-scale pre-trained language models from the source domain to the target domain with high quality in low-computing-resource scenarios, significantly reduce the need for text training data in the target domain during domain transfer, reduce the computing resource cost, and improve the adaptability of large-scale pre-trained language models to recognize texts in different domains.
[0106] The training device of the Transformer model provided by the present invention will be described below. The training device of the Transformer model described below can be correspondingly referred to the training method of the Transformer model described above.
[0107] Figure 3 It is a schematic structural diagram of a training device of a Transformer model provided by the present invention. The training device of the Transformer model is applied to devices such as servers, desktops, and laptops. Referring to Figure 3 Figure, the training device 30 of the Transformer model includes a processing module 301, a feature extraction module 302, a feature fusion module 303, and a training module 304.
[0108] The processing module 301 is configured to divide the attention heads in each layer of the original Transformer model into self-attention heads and cross-attention heads; the original Transformer model is trained based on the source domain text;
[0109] The feature extraction module 302 is configured to extract features from the target domain text to obtain target domain text feature vectors;
[0110] The feature fusion module 303 is configured to introduce the target domain text feature vectors in the cross-attention heads for feature fusion to obtain an intermediate Transformer model;
[0111] The training module 304 is configured to use the target domain text to train the intermediate Transformer model to obtain a target Transformer model.
[0112] Optionally, the training module 304 is further configured to freeze all other parameters except the LoRA parameters in the cross-attention heads of the intermediate Transformer model, and train the LoRA parameters using the target domain text; the LoRA parameters are adjustment parameters introduced through low-rank factorization;
[0113] Unfreeze the self-attention parameters of the specified key layers of the intermediate Transformer model for fine-tuning to obtain the target Transformer model.
[0114] Optionally, the training module 304 is further configured to unfreeze the self-attention parameters of the last 3 Transformer layers of the intermediate Transformer model for fine-tuning to obtain the target Transformer model.
[0115] Optionally, the feature extraction module 302 is further configured to vectorize the target domain text respectively using a vector model with the same structure as the original Transformer model to obtain multiple text vectors;
[0116] Calculate the arithmetic mean of the multiple text vectors to obtain the target domain average feature vector;
[0117] Normalize the target domain average feature vector to obtain the target domain text feature vector.
[0118] Optionally, the calculation formula of the cross-attention head is:
[0119]
[0120] where, is the cross-attention head, Q is the query matrix of the input sequence, K is the key matrix of the input sequence, V is the value matrix of the input sequence, is the average feature vector of the input sequence, is the dimension of the self-attention head, is the key matrix K is the transposed matrix of, is the average feature vector after normalization calculation, is is the transposed matrix of.
[0121] Optionally, the original Transformer model uses LLaMA-7B as the base model.
[0122] Optionally, the processing module 301 is further configured to divide the 32 attention heads of each layer of the original Transformer model into 24 self-attention heads and 8 cross-attention heads.
[0123] Optionally, the processing module 301 is further configured to adjust the quantity ratio of the self-attention heads to the cross-attention heads based on actual task requirements.
[0124] Figure 4 An entity structure diagram of an electronic device is exemplified, as Figure 4 shown. The electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communications interface 420, and the memory 430 complete communication with each other through the communication bus 440. The processor 410 may call the logical instructions in the memory 430 to execute a method for training a Transformer model. The method includes: dividing the attention heads in each layer of the original Transformer model into self-attention heads and cross-attention heads; the original Transformer model is trained based on source domain texts; extracting feature vectors of target domain texts to obtain target domain text feature vectors; introducing the target domain text feature vectors into the cross-attention heads for feature fusion to obtain an intermediate Transformer model; using the target domain texts to train the intermediate Transformer model to obtain a target Transformer model.
[0125] In addition, when the logical instructions in the foregoing memory 430 can be implemented in the form of software functional units and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0126] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the training method of the Transformer model provided by each of the above methods. The method includes: dividing the attention heads in each layer of the original Transformer model into self-attention heads and cross-attention heads; the original Transformer model is trained based on source domain texts; extracting features from target domain texts to obtain target domain text feature vectors; introducing the target domain text feature vectors into the cross-attention heads for feature fusion to obtain an intermediate Transformer model; using the target domain texts to train the intermediate Transformer model to obtain a target Transformer model.
[0127] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the training method of the Transformer model provided by each of the above methods. The method includes: dividing the attention heads in each layer of the original Transformer model into self-attention heads and cross-attention heads; the original Transformer model is trained based on source domain texts; extracting features from target domain texts to obtain target domain text feature vectors; introducing the target domain text feature vectors into the cross-attention heads for feature fusion to obtain an intermediate Transformer model; using the target domain texts to train the intermediate Transformer model to obtain a target Transformer model.
[0128] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0129] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0130] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A training method for a Transformer model, characterized in that: include: The attention heads in each layer of the original Transformer model are divided into self-attention heads and cross-attention heads; The original Transformer model is trained based on source domain text; Extract features from the target domain text to obtain the target domain text feature vector; Introducing the target domain text feature vector into the cross-attention head to perform feature fusion to obtain an intermediate Transformer model; Using the target domain text to train the intermediate Transformer model to obtain a target Transformer model includes: Freeze all parameters except the LoRA parameters in the cross-attention head of the intermediate Transformer model, and train the LoRA parameters using the target domain text; the LoRA parameters are adjustment parameters introduced by low-rank decomposition; The self-attention parameters of the specified key layers of the intermediate Transformer model are unfrozen and fine-tuned to obtain the target Transformer model.
2. The method according to claim 1, characterized in that The unfreezing of the self-attention parameters of the designated key layer of the intermediate Transformer model is fine-tuned to obtain the target Transformer model, including: The self-attention parameters of the last three Transformer layers of the intermediate Transformer model are unfrozen and fine-tuned to obtain the target Transformer model.
3. The method according to claim 1, characterized in that The feature extraction of the target domain text to obtain the target domain text feature vector includes: Using a vector model consistent with the structure of the original Transformer model to vectorize the target domain texts respectively, to obtain multiple text vectors; Calculating the arithmetic mean of the multiple text vectors to obtain an average feature vector of the target domain; The target domain average feature vector is normalized to obtain the target domain text feature vector.
4. The method according to claim 1, characterized in that: The calculation formula for the cross-attention head is: CrossAttention(Q,K,V,D)=softmax((Q K^T) / √(d_k )+(Q D_norm^T) / √(d_k )) (V+D_norm ) Among them, CrossAttention(Q,K,V,D) is the cross-attention head, Q is the query matrix of the input sequence, K is the key matrix of the input sequence, V is the value matrix of the input sequence, D is the average eigenvector of the input sequence, d_k is the dimension of the self-attention head, K^T is the transposed matrix of the key matrix K, D_norm is the average eigenvector after normalization calculation, and D_norm^T is the transposed matrix of D_norm.
5. The method according to claim 1, characterized in that The original Transformer model uses LLaMA-7B as the base model.
6. The method according to claim 5, characterized in that The attention heads in each layer of the original Transformer model are divided into self-attention heads and cross-attention heads, including: The 32 attention heads of each layer of the original Transformer model are divided into 24 self-attention heads and 8 cross-attention heads.
7. The method according to claim 1, characterized in that The attention heads in each layer of the original Transformer model are divided into self-attention heads and cross-attention heads, including: The ratio of the number of the self-attention heads to the cross-attention heads is adjusted based on actual task requirements.
8. A training device for a Transformer model, characterized in that: include: A processing module configured to separate the attention heads in each layer of the original Transformer model into self-attention heads and cross-attention heads; The original Transformer model is trained based on source domain text; A feature extraction module is configured to extract features from the target domain text to obtain a target domain text feature vector; A feature fusion module is configured to introduce the target domain text feature vector into the cross-attention head to perform feature fusion to obtain an intermediate Transformer model; A training module is configured to train the intermediate Transformer model using the target domain text to obtain a target Transformer model, including: freezing all parameters except the LoRA parameters in the cross-attention head of the intermediate Transformer model, and training the LoRA parameters using the target domain text; the LoRA parameters are adjustment parameters introduced by low-rank decomposition; unfreezing the self-attention parameters of the specified key layer of the intermediate Transformer model for fine-tuning to obtain the target Transformer model.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, it implements the training method of the Transformer model as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the training method of the Transformer model as described in any one of claims 1 to 7 is implemented.
11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the training method of the Transformer model as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Transfer learning entity relationship extraction method, device and equipment based on domain self-adaption and storage medium
CN119357378A
Fine tuning training method and device for domain large language model, electronic equipment and medium
CN119538981A