A pre-training method and device for fusing multi-layer feedforward representations
By introducing multi-task learning and Self-train strategies into the pre-trained language model, effectively utilizing training data and unsupervised data, the problem of data value not being mined and model optimization target deviation in the existing technology is solved, and effective fusion of multi-layer vector representation and improvement of model generalization capabilities are achieved.
Patent Information
- Application Number
- CN202210433291.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-24
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-04-24
AI Technical Summary
In the prior art, the value of training data has not been fully mined, unsupervised data in specific classification tasks has not been effectively utilized, text semantic representations of each layer of the network have not been fully used and expressed in extreme terms, and there are deviations in the selection and design of loss functions, resulting in deviations in the model optimization target, and the traditional gradient update method has caused the model to fall into local optimization.
A pre-training method that integrates multi-layer feedforward representation is adopted. By setting specific task categories for multi-task learning, including NSP next sentence prediction and SQuAD reading comprehension tasks, the self-train strategy is used to effectively utilize unsupervised data, and by designing reasonable network structure and loss functions, the effective fusion and optimization of text vector representations at each layer is achieved.
By effectively utilizing training data and unsupervised data, the global text representation ability of word vectors is improved, the effective fusion of multi-layer vector representation is achieved, the learning ability of the model is enhanced, the problems of high deviation and high variance are avoided, and the generalization ability of the model is improved.
Smart Images

Figure CN114912606B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning, and specifically provides a pre-training method and device that fuse multi-layer feedforward representations. Background Art
[0002] In recent years, the two major problems faced by DL (Deep Learning) are the contradiction between the extremely small amount of data in vertical domain-specific tasks and the demand of DL for massive data, and the contradiction between the demand of SL (Supervised Learning) for a large amount of labeled data and the extremely high cost of labeling massive data.
[0003] Looking at the development of the NLP (Natural Language Processing) field in the past few decades, it can be generally summarized as certain forms of language modeling, so as to achieve a more effective representation of natural language text. The development of the representation of natural language text, restricted by the sparsity of language, mostly starts from the representation at the word granularity to construct a language model. From the development of statistical language models to neural language models to some current relatively good pre-trained semantic (language) models such as ELMo (Embeddings from Language Models) / ULMFit (Universal Language Model Fine-tuning), GPT (Generative Pre-trained Transformer), BERT, etc., have achieved SOTA (State-of-the-Art) on multiple GLUE task leaderboards. The underlying idea is Transfer (transfer learning) and language model modeling.
[0004] Pre-train in Transfer plays an effective role in the utilization of massive open-source big datasets, so that only a small amount of Specific-task (specific task) data is required during Finetune (fine-tuning) to achieve better results. And language model modeling, based on the distribution hypothesis and through network structuring to achieve long-range dependencies in context, as well as task transformations on Speific-task such as MLM (Masked Language Model), NSP (Next Sentence Prediction Task), etc., truly realizes getting rid of the limitation of data labeling, and ultimately realizes the effective understanding of text semantics based on this.
[0005] However, in the prior art, the value of training data has not been fully exploited, unsupervised data in specific classification tasks has not been effectively utilized, and the text semantic representations of each layer of the network have not been fully used and optimally expressed. And in the selection and design of the loss function, the real problem is not reflected, resulting in a deviation in the model optimization goal, and the traditional gradient update method causes the model to fall into the dilemma of local optimality. Summary of the Invention
[0006] In view of the above deficiencies of the prior art, the present invention provides a practical pre-training method that fuses multi-layer feedforward representations.
[0007] A further technical task of the present invention is to provide a pre-training device that is reasonably designed, safe and applicable and integrates multi-layer feedforward representations.
[0008] The technical solution adopted by the present invention to solve its technical problems is as follows:
[0009] A pre-training method that integrates multi-layer feedforward representations has the following steps:
[0010] S1. Collect text data;
[0011] S2. Set specific task categories for multi-task learning, including sentence pair tasks for NSP next sentence prediction and SQuAD reading comprehension tasks;
[0012] S3. According to the selected task types, preprocess the corresponding texts respectively, including supervised labeling tasks and Self-train strategy customization for unlabeled data;
[0013] S4. Set the network structure and write code;
[0014] S5. Develop and write code to achieve the fusion of flattened text vectors between layers;
[0015] S6. Design and implement the MLP for Specific-task;
[0016] S7. Develop data strategies and algorithms and write code;
[0017] S8. Integrate the code in steps S4 to S7 and perform an End-to-End full network feedforward process;
[0018] S9. Use the preprocessed text data to train the encoder network that integrates multi-layer feedforward representations to achieve global optimality;
[0019] S10. Serialize the pre-trained language model that integrates multi-layer feedforward representations;
[0020] S11. Connect the Encoder to the Specific-task post-processing model respectively, and use the test data to evaluate the performance of the encoder network that integrates multi-layer feedforward representations.
[0021] Further, in step S3, in terms of Data Augmentation, for the corpus of specific tasks for post-position word classification, perform Word Mixup based on Skip-Gram Word Embedding, and at the same time perform LabelSmoothing for the labels;
[0022] Integrate Self-training weak supervision learning and Pure Semi-supervised Learning to effectively utilize unsupervised data.
[0023] Furthermore, in step S4, the Encoder part uses a 14-head Multi-headed Attention mechanism and Position Embedding to actively amplify the Sequence Mask. For the global vector representation of multiple layers of Encoder, a 12-layer Feed Forward structure of BERT base-chinese is used.
[0024] Furthermore, in step S5, for the fusion of multiple-layer vector representations, two fusion strategies are adopted. One fusion strategy is to refer to SENet to perform LN operation on each layer of representation, and perform one-dimensional global maximum pooling, and then connect to a 2-layer FC to obtain the importance degree of each layer of vector representation, and finally perform weighted fusion on multiple-layer vector representations.
[0025] Furthermore, the second fusion strategy in the two fusion strategies is to regard the layer relationship of multiple-layer vector representations as the Channel depth relationship. First, reduce the channels through Point-wise Convolution with fewer channels than the number of Channels and alleviate aliasing, and then perform Point-wise Convolution with a single filter to Flatten the features into a 1d vector. Immediately connect to the output layer to construct an FC network. The output dimension of the FC network is the same as the dimension of the input 1d vector, so as to achieve the fusion of multiple-layer vector representations through the special setting of the network structure.
[0026] Furthermore, in step S5, for the Feed Forward part, refer to CSPDarknet-53 to adjust the ResNet-shortcut of BERTbase to a CSP structure, set the number of Bottleneck modules to 6, replace CSP with 1d convolution, and retain the BN operation. At the same time, use the GELU activation function.
[0027] Furthermore, in step S7, construct a Multi-task learning training objective, implement and verify the comparison of the loss functions corresponding to various GLUE tasks through experimental Coding, and finally select Soft F1 Loss to replace the cross-entropy loss in the original network as the final policy element.
[0028] Further, in step S8, Adam that introduces the exponentially weighted moving average and Momentum is used, and a network is designed on the Specific-task layer. An 8-layer FC is connected after BERT to form an MLP, where the number of network layers of the FC is used as a hyperparameter for GridSearch / RandomSearch tuning.
[0029] A pre-training device that fuses multi-layer feed-forward representations, characterized by comprising: at least one memory and at least one processor;
[0030] The at least one memory is used to store machine-readable programs;
[0031] The at least one processor is used to call the machine-readable program to execute a pre-training method that fuses multi-layer feed-forward representations.
[0032] Compared with the prior art, the pre-training method and device that fuse multi-layer feed-forward representations of the present invention have the following prominent beneficial effects:
[0033] (1) By introducing a Dense Block into the encoder Feed Forward structure, the present invention realizes the effective utilization of text vector representations of each encoder layer, and effectively improves the global text representation ability of word vectors through a fusion means.
[0034] (2) The fusion strategy realized through the design of a pure network structure and the vector representation fusion realized through the squeeze-and-excitation weighting mechanism can more effectively reflect the deep relationship between n text vector representations and the flattened word vector representations compared with the simple ADD method, and finally give reliable depth correlation information, improving the representation learning ability of the network.
[0035] (3) By designing a strategy and optimization algorithm that are more suitable for actual specific problems, the learning (fitting) ability of the model is enhanced, the problems of high bias and high variance are effectively avoided, and the generalization ability of the model is enhanced. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0037] Attached Figure 1 is a schematic flowchart of FeedForward adding a Dense Block in a pre-training method that fuses multi-layer feed-forward representations;
[0038] Appendix Figure 2 It is a schematic flow chart for obtaining the weight coefficients of the representations of each layer by the SE mechanism in a pre-training method that fuses multi-layer feed-forward representations;
[0039] Appendix Figure 3 It is a schematic diagram for fusing the features of each layer by point-wise convolution connected to FC in a pre-training method that fuses multi-layer feed-forward representations;
[0040] Appendix Figure 4 A schematic flow chart of a pre-training method that fuses multi-layer feed-forward representations. Specific implementation manners
[0041] In order to enable those skilled in the art to better understand the solution of the present invention, the present invention will be further described in detail below in conjunction with specific implementation manners. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the protection scope of the present invention.
[0042] The following gives a best embodiment:
[0043] A pre-training method that fuses multi-layer feed-forward representations in this embodiment is characterized by the following steps:
[0044] S1. Collect text data:
[0045] Collect text data such as open-source Wikipedia / Books corpus, address information of third-party statistical bureaus, commodity names in the internal tax field, and production and operation addresses of taxpayers.
[0046] S2. Set specific task categories for multi-task learning, including sentence pair tasks for next sentence prediction under NSP and SQuAD reading comprehension tasks:
[0047] Set the task categories of multi-task learning as Cloze cloze (including 15% Mask, where 70% are real Masks, 20% retain the original words, and 10% are randomly replaced with other words), classification prediction of true and false production and operation addresses of taxpayers, same-site verification classification of address sentence pairs, and matching classification of commodity names and codes (set the coding value granularity to the first 9 digits), and classify through the sentence pair task of next sentence prediction.
[0048] S3. According to the selected task types, preprocess the corresponding texts respectively, including supervised labeling tasks and Self-train (Pure semi-supervised learning) strategy customization for unlabeled data.
[0049] In Data Augmentation, for the corpus of the specific task of post-word classification, Word Mixup is performed based on Skip-Gram's Word Embedding, and Label Smoothing is performed on the labels, so as to achieve the effect of data augmentation to enhance the robustness and generalization ability of the model.
[0050] Integrate Self-training weakly supervised learning and Pure Semi-supervised Learning to achieve effective utilization of unsupervised data.
[0051] S4. Setting the network structure and writing code, including adding Dense Block between Encoder layers and adding 6-branch Bottleneck in the Residual part.
[0052] like Figure 1 As shown in the figure, the Encoder part uses a 14-head Multi-headed Attention mechanism and Position Embedding, so that the semantic information in the global context can be better extracted, which in turn relaxes the structure of the network.
[0053] Actively enlarge the Sequence Mask to allow longer text to enter the network, better enhance the long-range dependency on the context, and make more contextual contexts effectively utilized.
[0054] For the global vector representation of multi-layer encoders, the 12-layer Feed Forward structure of BERT base-chinese is used, and the Dense block module is introduced in each layer of encoder by referring to Densenet.
[0055] S5. Develop and write code to achieve the fusion of inter-layer flattened text vectors, including SE structure fusion and Point-wise Conv+FC structure fusion.
[0056] like Figure 2 , 3 As shown, this embodiment proposes two fusion strategies for the fusion of multi-layer vector representations through experimental verification. The starting point of the two strategies is to concatenate / flatten the vector representation of each layer of Sequence into a 1d vector. The physical meaning of this operation is to directly retain the word order through the index of the vector, which enhances OptionalEmbedding in a certain sense.
[0057] The first method (Figure 2 ), refer to SENet to perform LN (Layer Nornalization) operation on each layer of representation, and perform one-dimensional global maximum pooling, and then access 2 layers of FC (Full Connection) to obtain the importance of each layer of vector representation, and finally perform weighted fusion of multi-layer vector representation.
[0058] The second method ( Figure 3 ), the layer relationship of the multi-layer vector representation is regarded as the channel depth relationship, and the channel is first reduced and the aliasing problem is alleviated through Point-wise Convolution (1×1Conv) with less than the number of channels, and then the single-filter Point-wise Convolution (1×1Conv) is performed to Flatten the features into 1d vectors, and then the output layer is connected to build an FC network. The output dimension of the FC network is equivalent to the dimension of the input 1d vector, thereby realizing the fusion of multi-layer vector representations through the special setting of the network structure.
[0059] In the feed forward part, we refer to CSPDarknet-53 to adjust the ResNet-shortcut of BERT base to CSP (Cross Stage Partial) structure. The number of Bottleneck modules in the model is explicitly set to 6. The 2 / 3d convolution operation of CSP is replaced with 1d convolution, and the BN (Batch Normalization) operation is retained. At the same time, the Mish activation function is replaced with the GELU activation function.
[0060] S6. Specific-task MLP design and programming implementation.
[0061] S7. Develop data strategies and algorithms, and write code.
[0062] Construct a multi-task learning training objective, implement and verify the loss functions corresponding to various GLUE tasks through experimental coding, and finally choose Soft F1 Loss to replace the cross entropy loss in the original network as the final strategy element.
[0063] S8. Integrate the codes from steps S4 to S7 to perform an end-to-end full-network feedforward process.
[0064] like Figure 4As shown in the figure, by introducing Adam with exponentially weighted moving average and Momentum, and designing a relatively complex network on the Specific-task layer, for the word / MASK category output problem of MLM, an 8-layer FC is connected after BERT to form an MLP, which improves the complexity of the classification model, sets a relatively small learning rate, and effectively avoids overfitting and underfitting problems through the robust update of the gradient, thereby achieving the improvement of the model effect. The number of network layers of FC can be used as a hyperparameter for GridSearch / RandomSearch tuning.
[0065] S9. Train the encoder network that fuses multi-layer feedforward representations using the preprocessed text data to reach the global optimum.
[0066] S10. Serialize the pre-trained language model that fuses multi-layer feedforward representations.
[0067] S11. The Encoder is respectively followed by a Specific-task post-processing model, and the performance of the encoder network that fuses multi-layer feedforward representations is evaluated using the test data.
[0068] Based on the above method, a pre-training device that fuses multi-layer feedforward representations in this embodiment is characterized by including: at least one memory and at least one processor;
[0069] The at least one memory is used to store machine-readable programs;
[0070] The at least one processor is used to call the machine-readable program and execute a pre-training method that fuses multi-layer feedforward representations.
[0071] The above specific implementation manners are only specific cases of the present invention. The patent protection scope of the present invention includes but is not limited to the above specific implementation manners. Any implementation that conforms to the claims of a pre-training method and device that fuses multi-layer feedforward representations of the present invention and any appropriate changes or substitutions made by those of ordinary skill in the art in any technical field shall fall within the patent protection scope of the present invention.
[0072] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirits of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A pre-training method that fuses multi-layer feed-forward representations, characterized in that, it has the following steps: S1. Collect text data; S2. Set specific task categories for multi-task learning, including the sentence pair task of next sentence prediction under NSP and the SQuAD reading comprehension task; S3. According to the selected task types, preprocess the corresponding texts respectively, including supervised labeling tasks and Self-train strategy customization for unlabeled data; In terms of Data Augmentation, for the corpus of the specific task of post-position word classification, perform Word Mixup based on Skip-Gram's WordEmbedding, and at the same time perform Label Smoothing for the labels; Fuse Self-training weak supervision learning Pure Semi-supervised Learning to effectively utilize unsupervised data; S4. Set the network structure and write the code; For the Encoder part, use a 14-head Multi-headed Attention mechanism and PositionEmbedding position embedding, actively amplify the Sequence Mask, and for the global vector representation of multiple layers of Encoder, use the 12-layer Feed Forward structure of BERTbase-chinese; S5. Develop and write the code to achieve the fusion of flattened text vectors between layers; In the fusion of multi-layer vector representations, two fusion strategies are adopted. One fusion strategy is to refer to SENet to perform LN operation on each layer of representation, and perform one-dimensional global maximum pooling, and then connect to a 2-layer FC to obtain the importance degree of each layer of vector representation, and finally perform weighted fusion on the multi-layer vector representations; The second fusion strategy in the two fusion strategies described above is to regard the layer relationship of multi-layer vector representations as the Channel depth relationship. First, perform channel reduction through Point-wise Convolution with fewer than the number of Channels to alleviate aliasing, and then perform Point-wise Convolution with a single filter to Flatten the features into a 1d vector, and then connect to the output layer to construct an FC network. The output dimension of the FC network is the same as the dimension of the input 1d vector, so as to achieve the fusion of multi-layer vector representations through the special setting of the network structure; For the Feed Forward part, refer to CSPDarknet-53 to adjust the ResNet-shortcut of BERT base to a CSP structure, set the number of Bottleneck modules to 6, replace CSP with 1d convolution, and retain the BN operation, and at the same time adopt the GELU activation function; S6. Design and program the MLP for Specific-task; S7. Develop data strategies and algorithms and write the code; Construct the training objective of multi-task learning, implement and verify the comparison of the loss functions corresponding to various tasks in GLUE through experimental coding, and finally select Soft F1 Loss to replace the cross-entropy loss in the original network as the final policy element; S8. Integrate the codes in steps S4 to S7 to perform the end-to-end full network feedforward process; By introducing Adam with exponential weighted moving average and Momentum, and designing a network on the specific-task layer, connect an 8-layer FC after BERT to form an MLP, where the number of network layers of FC is used as a hyperparameter for GridSearch / RandomSearch tuning; S9. Use the preprocessed text data to train the encoder network that fuses multi-layer feedforward representations to achieve global optimality; S10. Serialize the pre-trained language model that fuses multi-layer feedforward representations; S11. The encoder is respectively followed by a specific-task post-processing model, and the test data is used to evaluate the performance of the encoder network that fuses multi-layer feedforward representations.
2. A pre-training device that fuses multi-layer feedforward representations, characterized in that, it includes: at least one memory and at least one processor; the at least one memory is used to store machine-readable programs; the at least one processor is used to call the machine-readable program and execute the method described in claim 1.
Citation Information
Patent Citations
Text abstract automatic generation method and system fused with pre-training model
CN112765345A