A lightweight semi-supervised model framework for few-sample domains

By introducing a lightweight semi-supervised model framework of excitation networks and target networks in natural language processing, the problems of high dependence on pre-trained models and large number of model parameters in the prior art are solved, and a technical solution to improve prediction speed and effect with few sample data is realized.

CN113920395BActive Publication Date: 2025-05-13BEIJING SHANGJIAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111166569.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-30
Publication Date
2025-05-13
Estimated Expiration
2041-09-30

AI Technical Summary

Technical Problem

When the prior art uses a small amount of labeled data for natural language processing, the cost of relying on pre-trained models is high and the number of model parameters is large, resulting in slow prediction speed.

Method used

A lightweight semi-supervised model framework for the small sample field is proposed. By introducing motivation networks and target networks, using consistency regularization and data augmentation technology, information and features are mined from unsupervised data and supervised data, providing multi-level regularization constraints for target networks, reducing the amount of model parameters and improving prediction speed.

Benefits of technology

It is realized that while reducing the dependence on pre-trained models and reducing the amount of model parameters, the prediction speed of the model is improved, and only a small amount of labeled data is required in the text classification task to achieve good results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113920395B_ABST
    Figure CN113920395B_ABST
Patent Text Reader

Abstract

The present invention discloses a lightweight semi-supervised model framework for a few-sample field, which relates to the technical field of deep learning natural language processing, including: an excitation network as pre-training; a target network as training, the target network is connected to the excitation network, and completes several feature distillations with the excitation network; wherein the excitation network uses consistency regularization and data enhancement technology to mine information and features from unsupervised data and supervised data, and provides multi-level regularization constraints for subsequent training of the target network. The above framework provided by the present invention adopts a semi-supervised learning method, and only a small amount of labeled data is needed in text classification to achieve good results; the above framework provided by the present invention adopts a two-stage training method of using an excitation network to guide the target network, because the final target network is a lightweight convolutional neural network, therefore, the number of parameters is smaller, fewer resources are required during operation, and the speed is faster.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning natural language processing technology, and in particular to a lightweight semi-supervised model framework for a few-sample field. Background Art

[0002] In recent years, with the widespread application of deep learning, natural language processing technology has also transitioned from traditional methods such as statistics and machine learning to deep learning methods, and has already had very mature applications, such as text classification, information extraction, machine translation, sentiment analysis, and reading comprehension. These successful applications benefit from the rapid development of deep learning models, such as BERT, XLNET, etc.

[0003] However, the above progress depends largely on large-scale and high-quality manually annotated data. It is very expensive to obtain a large amount of high-quality labeled data, especially in the fields of finance, medicine, law, etc. Text annotation relies on the in-depth participation of domain experts. How to achieve good results with a small amount of annotated data is a new direction in the development of natural language processing.

[0004] Among them, the semi-supervised method is a technical route to solve the above problems. The Unsupervised Data Augmentation for Consistency Training (UDA) model is a semi-supervised learning method based on the BERT pre-training model. The training data includes a small amount of supervised data and a large amount of unsupervised data. The unsupervised data includes the unsupervised data itself and the corresponding enhanced data. The loss function consists of two parts, one of which is the cross entropy loss function of the supervised data, and the other is the unsupervised loss. After data enhancement by unsupervised data, the consistency regularization of the two is used as the loss function. On IMDB, UDA achieves an accuracy of 91% by using 20 supervised data. However, there are at least two defects in using the UDA model: 1. It is more dependent on the previous pre-training model, and the cost of the pre-training model is relatively high; 2. The UDA model is based on the BERT model, so the number of parameters is relatively large and the prediction speed is not fast.

[0005] Therefore, technical personnel in this field are committed to developing a lightweight semi-supervised model framework for the field of few samples, reducing the dependence on pre-trained models in existing technical solutions, reducing the number of model parameters, and improving the prediction speed of the model. Summary of the invention

[0006] In view of the above-mentioned defects of the prior art, the technical problem to be solved by the present invention is how to improve the prediction speed of the model while reducing the dependence on the pre-trained model and reducing the number of parameters of the model.

[0007] To achieve the above objectives, the present invention provides a lightweight semi-supervised model framework for the few-sample domain, including:

[0008] As a pre-trained incentive network;

[0009] As a target network for training, the target network is connected to the excitation network and performs several feature distillations with the excitation network;

[0010] in,

[0011] The excitation network utilizes consistency regularization and data enhancement technology to mine information and features from unsupervised data and supervised data, and provides multi-level regularization constraints for subsequent training of the target network.

[0012] Further, the excitation network includes a first text encoder, a first text classifier and a first feature projection which are connected in sequence;

[0013] in,

[0014] The first text encoder is a pre-trained language model, including a plurality of stacked transformer layers, the unsupervised data and the supervised data are taken as input, and are pre-trained by the first text editor, and output is a first vector;

[0015] The first text classifier includes a first multi-layer perceptron that receives the first vector and fine-tunes a downstream classification task;

[0016] The first feature projection includes a single-layer perceptron and a nonlinear activation function, which can align feature dimensions of the excitation network and the target network.

[0017] Further, the target network includes a second text encoder, a second text classifier, and a second feature projection connected in sequence;

[0018] in,

[0019] The second text encoder includes a plurality of convolutional neural networks TextCNN, the unsupervised data and the supervised data are used as input, and are trained by the second text editor, and output as a second vector;

[0020] The second text classifier includes a second multilayer perceptron and the received second vector;

[0021] The structure of the second feature projection is the same as that of the first feature projection, and further includes a feature mapping layer.

[0022] Furthermore, the first text encoder is a BERT model, and the BERT model includes a low-layer BERT, a middle-layer BERT, and a high-layer BERT;

[0023] in,

[0024] The low-level BERT captures low-level word-level features;

[0025] The middle-layer BERT captures intermediate grammatical level features;

[0026] The high-level BERT captures high-level semantic-level features.

[0027] Furthermore, the convolutional neural network TextCNN includes CNN convolution kernels of different sizes, namely, a CNN small filter, a CNN medium filter, and a CNN large filter;

[0028] in,

[0029] The CNN small filters capture the word-level features;

[0030] The CNN medium filter captures the grammatical level features;

[0031] The CNN large filters capture the semantic level features.

[0032] Furthermore, the CNN small filter is aligned with the low-layer BERT; the CNN medium filter is aligned with the middle-layer BERT; and the CNN large filter is aligned with the high-layer BERT.

[0033] Furthermore, the loss function of the excitation network training includes a first cross entropy loss of the supervised data and a first consistent regularization loss of the unsupervised data.

[0034] Furthermore, the loss function of the target network training includes the distillation loss of the model output layer, the distillation loss of the latent space feature layer and the second consistent regularization loss.

[0035] Furthermore, the distillation loss of the latent space feature layer uses the first feature projection and the second feature projection to match the hidden state and feature map of the transformer layer, and minimizes the mean square error between the first feature projection and the second feature projection to complete knowledge distillation.

[0036] Furthermore, the first multilayer perceptron has two layers.

[0037] The lightweight semi-supervised model framework for the few-sample field provided by the present invention has at least the following technical effects:

[0038] 1. The present invention adopts a semi-supervised learning method. In text classification, only a small amount of labeled data is needed to achieve good results;

[0039] 2. The present invention adopts a two-stage training method of using an excitation network to guide the target network. Because the final target network is a lightweight convolutional neural network, the number of parameters is smaller, fewer resources are required during operation, and the speed is faster.

[0040] The concept, specific structure and technical effects of the present invention will be further described below in conjunction with the accompanying drawings to fully understand the purpose, characteristics and effects of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is a schematic diagram of the overall model structure of a preferred embodiment of the present invention;

[0042] Figure 2 yes Figure 1 A schematic diagram of the structure of latent space feature distillation in the illustrated embodiment. DETAILED DESCRIPTION

[0043] The following describes several preferred embodiments of the present invention with reference to the drawings in the specification, so that the technical content is clearer and easier to understand. The present invention can be embodied in many different forms of embodiments, and the protection scope of the present invention is not limited to the embodiments mentioned in the text.

[0044] Figure 1 It is a schematic diagram of the overall model structure of a preferred embodiment of the present invention, wherein FD represents the distillation loss of the intermediate feature layer, and OD represents the distillation loss of the output layer.

[0045] The biggest difference between the technical solution provided by the present invention and the previous semi-supervised solution is that an activation network, or an excitation network, is introduced outside the lightweight target network. The excitation network fully mines information and features from unlabeled data and limited labeled data using consistency regularization and data enhancement techniques, and then provides regularization constraints on the output layer and the intermediate hidden layer to guide the training of the lightweight target network.

[0046] Specifically, the present invention provides a lightweight semi-supervised model framework for the field of few samples, including:

[0047] As a pre-trained incentive network;

[0048] As the target network for training, the target network is connected to the excitation network and performs several feature distillations with the excitation network;

[0049] in,

[0050] The incentive network uses consistency regularization and data augmentation techniques to mine information and features from unsupervised and supervised data, providing multi-level regularization constraints for the subsequent training of the target network.

[0051] The excitation network includes a first text encoder, a first text classifier and a first feature projection which are connected in sequence.

[0052] in,

[0053] The first text encoder is a pre-trained language model, including a number of transformer layers in a stacked setting, with unsupervised data and supervised data as input, pre-trained by the first text editor, and output as a first vector.

[0054] Unsupervised and supervised data include different languages, such as Figure 2 As shown, the unsupervised data includes French unsupervised data and English unsupervised data, and the supervised data includes English supervised data, etc.

[0055] The first text classifier includes a first multilayer perceptron that receives the first vector output by the first text encoder and fine-tunes a downstream classification task.

[0056] The first multilayer perceptron is a two-layer perceptron.

[0057] The first feature projection includes a single-layer perceptron and a nonlinear activation function, which can align the feature dimensions of the excitation network and the target network.

[0058] The pre-trained language model in the first text encoder may be a BERT model, including a low-layer BERT, a middle-layer BERT, and a high-layer BERT;

[0059] in,

[0060] The low-level BERT captures low-level word-level features;

[0061] The middle-layer BERT captures intermediate grammatical features;

[0062] The high-level BERT captures high-level semantic features.

[0063] The target network also includes three parts, namely, a second text encoder, a second text classifier, and a second feature projection, which are connected in sequence.

[0064] in,

[0065] Because the model parameters of the convolutional neural network are relatively small and can be calculated in parallel, the convolutional neural network TextCNN is used as the second text encoder in the embodiment of the present invention, and the unsupervised data and the supervised data are used as the input of the second text encoder. After being trained by the second text encoder, the output is the second vector;

[0066] A second text classifier includes a second multilayer perceptron and the received second vector;

[0067] The structure of the second feature projection is the same as that of the first feature projection, and also includes a feature mapping layer to replace the output of the transformer layer in the excitation network.

[0068] The convolutional neural network TextCNN includes CNN convolution kernels of different sizes, namely CNN small filters, CNN medium filters and CNN large filters;

[0069] in,

[0070] CNN small filters capture word-level features;

[0071] CNN medium filters capture grammatical level features;

[0072] CNN large filters capture semantic level features.

[0073] CNN small filters are aligned with low-layer BERT; CNN medium filters are aligned with middle-layer BERT; CNN large filters are aligned with high-layer BERT.

[0074] The embodiment of the present invention includes two training stages: the first stage is the pre-training of the excitation network, and the second stage is the training of the target network. In the first stage, the embodiment of the present invention introduces advanced semi-supervised concepts to complete the training of the excitation network. In the second stage, the embodiment of the present invention keeps the parameters of the excitation network unchanged, and guides the training of the target network in downstream tasks through the multi-level regularization constraints provided by the excitation network, and finally realizes efficient semi-supervised distillation learning.

[0075] The loss function of the two-stage training is as follows:

[0076] The training of the incentive network is inspired by the consistent regularization framework, and the loss function consists of two parts: the first cross entropy loss of the labeled data and the first consistent regularization loss of the unsupervised data. Because the gap between the labeled data and the unlabeled data is large, it is easy to cause overfitting on the labeled data, so the training signal annealing technology is also used in the implementation of the present invention to balance the participation of the labeled data in the training process.

[0077] The loss function for training the target network uses two types of distillation losses, specifically the distillation loss of the model output layer and the distillation of the latent space feature layer, and the second consistency regularization loss. Among them, the distillation loss of the model output layer of the target network uses the mean square error loss of the hard labels and soft labels of the output layer. In addition, because the distillation loss of the model output layer does not include the learning process of the intermediate layer, the distillation of the latent space feature layer is introduced.

[0078] The distillation loss of the latent space feature layer uses the first feature projection and the second feature projection to match the hidden state and feature map of the transformer layer, and minimizes the mean square error between the first feature projection and the second feature projection to complete the knowledge distillation.

[0079] Specifically, because the BERT model can capture surface, syntactic, and semantic representations from low to high levels, and different sizes of CNN convolution kernels extract different text features, and the CNN convolution kernel increases with the complexity of language features. For example, the CNN convolution kernel with a window size of 4 mainly focuses on word-level features, while the CNN convolution kernel with a window size of 15 can capture semantic-level features. Therefore, the extraction scheme based on latent space features can achieve knowledge transfer from BERT to TextCNN. The specific operation is as follows: align the CNN small filter with the low-level BERT, align the CNN medium filter with the middle-level BERT, and align the CNN large filter with the high-level BERT. Among them, the CNN small filter needs to capture word-level features, the CNN medium filter needs to capture grammatical features, and the CNN large filter needs to capture semantic features.

[0080] Because of the differences in parameter space and network structure between the target network and the excitation network, there is a problem of knowledge loss in the learning process. If only the knowledge distillation method is used, the target network will not be able to understand certain functional characteristics of the excitation network. Therefore, consistency regularization is introduced in the embodiment of the present invention to constrain the target network so that it remains sufficiently smooth in the function space, while also ensuring a stronger generalization ability of the model. Even if the input data changes slightly or its form changes, but the semantics remain unchanged, the output of the model can also remain basically unchanged.

[0081] Figure 2 A schematic diagram of feature distillation from BERT to TextCNN is shown, with BERT's excitation network on the left and TextCNN's target network on the right.

[0082] The BERT excitation network on the left includes a multi-head attention layer and preprocessing of summation and standardization. It then enters the feedforward neural network for training, sums and standardizes it, and inputs the output of the model into the hidden layer unit, outputs it to the linear layer, and then refines it into specific features through the hyperbolic true activation function.

[0083] The TextCNN target network on the right includes the convolutional layer of TextCNN, which is output to the maximum pooling layer, then input to the feature mapping layer, and then output to the linear layer. It is then refined into specific features through the hyperbolic true activation function, and the features output by the BERT excitation network on the left are distilled at the feature layer to obtain the final feature output.

[0084] In a specific embodiment, the maximum text length can be set to 256, the dropout rate can be set to 0.5, the Adam method can be used to update the parameters, and the training cycle can be set to 10.

[0085] The incentive network uses Google's open source pre-trained BERT-base-uncased model as the encoder. The two-layer perceptron contains 768 hidden units, tanh is used as the activation function, the learning rate of the BERT encoder is 2e-5, and the learning rate of the multi-layer perceptron is 1e-3.

[0086] In the target network, the parameters of the word embedding layer are initialized with a 300-dimensional glove word vector. TextCNN is used as the encoder, and the sizes of the convolution kernels are 2, 3, 5, 7, 9, and 11. The dimension of the output layer is 200 dimensions, and the maximum pooling method is used to extract the main information. The feature mapping layer uses a single-layer perceptron with 256 hidden units and Relu as the activation function.

[0087] The preferred specific embodiments of the present invention are described in detail above. It should be understood that ordinary technicians in the field can make many modifications and changes based on the concept of the present invention without creative work. Therefore, all technical solutions that can be obtained by technicians in the technical field based on the concept of the present invention through logical analysis, reasoning or limited experiments on the basis of the prior art should be within the scope of protection determined by the claims.

Claims

1. A lightweight semi-supervised model system for text classification in a few-sample domain, characterized in that: include: As a pre-trained incentive network; As a target network for training, the target network is connected to the excitation network and performs several feature distillations with the excitation network; in, The excitation network mines information and features from unsupervised data and supervised data using consistency regularization and data enhancement techniques, providing multi-level regularization constraints for subsequent training of the target network; The excitation network includes a first text encoder, a first text classifier and a first feature projection which are connected in sequence; in, The first text encoder is a pre-trained language model, including a plurality of stacked transformer layers, the unsupervised data and the supervised data are used as input, and after pre-training provided by the first text encoder, the output is a first vector; The first text classifier includes a first multi-layer perceptron that receives the first vector and fine-tunes a downstream classification task; The first feature projection includes a single-layer perceptron and a nonlinear activation function, which is used to align the feature dimensions of the excitation network and the target network; The target network includes a second text encoder, a second text classifier, and a second feature projection connected in sequence; in, The second text encoder includes a plurality of convolutional neural networks TextCNN, the unsupervised data and the supervised data are used as input, and after training provided by the second text encoder, the output is a second vector; The second text classifier includes a second multilayer perceptron and the received second vector; The structure of the second feature projection is the same as that of the first feature projection, and further includes a feature mapping layer.

2. The lightweight semi-supervised model system for text classification in a few-sample domain as claimed in claim 1, characterized in that: The first text encoder is a BERT model, and the BERT model includes a low-layer BERT, a middle-layer BERT, and a high-layer BERT; in, The low-level BERT captures low-level word-level features; The middle-layer BERT captures intermediate grammatical level features; The high-level BERT captures high-level semantic-level features.

3. The lightweight semi-supervised model system for text classification in a few-sample domain as claimed in claim 2, characterized in that: The convolutional neural network TextCNN includes CNN convolution kernels of different sizes, namely CNN small filter, CNN medium filter and CNN large filter; in, The CNN small filters capture the word-level features; The CNN medium filter captures the grammatical level features; The CNN large filters capture the semantic level features.

4. The lightweight semi-supervised model system for text classification in a few-sample domain as claimed in claim 3, characterized in that: The CNN small filters are aligned with the low-layer BERT; the CNN medium filters are aligned with the middle-layer BERT; the CNN large filters are aligned with the high-layer BERT.

5. The lightweight semi-supervised model system for text classification in a few-sample domain as claimed in claim 4, characterized in that: The loss function of the excitation network training includes a first cross entropy loss of the supervised data and a first consistent regularization loss of the unsupervised data.

6. The lightweight semi-supervised model system for text classification in a few-sample domain as claimed in claim 5, characterized in that: The loss function of the target network training includes the distillation loss of the model output layer, the distillation loss of the latent space feature layer and the second consistent regularization loss.

7. The lightweight semi-supervised model system for text classification in a few-sample domain as claimed in claim 6, characterized in that: The distillation loss of the latent space feature layer uses the first feature projection and the second feature projection to match the hidden state and feature map of the transformer layer, and minimizes the mean square error between the first feature projection and the second feature projection to complete the knowledge distillation.

8. The lightweight semi-supervised model system for text classification in a few-sample domain as claimed in claim 1, characterized in that: The first multilayer perceptron has two layers.