Context data enhancement method based on label invariance
Through the tag invariance context data enhancement method based on the CBert-BiLSTM model, the problem of sparse data samples of patent value evaluation model under low resources is solved, and the generated enhanced text significantly improves classification accuracy under small sample data.
Patent Information
- Application Number
- CN202411597802.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-11
- Publication Date
- 2025-05-13
AI Technical Summary
In the case of low resources, developing high-performance patent value evaluation models faces the problem of sparse data samples, and existing data augmentation methods are prone to destroying the original text structure, resulting in category induction errors.
A context data augmentation method based on label invariance is proposed. The target text encoding and context feature extraction is performed using the CBert-BiLSTM model, feature transformation is performed through pooling operations, autoencoder and denoising autoencoder, and enhanced text is generated through reverse decoding.
It effectively improves the accuracy of patent value evaluation under small sample data, and the generated enhanced text keeps the label information unchanged, integrates rich context semantic information, significantly improving the effectiveness of text classification tasks.
Smart Images

Figure CN119990068A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of natural language processing, and in particular relates to a context data enhancement method based on label invariance. Background Art
[0002] Patent value assessment is the premise for all patent layout. Due to the lack of patent data with value annotation, it is very challenging to develop a high-performance value assessment model under low resource conditions. How to effectively manage and use high-value patents has become a dilemma faced by various industries. Whether it is the traditional statistical analysis method or the current mainstream deep learning method, or the combination of different methods, a certain number of data sets are required as support so that the model can fully learn the inherent effective laws of the value assessment data and further realize the reasoning and prediction of unknown data. In order to solve the problem of scarce data samples, data enhancement methods are usually used to generate more training samples.
[0003] Data augmentation is an effective strategy to alleviate the overfitting problem of deep learning in the case of scarce training data. It can improve the generalization ability of the model and enable it to predict unknown data more accurately. Data augmentation methods were initially widely used in the field of computer vision and then expanded to the field of natural language processing. Data augmentation methods for text usually operate at the word or sentence level, but most methods will destroy the structure of the original text, resulting in incorrect category induction in classification tasks. Summary of the invention
[0004] In order to improve the accuracy of patent value assessment under small sample data, the present invention proposes a contextual data enhancement scheme based on label invariance.
[0005] The first aspect of the present invention discloses a context data enhancement method based on label invariance, the method comprising:
[0006] Step S1: ID mapping is performed on the input original text, and then the order is randomly disrupted, and a specified number of characters are selected as the target text for subsequent data enhancement;
[0007] Step S2, using the Bert model and the bidirectional LSTM model that change the embedding layer vector, perform text encoding processing and context feature extraction on the target text while retaining the classification label information;
[0008] Step S3: transform and concatenate the extracted feature vectors through pooling operation, autoencoder and denoising autoencoder respectively, and generate enhanced text as output through reverse decoding.
[0009] According to the method of the first aspect of the present invention, in step S1:
[0010] Segment the input raw text and build a vocabulary, converting each character into an id;
[0011] The converted token array is randomly shuffled in order without setting a random seed for the operation, so that the order of each shuffle is different and the characters selected when multiple data enhancements are performed on the same text are not completely uniform;
[0012] In the random token array, select K characters, and K is calculated as follows:
[0013] K = min(length, len(tokens)*p)
[0014] Among them, length represents the maximum number of replacement characters, len(tokens) represents the number of characters in the current text, and p represents the probability of the number of replaced characters;
[0015] The following strategy is adopted for the first K characters of the selected token array:
[0016]
[0017] Among them, token represents the characters in the original text, [MASK] represents the mask mark, synonym represents the synonym of token, random represents the random word, and rp represents the random probability.
[0018] According to the method of the first aspect of the present invention, in step S2:
[0019] The Bert model of the embedding layer vector is changed to the CBert model. The structure of the CBert model is the same as that of the Bert model. The embedding layer vector of the CBert model converts the sentence discrimination vector in the embedding layer of the Bert model into a label vector to make predictions based on the context and label of [MASK];
[0020] Before performing the data augmentation task, CBert is trained and fine-tuned on the domain dataset for the MLM task. The bidirectional LSTM model is a BiLSTM model, and the CBert model is concatenated with the BiLSTM model. The BiLSTM model processes the text sequence bidirectionally at each time step, captures semantic information through a recursive structure, and learns the dependencies in the sequence.
[0021] According to the method of the first aspect of the present invention, in step S3, for the pooling operation: the embedding vectors of invalid tags in the feature vector are set to zero, the embedding vectors of valid tags are summed and averaged, so as to capture the key information in the sentence, eliminate the influence of the sentence length on the vector, and map sentences of different lengths to embedded vector representations of fixed length.
[0022] According to the method of the first aspect of the present invention, in step S3, the autoencoder is a neural network model for unsupervised learning, consisting of an encoder and a decoder, compressing and quantizing the feature vector, converting it into a low-dimensional vector representation, and learning a similar embedding representation of the feature vector; the learning process is:
[0023] S=Sbert(sentence1)
[0024] U=Sbert(sentence2)
[0025]
[0026] Among them, sentence1 is sentence 1, sentence2 is sentence 2, S and U are the embedded representations of sentence 1 and sentence 2 encoded by the SBert model, respectively, and f autoencoder For the autoencoder.
[0027] According to the method of the first aspect of the present invention, in step S3, the denoising autoencoder is an extension of the autoencoder, which introduces data with added noise for training, encodes sentence 1 through the SBert model to obtain an embedded representation S, and introduces Gaussian noise to form noisy data SN as the input of the denoising autoencoder, and trains the mapping from SN to S:
[0028] SN=Noise(S)
[0029]
[0030] Among them, Noi se represents noise, f denoising-autoencoder represents a denoising autoencoder.
[0031] According to the method of the first aspect of the present invention, in step S3, the extracted feature vector V1, the pooling operation, the autoencoder and the denoising autoencoder are transformed to obtain the embedding vector V 11 、V 21 、V 31 Perform concatenation and generate an embedded representation of the fused information through a fully connected neural network:
[0032] T:V→V'
[0033] V′=T(V)=T(concat(V1,V 11 ,V 21 ,V 31 ))
[0034] Among them, V represents concatenation features, concat represents cascade, T represents the fusion process through a fully connected neural network, and V' represents the embedded representation of fusion information generated by a fully connected neural network;
[0035] Decode V', reverse map the id back to the character, and extract the category label value of the label vector layer to obtain enhanced text based on label semantics and context-rich information.
[0036] A second aspect of the present invention discloses a context data enhancement system based on label invariance, characterized in that the system comprises:
[0037] The first processing unit is configured to: perform ID mapping on the input original text, then shuffle the order randomly, and select a specified number of characters as the target text for subsequent data enhancement;
[0038] The second processing unit is configured to: perform text encoding processing and context feature extraction on the target text by using a Bert model and a bidirectional LSTM model that change the embedding layer vector;
[0039] The third processing unit is configured to: transform and concatenate the extracted feature vectors through pooling operation, autoencoder and denoising autoencoder respectively, and generate enhanced text as output through reverse decoding.
[0040] The third aspect of the present invention discloses an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the context data enhancement method based on label invariance described in the first aspect of the present disclosure is implemented.
[0041] The fourth aspect of the present invention discloses a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the context data enhancement method based on label invariance described in the first aspect of the present disclosure is implemented.
[0042] The technical solution proposed in the present invention is based on the Bert model and the BiLSTM model, integrates label information, transforms and combines sentence embeddings, and generates a new sentence embedding function; a large number of experiments are carried out based on multiple data sets to verify the effectiveness of the data enhancement method.
[0043] Based on the CBert-BiLSTM model, the present invention generates vectors with different angle information through pooling operation, autoencoder and denoising autoencoder model, and then splices and linearly maps them to obtain new texts with rich context information while maintaining label invariance. The present invention uses patents and online shopping reviews in the field of artificial intelligence as classification task datasets, which has a positive effect on both small sample datasets and multi-sample datasets, and has significant advantages over current mainstream algorithms. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0045] Figure 1 This is a diagram of the data enhancement model framework;
[0046] Figure 2 Training graph for the autoencoder;
[0047] Figure 3 Training graph for denoising autoencoder;
[0048] Figure 4 Schematic diagram of fine-tuning loss values for patent data;
[0049] Figure 5 Fine-tuning loss values for online shopping review data. DETAILED DESCRIPTION
[0050] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0051] Data augmentation refers to a method of obtaining new data that is semantically invariant to some extent by modifying or combining existing data. The use of data augmentation technology can significantly improve the performance of the model without increasing the cost of collecting and annotating data. Unlike data augmentation in the visual field, natural language text is discrete, which makes text data augmentation more challenging. Currently, representative data augmentation methods can be divided into two categories: rule-based and model-based.
[0052] Research status of data enhancement methods based on rule matching:
[0053] The rule-based data augmentation method is a method of obtaining new data by transforming text by defining rules. These rules can be used to modify, delete, replace or insert text based on the grammar, semantics or structure of the sentence to generate new text. At the same time, rule-based data augmentation methods are usually easy to implement.
[0054] Inspired by the cropping and rotation of images in the field of computer vision to achieve data enhancement, for example, it is proposed to use dependency trees for data enhancement, and sentences are cropped and rotated by moving tree fragments. This method has achieved great success in part-of-speech tagging tasks. For another example, the EDA method is proposed. EDA consists of four components, including synonym replacement, random insertion, random exchange, and random deletion. This method is simple to implement and shows good performance in many text classification tasks. For another example, the UDA method is proposed. It uses supervised data constraints to process unsupervised data through consistency training. This method has substantial improvements in both visual and text tasks. The above rule-based methods are simple and intuitive data enhancement methods, and they have achieved significant advantages in certain tasks. However, this type of method has some potential disadvantages. It depends on the quality and accuracy of the rules. Natural language has complex and diverse semantics and structures. If the rules are not designed well or accurately, it will introduce incorrect text transformations, affecting the quality of the enhanced data. In addition, rule-based data enhancement methods often replace or transform based on local rules, which cannot capture context and contextual information, and are prone to destroying semantics, affecting the performance improvement of downstream tasks that rely on semantics, such as text classification.
[0055] Research status of model-based data augmentation methods:
[0056] With the continuous development of deep learning technology, text data augmentation has also begun to use the inherent ability of the model to generate new texts that are similar but different, so as to expand the data set while retaining the original semantics.
[0057] Thanks to the great success of machine translation, with the open source of various advanced translation models, for example, a data enhancement method based on back translation is proposed, that is, using a machine translation model to translate the source language into other languages, which can be converted multiple times or only once, and finally translated back to the source language. The text is the new text obtained after data enhancement. Back translation technology has a positive effect on many natural language processing downstream tasks. With the introduction of pre-trained language models, data enhancement methods began to randomly replace some words in sentences by predicting more diverse replacement words through pre-trained language models. For another example, a method based on improved conditional Bert is proposed, which has achieved significant improvements in 6 different text classification tasks. For another example, a data enhancement model based on LSTM is proposed, which expands more data samples for the field of traditional Chinese medicine and proves the similarity between the generated text and the source data; a language-independent data enhancement LiDA method is proposed. This technology uses a multilingual model to create synthetic data from available training data sets and only works on the sentence embedding layer. Since the pre-trained language model is pre-trained on large-scale data and has strong language understanding and generation capabilities, the model-based data augmentation method can better ensure the semantic consistency between the generated samples and the original samples, and ensure a certain emotional tendency, which is crucial for text classification tasks.
[0058] Based on the above research, the present invention proposes a contextual data enhancement scheme based on label invariance. On the basis of the conditional Bert model with an added label embedding layer, the BiLSTM model is combined, and the embedded vectors are respectively subjected to pooling operations, autoencoders, and denoising autoencoder models, and then the three generated vectors are concatenated and reversely mapped, finally obtaining a new text that maintains label invariance and integrates contextual semantic information after data enhancement.
[0059] Contextual data augmentation method model based on label invariance:
[0060] Based on conditional Bert and BiLSTM, data enhancement is performed through three embedding vector transformations, such as Figure 1 As shown, the model mainly consists of three parts.
[0061] The first part is the data enhancement character screening part. First, the input text is id mapped, then randomly shuffled, and a specified number of characters are selected as the target for subsequent data enhancement. The second part is the basic model part, including the Bert model and the bidirectional LSTM model that change the embedding layer vector. The text data after [MASK] is used as input, and the text encoding and context feature extraction are performed to retain the classification label information. The third part is the embedding vector transformation part, including pooling operation, autoencoder and denoising autoencoder model transformation. The vectors generated in the second part are respectively transformed through these three transformations to obtain 3 new vectors, and are concatenated with the original vectors, and then the dimension is converted through the fully connected layer. Finally, the vector is reversely decoded to obtain the corresponding text to generate enhanced text data.
[0062] Enhanced character filtering:
[0063] The deep bidirectional language model is currently the best basic model for understanding and generating natural language. Among them, the self-attention-based Bert model has been widely used in natural language processing tasks and has achieved remarkable results. It combines the ideas of deep learning and bidirectional modeling and proposes the MLM (Masked Language Modeling) task. The specific steps of the MLM task are as follows: first, the input text is preprocessed by word segmentation, word list construction, etc., and then a certain proportion of words are randomly selected from the text for masking, that is, replaced with [MASK] tags, and the processed sentences are input into the Bert model for encoding. Finally, the masked words in the original text are predicted through its hidden representation. Through the MLM task, the Bert model can learn contextual information that contains rich knowledge.
[0064] The present invention draws on the principle of the above-mentioned MLM task, imitates the task to perform data enhancement, and fits the pre-training task of the Bert model to the greatest extent. At the same time, based on its ability to understand contextual relevance, it provides new samples that retain basic semantics and describe diverse words. The process of selecting masked characters is as follows Figure 1 As shown in the token selection section, the specific steps and selection strategies are as follows:
[0065] Perform word segmentation and vocabulary building on the input raw text, and convert each character into an ID;
[0066] The converted token array is randomly shuffled in order without setting a random seed for the operation, so that the order of each shuffle is different, ensuring that the characters selected when multiple data enhancements are performed on the same text sentence are not completely uniform;
[0067] In the random token array, select K characters. K is calculated as follows:
[0068] K=min(length,len(tokens)*p) (1)
[0069] Among them, length represents the maximum number of replacement characters set in the experiment, len(tokens) represents the number of characters in the current text, and p represents the probability of the number of characters to be replaced set in the experiment;
[0070] For the first K characters of the selected token array, the following strategy is adopted according to the settings in the MLM task:
[0071]
[0072] Among them, token represents the characters in the original text, [MASK] represents the mask mark, synonym represents the synonym of token, random represents the random word, and rp represents the random probability. Specifically, the selected characters have a 70% probability of being masked, a 10% probability of keeping the original characters, a 10% probability of being replaced with synonymous characters, and a 10% probability of being replaced with random characters. The purpose of this operation is to intentionally introduce some errors or noise to simulate the errors and noise in real data in order to improve the robustness of the model.
[0073] CBert-BiLSTM model:
[0074] The model structures of CBert and Bert are the same, except that the embedding layer vector of CBert converts the sentence distinguishing vector in the Bert embedding layer into a label vector. In the Bert model, the sentence distinguishing vector is used to mark the difference between two sentences when processing sentence pair tasks. However, in the classification task, there is usually only one sentence as input, which makes the vector of this layer lose its meaning to some extent. Therefore, the vector can be replaced with the category label of the classification result. In this way, the classification result can be integrated into the MLM task as a prompt information in the subsequent task to avoid the predicted words losing their original emotional tendency. For example, the input sentence is "The cutout algorithm proposed by the present invention has high accuracy and fast speed." When the character "high" is subjected to the [MASK] operation, it is very likely to be predicted as "bad", "bad" and other words with completely opposite emotions through the MLM task, which is obviously unfavorable. When the label is also embedded into the input vector as a kind of information, its goal is to predict it based on the context of [MASK] and the category to which the sentence belongs. The context and label content are considered at the same time, which makes the generated new text similar to the original text in context and compatible with the label, which can play a positive role in the expansion of the data set.
[0075] At the same time, Bert's pre-training is based on a large-scale general text corpus. In order to enable the MLM task to better predict and generate data in the domain, we need to fine-tune Bert on the domain dataset before performing the data augmentation task.
[0076] In addition, the Bert model is based on self-attention, in which the sinusoidal position encoding used weakens the sequence distance information and direction information. In order to further enhance the model's ability in sequence modeling, the present invention splices the CBert model with the BiLSTM model. The BiLSTM model is a recurrent neural network that can process text sequences bidirectionally at each time step. It can capture more semantic details through its recursive structure and better learn the long-term dependencies in the sequence. Combining the advantages of the two provides a richer and more accurate semantic representation.
[0077] Embedding vector transformation:
[0078] The SBert model is a learning method for sentence embedding vectorization based on the Bert model. Its advantage is that it is significantly better than other models in constructing semantically similar sentence embeddings. Therefore, the SBert model can be used to generate sentence embeddings that are similar to the original semantics. Based on the SBert model, the following three vector transformations are performed on the sentence embedding vector to generate new sentence embeddings that are highly similar to the original sentence embeddings, and then these embeddings are concatenated and mapped to create synthetic sentence embeddings.
[0079] Pooling operation:
[0080] Pooling is an operation used for feature extraction and dimensionality reduction. It extracts the main features by aggregating data in a fixed area. Pooling is robust to a certain extent to small shifts in the input vector, that is, the feature representation after pooling remains unchanged for small translations of the input vector, which makes the model invariant to changes in the input position. The input vector of the model includes valid text word embeddings and invalid padding token embeddings that have no practical meaning. Average pooling only calculates the average value of all valid tokens for the input embedding vector, which can capture the semantic information of the entire sentence.
[0081] Therefore, the present invention performs average pooling operation on sentence embedding. The embedding vectors of invalid tags are set to zero, and the valid tags are summed and averaged, which effectively captures the key information in the sentence and eliminates the influence of sentence length on the vector. Sentences of different lengths are mapped to fixed-length embedding vector representations to facilitate further processing of subsequent models.
[0082] Autoencoder Model:
[0083] An autoencoder is a neural network model for unsupervised learning, consisting of an encoder and a decoder, which compresses and quantizes the input data into a low-dimensional vector representation for feature extraction tasks. The purpose of the autoencoder model is to learn similar embedding representations of input vectors, so the model needs to learn from a variety of sentence pairs with similar semantics. The present invention uses the Peking University Chinese text retelling dataset PKU-Paraphrase-Bank as an autoencoder training dataset, which contains more than 500,000 pairs of semantically similar sentences. Table 1 shows some examples of the dataset.
[0084] Table 1 Examples of Peking University Chinese text retelling dataset
[0085]
[0086] In the experiment, we first encode sentences 1 and 2 through the SBert model to obtain embedding representations S and U, and then train the autoencoder to map from the input sentence embedding U to the output embedding S. The learning process of the autoencoder is as follows:
[0087] S=Sbert(sentence1)(3)
[0088] U=Sbert(sentence2)(4)
[0089]
[0090] Among them, f autoencoder Represents the autoencoder. The implementation process is as follows Figure 2 shown.
[0091] Denoising Autoencoder Model:
[0092] Denoising autoencoder is an extended form of autoencoder. The difference is that it introduces noisy input data to train the model, so that the denoising autoencoder can learn more robust and generalized feature representations. Use the same dataset as the autoencoder, but only use sentence 1. Encode sentence 1 through the SBert model to get the embedding S, and introduce Gaussian noise to form noisy data SN as the input of the denoising autoencoder, and train its mapping to the original embedding S. The denoising autoencoder learning process is as follows:
[0093] SN=Noise(S)(6)
[0094]
[0095] Among them, f denoising-autoencoder represents the denoising autoencoder. The implementation process is as follows Figure 3 shown.
[0096] After all the above steps, after inputting the sentence, we can get the embedding vector V1 after the CBert-BiLSTM model, and the embedding vectors V11, V21, and V31 after three transformations, a total of 4 embedded coding representations containing information from different angles. After connecting these 4 vectors, a 768-dimensional embedded representation of fused information is generated through a fully connected neural network.
[0097] T:V→V′
[0098] V'=T(V)=T(concat(V1,V 11 ,V 21 ,V 31 )) (8)
[0099] Finally, perform a decoding operation on V', reversely map the id back to the character, and take out the category label value of the label vector layer to obtain a new sample sentence after data enhancement based on label semantics and rich context information.
[0100] Experimental data:
[0101] In order to verify the effectiveness of the proposed data enhancement method, the experiment was evaluated on the online shopping review dataset in addition to the patent value assessment task. The patent value data comes from 2,000 patents in the field of artificial intelligence from multiple platforms such as the China National Intellectual Property Administration database and Zhihuiya, including the application of artificial intelligence technology in various directions. The value level is divided into 5 levels from 0 to 4, with 0 representing the lowest value and 4 representing the highest value. The online shopping review dataset is publicly available on the Internet, including more than 60,000 evaluations of 10 items, of which the emotional evaluation is divided into two categories, 0 and 1, with 0 representing negative evaluation and 1 representing positive evaluation. The summary statistics of the two datasets are shown in Table 2.
[0102] Table 2 Dataset summary statistics
[0103] Dataset C L |V| Num Test Patent value 5 264 21128 2000 800 Online shopping reviews 2 59 21128 62774 12556
[0104] Among them, C represents the number of category labels, L represents the average sentence length of the data set text, |V| represents the vocabulary size of the data enhancement base dictionary, Num represents the data set size, and Test represents the test set size.
[0105] In the experiment, the training data set for the patent value rating task is set to 1,200 data items and the test set is set to 800 data items. The training data set for the online shopping review sentiment analysis task is set to 12,556 data items and the test set is set to 50,218 data items.
[0106] Experimental setup:
[0107] For the data enhancement experiment, the basic models used in the present invention include the Chinese basic BERT model, including a 12-layer Transformer structure, a hidden layer dimension of 768 dimensions, and 12 self-attention heads; a bidirectional LSTM model, whose hidden layer dimension is 768 dimensions; an autoencoder and a denoising autoencoder model, including a decoder and an encoder, both of which are composed of a fully connected feedforward neural network, and the input and output dimensions are both 768 dimensions. The experiment first fine-tunes the overall model based on the original data set, and then uses the well-trained model to expand the fusion label and context semantics of each text sentence in the data set. The relevant parameter settings for the data enhancement experiment part of the experiment are shown in Table 3. The maximum text length (length) is 512, the batch size (batchsize) is 8, the number of learning rounds (epoch) is 20, the initial learning rate (learning rate) is 5e-5, the learning rate decay rate (weight decay) is 0.01, the number of learning rate scheduling warmup rounds is 4, and the feature vector dimension (dim) is 768.
[0108] After data enhancement of the original data set, the present invention sets a classification task for two data sets to compare the classification accuracy effect before and after data enhancement. For the patent value assessment task, the experiment adopts the Chinese patent value assessment classification method of multi-feature fusion in the previous study. This method integrates the three dimensions of patent technical value, legal value and economic value, and predicts the value level based on the XGBoost model. For the online shopping review sentiment analysis task, the experiment adopts a basic single-layer LSTM model, and maps the output to the category label through the softmax function. The relevant parameter settings of the patent value classification experiment are shown in Table 4. The maximum text length (length) is 512, the batch size (batchsize) is 4, the number of learning rounds (epoch) is 50, the initial learning rate (learning rate) is 5e-4, the learning rate decay rate (weight decay) is 0.01, the number of learning rate scheduling warmup rounds is 4, the feature vector dimension (dim) is 768, the number of early stop rounds (early stop) is 10, and the random inactivation probability (dropout) is 0.2. The relevant parameter settings for the online shopping review classification experiment are shown in Table 5. The maximum text length (length) is 512, the batch size (batchsize) is 8, the number of learning rounds (epoch) is 20, the initial learning rate (learning rate) is 1e-5, the learning rate decay rate (weight decay) is 0.01, the number of learning rate scheduling warm-up rounds is 4, the feature vector dimension (dim) is 768, and the random inactivation probability (dropout) is 0.2.
[0109] Table 3 Data augmentation experiment parameter settings
[0110]
[0111] Table 4. Parameter settings for patent value classification experiment
[0112]
[0113]
[0114] Table 5. Experimental parameter settings for online shopping review classification
[0115]
[0116] Evaluation indicators:
[0117] The experiment uses commonly used evaluation indicators for text classification, including accuracy (Acc, accuracy), precision (P, precision), recall (R, recall), and F1 score (F1, f1-score) to evaluate and compare the experimental results. The calculation formulas for each indicator are as follows:
[0118]
[0119]
[0120] Among them, TP represents the number of samples that are actually positive and predicted as positive, FP represents the number of samples that are actually negative but predicted as positive, TN represents the number of samples that are actually negative and predicted as negative, and FN represents the number of samples that are actually positive but predicted as negative.
[0121] Experimental results and analysis:
[0122] In order to verify the effectiveness of the data enhancement method proposed in the present invention for classification tasks, the present invention conducted comparative experiments and ablation experiments, mainly including other text data enhancement methods and methods of combining different models.
[0123] In order to better distinguish the mentioned methods and enhance the readability of data comparison, the present invention names all the methods for conducting the experiments as follows.
[0124] NoAugment: No data augmentation is performed, and the classification task is only implemented on the original dataset. The results of this experiment are used as the baseline comparison for the data augmentation experiment.
[0125] Model-1: EDA data augmentation method based on synonym replacement, random insertion, deletion and exchange.
[0126] Model-2: Data augmentation method based on back translation. In the experiment, an open source translation software interface was used to translate the data text into five languages: English, Japanese, French, Korean, and Russian, and then translated back into Chinese.
[0127] Model-3: X is based on the improved conditional Bert model. This method has achieved significant improvements in text classification tasks.
[0128] Model-4: Language-independent data augmentation LiDA method. This technique uses a multilingual model to create synthetic data from available training datasets and works only on the sentence embedding layer.
[0129] Model-5: The contextual data enhancement method based on label invariance proposed in this paper is based on the CBert-BiLSTM model, which integrates the embedding vector transformation and combination to generate new sample sentences.
[0130] The experiment first fine-tuned the overall model based on the original data set, and then conducted data enhancement experiments and subsequent comparative experiments and ablation experiments. By fine-tuning parameters by imitating the MLM task on different text corpora, the data in the field can be better predicted and generated. The loss value changes during the fine-tuning process are as follows: Figure 4 , Figure 5 shown.
[0131] All the above models are subjected to data augmentation on two different datasets, and the final classification is performed based on the classification model described in the experimental setting. The comparative experimental results are shown in Table 6.
[0132] Table 6 Comparative experimental results
[0133]
[0134] Through the experimental results of each method in Table 6 on two different data sets, it can be concluded that the data enhancement method proposed in the present invention has the best overall effect on the text classification task. Common data enhancement methods include sample expansion and vector representation enhancement. Model-1 and Model-2 both belong to the sample expansion type. Compared with the classification results of the original data set, the classification results after data enhancement by these two models have increased the F1 value by 8.97% and 9.68% respectively in the patent value assessment task; in the online shopping review sentiment analysis task, the F1 value of the Model-1 method decreased by 0.36%, and the F1 value of the Model-2 method increased by 0.19%. From the experimental results, it can be seen that in the small sample data set, Model-1 and Model-2 both play a certain positive role. This is because the data of the small sample data itself is too scarce, so that the classification model cannot effectively learn the data features and patterns. After data enhancement, the diversity of the data will be increased from a certain angle, resulting in improved classification effect. In the case of sufficient sample quantity, the data is relatively sufficient and the diversity is high, which can enable the classification model to learn certain internal patterns and laws. Therefore, simple data enhancement methods such as EDA and back translation methods cannot bring obvious performance improvements. In fact, these two methods are too strong in the transformation of sentence structure and semantics, which is out of the inherent rules of the original text, resulting in too much noise and causing the classification accuracy to decrease instead of increase. After analysis, Model-1 and Model-2 generate new data samples based on the original data, but both change the semantics and sentence structure of the original text, which easily introduces noise and leads to uncertain classification. Semantic consistency is very important in text classification tasks because the model needs to pay attention to the subtle differences in the text to make the final category delineation. Model-3 is a conditional Bert model based on the context-based change of the embedding layer representation. Compared with the classification results of the original data set, the classification results after data enhancement by this model have an F1 value increase of 21.52% in the patent value assessment task and an F1 value increase of 1.36% in the online shopping review sentiment analysis task. Model-4 is a LiDA method that only transforms sentence embeddings without generating new text sentences for data augmentation. Compared with the classification results of the original data set, the classification results after data augmentation by this method increased the F1 value by 20.21% in the patent value assessment task and by 1.96% in the online shopping review sentiment analysis task. Model-3 increased the amount of training data by context-based sample augmentation, and Model-4 increased the diversity of data and the generalization ability of the model by changing the vector representation of the input data.The method proposed in the present invention, namely Model-5, combines the characteristics of the above-mentioned sample expansion and vector representation enhancement to perform data enhancement. Compared with the classification results of the original data set, the classification results after data enhancement by this method have an F1 value improvement of 30.97% in the patent value assessment task and an F1 value improvement of 2.07% in the online shopping review sentiment analysis task. The reason is that Model-5 combines the respective advantages of Model-3 and Model-4, splices the BiLSTM model on the basis of the CBert model to enhance the contextual semantic understanding ability, and makes up for the defects of the Bert model in sequence distance information and direction information. At the same time, vector transformation is added to the basic embedding representation. Through different types of linear transformations and encoders, a certain degree of data noise is introduced, which not only provides richer feature information, but also enhances the diversity and robustness of data text. The new samples obtained can play a positive role in the classification task.
[0135] Analysis of ablation experiment results:
[0136] The ablation experiment combines and ablates the various parts of the model proposed in the present invention, applies the model to perform data enhancement experiments on two different data sets, and performs final classification based on the classification model described in the experimental setting. The method of the present invention is based on CBert, i.e., the Model-3 method, so the ablation experiment uses the Model-3 method as the baseline model. The ablation experiment results are shown in Table 7.
[0137] Table 7 Ablation experiment results
[0138]
[0139] Through the analysis of the results of the ablation experiment in Table 7, it can be seen that each part included in the method proposed in the present invention has a positive impact on data enhancement. For the data enhancement method with the addition of the BiLSTM model, compared with the classification results after the Model-3 data enhancement, the F1 value on the patent value assessment task increased by 7.09%, and the F1 value on the online shopping review sentiment analysis task increased by 0.40%. For the data enhancement method with the addition of vector transformation, compared with the classification results after the Model-3 data enhancement, the F1 value on the patent value assessment task increased by 5.52%, and the F1 value on the online shopping review sentiment analysis task increased by 0.54%. Analyzing its internal reasons, it can be seen that the Bert model is based on self-attention, and the sinusoidal position encoding used in the attention mechanism weakens the sequence distance information and direction information. Sequence distance information refers to the distance relationship between different words or characters in the text. Direction information refers to the arrangement direction of words or characters in the text, such as sequential or reverse order. In Chinese text, both sequence distance information and direction information are of certain importance, and they can provide additional semantic information and context information. By splicing the BiLSTM model, the model's ability to model the text sequence and capture directional information can be enhanced, and new samples with different character orders and distances, as well as reversed or rearranged text samples, can be generated. However, relying solely on the modeling ability of deep neural networks will result in a relatively simple form of predicted enhanced characters. Vector transformation can introduce a variety of vector representations, allowing the model to better adapt to different semantic changes and contexts, and by enhancing the noise and variation of the data, the model's anti-interference ability can be improved, which is important for dealing with the ambiguity and complex contexts in Chinese text. Therefore, the classification experiment was carried out using the method proposed by the present invention that combines the above-mentioned model with vector representation enhancement. Compared with the classification results after Model-3 data enhancement, the F1 value in the patent value assessment task increased by 9.45%, and the F1 value in the online shopping review sentiment analysis task increased by 0.71%, both of which were improved to the greatest extent.
[0140] The experimental results above show that the contextual data enhancement method based on label invariance proposed in this paper has significant advantages in improving the effect of text classification tasks, and solves the problem that a small sample data set may not fully represent the entire data distribution, resulting in insufficient model generalization ability. At the same time, it still has a certain positive effect when the sample data volume is sufficient.
[0141] In order to solve the problem that the patent value level cannot be effectively evaluated when the data samples are scarce, the present invention proposes a contextual data enhancement method based on label invariance. The method is based on the CBert-BiLSTM model, and the embedded vector is subjected to pooling operation, autoencoder and denoising autoencoder model, and then all the generated vectors with different angle encodings are spliced and reverse mapped to obtain the diversified new text data that maintains label invariance and integrates context after data enhancement. It provides a powerful data expansion support means for subsequent tasks such as patent value assessment. At the same time, it provides great convenience for the development of high-performance text classification models under small sample data sets. In subsequent work, more forms of vector embedding representation conversion combinations can be tried in order to obtain better results.
[0142] Please note that the technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, all possible combinations of the technical features in the above embodiments are not described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification. The above-mentioned embodiments only express several implementation methods of the present application, and their descriptions are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that for ordinary technicians in this field, without departing from the concept of the present application, several variations and improvements can be made, which all belong to the scope of protection of the present application. Therefore, the scope of protection of the patent in this application shall be based on the attached claims.
Claims
1. A context data enhancement method based on label invariance, characterized in that: The method comprises: Step S1: ID mapping is performed on the input original text, and then the order is randomly disrupted, and a specified number of characters are selected as the target text for subsequent data enhancement; Step S2, using the Bert model and the bidirectional LSTM model that change the embedding layer vector, perform text encoding processing and context feature extraction on the target text while retaining the classification label information; Step S3: transform and concatenate the extracted feature vectors through pooling operation, autoencoder and denoising autoencoder respectively, and generate enhanced text as output through reverse decoding.
2. The context data enhancement method based on label invariance according to claim 1, characterized in that: In step S1: Segment the input raw text and build a vocabulary, converting each character into an id; The converted token array is randomly shuffled in order without setting a random seed for the operation, so that the order of each shuffle is different and the characters selected when multiple data enhancements are performed on the same text are not completely uniform; In the random token array, select K characters, and K is calculated as follows: K = min(length, len(tokens)*p) Among them, length represents the maximum number of replacement characters, len(tokens) represents the number of characters in the current text, and p represents the probability of the number of replaced characters; The following strategy is adopted for the first K characters of the selected token array: Among them, token represents the characters in the original text, [MASK] represents the mask mark, synonym represents the synonym of token, random represents the random word, and rp represents the random probability.
3. The context data enhancement method based on label invariance according to claim 2 is characterized in that: In step S2: The Bert model of the embedding layer vector is changed to the CBert model. The structure of the CBert model is the same as that of the Bert model. The embedding layer vector of the CBert model converts the sentence discrimination vector in the embedding layer of the Bert model into a label vector to make predictions based on the context and label of [MASK]; Before performing the data augmentation task, CBert is trained and fine-tuned on the domain dataset for the MLM task. The bidirectional LSTM model is a BiLSTM model, and the CBert model is concatenated with the BiLSTM model. The BiLSTM model processes the text sequence bidirectionally at each time step, captures semantic information through a recursive structure, and learns the dependencies in the sequence.
4. The context data enhancement method based on label invariance according to claim 3 is characterized in that: In step S3, for the pooling operation: the embedding vectors of invalid tags in the feature vector are set to zero, the embedding vectors of valid tags are summed and averaged, so as to capture the key information in the sentence and eliminate the influence of sentence length on the vector, and map sentences of different lengths to fixed-length embedding vector representations.
5. The context data enhancement method based on label invariance according to claim 4 is characterized in that: In step S3, the autoencoder is an unsupervised learning neural network model, which consists of an encoder and a decoder, compresses and quantizes the feature vector, converts it into a low-dimensional vector representation, and learns a similar embedding representation of the feature vector; The learning process is: S=Sbert(sentence1) U=Sbert(sentence2) Among them, sentence1 is sentence 1, sentence2 is sentence 2, S and U are the embedded representations of sentence 1 and sentence 2 encoded by the SBert model, respectively, and f autoencoder For the autoencoder.
6. The context data enhancement method based on label invariance according to claim 5, characterized in that: In step S3, the denoising autoencoder is an extension of the autoencoder, which introduces noisy data for training. Sentence 1 is encoded by the SBert model to obtain the embedded representation S, and Gaussian noise is introduced to form noisy data SN as the input of the denoising autoencoder, and the mapping from SN to S is trained: SN=Noise(S) Among them, Noise represents noise, f denoising-autoencoder represents a denoising autoencoder.
7. The context data enhancement method based on label invariance according to claim 6, characterized in that: In step S3, the extracted feature vector V1, the pooling operation, the autoencoder and the denoising autoencoder are transformed to obtain the embedding vector V 11 、V 21 、V 31 Perform concatenation and generate an embedded representation of the fused information through a fully connected neural network: T:V→V' V′=T(V)=T(concat(V1,V 11 ,V 21 ,V 31 )) Among them, V represents concatenation features, concat represents cascade, T represents the fusion process through a fully connected neural network, and V' represents the embedded representation of fusion information generated by a fully connected neural network; Decode V', reverse map the id back to the character, and extract the category label value of the label vector layer to obtain enhanced text based on label semantics and context-rich information.
8. A contextual data augmentation system based on label invariance, characterized in that: The system comprises: The first processing unit is configured to: perform ID mapping on the input original text, then shuffle the order randomly, and select a specified number of characters as the target text for subsequent data enhancement; The second processing unit is configured to: perform text encoding processing and context feature extraction on the target text by using a Bert model and a bidirectional LSTM model that change the embedding layer vector; The third processing unit is configured to: transform and concatenate the extracted feature vectors through pooling operation, autoencoder and denoising autoencoder respectively, and generate enhanced text as output through reverse decoding.
9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the context data enhancement method based on label invariance as described in any one of claims 1-4.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the context data enhancement method based on label invariance described in any one of claims 1 to 4 is implemented.