A service text classification method based on domain BERT model
By introducing domain vocabulary and BiLSTM models into the BERT model and adopting the zoom loss function, the problems of limited accuracy and unbalanced data in text classification are solved, and higher classification accuracy and adaptability are achieved.
Patent Information
- Application Number
- CN202210890080.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-27
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-07-27
AI Technical Summary
There are two problems with the current text classification method based on the BERT model: one is the lack of integration with vocabulary in related professional fields, resulting in limited classification accuracy; the other is the failure to effectively deal with the imbalance of the service text corpus.
A service text classification method based on the domain BERT model is proposed. Domain vocabulary is extracted through the TF-IDF algorithm and the vocabulary list of the BERT model is expanded, and the BiLSTM model is trained, and the zoom loss function is used to equalize the data set.
It improves the accuracy and domain adaptability of service text classification, effectively deals with the imbalance of data sets, and improves the flexibility and classification performance of the model.
Smart Images

Figure CN115344695B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network service text, and specifically relates to a service text classification method based on a domain BERT model. Background Art
[0002] With the continuous development of SOA architecture, the number of services in the network platform has exploded, and the difficulty of service management has also increased. Service classification is an important method of service management. By classifying and managing services, service retrieval and service discovery can be quickly realized according to service categories, thereby ensuring the efficient use of services in the service-oriented architecture (SOA). Among them, service text classification is one of the implementation methods of service classification. Since the service description text has rich semantic information and is easy to edit and extract, service text classification has become a research hotspot in the field of service classification.
[0003] At present, service text classification methods can be divided into feature engineering-based methods and deep learning-based methods. Feature engineering-based methods often generate text feature vectors through artificial feature engineering, and implement text classification through classification algorithms. However, the classification effect based on feature engineering directly depends on the effect of feature engineering, and requires high experience in artificial feature engineering processing, and it is difficult to quickly reach the optimal classification level. Deep learning-based methods can automatically extract features from data without the need for manual feature extraction and selection, greatly reducing the difficulty of operation and better meeting the current research needs in the field of text classification. Among them, the introduction of the BERT (Bidirectional Encoder Representations from Transformers) model has brought a milestone change to the field of text classification. Its multi-head attention mechanism and convenient general framework have significantly improved the model's text classification ability compared to traditional neural networks. At present, the use of BERT pre-trained models for text classification is gradually becoming a research hotspot.
[0004] However, there are two problems with the current text classification method based on the BERT model:
[0005] (1) Text classification algorithms based on the BERT model often generate word vectors based on the model’s own vocabulary. Although the model’s own vocabulary is highly compatible with the model, the lack of integration with related professional vocabulary means that the BERT model has no room to further improve classification accuracy.
[0006] (2) Insufficient consideration is given to the imbalance of service text corpora. Classification datasets often experience data imbalance, and traditional data balancing methods often balance data based on methods such as sampling method adjustment and category weight adjustment. There is no comprehensive consideration of the number of categories and the difficulty of sample classification. Summary of the invention
[0007] In view of the above-mentioned problems, the present invention proposes a service text classification method based on the domain BERT model. The method is a new service text classification framework. The framework mainly includes three steps. First, the proprietary vocabulary in the Web service field is extracted and added to the original vocabulary of the BERT model to expand the coverage of the vocabulary in the Web service field. Secondly, the service text corpus is input into the BERT model with the expanded vocabulary for training. Finally, according to the characteristics and classification results of the service data set, the optimal loss function is selected to balance the data set. In order to verify the performance of the model in a real environment, the present invention uses a crawler algorithm to obtain 437 categories and a total of 22,205 service description texts in the network as the experimental data set of this article. During the experiment, this article compares the proposed method with multiple existing deep learning methods. The results show that the model proposed in this article has better accuracy.
[0008] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0009] A service text classification method based on a domain BERT model, comprising:
[0010] Step 1: Use the TF-IDF algorithm to extract domain vocabulary from the service text corpus;
[0011] Step 2: Based on step 1, a BERT-BiLSTM model is established. After the domain vocabulary extracted in step 1 is input into the BERT vocabulary of the BERT-BiLSTM model, the service text corpus is input into the BERT-BiLSTM model for training to achieve service text classification.
[0012] Step 3: According to the service text corpus characteristics and classification results in step 2, select the best loss function to balance the dataset.
[0013] Preferably, the step 1 comprises:
[0014] Step 1.1: Use the TF-IDF algorithm to extract vocabulary from the service text corpus. The calculation formula is as follows:
[0015]
[0016] Among them, TF represents word frequency, IDF represents the inverse text frequency index of word i in article j, and TF i,jrepresents the frequency of word i appearing in article j, N represents the total number of articles in the corpus, and n i represents the number of articles in the corpus containing word i;
[0017] Step 1.2: Use the TF-IDF algorithm in step 1.1 to calculate the weights of all words and sort them to finally obtain the domain words extracted from the service text corpus.
[0018] Preferably, the step 2 is specifically:
[0019] Step 2.1: Establish the BERT-BiLSTM model structure. The BERT-BiLSTM model structure is composed of the BERT model and the BiLSTM model. After the BERT model operation is completed, it enters the BiLSTM model. The BERT model includes an embedding layer, multiple encoder layers, and a pooler layer in sequence. The BiLSTM layer includes multiple LSTM layers. The BERT model generates text word vectors through the embedding layer, captures text vocabulary features through the multi-head attention mechanism and feedforward neural network layer in the encoder layer, and finally enters the BiLSTM layer through the fully connected layer in the pooler layer. The BiLSTM layer is responsible for obtaining the contextual features between word vectors, and finally classification is performed through the fully connected layer;
[0020] Step 2.2: Input a text sentence into the BERT model structure of step 2.1, and convert the words in the sentence into embeddings. The input of BERT consists of character embeddings, segment embeddings, and position embeddings. Character embeddings represent the embeddings of text words, which perform vocabulary matching according to the set domain vocabulary according to the greedy principle. Segment embeddings indicate the sentence to which the word belongs, and position embeddings identify the specific position of the word in the input text.
[0021] Preferably, in step 2.1, the encoder layer in the BERT model is composed of a transformer encoder structure, which sequentially includes an input embedding layer, a multi-head self-attention mechanism, a residual connection / layer normalization, a feedforward neural network, and a residual connection / layer normalization.
[0022] Preferably, the multi-head self-attention mechanism layer mainly emphasizes different parts of the text by updating weights, and the multi-head self-attention mechanism is calculated as follows:
[0023] MultiHead(Q,K,V)=Concat(head 1 ,…head h )W
[0024]
[0025] Among them, Concat is a concatenation operation, that is, each self-attention result is connected end to end. The matrices Q, K, and V represent the Query, Key, and Value matrices respectively. i Q , W i K , W i V , W are matrices Q, K, V and the coefficient matrix of the Concat operation, head i represents the attention result of the i-th head, is the square root of the Key dimension, Attention represents the calculation process of the attention mechanism, and softmax is a function.
[0026] Preferably, the residual connection and normalization layer is to add the input and output of the multi-head self-attention mechanism layer and normalize it to a standard normal distribution. The calculation formula of the residual connection and normalization layer is as follows:
[0027] output=LayerNorm(input+Sublayer(input))
[0028] Among them, input represents input data, output represents the output of residual connection and normalization layer, Sublayer represents the corresponding network layer of input, and LayerNorm represents data normalization operation;
[0029] Preferably, the feedforward network layer processes the input data of the residual connection and normalization layer through two layers of linear mapping and activation function to improve the nonlinear fitting ability of the network. The calculation formula is:
[0030] output=Relu(input*W 1 *W 2 )
[0031] Among them, input and output represent the input and output of the feedforward network layer respectively, Relu represents the activation function used by the feedforward network layer, and W 1 is the weight matrix used in the first layer linear mapping, W 2 The weight matrix used for the second-layer linear mapping.
[0032] Preferably, the structure of the BiLSTM layer in step 2.1 is composed of two LSTM networks, a forward LSTM network and a backward LSTM network, the input sequence is respectively connected to the forward LSTM network and the backward LSTM network, and then the two LSTM networks are connected to each other and connected to the output layer together;
[0033] represents the forward output of the LSTM network at time n, represents the reverse output of the LSTM network at time n, x n represents the input of the BiLSTM network at time n, h n Represents the output of the BiLSTM network at time n. The output of the BiLSTM network at time n is represented by x n , Together, we decided that the BiLSTM network update formula is:
[0034]
[0035]
[0036]
[0037] Among them, w n 、v n are the weight matrices representing the output of the forward LSTM network at time n and the weight matrices representing the output of the reverse LSTM network at time n; b n h n Bias term used in the calculation.
[0038] Preferably, step 3 comprises:
[0039] Step 3.1: After training the BERT-BiLSTM model in step 2, obtain the service dataset characteristics of the domain vocabulary in step 1 and the classification results in step 2, change the loss function to obtain the optimal result, and for the imbalance of the quantity of service corpus, it is necessary to list the number of services under each service category to observe the imbalance of quantity; for the difficulty imbalance of service corpus, it is necessary to list the probability of correct classification of each service sample after training, that is, the difficulty of classification. Finally, compare the coefficient of variation of the two sets of data to measure the imbalance of the service corpus. The coefficient of variation calculation method is as follows:
[0040]
[0041] Among them, CV is the coefficient of variation, σ is the standard deviation of the data, and μ is the mean value of the data.
[0042] Step 3.2: According to the coefficient of variation obtained in step 3.1, the imbalance characteristics of the service text corpus are qualitatively determined, and according to the quality of the classification results, the coefficient of the zoom loss function is quantitatively adjusted to obtain the optimal result. The formula of the zoom loss function is:
[0043]
[0044] Among them, α t is the category quantity balance weight, θ is the quantity modulation coefficient, and pt Represents the probability that the sample belongs to the real sample, γ is the difficulty modulation coefficient, and the sample imbalance property focused on by the loss function is adjusted by controlling θ and γ. When θ = 1 and γ = 0, the function form is equivalent to the cross entropy loss function. When θ = 1 and γ > 0, the function form is equivalent to the focus loss function.
[0045] Compared with the prior art, the present invention has the following beneficial effects:
[0046] (1) A service text classification model WBBI is designed. Compared with traditional service text classification methods, its advantages are higher domain adaptability and adaptability to the imbalanced characteristics of data sets;
[0047] (2) A BERT-BiLSTM model based on domain vocabulary enhancement is proposed. This model uses the TF-IDF algorithm to obtain domain vocabulary in the domain corpus and expand the original vocabulary of the BERT model. Finally, the word vectors generated by the BERT model are combined with the context feature capture capability of the Bidirectional Long Short-Term Memory networks (BiLSTM) model to achieve service text classification.
[0048] (3) A zoom loss function that can be dynamically adjusted according to the imbalance of the data set is proposed. Its advantage is that it comprehensively considers the imbalance of data categories and the imbalance of difficulty, and has greater flexibility than traditional loss functions;
[0049] (4) In order to verify the effectiveness of the proposed method, a large number of experiments were conducted on real datasets crawled from the Internet. The experimental results show that compared with the baseline model, the classification effect of our proposed model is better. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.
[0051] In the attached picture:
[0052] Figure 1 is the WBBI model structure;
[0053] Figure 2 Input representation for the BERT model;
[0054] Figure 3 It is the transformer encoder unit structure;
[0055] Figure 4 It is the LSTM structure diagram;
[0056] Figure 5 This is the BiLSTM structure diagram;
[0057] Figure 6 Comparison chart between CE and Focal loss;
[0058] Figure 7 This is a three-dimensional comparison chart of CE and Focal loss;
[0059] Figure 8 is the zoom loss function image;
[0060] Fig. 9 It is the curve of Loss changing with epoch;
[0061] Fig.10 This is the curve of Macro-F1 changing with epoch;
[0062] Fig.11 is a flow chart of the method of the present invention; DETAILED DESCRIPTION
[0063] The preferred embodiments of the present invention are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0064] Example:
[0065] See attached Figure 1-11 As shown, first of all, the present invention is a service classification method based on the WBBI model. The WBBI service text classification model is divided into three parts: domain vocabulary extraction, model training, and zoom loss function optimization. The WBBI model structure is as follows Figure 1 shown.
[0066] A service text classification method based on a domain BERT model, comprising:
[0067] Step 1: Use the TF-IDF algorithm to extract domain vocabulary from the service text corpus;
[0068] Step 3: Based on the serving dataset characteristics in step 1 and the classification results in step 2, select the best loss function to balance the dataset.
[0069] Furthermore, the step 1 comprises:
[0070] Step 1.1: To extract vocabulary from the service text corpus, the TF-IDF (Term Frequency-Inverse Document Frequency) algorithm is used to extract vocabulary from the service text corpus and expand the BERT vocabulary. The basic idea of this algorithm is that if a vocabulary has a high frequency of occurrence in one article and rarely appears in other articles, it is considered that the vocabulary has good representative ability. The calculation formula of the TF-IDF algorithm is as follows:
[0071]
[0072] Among them, TF represents word frequency, IDF represents the inverse text frequency index of word i in article j, and TF i,j represents the frequency of word i appearing in article j, N represents the total number of articles in the corpus, and n i represents the number of articles in the corpus containing word i;
[0073] Step 1.2: The present invention performs word segmentation on the articles in the crawled database, removes punctuation marks, and lowercases them uniformly. The TF-IDF algorithm of step 1.1 is used to calculate the weights of all words and sort them. The BERT model used in the present invention is based on the BERT-Base-uncased model publicly released by Google, whose vocabulary contains 28,996 words and provides 101 placeholders [unused] to expand the vocabulary. The first 101 words with the largest vocabulary weights in step 1.2 are selected to fill the original vocabulary, and finally the expanded domain vocabulary is obtained.
[0074] Step 2: Based on step 1, build a BERT-BiLSTM model, input the domain vocabulary extracted in step 1 into the BERT vocabulary of the model, and then input the service text corpus into the BERT-BiLSTM model for training to achieve service text classification. Specifically:
[0075] Step 2.1: Figure 1 , establish the BERT-BiLSTM model structure. The BERT-BiLSTM model structure is composed of the BERT model and the BiLSTM model. After the BERT model operation is completed, it enters the BiLSTM model. The BERT model includes an embedding layer, multiple encoder layers, and a pooler layer in sequence. The BiLSTM layer includes multiple LSTM layers. The BERT model generates text word vectors through the embedding layer, captures text vocabulary features through the multi-head attention mechanism and the feedforward neural network layer in the encoder layer, and finally enters the BiLSTM layer through the fully connected layer in the pooler layer. The BiLSTM layer is responsible for obtaining the contextual features between word vectors, and finally classification is performed through the fully connected layer;
[0076] Step 2.2: Input a text sentence into the BERT model structure of step 2.1, and convert the words in the sentence into embeddings. The input of BERT consists of token embeddings, segment embeddings, and position embeddings. Token embeddings represent the embeddings of text words, which are matched greedily according to the set domain vocabulary. For example, the word "embeddings" that is not in the vocabulary will be decomposed into ['em', '##bed', '##ding', '##s'] when input. In this way, the model can find the corresponding embedding of any word in the vocabulary.
[0077] Segment embedding indicates the sentence to which the word belongs, and position embedding indicates the specific position of the word in the input text. In addition, the vocabulary also presets flags with special functions, among which: the [CLS] flag is located at the first position of the input text as the flag for the classification task; the [SEP] flag is used as a separator between sentences. If the input text is a pair of sentences, it is added after the two sentences respectively; the [Mask] flag is used to mask the words in the sentence. Sentence input method is as follows: Figure 2 shown.
[0078] Furthermore, in step 2.1, the encoder layer in the BERT model is composed of a transformer encoder structure, and its basic unit structure is as follows: Figure 3 As shown in Figure 2, the transformer encoder structure includes an input embedding layer, a multi-head self-attention mechanism, a residual connection / layer normalization, a feedforward neural network, and a residual connection / layer normalization.
[0079] Furthermore, the multi-head self-attention mechanism layer mainly emphasizes different parts of the text by updating the weights, where the single-head attention mechanism is calculated as follows:
[0080]
[0081] Among them, matrices Q, K, and V represent Query, Key, and Value matrices respectively. The goal is to transform the attention matrix into a standard normal distribution.
[0082] The multi-head attention mechanism is a parallel operation of the single-head attention mechanism. The calculation method of the multi-head self-attention mechanism is:
[0083] MultiHead(Q,K,V)=Concat(head 1 ,…head h)W
[0084] head i =Attention(QW i Q ,KW i K ,VW i V )
[0085] Among them, Concat is a concatenation operation, that is, each self-attention result is connected end to end. The matrices Q, K, and V represent the Query, Key, and Value matrices respectively. i Q , W i K , W i V , W are matrices Q, K, V and the coefficient matrix of the Concat operation, head i represents the attention result of the i-th head, is the square root of the Key dimension. Attention represents the calculation process of the attention mechanism. The attention mechanism calculation process used in this model is calculated through the softmax function. The softmax calculation method is:
[0086]
[0087] Among them, z i is the attention parameter of the i-th head to be calculated, and C is the sum of all attention heads.
[0088] Furthermore, the residual connection and normalization layer adds the input and output of the multi-head self-attention mechanism layer and normalizes them to a standard normal distribution, the purpose of which is to prevent the gradient vanishing problem in deep networks. The residual connection and normalization layer calculation formula is as follows:
[0089] output=LayerNorm(input+Sublayer(input))
[0090] Among them, input represents input data, output represents the output of residual connection and normalization layer, Sublayer represents the corresponding network layer of input, and LayerNorm represents data normalization operation;
[0091] Preferably, the feedforward network layer processes the input data through two layers of linear mapping and activation function to improve the nonlinear fitting ability of the network. The calculation formula is:
[0092] output=Relu(input*W 1 *W 2 )
[0093] Among them, input and output respectively represent the input and output of the feedforward network layer, Relu represents the activation function adopted by the feedforward network layer, and W 1 is the weight matrix used for the first-layer linear mapping, and W 2 is the weight matrix used for the second-layer linear mapping.
[0094] Furthermore, the Long Short-Term Memory networks (LSTM) model is an improved structure based on the Recurrent Neural Network (RNN) to solve the problems of gradient vanishing and gradient explosion that occur in traditional RNN models. The LSTM cell structure is as Figure 4 shown.
[0095] Among them, x t is the input text word vector matrix, h t-1 is the output value of the previous cell, and C t-1 is the state of the previous cell. Then the update formula of the LSTM cell is:
[0096] f t = σ(W f · [h t-1 , x t + b f )
[0097]
[0098] o t = σ(W o · [h t-1 , x t + b o )
[0099] h t = o t * tanh(C t )
[0100] Among them, σ(·) is the Sigmoid function; tanh(·) is the hyperbolic tangent function; W f , W o are the weight matrices used to calculate f t and o t respectively; b f , b o are the bias terms used to calculate f t and o t respectively.
[0101] In the text classification task, it is more accurate to obtain feature information by analyzing features from the context of the current word. Therefore, this paper further adopts the BiLSTM model to enrich the model's word vector feature acquisition capabilities. BiLSTM is composed of forward LSTM and reverse LSTM. The model is as follows Figure 5 shown.
[0102] The structure of the BiLSTM layer in step 2.1 is composed of two LSTM networks, a forward LSTM network and a backward LSTM network. The input sequence is connected to the forward LSTM network and the backward LSTM network respectively, and then the two LSTM networks are connected to each other and connected to the output layer together.
[0103] represents the forward output of the LSTM network at time n, represents the reverse output of the LSTM network at time n, x n represents the input of the BiLSTM network at time n, h n Represents the output of the BiLSTM network at time n. The output of the BiLSTM network at time n is represented by x n , Together, we decided that the BiLSTM network update formula is:
[0104]
[0105]
[0106]
[0107] Among them, w n 、v n are the weight matrices representing the output of the forward LSTM network at time n and the weight matrices representing the output of the reverse LSTM network at time n; b n h n Bias term used in the calculation.
[0108] Step 3: According to the service text corpus characteristics and classification results in step 2, select the best loss function to balance the dataset.
[0109] After the BERT-BiLSTM model training in step 2, the loss function is changed to obtain the optimal result by analyzing the service text corpus characteristics and classification results in step 2. For the imbalance of the quantity of the service corpus, it is necessary to list the number of services under each service category to observe the imbalance of the quantity; for the imbalance of the difficulty of the service corpus, it is necessary to list the probability of correct classification of each service sample after training, that is, the difficulty of classification. Finally, by comparing the coefficient of variation of the two groups of data to measure the imbalance of the service corpus, determine the strength of the imbalance of quantity and difficulty, and provide a preliminary basis for the value of the zoom loss function in the next step. The coefficient of variation calculation method is as follows:
[0110]
[0111] Among them, CV is the coefficient of variation, σ is the standard deviation of the data, and μ is the mean value of the data.
[0112] Step 3.2:
[0113] In view of the imbalance of data sets in classification tasks, neural networks usually use a weighted cross entropy loss function to balance the categories in the data set, that is, to weight the sum of the cross entropy (CE) of each training sample. The weighted cross entropy calculation formula is as follows:
[0114] CE(p t )=-α t log(p t )
[0115] Among them, p t Represents the probability of the sample being correctly classified, α t Represents the category quantity balance weight, and its value is usually determined based on the inverse category frequency.
[0116] However, the difficulty of classifying categories in classification tasks often has a great impact on the training results. For example, in image recognition tasks, objects with relatively simple features (such as grass and ocean) are easier to recognize, while objects with complex features (such as animals and buildings) are more difficult to recognize. If the weight of difficult-to-classify samples is increased, the network will pay more attention to difficult-to-classify samples and improve the classification effect.
[0117] In order to increase the weight of difficult-to-classify samples, a focal loss function (FL) is used. The calculation formula is as follows:
[0118] FL(p t )=-α t (1-p t ) γ log(p t )
[0119] Among them, α t represents the category quantity balance weight, p t Represents the probability of the sample being correctly classified, (1-p t ) γ It is called the modulation coefficient, which controls the weight of difficult and easy samples. Specifically, when γ>0, the difficult samples (p t <0.5) weight will increase, easy to classify samples (p t >0.5) The weight will be reduced, making the model more focused on training difficult-to-classify samples. Figure 6 .
[0120] Figure 6 In the example, the horizontal axis represents the sample difficulty index p t , the vertical axis represents when α t =1, the values of various loss functions. It can be seen that the focus loss function reduces the weight of easy-to-classify samples in the loss, making the model pay more attention to training difficult-to-classify samples. Based on the above analysis, we can classify the sample properties as follows:
[0121] Table 1 Sample properties and weight changes
[0122] Sample properties Weight changes Sample properties Weight changes many reduce Less difficult Increase few Increase Duoyi reduce Disaster Increase How difficult Unknown changes easy reduce Shao Yi Unknown changes
[0123] It can be seen that considering the number characteristics and difficulty characteristics of the samples, the sample properties can be divided into four types: small number and high classification difficulty (small number and difficult), large number and low classification difficulty (large number and easy), large number and high classification difficulty (large number and difficult), small number and low classification difficulty (small number and easy). In order to achieve the purpose of balancing the data set, the weight change direction of small number and large number of easy samples is clear, while the weight change of large number and small number of easy samples conflicts, which makes it difficult for the loss function to assign appropriate weights to these two types of samples. In order to intuitively analyze the changes in the loss function under the two types of properties, the category quantity balance weight α can be drawn t , sample difficulty index p t 3D image of the variation with the loss function.
[0124] like Figure 7 , the x-axis represents the category quantity balance weight α t , the y-axis represents the sample difficulty index p t, the z-axis is the loss function value. The closer the image is to the origin, the stronger the difficulty attribute of the sample is, and the farther it is from the origin, the stronger the sparseness attribute of the sample is. Compared with the cross entropy loss function, the focal loss function moves toward the origin, that is, compared with the number of samples, the focal loss function tends to pay more attention to the imbalance of the difficulty of the samples and give higher weights to difficult samples. However, if the imbalance of the number of categories in the data set is higher than the imbalance of the difficulty, the focal loss function will assign higher weights to the samples with a large number, which will reduce the balance ability of the data set.
[0125] In order to obtain a loss function that can adapt to the imbalanced nature of the data set, this paper proposes a zoom loss function (Zoom loss, ZL) based on the original focus loss function. Its function form is:
[0126]
[0127] Among them, α t is the category quantity balance weight, θ is the quantity modulation coefficient, and p t Represents the probability that the sample belongs to the real sample, γ is the difficulty modulation coefficient, and the sample imbalance property focused on by the loss function is adjusted by controlling θ and γ. When θ = 1 and γ = 0, the function form is equivalent to the cross entropy loss function. When θ = 1 and γ > 0, the function form is equivalent to the focus loss function.
[0128] It can be seen that the characteristic of the zoom loss function is that the focus of the loss function can be adjusted according to the imbalance of the number and the imbalance of difficulty of the data set. After determining the preliminary properties of the imbalance of the number and the imbalance of difficulty of the service text corpus in step 3.1, the size of the zoom loss function parameters can be specifically adjusted according to the quality of the experimental results.
[0129] Finally, we can draw a conclusion from the experimental results: when the imbalance of the number of data sets is stronger than the difficulty, θ can be increased so that the loss function tends to assign larger weights to the few samples; when the difficulty of the data set is stronger than the imbalance, γ can be increased so that the loss function tends to assign larger weights to the difficult samples. This allows the zoom loss function to make adaptive adjustments to the imbalance characteristics of the data set and improve the classification performance of the model.
[0130] Simulation experiment and analysis
[0131] 1. Experimental Preparation
[0132] 1.1 Experimental Dataset Acquisition and Processing
[0133] In order to obtain real data of service description texts, this paper uses crawler tools to crawl 437 categories and a total of 22,205 service description texts from the Programmable Web (https: / / www.programmableweb.com) website, which are used as the original data set for this experiment.
[0134] In order to ensure the model training effect, invalid services need to be removed from the data set. By analyzing the data set, invalid services with problems such as repeated registration of the same service, identical service description text, invalid service, empty service label or empty service text in the original data are removed. After the service screening data is segmented, punctuation marks are removed, and stop words are removed, 51 categories and a total of 13,694 service texts are finally obtained as the final data set of this experiment.
[0135] 1.2 Experimental environment construction
[0136] The experimental equipment used in this experiment is Intel(R) Core(TM) i7-10875H CPU @ 2.30GHz; GeForce GTX 2060 graphics processor, CUDA version 11.1.114. The code is programmed on the pycharm platform, and the 1.9.0 version of the pytorch deep learning framework is used to build the model. The BERT model version uses the Google pre-trained BERT-Base-uncased model, which contains 12 multi-head attention layers, 768 hidden layers, 12 attention heads per layer, and a total of 110 million parameters.
[0137] In terms of data set processing, this paper divides the training set and the validation set in a ratio of 10:1, sets the batch_size to 8, sets the maximum number of iterations to 100, and sets the early stopping method to stop when the evaluation indicators of each type of samples no longer increase within 6 generations.
[0138] 1.3 Model Evaluation Method
[0139] In classification problems, precision, recall, and F1 value are often used as evaluation methods for multi-classification tasks under imbalanced data. The confusion matrix is shown in Table 2.
[0140] Table 2 Confusion matrix table
[0141] Positive Category Negative True Class True class (Tp) True Negative (TN) Fake False Positive (FP) False Negative (FN)
[0142] According to the confusion matrix table, each evaluation index can be expressed as follows:
[0143] (1) Accuracy
[0144]
[0145] (2) Recall rate
[0146]
[0147] (3) F1 value
[0148]
[0149] In multi-classification problems, there are many ways to calculate the F1 value. This article uses the Macro-F1 value as the measurement standard in multi-classification problems. The calculation method is:
[0150]
[0151] Where n represents the number of classification categories; F1 i Represents the F1 value of the i-th class.
[0152] The model also often uses accuracy as an indicator to evaluate the classification performance of the model. The calculation formula is:
[0153]
[0154] 2. Experimental design and results analysis
[0155] 2.1 Domain vocabulary extraction experiment
[0156] To verify the advantages of the model after adding domain vocabulary, this paper introduces a control group BERT-Random, which randomly obtains the same number of words from the data set and adds them to the vocabulary in the same way for training. By comparing the classification effect of the BERT model, the BERT-Random model and the domain vocabulary used in this paper to enhance the BERT model, the results are shown in Table 3:
[0157] Table 3 Domain vocabulary enhancement results
[0158]
[0159] According to the experimental results, the model used in this paper has improved accuracy and Macro-F1 value compared with the original BERT model and BERT-Random model, and the training time is significantly reduced. According to the experimental comparative analysis, vocabulary enhancement for the original vocabulary can effectively improve the model performance, and the enhancement for domain vocabulary has achieved the best effect.
[0160] 2.2 Loss function improvement experiment
[0161] To verify the effectiveness of the zoom loss function, this paper conducted experiments on the WBBI model and compared it with the weighted cross entropy and focal loss functions. The experiment was conducted in the range of θ∈[1.0,2.5], γ∈[0.0,2.0]. The experimental results are shown in Table 4:
[0162] Table 4. Experimental results of loss function
[0163]
[0164] Experiments show that the zoom loss function is better than CE and Focal loss when θ=2.0 and γ=0.0. It can also be seen that the experimental effect of CE is better than Focal loss. This is because the number of this dataset is highly unbalanced, and the focal loss function mainly focuses on difficult-to-classify samples, and its performance cannot be fully utilized. When the zoom loss function θ=2.0 and γ=0.0, the loss function focuses on the number imbalance, which is more suitable for the dataset experiment in this paper, and finally obtains the best effect. At the same time, it is also noted that the value of θ should not be too high. Too high a value of θ will increase the loss value of easy-to-classify samples, which will greatly increase the model convergence time. The comparison of model training time is shown in Table 5.
[0165] Table 5 Comparison of loss function training time
[0166]
[0167] It can be seen from Table 5 that when γ remains unchanged and the value of θ continues to increase, the model training time will continue to increase. On the contrary, when the value of θ remains unchanged and γ increases, the training time will continue to decrease. This is because in the dataset of this article, when the focus of the zoom loss function gradually shifts to the majority of easy samples, the weight of the easy sample loss value increases, which increases the overall loss value of the model and makes it difficult for the model to converge.
[0168] 2.3 Model comparison experiment
[0169] To verify the superiority of the WBBI service classification model, it is compared with TextCNN, BiLSTM-Attention, RCNN, and Transformer text classification models under the same experimental environment. Among them, TextCNN uses pre-trained word vectors to initialize the word embedding matrix to generate word vectors, and the convolution kernel size uses four sizes [2, 3, 4, 5], and finally outputs through the maximum pooling layer and softmax layer. The model parameters are:
[0170] Table 6 TextCNN model parameters
[0171]
[0172] The BiLSTM-attention model has a two-layer bidirectional LSTM structure. The forward output and the backward output are added in the output layer and output after passing through the attention mechanism layer and the softmax layer.
[0173] Table 7 BiLSTM-Attention model parameters
[0174]
[0175] The RCNN model also has a two-layer bidirectional LSTM structure, and in the output layer, the forward output, backward output and word vector are concatenated to obtain the final word representation, and finally output through the maximum pooling layer. The model parameters are:
[0176] Table 8 RCNN model parameters
[0177]
[0178] The Transformer model as a whole consists of an encoder and a decoder, which captures text feature information through a multi-head attention mechanism. The model parameters are:
[0179] Table 9 Transformer model parameters
[0180]
[0181] The results of the model comparison experiment are shown in Table 10. As can be seen from Table 10, the WBBI model experimental results are better than the baseline model in both the training set and the validation set evaluation indicators, and the classification effect is significantly improved. Among them, compared with the TextCNN model Macro-F1 increased by 4.29 percentage points. This is because the TextCNN model can only obtain local features of the text, while the bidirectional LSTM layer of WBBI can better obtain contextual information and has a strong feature extraction capability. Compared with the BiLSTM-Attention and RCNN models, Macro-F1 increased by 6.59 percentage points and 5.3 percentage points respectively. This is because in word vector processing, the word vector features obtained by the BERT model have better representation capabilities than other deep learning models. At the same time, it can be seen that the classification effect of RCNN is better than BiLSTM-Attention. This is because RCNN combines the bidirectional LSTM information while also integrating the original word vector information, and the information feature acquisition is more comprehensive. Compared with the Transformer model, the Macro-F1 value of the WBBI model increased by 43 percentage points. This is because the Transformer model requires more iterations for training. Under the same training conditions, the convergence speed of the Transformer model is slower than that of the WBBI model, so the training effect is not good.
[0182] Table 10 Comparative experimental results
[0183]
[0184] Fig. 9 and Fig.10 The loss values and Macro-F1 changes of different models during the iteration process are shown respectively. Fig. 9 It can be seen that compared with the baseline model, the WBBI model training process is more stable and has a lower loss value, and the overall curve fluctuation is small. Fig.10 It can be seen that WBBI has the highest Macro-F1 value, the curve fluctuates less, and the training process is stable, which well reflects the advantages of the model.
[0185] In summary, this paper proposes a novel service text classification model to solve the service classification problem. The model first obtains word vector context features for classification through the BERT-BiLSTM model enhanced by domain vocabulary, and finally proposes a zoom loss function. By analyzing the imbalance of the number of categories and the imbalance of difficulty in the data set, the corresponding parameter adjustment is made to improve the classification performance of the model. Experiments show that the framework proposed in this paper has the best classification effect.
[0186] The above shows and describes the basic principles, main features and advantages of the present invention. It should be understood by those skilled in the art that the present invention is not limited to the above embodiments. The above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention. The scope of protection of the present invention is defined by the attached claims and their equivalents.
Claims
1. A service text classification method based on the domain BERT model, Features: include: Step 1: Use the TF-IDF algorithm to extract domain vocabulary from the service text corpus; Step 2: Based on step 1, a BERT-BiLSTM model is established. After the domain vocabulary extracted in step 1 is input into the BERT vocabulary of the BERT-BiLSTM model, the service text corpus is input into the BERT-BiLSTM model for training to realize service text classification. Step 2 is specifically as follows: Step 2.1: Establish the BERT-BiLSTM model structure. The BERT-BiLSTM model structure is composed of the BERT model and the BiLSTM model. After the BERT model operation is completed, it enters the BiLSTM model. The BERT model includes an embedding layer, multiple encoder layers, and a pooler layer in sequence. The BiLSTM layer includes multiple LSTM layers. The BERT model generates text word vectors through the embedding layer, captures text vocabulary features through the multi-head attention mechanism and feedforward neural network layer in the encoder layer, and finally enters the BiLSTM layer through the fully connected layer in the pooler layer. The BiLSTM layer is responsible for obtaining the contextual features between word vectors, and finally classification is performed through the fully connected layer; Step 2.2: Input a text sentence into the BERT model structure of step 2.1, and convert the words in the sentence into embeddings. The input of BERT consists of word embedding, segment embedding, and position embedding. Word embedding represents the embedding of text words, which performs vocabulary matching according to the set domain vocabulary according to the greedy principle. Segment embedding indicates the sentence to which the word belongs, and position embedding identifies the specific position of the word in the input text. Step 3: According to the service text corpus characteristics and classification results of step 2, select the best loss function to balance the data set; step 3 includes: Step 3.1: After training the BERT-BiLSTM model in step 2, obtain the service dataset characteristics of the domain vocabulary in step 1 and the classification results in step 2, change the loss function to obtain the optimal result, and for the imbalance of the quantity of service corpus, it is necessary to list the number of services under each service category to observe the imbalance of quantity; for the difficulty imbalance of service corpus, it is necessary to list the probability of correct classification of each service sample after training, that is, the difficulty of classification. Finally, compare the coefficient of variation of the two sets of data to measure the imbalance of the service corpus. The coefficient of variation calculation method is as follows: Among them, CV is the coefficient of variation, σ is the standard deviation of the data, and μ is the mean value of the data; Step 3.2: According to the coefficient of variation obtained in step 3.1, the imbalance characteristics of the service text corpus are qualitatively determined, and according to the quality of the classification results, the coefficient of the zoom loss function is quantitatively adjusted to obtain the optimal result. The formula of the zoom loss function is: Among them, α t is the category quantity balance weight, θ is the quantity modulation coefficient, and p t Represents the probability that the sample belongs to the real sample, γ is the difficulty modulation coefficient, and the sample imbalance property focused on by the loss function is adjusted by controlling θ and γ. When θ = 1 and γ = 0, the function form is equivalent to the cross entropy loss function. When θ = 1 and γ > 0, the function form is equivalent to the focus loss function.
2. A service text classification method based on a domain BERT model according to claim 1, Features: The step 1 comprises: Step 1.1: Use the TF-IDF algorithm to extract vocabulary from the service text corpus. The calculation formula is as follows: Among them, TF represents word frequency, IDF represents the inverse text frequency index of word i in article j, and TF i,j represents the frequency of word i appearing in article j, N represents the total number of articles in the corpus, and n i represents the number of articles in the corpus containing word i; Step 1.2: Use the TF-IDF algorithm in step 1.1 to calculate the weights of all words and sort them to finally obtain the domain words extracted from the service text corpus.
3. A service text classification method based on a domain BERT model according to claim 2, Features: In step 2.1, the encoder layer in the BERT model is composed of a transformer encoder structure, and the transformer encoder structure includes an input embedding layer, a multi-head self-attention mechanism, a residual connection / layer normalization, a feedforward neural network, and a residual connection / layer normalization in sequence.
4. A service text classification method based on a domain BERT model according to claim 3, Features: The calculation method of the multi-head self-attention mechanism is: MultiHead(Q,K,V)=Concat(head 1 ,…head h )W Among them, Concat is a concatenation operation, that is, connecting each self-attention result end to end. The matrices Q, K, and V represent the Query, Key, and Value matrices respectively. W is the coefficient matrix of the matrix Q, K, V and Concat operation, head i represents the attention result of the i-th head, is the square root of the Key dimension, Attention represents the calculation process of the attention mechanism, and softmax is a function.
5. A service text classification method based on a domain BERT model according to claim 4, Features: The residual connection and normalization layer adds the input and output of the multi-head self-attention mechanism layer and normalizes them to a standard normal distribution. The calculation formulas of the residual connection and normalization layer are as follows: output=LayerNorm(input+Sublayer(input)) Among them, input represents input data, output represents the output of residual connection and normalization layer, Sublayer represents the corresponding network layer of input, and LayerNorm represents data normalization operation.
6. A service text classification method based on a domain BERT model according to claim 5, Features: The feedforward network layer processes the input data of the residual connection and normalization layer through two layers of linear mapping and activation function to improve the nonlinear fitting ability of the network; its calculation formula is: output=Relu(input*W 1 *W 2 ) Among them, input and output represent the input and output of the feedforward network layer respectively, Relu represents the activation function used by the feedforward network layer, and W 1 is the weight matrix used in the first layer linear mapping, W 2 The weight matrix used for the second-layer linear mapping.
7. A service text classification method based on a domain BERT model according to claim 6, Features: The structure of the BiLSTM layer in step 2.1 is composed of two LSTM networks, a forward LSTM network and a backward LSTM network. The input sequence is connected to the forward LSTM network and the backward LSTM network respectively, and then the two LSTM networks are connected to each other and connected to the output layer together. represents the forward output of the LSTM network at time n, represents the reverse output of the LSTM network at time n, x n represents the input of the BiLSTM network at time n, h n Represents the output of the BiLSTM network at time n. The output of the BiLSTM network at time n is represented by x n , Together, we decided that the BiLSTM network update formula is: Among them, w n 、v n are the weight matrices representing the output of the forward LSTM network at time n and the weight matrices representing the output of the reverse LSTM network at time n; b n h n Bias term used in the calculation.
Citation Information
Patent Citations
Judicial public opinion sensitive information identification method integrated with domain term dictionary
CN112231472A
Systems and Methods for Distilled BERT-Based Training Model for Text Classification
US20210150340A1