Chemical Emergency News Classification Method Based on ChineseBERT Model and Attention Mechanism
By improving the ChineseBERT model and the attention mechanism, the problems of high time complexity and insufficient feature extraction in the text classification of chemical emergencies were solved, and more accurate and efficient text classification was achieved.
Patent Information
- Application Number
- CN202210030824.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-12
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-01-12
AI Technical Summary
The existing technology has problems such as high time complexity, large spatial complexity, poor model robustness, insufficient feature extraction, and insufficient global context information in the classification of chemical emergencies.
The chemical emergencies news classification method based on the ChineseBERT model and attention mechanism is adopted. By improving the architecture of the ChineseBERT model, pinyin vectors and character vectors are extracted for fusion, and position vectors are added for integration, input into the Bert model for training, sharing Bert parameters to reduce time complexity, and using the continuous attention mechanism to enhance context feature information.
It realizes multi-level precise characterization of text data features, reduces time complexity, decouples different semantics in the same character form, enhances context feature information, and improves text classification accuracy of chemical emergencies news.
Smart Images

Figure CN114510569B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of text classification and natural language processing, and particularly relates to a method for classifying chemical accident news based on the ChineseBERT model and the attention mechanism. Background Art
[0002] The ChineseBERT model is mainly a Chinese pre-trained model that integrates glyph and pinyin information. The model concatenates character embedding, glyph embedding, and pinyin embedding; then through a fusion layer, a d-dimensional fusion embedding is obtained; finally, it is added to the position embedding and segment embedding to form the input of the Transformer-Encoder layer. Since the NSP task is not used during pre-training, the segment embedding is omitted in the model structure.
[0003] The MLP multi-layer perceptron, also called an artificial neural network, can have multiple hidden layers in the middle in addition to the input and output layers. The simplest MLP only contains one hidden layer, that is, a three-layer structure, and the layers of the multi-layer perceptron are fully connected. The bottom layer of the multi-layer perceptron is the input layer, the middle is the hidden layer, and the last is the output layer.
[0004] The Attention mechanism considers different weight parameters for each element of the input, so as to pay more attention to the parts similar to the input elements and suppress other useless information. Its greatest advantage is that it can consider global and local connections in one step and can be calculated in parallel, which is particularly important in the context of big data.
[0005] When facing the problem of news text classification, researchers will choose to integrate sentence similarity, neural networks, etc. into text classification, ignoring the time complexity during text data training, the pinyin information of Chinese characters, the extraction of deep text features, and the semantic information of the corresponding data. Therefore, by improving the architecture of the ChineseBERT pre-trained model and sharing the parameters of the Bert model, the robustness of the model is improved and the time complexity is reduced. At the same time, combined with the cascaded attention mechanism, the character-to-subsequence context feature information is obtained, so as to solve the problem of classifying Chinese chemical accident news texts, and further improve the accuracy of text classification.
[0006] In existing text classification methods, some only focus on the similarity between the feature vectors of classified short texts and the central vectors of feature vector clusters in the preset feature vector cluster set, without considering the entity feature information of text information; some only focus on topic semantic features, without considering the global feature information of texts. There are also some methods that mainly perform simple extraction of features, without considering the use of pre-trained models and the relationships of long dependencies.
[0007] When facing the problem of chemical emergency news text classification, existing papers mainly rely on traditional feature extraction methods and topic recognition methods, and secondly on deep neural network classification models, etc. However, there are still many problems to be solved regarding text classification: the time complexity, space complexity, and model robustness issues during the training of chemical news information; the information extracted by feature extraction cannot fully depict the full text information of the text, and some semantics are different, such as homographs with different meanings, and the global context information is not comprehensive enough; for the Chinese pre-trained model ChineseBERT, during pre-training, for glyph information, it needs to be processed through instantiated images of different fonts, and then recognition learning and flattening operations are required, occupying a lot of space complexity; and the model is trained from scratch. It is required in the vector layer, but it is also trained from scratch in the transformer-encoder layer, resulting in an increase in time complexity. Summary of the Invention
[0008] Object of the Invention: Aiming at the problems existing in the prior art, the present invention proposes a chemical emergency news classification method based on the ChineseBERT model and the attention mechanism, which can accurately depict the characteristics of text data at multiple levels. By improving the architecture of the ChineseBERT model, that is, extracting pinyin vectors and character vectors for fusion, and then adding position vectors for integration, and inputting them into the Bert model for training, where the Bert parameters are shared to reduce the time complexity, decouple different semantics belonging to the same character form, and at the same time use the concatenated attention mechanism to enhance the context feature information to make up for the loss of traditional news text feature information, improve the actual application efficiency of chemical emergency news, and achieve accurate text classification.
[0009] Technical Solution: The present invention proposes a chemical emergency news classification method based on the ChineseBERT model and the attention mechanism, specifically including the following steps:
[0010] (1) Perform text preprocessing on the chemical emergency news text data D to obtain the news text data D1;
[0011] (2) Process the chemical industry emergency text data D1 through the word2vec model to obtain the text feature vector R1. Input the word vector R1 into the Word Attention model to obtain the new word dependency feature information H1, and then input the word dependency feature information H1 into the Seq Attention model to obtain the subsequence feature information H2;
[0012] (3) Process the text data D1 through an open-source pinyin package to obtain the corresponding pinyin sequence, and then input it into the MLP. After passing through the max-pooling layer, output the pinyin vector H3. Perform one-hot encoding on the preprocessed text to obtain the character vector H4, and perform matrix embedding with the pinyin vector H3 to obtain the 2D matrix vector R3;
[0013] (4) Integrate the matrix feature information R3 and the position vector information R4 to obtain the feature information H5, and input H5 into the Bert pre-trained model to obtain the corresponding feature information H6;
[0014] (5) Integrate the context feature information H2 in step (2) and the semantic feature information H6 in step (4), and input them into the CNN model to obtain the final text classification result.
[0015] Further, step (1) includes the following steps:
[0016] (11) Define the chemical industry emergency news text data set as D, define Text as a single text data, and define id, title, and label as the single text serial number, the title of the data, and the text label respectively, and satisfy the relationship Text = {id, title, label}, D = {Text1, Text2,..., Text i ,…,Text n}, where Text i is the i-th text information data in D, where n = len(D) is the number of texts in D, and the variable i ∈ [1, n];
[0017] (12) Define the processed chemical industry emergency text data set as D1, D1 = {Text1, Text2,..., Text j ,…,Text m}, where Text j is the j-th text information data in D1, where m = len(D1) is the number of texts in D1 respectively, and the variable j ∈ [1, m];
[0018] (13) Read the data set D and traverse the entire data set;
[0019] (14) If title == null, execute (15); otherwise, execute (16).
[0020] (15) Delete the corresponding row data.
[0021] (16) Remove some useless characters according to the stop word list.
[0022] (17) Save the preprocessed text dataset D1.
[0023] Further, the step (2) includes the following steps:
[0024] (201) Read the preprocessed text dataset D1.
[0025] (202) Define the word feature vector set R1.
[0026] (203) Perform data word segmentation through the word2vec model, and obtain the text word feature vectors by training with the word2vec model.
[0027] (204) Save the word feature vector R1, and satisfy is the i-th word feature vector in the data vector set, where the variable i ∈ [1, a], and a is the number of word vectors after word segmentation.
[0028] (205) Define the word dependency feature vector H1 based on the attention mechanism.
[0029] (206) Input the word feature vector R1 into the Attention mechanism to obtain the word dependency feature vector based on attention. Where represents the j-th word dependency feature vector in the text, satisfying The variable j ∈ [1, b], b is the number of word dependency feature vectors, and the input and adjustment method of the Attention mechanism is to use softmax normalization to perform weight matrix W f Adjustment, and then multiply by V. Where d k is the dimension of a Q and K vector. is the scale scalar factor, representing query, key, and value respectively.
[0030] (207) Define the loop variable k to learn the word feature vector H1 of the first-level attention mechanism, and the initial value of k is 1.
[0031] (208) Define the subsequence dependency feature vector H2 based on the attention mechanism.
[0032] (209) If k ≤ b, then execute (210); otherwise, execute (212);
[0033] (210) Input the word dependency feature vector H1 into the Attention mechanism to obtain the attention-based subsequence dependency feature vector where represents the t-th subsequence dependency feature vector in the text, satisfying The variable t ∈ [1, c], and c is the number of subsequence dependency feature vectors;
[0034] (211) k = k + 1;
[0035] (212) Output and save the feature vector H2 of the secondary attention mechanism.
[0036] Furthermore, the step (3) includes the following steps:
[0037] (31) Define the pinyin feature vector H3, define the one-hot character vector H4, and define the fusion embedding matrix R3;
[0038] (32) Read the text data D1 into the open-source pinyin package to obtain the pinyin representation, input it into the MLP. There are 3 hidden layers in the neural network, with 64 nodes in each hidden layer, and then obtain the pinyin vector through the max pooling layer satisfying is the pinyin vector corresponding to the i-th character in the data vector set, where the variable i ∈ [1, d], and d is the number of pinyin vectors;
[0039] (33) Read the preprocessed data D1, and obtain the character vector through one-hot encoding the character vector satisfying is the character feature vector of the j-th character in the data vector set, where the variable j ∈ [1, e];
[0040] (34) Fuse the pinyin vector H3 and the character vector H4 to obtain the fusion embedding vector Mainly use the fully connected layer with a learnable matrix to induce the embedding of the matrix vector and fuse the matrix vector where represents the fusion feature vector corresponding to the t-th character in the text, and the variable t ∈ [1, s].
[0041] Furthermore, the step (4) includes the following steps:
[0042] (41) Define the position vector R4, define the feature vector matrix H5 that fuses the position vector, and define the feature vector H6 after Bert pre-training;
[0043] (42) Add the fusion matrix vector R3 to the positional Embedding to obtain the integrated feature vector matrix where the variable h ∈ [1, f];
[0044] (43) Read the integrated feature vector matrix H5 and input it into the Bert model for training to obtain the final feature information vector H6, where is the p-th feature vector of the vector after Bert training, where the variable p ∈ [1, g]. Share the training parameters of the Bert model to obtain the corresponding training feature vectors.
[0045] Furthermore, step (5) includes the following steps:
[0046] (51) Read the context feature information H2 and read the semantic information H6;
[0047] (52) Input the feature vector obtained by integrating H2 and H6 into the convolutional layer of the CNN classification model. Convolve the feature map of the previous layer with the convolutional kernel and add the corresponding correction bias b1 as the correction hyperparameter of the weight;
[0048] (53) Through the relevant operations of the hidden layer activation function, output the feature map. Use the Leaky-ReLU activation function as the activation function of the hidden layer. The following formula shows that Leaky-ReLU assigns a non-zero slope to all negative values:
[0049]
[0050] where a i is a fixed hyperparameter, and i represents a corresponding to the i-th feature information i ;
[0051] (54) Define the prediction label set L, use the max pooling layer for processing, and then perform a fully connected operation for text classification L = {label} to obtain the final text classification result S.
[0052] Beneficial effects: Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention is based on the improvement of the ChineseBert model, integrates and embeds pinyin and character vector information, and adds position vector information for integration, and then inputs it into the Bert model for training. The Bert parameters use a sharing mechanism to decouple different semantics belonging to the same character form, obtaining the corresponding context semantic information while saving resource consumption; at the same time, the Word2Vec model is used to process the preprocessed data, and then the cascaded Attention mechanism is used for information learning to obtain the feature information from words to sequences and context associations; finally, the feature vectors of the above two parts are fused and input into the CNN classification model to obtain the final text classification result. Description of the Drawings
[0053] Figure 1 is the flowchart of the present invention;
[0054] Figure 2 is the flowchart of the preprocessing of news text data;
[0055] Figure 3 is the flowchart of feature information extraction of the Word2Vec module and the cascaded Attention mechanism;
[0056] Figure 4 is the flowchart of pinyin and character vector embedding;
[0057] Figure 5 is the flowchart of feature fusion embedding and Bert model training;
[0058] Figure 6 is the flowchart of multi-feature fusion text classification. Detailed Implementation Manner
[0059] The present invention will be further described in detail below with reference to the accompanying drawings.
[0060] The present invention proposes a chemical accident news classification method based on the ChineseBERT model and the attention mechanism, as Figure 1 shown, which specifically includes the following steps:
[0061] The variables involved in the present invention are shown in Table 1:
[0062] Table 1 Variable Description Table
[0063]
[0064]
[0065] Step 1: By traversing and screening the chemical accident news dataset D, the preprocessed chemical accident news set D1 is obtained. AsFigure 2 As shown below, the specific method is as follows:
[0066] Step 1.1: Define the chemical emergency news text dataset as D, define Text as a single text data, and define id, title, and label as the serial number of a single text, the title of the data, and the text label respectively, and satisfy the relationship Text = {id, title, label}, D = {Text1, Text2, …, Text i , …, Text n},Text i is the i-th text information data in D, where n = len(D) is the number of texts in D, and the variable i ∈ [1, n];
[0067] Step 1.2: Define the processed chemical emergency text dataset as D1, D1 = {Text1, Text2, …, Text j , …, Text m},Text j is the j-th text information data in D1, where m = len(D1) is the number of texts in D1 respectively, and the variable j ∈ [1, m];
[0068] Step 1.3: Read the dataset D and traverse the entire dataset;
[0069] Step 1.4: If title == null, execute Step 1.5, otherwise execute Step 1.6;
[0070] Step 1.5: Delete the corresponding row of data;
[0071] Step 1.6: Remove some useless characters according to the stop word list;
[0072] Step 1.7: Save the preprocessed text dataset D1.
[0073] Step 2: Read the preprocessed dataset D1, train it through the word2vec model to obtain text word vectors, use them as the input of the first-level attention mechanism, and then use them as the input of the second-level attention mechanism to obtain the final context feature vectors. As Figure 3 shown below, the specific method is as follows:
[0074] Step 2.1: Read the preprocessed text dataset D1;
[0075] Step 2.2: Define the word feature vector set R1;
[0076] Step 2.3: Perform data word segmentation through the word2vec model, and train it through the word2vec model to obtain text word feature vectors
[0077] Step 2.4: Save the word feature vector R1, and satisfy is the i-th word feature vector in the data vector set, where the variable i ∈ [1, a], and a is the number of word vectors after word segmentation;
[0078] Step 2.5: Define the word dependency feature vector H1 based on the attention mechanism;
[0079] Step 2.6: Input the word feature vector R1 into the Attention mechanism to obtain the attention-based word dependency feature vector where represents the j-th word dependency feature vector in the text, satisfying The variable j ∈ [1, b], and b is the number of word dependency feature vectors. The input and adjustment method of the Attention mechanism is to use softmax normalization to adjust the weight matrix W f Adjustment, and then multiply by V, where, d k is the dimension of a Q and K vector, is the scale scalar factor, and Q, K, V are tensors, representing query, key, and value respectively;
[0080] Step 2.7: Define the loop variable k to learn the word feature vector H1 of the first-level attention mechanism, and the initial value of k is 1;
[0081] Step 2.8: Define the subsequence dependency feature vector H2 based on the attention mechanism;
[0082] Step 2.9: If k ≤ b, then execute Step 2.10, otherwise execute 2.12;
[0083] Step 2.10: Input the word dependency feature vector H1 into the Attention mechanism to obtain the attention-based subsequence dependency feature vector where represents the t-th subsequence dependency feature vector in the text, satisfying The variable t ∈ [1, c], and c is the number of subsequence dependency feature vectors;
[0084] Step 2.11: k = k + 1;
[0085] Step 2.12: Output and save the feature vector H2 of the secondary attention mechanism.
[0086] Step 3: Read the preprocessed news dataset D1, process it with an open-source pinyin package, and then input it into the MLP for vectorization. At the same time, perform one-hot encoding on the news dataset D1, and fuse the obtained character vectors and pinyin vectors through matrix embedding to obtain a 2D matrix vector R3. As Figure 4 shown, the specific method is as follows:
[0087] Step 3.1: Define the pinyin feature vector H3, define the one-hot character vector H4, and define the fusion embedding matrix R3;
[0088] Step 3.2: Read the text data D1 into the open-source pinyin package to obtain the pinyin representation, and input it into the MLP. There are 3 hidden layers in the neural network, with 64 nodes in each hidden layer, and then obtain the pinyin vector through the max pooling layer satisfying is the pinyin vector corresponding to the i-th character in the data vector set, where the variable i ∈ [1, d], and d is the number of pinyin vectors;
[0089] Step 3.3: Read the preprocessed data D1, and obtain the character vector through one-hot encoding of the character vector satisfying is the character feature vector of the j-th character in the data vector set, where the variable j ∈ [1, e];
[0090] Step 3.4: Fuse the pinyin vector H3 and the character vector H4 to obtain the fusion embedding vector Mainly use the fully connected layer with a learnable matrix to induce the embedding of the matrix vector, and fuse the matrix vector where represents the fusion feature vector corresponding to the t-th character in the text, and the variable t ∈ [1, s].
[0091] Step 4: Fuse the matrix feature information R3 with the position vector to obtain the feature information H5, and input it into the Bert model for vectorization training to obtain the final semantic feature information H6. As Figure 5 shown, the specific method is as follows:
[0092] Step 4.1: Define the position vector R4, define the feature vector matrix H5 that fuses the position vector, and define the feature vector H6 after Bert pre-training;
[0093] Step 4.2: Add the fusion matrix vector R3 and the positional Embedding to obtain the integrated feature vector matrix where the variable h ∈ [1, f];
[0094] Step 4.3: Read the integrated feature vector matrix H5 and input it into the Bert model for training to obtain the final feature information vector H6, where is the p-th feature vector of the vector after Bert training, where the variable p ∈ [1, g]. The training parameters of the Bert model are shared to obtain the corresponding training feature vectors.
[0095] Step 5: Integrate the feature information obtained in Steps 2 and 4, perform a fully connected process, input it into the CNN model for classification processing, and obtain the final text classification result. As Figure 6 shown, the specific method is as follows:
[0096] Step 5.1: Read the context feature information H2 and read the semantic information H6;
[0097] Step 5.2: Input the feature vector obtained by integrating H2 and H6 into the convolutional layer (hidden unit) of the CNN classification model, convolve the feature map of the previous layer with the convolutional kernel, and add the corresponding correction bias b1 as the correction hyperparameter of the weight;
[0098] Step 5.3: Through the relevant operations of the hidden layer activation function, output the feature map, and use the Leaky-ReLU activation function as the activation function of the hidden layer. The following formula shows that Leaky-ReLU assigns a non-zero slope to all negative values:
[0099]
[0100] where a i is a fixed hyperparameter, and i represents a i corresponding to the i-th feature information;
[0101] Step 5.4: Define the prediction label set L, perform processing using the max pooling layer, and then perform a fully connected operation for text classification L = {label} to obtain the final text classification result S.
[0102] The present invention can be combined with chemical emergency news to complete the extraction of text context features through learning based on the cascaded Attention mechanism, and use the ChineseBERT pre-trained model to add position information on the basis of using pinyin and character information, and input it into the Bert model for training to obtain the final semantic feature information. The two are fused and embedded for text classification operations through the CNN model. For chemical safety news, according to the classification of emergencies in the "Overall Emergency Plan for National Public Emergencies", a part of them is classified and summarized to obtain the categories of chemical emergency events (such as fires, explosions, flammable, explosive, and toxic gas leaks) for the classification of chemical news emergencies.
[0103] The present invention can be used in aspects such as classification in natural language processing, extraction of feature information, and pre-training of pinyin character information to obtain semantic feature information, as well as classification of various chemical industry news texts.
Claims
1. A chemical emergency news classification method based on the ChineseBERT model and the attention mechanism, characterized in that It includes the following steps: (1) Preprocess the chemical emergency news text dataset D to obtain the preprocessed news text dataset D1; (2) Process D1 through the word2vec model to obtain the word feature vector R1. Input the word feature vector R1 into the WordAttention model to obtain the new word dependency feature vector H1, and then input the word dependency feature vector H1 into the SeqAttention model to obtain the subsequence feature vector H2; (3) Process D1 through the open-source pinyin package to obtain the corresponding pinyin sequence, and then input it into the MLP. After passing through the max pooling layer, output the pinyin vector H3. Perform one-hot encoding on the preprocessed text to obtain the character vector H4, and perform matrix embedding with the pinyin vector H3 to obtain the fusion feature vector R3; (4) Integrate the fusion feature vector R3 with the position feature vector R4 to obtain the feature vector H5, and input H5 into the Bert pre-trained model to obtain the pre-trained semantic feature vector H6; (5) Integrate the subsequence feature vector H2 in step (2) with the semantic feature vector H6 in step (4), and input it into the CNN model to obtain the final text classification result.
2. The chemical accident news classification method based on the ChineseBERT model and the attention mechanism according to claim 1, characterized in that, The step (1) includes the following steps: (11) Define the chemical emergency news text dataset as D, define Text as a single text data, and define id, title, and label as the serial number of a single text, the title of the data, and the text label respectively, and satisfy the relationship Text = {id, title, label}, D = {Text1, Text2, …, Text i , …, Text n}, Text i is the i-th text information data in D, where n = len(D) is the number of texts in D, and the variable i ∈ [1, n]; (12) Define the processed news text dataset as D1, D1 = {Text1, Text2, …, Text j , …, Text m}, where Text j is the j-th text information data in D1. Here, m = len(D1) represents the number of texts in D1, and the variable j ∈ [1, m]; (13) Read D and traverse the entire dataset; (14) If title == null, execute (15), otherwise execute (16); (15) Delete the corresponding row data; (16) Remove useless characters according to the stop word list; (17) Save the preprocessed news text dataset D1.
3. The chemical accident news classification method based on the ChineseBERT model and the attention mechanism according to claim 1, wherein The step (2) includes the following steps: (201) Read the preprocessed news text dataset D1; (202) Define the word feature vector R1; (203)Perform data word segmentation processing through the word2vec model, and obtain text word feature vectors through training with the word2vec model (204) Save the word feature vector R1 and satisfy is the i-th word feature vector in the data vector set, where the variable i ∈ [1, a], and a is the number of word vectors after word segmentation; (205) Define the word dependency feature vector H1 based on the attention mechanism; (206) Input the word feature vector R1 into the Attention mechanism to obtain the attention-based word dependency feature vector Among them represents the j-th word dependency feature vector in the text, satisfying The variable j ∈ [1, b], where b is the number of word dependency feature vectors. The input and adjustment method of the Attention mechanism is to use softmax normalization to adjust the weight matrix W f and then multiply by V where d k is the dimension of a Q and K vector is the scale scalar factor, and Q, K, and V represent query, key, and value respectively (207) Define the loop variable k to learn the word dependency feature vector H1 of the first-level attention mechanism, and the initial value of k is 1; (208) Define the subsequence feature vector H2 based on the attention mechanism; (209) If k ≤ b, execute (210), otherwise execute (212); (210) Input the word dependency feature vector H1 into the Attention mechanism to obtain the attention-based subsequence dependency feature vector where represents the t-th subsequence dependency feature vector in the text, satisfying the variable t ∈ [1, c], where c is the number of subsequence dependency feature vectors; (211) k = k + 1; (212) Output and save the subsequence feature vector H2 of the second-level attention mechanism.
4. The chemical accident news classification method based on the ChineseBERT model and the attention mechanism according to claim 1, wherein The step (3) includes the following steps: (31) Define the pinyin vector H3, define the character vector H4, and define the fusion feature vector R3; (32) Read D1 into the open-source pinyin package to obtain the pinyin representation, and input it into the MLP. There are 3 hidden layers in the neural network, with 64 nodes in each hidden layer, and then obtain the pinyin vector through the max pooling layer Meet is the pinyin vector corresponding to the i-th character in the data vector set, where the variable i ∈ [1, d], and d is the number of pinyin vectors; (33) Read D1, and obtain a character vector by one-hot encoding the character vector Satisfy is the j-th character feature vector in the data vector set, where the variable j ∈ [1, e]; (34)Fuse the pinyin vector H3 and the character vector H4 to obtain a fused embedding vector Mainly use a fully connected layer with a learnable matrix to induce the embedding of the matrix vector and fuse the feature vectors Where represents the fused feature vector corresponding to the t-th character in the text, and the variable t ∈ [1, s].
5. The chemical accident news classification method based on the ChineseBERT model and the attention mechanism according to claim 1, characterized in that The step (4) includes the following steps: (41) Define the position feature vector R4, define the feature vector H5 that fuses the position vector, and define the semantic feature vector H6 after Bert pre-training; (42) Add the fused feature vector R3 to the positional Embedding to obtain the feature vector where the variable h ∈ [1, f]; (43) Input the feature vector H5 into the Bert model for training to obtain the pre-trained semantic feature vector H6, where is the p-th feature vector of the vector after Bert training, where the variable p ∈ [1, g], and the training parameters of the Bert model are shared to obtain the corresponding training feature vector.
6. The chemical emergency news classification method based on the ChineseBERT model and the attention mechanism according to claim 1, wherein The step (5) includes the following steps: (51) Read the subsequence feature vector H2 and read the semantic feature vector H6; (52) Input the feature vector obtained by integrating H2 and H6 into the convolutional layer of the CNN classification model, perform convolution on the feature map of the previous layer with the convolutional kernel, and add the corresponding correction bias b1 as the correction hyperparameter of the weight; (53)Output the feature map through the relevant operations of the hidden layer activation function. Use the Leaky-ReLU activation function as the activation function of the hidden layer. The following formula shows that Leaky-ReLU assigns a non-zero slope to all negative values: where a i is a fixed hyperparameter, and i represents a i corresponding to the i-th feature information; (54)Define the prediction label set L, process it using the max pooling layer, and then perform a fully connected operation for text classification L = {label}, obtaining the final text classification result S.
Citation Information
Patent Citations
Entity relation joint extraction method and system based on active deep learning
CN113901825A
Method and apparatus for performing entity linking
US20210406706A1