A Knowledge Entity Recognition Method and System for the Chinese Industrial Field

By manually labeling characteristic industrial text corpus and data enhancement, combined with BERT and expanded convolutional neural network, the problem of lack of corpus labeling sets in small Chinese industries is solved, and efficient recognition of knowledge entities and improved model training capabilities are achieved.

CN114780729BActive Publication Date: 2025-06-13HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210463257.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-28
Publication Date
2025-06-13
Estimated Expiration
2042-04-28

AI Technical Summary

Technical Problem

The knowledge entity recognition construction method used in the prior art for small Chinese fields industry lacks a public production-related corpus labeling set, resulting in insufficient model training and prediction capabilities.

Method used

Characteristic industrial text corpus is used for manual annotation, and the training text is multi-featured through two data augmentation methods (entity replacement and entity position exchange). Then, the processed text is input into the BERT long sentence multi-feature embedding layer, and the combined text embedding vector is formed by superposition to find the mean, and the expanded convolutional neural network and attention mechanism are trained.

Benefits of technology

It realizes efficient identification of knowledge entities in small Chinese industries, solves the problem of lack of labeled data, and improves the training and prediction capabilities of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114780729B_ABST
    Figure CN114780729B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for knowledge entity recognition in the Chinese industrial field. The present invention uses characteristic industrial text corpus and makes manual annotation on it as the input of the model. After multi-feature processing of the training text by two different data augmentation methods, it is respectively input into the word embedding layer. For different embedded feature vectors, the mean value is obtained by superposition. A convolutional neural network is adopted, and in combination with the advantage of processing long texts, a dilated convolutional neural network is used, and attention is added after the convolutional layer to set training weights. Among them, the network structure of the improved dilated convolution is used to better expand the receptive field when extracting features of long texts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer artificial intelligence, and particularly to a method and system for identifying knowledge entities for the Chinese industrial field. Background Art

[0002] In recent years, with the continuous development of artificial intelligence technology, intelligent computers have been integrated into all aspects of life, which has of course promoted the informatization construction work of the manufacturing industry. With the continuous development of industrial automation in industrial workshops, industrial automation has become increasingly important. At the same time, the amount of data generated in industries related to industrial production is countless. Among them, manufacturing industry data is a complete and comprehensive record of industrial supply and production processes, containing a large amount of information. People have begun to use natural language processing technology to mine this manufacturing industry data information, so as to obtain industrial knowledge that is structured and closely related to industrial production and manufacturing supply.

[0003] Named Entity Recognition (NER) refers to identifying specific entities in text, and commonly used ones include person names, place names, organization names, etc. In the industrial field, it aims to automatically identify, classify, and process entities in the process of industrial production and development, such as parts, technical means, etc. NER is the basis for structuring industrial data and the premise for conducting research on industrial data text. Due to the complexity of Chinese text processing, NER for Chinese text is more difficult than that for English text. Currently, the commonly used methods for named entity recognition mainly include: dictionary and rule-based methods, traditional machine learning-based methods, and deep learning-based methods.

[0004] The dictionary-based method searches for strings by fuzzy matching or exact matching, but it cannot retrieve entities that do not exist in the dictionary. The rule-based method formulates a rule set based on entity features and their common collocations, but it takes a long time, requires domain experts to write rules, and cannot be applied to new domains.

[0005] In recent years, with the development and application of machine learning technology, machine learning-based methods have gradually become the mainstream methods. Although this method has strong portability, it depends on the quality and scale of labeled data, and the feature engineering is complex. With the further development of machine learning, deep learning-based methods have received further attention. Although this method no longer requires artificially selecting a complex feature set as the model training set like traditional machine learning methods, it requires a larger-scale corpus.

[0006] Since each manufacturing factory has different manufactured products, manufacturing processes, manufacturing industries, and supply chains, each physical factory is unique, and there is also a lack of publicly available corpus annotation sets related to production and manufacturing. Therefore, the promotion of intelligent industry by applying different manually annotated sets for different intelligent manufacturing factories is quite significant. Summary of the Invention

[0007] The technical problem to be solved by the present invention is that in the prior art, the knowledge entity recognition and construction method for the industrial field of Chinese small domains lacks publicly available corpus annotation sets related to production and manufacturing due to the different characteristics of manufacturing industry data. The present invention proposes a knowledge entity recognition and construction method and system for the industrial field of Chinese small domains. The present invention uses characteristic industrial text corpus and makes manual annotations on it as the input of the model. After multi-feature processing of the training text by using two different data augmentation methods, it is respectively input into the word embedding layer. For different embedding feature vectors, the mean value is obtained by superposition. A convolutional neural network is used, and in combination with the advantage of processing long texts, a dilated convolutional neural network is combined, and attention is added after the convolutional layer to set training weights. Among them, the network structure of the improved dilated convolution is used to better expand the receptive field when extracting features of long texts.

[0008] To achieve the above object, the technical solution of the present invention is as follows: In the first aspect, the present invention provides a knowledge entity recognition method for the industrial field of Chinese, including the following steps:

[0009] Step S1: Collect sentence texts in the industrial field and preprocess the text data, including: selecting literature and periodicals in the specified industrial field to construct a data set, extracting the text in the data set by paragraph, splitting the text paragraphs into sentences, and cleaning the text sentences;

[0010] Step S2: Perform label annotation on the text preprocessed in Step 1; specific label classification is given according to actual needs; data augmentation is performed on the annotated text data, and the augmentation methods are divided into two categories: entity replacement and entity position exchange;

[0011] Step S3: Input the original text after annotation and its two types of enhanced new texts into the named entity recognition model. The named entity recognition model includes a BERT long sentence multi-feature embedding layer, an entity label training layer, an attention layer, a bidirectional long short-term memory network layer, and a conditional random field connected in sequence;

[0012] The BERT long sentence multi-feature embedding layer outputs three types of long sentence text embedding vectors respectively, and the mean value is obtained by superposition of the three types of embedding vectors to form a combined text embedding vector;

[0013] Input the combined text embedding vector into the entity label training layer to train the labels; the entity label training layer is composed of 4 dilated convolution modules, and each module is provided with 1 convolutional network and 2 two-dilated convolution neural networks (2-dilated convolution neural network);

[0014] After performing weighted emphasis training on the label training results through the attention layer, output them to the bidirectional long short-term memory network layer and the conditional random field for label prediction, and train the named entity recognition model to obtain a trained named entity recognition model;

[0015] Step S4: After preprocessing the sentence text in the Chinese industrial field to be recognized, input it into the trained named entity recognition model for entity recognition.

[0016] Further, in the step S1, the following sub-steps are included:

[0017] S11, search for relevant analysis literature on industrial data and intelligent factories, and correspondingly intercept text paragraphs with relevance and containing industrial entities in related fields;

[0018] S12, clean the sentences to remove punctuation, arrange the selected text paragraphs with one character per line plus a space character and "\n", and replace the full stop character with "\n" to represent sentence segmentation for subsequent sentence input into the network.

[0019] Further, in the step S2, the following sub-steps are included:

[0020] S21, classify the industrial field text data entities, and the specific categories depend on specific requirements. After classification, label the entity category labels for the entities in the text data;

[0021] S22, perform data augmentation on the text sentences with completed annotation. Use the nlpcda toolkit to perform two different text entity augmentation methods on the text, including entity replacement and entity position exchange. Among them, entity replacement randomly replaces the entities in the text sentence with other entities included in the dataset; entity position exchange exchanges the positions of multiple entities in the sentence in space without changing the entities.

[0022] Further, in the step S21, the industrial field text data categories include three types: physical objects, technologies, and concepts.

[0023] Further, in the step S3, the following sub-steps are included:

[0024] S31. Convert each type of tag for the entities in the text using the BIOE tagging method to generate annotation information. In the annotated industrial text data information, each character corresponds to an annotated BIOE.

[0025] S32. Use the annotation information as the named entity recognition tags for the industrial text data, and construct a sample data set with named entity recognition tags.

[0026] S33. Input the original text and the two types of texts after text augmentation into the BERT long sentence multi-feature embedding layer sentence by sentence. In the high-dimensional feature vector sequence output by the BERT long sentence multi-feature embedding layer, each character corresponds to a feature vector.

[0027] S34. Stack and average the combined text embedding vectors output by the BERT long sentence multi-feature embedding layer, and then use the processed result vector as the input to the entity label training layer to perform global feature extraction on the combined text embedding vectors. Input the text features trained by the entity label training layer into the attention layer for weight training, train different weight matrices for different text features to form preferences for the features, and after the attention-based weight training, input them into the bidirectional long short-term memory network layer and the conditional random field for further feature refinement to obtain the label probabilities corresponding to the text entities.

[0028] Further, the attention layer calculates the attention weight a t according to the output vector h of the BERT long sentence multi-feature embedding layer i , and the output vector x k of the attention layer is:

[0029] x k =∑ i (a i ·h t ).

[0030] Further, the entity label training layer calculates the loss value based on the predicted label and the entity type annotation corresponding to the predicted label, and updates the parameters using the backpropagation algorithm and the gradient descent algorithm.

[0031] The present invention also provides a knowledge entity recognition system for the Chinese industrial field, which includes: an industrial data text processing unit, a text embedding unit, an entity label training unit, and an entity label prediction unit;

[0032] The industrial data text processing unit is used to clean the selected industrial literature text into a unified format, perform label annotation on the text, and then perform data augmentation on the annotated text data. The augmentation methods are divided into 2 categories: entity replacement and entity position exchange;

[0033] The text embedding unit is used to input the original text with completed annotation and its two types of enhanced new texts into the BERT long sentence multi-feature embedding layer, and output 3 types of long sentence text embedding vectors respectively. The 3 types of embedding vectors are superimposed and averaged to form a combined text embedding vector;

[0034] The entity label training unit is used to input the combined text embedding vector into the entity label training layer for training, and introduce attention to train the preference of the weight of the text vector;

[0035] The entity label prediction unit is used to input the text vector output by the entity label training unit into the bidirectional long short-term memory network layer and the conditional random field for label prediction to obtain industrial entities in the text.

[0036] In a second aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the knowledge entity recognition method for the Chinese industrial field is implemented.

[0037] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the knowledge entity recognition method for the Chinese industrial field is implemented.

[0038] The beneficial effects of the present invention: The knowledge entity recognition construction method and system for the Chinese small industrial field provided by the method embodiment of the present invention independently select characteristic industrial text corpora and make manual annotations on them as the input of the model. The most classic convolutional neural network in the deep neural network is adopted, and the dilated convolutional neural network is combined for the advantage of processing long texts. The network structure of the dilated convolution is improved by using the supervised learning network model to better expand the receptive field when extracting the features of long texts, realizing the named entity recognition in the small industrial field, solving the problem that there is a general lack of labeled data in the industrial field, and improving the model training and prediction capabilities. Description of the Drawings

[0039] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0040] Figure 1 It is a flowchart of the knowledge entity recognition construction method for the Chinese small industrial field.

[0041] Figure 2Schematic diagram of data augmentation for building a method for knowledge entity recognition for Chinese small-domain industries.

[0042] Figure 3 A network diagram of the method for building knowledge entity recognition for Chinese small-domain industries. DETAILED DESCRIPTION

[0043] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0044] Finally, it is explained that the above are the preferred specific implementation modes of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the field can easily modify or replace the technical solutions of the present invention within the technical scope disclosed by the present invention, which should be included in the scope of the claims of the present invention.

[0045] like Figure 1 , 3 As shown, the present invention proposes a knowledge entity recognition construction method for Chinese small field industry, which includes the following steps:

[0046] Step 1: Collect sentence texts in the industrial field and preprocess the text data, including: selecting literature journals in the specified industrial field to build a data set, extracting texts in the data set by paragraphs, segmenting text paragraphs by sentences, and cleaning text sentences; the details are as follows:

[0047] 1.1. Find relevant analytical literature on industrial data and smart factories, and correspondingly intercept text paragraphs that are relevant and contain industrial entities in related fields; 1.2. Clean sentences and remove punctuation. Arrange the selected text paragraphs with one word per line plus a space character and "\n", and replace the period character with "\n" to represent a sentence, so that the subsequent sentence can be input into the network;

[0048] 1.3, the plain text sample data that has not been labeled by entity recognition classification is used to form the unlabeled sample data set.

[0049] Step 2: Label the text preprocessed in step 1; the specific label classification is given according to actual needs; data enhancement is performed on the labeled text data, and the enhancement methods are divided into two categories: entity replacement and entity position exchange; specifically:

[0050] 2.1, classify the text data entities in the industrial field. The specific categories depend on the specific needs. After classification, the entities in the text data are labeled with entity categories;

[0051] 2.2. Perform data augmentation on the text sentences that have been marked. Use the nlpcda toolkit to perform two different text entity augmentation methods on the text, including entity replacement and entity position swapping. Among them, entity replacement randomly replaces the entities in the text sentence with other entities included in the dataset; entity position swapping is to spatially swap multiple entities in the sentence without changing the entities.

[0052] Step 3: Input the original text that has been marked and its two types of augmented new texts into the named entity recognition model. The named entity recognition model includes a BERT long sentence multi-feature embedding layer, an entity label training layer, an attention layer, a bidirectional long short-term memory network layer, and a conditional random field connected in sequence;

[0053] The BERT long sentence multi-feature embedding layer outputs 3 types of long sentence text embedding vectors respectively, and the 3 types of embedding vectors are superimposed and averaged to form a combined text embedding vector;

[0054] Input the combined text embedding vector into the entity label training layer to train the labels; the entity label training layer is composed of 4 dilated convolution modules, and each module is provided with 1 convolution network and 2 two-dilated convolution networks (2-dilated convolution);

[0055] After weight focusing training on the label training results through the attention layer, output to the bidirectional long short-term memory network layer and the conditional random field for label prediction to extract entities. The specific process of this step is as follows:

[0056] 3.1. Perform conversion on each type of label for the entities of the text based on the BIOE label marking method to generate annotation information; in the annotated industrial text data information, each character corresponds to an annotated BIOE; for example, corresponding to the B annotation or I annotation or O annotation or E annotation in the BIOE annotation;

[0057] 3.2. Use the annotation information as the named entity recognition label of the industrial text data to construct a sample dataset with named entity recognition labels;

[0058] 3.3. Input the original text and its two types of texts after text augmentation into the BERT long sentence multi-feature embedding layer sentence by sentence; in the high-dimensional feature vector sequence output by the BERT long sentence multi-feature embedding layer, each character corresponds to a feature vector;

[0059] Use the token embedding layer to read the text, add additional tokens at the beginning [CLS] and end [SEP] of the tokens, and convert the words into fixed vector representations, where each word is represented as a 768-dimensional vector, and the longest sentence length is set to 256, and the insufficient tokens are filled with [PAD];

[0060] Use the segment embedding layer to read the text and use the [SEP] tag to separate the text into sentences;

[0061] Use the position embedding layer to read the text, understand the order of characters in the text, and use vectors to represent the position of each word to represent the sequential features of the input sequence. BERT is designed to process input sequences with a maximum length of 256. The position embedding layer is a lookup table of size (256, 768), where the first row is the vector representation of any word in the first position, and the second row is the vector representation of any word in the second position.

[0062] The input text sequence has three different representations: vector representation of words: token embedding, shape (256, 768); vector representation to help BERT distinguish between pairs of input sequences, segment embedding, shape (256, 768); text input has time attributes, position embedding, shape (256, 768);

[0063] Sum the three representations and output a high-dimensional vector with a shape of (256, 768);

[0064] 3.4. The combined text embedding vectors output by the BERT long sentence multi-feature embedding layer are superimposed and averaged, and then the processed result vector is used as the input of the entity label training layer to perform global feature extraction on the combined text embedding vector; the text features trained by the entity label training layer are input into the attention layer for weight training, and different weight matrices are trained for different text features to form feature preferences. After the attention weight training, they are input into the bidirectional long short-term memory network layer and the conditional random field for further feature extraction to obtain the label probability corresponding to the text entity.

[0065] The attention layer is based on the output vector h of the BERT long sentence multi-feature embedding layer t And the attention mechanism calculates the attention weight a i , the output vector x of the attention layer k for:

[0066] x k =∑ i (a i ·h t ).

[0067] The entity label training layer calculates the loss value according to the predicted label and the entity type annotation corresponding to the predicted label, and updates the parameters using the back propagation algorithm and the gradient descent algorithm to obtain a trained named entity recognition model.

[0068] Step S4: After preprocessing the sentence text in the Chinese industrial field to be recognized, input it into the trained named entity recognition model for entity recognition.

[0069] The present invention also provides a knowledge entity recognition system for the Chinese industrial field, which includes: an industrial data text processing unit, a text embedding unit, an entity label training unit, and an entity label prediction unit;

[0070] The industrial data text processing unit is used to clean the selected industrial literature text into a unified format, label the text, and then perform data augmentation on the labeled text data. The augmentation methods are divided into two categories: entity replacement and entity position exchange;

[0071] The text embedding unit is used to input the original text after annotation and its two types of augmented new texts into the BERT long sentence multi-feature embedding layer, respectively output three types of long sentence text embedding vectors, and perform superposition and mean calculation on the three types of embedding vectors to form a combined text embedding vector;

[0072] The entity label training unit is used to input the combined text embedding vector into the entity label training layer for training, and introduce attention to train the preference for the weights of the text vectors;

[0073] The entity label prediction unit is used to input the text vector output by the entity label training unit into the bidirectional long short-term memory network layer and the conditional random field for label prediction to obtain the industrial entities in the text.

[0074] Among them, the named entity recognition model is trained according to the industrial journal text data set with named entity recognition labels and the unlabeled text data set.

[0075] Figure 2 It is a schematic diagram of text data augmentation. The nlpcda package is used to perform data augmentation on the labeled text. The augmentation methods are divided into two categories: entity replacement and entity position exchange. The combined text data composed of two sets of augmented text data and one set of original text data is respectively input into the embedding layer for word vector conversion, and the three sets of output word vectors are accumulated and averaged to obtain the word features of the combined vector.

[0076] Figure 3 It shows a schematic diagram of the network structure of the above-mentioned knowledge entity recognition construction method and system for the Chinese small-field industry. As Figure 3As shown, assume that the industrial data information to be recognized is "hydraulic support push control system using sensors". Without considering the input format requirements of the named entity recognition model, the input sentence for the named entity recognition model is: "[CLS]Hydraulic support retires the control system using sensors[SEP]". The predicted result output by the named entity recognition model is: [CLS][O]Liquid[B-Ent]Pressure[I-Ent]Support[I-Ent]Frame[E-Ent]Utilize[O]Use[O]Transmission[B-Ent]Sensor[I-Ent]Device[E-Ent]Push[O]Shift[O]Control[B-Tec]System[I-Tec]Series[I-Tec]System[E-Tec][SEP][O]. Based on this predicted result, the named entities (entities) can be obtained, which are: Entity: Hydraulic support, Sensor; Technology: Control system.

[0077] The above embodiments are used to explain the present invention, rather than limit the present invention. Any modifications and changes made to the present invention within the spirit and scope of the claims of the present invention fall within the protection scope of the present invention.

Claims

1. A method for identifying knowledge entities in the Chinese industrial field, characterized in that, it includes the following steps: Step S1: Collect sentence texts in the industrial field and preprocess the text data, including: selecting literature journals in the specified industrial field to construct a data set, extracting text by paragraphs in the data set, splitting text paragraphs into sentences, and cleaning text sentences; Step S2: Perform label annotation on the text preprocessed in Step 1; specific label classification is given according to actual needs; perform data enhancement on the labeled text data, and the enhancement methods are divided into two categories: entity replacement and entity position exchange; Step S3: Input the original text after annotation and its two types of enhanced new texts into the named entity recognition model. The named entity recognition model includes a BERT long sentence multi-feature embedding layer, an entity label training layer, an attention layer, a bidirectional long short-term memory network layer, and a conditional random field connected in sequence; The BERT long sentence multi-feature embedding layer outputs three types of long sentence text embedding vectors respectively, and the three types of embedding vectors are superimposed and averaged to form a combined text embedding vector; Input the combined text embedding vector into the entity label training layer to train the labels; the entity label training layer is composed of 4 dilated convolution modules, and each module is provided with 1 convolutional network and 2 dilated convolutional networks; After performing weighted emphasis training on the label training results through the attention layer, output them to the bidirectional long short-term memory network layer and the conditional random field for label prediction. After training the named entity recognition model, obtain the trained named entity recognition model; Step S4: Preprocess the sentence text in the Chinese industrial field to be recognized, and input it into the trained named entity recognition model for entity recognition.

2. The method for identifying knowledge entities in the Chinese industrial field according to claim 1, characterized in that, in the said Step S1, it includes the following sub-steps: S11, search for relevant analysis literature on industrial data and intelligent factories, and correspondingly intercept text paragraphs with relevance and containing industrial entities in related fields; S12, clean the sentences to remove punctuation, arrange the selected text paragraphs with one character per line plus a space character and "\n", and replace the full stop character with "\n" to represent sentence splitting for subsequent sentence input into the network.

3. The method for identifying knowledge entities in the Chinese industrial field according to claim 1, characterized in that, in the said Step S2, it includes the following sub-steps: S21, classify the entity of the industrial field text data, and the specific categories depend on specific needs. After classification, perform entity category label annotation on the entities of the text data; S22, perform data enhancement on the text sentences that have been labeled. Use the nlpcda toolkit to perform two different text entity enhancement methods on the text, including entity replacement and entity position exchange. Among them, entity replacement randomly replaces the entities in the text sentences with other entities contained in the data set; entity position exchange exchanges the positions of multiple entities in the sentence in space without changing the entities.

4. The method for identifying knowledge entities in the Chinese industrial field according to claim 3, characterized in that, In the step S21, the text data categories in the industrial field include three types: physical objects, technologies, and concepts.

5. The method for identifying knowledge entities in the Chinese industrial field according to claim 1, characterized in that, in the step S3: S31, convert each type of label for the entities in the text based on the BIOE tagging method to generate annotation information; in the annotated industrial text data information, each character corresponds to an annotated BIOE; S32, use the annotation information as the named entity recognition label for the industrial text data, and construct a sample data set with named entity recognition labels; S33, input the original text and the two types of texts enhanced by text enhancement into the BERT long sentence multi-feature embedding layer sentence by sentence; in the high-dimensional feature vector sequence output by the BERT long sentence multi-feature embedding layer, each character corresponds to a feature vector; S34, perform superposition and averaging processing on the combined text embedding vectors output by the BERT long sentence multi-feature embedding layer, and then use the processed result vector as the input of the entity label training layer to perform global feature extraction on the combined text embedding vectors; input the text features trained by the entity label training layer into the attention layer for weight training, train different weight matrices for different text features to form preferences for the features, and after the attention-based weight training, input them into the bidirectional long short-term memory network layer and the conditional random field for further feature refinement to obtain the label probability corresponding to the text entity.

6. The method for identifying knowledge entities in the Chinese industrial field according to claim 1, characterized in that, The attention layer calculates the attention weight a based on the output vector h of the BERT long sentence multi-feature embedding layer t and the attention mechanism i , and the output vector x of the attention layer k is as follows: x k = ∑ i (a i ·h t ).

7. The method for identifying knowledge entities in the Chinese industrial field according to claim 1, characterized in that, the entity label training layer calculates the loss value according to the predicted label and the entity type annotation corresponding to the predicted label, and updates the parameters using the backpropagation algorithm and the gradient descent algorithm.

8. A knowledge entity recognition system for the Chinese industrial field, characterized in that, the system includes: an industrial data text processing unit, a text embedding unit, an entity label training unit, and an entity label prediction unit; the industrial data text processing unit is used to clean the selected industrial literature text into a unified format, perform label annotation on the text, and then perform data enhancement on the annotated text data. The enhancement methods are divided into 2 categories: entity replacement and entity position exchange; the text embedding unit is used to input the original text that has been annotated and its two types of enhanced new texts into the BERT long sentence multi-feature embedding layer, respectively output 3 types of long sentence text embedding vectors, and perform superposition and averaging on the 3 types of embedding vectors to form a combined text embedding vector; the entity label training unit is used to input the combined text embedding vector into the entity label training layer for training, and introduce attention to train the preferences for the weights of the text vectors; the entity label prediction unit is used to input the text vector output by the entity label training unit into the bidirectional long short-term memory network layer and the conditional random field for label prediction to obtain the industrial entities in the text.

9. An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, Characterized in that, when the processor executes the program, it implements the knowledge entity recognition method for the Chinese industrial field according to any one of claims 1 to 7.

10. A computer-readable storage medium, on which a computer program is stored, Characterized in that, when the computer program is executed by a processor, it implements the knowledge entity recognition method for the Chinese industrial field according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Judicial text named entity recognition-oriented method and system

    CN113869053A

  • Coal mine safety field-oriented entity recognition method

    CN113988054A