Text entity nesting method based on neural network

Through the text nesting entity method based on neural networks, combined with field dictionaries and multiple neural network models, the problem of fragmentation and entity nesting extraction of equipment maintenance guarantee business data information is solved, and efficient standardization conversion of text information and effective construction of big data resources are realized.

CN120045717APending Publication Date: 2025-05-27DEPARTMENT OF JOINT EQUIPMENT SUPPORT JOINT SERVICE COLLEGE NATIONAL DEFENSE UNIVERSITY PEOPLES LIBERATION ARMY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510219139.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Due to the traditional and backward data management methods, data information is seriously fragmented, especially in the text, there is serious entity nesting phenomenon, which is difficult to extract efficiently.

Method used

The text nested entity method based on neural network is adopted, combined with domain dictionary, sequence annotation, feature extraction and model training, text features are extracted through the Bert model, BiLSTM and TextCNN models, and label prediction is used for label prediction to achieve efficient extraction of nested entities.

Benefits of technology

Effectively support the standardization and standardization conversion of equipment maintenance guarantee text information, support the construction of big data resource system, and realize the transformation from data advantages to application advantages and decision-making advantages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045717A_ABST
    Figure CN120045717A_ABST
Patent Text Reader

Abstract

The invention relates to a neural network-based text entity nesting method, which comprises the following steps of: S1, determining a domain dictionary, and performing word segmentation processing on text information according to the domain dictionary to obtain a word segmentation sequence; s2, inputting the word segmentation sequence into a Bert model, and converting the word segmentation sequence into a characteristic vector sequence; s3, respectively inputting the characteristic vector sequences into a BiLSTM model and a TextCNN model, respectively calculating semantic vector sequences corresponding to the characteristic vector sequences in the two models, and fusing the two semantic vector sequences to obtain a characteristic vector sequence; s4, the feature vector sequence passes through a softmax classifier, and the label prediction probability of each sequence fragment is obtained through calculation. The method effectively supports conversion and loading of equipment maintenance support information to standardized and normalized value information, supports construction of an equipment maintenance support big data resource system, and provides a new technical means for conversion of equipment maintenance support from a data advantage and an information advantage to an application advantage and a decision advantage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data processing, in particular to an equipment maintenance and support big data ETL (Extract Transform Load) technology, and in particular to a text nested entity method based on a neural network. Background Art

[0002] At present, equipment maintenance and support business data are scattered in different equipment management departments, equipment using troops, etc., and due to the relatively traditional and backward data management methods of the troops, these data information are usually recorded in text form, which directly leads to serious fragmentation of equipment maintenance and support business data information. In particular, there is a serious phenomenon of entity nesting in the equipment maintenance and support text, for example: "The Army Equipment Department has the Supply Support Bureau, Maintenance Support Bureau, Aviation Equipment Bureau and other units." Due to the complexity and high professionalism of equipment maintenance and support business, the extraction of text nested entities is more challenging, and has become a key issue hindering the construction and development of the equipment maintenance and support data resource system. Summary of the invention

[0003] The embodiment of the present invention provides a text nested entity method based on a neural network, which aims to efficiently extract nested entities in equipment maintenance and support text records in accordance with relevant standards and system establishment of equipment maintenance and support through a neural network, in combination with methods such as sequence labeling, feature extraction and model training, thereby effectively supporting the conversion and loading of equipment maintenance and support text information into standardized and normalized value information, supporting the construction of a big data resource system for equipment maintenance and support, and realizing the transformation of equipment maintenance and support from data advantages and information advantages to application advantages and decision-making advantages.

[0004] In order to achieve the above objectives, the embodiments of the present invention provide the following technical solutions:

[0005] A text embedding entity method based on a neural network comprises the following steps:

[0006] S1. Determine the domain dictionary, perform word segmentation on the text information according to the domain dictionary, and obtain a word segmentation sequence;

[0007] S2, input the word segmentation sequence into the Bert model to convert the word segmentation sequence into a feature vector sequence;

[0008] S3, input the feature vector sequence into the BiLSTM model and the TextCNN model respectively, calculate the semantic vector sequence corresponding to the feature vector sequence in the two models respectively, and then fuse the two semantic vector sequences to finally obtain the feature vector sequence;

[0009] S4. Pass the feature vector sequence through the softmax classifier to calculate the label prediction probability of each sequence segment in the feature vector sequence.

[0010] Furthermore, in S1, the text information is segmented using words as basic units.

[0011] Furthermore, the word segmentation sequence is filtered in S1, wherein when the scanning window size is 1, non-noun sequences obtained by scanning are filtered out; when the scanning window size is greater than 1, sequences with specific name suffixes and sequences starting or ending with adverbs obtained by scanning are filtered out.

[0012] Furthermore, in S3, the characteristic vector sequence E bert ={E 1 , E 2 , E 3 …, E N} is input into BiLSTM, and the forward LSTM output and the reverse LSTM output Splicing That is, the semantic vector sequence of the BiLSTM module is H out =[σ 1 , σ 2 , σ 3 ,…,h t ]; In this technology, the LSTM unit controls the feature vector sequence through the input gate, output gate, and forget gate, and its calculation process is:

[0013] f t =σ[W f (σ t-1 +x t )+b f ]

[0014] i t =σ[W i (h t-1 +x t )+b i ]

[0015] o t =σ[W o (h t-1 +x t )+b o ]

[0016] c t =tanh[W c (h t-1 +x t )+b c ]

[0017] Ct =f t ·C t-1 +i·C t

[0018] σ t =o t tanh(C t )

[0019] In the formula, f t 、i t , o t They represent forget gate, input gate, and output gate respectively;

[0020] c t represents the excess vector;

[0021] C t represents the unit state at time t;

[0022] h t is a hidden state;

[0023] W * and b * represents the weight matrix and bias;

[0024] σ and tanh are nonlinear activation functions.

[0025] Furthermore, in S3, the characteristic vector sequence E bert ={E 1 , E 2 , E 3 ,…,E N} is input into the TextCNN model, where E bert It is expressed as:

[0026]

[0027] In the formula, Represents vector concatenation, and its calculation process is as follows:

[0028] c i = relu(W·E ii+h-1 +b)

[0029] C=[c 1 , c 1 , c 1 …c N-h+1 ] h∈[1,N]

[0030]

[0031] Where W represents the weight matrix of the convolution kernel,

[0032] h represents the height of the convolution kernel;

[0033] b represents bias;

[0034] Ci represents the value obtained by the i-th convolution of the convolution kernel;

[0035] C represents the feature vector obtained by convolution,

[0036] Perform the maximum pooling on C to get q represents the number of convolution kernels.

[0037] Furthermore, the fusion of the two semantic vector sequences in S3 includes fusing the semantic vector sequence H output by the BiLSTM model out =[h 1 ,h 1 ,h 1 …h t ] and the semantic vector sequence output by the TextCNN model Splicing to get the feature vector sequence H out =[H out ·C out ].

[0038] Furthermore, after S4, the true label y of the sequence segment is one-hot encoded, and the training objective function is defined as the cross entropy loss between the label prediction probability and the true label. The calculation process formula is:

[0039]

[0040] Among them, y i Represents h′ i The corresponding true labels are 1 for the positive class and 0 for the negative class.

[0041] The embodiments of the present invention have the following advantages:

[0042] The present invention proposes a text embedding entity method based on neural network. The method first combines the domain dictionary to segment the equipment maintenance and support text information to obtain the word sequence; secondly, the sliding window is used to scan the sampled word sequence to obtain the sequence fragment, and then the candidate entity sequence is screened by limiting the size of the scanning window and formulating the filtering rules, so as to reduce the number of negative samples and improve the model recognition efficiency; then the classification model uses the Bert model as the text representation layer, and extracts the text feature vector by the fusion of BiLSTM and TextCNN neural network; finally, the sequence entity label is decoded by softmax. This method verifies the superiority of the model in the fault corpus by comparing with the commonly used baseline model, and proves that the performance of the model using the dual neural network is better than that of the single neural model by ablation experiments, which provides a new technical means for extracting text embedding entities to effectively support the conversion and loading of equipment maintenance and support text information into standardized and normalized value information, support the construction of equipment maintenance and support big data resource system, and realize the transformation of equipment maintenance and support from data advantage and information advantage to application advantage and decision advantage. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the implementation methods of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the implementation methods or the description of the prior art. Obviously, the drawings in the following description are only exemplary, and for ordinary technicians in this field, other implementation drawings can be derived from the provided drawings without creative work.

[0044] The structures, proportions, sizes, etc. illustrated in this specification are only used to match the contents disclosed in the specification so as to facilitate understanding and reading by persons familiar with the technology. They are not used to limit the conditions under which the present invention can be implemented, and therefore have no substantial technical significance. Any structural modification, change in proportion or adjustment of size shall still fall within the scope of the technical contents disclosed in the present invention without affecting the effects and purposes that can be achieved by the present invention.

[0045] Figure 1 A method flow chart of a text nesting entity method based on a neural network provided by an embodiment of the present invention;

[0046] Figure 2 A specific flow chart of word segmentation processing in a text nested entity method based on a neural network provided by an embodiment of the present invention;

[0047] Figure 3 A specific flow chart of word segmentation sequence nested entities in a neural network-based text nested entity method provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0048] The following is a description of the implementation of the present invention by specific embodiments. People familiar with the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0049] like Figure 1-3 As shown, a text nested entity method based on a neural network includes the following steps:

[0050] S1. Determine the domain dictionary, perform word segmentation on the text information according to the domain dictionary, and obtain a word segmentation sequence;

[0051] S2, input the word segmentation sequence into the Bert model, and convert the word segmentation sequence into a feature vector sequence;

[0052] S3, input the feature vector sequence into the BiLSTM model and the TextCNN model respectively, calculate the semantic vector sequence corresponding to the feature vector sequence in the two models respectively, and then fuse the two semantic vector sequences to finally obtain the feature vector sequence;

[0053] S4. Pass the feature vector sequence through the softmax classifier to calculate the label prediction probability of each sequence segment in the feature vector sequence.

[0054] 1. Domain Dictionary

[0055] A domain dictionary is a collection of professional terms and common words required by a specific industry or field. This technology is used to effectively support the conversion and loading of equipment maintenance and support text information into standardized and normalized value information, support the construction of a big data resource system for equipment maintenance and support, and realize the transformation of equipment maintenance and support from data advantages and information advantages to application advantages and decision-making advantages. Therefore, the domain dictionary is a dictionary for the equipment maintenance and support field.

[0056] Methods for constructing domain dictionaries include manual construction, text mining, corpus analysis, machine learning, and other methods. Manual construction refers to manually constructing a domain dictionary through the experience and knowledge of experts; text mining refers to using natural language processing technology to analyze a large amount of text data, extract key words in a certain field, and then construct a domain dictionary; machine learning refers to using machine learning algorithms to train a large amount of data, identify key words and grammatical rules in the field, and then construct a domain dictionary. Determining the domain dictionary is the preliminary preparation work of this technology. As long as the existing methods can construct a domain dictionary, they fall within the meaning of the features of this technology.

[0057] 2. Word segmentation

[0058] The text information is segmented according to the determined domain dictionary to obtain a segmentation sequence. This embodiment takes the equipment maintenance and support text record (hereinafter referred to as "text information") as an example to specifically explain the implementation details.

[0059] If the relevant equipment maintenance and support entities are directly identified from all the sequence fragments, the training difficulty of the classification model will increase and the classification efficiency will be reduced. Assuming that the text information is X = {x1, x2, x3, ... xn}, where xi represents the i-th character in the text information and n represents the text length of the text information, the number of sequence fragments of X is n(n+1) / 2. In order to reduce the difficulty of training, this technology uses words as basic units to annotate the sampling sequence. First, the Elasticsearch word segmentation tool of the equipment maintenance and support big data management and analysis platform is used to segment the text information X to obtain the word segmentation sequence C = {c1, c2, c3, ... cm}, where cj represents the jth word in the word segmentation sequence and m represents the number of words in the word segmentation sequence. Using words as basic units increases the granularity of the equipment maintenance and support text and reduces the acquisition density. After word segmentation, the number of sequence fragments of C is m(m+1) / 2, and m must be less than n. For example, after word segmentation, the sequence fragments of "Army Equipment Department has Supply and Support Bureau, Maintenance and Support Bureau, Aviation Equipment Bureau, etc." can be obtained as "Army / Equipment Department / Subordinate / Supply and Support Bureau / Maintenance and Support Bureau / Aviation Equipment Bureau / etc.". Although sampling sequence annotation cannot directly reduce the spatial complexity of model classification, this processing will greatly reduce the number of sequence fragments, and to a certain extent can reduce the model classification pressure and the number of negative samples, providing input for subsequent scanning and filtering of sequence fragments.

[0060] 3. Scanning and filtering of word segmentation sequences

[0061] Since the sequence fragments obtained by sampling sequence annotation still contain a large number of negative samples, if all sequence fragments are classified, it will not only increase the model operation time, but also cause the model to overfit. Therefore, this technology reasonably sets the scanning window size (the number of words in the word segmentation sequence m), and then scans and filters the sequence fragments according to the word parts of speech, grammatical rules, semantic similarity, etc. in the sequence fragments, so as to obtain candidate sequences as the input of the natural language training model.

[0062] Specifically, when the scanning window size is 1, the sequence fragment is a single word, and the possibility that a non-noun word is an equipment maintenance and support entity is 0, so it can be directly filtered out.

[0063] When the scanning window size is greater than 1, the sequence fragment contains words of multiple parts of speech, and filtering rules need to be adopted for further processing. Although the filtering rules here cannot completely cover all non-entity sequence fragments, they can still filter out most of the meaningless sequence fragments. The filtering rules are shown in Table 1.

[0064] Table 1 Filtering rules table

[0065]

[0066] 4. Text representation based on Bert model

[0067] The Bert model transforms the "input word segmentation sequence" into a "feature vector sequence" and has strong semantic representation capabilities. The BERT model (Bidirectional Encoder Representations from Transformers) is a deep learning model for natural language processing (NLP). The core of the BERT model lies in its bidirectional encoding feature, which can simultaneously consider the left and right contexts of words in a sentence, thereby providing richer semantic understanding. The training of the BERT model is divided into two stages: pre-training and fine-tuning. The pre-training uses MLM (Masked Language Model) and NSP (Next SentencePrediction) target losses and is trained on a large amount of unlabeled text data, so that BERT can learn contextualized word representations. Then the fine-tuning stage uses task-specific labeled data to optimize the training objectives of specific tasks, so that the pre-trained BERT model can adapt to specific downstream tasks.

[0068] This technology uses the NSP task (Next Sentence Prediction) to implement the pre-training process of Bert. The NSP task is used to determine whether two input texts are adjacent. The principle of the NSP task is to regard the input as the first sentence, and then add a specific prompt to each candidate category as the second sentence, judge the coherence of the first sentence and the second sentence one by one, and select the second sentence with the highest coherence. Conventional technology usually inserts a classification character [CLS] before the two sentences and a segmentation character [SEP] between the two sentences. The semantic vector C contains the semantic information of the entire input sample and can be used to determine whether the two sentences have a contextual relationship. Since this technology inputs a single text sequence into the Bert model, it is not necessary to insert the segmentation character [SEP] during model training, and the semantic vector output by the Bert model is selected as the input of the subsequent model.

[0069] Specifically, the Bert model is regarded as a nonlinear mapping function from a character sequence to a semantic vector. The output of the Bert model is:

[0070] E bert = Bert(X′)

[0071] In the formula, X′={x′ 1 , x′ 2 , x′ 3 , …x′ N},x′ i Represents the i-th character in the output sequence;

[0072] N represents the length of the input sequence;

[0073] E bert represents the output feature vector;

[0074] d bert Represents the Bert feature dimension.

[0075] 5. BiLSTM model

[0076] The BiLSTM (Bidirectional Long Short-Term Memory) model is a neural network model that combines forward and backward LSTM (Long Short-Term Memory). It is mainly used to process sequence data, so that it can better capture the contextual information in the sequence. The forward LSTM processes data from the beginning to the end of the sequence, while the backward LSTM processes data from the end to the beginning of the sequence. This bidirectional processing mechanism enables BiLSTM to consider the contextual information of the sequence at the same time, thereby improving the prediction accuracy of the model.

[0077] The feature vector sequence E bert ={E 1 , E 2 , E 3 …, E N} is input into BiLSTM, and the forward LSTM output and the reverse LSTM output Splicing That is, the semantic vector sequence of the BiLSTM module is H out =[h 1 ,h 2 ,h 3 ,…,h t ].

[0078] In this technology, the LSTM unit controls the feature vector sequence through the input gate, output gate, and forget gate. The calculation process is:

[0079] f t =σ[W f (ht-1 +x t )+b f ]

[0080] i t =σ[W i (h t-1 +x t )+b i ]

[0081] o t =σ[W o (h t-1 +x t )+b o ]

[0082] c t =tanh[W c (h t-1 +x t )+b c ]

[0083] C t =f t ·C t-1 +i·C t

[0084] h t =o t tanh(C t )

[0085] In the formula, f t 、i t , o t They represent forget gate, input gate, and output gate respectively;

[0086] c t represents the excess vector;

[0087] C t represents the unit state at time t;

[0088] h t is a hidden state;

[0089] W * and b * represents the weight matrix and bias;

[0090] σ and tanh are nonlinear activation functions.

[0091] 6. TextCNN Model

[0092] TextCNN (Text Convolutional Neural Network) is a convolutional neural network (CNN) model for text classification. TextCNN extracts text information features through structures such as word embedding, convolution layer, pooling layer and fully connected layer. Word embedding is used to map words to vectors in high-dimensional space. These vectors can capture the semantic information of words; the convolutional layer is used to slide on the text using convolution kernels of different sizes to capture local features of different lengths. Each convolution kernel corresponds to a feature map, which can capture n-gram features of different sizes; the pooling layer usually uses max pooling to reduce the dimension of the feature map and extract the most important features; the fully connected layer is used to connect the outputs of the convolution layer and the pooling layer to form the final classification result.

[0093] In this technology, the feature vector sequence E bert ={E 1 , E 2 , E 3 ,…,E N} is input into the TextCNN model, where E bert It is expressed as:

[0094]

[0095] In the formula, Represents vector concatenation, and its calculation process is as follows:

[0096] c i = relu(W·E ii+h-1 +b)

[0097] C=[c 1 , c 1 , c 1 …c N-h+1 ] h∈[1,N]

[0098]

[0099] Where W represents the weight matrix of the convolution kernel,

[0100] h represents the height of the convolution kernel;

[0101] b represents bias;

[0102] Ci represents the value obtained by the i-th convolution of the convolution kernel;

[0103] C represents the feature vector obtained by convolution,

[0104] Then perform maximum pooling on C to obtain q represents the number of convolution kernels.

[0105] 7. Semantic vector sequence fusion

[0106] The semantic vector sequence output by the BiLSTM model is fused with the semantic vector sequence output by the TextCNN model, that is, the semantic vector sequence H output by the BiLSTM model is fused with the semantic vector sequence H output by the TextCNN model. out =[h 1 ,h 1 ,h 1 …h t ] and the semantic vector sequence output by the TextCNN model Splicing to get the feature vector sequence H out =[H out ·C out ].

[0107] 8. Label prediction probability

[0108] In order to predict the sequence segment label of the feature vector sequence, this technology uses a softmax classifier to calculate each sequence segment h in the feature vector sequence. i The label probability is P i Indicates h i The label prediction probability is calculated as follows:

[0109] 9. Cross Entropy Loss

[0110] In order to measure the difference between the predicted label and the true label during model training, the true label y of the sequence segment of the feature vector sequence is one-hot encoded, and the training objective function is defined as the cross entropy loss between the label prediction probability and the true label. The calculation process formula is:

[0111]

[0112] Among them, y i Represents h′ i The corresponding true labels are 1 for the positive class and 0 for the negative class.

[0113] The following experimental data shows the difference between this method and existing ideas:

[0114] Step 1: Preparation of experimental data and evaluation indicators

[0115] The experiment was conducted with the text records of equipment maintenance and support of a certain military service as the object, and 850 text information about the organization system of equipment maintenance and support were manually sorted out. For example, "The equipment maintenance department of the XX Group Army of the Eastern Theater Command of the Army is located in XX City, XX Province, China". The Elasticsearch word segmentation tool of the equipment maintenance and support big data management and analysis platform was used to segment and annotate the text content. The results are shown in Table 2.

[0116] Table 2 Schematic diagram of the equipment maintenance support system text annotation

[0117]

[0118] The labeled data is divided into training set, validation set and test set in the ratio of 7:1:2. The amount of entity data in each data set is shown in Table 3.

[0119] Table 3 Number of institutional entities in the dataset

[0120]

[0121] In summary, precision, recall and F1 value are used as evaluation indicators. Precision indicates the percentage of correctly recognized entities to the total number of recognized entities, recall indicates the percentage of recognized entities to the total number of entities, and F1 value is a comprehensive performance of precision and recall. The overall precision and recall of the model are the average of the precision and recall of different types of entities. The calculation formula is:

[0122]

[0123]

[0124] in:

[0125] M represents the number of entity categories;

[0126] TP i is the number of entities of the i-th category that are correctly classified in the recognition results;

[0127] FP i is the number of entities of the i-th category that are misidentified in the classification results;

[0128] FN i is the number of entities of the i-th category that are misclassified as other entities.

[0129] Step 2: Parameter analysis and selection

[0130] The experimental hyperparameters are set using the inherent parameters of the Bert model. The embedding dimension is 768, the character embedding dimension is 128, the BiLSTM hidden layer unit is 256, the model input batch is 16, the learning rate is 1×10-5, and the training round is 50. Three different convolution kernels are set for TextCNN, namely 2×768, 3×768, and 4×768, and the number of each convolution kernel is set to 256. The maximum value of the scanning window is set to 6. The specific implementation parameter settings are shown in Table 4.

[0131] Table 4 Experimental parameter settings

[0132]

[0133]

[0134] Step 3: Comparative experiment

[0135] TextCNN, RCNN, DPCNN, BiLSTM, Transformer and other neural network models are used as baseline models for experimental comparison to fully verify the performance of the model of the present invention. The experimental results are shown in Table 5.

[0136] Table 5 Comparative experimental results

[0137]

[0138] As shown in Table 5 above, using the F1 value as the evaluation criterion for the test set, TextCNN performs the worst in the test set, with an F1 value of only 79.50%, while BiLSTM performs better in the test set, with an F1 value of 90.34%. The F1 values ​​of DPCNN and RCNN are similar, but it can be seen from the overall results that the F1 value of the model proposed in the present invention is the highest, and the recall rate and precision of the model are also relatively high.

[0139] Step 4: Ablation experiment

[0140] The model for extracting nested entities in equipment maintenance and support text based on neural networks is composed of a variety of algorithms and models. Each component has a great influence on the performance of the model. The influence of each model component is tested by ablation method. The ablation test results are shown in Table 6.

[0141] Table 6 Ablation experiment results

[0142]

[0143]

[0144] As shown in Table 6 above, when the Bert embedding layer is removed, the performance of this model is the worst, with the F1 value reduced by 4.78% and the recall rate reduced by 7.69%. However, compared with the single-channel model, it performs better, proving that the dual-channel feature fusion model can better extract the equipment maintenance and support text features. When BiLSTM and TextCNN are removed at the same time, only the Bert embedding layer is used to classify the equipment maintenance and support organization entities, the performance loss of the model is large, the F1 value is reduced by 4.65%, and the recall rate is reduced by 7.52%. After removing TextCNN, the performance of the model is also reduced, with the F1 value reduced by 2.52%. After removing BiLSTM, the F1 value of the model is only reduced by 1.04%, indicating that TextCNN has a greater impact on the performance of the model than BiLSTM.

[0145] Although the present invention has been described in detail above by general description and specific embodiments, it is obvious to those skilled in the art that some modifications or improvements can be made to the present invention. Therefore, these modifications or improvements made without departing from the spirit of the present invention all belong to the scope of protection claimed by the present invention.

Claims

1. A text-embedded entity method based on neural network, characterized in that: The following steps are involved: S1. Determine the domain dictionary, perform word segmentation on the text information according to the domain dictionary, and obtain a word segmentation sequence; S2, input the word segmentation sequence into the Bert model to convert the word segmentation sequence into a feature vector sequence; S3, input the feature vector sequence into the BiLSTM model and the TextCNN model respectively, calculate the semantic vector sequence corresponding to the feature vector sequence in the two models respectively, and then fuse the two semantic vector sequences to finally obtain the feature vector sequence; S4. Pass the feature vector sequence through the softmax classifier to calculate the label prediction probability of each sequence segment in the feature vector sequence.

2. A text-embedded entity method based on a neural network according to claim 1, characterized in that: In S1, the text information is segmented using words as basic units.

3. A text-embedded entity method based on a neural network according to claim 2, characterized in that: In S1, the word segmentation sequence is filtered, wherein when the scanning window size is 1, the non-noun sequence obtained by scanning is filtered out; When the scanning window size is greater than 1, the sequences obtained by filtering out the sequences with specific name suffixes and the sequences starting or ending with adverbs are eliminated.

4. A text-embedded entity method based on a neural network according to claim 1, characterized in that: In S3, the characteristic vector sequence E bert ={E1, E2, E3..., E N } is input into BiLSTM, and the forward LSTM output of the LSTM unit is and the reverse LSTM output Splicing That is, the semantic vector sequence of the BiLSTM module is H out =[h1, h2, h3, …, h t ]; The LSTM unit controls the feature vector sequence through the input gate, output gate, and forget gate, and its calculation process is: f t =σ[W f (h t-1 +x t )+b f ] I t =σ[W i (s t-1 +x t )+b i ] the t =σ[W o (h t-1 +x t )+b o ] c t =tanh[W c (h t-1 +x t )+b c ] C t =f t ·C t-1 +i·C t h t =o t ·tanh(C t ) In the formula, f t 、i t , o t They represent forget gate, input gate, and output gate respectively; c t represents the excess vector; C t represents the unit state at time t; h t is a hidden state; W * and b * represents the weight matrix and bias; σ and tanh are nonlinear activation functions.

5. A text-embedded entity method based on a neural network according to claim 1, characterized in that: In S3, the characteristic vector sequence E bert ={E1, E2, E3, ..., E N } is input into the TextCNN model, where E bert It is expressed as: In the formula, Represents vector concatenation, and its calculation process is as follows: c i =relu(W·E ii+h-1 +b) C=[c1,c1,c1…c N-h+1 ] h∈[1,N] Where W represents the weight matrix of the convolution kernel, h represents the height of the convolution kernel; b represents bias; Ci represents the value obtained by the i-th convolution of the convolution kernel; C represents the feature vector obtained by convolution, Perform the maximum pooling on C to get q represents the number of convolution kernels.

6. A text-embedded entity method based on a neural network according to claim 1, characterized in that: In S3, the two semantic vector sequences are fused, including the semantic vector sequence H output by the BiLSTM model. out =[h1,h1,h1…h t ] and the semantic vector sequence output by the TextCNN model Splicing to get the feature vector sequence H out =[H out ·C out ].

7. A text-embedded entity method based on a neural network according to claim 1, characterized in that: After S4, the true label y of the sequence segment is one-hot encoded, and the training objective function is defined as the cross entropy loss between the label prediction probability and the true label. The calculation process formula is: Among them, y i Represents σ ′ i The corresponding true labels are 1 for the positive class and 0 for the negative class.

Citation Information

Patent Citations

  • Double neural network equipment fault domain entity identification method based on scanning filtering

    CN116011452A

  • BERT and pooling-free convolutional neural network-based text classification method

    CN116340506A