A Chinese Named Entity Recognition Method, Device, and Medium Integrating Multi-Granularity Information

By fusing characters, soft words and radical information, a Chinese named entity recognition model for multi-grained information fusion is constructed, which solves the problem that existing methods are difficult to utilize sequence information and improves the accuracy of named entity recognition.

CN114781380BActive Publication Date: 2025-06-24HARBIN ENG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210277553.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-21
Publication Date
2025-06-24
Estimated Expiration
2042-03-21

AI Technical Summary

Technical Problem

The existing Chinese naming entity recognition method is difficult to make full use of the information in the sequence, is susceptible to word segmentation errors, and fails to effectively utilize word and radical semantic information, resulting in poor recognition effect.

Method used

A Chinese named entity recognition method that fuses multi-grained information is proposed. By obtaining and preprocessing corpus data, extracting characters, soft words and radical pre-trained vectors, performing vector fusion, and building a model that fuses multi-grained information, using BiLSTM and CRF layers for training and recognition.

Benefits of technology

By mining the radical semantic information in the sequence and expanding soft word methods, the word, character and radical information in the sequence are effectively utilized, the accuracy of naming entity recognition is improved and the ability to recognize long entities is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114781380B_ABST
    Figure CN114781380B_ABST
Patent Text Reader

Abstract

The present invention proposes a Chinese named entity recognition method, device, and medium that fuse multi-granularity information. The steps of the method are as follows: (1) Obtain a domain corpus dataset, preprocess the dataset, and divide it into a training set, a test set, and a validation set; (2) Extract character, soft word, and radical-level pre-training vectors from the preprocessed corpus data in (1) and fuse them; (3) Construct a Chinese named entity recognition model that fuses multi-granularity information; (4) Input the data obtained in (2) into the model for training; (5) Use the recognition model obtained in (4) to process and calculate the data to be recognized, and obtain the named entity recognition result. In view of the deficiencies in Chinese named entity recognition, the present invention utilizes the inherent semantic information within characters in the sequence by fusing radical-level information, obtains word-level semantic information using an extended soft word module, and integrates the two into the character embedding vector, thereby improving the accuracy of Chinese named entity recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of named entity recognition, and particularly relates to a Chinese named entity recognition method, device and medium that fuse multi-granularity information. Background Technique

[0002] With the continuous development of the economic level and computer technology, a vast amount of text emerges from the Internet every moment. These texts cover information in various aspects such as society, economy, life, and technology. However, due to their large quantity and diverse types, the information contained in the texts is often difficult to be effectively utilized.

[0003] Named entity recognition (NER) has, to some extent, solved this problem. Its core task is to identify and extract named entities containing key information such as person names, place names, and organization names in a given text. NER is a fundamental task in the field of natural language processing. Many downstream tasks, such as question answering systems, knowledge graph construction, and information extraction, rely on NER. Since Chinese has no obvious word boundaries like spaces, characters are closely arranged, and Chinese NER is usually more difficult than English NER. NER is typically regarded as a sequence labeling problem. Most traditional NER models are linear statistical models, such as hidden Markov models, maximum entropy models, maximum entropy hidden Markov models, conditional random fields, and support vector machines. In recent years, deep learning methods have achieved good results in NER tasks due to their powerful capabilities and have gradually become the mainstream methods for NER. Existing Chinese NER models can be divided into character-based models and word-based models. In word-based models, a Chinese word segmentation system is first required to segment the input sequence, which is then used as the input to the model. However, due to the complexity of Chinese, the word segmentation system cannot avoid segmentation errors, and these errors will continue to propagate to the end of the sequence, resulting in poor model performance. Character-based models, although avoiding this problem, have difficulty leveraging word information in the sequence. Named entities usually consist of one or more consecutive words, and word boundaries often coincide with entity boundaries. Effectively using word information in the sequence can greatly improve the performance of NER. To incorporate word information into character-based models, many scholars have tried to use an external dictionary to match potential words in the sequence. Representatively, Zhang and Yang et al. proposed the lattice model, which uses a gating mechanism to control the weight of the potential word information obtained by matching the sequence with the dictionary and incorporates it into the character-based model. After that, Ma et al. proposed a simplified lattice model for the inefficiency of batch training due to the directed acyclic graph structure of the lattice model and the degradation problem of the lattice model. The simplified lattice model uses a soft word strategy and a weight fusion mechanism to replace the gating mechanism of the lattice model, effectively avoiding the model degradation problem. At the same time, by fixing the sentence length, the efficiency has also been greatly improved. Although the simplified lattice model directly and effectively utilizes word information, there are still the following problems: On the one hand, the soft word method loses some information of the middle groups for longer words. For example, for the input sequence: "Chinese football team", the middle group dictionary candidates corresponding to the character "ball" are "Chinese football team", "national football team", and "football team" (the character "ball" is in the middle position in all three words). The soft word method indiscriminately classifies these three words into the middle group without distinguishing their specific positions, and this relative position information is very important for NER. This problem becomes more and more serious as the entity length increases.On the other hand, the radical-level semantic information within characters in the sequence has not been explored and utilized. Although pre-trained language models can effectively capture the contextual semantic information in the sequence, they cannot obtain the inherent semantic information contained in the pictographic nature of characters. This semantic information is specifically reflected in: radicals, character structures, and writing order sequences. To sum up, the main problems in current research work are that the model is vulnerable to word segmentation errors or insufficient utilization of word information, and the radical-level semantic information within characters has not been considered. The lattice series of models have not fully utilized the semantic information at three different granularities of words, characters, and radicals in the sequence, and the recognition accuracy still needs to be improved. Summary of the Invention

[0004] The purpose of the present invention is to solve the problem that traditional Chinese named entity recognition methods are difficult to fully utilize the information in the sequence and have poor recognition effects, and a Chinese named entity recognition method, device, and medium that fuse multi-granularity information are proposed.

[0005] The present invention is implemented through the following technical solutions. The present invention proposes a Chinese named entity recognition method that fuses multi-granularity information, specifically including the following steps:

[0006] Step 1: Obtain a domain corpus dataset, preprocess the dataset, and divide it into a training set, a test set, and a validation set;

[0007] Step 2: Extract character, soft word, and radical-level pre-trained vectors from the corpus data preprocessed in Step 1 for vector fusion, and construct a Chinese named entity recognition model that fuses multi-granularity information;

[0008] Step 3: Input the data obtained in Step 2 into the model for training;

[0009] Step 4: Use the Chinese named entity recognition model that fuses multi-granularity information obtained in Step 3 to process and calculate the data to be recognized, and obtain the named entity recognition result.

[0010] Furthermore, Step 1 specifically includes the following steps:

[0011] Step 1.1: Identify the named entities in the sentence-level corpus data and label them with predefined types, and the types include person names, place names, and organization names;

[0012] Step 1.2: Divide the labeled results into character-level corpus data in the BMESO marking method, and its form is: character entity position - belonging to the predefined type;

[0013] Step 1.3: Divide the preprocessed dataset into a training set, a test set, and a validation set at a certain ratio.

[0014] Furthermore, Step 2 specifically includes the following steps:

[0015] Step 2.1: For the characters in the sequence, use the pre-trained language model to perform character mapping on the character sequence one by one, and encode each character in the input sequence into a low-dimensional dense embedding vector;

[0016] Step 2.2: For the candidate words corresponding to the characters in the sequence: build a vocabulary search tree based on the external dictionary, match the candidate words corresponding to the characters in the sentence, and build an extended soft word set. Then use the weight fusion strategy to weight the extended soft word set corresponding to the characters to obtain the word-level vector corresponding to the characters.

[0017] Step 2.3: For the radical-level features corresponding to the characters in the sequence: construct a radical-level feature lookup table for commonly used Chinese characters, represent the features as pre-trained embedding vectors, and use a convolutional neural network to extract the radical-level feature embedding vectors;

[0018] Step 2.4: sequentially concatenate character, soft word, and radical-level feature vectors;

[0019] Step 2.5: Perform padding / truncation operations on each sentence in the data set to a fixed length; for sentences that exceed the specified length, discard the part that exceeds the specified length; for sentences that are less than the specified length, perform padding operations to fill them up to the specified length;

[0020] Step 2.6: Take fixed-length sentences in batches of size Batch_Size as the input of the model. Each subsequence in the batch is a sentence.

[0021] Step 2.7: Perform hidden layer forward LSTM encoding and reverse LSTM encoding on the feature vectors in the Batch, and concatenate the forward and reverse hidden vectors to obtain a bidirectional feature vector of the data.

[0022] Furthermore, the step 2.2 specifically includes the following steps:

[0023] Step 2.2.1: Traverse the external dictionary and build a vocabulary prefix lookup tree;

[0024] Step 2.2.2: Use the vocabulary search tree to match the candidate words in the sentence, and build a soft word set for the character according to the position of the character in the candidate word;

[0025] Step 2.2.3: Count the total number of times the candidate word appears in the corpus data, as well as the number of times the candidate word appears in each position in the soft word set, and obtain its weight in each position of the soft word set;

[0026] Step 2.2.4: Weight the candidate words at all positions corresponding to the characters and concatenate the soft word level vectors.

[0027] Further, step 2.3 specifically includes the following steps:

[0028] Step 2.3.1: Construct a radical-level feature lookup table for common Chinese characters, where the radical-level features include: the simplified / traditional radical of the character, the structural composition of the character, and the writing order sequence of the character, in the form of: character - radical - structural composition - writing order sequence;

[0029] Step 2.3.2: Search the pre-trained embedding vector lookup table, and represent each radical-level feature corresponding to the character as an embedding vector of dimension d. At this time, the radical-level features corresponding to the character are represented as an embedding vector matrix;

[0030] Step 2.3.3: Fix the dimension of the embedding matrix to For a matrix with a length exceeding k, perform a truncation operation to take the first k features; for a matrix with a length less than k, perform random initialization to pad the length to k;

[0031] Step 2.3.4: Perform x consecutive one-dimensional convolutions on the radical-level feature embedding matrix with a fixed dimension and perform a max pooling operation to obtain a d-dimensional embedding vector representing the radical-level features corresponding to the character.

[0032] Further, step 3 specifically includes the following steps:

[0033] Step 3.1: Perform iterative update calculations on the bidirectional feature vectors in the hidden layer;

[0034] Step 3.2: Input the result into the CRF layer, iteratively update the emission probability and the transition probability, and calculate the maximum score sequence;

[0035] Step 3.3: Update and save the parameters of the trained model.

[0036] Further, step 4 specifically includes the following steps:

[0037] Step 4.1: Use the Chinese text sequence to be recognized as the input of the model in units of characters;

[0038] Step 4.2: Calculate and output the entity recognition result.

[0039] Further, the corpus data in the total number of times the candidate word appears in the corpus data refers to the training set + the test set.

[0040] The present invention provides an electronic device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the Chinese named entity recognition method for fusing multi-granularity information are implemented.

[0041] The present invention provides a computer-readable storage medium for storing computer instructions, and when the computer instructions are executed by a processor, the steps of the Chinese named entity recognition method for fusing multi-granularity information are implemented.

[0042] Compared with the prior art, the beneficial effect of the present invention is that on the basis of the BiLSTM model, the potential information in the sequence is fully utilized: the radical-level semantic information in the sequence is mined, and at the same time, the original soft word method is extended to better cope with the challenges brought by the increase in entity length, and the accuracy of named entity recognition is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is a flowchart of the Chinese named entity recognition method for fusing multi-granularity information;

[0044] Figure 2 is a model framework diagram of the Chinese named entity recognition method for fusing multi-granularity information;

[0045] Figure 3 is a schematic diagram of the extended soft word method;

[0046] Figure 4 is a schematic diagram of the detailed explanation of the radical-level information of the character "tang";

[0047] Figure 5 is a diagram of the radical-level feature extraction module. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0049] As Figures 1 to 5 shown, the present invention provides a Chinese named entity recognition method for fusing multi-granularity information, which specifically includes the following steps:

[0050] Step 1: Obtain a domain corpus dataset, preprocess the dataset and divide it into a training set, a test set, and a validation set;

[0051] The specific steps of Step 1 are as follows:

[0052] Step 1.1: Identify the named entities in the sentence-level corpus data and label them with predefined entity types, such as: person names, place names, organization names, etc.;

[0053] Step 1.2: Divide the labeled results into character-level corpus data in the BMESO tagging format, which is in the form of: character entity position - predefined type to which it belongs;

[0054] Step 1.3: Divide the preprocessed dataset into a training set, a test set, and a validation set in a ratio of 6:2:2.

[0055] The character sequence s in the preprocessed dataset is in the form of:

[0056] s = [c1, c2, c3, …, c n

[0057] where, c i represents the i-th character in the character sequence, and i ∈ [1, n]; c i,j represents the word composed of the i-th character to the j-th character in the character sequence, i, j ∈ [1, n] and i <= j (some words may consist of a single character);

[0058] Step 2: Extract the character, soft word, and radical-level pre-trained vectors from the corpus data preprocessed in Step 1, perform vector fusion, and build a Chinese named entity recognition model that fuses multi-granularity information;

[0059] Step 2.1: For the characters in the sequence, use the pre-trained language model BERT-wwm to perform character mapping on the character sequence one by one, and encode each character c i in the input sequence into a 768-dimensional embedding vector, as shown in the following formula:

[0060]

[0061] where, e c is the character embedding vector lookup table.

[0062] As Figure 2 shown, for the input sequence, first use BERT-wwm to obtain the TokenEmbedding, SegmentEmbedding, and PositionEmbedding of each character c i , and combine them as the pre-trained vector of the character.

[0063] Step 2.2: For the candidate words corresponding to the characters in the sequence: Build a lexical lookup tree based on an external dictionary, match the candidate words corresponding to the characters in the sentence, and construct an extended soft word set. Then, use a weight fusion strategy to weight the extended soft word set corresponding to the characters to obtain the word-level vector corresponding to the characters, which specifically includes the following steps:

[0064] Step 2.2.1: Traverse the external dictionary and build a lexical prefix lookup tree; ​

[0065] Step 2.2.2: Use the lexical search tree to match the candidate words in the sentence, and construct an extended soft word set for each character in the sequence according to the position of the character in the candidate word (including: start, middle group (the first middle position, the second middle position, the remaining middle positions), end, single-character word formation).

[0066] As Figure 3 shown, for the input sequence "Chinese football team", first find all potential words in the sequence: "Chinese football", "Chinese football team", "national football team", "football", "team". For the character "ball", the words containing it include "team", "football team", "Chinese football team", "national football team", "Chinese football", "football", "ball" (considering the case of single-character word formation). The character "ball" is in the starting position in the word "team", so add the word "team" to the Begin position word set of the extended soft word set of "ball". The character "ball" is in the M1 position in the word "football team", so add the word "football team" to the M1 position word set of the extended soft word set of "ball". The character "ball" is in the M2 position in the word "national football team", so add the word "national football team" to the M2 position word set of the extended soft word set of "ball". The character "ball" is in the M o position in the word "Chinese football team", so add the word "Chinese football team" to the M o position word set of the extended soft word set of "ball". After performing the above operations, obtain the extended soft word set as Figure 3 shown.

[0067] Step 2.2.3: Count the number of times z(w) that a candidate word appears in a certain position in the soft word set in the corpus data (training set + test set), and count the total number of times the soft word set appears in the data, and perform weighting to obtain the weighted word embedding representation v w (W):

[0068]

[0069]

[0070] Among them, W represents the "BM1M2M o ES" soft word set, and w represents a certain candidate word in W.

[0071] Step 2.2.4: Concatenate the soft word vectors at different positions of the character to obtain the soft word-level vector z w of the character.

[0072] Specifically, z w is calculated by the following formula:

[0073]

[0074] in Represents the concatenation operation of vectors, v w () represents the word embedding vector of a soft word position.

[0075] Step 2.3: For the radical-level features corresponding to the characters in the sequence: Build a radical-level feature lookup table for commonly used Chinese characters, and represent the features as pre-trained embedding vectors. Use a convolutional neural network to extract the radical-level feature embedding vectors. The specific steps include the following:

[0076] Step 2.3.1: Build a radical-level feature lookup table for commonly used Chinese characters r , its radical-level features include: simplified / traditional radical of the character, structural composition of the character, and writing order sequence of the character, in the form of: character-radical-structural composition-writing order sequence;

[0077] like Figure 4 As shown, for the character "涼", its radical is: "火", its structural components are: "氵", "扬", "火", and its writing order is: "汤", "火", and its radical-level features reflect the original meaning of the character.

[0078] Step 2.3.2: Find the pre-trained embedding vector lookup table and convert the character c i ={r1,r2,…,r m Each radical-level feature r corresponding to j , j∈[1,m] is represented as an embedding vector with dimension d. At this time, the radical-level features corresponding to the characters can be represented as an embedding vector matrix O, as shown in the following formula:

[0079]

[0080]

[0081] where e r (r j ) represents the radical-level embedding vector lookup table, which is obtained by training ChineseGiga-Word using word2vec.

[0082] Step 2.3.3: Fix the dimension of the embedding matrix to For matrices with a length exceeding k, truncation is performed to obtain the first k features; for matrices with a length less than k, random initialization is performed to fill the length to k; after this step, the embedded vector matrix becomes:

[0083] Step 2.3.4: If Figure 5As shown in FIG. 1 , a fixed-dimensional radical-level feature embedding matrix is ​​subjected to x consecutive one-dimensional convolutions and a maximum pooling operation to obtain a d-dimensional embedding vector representing the radical-level feature corresponding to the character.

[0084] Step 2.4: Concatenate the character, soft word, and radical level feature vectors in sequence, as shown below:

[0085]

[0086] Among them, x c is the character embedding vector, y r is the radical level embedding vector, z w is the soft word level embedding vector, Represents the concatenation operation of vectors. The concatenated result is a character embedding vector containing word-level information and radical-level information.

[0087] Step 2.5: Perform padding / truncation operations on each sentence in the data set to a fixed length. Specifically, for sentences that exceed the specified length, the part that exceeds the specified length is discarded; for sentences that are less than the specified length, perform padding operations to fill them up to the specified length;

[0088] Step 2.6: Take fixed-length sentences in batches of size Batch_Size as the input of the model. Each subsequence in the batch is a sentence.

[0089] Step 2.7: Perform hidden layer forward LSTM encoding and reverse LSTM encoding on the feature vectors in the Batch, and concatenate the forward and reverse hidden vectors to obtain a bidirectional feature vector of the data.

[0090]

[0091]

[0092] h t =o t ⊙tanh(c t ).

[0093] Where σ is the sigmoid function of the element, ⊙ represents the product of the element, and W and b are trainable parameters. The memory cell c can be regarded as long-term memory, and the hidden state h is short-term memory. The reverse LSTM shares the same definition as the forward LSTM, but models the sequence in reverse order. The hidden state of the i-th step concatenated by the forward and reverse LSTM forms c i Context-dependent representation of .

[0094] Step 3: Input the data obtained in step 2 into the model for training, which includes the following steps:

[0095] Step 3.1: Iteratively update and calculate the bidirectional feature vectors in the hidden layer;

[0096] Step 3.2: Input the result into the CRF layer, iteratively update the emission probability and transition probability, and calculate the maximum score sequence;

[0097]

[0098] where P is the output of the BiLSTM, representing the emission score of label y i and the transition matrix T represents the transition probability from label y i to y i+1

[0099] Step 3.3: Update and save the parameters of the trained model.

[0100] Specifically, the model is trained using the negative log-likelihood loss function, and L2 regularization is used to alleviate the overfitting problem, as shown in the following formula:

[0101]

[0102] where θ represents the set of parameters and λ is the regularization parameter.

[0103] Step 4: Use the Chinese named entity recognition model that fuses multi-granularity information obtained in Step 3 to process and calculate the data to be recognized, and obtain the named entity recognition result, which specifically includes the following steps:

[0104] Step 4.1: Use the Chinese text sequence to be recognized as the input of the model in units of characters;

[0105] Step 4.2: Calculate and output the entity recognition result.

[0106] The present invention provides an electronic device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the Chinese named entity recognition method that fuses multi-granularity information are implemented.

[0107] The present invention provides a computer-readable storage medium for storing computer instructions, and when the computer instructions are executed by a processor, the steps of the Chinese named entity recognition method that fuses multi-granularity information are implemented.

[0108] ​The above has introduced in detail a Chinese named entity recognition method, device, and medium that fuse multi-granularity information. In this article, specific examples are used to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation on the present invention.

Claims

1. A Chinese named entity recognition method that integrates multi-granularity information, characterized in that, The specific steps include: Step 1: Obtain the domain corpus dataset, preprocess the dataset and divide it into training set, test set and validation set; Step 2: extract the character, soft word and radical level pre-trained vectors from the corpus data pre-processed in step 1, perform vector fusion, and build a Chinese named entity recognition model that integrates multi-granularity information; The step 2 specifically includes the following steps: Step 2.1: For the characters in the sequence, use the pre-trained language model to perform character mapping on the character sequence one by one, and encode each character in the input sequence into a low-dimensional dense embedding vector; Step 2.2: For the candidate words corresponding to the characters in the sequence: build a vocabulary search tree based on the external dictionary, match the candidate words corresponding to the characters in the sentence, and build an extended soft word set. Then use the weight fusion strategy to weight the extended soft word set corresponding to the characters to obtain the word-level vector corresponding to the characters. Step 2.3: For the radical-level features corresponding to the characters in the sequence: construct a radical-level feature lookup table for commonly used Chinese characters, represent the features as pre-trained embedding vectors, and use a convolutional neural network to extract the radical-level feature embedding vectors; Step 2.4: sequentially concatenate character, soft word, and radical level feature vectors; Step 2.5: Perform padding / truncation operations on each sentence in the data set to a fixed length; for sentences that exceed the specified length, discard the part that exceeds the specified length; for sentences that are less than the specified length, perform padding operations to fill them up to the specified length; Step 2.6: Take fixed-length sentences in batches of size Batch_Size as the input of the model. Each subsequence in the batch is a sentence. Step 2.7: Perform hidden layer forward LSTM encoding and reverse LSTM encoding on the feature vectors in the Batch, and concatenate the forward and reverse hidden vectors to obtain a bidirectional feature vector of the data; Step 3: Input the data obtained in step 2 into the model for training; Step 4: Use the Chinese named entity recognition model that integrates multi-granularity information obtained in step 3 to process and calculate the data to be recognized to obtain the named entity recognition result.

2. The method according to claim 1, wherein The step 1 specifically comprises the following steps: Step 1.1: Identify named entities in sentence-level corpus data and annotate them into predefined types, including names of people, places, and organizations; Step 1.2: Divide the annotated results into character-level corpus data using BMESO tags in the form of: character entity position-predefined type; Step 1.3: Divide the preprocessed dataset into training set, test set and validation set in a certain ratio.

3. The method according to claim 1, characterized in that The step 2.2 specifically includes the following steps: Step 2.2.1: Traverse the external dictionary and build a vocabulary prefix search tree; Step 2.2.2: Use the vocabulary search tree to match the candidate words in the sentence, and build a soft word set for the character according to the position of the character in the candidate word; Step 2.2.3: Count the total number of times the candidate word appears in the corpus data, as well as the number of times the candidate word appears in each position in the soft word set, and obtain its weight in each position of the soft word set; Step 2.2.4: Weight the candidate words at all positions corresponding to the characters and splice the soft word-level vectors.

4. The method according to claim 3, wherein The specific steps of step 2.3 are as follows: Step 2.3.1: Construct a radical-level feature lookup table for common Chinese characters, and its radical-level features include: the simplified / traditional radical of the character, the structural composition of the character, and the writing order sequence of the character, and its form is: character - radical - structural composition - writing order sequence; Step 2.3.2: Look up the pre-trained embedding vector lookup table, and represent each radical-level feature corresponding to the character as an embedding vector with a dimension of d. At this time, the radical-level features corresponding to the character are represented as an embedding vector matrix; Step 2.3.3: Fix the dimension of the embedding matrix to be For matrices with a length exceeding k, perform a truncation operation to take the first k features; for matrices with a length less than k, perform random initialization to pad the length to k. Step 2.3.4: Perform x consecutive one-dimensional convolutions on the radical-level feature embedding matrix with a fixed dimension and perform a max pooling operation to obtain a d-dimensional embedding vector representing the radical-level features corresponding to the character.

5. The method according to claim 4, characterized in that The specific steps of step 3 are as follows: Step 3.1: Iteratively update and calculate the bidirectional feature vectors in the hidden layer; Step 3.2: Input the result into the CRF layer, iteratively update the emission probability and the transition probability, and calculate the maximum score sequence; Step 3.3: Update and save the parameters of the trained model.

6. The method according to claim 5, characterized in that The specific steps of step 4 are as follows: Step 4.1: Use the Chinese text sequence to be recognized as the input of the model in units of characters; Step 4.2: Calculate and output the entity recognition result.

7. The method according to claim 3, characterized in that, The corpus data in the total number of occurrences of the statistical candidate words in the corpus data refers to the training set + the test set.

8. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1-7.

9. A computer-readable storage medium for storing computer instructions, characterized in that, When the computer instructions are executed by the processor, it implements the steps of the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Chinese entity identification method based on BERT and Word2Vec vector fusion

    CN112632997A

  • Chinese sentence keyboard input system

    CN1159028A