Entity naming recognition method, device, equipment and storage medium

CN113901821BActive Publication Date: 2025-09-05PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111217330.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-19
Publication Date
2025-09-05
Estimated Expiration
2041-10-19

AI Technical Summary

Technical Problem

[0003]本申请提供了一种实体命名识别方法、装置、设备及存储介质,以解决现有技术中某些词一词多义,导致对实体命名的识别不准确的问题

Benefits of technology

[0042] By obtaining the text to be processed, and then randomly selecting a preset parallel corpus from a preset corpus, and replacing the text to be processed according to a certain ratio, a parallel text is obtained, so that the comparison between the parallel text and the text to be processed can be easily realized. The text to be processed and the parallel text are then input into the vector conversion layer of the recognition model for vector conversion to obtain corresponding word vectors, and the similarity between the text to be processed and the parallel text is calculated based on the word vector to determine whether the entity naming is polysemous and whether the meaning of the entity naming is the same as that in the parallel text. When the similarity is greater than a preset value, it proves that there may be an entity in the text to be processed. The word vector corresponding to the text to be processed is then input into the LSTM layer in the recognition model to obtain the type distribution probability corresponding to each word in the text to be processed, and then processed through the conditional random field layer to obtain the entity in the text to be processed. Therefore, even if the entity naming is polysemous, the entity naming can be accurately determined, thereby improving the accuracy of entity naming recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113901821B_ABST
    Figure CN113901821B_ABST
Patent Text Reader

Abstract

The present application relates to the fields of artificial intelligence and digital medicine, and discloses a method, apparatus, device, and storage medium for entity naming recognition, the method comprising: obtaining a text to be processed; randomly selecting a preset parallel corpus from a preset corpus, replacing the text to be processed according to a preset ratio, and obtaining a parallel text; inputting the parallel text and the text to be processed into a vector conversion layer in a recognition model for vector conversion, obtaining word vectors corresponding to the text to be processed and the parallel text, and calculating the similarity between the word vectors corresponding to the parallel text and the text to be processed; when the similarity is greater than a preset value, inputting the word vector corresponding to the text to be processed into the LSTM layer in the recognition model, obtaining the type distribution probability corresponding to each word in the text to be processed; inputting the type distribution probability into the conditional random field layer in the recognition model, and obtaining the entity name. The present application also relates to blockchain technology, and the text data to be processed is stored in the blockchain. The present application can improve the accuracy of entity naming recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and in particular to a method, apparatus, device and storage medium for entity naming recognition. Background Art

[0002] With the continuous development of artificial intelligence technology and the emergence of a large number of demand scenarios in daily work and life, entity recognition (NER) has been widely used in our work and life needs, such as identifying product names, medicine names, etc. in the text entered on e-commerce websites, and identifying names of people and places in a description. In the existing technology, the main technical methods for entity recognition are divided into: rule-based and dictionary-based methods, statistical-based methods, hybrid methods of the two, neural network methods, etc. The statistical-based methods mainly include Hidden Markov Model (HMM), Maximum Entropy (Maximum Entropy), Support Vector Machine (SVM), etc. However, the solutions in the existing technology have the problem that the recognition accuracy of some entity words is not high when they have multiple meanings. Therefore, how to solve the problem of low recognition accuracy of some entity words when they have multiple meanings has become an urgent problem to be solved. Summary of the Invention

[0003] The present application provides an entity naming recognition method, apparatus, device and storage medium to solve the problem in the prior art that some words have multiple meanings, resulting in inaccurate recognition of entity names.

[0004] To solve the above problems, this application provides an entity naming recognition method, including:

[0005] Get the text to be processed;

[0006] Randomly selecting a preset parallel corpus from a preset corpus, replacing the text to be processed according to a preset ratio, and obtaining a parallel text;

[0007] Inputting the text to be processed and the parallel text into a vector conversion layer in a recognition model for vector conversion, obtaining word vectors corresponding to the text to be processed and the parallel text, and calculating the similarity between the word vectors corresponding to the text to be processed and the parallel text;

[0008] When the similarity is greater than a preset value, the word vector corresponding to the text to be processed is input into the LSTM layer in the recognition model to obtain the type distribution probability corresponding to each word in the text to be processed;

[0009] The type distribution probability is input into the conditional random field layer in the recognition model to obtain the entity names in the text to be processed.

[0010] Furthermore, the randomly selecting a preset parallel corpus from a preset corpus and replacing the to-be-processed text according to a preset ratio includes:

[0011] Determining keywords of the text to be processed based on a preset keyword set and the text to be processed;

[0012] According to the text to be processed and the keywords, a preset proportion of the content in the text to be processed is replaced with a preset parallel corpus to obtain a parallel text.

[0013] Furthermore, before inputting the text to be processed and the parallel text into the vector conversion layer in the recognition model for vector conversion, the method further includes:

[0014] The text to be processed and the parallel text are processed by the embedding layer in the recognition model to obtain corresponding first vectors and second vectors;

[0015] Inputting the text to be processed and the parallel text into a vector conversion layer in a recognition model for vector conversion to obtain word vectors corresponding to the text to be processed and the parallel text includes:

[0016] The attention layer in the vector conversion layer calculates the attention vectors of each word in the to-be-processed text and the parallel text for the keyword based on the first vector and the second vector;

[0017] The attention vector is processed by two fully connected layers in the vector conversion layer to obtain word vectors corresponding to each word in the parallel text and the text to be processed.

[0018] Furthermore, the attention layer in the vector conversion layer calculates the attention vector of each word in the to-be-processed text for the keyword based on the first vector, including:

[0019] Multiplying the first vector by the parameter matrix obtained after pre-training to obtain the corresponding Q matrix, K matrix and V matrix;

[0020] A weight matrix is ​​obtained by performing a dot product of the Q matrix and the K matrix, dividing a first result obtained by the dot product by the square root of the corresponding dimension of the Q matrix to obtain a second result, and performing a Softmax calculation on the second result.

[0021] Multiplying the V matrix by the weight matrix to obtain a first matrix;

[0022] The first matrix is ​​processed by the fully connected layer in the vector conversion layer to obtain the attention vector corresponding to the text to be processed.

[0023] Furthermore, the calculating of the similarity between the word vectors corresponding to the to-be-processed text and the parallel text includes:

[0024] Obtaining a first word vector corresponding to the replacement text in the parallel text and a second word vector corresponding to the replaced text in the to-be-processed text;

[0025] The cosine similarity between the first word vector and the second word vector is calculated to obtain the similarity between the word vectors corresponding to the to-be-processed text and the parallel text.

[0026] Furthermore, before inputting the word vector corresponding to the to-be-processed text into the LSTM layer in the recognition model, the method further includes:

[0027] The word vector corresponding to the text to be processed is processed through a Dropout layer in the recognition model to suppress overfitting of the recognition model.

[0028] Furthermore, before inputting the type distribution probability into the conditional random field layer in the recognition model, the method further includes:

[0029] The type distribution probability is also processed by the Dropout layer and the linear transformation layer in the recognition model.

[0030] In order to solve the above problems, the present application further provides an entity naming recognition device, the device comprising:

[0031] Acquisition module, used to obtain the text to be processed;

[0032] A replacement module is used to randomly select a preset parallel corpus from a preset corpus, replace the text to be processed according to a preset ratio, and obtain a parallel text;

[0033] a similarity calculation module, configured to input the text to be processed and the parallel text into a vector conversion layer in a recognition model for vector conversion, obtain word vectors corresponding to the text to be processed and the parallel text, and calculate the similarity between the word vectors corresponding to the text to be processed and the parallel text;

[0034] A probability calculation module, configured to input the word vector corresponding to the text to be processed into the LSTM layer in the recognition model when the similarity is greater than a preset value, to obtain a type distribution probability corresponding to each word in the text to be processed;

[0035] The entity extraction module is used to input the type distribution probability into the conditional random field layer in the recognition model to obtain the entity naming in the text to be processed.

[0036] In order to solve the above problem, the present application further provides a computer device, comprising:

[0037] at least one processor; and,

[0038] a memory communicatively connected to the at least one processor; wherein,

[0039] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the entity naming recognition method as described above.

[0040] In order to solve the above problems, the present application also provides a non-volatile computer-readable storage medium, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor, the entity naming recognition method as described above is implemented.

[0041] The entity naming recognition method, apparatus, device, and storage medium provided in the embodiments of the present application have at least the following advantages compared to the prior art:

[0042] By obtaining the text to be processed, and then randomly selecting a preset parallel corpus from a preset corpus, and replacing the text to be processed according to a certain ratio, a parallel text is obtained, so that the comparison between the parallel text and the text to be processed can be easily realized. The text to be processed and the parallel text are then input into the vector conversion layer of the recognition model for vector conversion to obtain corresponding word vectors, and the similarity between the text to be processed and the parallel text is calculated based on the word vector to determine whether the entity naming is polysemous and whether the meaning of the entity naming is the same as that in the parallel text. When the similarity is greater than a preset value, it proves that there may be an entity in the text to be processed. The word vector corresponding to the text to be processed is then input into the LSTM layer in the recognition model to obtain the type distribution probability corresponding to each word in the text to be processed, and then processed through the conditional random field layer to obtain the entity in the text to be processed. Therefore, even if the entity naming is polysemous, the entity naming can be accurately determined, thereby improving the accuracy of entity naming recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the solutions in this application, a brief introduction will be given below to the drawings required for use in the description of the embodiments of this application. Obviously, the drawings described below are some embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0044] Figure 1 A flowchart of an entity naming recognition method provided in one embodiment of the present application;

[0045] Figure 2 A schematic diagram of a module of an entity naming recognition device provided in one embodiment of the present application;

[0046] Figure 3 This is a schematic structural diagram of a computer device according to an embodiment of the present application. DETAILED DESCRIPTION

[0047] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of the application are for the purpose of describing specific embodiments only and are not intended to limit this application. The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first" and "second" in the specification and claims of this application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.

[0048] References to "embodiments" herein mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various locations in the specification does not necessarily refer to the same embodiment, nor does it refer to independent or alternative embodiments that are mutually exclusive of other embodiments. It is understood, either explicitly or implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0049] This application provides a method for entity naming recognition. Figure 1 , which is a flow chart of an entity naming recognition method provided in one embodiment of the present application.

[0050] In this embodiment, the entity naming recognition method includes:

[0051] S1. Get the text to be processed;

[0052] In this application, the text to be processed, ie, the text to be recognized, is obtained from a database, or the text to be processed is directly input by a user.

[0053] Furthermore, obtaining the text to be processed includes:

[0054] Send a call request to the preset knowledge base, the call request carries a signature verification token;

[0055] Receive the signature verification result returned by the knowledge base, and when the signature verification result is passed, call the text to be processed in the knowledge base, wherein the signature verification result is obtained by the knowledge base through RSA asymmetric encryption verification based on the signature verification token.

[0056] Since the text to be processed may involve the user's private data, it will be saved in the preset database. Therefore, when obtaining the text to be processed, the database will perform a signature verification step to ensure the security of the data and avoid data leakage.

[0057] The whole process is that the client calculates the first message digest of message m, and encrypts the first message digest with RSA asymmetric encryption (using the client's private key) to obtain signature s, and then uses the public key of the knowledge base to obtain ciphertext c for message m and signature s, and sends it to the knowledge base. The knowledge base uses its own private key to decrypt the ciphertext c to obtain message m and signature s. The knowledge base uses the client's public key to decrypt signature s to obtain the first message digest; at the same time, the knowledge base uses the same method to extract the digest of message m to obtain the second message digest, and determines whether the first message digest and the second message digest are the same. If they are the same, the verification is successful; if they are different, the verification fails.

[0058] By requiring signature verification when retrieving data, the security of data stored in the database is guaranteed and data leakage is avoided.

[0059] S2. Randomly select a preset parallel corpus from a preset corpus, and replace the text to be processed according to a preset ratio to obtain a parallel text;

[0060] Specifically, a parallel text is obtained by randomly selecting a preset parallel corpus from a preset corpus and replacing the text to be processed with the parallel corpus at a preset ratio. The parallel text includes both part of the text to be processed and the parallel corpus, thereby obtaining the parallel text. The preset corpus stores a large amount of parallel corpus, and the parallel corpus is sentences containing entity names.

[0061] Furthermore, the randomly selecting a preset parallel corpus from a preset corpus and replacing the to-be-processed text according to a preset ratio includes:

[0062] Determining keywords of the text to be processed based on a preset keyword set and the text to be processed;

[0063] According to the text to be processed and the keywords, a preset proportion of the content in the text to be processed is replaced with a preset parallel corpus to obtain a parallel text.

[0064] Before replacing the text to be processed, keywords are first extracted from the text to be processed. Specifically, since this application primarily identifies corporate entities or trademarks, a pre-defined keyword library is provided, which stores company names or their abbreviations, as well as trademark names. Keywords are obtained by matching the text to be processed with data in the keyword library. When replacing the text to be processed using a parallel corpus, a predetermined proportion of the text to be processed is replaced, wherein the replaced portion of the text to be processed does not include keywords. Furthermore, the replacement operation is based on sentences. For example, if "Red Fuji apples are indeed delicious, and the peaches in Changping, Beijing are also good" is replaced, the result is "Red Fuji apples are indeed delicious, and Honor phones are actually quite useful." This uses sentences as the minimum replacement unit, and does not directly replace individual words within a sentence. Furthermore, when the text to be processed consists of only one sentence, a pre-defined parallel corpus is randomly selected and appended directly to the text to be processed, thereby generating a parallel text.

[0065] By replacing or adding part of the text to be processed, a parallel text is obtained, which facilitates subsequent comparison between the text to be processed and the parallel text to pre-judge the keywords in the text to be processed.

[0066] S3, inputting the text to be processed and the parallel text into a vector conversion layer in a recognition model for vector conversion, obtaining word vectors corresponding to the text to be processed and the parallel text, and calculating the similarity between the word vectors corresponding to the text to be processed and the parallel text;

[0067] Specifically, the parallel text and the text to be processed are input into the vector conversion layer in the recognition model for processing to obtain word vectors corresponding to the text to be processed and the parallel text, each word vector contains the word vectors of other words in the sentence, and the vector conversion layer is trained based on the BERT (Bidirectional Encoder Representation from Transformers, language representation model) model; by calculating the similarity between the corresponding word vectors of the parallel text and the text to be processed, a preliminary judgment is made as to whether the keywords in the text to be processed belong to a company entity or a trademark entity.

[0068] Furthermore, before inputting the text to be processed and the parallel text into the vector conversion layer in the recognition model for vector conversion, the method further includes:

[0069] The text to be processed and the parallel text are processed by the embedding layer in the recognition model to obtain corresponding first vectors and second vectors;

[0070] Inputting the text to be processed and the parallel text into a vector conversion layer in a recognition model for vector conversion to obtain word vectors corresponding to the text to be processed and the parallel text includes:

[0071] The attention layer in the vector conversion layer calculates the attention vectors of each word in the to-be-processed text and the parallel text for the keyword based on the first vector and the second vector;

[0072] The attention vector is processed by two fully connected layers in the vector conversion layer to obtain word vectors corresponding to each word in the text to be processed and the parallel text.

[0073] Specifically, the embedding layer in the recognition model is first processed to obtain a first vector and a second vector corresponding to the text to be processed and the parallel text. The first vector and the second vector are ordinary word vectors. After obtaining the ordinary vectors, the attention layer in the vector conversion layer is used to calculate the attention vector of each word in the text to be processed and the parallel text for the keyword. The attention vector is then processed through the two fully connected layers to obtain the word vectors corresponding to each word in the text to be processed and the parallel text. The word vector corresponding to each word contains information about all word vectors in the current sentence.

[0074] Furthermore, the attention layer and the two fully connected layers in the recognition model can be repeatedly set in multiple groups to obtain better word vectors corresponding to each word in the parallel text and the text to be processed. In this application, the attention layer and the two fully connected layers in the recognition model are repeated in 12 groups.

[0075] By first performing ordinary embedding layer processing on the text to be processed and the parallel text to obtain a first vector and a second vector, the first vector and the second vector are then processed by an attention layer and a fully connected layer to obtain word vectors that can better represent the text to be processed and the parallel text, thereby improving the accuracy of subsequent similarity judgment and final result output.

[0076] Furthermore, the attention layer in the vector conversion layer calculates the attention vector of each word in the to-be-processed text for the keyword based on the first vector, including:

[0077] Multiplying the first vector by the parameter matrix obtained after pre-training to obtain the corresponding Q matrix, K matrix and V matrix;

[0078] A weight matrix is ​​obtained by performing a dot product of the Q matrix and the K matrix, dividing a first result obtained by the dot product by the square root of the corresponding dimension of the Q matrix to obtain a second result, and performing a Softmax calculation on the second result.

[0079] Multiplying the V matrix by the weight matrix to obtain a first matrix;

[0080] The first matrix is ​​processed by the fully connected layer in the vector conversion layer to obtain the attention vector corresponding to the text to be processed.

[0081] Specifically, before the first training of the recognition model, the parameter matrix is ​​randomly generated, or can be randomly generated according to a normal distribution or uniform distribution. During the continuous training of the recognition model, the parameter matrix continuously changes and converges until the recognition model training is completed and the parameter matrix tends to be stable. In subsequent use of the trained recognition model, the stable multi-batch parameter matrix can be directly used.

[0082] By multiplying the first vector by the parameter matrix obtained after pre-training, the three matrices Q, K, and V are obtained, and the formula is Q = x4·W Q K=x4·W K ,V=x4·W V Where W Q , W K , W V That is the parameter matrix. Then use the three matrices Q, K, and V to calculate the proportion of each word in the input text. The weight of the word that only receives the above information is 0. The specific calculation formula is

[0083] Where A represents the weight matrix, d k The weight matrix represents the Q, K, or V matrix dimension. The weight matrix is ​​then multiplied by the V matrix of the corresponding batch to obtain a first matrix. This first matrix is ​​processed by the fully connected layer in the vector conversion layer to obtain the attention vector corresponding to the text to be processed. The fully connected layer includes two basic fully connected networks. Correspondingly, the second vector corresponding to the parallel text is processed according to the above method to obtain the attention vector corresponding to the parallel text.

[0084] By introducing the attention mechanism, we can obtain word vectors that better represent the text to be processed and the parallel text, thereby improving the accuracy of subsequent similarity judgment and final result output.

[0085] Furthermore, the calculating of the similarity between the word vectors corresponding to the to-be-processed text and the parallel text includes:

[0086] Obtaining a first word vector corresponding to the replacement text in the parallel text and a second word vector corresponding to the replaced text in the to-be-processed text;

[0087] The cosine similarity between the first word vector and the second word vector is calculated to obtain the similarity between the word vectors corresponding to the to-be-processed text and the parallel text.

[0088] Specifically, after obtaining the word vectors of the parallel text and the text to be processed, the cosine similarity between the first word vector and the second word vector is calculated by obtaining the first word vector of the replacement text in the parallel text and the second word vector of the replaced text in the text to be processed, that is, extracting the word vectors corresponding to different parts of the parallel text and the text to be processed. For example, for the text to be processed "Red Fuji apples are indeed delicious, and the peaches in Changping, Beijing are also good" and the parallel text "Red Fuji apples are indeed delicious, and Honor phones are actually quite good", the similarity is calculated by extracting "The peaches in Changping, Beijing are also good" in the text to be processed and "Honor phones are actually quite good" in the parallel text. More specifically, the similarity is calculated between "Peaches, not bad" and "Honor phones, quite good".

[0089] By calculating the similarity, it is pre-determined whether the keywords in the text to be processed are entity names, thereby improving the judgment speed, and entity name recognition is performed by combining pre-determination and subsequent judgment, thereby improving the recognition accuracy.

[0090] S4. When the similarity is greater than a preset value, inputting the word vector corresponding to the text to be processed into the LSTM (Long-Short Term Memory, long short-term memory model recurrent neural network) layer in the recognition model to obtain the type distribution probability corresponding to each word in the text to be processed;

[0091] Specifically, when the similarity is less than a preset value, it proves that the keyword is not a company entity or a trademark entity, and the result can be directly output, that is, the text to be processed does not contain a company entity or a trademark entity; when the similarity is greater than a preset value, the vector corresponding to the text to be processed is input into the LSTM layer in the recognition model. The LSTM layer is a bidirectional LSTM neural network, which outputs the probability distribution of the text to be processed being labeled into various types, and obtains the type distribution probability corresponding to each word in the text to be processed, and the LSTM layer can learn the semantic relationship of the text.

[0092] Furthermore, before inputting the word vector corresponding to the to-be-processed text into the LSTM layer in the recognition model, the method further includes:

[0093] The word vector corresponding to the text to be processed is processed through a Dropout layer in the recognition model to suppress overfitting of the recognition model.

[0094] In the present application, a Dropout layer is provided between the vector conversion layer and the LSTM layer in the recognition model to suppress overfitting of the recognition model.

[0095] The Dropout layer is set to suppress model overfitting.

[0096] S5. Input the type distribution probability into the conditional random field layer in the recognition model to obtain the entity names in the text to be processed.

[0097] Specifically, the Conditional Random Field (CRF) layer adds constraints to the resulting class distribution to ensure compliance. These constraints are automatically learned during the training of the recognition model.

[0098] Furthermore, before inputting the type distribution probability into the conditional random field layer in the recognition model, the method further includes:

[0099] The type distribution probability is also processed by the Dropout layer and the linear transformation layer in the recognition model.

[0100] Specifically, a Layer Norm layer, a Dropout layer, and a linear transformation layer are set between the LSTM layer and the conditional random field layer to perform normalization and suppress model overfitting.

[0101] By setting the Layer Norm layer, Dropout layer and linear transformation layer to perform normalization, it also suppresses model overfitting.

[0102] It should be emphasized that in order to further ensure the privacy and security of the data, the text data to be processed can also be stored in a node of a blockchain.

[0103] The blockchain referred to in this application is a new application model for computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Blockchain is essentially a decentralized database, a series of data blocks generated using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of this information (to prevent counterfeiting) and generate the next block. Blockchain can include the underlying blockchain platform, the platform product service layer, and the application service layer.

[0104] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.

[0105] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0106] By obtaining the text to be processed, and then randomly selecting a preset parallel corpus from a preset corpus, and replacing the text to be processed according to a certain ratio, a parallel text is obtained, so that the comparison between the parallel text and the text to be processed can be easily realized. The text to be processed and the parallel text are then input into the vector conversion layer of the recognition model for vector conversion to obtain corresponding word vectors, and the similarity between the text to be processed and the parallel text is calculated based on the word vector to determine whether the entity naming is polysemous and whether the meaning of the entity naming is the same as that in the parallel text. When the similarity is greater than a preset value, it proves that there may be an entity in the text to be processed. The word vector corresponding to the text to be processed is then input into the LSTM layer in the recognition model to obtain the type distribution probability corresponding to each word in the text to be processed, and then processed through the conditional random field layer to obtain the entity in the text to be processed. Therefore, even if the entity naming is polysemous, the entity naming can be accurately determined, thereby improving the accuracy of entity naming recognition.

[0107] This embodiment also provides an entity naming recognition device, such as Figure 2 , which is a functional module diagram of the entity naming recognition device of the present application.

[0108] The entity naming recognition device 100 described in this application can be installed in an electronic device. Depending on the functions implemented, the entity naming recognition device 100 may include an acquisition module 101, a replacement module 102, a similarity calculation module 103, a probability calculation module 104, and an entity extraction module 105. The module described in this application can also be referred to as a unit, which refers to a series of computer program segments that can be executed by an electronic device processor and can perform fixed functions, which are stored in the memory of the electronic device.

[0109] In this embodiment, the functions of each module / unit are as follows:

[0110] An acquisition module 101 is used to acquire the text to be processed;

[0111] Specifically, the acquisition module 101 acquires the text to be processed, ie, the text to be recognized, from a database, or directly allows the user to input the text to be processed.

[0112] Furthermore, the acquisition module 101 includes a request sending submodule and a data calling submodule;

[0113] The request sending submodule is used to send a call request to a preset knowledge base, wherein the call request carries a signature verification token;

[0114] The data calling submodule is used to receive the verification result returned by the knowledge base, and when the verification result is passed, call the text to be processed in the knowledge base, wherein the verification result is obtained by the knowledge base through RSA asymmetric encryption verification based on the verification token.

[0115] Since the text to be processed may involve the user's private data, the text data to be processed will be saved in the preset database. Therefore, when obtaining the text to be processed, the database will perform a signature verification step to ensure the security of the data and avoid data leakage and other problems.

[0116] By cooperating with the request sending submodule and the data calling submodule, signature verification is required when retrieving data, which ensures the security of the data stored in the database and avoids data leakage.

[0117] The replacement module 102 is configured to randomly select a preset parallel corpus from a preset corpus and replace the text to be processed according to a preset ratio to obtain a parallel text;

[0118] Specifically, the replacement module 102 randomly selects preset parallel corpora from a preset corpus, and replaces the text to be processed with the parallel corpora in a preset ratio to obtain a parallel text, wherein the parallel text includes both part of the text to be processed and the parallel corpora, thereby obtaining a parallel text.

[0119] Furthermore, the replacement module 102 includes a keyword determination submodule and a ratio replacement submodule;

[0120] The keyword determination submodule is used to determine the keywords of the text to be processed based on a preset keyword set and the text to be processed;

[0121] The proportion replacement submodule is used to replace the content of the to-be-processed text in a preset proportion with a preset parallel corpus according to the to-be-processed text and the keywords, so as to obtain a parallel text.

[0122] Specifically, before replacing the text to be processed, keywords are first extracted from the text to be processed. Specifically, since this application primarily identifies corporate entities or trademarks, a keyword library is pre-set, storing company names or their abbreviations, trademark names, etc. The keyword determination submodule matches the text to be processed with data in the keyword library to obtain keywords from the text to be processed. When replacing the text to be processed using parallel corpus, the proportional replacement submodule replaces a preset proportion of the text to be processed, wherein the replaced portion of the text to be processed does not include keywords. Furthermore, the replacement operation uses a sentence as the minimum replacement unit.

[0123] By cooperating with the keyword determination submodule and the proportional replacement submodule, part of the text to be processed is replaced or added to obtain a parallel text, which is convenient for subsequent comparison between the text to be processed and the parallel text to pre-judge the keywords in the text to be processed.

[0124] A similarity calculation module 103 is configured to input the text to be processed and the parallel text into a vector conversion layer in a recognition model for vector conversion, obtain word vectors corresponding to the text to be processed and the parallel text, and calculate the similarity between the word vectors corresponding to the text to be processed and the parallel text;

[0125] Specifically, the similarity calculation module 103 inputs the parallel text and the text to be processed into the vector conversion layer in the recognition model for processing, and obtains the word vectors corresponding to the text to be processed and the parallel text, each of which contains the word vectors of other words in the sentence, and the vector conversion layer is obtained based on BERT model training; by calculating the similarity between the corresponding word vectors of the parallel text and the text to be processed, it is pre-judged whether the keywords in the text to be processed belong to a company entity or a trademark entity.

[0126] Furthermore, the entity naming recognition device 100 includes a vectorization module; the similarity calculation module 103 includes an attention vector submodule and a connection submodule;

[0127] The vectorization module is configured to process the text to be processed and the parallel text through an embedding layer in the recognition model to obtain corresponding first and second vectors;

[0128] The attention vector submodule is used for the attention layer in the vector conversion layer to calculate the attention vector of each word in the to-be-processed text and the parallel text for the keyword based on the first vector and the second vector;

[0129] The connection submodule is used to process the attention vector through two fully connected layers in the vector conversion layer to obtain word vectors corresponding to each word in the text to be processed and the parallel text.

[0130] Specifically, the vectorization module obtains the first vector and the second vector corresponding to the text to be processed and the parallel text through the embedding layer in the recognition model, and the first vector and the second vector are ordinary word vectors. After obtaining the ordinary vectors, the attention vector sub-module calculates the attention vector of each word in the text to be processed and the parallel text for the keyword through the attention layer in the vector conversion layer. The connection sub-module obtains the word vector corresponding to each word in the parallel text and the text to be processed through the two fully connected layers.

[0131] Through the cooperation of the vectorization module, the attention vector submodule and the connection submodule, the text to be processed and the parallel text are first processed by the ordinary embedding layer to obtain the first vector and the second vector. The first vector and the second vector are then processed by the attention layer and the fully connected layer to obtain word vectors that can better represent the text to be processed and the parallel text, thereby improving the accuracy of subsequent similarity judgment and final result output.

[0132] Furthermore, the attention vector submodule includes a matrix multiplication unit, a first calculation unit, a second calculation unit and a fully connected unit;

[0133] The matrix multiplication unit is used to multiply the first vector by the parameter matrix obtained after pre-training to obtain corresponding Q matrix, K matrix and V matrix;

[0134] The first calculation unit is configured to obtain a second result by performing a dot product between the Q matrix and the K matrix, dividing the first result obtained by the dot product by the square root of the corresponding dimension of the Q matrix, and performing a Softmax calculation on the second result to obtain a weight matrix;

[0135] The second calculation unit is configured to multiply the V matrix of the weight matrix to obtain a first matrix;

[0136] The fully connected unit is used to process the first matrix through the fully connected layer in the vector conversion layer to obtain the attention vector.

[0137] Specifically, the matrix multiplication unit multiplies the first vector by the parameter matrix obtained after pre-training to obtain three matrices Q, K, and V, whose formula is Q=x4·W Q K=x4·W K ,V=x4·W V Where W Q , W K , WV That is, the parameter matrix. The first calculation unit uses the Q, K, and V matrices to calculate the proportion of each word in the input text. The weight of the word that only receives the above information is 0. The specific calculation formula is

[0138]

[0139] Where A represents the weight matrix, d k Represents the Q, K, or V matrix dimension. The second computing unit then multiplies the weight matrix with the V matrix of the corresponding batch to obtain a first matrix. The fully connected unit processes the first matrix through the fully connected layer in the vector conversion layer to obtain the attention vector. The fully connected layer includes two basic fully connected networks.

[0140] By integrating the matrix multiplication unit, the first computing unit, the second computing unit, and the fully connected unit, an attention mechanism is introduced to obtain word vectors that better represent the text to be processed and the parallel text, thereby improving the accuracy of subsequent similarity judgments and the final result output.

[0141] Furthermore, the similarity calculation module 103 includes a vector acquisition submodule and a cosine similarity calculation submodule;

[0142] The vector acquisition submodule is configured to acquire a first word vector corresponding to the replacement text in the parallel text, and a second word vector corresponding to the replaced text in the to-be-processed text;

[0143] The cosine similarity calculation submodule is used to calculate the cosine similarity between the first word vector and the second word vector to obtain the similarity between the word vectors corresponding to the text to be processed and the parallel text.

[0144] Specifically, after obtaining the word vectors of the parallel text and the text to be processed, the vector acquisition submodule obtains the first word vector of the replacement text in the parallel text and the second word vector of the replaced text in the text to be processed, that is, extracts the word vectors corresponding to different parts of the parallel text and the text to be processed respectively, and the cosine similarity calculation submodule calculates the cosine similarity between the first word vector and the second word vector.

[0145] By cooperating with the vector acquisition submodule and the cosine similarity calculation submodule, the similarity is calculated to pre-judge whether the keywords in the text to be processed are entity names, thereby improving the judgment speed, and by combining pre-judgment and subsequent judgment to perform entity name recognition, the recognition accuracy is improved.

[0146] A probability calculation module 104 is configured to input the word vector corresponding to the text to be processed into the LSTM layer in the recognition model when the similarity is greater than a preset value, to obtain a type distribution probability corresponding to each word in the text to be processed;

[0147] Specifically, when the similarity is less than a preset value, the probability calculation module 104 proves that the keyword is not a company entity or a trademark entity, and can directly output the result, that is, the text to be processed does not contain a company entity or a trademark entity; when the similarity is greater than the preset value, the vector corresponding to the text to be processed is input into the LSTM layer in the recognition model. The LSTM layer is a bidirectional LSTM neural network, which outputs the probability distribution of the text to be processed being labeled into various types, and obtains the type distribution probability corresponding to each word in the text to be processed, and the LSTM layer can learn the semantic relationship of the text.

[0148] Furthermore, the entity naming recognition apparatus 100 includes an overfitting prevention module;

[0149] The overfitting prevention module is used to process the word vector corresponding to the text to be processed through the Dropout layer in the recognition model to suppress overfitting of the recognition model.

[0150] The Dropout layer is set to suppress model overfitting.

[0151] The entity extraction module 105 is used to input the type distribution probability into the conditional random field layer in the recognition model to obtain the entity naming in the text to be processed.

[0152] Furthermore, the entity naming recognition device 100 includes a linear transformation module;

[0153] The linear transformation module is used for processing the type distribution probability through the Dropout layer and the linear transformation layer in the recognition model.

[0154] By setting the Layer Norm layer, Dropout layer and linear transformation layer to perform normalization, it also suppresses model overfitting.

[0155] By adopting the above-mentioned device, the entity naming recognition device 100 obtains the text to be processed through the coordinated use of the acquisition module 101, the replacement module 102, the similarity calculation module 103, the probability calculation module 104 and the entity extraction module 105, and then randomly selects a preset parallel corpus from the preset corpus, and replaces the text to be processed according to the ratio to obtain a parallel text, so as to facilitate the comparison between the parallel text and the text to be processed, and then the text to be processed and the parallel text are input into the vector conversion layer of the recognition model for vector conversion to obtain the corresponding word vector, and the word vector is calculated according to the word vector. The similarity between the text to be processed and the parallel text is used to determine whether the entity name is polysemous and whether the meaning of the entity name is the same as that in the parallel text. When the similarity is greater than a preset value, it proves that there may be an entity in the text to be processed. The word vector corresponding to the text to be processed is then input into the LSTM layer in the recognition model to obtain the type distribution probability corresponding to each word in the text to be processed. Then, the conditional random field layer is used to process the entity in the text to be processed. In this way, even if the entity name is polysemous, the entity name can be accurately judged, thereby improving the accuracy of entity name recognition.

[0156] The present application also provides a computer device. Figure 3 , Figure 3 This is a basic structural block diagram of the computer device in this embodiment.

[0157] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected through a system bus. It should be noted that the figure only shows a computer device 4 with components 41-43, but it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Among them, those skilled in the art can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application-specific integrated circuits (ASICs), programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0158] The computer device may be a desktop computer, notebook computer, PDA, cloud server, etc. The computer device may interact with the user via a keyboard, mouse, remote control, touchpad, or voice control device.

[0159] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 can be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 can also be an external storage device of the computer device 4, such as a plug-in hard disk equipped on the computer device 4, a smart memory card (SMC), a secure digital (SD) card, a flash memory card, etc. Of course, the memory 41 can also include both the internal storage unit of the computer device 4 and its external storage device. In this embodiment, the memory 41 is generally used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions of the entity naming recognition method. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or are to be output.

[0160] In some embodiments, the processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 42 is generally used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to execute computer-readable instructions stored in the memory 41 or process data, such as computer-readable instructions for executing the entity naming recognition method.

[0161] The network interface 43 may include a wireless network interface or a wired network interface. The network interface 43 is generally used to establish a communication connection between the computer device 4 and other electronic devices.

[0162] This embodiment implements the steps of the entity naming recognition method as described in the above embodiment when a processor executes computer-readable instructions stored in a memory. The method obtains a text to be processed, randomly selects a preset parallel corpus from a preset corpus, and replaces the text to be processed according to a certain ratio to obtain a parallel text, thereby facilitating comparison between the parallel text and the text to be processed. The text to be processed and the parallel text are then input into a vector conversion layer of a recognition model for vector conversion to obtain corresponding word vectors. The similarity between the text to be processed and the parallel text is calculated based on the word vectors to determine whether the entity naming is polysemous and whether the meaning of the entity naming is the same as that in the parallel text. When the similarity is greater than a preset value, it is proved that an entity may exist in the text to be processed. The word vector corresponding to the text to be processed is then input into an LSTM layer of a recognition model to obtain a type distribution probability corresponding to each word in the text to be processed. The word vector is then processed through a conditional random field layer to obtain the entity in the text to be processed. Therefore, even if the entity naming is polysemous, the entity naming can be accurately determined, thereby improving the accuracy of entity naming recognition.

[0163] The embodiment of the present application also provides a computer-readable storage medium, which stores computer-readable instructions. The computer-readable instructions can be executed by at least one processor to enable the at least one processor to perform the steps of the entity naming recognition method as described above, by obtaining the text to be processed, and then randomly selecting a preset parallel corpus from a preset corpus, and replacing the text to be processed according to a certain ratio to obtain a parallel text, so as to facilitate the comparison between the parallel text and the text to be processed, and then inputting the text to be processed and the parallel text into the vector conversion layer of the recognition model for vector conversion to obtain the corresponding word vector, and then replacing the text to be processed and the parallel text according to a certain ratio to obtain a parallel text. The similarity between the text to be processed and the parallel text is calculated based on the word vector to determine whether the entity name is polysemous and whether the meaning of the entity name is the same as that in the parallel text. When the similarity is greater than a preset value, it proves that there may be an entity in the text to be processed. The word vector corresponding to the text to be processed is then input into the LSTM layer in the recognition model to obtain the type distribution probability corresponding to each word in the text to be processed. Then, the conditional random field layer is used for processing to obtain the entity in the text to be processed. In this way, even if the entity name is polysemous, the entity name can be accurately judged, thereby improving the accuracy of entity name recognition.

[0164] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0165] The entity naming recognition apparatus, computer device, and computer-readable storage medium of the above-mentioned embodiments of the present application have the same technical effects as the entity naming recognition method of the above-mentioned embodiments, and will not be elaborated here.

[0166] Obviously, the embodiments described above are only some of the embodiments of the present application, rather than all of the embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions described in the aforementioned specific embodiments, or to make equivalent replacements for some of the technical features therein. Any equivalent structure made using the contents of the present application specification and the accompanying drawings, directly or indirectly used in other related technical fields, is also within the scope of patent protection of the present application.

Claims

1. A method for entity naming recognition, characterized in that: The method comprises: Get the text to be processed; Randomly selecting a preset parallel corpus from a preset corpus, replacing the text to be processed according to a preset ratio, and obtaining a parallel text; Inputting the text to be processed and the parallel text into a vector conversion layer in a recognition model for vector conversion, obtaining word vectors corresponding to the text to be processed and the parallel text, and calculating the similarity between the word vectors corresponding to the text to be processed and the parallel text; When the similarity is greater than a preset value, the word vector corresponding to the text to be processed is input into the LSTM layer in the recognition model to obtain the type distribution probability corresponding to each word in the text to be processed; Inputting the type distribution probability into the conditional random field layer in the recognition model to obtain entity names in the text to be processed; The randomly selecting a preset parallel corpus from the preset corpus and replacing the to-be-processed text according to a preset ratio includes: Determining keywords of the text to be processed based on a preset keyword set and the text to be processed; Based on the text to be processed and the keywords, a preset proportion of the content in the text to be processed is replaced with a preset parallel corpus to obtain a parallel text, wherein the replaced part of the text to be processed does not include keywords, and the replacement operation is based on sentences.

2. The entity naming recognition method according to claim 1, characterized in that: Before inputting the to-be-processed text and the parallel text into the vector conversion layer of the recognition model for vector conversion, the method further includes: The text to be processed and the parallel text are processed by the embedding layer in the recognition model to obtain corresponding first vectors and second vectors; Inputting the text to be processed and the parallel text into a vector conversion layer in a recognition model for vector conversion to obtain word vectors corresponding to the text to be processed and the parallel text includes: The attention layer in the vector conversion layer calculates the attention vectors of each word in the to-be-processed text and the parallel text for the keyword based on the first vector and the second vector; The attention vector is processed by two fully connected layers in the vector conversion layer to obtain word vectors corresponding to each word in the parallel text and the text to be processed.

3. The entity naming recognition method according to claim 2, characterized in that: The attention layer in the vector conversion layer calculates the attention vector of each word in the to-be-processed text for the keyword based on the first vector, including: Multiplying the first vector by the parameter matrix obtained after pre-training to obtain the corresponding Q matrix, K matrix and V matrix; By performing a dot product of the Q matrix and the K matrix, dividing the first result obtained by the dot product by the square root of the corresponding dimension of the Q matrix to obtain a second result, and then performing a Softmax calculation on the second result to obtain a weight matrix; Multiplying the V matrix by the weight matrix to obtain a first matrix; The first matrix is ​​processed by the fully connected layer in the vector conversion layer to obtain the attention vector corresponding to the text to be processed.

4. The entity naming recognition method according to claim 1, characterized in that: Calculating the similarity between the word vectors corresponding to the text to be processed and the parallel text includes: Obtaining a first word vector corresponding to the replacement text in the parallel text and a second word vector corresponding to the replaced text in the to-be-processed text; The cosine similarity between the first word vector and the second word vector is calculated to obtain the similarity between the word vectors corresponding to the to-be-processed text and the parallel text.

5. The entity naming recognition method according to claim 1, characterized in that: Before inputting the word vector corresponding to the to-be-processed text into the LSTM layer in the recognition model, the method further includes: The word vector corresponding to the text to be processed is processed through a Dropout layer in the recognition model to suppress overfitting of the recognition model.

6. The entity naming recognition method according to claim 5, characterized in that: Before inputting the type distribution probability into the conditional random field layer in the recognition model, the method further includes: The type distribution probability is also processed by the Dropout layer and the linear transformation layer in the recognition model.

7. An entity naming recognition device, characterized in that: The device comprises: Acquisition module, used to obtain the text to be processed; A replacement module is used to randomly select a preset parallel corpus from a preset corpus, replace the text to be processed according to a preset ratio, and obtain a parallel text; a similarity calculation module, configured to input the text to be processed and the parallel text into a vector conversion layer in a recognition model for vector conversion, obtain word vectors corresponding to the text to be processed and the parallel text, and calculate the similarity between the word vectors corresponding to the text to be processed and the parallel text; A probability calculation module, configured to input the word vector corresponding to the text to be processed into the LSTM layer in the recognition model when the similarity is greater than a preset value, to obtain a type distribution probability corresponding to each word in the text to be processed; An entity extraction module, configured to input the type distribution probability into the conditional random field layer in the recognition model to obtain entity names in the text to be processed; The replacement module includes a keyword determination submodule and a ratio replacement submodule, wherein: The keyword determination submodule is used to determine the keywords of the text to be processed based on a preset keyword set and the text to be processed; The proportion replacement submodule is used to replace a preset proportion of the content in the text to be processed with a preset parallel corpus based on the text to be processed and the keywords to obtain a parallel text, wherein the replaced part of the text to be processed does not include keywords, and the replacement operation is based on sentences.

8. A computer device, characterized in that: The computer device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, the entity naming recognition method according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, the entity naming recognition method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Named entity identification method and device, storage medium and electronic equipment

    CN112749562A

  • Adversarial text generation method and device, equipment and storage medium

    CN113204974A