Entity recognition method and device, electronic equipment, storage medium and program product
By combining the fusion of character vector representation, vocabulary collection and dependent information representation, using MacBERT and dependent syntax information parser, the problem of unclear entity boundaries is solved, and the multi-scene recognition capability and accuracy of the entity recognition model are improved.
Patent Information
- Application Number
- CN202510513741.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-29
AI Technical Summary
The existing entity recognition technology has the problem of unclear entity boundaries, especially in many scenarios, the ability to identify entity categories is insufficient.
By obtaining the character vector representation and vocabulary collection of the statement to be recognized, combining the dependency information representation, the vocabulary information fusion unit and feature fusion unit are used to enrich the boundary characteristics of characters, use the MacBERT pre-trained language model for character encoding, and extract more comprehensive dependency information through the dependent syntax information parser, and combine it with convolutional neural network, two-way long and short-term memory network and conditional random field for entity recognition.
It improves the entity recognition capability of the entity recognition model in multiple scenarios, solves the problem of unclear entity boundaries, and improves the accuracy of entity recognition.
Smart Images

Figure CN120387453A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly relates to a method, device, electronic device, storage medium and program product for entity recognition. Background Art
[0002] Named Entity Recognition (NER) is an important task in Natural Language Processing (NLP), aiming to identify specific entities in text, such as person names, place names, organization names, time, dates, etc., to help computers understand the structure and meaning of text, and is widely used in fields such as information extraction, question answering systems, and text classification.
[0003] Existing entity recognition technologies are divided into two categories. One category relies on methods such as dictionary rules for recognition, and the other category constructs feature vectors using BERT series models and then combines downstream sequence labeling tasks for entity recognition. However, using methods such as dictionary rules to match and recognize entities in sentences is prone to misrecognition, and cannot solve the problem of nested entities and cannot handle entity category problems in multiple scenarios. The method of constructing feature vectors using BERT series models and then combining downstream sequence labeling tasks for entity recognition can make up for the defects of the dictionary rule method, but there is a defect that the entity boundary is not clear. Summary of the Invention
[0004] This application provides a method, device, electronic device, storage medium and program product for entity recognition, which solves the problem of unclear boundaries in existing named entity recognition.
[0005] This application provides a method for entity recognition, including: Obtain the sentence to be recognized; Input the sentence to be recognized into the entity recognition model to obtain the entity recognition result output by the entity recognition model; Wherein, the entity recognition model includes: A lexical information fusion unit, configured to determine an associated feature representation of a character and a vocabulary based on the vector representation of each character in the sentence to be recognized and the vocabulary set associated with each character; A feature fusion unit, configured to fuse the associated feature representation and the dependency information representation of the sentence to be recognized to obtain a vector feature; the dependency information representation is determined based on multiple types of dependency information of the central entity of the sentence to be recognized; An entity recognition unit, configured to perform entity recognition based on the vector feature to obtain an entity recognition result.
[0006] In one embodiment, determining the association feature representation between characters and vocabulary based on the vector representation of each character in the statement to be recognized and the vocabulary set associated with each character includes: Determining a vocabulary subset of each character at different vocabulary positions according to the vocabulary set associated with each character in the statement to be recognized; Determining the correlation coefficient between each vocabulary in the vocabulary subset and the character according to the vector representation; Determining the sub - association feature representation of each character at different vocabulary positions according to the product sum of the correlation coefficient and the vocabulary; Merging the sub - association feature representations of each character at different vocabulary positions to obtain the association feature representation between characters and vocabulary.
[0007] In one embodiment, determining the correlation coefficient between each vocabulary in the vocabulary subset and the character according to the vector representation includes: Substituting the vector representation and each vocabulary in the vocabulary subset into an activation function to determine the correlation between each vocabulary in the vocabulary subset and the character corresponding to the vector representation; Determining the total correlation between all vocabularies in the vocabulary subset and the character according to the correlation between each vocabulary in the vocabulary subset and the character; Taking the quotient of the correlation between each vocabulary in the vocabulary subset and the character and the total correlation as the correlation coefficient between the vocabulary and the character.
[0008] In one embodiment, the entity recognition based on the vector feature to obtain an entity recognition result includes: Performing local feature extraction on the vector feature to obtain local features; Performing context - dependent feature extraction on the merged vector feature and the local features to obtain dependent features; Generating a label for each character in the statement to be recognized according to the dependent features, and taking the labels of all characters as the entity recognition result of the statement to be recognized.
[0009] In one embodiment, the steps for determining the dependency information representation include: Extracting the part - of - speech of the central entity in the statement to be recognized, the part - of - speech of the dependent entity with the central entity as the head entity, the part - of - speech of the dependent entity with the central entity as the tail entity, the feature representation of the central entity, the feature representation of the dependent entity with the central entity as the head entity, the feature representation of the dependent entity with the central entity as the tail entity, the relationship category with the central entity as the head entity, and the relationship category with the central entity as the tail entity; Fusing the extracted information to obtain the dependency information representation.
[0010] In one embodiment, the entity recognition model further includes: An encoding unit, configured to perform character encoding based on whole-word masking and the position features of each character in the statement to be recognized, so as to obtain a vector representation of each character in the statement to be recognized.
[0011] This application also provides an entity recognition device, including: Obtain a statement to be recognized; Input the statement to be recognized into the entity recognition model to obtain an entity recognition result output by the entity recognition model; Wherein, the entity recognition model includes: A lexical information fusion unit, configured to determine an association feature representation between a character and a vocabulary based on the vector representation of each character in the statement to be recognized and the vocabulary set associated with each character; A feature fusion unit, configured to fuse the association feature representation with the dependency information representation of the statement to be recognized to obtain vector features; the dependency information representation is determined based on multiple types of dependency information of the central entity of the statement to be recognized; An entity recognition unit, configured to perform entity recognition based on the vector features to obtain an entity recognition result.
[0012] This application also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the entity recognition method described in any one of the above is implemented.
[0013] This application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the entity recognition method described in any one of the above is implemented.
[0014] This application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the entity recognition method described in any one of the above is implemented.
[0015] The entity recognition method, device, electronic device, storage medium, and program product provided by this application determine an association feature representation between a character and a vocabulary based on the vector representation of each character in the statement to be recognized and the vocabulary set associated with each character, fuse the association feature representation with the dependency information representation of the statement to be recognized to obtain vector features, and then perform entity recognition based on the vector features to obtain an entity recognition result. The vector features used for entity recognition simultaneously include the association feature representation between a character and a vocabulary and the dependency information representation of the statement to be recognized, enriching the boundary features and feature information of the character, and can solve the problem of unclear entity boundaries, effectively improving the multi-scenario entity recognition ability of the entity recognition model. Description of the Drawings
[0016] To more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0017] Figure 1 It is a schematic flowchart of the entity recognition method provided by the present application.
[0018] Figure 2 It is a schematic structural diagram of the vocabulary information fusion component provided by the present application.
[0019] Figure 3 It is a schematic structural diagram of the dependency syntax information parser component provided by the present application.
[0020] Figure 4 It is a schematic structural diagram of the feature extraction and label generation component provided by the present application.
[0021] Figure 5 It is a schematic structural diagram of the entity recognition model provided by the present application.
[0022] Figure 6 It is a schematic structural diagram of the entity recognition device provided by the present application.
[0023] Figure 7 It is a schematic structural diagram of the electronic device provided by the present application. Detailed implementation manners
[0024] To make the objectives, technical solutions and advantages of the present application clearer, the following will clearly and completely describe the technical solutions in the present application with reference to the drawings in the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present application belong to the scope of protection of the present application.
[0025] It should be noted that all actions of obtaining signals, information or data in the present application are carried out on the premise of complying with the corresponding data protection regulations and policies of the location and with the authorization given by the owner of the corresponding device.
[0026] Figure 1 It is a schematic flowchart of the entity recognition method provided by the present application. As Figure 1 shown, the present application provides an entity recognition method, including step S100-step S200.
[0027] Step S100, obtain the statement to be recognized.
[0028] The statement to be recognized includes Chinese statements and foreign language statements. In the embodiments of the present application, Chinese statements are taken as an example for illustration. The Chinese statement to be recognized contains several Chinese characters.
[0029] Step S200: Input the statement to be recognized into the entity recognition model to obtain the entity recognition result output by the entity recognition model. The entity recognition result is a sequence tag of the statement to be recognized, and the sequence tag contains the tag corresponding to each Chinese character.
[0030] Among them, the entity recognition model includes a lexical information fusion unit, a feature fusion unit, and an entity recognition unit.
[0031] The lexical information fusion unit is used to determine the associated feature representation of the character and the vocabulary based on the vector representation of each character in the statement to be recognized and the vocabulary set associated with each character. Each character is in a different position in the vocabulary in the vocabulary set, and the associated feature representation can enrich the boundary features of the character.
[0032] The feature fusion unit is used to fuse the associated feature representation with the dependency information representation of the statement to be recognized to obtain a vector feature. The dependency information representation is determined based on various types of dependency information of the central entity of the statement to be recognized. Specifically, the dependency information representation in the present application may include the part of speech and feature representation of the central entity of the statement to be recognized, as well as the part of speech, feature representation, and relationship category of the dependency entity with the central entity as an entity in different positions, thereby enriching the feature information of the characters in the statement to be recognized.
[0033] The entity recognition unit is used to perform entity recognition based on the vector feature to obtain the entity recognition result.
[0034] It can be understood that the present application determines the associated feature representation of the character and the vocabulary based on the vector representation of each character in the statement to be recognized and the vocabulary set associated with each character, fuses the associated feature representation with the dependency information representation of the statement to be recognized to obtain a vector feature, and then performs entity recognition based on the vector feature to obtain the entity recognition result. The vector feature used for entity recognition simultaneously includes the associated feature representation of the character and the vocabulary and the dependency information representation of the statement to be recognized, enriching the boundary features and feature information of the character, and can solve the problem of unclear entity boundaries, effectively improving the multi-scenario entity recognition ability of the entity recognition model.
[0035] Based on the above embodiments, as an optional embodiment, the entity recognition model further includes: The encoding unit is used to perform character encoding based on the whole-word mask and the position feature of each character in the statement to be recognized to obtain the vector representation of each character in the statement to be recognized.
[0036] Specifically, the encoding unit can be composed of a pre-trained language model. In the embodiments of the present application, the MacBERT pre-trained language model is taken as an example for illustration.
[0037] MacBERT uses the whole-word masking method for pre-training. That is, in the whole-word masking, if a Chinese character is masked, then all Chinese characters corresponding to the vocabulary of that Chinese character are respectively masked. Through the whole-word masking, the text "MacBERT is a large model trained by XXXX company itself" can be transformed into "[MASK] is a large [MASK][MASK] trained by XXXX company itself". Finally, by replacing [MASK] with the corresponding Chinese characters, it is used as the pre-training objective. The MacBERT-SDI-ML model uses the character set S={c1, c2, c3,..., ck}, where k is the sentence length, as the encoding unit. In addition, the model introduces the position information P={p1, p2, p3... pk} of the characters to enhance the position features of the characters. Fuse S and P and input them into MacBERT to obtain H1={ }, .
[0038] It can be understood that the present application uses the whole-word masking method and character position features for character encoding, which is beneficial to the improvement of the task accuracy rate under Chinese sentences. It can directly replace the traditional dictionary method, make full use of the advantages of the language model, and improve the accuracy rate of entity recognition, etc.
[0039] Based on the above embodiments, as an optional embodiment, determining the association feature representation between characters and vocabulary based on the vector representation of each character and the vocabulary set associated with each character in the to-be-recognized statement includes: Determine the vocabulary subset of each character at different vocabulary positions according to the vocabulary set associated with each character in the to-be-recognized statement; Determine the correlation coefficient between each vocabulary in the vocabulary subset and the character according to the vector representation; Determine the sub-association feature representation of each character at different vocabulary positions according to the product sum of the correlation coefficient and the vocabulary; Merge the sub-association feature representations of each character at different vocabulary positions to obtain the association feature representation between characters and vocabulary.
[0040] Optionally, determining the correlation coefficient between each vocabulary in the vocabulary subset and the character according to the vector representation includes: Substitute the vector representation and each vocabulary in the vocabulary subset into the activation function to determine the correlation between each vocabulary in the vocabulary subset and the character corresponding to the vector representation; Determine the total correlation between all vocabularies in the vocabulary subset and the character according to the correlation between each vocabulary in the vocabulary subset and the character; The quotient of the correlation between each word in the vocabulary subset and the character and the total correlation is used as the correlation coefficient between the word and the character.
[0041] As Figure 2 shown, considering that the character is in different positions in different sentences, this application adopts an attention mechanism and combines the association features between the character and the word from multiple perspectives to enrich the boundary features of the character. Specifically, a multi-view lexical information fusion component (Multi-view Lexcical Information FusionBased Self-attention Mechanism, MLIF) is used to construct a lexical information fusion unit.
[0042] The MLIF component uses an external dictionary to screen all combinations of the text to obtain a character and vocabulary set D = {c1, (c1, c2), (c1, c2, c3),..., c2, (c2, c3),...., ck-1, (ck-1, ck), ck}. Therefore, any character and the set related to this character are subsets of D. It should be noted that Chinese is a composite language, and each Chinese character has its basic meaning and cannot be ignored. Therefore, the character should also be put into D as an element.
[0043] This application takes , as the output of the MLIF component. Among them . are the representations of the character at three positions B, M, and E of the word respectively. The specific calculation formula is as follows.
[0044] Among them, use to represent the vocabulary set related to in the D set. respectively represent the sets where the character is at the beginning, middle, and end of the word. m , n , p represent the sizes of the above sets. In addition, the calculation process of the coefficient is as follows.
[0045] This application respectively adopts , , to construct characters and the vocabulary set , and the relevance of all words in it. The specific representation is shown in Formula 7.
[0046] where represents the parameter related to the character position, represents the ReLU(.) function.
[0047] It can be understood that this application adopts the attention mechanism and combines the association features of characters and words from multiple perspectives, which can effectively enrich the boundary features of characters, and fully considers the position features of characters in different contexts, improving the model's ability to recognize entity boundaries and effectively solving the situation of low entity recognition ability in multiple scenarios caused by simple methods such as traditional dictionary rules.
[0048] On the basis of the above embodiments, as an optional embodiment, the determining step of the represented dependency information includes: extracting the part-of-speech of the central entity of the to-be-recognized sentence, the part-of-speech of the dependent entity with the central entity as the head entity, the part-of-speech of the dependent entity with the central entity as the tail entity, the feature representation of the central entity, the feature representation of the dependent entity with the central entity as the head entity, the feature representation of the dependent entity with the central entity as the tail entity, the relationship category with the central entity as the head entity, and the relationship category with the central entity as the tail entity; fusing the extracted information to obtain the represented dependency information.
[0049] As Figure 3 shown, in order to make up for the problem of imperfect dependency information in the existing methods, this application incorporates a Syntactic Dependency Information Parser (SDIP) component into the model to extract more comprehensive dependency information.
[0050] The SDIP component is constructed using a Chinese dependency syntax analysis tool. In this application embodiment, the DDParser tool is taken as an example for illustration. DDParser is trained based on a large-scale labeled data, adopts a more simple and understandable labeling relationship, and can extract the dependency information of Chinese entities more accurately and efficiently. Therefore, this paper uses the DDParser tool to respectively analyze the part-of-speech (pos) of the central entity, the part-of-speech of the dependent entity with the central entity as the head entity ( ), the part-of-speech (pos) of the dependent entity with the central entity as the tail entity, the feature representation (ent) of the central entity, and the feature representation of the dependent entity with the central entity as the head entity ( ), the feature representation (ent) of the dependent entity with the central entity as the tail entity, and the relationship category with the central entity as the head entity ( ), and the relationship category with the central entity as the tail entity ( ) are extracted. Then, the extracted information is input into a multi-layer perceptron (MLP) neural network to fully integrate various types of dependency information of the entity, and the final dependency information representation T = { , , ..., } is obtained. . Finally, T is integrated into the character representation to enhance the features of the character representation. The calculation process of the dependency feature of entity i is shown in Equation 8.
[0051] Among them, , . , , and represent the trainable parameters related to part-of-speech, entity representation, and dependency relationship respectively. Define the operator as the inner product of the three parameters and the above three lists respectively. MLP(.) represents the multi-layer perceptron network. f(.) represents ReLU(.).
[0052] It should be noted that the model uses characters as the encoding unit, which is different from the word segmentation granularity of the SDIP component. In order to integrate the information extracted by SDIP into the character representation, the features extracted by SDIP are respectively integrated into each character related to the vocabulary. Let H3 = { , , , ..., } represent the features after fusion.
[0053] It can be understood that this application uses a syntactic dependency information parser (SDIP) to extract richer dependency information, including but not limited to dependency type information, dependent entity information, dependency relationship information, etc., further enriching the features of the entity representation, thereby effectively solving the problem of unclear entity boundaries that cannot be solved by traditional solutions.
[0054] On the basis of the above embodiments, as an optional embodiment, the entity recognition unit includes a convolutional neural network layer, a BILSTM layer, and a CRF layer; The convolutional neural network layer is used to perform local feature extraction on the vector features after merging the associated feature representation and the dependency information representation, so as to obtain local features; The BILSTM layer is used to perform context-dependent feature extraction on the merged vector features and the local features, so as to obtain dependency features; The CRF layer is used to generate labels for each character of the sentence to be recognized according to the dependency features, and use the labels of all characters as the entity recognition result of the sentence to be recognized.
[0055] As Figure 4 shown, due to the different lengths and positions of Chinese words, their semantics are also different. In this application, a CNN neural network is introduced before the BILSTM layer to extract local context information of the sentence. First, the associated feature representation and the dependency information representation are merged to maintain the integrity of the lexical information features; secondly, the convolutional neural network layer (CNN) is used to complete the local feature extraction of the merged vector features; the local features and the vector features are merged, aiming to add local features to the vector features in the previous step; then, the BILSTM network is used to take advantage of long-distance information extraction to perform dependency feature extraction; finally, since there is a strong correlation between entity category labels, the CRF layer is used to ensure the accuracy of the output label sequence.
[0056] It can be understood that this application uses CNN to take into account the local features of the sentence and uses BILSTM to take into account the long-term context features, and uses CRF to complete the extraction of the final label entity, which can comprehensively integrate the two types of features and improve the accuracy of entity recognition, etc.
[0057] As Figure 5As shown in the figure, the entity recognition model provided by this application can achieve entity recognition by integrating dependency syntactic information and multi-view lexical information of MacBERT. The entity recognition model extracts richer dependency information through a Syntactic Dependency Information Parser (SDIP). The entity recognition model also includes a multi-view lexical information fusion component (Multi-view Lexical Information Fusion Based Self-attention Mechanism, MLIF), which represents characters as a combination of three parts: B (beginning of the word), M (middle of the word), and E (end of the word) according to the position of the characters in the word, fully considering the position features of the characters in different contexts and improving the model's ability to recognize entity boundaries. The pre-trained language model MacBERT is used as the encoder for Chinese characters, and the advantage of MacBERT in extracting Chinese features is utilized to improve the quality of Chinese character representation. Finally, the BILSTM+CRF component is used to generate the sequence tags for the recognition task. The BILSTM+CRF component takes into account both the local features and long-term features of the sentence to be recognized, can comprehensively integrate the two types of features, improve the accuracy of entity recognition, etc., and has high practical value and application prospects. It can also be extended to scenarios such as text question answering and text writing.
[0058] The entity recognition device provided by this application will be described below. The entity recognition device described below can be correspondingly referred to the entity recognition method described above.
[0059] Figure 6 is a schematic structural diagram of the entity recognition device provided by this application, as Figure 6 shown, this application also provides an entity recognition device, including the following modules: An acquisition module 610, configured to acquire a sentence to be recognized; A recognition module 620, configured to input the sentence to be recognized into the entity recognition model to obtain the entity recognition result output by the entity recognition model; Among them, the entity recognition model includes: A lexical information fusion unit, configured to determine the associated feature representation of characters and words based on the vector representation of each character in the sentence to be recognized and the set of words associated with each character; A feature fusion unit, configured to fuse the associated feature representation with the dependency information representation of the sentence to be recognized to obtain vector features; the dependency information representation is determined based on multiple types of dependency information of the central entity of the sentence to be recognized; An entity recognition unit, configured to perform entity recognition based on the vector features to obtain an entity recognition result.
[0060] In one embodiment, determining the association feature representation between characters and vocabulary based on the vector representation of each character in the statement to be recognized and the vocabulary set associated with each character includes: Determine the vocabulary subsets of each character at different vocabulary positions according to the vocabulary set associated with each character in the statement to be recognized; Determine the correlation coefficient between each vocabulary in the vocabulary subset and the character according to the vector representation; Determine the sub - association feature representation of each character at different vocabulary positions according to the product sum of the correlation coefficient and the vocabulary; Merge the sub - association feature representations of each character at different vocabulary positions to obtain the association feature representation between characters and vocabulary.
[0061] In one embodiment, determining the correlation coefficient between each vocabulary in the vocabulary subset and the character according to the vector representation includes: Substitute the vector representation and each vocabulary in the vocabulary subset into the activation function to determine the correlation between each vocabulary in the vocabulary subset and the character corresponding to the vector representation; Determine the total correlation between all the vocabularies in the vocabulary subset and the character according to the correlation between each vocabulary in the vocabulary subset and the character; Take the quotient of the correlation between each vocabulary in the vocabulary subset and the character and the total correlation as the correlation coefficient between the vocabulary and the character.
[0062] In one embodiment, the entity recognition based on the vector features to obtain the entity recognition result includes: Extract local features from the vector features to obtain local features; Extract context - dependent features from the merged vector features and the local features to obtain dependent features; Generate labels for each character in the statement to be recognized according to the dependent features, and use the labels of all characters as the entity recognition result of the statement to be recognized.
[0063] In one embodiment, the steps for determining the dependency information representation include: Extract the part - of - speech of the central entity in the statement to be recognized, the part - of - speech of the dependent entity with the central entity as the head entity, the part - of - speech of the dependent entity with the central entity as the tail entity, the feature representation of the central entity, the feature representation of the dependent entity with the central entity as the head entity, the feature representation of the dependent entity with the central entity as the tail entity, the relationship category with the central entity as the head entity, and the relationship category with the central entity as the tail entity; Fuse the extracted information to obtain the dependency information representation.
[0064] In one embodiment, the entity recognition model further includes: An encoding unit, configured to perform character encoding based on whole-word masking and the position features of each character in the statement to be recognized, so as to obtain a vector representation of each character in the statement to be recognized.
[0065] Figure 7 The figure illustrates a schematic diagram of the entity structure of an electronic device, as Figure 7 shown. The electronic device may include: a processor 710, a communication interface 720, a memory 730, and a communication bus 740. Among them, the processor 710, the communication interface 720, and the memory 730 complete mutual communication through the communication bus 740. The processor 710 can call the logical instructions in the memory 730 to execute the entity recognition method, and the method includes: Obtain a statement to be recognized; Input the statement to be recognized into the entity recognition model to obtain the entity recognition result output by the entity recognition model; Wherein, the entity recognition model includes: A lexical information fusion unit, configured to determine the associated feature representation of a character and a vocabulary based on the vector representation of each character in the statement to be recognized and the vocabulary set associated with each character; A feature fusion unit, configured to fuse the associated feature representation and the dependency information representation of the statement to be recognized to obtain vector features; the dependency information representation is determined based on various types of dependency information of the central entity of the statement to be recognized; An entity recognition unit, configured to perform entity recognition based on the vector features to obtain an entity recognition result.
[0066] In addition, when the logical instructions in the foregoing memory 730 are implemented in the form of a software functional unit and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0067] On the other hand, the present application also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the entity recognition method provided by each of the above methods. The method includes: Obtain the statement to be recognized; Input the statement to be recognized into the entity recognition model to obtain the entity recognition result output by the entity recognition model; Wherein, the entity recognition model includes: A vocabulary information fusion unit, configured to determine the associated feature representation of a character and a vocabulary based on the vector representation of each character in the statement to be recognized and the vocabulary set associated with each character; A feature fusion unit, configured to fuse the associated feature representation and the dependency information representation of the statement to be recognized to obtain a vector feature; the dependency information representation is determined based on multiple types of dependency information of the central entity of the statement to be recognized; An entity recognition unit, configured to perform entity recognition based on the vector feature to obtain an entity recognition result.
[0068] On another aspect, the present application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the entity recognition method provided by each of the above methods. The method includes: Obtain the statement to be recognized; Input the statement to be recognized into the entity recognition model to obtain the entity recognition result output by the entity recognition model; Wherein, the entity recognition model includes: A vocabulary information fusion unit, configured to determine the associated feature representation of a character and a vocabulary based on the vector representation of each character in the statement to be recognized and the vocabulary set associated with each character; A feature fusion unit, configured to fuse the associated feature representation and the dependency information representation of the statement to be recognized to obtain a vector feature; the dependency information representation is determined based on multiple types of dependency information of the central entity of the statement to be recognized; An entity recognition unit, configured to perform entity recognition based on the vector feature to obtain an entity recognition result.
[0069] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0070] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0071] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.
Claims
1. An entity recognition method, characterized in that, including: Obtain the statement to be recognized; Input the statement to be recognized into the entity recognition model to obtain the entity recognition result output by the entity recognition model; Wherein, the entity recognition model includes: A lexical information fusion unit for determining the associated feature representation of characters and vocabulary based on the vector representation of each character in the statement to be recognized and the vocabulary set associated with each character; A feature fusion unit for fusing the associated feature representation and the dependency information representation of the statement to be recognized to obtain vector features; the dependency information representation is determined based on multiple types of dependency information of the central entity of the statement to be recognized; An entity recognition unit for performing entity recognition based on the vector features to obtain an entity recognition result.
2. The entity recognition method according to claim 1, wherein The determining the associated feature representation of characters and vocabulary based on the vector representation of each character in the statement to be recognized and the vocabulary set associated with each character includes: Determine the vocabulary subset of each character at different vocabulary positions according to the vocabulary set associated with each character in the statement to be recognized; Determine the correlation coefficient between each vocabulary in the vocabulary subset and the character according to the vector representation; Determine the sub-associated feature representation of each character at different vocabulary positions according to the product sum of the correlation coefficients and the vocabulary; Merge the sub-associated feature representations of each character at different vocabulary positions to obtain the associated feature representation of characters and vocabulary.
3. The entity recognition method according to claim 2, wherein The determining the correlation coefficient between each vocabulary in the vocabulary subset and the character according to the vector representation includes: Substitute the vector representation and each vocabulary in the vocabulary subset into the activation function to determine the correlation between each vocabulary in the vocabulary subset and the character corresponding to the vector representation; Determine the total correlation between all the vocabulary in the vocabulary subset and the character according to the correlation between each vocabulary in the vocabulary subset and the character; Take the quotient of the correlation between each vocabulary in the vocabulary subset and the character and the total correlation as the correlation coefficient between the vocabulary and the character.
4. The entity recognition method according to claim 1, wherein The performing entity recognition based on the vector features to obtain an entity recognition result includes: Perform local feature extraction on the vector features to obtain local features; Perform context-dependent feature extraction on the merged vector features and the local features to obtain dependency features; Generate labels for each character of the statement to be recognized according to the dependency features, and use the labels of all characters as the entity recognition result of the statement to be recognized.
5. The entity recognition method according to claim 1, wherein The determining step of the dependency information representation includes: Extract the part-of-speech of the central entity of the statement to be recognized, the part-of-speech of the dependent entity with the central entity as the head entity, the part-of-speech of the dependent entity with the central entity as the tail entity, the feature representation of the central entity, the feature representation of the dependent entity with the central entity as the head entity, the feature representation of the dependent entity with the central entity as the tail entity, the relationship category with the central entity as the head entity, and the relationship category with the central entity as the tail entity; Fuse the extracted information to obtain the dependency information representation.
6. The entity recognition method according to any one of claims 1-5, characterized in that The entity recognition model further includes: An encoding unit for performing character encoding based on whole-word masking and the position features of each character in the statement to be recognized to obtain the vector representation of each character in the statement to be recognized.
7. An entity recognition device, characterized in that, including; An acquisition module, configured to acquire a statement to be recognized; A recognition module, configured to input the statement to be recognized into an entity recognition model, and obtain an entity recognition result output by the entity recognition model; Wherein, the entity recognition model includes: A lexical information fusion unit, configured to determine an associated feature representation of a character and a vocabulary based on a vector representation of each character in the statement to be recognized and a vocabulary set associated with each character; A feature fusion unit, configured to fuse the associated feature representation and a dependency information representation of the statement to be recognized to obtain a vector feature; the dependency information representation is determined based on multiple types of dependency information of a central entity of the statement to be recognized; An entity recognition unit, configured to perform entity recognition based on the vector feature to obtain an entity recognition result.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, the entity recognition method according to any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the entity recognition method according to any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the entity recognition method according to any one of claims 1 to 6 is implemented.