Entity recognition method, device, electronic device, and computer-readable storage medium

Through dependency syntactic analysis and part-of-speech matching methods, the problem of place name entities being split and misidentified in named entity recognition is solved, and the accuracy of entity recognition is improved.

CN114417869BActive Publication Date: 2025-09-23ASIAINFO TECH CHINA INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202011176742.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-28
Publication Date
2025-09-23
Estimated Expiration
2040-10-28

AI Technical Summary

Technical Problem

Existing named entity recognition methods have low accuracy in specific scenarios, especially place name entities are easily split, leading to misrecognition.

Method used

By performing dependency syntactic analysis on the sentence to be processed, the candidate position of the target entity type is determined according to the dependency relationship, and whether the part of speech of the candidate entity matches the preset part of speech is judged. Only when it matches, the candidate entity is determined as the target entity.

Benefits of technology

This reduces the misidentification problem of specific entities being split up and improves the accuracy of entity recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114417869B_ABST
    Figure CN114417869B_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide an entity recognition method, device, electronic device and computer-readable storage medium, which relate to the field of text processing technology. The method comprises: obtaining a sentence to be processed; analyzing the sentence to be processed to obtain the dependency relationship contained in the sentence to be processed, and determining the candidate position of the target entity type to be identified in the sentence to be processed based on the dependency relationship, and obtaining the candidate entity corresponding to the candidate position; if the part of speech of the candidate entity corresponding to the candidate position matches the preset part of speech, the candidate entity is determined as the target entity corresponding to the target entity type. The implementation of the present application can solve the problem of low entity recognition accuracy in specific scenarios and improve the entity recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of text processing technology. Specifically, the present application relates to an entity recognition method, device, electronic device and computer-readable storage medium. Background Art

[0002] Named entity recognition (NER) is a fundamental research area in natural language processing (NLP) tasks, with a wide range of applications, including keyword extraction, information retrieval, information extraction, event analysis, machine translation, and intelligent dialogue. NER generally involves identifying noun entities such as names of people, places, and organizations. Within specific domains, corresponding named entities are defined. However, existing NER methods have low accuracy for entity recognition in specific scenarios. Summary of the Invention

[0003] This application provides an entity recognition method, device, electronic device, and computer-readable storage medium that can improve the accuracy of entity recognition. The technical solution includes:

[0004] In the first aspect, an embodiment of the present application provides an entity recognition method, which includes: obtaining a sentence to be processed; analyzing the sentence to be processed to obtain the dependency relationship contained in the sentence to be processed, and determining the candidate position of the target entity type to be identified in the sentence to be processed based on the dependency relationship, and obtaining the candidate entity corresponding to the candidate position; if the part of speech of the candidate entity corresponding to the candidate position matches the preset part of speech, determining the candidate entity as the target entity corresponding to the target entity type.

[0005] In the second aspect, an embodiment of the present application provides an entity recognition device, which includes: a sentence acquisition module for acquiring a sentence to be processed; an entity acquisition module for analyzing the sentence to be processed to obtain the dependency relationship contained in the sentence to be processed, and determining the candidate position of the target entity type to be identified in the sentence to be processed based on the dependency relationship, and acquiring the candidate entity corresponding to the candidate position; an entity recognition module for determining the candidate entity as the target entity corresponding to the target entity type if the part of speech of the candidate entity corresponding to the candidate position matches the preset part of speech.

[0006] In a third aspect, an embodiment of the present application provides an electronic device, comprising: one or more computer programs, wherein the one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and the one or more computer programs are configured to: execute the method described in the first aspect above.

[0007] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when called and executed by a processor, implements the method described in the first aspect above.

[0008] The embodiment of the present application provides an entity recognition method, device, electronic device and computer-readable storage medium, which obtains a sentence to be processed, analyzes the sentence to be processed, obtains the dependency relationship contained in the sentence to be processed, and determines the candidate position of the target entity type to be identified in the sentence to be processed based on the dependency relationship, obtains the candidate entity corresponding to the candidate position, and if the part of speech of the candidate entity corresponding to the candidate position matches the preset part of speech, determines the candidate entity as the target entity corresponding to the target entity type. Thus, the embodiment of the present application can perform dependency syntactic analysis on the sentence to be processed, determine the candidate position where the target entity type may exist in the sentence to be processed based on the dependency relationship contained therein, and then judge whether the part of speech of the candidate entity corresponding to the candidate position matches the preset part of speech, and only when it matches, use the candidate entity as the target entity corresponding to the target entity type, which can reduce the occurrence of the misidentification problem of splitting a specific entity, reduce the recognition error rate, and thus improve the accuracy of entity recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments of the present application.

[0010] Figure 1 A flow chart of an entity recognition method provided in one embodiment of the present application is shown.

[0011] Figure 2 A dependency diagram provided by an exemplary embodiment of the present application is shown.

[0012] Figure 3 A flow chart of an entity recognition method provided in another embodiment of the present application is shown.

[0013] Figure 4 An exemplary embodiment of the present application provides Figure 3 Detailed flowchart of step S220.

[0014] Figure 5 A state transition diagram provided by an exemplary embodiment of the present application is shown.

[0015] Figure 6 A state transition diagram provided by another exemplary embodiment of the present application is shown.

[0016] Figure 7 A flow chart of an entity recognition method provided in yet another embodiment of the present application is shown.

[0017] Figure 8 An exemplary embodiment of the present application provides Figure 7 Detailed flowchart of step S340.

[0018] Figure 9 A flowchart of a method for obtaining the probability that a dependency relationship includes a target entity, provided by an exemplary embodiment of the present application, is shown.

[0019] Figure 10 A module block diagram of an entity recognition device provided by an embodiment of the present application is shown.

[0020] Figure 11 The figure shows a structural block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0021] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and are not to be construed as limiting the present invention.

[0022] It will be understood by those skilled in the art that, unless expressly stated otherwise, the singular forms "a", "an", "said" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" used herein includes all or any units and all combinations of one or more associated listed items.

[0023] In order to make the purpose, technical solutions and advantages of this application clearer, the implementation methods of this application will be described in further detail below.

[0024] At present, named entity recognition technology is relatively complete, and many new technologies have gradually been eliminated and well applied. Different solutions have also been given for different entity types. Existing entity recognition methods mainly include rule-based matching methods, feature template-based methods, and neural network-based methods.

[0025] First, the rule-matching method manually defines specific scenarios of named entities according to Chinese writing standards. For example, occupational names (such as "doctor", "teacher", etc.) are related to personal names and can be used as suffixes to identify personal names; organizational suffixes (such as "company", "university", etc.) can be used to identify organizational names; and place name signs (such as "province", "city", "district", etc.) can identify related place name entities. The rule-matching method defines entity recognition rules and matches the rules with the text to achieve named entity recognition. This method uses Chinese writing standards to identify standard named entities with a high recognition accuracy. However, the scope of entity recognition using manually defined rules is still limited, the generalization ability is weak, and it requires a lot of manpower costs, resulting in low recognition efficiency.

[0026] Second, feature template-based methods first require manual annotation of large-scale corpora to form feature templates. Machine learning methods are then used to learn a labeling model, which is then used to annotate the positions of each word in the sentence to achieve named entity extraction. Commonly used models include Hidden Markov Models (HMMs) and Conditional Random Fields (CRFs). Feature templates are generally based on the positional relationship of entity words in the context of the training corpus. Depending on whether the entity words meet the positional characteristics, they are converted into numerical vectors of 0 or 1. The CRF training model is then used to achieve named entity recognition. This method builds on the rule-matching method and uses machine learning methods to construct a statistical model for named entity recognition, thereby reducing the cost of manually defining rules and improving the model's generalization ability. However, model training relies on a large amount of annotated corpus, which is costly. In addition, feature engineering significantly affects the accuracy and recall of model recognition, resulting in poor scalability.

[0027] Third, neural network-based methods first use word embedding to identify words and then use neural networks to automatically extract features. Currently commonly used methods include the Long Short-Term Memory-Conditional Random Field (LSTM-CRF) model and its improved algorithm, the Bidirectional Long Short-Term Memory-Conditional Random Field (BiLSTM-CRF) model. The BIO annotation set is mainly used (B-PER and I-PER represent the first and non-first characters of people's names, B-LOC and I-LOC represent the first and non-first characters of place names, B-ORG and I-ORG represent the first and non-first characters of organizational names, and O represents that the word is not part of the named entity). Then, a multi-layer neural network training feature is constructed, and a CRF model is used for sequence labeling to predict the labeling result of each word. The optimal path is solved through the Viterbi algorithm to achieve named entity recognition. This method is based on machine learning methods and uses neural networks to self-identify entity feature templates. It can self-learn text content, construct entity features, reduce feature engineering workload, and improve execution efficiency. However, neural networks are relatively complex, have poor interpretability and flexibility, and model optimization is also relatively complex.

[0028] Currently, existing entity recognition methods have low accuracy for noun entity recognition in specific scenarios, and some specific entities, such as place names, are easily split during recognition. For example, in the text "The user lives in Shanshui Huafu and has poor mobile phone signal," existing place name entity recognition methods fail to recognize the actual place name "Shanshui Huafu," only recognizing "Huafu" or "Shanshui," leading to misidentification.

[0029] Based on the problems existing in the above-mentioned existing entity recognition methods, the embodiments of the present application provide an entity recognition method, device, electronic device and computer-readable storage medium. By performing dependency syntactic analysis on the sentence to be processed, the candidate position where the target entity type may exist in the sentence to be processed is determined based on the dependency relationship contained therein, and then it is judged whether the part of speech of the candidate entity corresponding to the candidate position matches the preset part of speech. Only when it matches, the candidate entity is used as the target entity corresponding to the target entity type. This can reduce the occurrence of misidentification problems caused by splitting specific entities, reduce the recognition error rate, and thus improve the accuracy of entity recognition.

[0030] The entity recognition method, device, electronic device, and computer-readable storage medium provided in the embodiments of the present application will be described in detail below through specific embodiments.

[0031] The entity recognition method and entity recognition device of the present application can be set in various electronic devices that can process text, including but not limited to wearable devices, head-mounted devices, medical and health platforms, personal computers, server computers, handheld or laptop devices, mobile terminals (such as mobile phones, personal digital assistants (PDAs), media players, etc.), multi-processor systems, consumer electronic devices, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like.

[0032] See also Figure 1 , Figure 1 The following is a flow chart of an entity recognition method provided by an embodiment of the present application, which can be applied to the above electronic devices. Figure 1 The process shown is described in detail. The entity recognition method may include the following steps:

[0033] S110: Obtain statements to be processed.

[0034] The electronic device may segment the original sentence, obtain the segmented sub-sentences, and use the sub-sentences as the statements to be processed, or directly obtain the original sentence as the statement to be processed, without limitation herein. It should be noted that the original sentence described in the embodiment of the present application may be a sentence that has not been segmented, or a sentence that has not been processed in any way. In addition, the original sentence may be a sentence or a sentence composed of multiple sentences, without limitation herein.

[0035] In some embodiments, when the electronic device segments the original sentence, it may first use a word segmentation tool to segment the original sentence to obtain a word sequence w1, w2, ..., w n , construct the corpus dictionary D, define the punctuation marks and special characters in the original sentence as display separators, and then segment the sentence according to the display separators to obtain n short sentences s1, s2, ... s n , and then obtain one or more short sentences from the n short sentences as sentences to be processed. Among them, the word segmentation tool can be but is not limited to SnowNLP, Thulac, HanLP, and this embodiment of the application does not limit this.

[0036] In one example, for the original sentence "A customer reported that they live in Shanshui Huafu and have no signal. They requested the company to address the issue as quickly as possible. Please assist.", the electronic device can use the HanLP word segmentation tool to segment the original sentence. Based on the HanLP word segmentation results, the electronic device can segment the original sentence using the display separator to obtain five short sentences s1, s2, ..., s5, as shown below:

[0037] s1:[Customer / n, feedback / v]

[0038] s2: [I live in / p, Shanshui / n, Washington / n]

[0039] s3:[no / v,signal / n]

[0040] s4: [require / n, company / nis, to process / vn as soon as possible / d]

[0041] s5:[Please / v, assist / v, handle / vn]

[0042] Taking sentence s1 as an example, "customer" and "feedback" represent the word segmentation results, " / n" and " / v" indicate the noun and verb parts of speech, respectively. For detailed explanations of part-of-speech tagging, please refer to the HanLP part-of-speech tagging comparison table and will not be explained here. The electronic device can then obtain one or more of the five sentences mentioned above as sentences to be processed.

[0043] S120: Analyze the statement to be processed to obtain the dependency relationship contained in the statement to be processed, determine the candidate position of the target entity type to be identified in the statement to be processed according to the dependency relationship, and obtain the candidate entity corresponding to the candidate position.

[0044] Electronic devices can analyze the sentence to be processed using a dependency syntax recognition algorithm and extract dependency relationships within the sentence. These dependency relationships include, but are not limited to, head relationships (HED), subject-verb relationships (SBV), adverbial structures (ADV), preposition-object relationships (POB), and attribute relationships (ATT). These relationships are also known as adjective relationships, which refer to the relationship between a specific term and a central term.

[0045] In an example, taking the sentence to be processed as the short sentence s2 in the above example, the dependency relationship contained in the sentence to be processed obtained by analyzing the sentence to be processed can be as follows: Figure 2 As shown, Figure 2 FIG. 1 shows a dependency relationship diagram provided by an exemplary embodiment of the present application, wherein Root refers to the root node and WP refers to punctuation. Figure 2 As shown, the dependency relations contained in sentence s2 include HED-core relation, SBV-subject-verb relation, ADV-adverbial-verb structure, POB-preposition-object relation, and AIT-attributive-predicate relation.

[0046] After the electronic device analyzes the sentence to be processed and obtains the dependency relationship contained in the sentence to be processed, it can determine the candidate position of the target entity type to be identified in the sentence to be processed based on the dependency relationship and obtain the candidate entity corresponding to the candidate position.

[0047] Among them, the type of the target entity to be recognized can be set according to actual needs. For example, the type of the target entity can be a geographical name entity, a personal name entity, etc., and this application does not limit this. It should be noted that the geographical name entity mentioned in some embodiments can be a broad geographical name entity, that is, it can include geographical name entities and entities related to organizations. For example, "Beijing City" and "Peking University" both belong to the geographical name entities mentioned in the following content of the embodiments of this application. In some other embodiments, the geographical name entity can also only refer to a narrow geographical name entity, and in this case, it does not include entities related to organizations. The embodiments of this application do not limit this.

[0048] In some implementation manners, a mapping relationship between a dependency relationship and an entity corresponding to an entity type included in the dependency relationship may be preset, so that according to the type of the target entity to be recognized, the corresponding dependency relationship can be determined, and according to this dependency relationship, the candidate position of the target entity type in the to-be-processed sentence can be determined, and the candidate entity corresponding to the candidate position can be obtained. For example, a group of words can be determined according to the dependency relationship, and this group of words can include 2 words. Then, the position between the positions of the 2 words in the to-be-processed sentence can be determined as the candidate position, and the word at this candidate position can be obtained as the candidate entity corresponding to the candidate position. Another example is that if a group of words determined according to the dependency relationship is continuous, that is, there are no other words or characters in the middle, this group of words can be used as the candidate entity, and at this time, the position of this group of words in the to-be-processed sentence is the candidate position. It should be noted that the foregoing are only 2 possible examples, and the embodiments of this application are not limited to the above two ways of determining the candidate position and the candidate entity.

[0049] In one example, taking the type of the target entity as a geographical name entity as an example, the electronic device may preset a dependency relationship that may include a geographical name entity. For example, in a sentence containing the preposition "in", the words included in the adverbial-medial structure (that is, the words located between the adverbial-medial structure) or related (such as the attributive-medial relationship) phrases can be extracted as geographical name entities. Then, the electronic device may preset a mapping relationship between the adverbial-medial structure, the attributive-medial relationship and the geographical name entity. Of course, according to actual needs, other or more dependency relationships than the foregoing examples may also be preset, and this is not limited here.

[0050] Further, in the example where the to-be-processed sentence is the short sentence s2 in the above example, the electronic device may use the position between the adverbial-medial structure ("in" - "live") in the short sentence s2 as the candidate position, and obtain the word "Shanshui Huafu" at the candidate position as the candidate entity; the electronic device may also obtain the word "Shanshui Huafu" corresponding to the attributive-medial relationship according to the attributive-medial relationship included in the short sentence s2 as the candidate entity. It should be noted that in the word segmentation result, "Shanshui" and "Huafu" are 2 words, and through the above method, the 2 words can be used as the candidate entity as a whole.

[0051] S130: If the part of speech of the candidate entity corresponding to the candidate position matches the preset part of speech, the candidate entity is determined as the target entity corresponding to the target entity type.

[0052] The electronic device may be pre-set with a preset part of speech, which is associated with the target entity type to be identified and can be set according to actual needs. For example, it can be set according to the part of speech that the target entity type may correspond to in the actual sentence, and the part of speech that the target entity type may correspond to is used as the preset part of speech. Taking the target entity type as a place name entity as an example, the part of speech corresponding to the place name entity is usually a noun, and the noun can be used as the preset part of speech. Among them, nouns can also specifically include proper nouns, common nouns, etc. in some examples, which are not limited here. Of course, according to actual needs, the preset part of speech can also include multiple parts of speech, and this embodiment does not limit their number.

[0053] After determining the candidate entity, the electronic device may obtain the part of speech of the candidate entity corresponding to the candidate position, determine whether the part of speech of the candidate entity matches the preset part of speech, and if so, determine the candidate entity as the target entity corresponding to the target entity type. The candidate entity may contain only one word, that is, only one segmentation result after word segmentation processing. In this case, if the part of speech of the candidate entity belongs to the preset part of speech, it can be determined that the part of speech of the candidate entity matches the preset part of speech.

[0054] In addition, the candidate entity may also contain multiple words. In this case, as an embodiment, if the part of speech of the candidate entity completely matches the preset part of speech (for example, the part of speech of all words contained in the candidate entity belongs to the preset part of speech), it can be determined that the part of speech of the candidate entity matches the preset part of speech. As another embodiment, if the part of speech of the candidate entity partially matches the preset part of speech (for example, the part of speech of a specified number of words contained in the candidate entity belongs to the preset part of speech; for example, the part of speech of the words located at a specified position in the candidate entity belongs to the preset part of speech, such as if the last word in the candidate entity belongs to the preset part of speech, this case also belongs to the partial match between the part of speech of the candidate entity and the preset part of speech), it can also be determined that the part of speech of the candidate entity matches the preset part of speech. The embodiments of the present application are not limited to this.

[0055] In one example, based on the foregoing example, taking the preset word type as a noun for example, according to the word segmentation result, the candidate entity "Shanshui Huafu" includes two words, namely "Shanshui" and "Huafu". The word types of these two words are both nouns, which match the preset word type. Then, the candidate entity "Shanshui Huafu" obtained by splicing and combining the two words can be determined as a geographical name entity, and there will be no misrecognition of splitting the two words "Shanshui" and "Huafu" into two entities and only taking one of the words as a geographical name entity or being unable to determine which word is a geographical name entity. Therefore, the entity recognition method provided in this embodiment can, based on the word segmentation result, perform splicing and combination on the word segmentation result according to the dependency relationship of the to-be-processed sentence, obtain the target entity, so as to accurately recognize the entity corresponding to the target entity type included in the sentence. If the target entity type is a geographical name entity, accurate recognition of the geographical name entity can be achieved. Compared with the existing entity recognition technology, the misrecognition rate of splitting entities can be reduced, and the recognition accuracy can be improved.

[0056] It should be noted that the data preset in the electronic device in the embodiments of the present application, such as the dependency relationship and the mapping relationship between the dependency relationship and the entity corresponding to the entity type, the preset word type, etc., can be either data pre-stored locally in the electronic device or data pre-stored in the server. The electronic device is associated with the server, and the corresponding data can be obtained from the server.

[0057] Therefore, the entity recognition method provided in this embodiment obtains the to-be-processed sentence, then analyzes the to-be-processed sentence to obtain the dependency relationship included in the to-be-processed sentence, and determines the candidate position of the target entity type to be recognized in the to-be-processed sentence according to the dependency relationship, obtains the candidate entity corresponding to the candidate position. If the word type of the candidate entity corresponding to the candidate position matches the preset word type, the candidate entity is determined as the target entity corresponding to the target entity type. Therefore, the embodiments of the present application can perform dependency syntactic analysis on the to-be-processed sentence, determine the possible candidate positions of the target entity type in the to-be-processed sentence according to the dependency relationship included therein, then judge whether the word type of the candidate entity corresponding to the candidate position matches the preset word type, and only when they match, take the candidate entity as the target entity corresponding to the target entity type, which can reduce the occurrence of misrecognition problems of splitting specific entities, reduce the recognition error rate, and thus improve the entity recognition accuracy.

[0058] Please refer to Figure 3 , which shows a schematic flowchart of an entity recognition method provided in another embodiment of the present application. In this embodiment, the method may include:

[0059] S210: Obtain the to-be-processed sentence.

[0060] S220: Construct a state machine for the to-be-processed sentence.

[0061] The state machine may be a finite state machine (FSM), which may also be called a finite automaton (FA). A finite automaton contains a finite set of states, each of which can transition to zero or more states. A finite state automaton is a mathematical model of a system with discrete inputs and outputs. In some embodiments, the state machine may be specifically a deterministic finite automaton (DFA). For a given state belonging to the state machine and a character belonging to the alphabet Σ of the state machine, the deterministic finite state automaton can transition to the next state (which may be the previous state) according to a predetermined state transition function.

[0062] In some embodiments, step S220 may specifically include steps S221-S222. Figure 4 , which shows an exemplary embodiment of the present application. Figure 3 Detailed flow diagram of step S220, step S220 may include:

[0063] S221: Perform word segmentation processing on the sentence to be processed and obtain corresponding word segmentation results.

[0064] In some embodiments, if step S210 is to obtain the sentence to be processed by performing word segmentation on the original sentence, the word segmentation result of the sentence to be processed can be obtained from the word segmentation result of the original sentence, and step S222 and subsequent steps can be continued. For example, the word segmentation result of the original sentence, such as the word sequence w1, w2, ..., w n , obtain the word sequence corresponding to the sentence to be processed as the word segmentation result of the sentence to be processed.

[0065] In other embodiments, after obtaining the sentence to be processed, a word segmentation tool may be used to segment the sentence to be processed and obtain a corresponding word segmentation result, wherein the word segmentation result may include one or more words obtained after the sentence to be processed is segmented.

[0066] S222: Construct a state machine for the sentence to be processed based on the word segmentation result.

[0067] In some embodiments, a DFA of the sentence to be processed can be constructed based on the word segmentation results. In one example, the definition of DFA can be as follows, and can be illustrated as follows: Figure 5 , where q0∈Q is the starting state of DFA, {q2, q3}∈F is the terminal state of DFA, q1∈Q is the intermediate state of DFA, and {a, c, d}∈∑ is the alphabet of DFA.

[0068] M=(Q,∑,δ,q0,F)

[0069] Here, M refers to the state machine and Q refers to the non-empty finite set of states. q is called a state of the state machine M.

[0070] Here, ∑ refers to the input alphabet, and the input character strings are all character strings on ∑, that is, the words obtained by word segmentation of the sentence to be processed. If a corpus dictionary D is constructed for the word sequence composed of these words, then ∑=D.

[0071] Among them, δ refers to the state transfer function, δ: Q×∑, δ(q, a)=p means: when the state machine M reads a character a in state q, it will enter the next state p.

[0072] Among them, q0 (q0∈Q) is the transition start of the state machine M, and the initialization DFA short sentence begins to enter the startup state.

[0073] in, is the terminal state set of the state machine M, It is called the terminal state of the state machine M.

[0074] In some embodiments, the electronic device may use a HanLP word segmentation tool to segment the original sentence and, based on the HanLP word segmentation results, segment the original sentence using display separators to obtain multiple short sentences, such as short sentences s1, s2, ..., s5. Then, based on the word segmentation results, a short sentence initialization state machine is constructed for each short sentence.

[0075] In an example, taking the sentence to be processed as short sentence s2: [self / rr, in / p, landscape / n, Huafu / n, residence / v] as an example, the state machine of short sentence s2 can be constructed based on the word segmentation result, such as Figure 6 As shown, q0 is the starting state of the state machine, q5 is the ending state of the state machine, q1, q2, q3, q4 are the intermediate states of the state machine, and the alphabet of state transition is {self / rr, in / p, landscape / n, Washington / n, residence / v}.

[0076] S230: Analyze the statement to be processed to obtain the dependency relationship contained in the statement to be processed.

[0077] In an example, taking the sentence to be processed as short sentence s2 as an example, the sentence to be processed can be analyzed by the dependency syntax recognition algorithm to extract the dependency relationship of each state in the sentence to be processed, such as Figure 2 As shown in the figure, the words in the alphabet are the input characters or strings of the state transition function. After each word, the state machine will enter the next state until the terminal state.

[0078] S240: Determine the starting position and the ending position of the state machine according to the dependency relationship, and use the position between the starting position and the ending position as the candidate position of the target entity type to be identified in the statement to be processed, and obtain the candidate entity corresponding to the candidate position.

[0079] The electronic device can determine the starting position and ending position of the state machine based on the dependency relationship, and use the position between the starting position and the ending position as the candidate position of the target entity type to be identified in the sentence to be processed, and obtain the words corresponding to the candidate position as the candidate entity.

[0080] Since the input parameters of the state transition function for a state machine include the state and the characters read in that state (characters corresponding to words), the state of the state machine can also be understood as a position in some embodiments, that is, the starting state is equivalent to the starting position, and the ending state is equivalent to the ending position. In some embodiments, the starting state and ending state of the state machine can be first determined based on the dependency relationship, and then the words corresponding to the states between the starting state and the ending state are used as candidate entities.

[0081] It should be noted that, depending on the number of words corresponding to the candidate position, the candidate entity may include one word or multiple (two or more) words. Similarly, depending on the number of words corresponding to the states between the starting state and the ending state, the candidate entity may include one word or multiple (two or more) words.

[0082] In addition, in some embodiments, if there are multiple sets of dependency relationships, the starting position and the ending position of the state machine can be determined according to each set of dependency relationships, and multiple candidate positions can be determined and corresponding multiple candidate entities can be obtained.

[0083] In other embodiments, the starting position and the ending position of the state machine may be determined based on only one set of dependency relationships among the multiple sets of dependency relationships, thereby determining the candidate position and obtaining the corresponding candidate entity. The specific implementation method can be found in the following embodiments and will not be described in detail here.

[0084] S250: If the part of speech of the candidate entity corresponding to the candidate position matches the preset part of speech, the candidate entity is determined as the target entity corresponding to the target entity type.

[0085] It should be noted that, for the parts not described in detail in this embodiment, reference may be made to the corresponding parts of the aforementioned embodiments, and no further details will be given here.

[0086] Therefore, the entity recognition method provided in this embodiment can analyze the dependency relationship in the sentence to be processed based on word segmentation technology, construct a state machine for the sentence to be processed, and automatically identify the state transition process of the state machine according to the dependency relationship. Specifically, the starting position and ending position of the state machine can be determined according to the dependency relationship, and then the position between the starting position and the ending position is used as the candidate position of the target entity type to be identified in the sentence to be processed, and the candidate entity corresponding to the candidate position is obtained. If the part of speech of the candidate entity corresponding to the candidate position matches the preset part of speech, the candidate entity is determined as the target entity to be identified, thereby extracting the target entity.

[0087] In some embodiments, the electronic device can determine the starting position and the ending position of the state machine based on the dependency relationship contained in the sentence to be processed and the probability of the target entity contained in the dependency relationship, and then determine the candidate position and obtain the corresponding candidate entity. Figure 7 , which shows a flow chart of an entity recognition method provided by another embodiment of the present application. In this embodiment, the method may include:

[0088] S310: Obtain statements to be processed.

[0089] S320: Build a state machine for the statement to be processed.

[0090] S330: Analyze the statement to be processed to obtain the dependency relationship contained in the statement to be processed.

[0091] S340: Determine a target dependency relationship from the dependency relationships.

[0092] In some embodiments, the electronic device may determine the target dependency relationship according to the position of the words corresponding to the dependency relationship in the sentence to be processed. In some embodiments, the electronic device may determine the target dependency relationship according to the context of the positions of the words corresponding to the dependency relationship in the sentence to be processed. For example, the dependency relationship corresponding to the first word in the sentence to be processed may be determined as the target dependency relationship. Taking the sentence to be processed as the short sentence s2: [self / rr, in / p, landscape / n, Huafu / n, living / v] as an example, the subject-predicate relationship corresponding to the first word "self" may be determined as the target dependency relationship. Of course, in other embodiments, the dependency relationships corresponding to words in other positions may also be determined as target dependencies according to actual needs.

[0093] Furthermore, in one embodiment, when returning to step S340, the dependency relationship corresponding to the second word can be determined as a new target dependency relationship. When returning to step S340 the next time, the dependency relationship corresponding to the third word can be determined as a new target dependency relationship, and so on. In this way, the target dependency relationships can be determined sequentially from front to back to verify whether the part of speech of the corresponding candidate entity matches the preset part of speech, thereby realizing the recognition of the target entity. Of course, the target dependency relationships can also be determined sequentially from back to front according to actual needs, which is not limited here.

[0094] In other embodiments, the electronic device may also determine the target dependency relationship based on the internal and external relationship of the positions of the words corresponding to the dependency relationship in the sentence to be processed. For example, the dependency relationship corresponding to the outermost word in the sentence to be processed may be determined as the target dependency relationship.

[0095] In an example, taking the sentence to be processed as short sentence s2 (s2: [self / rr, in / p, landscape / n, Huafu / n, residence / v]) as an example, the dependency relationship corresponding to the outermost word is the subject-predicate relationship (self / rr->residence / v), and then the next dependency relationship from the outside to the inside is the preposition-object relationship (in / p, landscape / n, Huafu / n).

[0096] Furthermore, in one embodiment, when returning to execute step S340, the second dependency relationship from outside to inside can also be determined as the new target dependency relationship. The next time the electronic device returns to execute step S340, the third dependency relationship from outside to inside can be determined as the new target dependency relationship, and so on. In this way, the target dependency relationship can be determined from outside to inside in sequence to check whether the part of speech of the corresponding candidate entity matches the preset part of speech in sequence, thereby gradually narrowing the scope of the candidate entities to be checked, so that the number of words contained in each candidate entity to be checked gradually decreases until the part of speech of the candidate entity is locked to match the preset part of speech, and the candidate entity at this time is determined as the target entity, thereby completing the recognition of the target entity. By narrowing the scope in sequence from outside to inside, and judging whether the part of speech of the intermediate candidate entity matches the preset part of speech after each clue, until a match is found, the narrowing can be stopped and the current candidate entity is determined as the target entity.

[0097] In other embodiments, to expand the recognition rules, target dependencies can be determined based on the probabilities of dependencies and target entities contained therein, in addition to the state machine and dependency syntax analysis. For example, the dependency with the highest probability among multiple sets of dependencies can be used as the target dependency, and the starting and ending positions of the state machine can be determined based on this target dependency, thereby determining candidate positions and obtaining corresponding candidate entities.

[0098] In some embodiments, step S340 may specifically include steps S341-342. Specifically, see Figure 8 , which shows an exemplary embodiment of the present application. Figure 7 Detailed flow diagram of step S340, step S340 may include:

[0099] S341: Based on the dependency relationships and the probability that the dependency relationships include the target entity, the dependency relationships are sorted from high to low according to their corresponding probabilities to obtain a dependency relationship sequence.

[0100] The probability that a dependency relationship includes a target entity can be pre-set or obtained through statistical analysis of a specified number of sample corpora. This embodiment is not limited to this. As an implementation method, the probability that a dependency relationship includes a target entity can be obtained through machine learning, as will be described in detail in the following embodiments.

[0101] S342: The dependency relationship that is first in the dependency relationship sequence is used as the target dependency relationship.

[0102] In some embodiments, based on the dependency and the probability that the dependency contains the target entity, the dependency can be sorted from high to low according to its corresponding probability to obtain a dependency sequence, thereby taking the dependency with the highest probability of containing the target entity as the target dependency. Based on this, through subsequent steps, the starting position and the ending position of the state machine can be determined first according to the target dependency that is most likely to contain the target entity, which can improve the efficiency of ultimately identifying the target entity in the statement to be processed, that is, improve the efficiency of entity recognition. In addition, in some embodiments, the electronic device can also determine the target dependency from the dependency in other ways, which are not limited here. For example, a preset dependency sequence can be preset, and the preset dependency sequence can contain one or more preset dependencies. Then, by matching the dependency contained in the statement to be processed with the preset dependency sequence, the preset dependency with the highest matching order can be taken as the target dependency.

[0103] S350: Determine the starting position and the ending position of the state machine based on the target dependency relationship. S360: Use the position between the starting position and the ending position as a candidate position of the target entity type to be identified in the sentence to be processed, and obtain the candidate entity corresponding to the candidate position.

[0104] S370: Determine whether the part of speech of the candidate entity corresponding to the candidate position matches the preset part of speech.

[0105] In some embodiments, if the词性 of the candidate entity corresponding to the candidate position does not match the预设词性, steps S340 - S370 may be repeatedly executed until the词性 of the candidate entity corresponding to the candidate position matches the预设词性. That is, if it does not match, step S340 may be returned to execute, determine a new target dependency relationship from the dependency relationships included in the to - be - processed statement, and execute subsequent steps based on the new target dependency relationship until the词性 of the candidate entity corresponding to the candidate position matches the预设词性.

[0106] In other embodiments, if the词性 of the candidate entity corresponding to the candidate position does not match the预设词性, step S340 may not be returned to execute. That is, only the candidate position and the candidate entity are determined once. If the词性 of the candidate entity corresponding to the candidate position does not match the预设词性, this method may be ended.

[0107] In still other embodiments, if the词性 of the candidate entity corresponding to the candidate position does not completely match the预设词性, the target entity may be determined according to the words in the candidate entity that completely match the预设词性.

[0108] In an example, taking the to - be - processed statement as the short sentence s2 (s2: [自己 / rr, 在 / p, 山水 / n, 华府 / n, 居住 / v]), the target entity type being a place name entity, and the预设词性 being a noun as an example, the electronic device may select the subject - predicate relationship (自己 / rr -> 居住 / v) in the dependency relationship as the target dependency relationship, and determine the start position and end position of the state machine based on the subject - predicate relationship. Specifically, the start position of the state machine may be determined according to the subject (自己 / rr), the end position may be determined according to the predicate (居住 / v), and then the words corresponding to the object - preposition relationship between the subject and the predicate (在 / p, 山水 / n, 华府 / n) are taken as the candidate entity. Since place name entities are often nouns, the预设词性 of place name entities may be a noun. For the前述 candidate entity, in addition to including nouns, it also includes the preposition "在", so it can be determined that the词性 of the candidate entity does not match the预设词性 and belongs to an incomplete match. At this time, the target entity may be determined according to the words in the candidate entity that completely match the noun, and the words composed of the nouns "山水" and "华府" may be used as the target entity. Thus, "山水华府" can be recognized as the target entity, rather than being split and recognizing "山水" or "华府" as the target entity, which may cause mis - recognition.

[0109] S380: Determine the candidate entity as the target entity corresponding to the target entity type.

[0110] It should be noted that for the parts not described in detail in this embodiment, reference may be made to the corresponding parts of the前述 embodiments, and details will not be repeated here.

[0111] It should be noted that the Chinese terms "词性" and "预设词性" do not have exact equivalents in English in the context of this text, so they are directly transliterated for the purpose of maintaining the integrity of the technical content. You may need to adjust them according to the specific technical definitions in your actual scenario.Therefore, the entity recognition method provided in this embodiment is based on the state machine principle, combines word segmentation and dependency syntactic analysis, and constructs a state machine for self-recognition state transition to achieve recognition of the target entity.

[0112] In addition, in some embodiments, the probability that the dependency relationship includes the target entity can be obtained through machine learning. Specifically, before step S340 or step S341, steps S410 to S440 may also be included, such as Figure 9 As shown, the method for obtaining the probability that the dependency relationship contains the target entity may include:

[0113] S410: Obtain sample corpus.

[0114] Among them, the sample corpus has been annotated with the target entity corresponding to the target entity type.

[0115] S420: Annotate the dependency relationships of the sample corpus.

[0116] In some embodiments, the dependency relationships contained in the sample corpus can be annotated, and different dependency relationships can be identified with different numerical values. For example, dependency relationships can be identified using numerical values, where the numerical value 0 identifies other relationships, the numerical value 1 is used to identify the core relationship, the numerical value 2 is used to identify the subject-predicate relationship, the numerical value 3 is used to identify the adverbial structure, the numerical value 4 is used to identify the preposition-object relationship, and the numerical value 5 is used to identify the attributive relationship, etc. Of course, the above is only an example of annotation and does not constitute a limitation to this embodiment. In addition, according to actual needs, letters or other characters can also be used to identify or annotate dependency relationships, and this embodiment does not limit this.

[0117] S430: Determine the dependency relationship of the target entity in the sample corpus.

[0118] Based on the annotated dependency relationships, a dependency relationship including the target entity can be determined. In some embodiments, if at least one word corresponding to the dependency relationship belongs to the target entity, the dependency relationship can be determined as a dependency relationship including the target entity.

[0119] In other embodiments, if a dependency relationship corresponds to at least two words, and at least one other word is contained between the at least two words, then if the at least one other word belongs to the target entity, then the dependency relationship can be determined as a dependency relationship that includes the target entity. In one example, taking a dependency relationship corresponding to two words (such as a subject-predicate relationship corresponding to the subject and predicate) as an example, if a dependency relationship corresponds to a first word and a second word, and a third word is contained between the first and second words, it can be determined whether the third word belongs to the target entity. If so, then the dependency relationship is determined as a dependency relationship that includes the target entity.

[0120] In some other embodiments, if any one of the prerequisite conditions of the foregoing two embodiments is satisfied, the corresponding dependency relationship can be determined as the dependency relationship that includes the target entity. For example, for a sample corpus such as "I live in Landscape Mansion", its word segmentation results include [I / rr, in / p, Landscape / n, Mansion / n, live / v], and the dependency relationships of this sample corpus include a subject-predicate relationship (I / rr -> live / v) and a prepositional-object relationship (in / p, Landscape / n, Mansion / n). If the target entity type is a place name entity, then the place name entity in this sample corpus corresponds to "Landscape Mansion". Since "Landscape Mansion" exists both between the subject and the predicate and is included in the subject-predicate relationship, and also exists as the object in the prepositional-object relationship, it can be determined that the dependency relationships including the place name entity in this sample corpus are the subject-predicate relationship and the prepositional-object relationship.

[0121] S440: Based on all sample corpora, count the number of times each dependency relationship includes the target entity, and obtain the probability of each dependency relationship including the target entity according to the ratio of the number of times to the total number of times that all dependency relationships include the target entity.

[0122] After determining the dependency relationships including the target entity in all sample corpora, the number of times each dependency relationship includes the target entity can be counted, and the probability of each dependency relationship including the target entity can be obtained according to the ratio of the number of times to the total number of times that all dependency relationships include the target entity.

[0123] In some embodiments, the number of times each dependency relationship includes the target entity and the sum of all the numbers of times (denoted as the total number of times) can be counted, and the ratio of the number of times each dependency relationship includes the target entity to the total number of times is used as the probability of each dependency relationship including the target entity. For example, among all the sample corpora, the dependency relationships including the target entity are only the subject-predicate relationship and the prepositional-object relationship. After statistics, the number of times the subject-predicate relationship includes the target entity is 2 times, and the number of times the prepositional-object relationship includes the target entity is 3 times. Then the total number of times is 5 times. The probability that the subject-predicate relationship includes the target entity is 2 / 5 = 40%, and the probability that the prepositional-object relationship includes the target entity is 3 / 5 = 60%.

[0124] It should be noted that the parts not described in detail in this embodiment can refer to the corresponding parts of the foregoing embodiment, and will not be elaborated here.

[0125] Thus, the probability that the dependency relationship includes the target entity can be obtained by machine learning through the aforementioned method, further improving the generalization ability and recognition accuracy of the entity recognition method provided by the embodiment of the present application. Then, on the basis of the aforementioned embodiment, a small amount of sample corpus (text) can be annotated on the basis of the DFA principle and the dependency syntax and word segmentation results to achieve target entity recognition. Thus, from the perspective of state transition of the state machine and dependency syntax analysis, entity recognition is first converted into a classification based on dependency relationships, and then disassembled according to the dependency relationships to extract the target entity, which greatly reduces the scope of analysis and can solve complex problems based on a small amount of text annotation.

[0126] Please refer to Figure 10 , an embodiment of the present application provides a module block diagram of an entity recognition device. The entity recognition device 1000 can be applied to the above-mentioned electronic device. The entity recognition device 1000 may specifically include: a statement acquisition module 1010, an entity acquisition module 1020, and an entity recognition module 1030, wherein:

[0127] The statement acquisition module 1010 is used to acquire the statement to be processed;

[0128] The entity acquisition module 1020 is configured to analyze the statement to be processed to obtain dependency relationships contained in the statement to be processed, determine candidate positions of the target entity type to be identified in the statement to be processed based on the dependency relationships, and obtain candidate entities corresponding to the candidate positions;

[0129] The entity recognition module 1030 is configured to determine the candidate entity as a target entity corresponding to the target entity type if the part of speech of the candidate entity corresponding to the candidate position matches a preset part of speech.

[0130] Furthermore, the entity acquisition module 1020 includes: a state machine construction submodule, a dependency syntax analysis submodule, and a candidate entity acquisition submodule, wherein:

[0131] A state machine construction submodule, used to construct a state machine for the statement to be processed;

[0132] A dependency syntax analysis submodule, configured to analyze the sentence to be processed to obtain dependency relations contained in the sentence to be processed;

[0133] The candidate entity acquisition submodule is used to determine the starting position and the ending position of the state machine according to the dependency relationship, and use the position between the starting position and the ending position as the candidate position of the target entity type to be identified in the statement to be processed, and obtain the candidate entity corresponding to the candidate position.

[0134] Furthermore, the candidate entity acquisition submodule includes: a target relationship acquisition unit, a start and end position determination unit, a candidate entity acquisition unit, and a part-of-speech matching unit, wherein:

[0135] a target relationship acquisition unit, configured to determine a target dependency relationship from the dependency relationships;

[0136] a start and end position determining unit, configured to determine a start position and an end position of the state machine based on the target dependency relationship;

[0137] A candidate entity acquisition unit, configured to take a position between the starting position and the ending position as a candidate position of the target entity type to be identified in the sentence to be processed, and acquire a candidate entity corresponding to the candidate position;

[0138] A part-of-speech matching unit is used to repeatedly perform the steps of determining the target dependency from the dependency, determining the starting position and the ending position of the state machine based on the target dependency, and using the position between the starting position and the ending position as the candidate position of the target entity type to be identified in the sentence to be processed, and obtaining the candidate entity corresponding to the candidate position, if the part-of-speech of the candidate entity corresponding to the candidate position does not match the preset part-of-speech, until the part-of-speech of the candidate entity corresponding to the candidate position matches the preset part-of-speech.

[0139] Furthermore, the target relationship acquisition unit includes: a probability ranking subunit and a target determination subunit, wherein:

[0140] a probability sorting subunit, configured to sort the dependency relationships from high to low according to their corresponding probabilities based on the dependency relationships and the probabilities that the dependency relationships contain target entities, to obtain a dependency relationship sequence;

[0141] The target determination subunit is used to take the dependency relationship located at the first position in the dependency relationship sequence as the target dependency relationship.

[0142] Furthermore, before determining the target dependency relationship from the dependency relationships, the entity recognition device may further include: a sample corpus acquisition module, a dependency relationship annotation module, a dependency relationship determination module, and a probability statistics module, wherein:

[0143] A sample corpus acquisition module is used to acquire sample corpus, wherein the sample corpus has been annotated with a target entity corresponding to the target entity type;

[0144] A dependency annotation module, used to annotate the dependency relationships of the sample corpus;

[0145] A dependency relationship determination module, configured to determine the dependency relationship of the target entity in the sample corpus;

[0146] The probability statistics module is used to count the number of times each dependency relationship contains the target entity based on all sample corpora, and obtain the probability of each dependency relationship containing the target entity based on the ratio of the number of times to the total number of times all dependency relationships contain the target entity.

[0147] It should be noted that the above-mentioned device provided in the embodiment of the present application can implement all the method steps implemented in the above-mentioned method embodiment and can achieve the same technical effect. The parts and beneficial effects of this embodiment that are the same as those in the method embodiment will not be described in detail here.

[0148] In an embodiment of the present application, an electronic device is provided, which includes: a memory and a processor; at least one program, stored in the memory, for being executed by the processor, and compared with the prior art, it can achieve: by obtaining a statement to be processed, then analyzing the statement to be processed to obtain the dependency relationship contained in the statement to be processed, and determining the candidate position of the target entity type to be identified in the statement to be processed based on the dependency relationship, obtaining the candidate entity corresponding to the candidate position, and if the part of speech of the candidate entity corresponding to the candidate position matches the preset part of speech, determining the candidate entity as the target entity corresponding to the target entity type. Thus, the embodiment of the present application can perform dependency syntactic analysis on the statement to be processed, determine the candidate position where the target entity type may exist in the statement to be processed based on the dependency relationship contained therein, and then determine whether the part of speech of the candidate entity corresponding to the candidate position matches the preset part of speech, and only when it matches, use the candidate entity as the target entity corresponding to the target entity type, which can reduce the occurrence of the misidentification problem of splitting a specific entity, reduce the recognition error rate, and thus improve the accuracy of entity recognition.

[0149] In an alternative embodiment, an electronic device is provided, such as Figure 11 As shown, Figure 11 The electronic device 1100 shown includes a processor 1101 and a memory 1103. The processor 1101 and the memory 1103 are connected, for example, via a bus 1102. Optionally, the electronic device 1100 may further include a transceiver 1104. It should be noted that in actual applications, the number of transceivers 1104 is not limited to one, and the structure of the electronic device 1100 does not constitute a limitation on the embodiments of the present application.

[0150] The processor 1101 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor 1101 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.

[0151] Bus 1102 may include a path for transmitting information between the aforementioned components. Bus 1102 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, for example. Bus 1102 may be divided into an address bus, a data bus, a control bus, and so on. For ease of illustration, a single thick line is used in the figure, but this does not indicate that there is only one bus or only one type of bus.

[0152] The memory 1103 may be a ROM (Read Only Memory) or other types of static storage terminals that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage terminals that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), a magnetic disk storage medium or other magnetic storage terminals, or any other medium that can be used to carry or store a desired computer program in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0153] The memory 1103 is used to store the computer program for executing the solution of the present application, and the execution is controlled by the processor 1101. The processor 1101 is used to execute the computer program stored in the memory 1103 to implement the content shown in the above method embodiment.

[0154] Among them, electronic equipment includes but is not limited to: servers, desktops, laptops, etc.

[0155] The embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer is run on a computer, the computer can execute the corresponding content in the above method embodiment. Compared with the prior art, the embodiment of the present application obtains a statement to be processed, analyzes the statement to be processed, obtains the dependency relationship contained in the statement to be processed, and determines the candidate position of the target entity type to be identified in the statement to be processed according to the dependency relationship, obtains the candidate entity corresponding to the candidate position, and if the part of speech of the candidate entity corresponding to the candidate position matches the preset part of speech, the candidate entity is determined as the target entity corresponding to the target entity type. Thus, the embodiment of the present application can perform dependency syntactic analysis on the statement to be processed, determine the candidate position where the target entity type may exist in the statement to be processed according to the dependency relationship contained therein, and then judge whether the part of speech of the candidate entity corresponding to the candidate position matches the preset part of speech, and only when it matches, use the candidate entity as the target entity corresponding to the target entity type, which can reduce the occurrence of the misidentification problem of splitting a specific entity, reduce the recognition error rate, and thus improve the accuracy of entity recognition.

[0156] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0157] The above descriptions are only partial embodiments of the present invention. It should be pointed out that ordinary technicians in this technical field can make several improvements and modifications without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. A method for entity recognition, characterized in that: The method comprises: Obtaining a word segmentation result of the sentence to be processed, wherein the word segmentation result includes multiple words and the part of speech of each word; Analyzing the word segmentation results of the sentence to be processed to obtain dependency relationships contained in the sentence to be processed, determining a candidate position of the target entity type to be identified in the sentence to be processed based on the dependency relationships, and obtaining a candidate entity corresponding to the candidate position, wherein the candidate entity includes at least one word; If the part of speech of the candidate entity corresponding to the candidate position matches the preset part of speech, all words in the candidate entity are determined as a whole as the target entity corresponding to the target entity type.

2. The method according to claim 1, characterized in that The step of analyzing the statement to be processed to obtain dependency relationships contained in the statement to be processed, determining a candidate position of the target entity type to be identified in the statement to be processed based on the dependency relationships, and obtaining a candidate entity corresponding to the candidate position includes: Constructing a state machine for the statement to be processed; Analyzing the statement to be processed to obtain dependency relationships contained in the statement to be processed; The starting position and the ending position of the state machine are determined according to the dependency relationship, and the position between the starting position and the ending position is used as the candidate position of the target entity type to be identified in the statement to be processed, and the candidate entity corresponding to the candidate position is obtained.

3. The method according to claim 2, characterized in that The starting position and the ending position of the state machine are determined according to the dependency relationship, and the position between the starting position and the ending position is used as a candidate position of the target entity type to be identified in the statement to be processed, and the candidate entity corresponding to the candidate position is obtained, including determining a target dependency from the dependencies; Determining a starting position and an ending position of the state machine based on the target dependency; Taking the position between the starting position and the ending position as a candidate position of the target entity type to be identified in the sentence to be processed, and obtaining a candidate entity corresponding to the candidate position; If the part of speech of the candidate entity corresponding to the candidate position does not match the preset part of speech, repeat the steps of determining the target dependency relationship from the dependency relationship, determining the starting position and ending position of the state machine based on the target dependency relationship, and taking the position between the starting position and the ending position as the candidate position of the target entity type to be identified in the sentence to be processed, and obtaining the candidate entity corresponding to the candidate position, until the part of speech of the candidate entity corresponding to the candidate position matches the preset part of speech.

4. The method according to claim 3, characterized in that The determining of the target dependency relationship from the dependency relationships includes: Based on the dependency relationships and the probabilities that the dependency relationships include the target entity, the dependency relationships are sorted from high to low according to their corresponding probabilities to obtain a dependency relationship sequence; The dependency relationship that is located first in the dependency relationship sequence is used as the target dependency relationship.

5. The method according to claim 4, characterized in that Before determining the target dependency relationship from the dependency relationships, the method further includes: Acquire a sample corpus, wherein the sample corpus has been annotated with a target entity corresponding to the target entity type; Annotating the dependency relationships of the sample corpus; Determining the dependency relationship of the target entity in the sample corpus; Based on all sample corpora, the number of times each dependency relationship includes the target entity is counted, and the probability of each dependency relationship including the target entity is obtained based on the ratio of the number of times to the total number of times all dependency relationships include the target entity.

6. The method according to claim 2, characterized in that The state machine for constructing the statement to be processed includes: Perform word segmentation processing on the sentence to be processed and obtain corresponding word segmentation results; A state machine for the sentence to be processed is constructed according to the word segmentation result.

7. The method according to any one of claims 1 to 6, characterized in that The target entity type is a place name entity.

8. An entity recognition device, characterized in that: The device comprises: A sentence acquisition module is used to obtain the word segmentation result of the sentence to be processed, wherein the word segmentation result includes multiple words and the part of speech of each word; an entity acquisition module, configured to analyze the word segmentation results of the sentence to be processed to obtain dependency relationships contained in the sentence to be processed, determine a candidate position of the target entity type to be identified in the sentence to be processed based on the dependency relationships, and obtain a candidate entity corresponding to the candidate position, wherein the candidate entity includes at least one word; The entity recognition module is configured to determine all words in the candidate entity as a whole as a target entity corresponding to the target entity type if the part of speech of the candidate entity corresponding to the candidate position matches a preset part of speech.

9. An electronic device, characterized in that: include: one or more processors; Memory; One or more computer programs, wherein the one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and the one or more computer programs are configured to: perform the entity recognition method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store a computer program, and the computer program is called by a processor to execute the entity recognition method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Non-named entity object extraction method and device, electronic equipment and storage medium

    CN110929520A