A method and device for quickly extracting organization names from text

By constructing the AC automata and using word feature information for scoring, the problems of high cost of institutional name extraction, poor real-time and accuracy, and difficult cross-language extraction in the prior art are solved, and efficient and accurate institutional name extraction is achieved.

CN114676214BActive Publication Date: 2025-06-10BEIJING KNOWLEDGE ATLAS TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210200301.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-01-14
Filing Date
2022-03-02
Publication Date
2025-06-10
Estimated Expiration
2042-03-02

AI Technical Summary

Technical Problem

The existing text organization name extraction methods are costly, have poor real-time and accuracy, and are difficult to extract across language organizations.

Method used

By obtaining the candidate organization list, scoring based on the word feature information of the organization name, an AC automata is constructed to match text, and the name of the organization with the highest score is selected.

Benefits of technology

It reduces the cost of institutional name extraction in text, improves the real-time and accuracy of extraction, and reduces the difficulty of cross-language institutional name extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114676214B_ABST
    Figure CN114676214B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and apparatus for quickly extracting organization names from text. The method includes: obtaining a candidate organization list, where the candidate organization list includes at least one organization name; scoring the organization names according to the word feature information constituting the organization names and obtaining a scoring result to calculate the importance degree of the organization names, and constructing an AC automaton according to the organization names constituting the candidate organization list; where the word feature information includes multiple of: the number of times a word appears, rarity, and length; inputting the text to be extracted into the AC automaton, and performing text matching through the constructed AC automaton to obtain the organization names included in the text to be extracted; and screening out the organization name with the highest score from the organization names included in the text to be extracted according to the importance degree of the organization names. The present invention reduces the cost of the extraction method, improves the real-time performance and accuracy, and reduces the difficulty of cross-language organization extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information retrieval, and in particular to a method and device for rapidly extracting institution names from text. Background Art

[0002] In the process of building the academic knowledge graph, the extraction of institution names is the basis for establishing institution entities and is also the core operation for entity association in the academic knowledge graph. Specifically, in the construction process, the extraction of the institution names of the paper authors can establish edges between the paper authors, papers and institutions; the extraction of the institution names of patent assignees can establish edges between patents and the institutions. The establishment of institution entities and the establishment of edge relationships can greatly facilitate users to quickly understand the institution and mine more useful information from the graph.

[0003] The traditional method of extracting institution names uses a sequence tagging-based method to extract institution names from text. This method first needs to collect training data, and combine word segmentation, part-of-speech tagging, and dependency trees to form features; then, the semantic features of the institution name are extracted through a certain model, and a sequence tagging model for the institution name is obtained through conditional random field training; finally, the model is used to perform sequence tagging on the text to extract the institution name information. This method relies on multiple steps such as part-of-speech tagging of the text and setting of sequence rules. It is slow, and its accuracy is greatly affected by the accuracy of word segmentation; in cross-language scenarios, since the sequence rules of each language are quite different, it is necessary to train the model specifically for each language, which is very costly to train and apply and inconvenient to use.

[0004] Another prior art uses a dictionary tree to extract organization names from text. Compared with the N-times KMP string matching method, this method reduces the time complexity of the algorithm from O(T*N*average word length) to O(T*average depth of the dictionary tree), where T is the length of the text and N is the number of organization strings in the dictionary. Since the common information of all organization strings is taken into account, this method is very fast when the match is successful. However, when the match fails, this method does not fully utilize the information obtained when there is a mismatch.

[0005] In summary, the existing methods for extracting institution names are costly, have poor real-time performance and accuracy, and are difficult to extract institutions across languages. Summary of the invention

[0006] The present invention aims to solve one of the technical problems in the related art at least to a certain extent.

[0007] To this end, the purpose of the present invention is to solve the problems of high cost, poor real-time performance and accuracy, and difficulty in cross-language organization extraction in existing text organization name extraction methods, and proposes a method for quickly extracting organization names from text.

[0008] Another object of the present invention is to provide an apparatus for quickly extracting organization names from text.

[0009] To achieve the above object, on the one hand, the present invention provides a method for quickly extracting organization names from text, including the following steps:

[0010] Obtain a candidate organization list, where the candidate organization list includes at least one organization name; score the organization names according to the word feature information constituting the organization names to obtain a scoring result, so as to calculate the importance degree of the organization names, and construct an AC automaton according to the organization names in the candidate organization list; wherein, the word feature information includes multiple ones of the number of word occurrences, rarity, and length; input the text to be extracted into the AC automaton, and perform text matching through the constructed AC automaton to obtain the organization names included in the text to be extracted; according to the importance degree of the organization names, screen out the organization name with the highest score from the organization names included in the text to be extracted by the AC automaton.

[0011] The method for quickly extracting organization names according to the embodiments of the present invention has high robustness when extracting organization names, and solves the problems of high cost, poor real-time performance and accuracy, and difficulty in cross-language organization extraction in the existing methods for extracting organization names.

[0012] In addition, the method for quickly extracting organization names according to the above embodiments of the present invention may further have the following additional technical features:

[0013] Further, the obtaining of the candidate organization list includes: obtaining the template type and organization-related organization pages from the general knowledge base as the seed set; based on the seed set, obtaining the corresponding pages of each organization in the general knowledge base according to the language link; and using the redirect entry of each organization in the general knowledge base as the name of each organization.

[0014] Further, the scoring of the organization names according to the word feature information constituting the organization names to obtain a scoring result includes: predefining a scoring rule for scoring the organization names. For a given name, if it contains double-byte characters, it is a double-byte name, otherwise it is a single-byte name.

[0015] Further, for the double-byte name, the score of the double-byte name is the number of characters constituting the double-byte name.

[0016] Further, for the single-byte name, obtain the total number of all single-byte names, and count the occurrence times of the words that make up all single-byte names; according to the total number and the occurrence times, calculate the importance degree of each word in each single-byte name; sort and list the importance degrees, and obtain the score of the single-byte name according to the list.

[0017] Further, constructing the AC automaton according to the institution names in the candidate institution list includes: using all the institution names in the candidate institution list as the first pattern strings of the AC automaton, and constructing the first pattern strings into a prefix tree; based on the prefix tree, construct the mismatch pointers.

[0018] Further, using all the institution names in the candidate institution list as the first pattern strings of the AC automaton, and constructing the first pattern strings into a prefix tree includes: starting from the root node, sequentially insert the second pattern strings into the prefix tree; wherein, if the second pattern string is a single-byte name, add spaces at the head and tail during indexing, and if it is a double-byte name, directly index; transfer along the current character in the second pattern string on the prefix tree, and create a node if the node does not exist; for the end point, mark the entity corresponding to the second pattern string and the score of the second pattern string.

[0019] Further, constructing the mismatch pointers based on the prefix tree includes: performing a breadth-first traversal on the prefix tree; for each child node i of the current node n, traverse k along the mismatch pointer of n, if k also has a child node i, then the mismatch pointer of the child node i of n points to the child node i of k; otherwise the mismatch pointer of the child node i of n points to the root node.

[0020] To achieve the above object, on the other hand, the present invention proposes a device for quickly extracting institution names in text, including:

[0021] An acquisition module, configured to acquire a candidate institution list, where the candidate institution list includes at least one institution name;

[0022] A scoring module, configured to score the institution name according to the word feature information that makes up the institution name and obtain a scoring result, so as to calculate the importance degree of the institution name, and construct an AC automaton according to the institution names in the candidate institution list; wherein, the word feature information includes multiple of: word occurrence times, rarity, and length;

[0023] A matching module, configured to input the text to be extracted into the AC automaton, and perform text matching through the constructed AC automaton to obtain the institution names included in the text to be extracted;

[0024] A screening module, configured to screen out the organization name with the highest score from the organization names included in the text to be extracted from the AC automaton according to the importance degree of the organization name.

[0025] When the apparatus for quickly extracting organization names in the embodiments of the present invention extracts organization names, it has high robustness, and solves the problems of high cost, poor real-time performance and accuracy, and difficulty in cross-language organization extraction in the existing organization name extraction methods.

[0026] Advantages of the present invention:

[0027] The present invention reduces the cost of the organization name extraction method in the text, improves the real-time performance and accuracy of the organization name extraction, and reduces the difficulty of cross-language organization name extraction.

[0028] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. Description of the Drawings

[0029] The above-mentioned and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, wherein:

[0030] Figure 1 is a flowchart of a method for quickly extracting organization names in the embodiments of the present invention;

[0031] Figure 2 is a schematic flowchart of the construction of an AC automaton in the embodiments of the present invention;

[0032] Figure 3 is a schematic flowchart of the extraction of organization names in the text in the embodiments of the present invention;

[0033] Figure 4 is a schematic structural diagram of an apparatus for quickly extracting organization names in the embodiments of the present invention. Detailed Embodiments

[0034] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present invention will be described in detail below with reference to the drawings and in conjunction with the embodiments.

[0035] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0036] The method and device for quickly extracting the organization names in the text according to the embodiments of the present invention will be described below with reference to the accompanying drawings. First, the method for quickly extracting the organization names in the text according to the embodiments of the present invention will be described with reference to the accompanying drawings.

[0037] Figure 1 It is a flowchart of the method for quickly extracting the organization names in the text according to an embodiment of the present invention.

[0038] As Figure 1 shown, the method for quickly extracting the organization names in the text includes the following steps:

[0039] Step S1, obtain a candidate organization list, where the candidate organization list includes at least one organization name.

[0040] It can be understood that the organization list is composed of an organization entity and a list of names representing the organization. The entities and names of the present invention can be generated through general knowledge bases such as Baidu Encyclopedia, Wikidata, Wikipedia or proprietary knowledge bases such as ROR (https: / / ror.org / ), GRID (https: / / www.grid.ac / ), etc.

[0041] As an example, the present invention obtains the organization list through Wikipedia, and this method includes but is not limited to the following steps:

[0042] The first step: Obtain all organization pages related to the template type and organization from the English Wikipedia as a seed set;

[0043] The second step: For each organization page, obtain the corresponding pages of the organization in all language Wikipedias according to the language links;

[0044] The third step: Take the redirect entries of the organization in all language Wikipedias as the names of the organization.

[0045] So far, the embodiments of the present invention have constructed a list of candidate organizations, and for each organization in the list, there is a corresponding set of multi-language organization names.

[0046] It should be understood that in actual operation, the organization entity and the name list can be supplemented and maintained from different data sources according to actual needs.

[0047] Step S2: Score the organization name based on the word feature information of the constituent organization name to obtain a scoring result, so as to calculate the importance degree of the organization name, and construct an AC automaton according to the organization names in the candidate organization list; wherein, the word feature information includes multiple of: the number of occurrences of the word, rarity, and length.

[0048] It can be understood that an organization name may be composed of multiple words. For example, "tsinghua university" contains the words "tsinghua" and "university". Score the organization name according to the information (number of occurrences, rarity, length, etc.) of the words that make up the organization name to measure the importance degree of this name.

[0049] Specifically, when scoring the organization name based on the word feature information of the organization name, considering that there may be an inclusion relationship between organization names, shorter organization strings will also be matched in a longer organization string. For example, "Northwestern University" and "Peking University". Therefore, the embodiments of the present invention need to score the importance degree of each organization name, and extract the organization name with the highest importance degree when multiple possible results are matched. The embodiments of the present invention define a scoring rule for the importance degree of the organization name. For a given name, if it contains double-byte characters, it is a double-byte name, otherwise it is a single-byte name. Then this rule may include the following two types of operations:

[0050] For the double-byte name N: The score Score(N) is the number of characters that make up the name N, that is, Score(N)=len(N). For example, Score("Northwestern University") = 4, Score("Peking University") = 2.

[0051] For the single-byte name:

[0052] Statistical total number T of all single-byte names;

[0053] Statistical number of occurrences of the words that make up all single-byte names, and record WC(w) as the number of occurrences of the word w, where the word is each substring formed after separating the single-byte name by spaces;

[0054] For each word w 1 , w 2 ,..., w n ) that makes up the single-byte name N=(w i (i∈n), its importance degree I(w i ) The calculation formula is

[0055] Sort the importance degrees of each word that makes up the name N to obtain a list L=(l 1 , l 2,...,l n )。Then the score of the single-byte name N where k is a hyperparameter representing the number of words participating in the scoring, and in this example, k = 3 is taken.

[0056] Furthermore, construct an Aho-Corasick automaton according to the organization names in the candidate organization list. The Aho-Corasick automaton is a character matching algorithm for multiple pattern strings, and its main advantage lies in that its matching time complexity is O(T), where T is the text length. In an embodiment of the present invention, the names of all organization entities are used as the pattern strings of the Aho-Corasick automaton. For example, assume that the pattern string set is P = {p 1 , p 2 ,..., p k}. In this embodiment, the Aho-Corasick automaton can be constructed according to the following steps:

[0057] Construct the pattern string P = {p 1 , p 2 ,..., p k} into a prefix tree. It should be noted that there are the following key points in constructing the prefix tree in this example: If p i is a single-byte name index, spaces need to be added at the head and tail. If it is a double-byte name, it is directly indexed. This can unify the processing methods of single-byte names and double-byte names while retaining the word boundaries of the two types of names. The prefix tree construction steps are as follows:

[0058] Starting from the root node, insert the pattern string p i into the prefix tree in turn;

[0059] Transfer along the current character in p i on the prefix tree. If the node does not exist, create a node;

[0060] For the end point, mark the entity corresponding to p i and the score Score(p i ). At this time, the transfer chain from the root node to the end point can form the pattern string p i . i

[0061] Construct the failure pointer. The failure pointer contains information on where to continue matching if the match fails. The failure pointer needs to be set for each node on the prefix tree, and this pointer points to the node on the prefix tree that satisfies the condition of "the longest suffix is the same as the prefix and is not the string itself". The construction steps of the failure pointer are as follows:

[0062] Perform a breadth-first traversal of the entire prefix tree;

[0063] For each child node i of the current node n, traverse k along the mismatch pointer of n. If k also has a child node i, the mismatch pointer of child node i of n points to child node i of k.

[0064] Otherwise, the mismatch pointer of child node i of n points to the root.

[0065] So far, the construction of the AC automaton is completed. Figure 2 This is the schematic diagram of the AC automaton construction process of the embodiment of the present invention.

[0066] Step S3, input the text to be extracted into the AC automaton, and perform text matching through the constructed AC automaton to obtain the organization names included in the text to be extracted.

[0067] It can be understood that the present invention uses the constructed AC automaton to extract organization names from the text to be extracted, as Figure 3 shown.

[0068] Specifically, the key points of the embodiment of the present invention in matching are as follows: Add spaces at the head and tail of the input text, and then use the above-mentioned AC automaton for matching to obtain all the organization names included in the text. Combining the above AC automaton construction logic, while unifying the processing methods of single-byte names and double-byte names, the word boundaries of the two types of names can be retained. The matching can be carried out according to the following steps:

[0069] (1) Starting from the root node, transfer along the current character in the input text in the prefix tree. The transfer is divided into three cases:

[0070] 1. There is a transfer path, directly transfer, and the transfer is successful;

[0071] 2. There is no transfer path, search along the mismatch pointer until there is a transfer path to transfer, and the transfer is successful;

[0072] 3. Until the root, there is no transfer path, the transfer fails, and directly proceed to the next character and start a new transfer from the root.

[0073] (2) If the transfer is successful, record all the endpoint information on the mismatch pointer chain of the current node into the result set.

[0074] Step S4, according to the importance degree of the organization names, select the organization with the highest score from the organization names included in the text to be extracted by the AC automaton.

[0075] Specifically, the result set generated in step S3 covers all the pattern strings and the scores of the pattern strings themselves included in the input text. The best organization can be selected according to the following situations:

[0076] If the result set is empty, the input string does not contain any organizations within the range, and return an empty set.

[0077] Sort the result set. If the results with the highest scores belong to the same institutional entity, return that entity;

[0078] Otherwise, determine the corresponding entity based on the remaining information in the input text string and the remaining information of the institutional entity. For example, for "Northeastern University", if "China" or "USA" can be extracted from the remaining text string, then it can be determined whether the institution in the text string is "Northeastern University (China)" or "Northeastern University (USA)".

[0079] Through the above steps, obtain a list of candidate institutions. The list of candidate institutions includes at least one institution name; according to the word feature information that makes up the institution name, score the institution name and obtain a scoring result to calculate the importance of the institution name, and construct an AC automaton based on the institution names that make up the candidate institution list; among them, the word feature information includes multiple of: the number of occurrences of the word, rarity, and length; input the text to be extracted into the AC automaton, and perform text matching through the constructed AC automaton to obtain the institution names included in the text to be extracted; according to the importance of the institution name, screen out the institution with the highest score from the institution names included in the text to be extracted. When extracting institution names, the present invention has high robustness, and solves the problems of high cost, poor real-time performance and accuracy, and difficulty in cross-language institution extraction of existing institution name extraction methods.

[0080] It should be noted that there are various implementation methods for the method of extracting institution names from text. However, regardless of the specific implementation method, as long as the method improves the real-time performance and accuracy of extraction, reduces the cost of extraction and the difficulty of cross-language extraction, it is a solution to the existing technical problems and has corresponding effects.

[0081] To implement the above embodiments, as Figure 4 shown, this embodiment also provides a device 10 for quickly extracting institution names from text. The device 10 includes: an acquisition module 100, a scoring module 200, a matching module 300, and a screening module 400.

[0082] The acquisition module 100 is used to acquire a list of candidate institutions. The list of candidate institutions includes at least one institution name;

[0083] The scoring module 200 is used to score the institution name according to the word feature information that makes up the institution name and obtain a scoring result to calculate the importance of the institution name, and construct an AC automaton based on the institution names that make up the candidate institution list; among them, the word feature information includes multiple of: the number of occurrences of the word, rarity, and length;

[0084] A matching module 300 for inputting the text to be extracted into an AC automaton, and performing text matching through the constructed AC automaton to obtain the organization names included in the text to be extracted;

[0085] A screening module 400 for screening out the organization with the highest score from the organization names included in the text to be extracted according to the importance of the organization names.

[0086] Furthermore, the acquisition module 100 includes but is not limited to the following sub-modules:

[0087] A first acquisition sub-module for acquiring the template type and the organization page related to the organization from the general knowledge base as a seed set;

[0088] A second acquisition sub-module for, based on the seed set, acquiring the corresponding page of each organization in the general knowledge base according to the language link; and,

[0089] A third acquisition sub-module for using the redirected entry of each organization in the general knowledge base as the name of each organization.

[0090] The apparatus for quickly extracting organization names according to an embodiment of the present invention acquires a candidate organization list, where the candidate organization list includes at least one organization name; scores the organization names according to the word feature information of the words constituting the organization names and obtains a scoring result to calculate the importance of the organization names, and constructs an AC automaton according to the organization names included in the candidate organization list; where the word feature information includes multiple ones of the number of word occurrences, rarity, and length; inputs the text to be extracted into the AC automaton, and performs text matching through the constructed AC automaton to obtain the organization names included in the text to be extracted; and screens out the organization with the highest score from the organization names included in the text to be extracted according to the importance of the organization names. When the present invention extracts organization names, it has high robustness, and solves the problems of high cost, poor real-time performance and accuracy, and difficulty in cross-language organization extraction of the existing organization name extraction methods.

[0091] It should be noted that the foregoing explanation of the embodiment of the method for quickly extracting organization names in the text also applies to the apparatus for quickly extracting organization names in the text of this embodiment, and will not be repeated here.

[0092] In addition, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of these features. In the description of the present invention, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0093] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0094] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for quickly extracting organization names from text, characterized in that, it includes the following steps: Obtain a candidate organization list, where the candidate organization list includes at least one organization name; According to the word feature information that makes up the organization name, score the organization name to obtain a scoring result, so as to calculate the importance degree of the organization name, and construct an AC automaton according to the organization names that make up the candidate organization list; wherein, the word feature information includes: multiple of the number of word occurrences, rarity, and length; Input the text to be extracted into the AC automaton, and perform text matching through the constructed AC automaton to obtain the organization names included in the text to be extracted; According to the importance degree of the organization name, screen out the organization with the highest score among the organization names included in the text to be extracted from the AC automaton; Among them, the scoring the organization name according to the word feature information that makes up the organization name to obtain a scoring result includes: Pre-define a scoring rule for scoring the organization name. For a given name, if it contains double-byte characters, it is a double-byte name, otherwise it is a single-byte name; For the double-byte name, the score of the double-byte name is the number of characters that make up the double-byte name; For the single-byte name, obtain the total number of all single-byte names, and count the number of occurrences of the words that make up all single-byte names; according to the total number and the number of occurrences, calculate the importance degree of each word in each single-byte name; sort and list the importance degrees, and obtain the score of the single-byte name according to the list; The constructing the AC automaton according to the organization names that make up the candidate organization list includes: Take all the organization names in the candidate organization list as the first pattern strings of the AC automaton, and construct the first pattern strings into a prefix tree; Based on the prefix tree, construct a mismatch pointer; Take all the organization names in the candidate organization list as the first pattern strings of the AC automaton, and construct the first pattern strings into a prefix tree, including: Starting from the root node, insert the second pattern string into the prefix tree in turn; wherein, if the second pattern string is a single-byte name, add spaces at the head and tail during indexing, and if it is a double-byte name, directly index; Transfer along the current character in the second pattern string on the prefix tree, and create a node if the node does not exist; For the end point, mark the entity corresponding to the second pattern string and the score of the second pattern string.

2. The method for quickly extracting organization names from text according to claim 1, characterized in that, the obtaining the candidate organization list includes: Obtain the template type and the organization page related to the organization from the general knowledge base as the seed set; Based on the seed set, obtain the corresponding page of each organization in the general knowledge base according to the language link; and, take the redirect entry of each organization in the general knowledge base as the name of each organization.

3. The method for quickly extracting organization names from text according to claim 1, characterized in that, Constructing a mismatch pointer based on the prefix tree includes: Performing a breadth-first traversal of the prefix tree; For each child node i of the current node n, traverse k along the mismatch pointer of n. If k also has a child node i, the mismatch pointer of the child node i of n points to the child node i of k; Otherwise, the mismatch pointer of the child node i of n points to the root node.

4. An apparatus for quickly extracting organization names from text, characterized in that, it includes: An acquisition module for acquiring a candidate organization list, where the candidate organization list includes at least one organization name; A scoring module for scoring the organization name according to the word feature information constituting the organization name and obtaining a scoring result to calculate the importance of the organization name, and constructing an AC automaton according to the organization names constituting the candidate organization list; wherein, the word feature information includes multiple of: the number of occurrences of the word, rarity, and length; A matching module for inputting the text to be extracted into the AC automaton and performing text matching through the constructed AC automaton to obtain the organization names included in the text to be extracted; A screening module for screening out the organization with the highest score from the organization names included in the text to be extracted according to the importance of the organization name; The scoring the organization name according to the word feature information constituting the organization name and obtaining a scoring result includes: Predefining a scoring rule for scoring the organization name. For a given name, if it contains double-byte characters, it is a double-byte name, otherwise it is a single-byte name; For the double-byte name, the score of the double-byte name is the number of characters constituting the double-byte name; For the single-byte name, obtain the total number of all single-byte names, and count the number of occurrences of the words constituting all single-byte names; according to the total number and the number of occurrences, calculate the importance of each word in each single-byte name; sort and list the importance, and obtain the score of the single-byte name according to the list; The constructing an AC automaton according to the organization names constituting the candidate organization list includes: Taking all the organization names in the candidate organization list as the first pattern strings of the AC automaton, and constructing the first pattern strings into a prefix tree; Based on the prefix tree, constructing a mismatch pointer; Taking all the organization names in the candidate organization list as the first pattern strings of the AC automaton, and constructing the first pattern strings into a prefix tree includes: Starting from the root node, sequentially insert the second pattern string into the prefix tree; wherein, if the second pattern string is a single-byte name, add spaces at the head and tail during indexing, and if it is a double-byte name, directly index; Transfer along the current character in the second pattern string on the prefix tree, and create a node if the node does not exist; For the end point, mark the entity corresponding to the second pattern string and the score of the second pattern string.

5. The apparatus for quickly extracting organization names from text according to claim 4, characterized in that, the acquisition module includes: The first acquisition sub-module is used to acquire the template type and the institutional page related to the institution from the general knowledge base as a seed set; The second acquisition sub-module is used to, based on the seed set, acquire the corresponding page of each institution in the general knowledge base according to the language link; and, The third acquisition sub-module is used to use the redirect entry of each institution in the general knowledge base as the name of each institution.

Citation Information

Patent Citations

  • Extraction method and device for company name in text

    CN108710671A

  • Text processing method and device and related equipment

    CN110209781A