An entity recognition method and device, an electronic device, and a storage medium
Patent Information
- Application Number
- CN202211509137.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-29
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-11-29
AI Technical Summary
[0004]本申请实施例提供一种实体识别方法、装置、电子设备及存储介质,用以解决相关技术中存在的实体识别准确度低的问题
[0049]本申请实施例中,对获取的待识别条款的文本内容进行分词,得到分词序列,从各预设实体的分词与各预设实体的索引之间的倒排索引表中,查询分词序列中每个分词的索引集合,基于分词序列中各分词的索引集合,从待识别条款的文本内容中确定候选实体,将各预设实体中与候选实体匹配的实体,作为待识别条款的实体识别结果,其中,各预设实体是基于待识别条款的文本内容包含的指定类型的实体确定的。这样,先对历史条款包含的指定类型的实体进行整理,得到多个预设实体,并建立这些预设实体中分词的倒排索引表,后续,在识别任一待识别条款中指定类型的实体时,借助于倒排索引表确定候选实体,并将与候选实体匹配的预设实体,确定为指定类型实体的识别结果,即便指定类型实体的长度较长,也可保证识别准确度。
Smart Images

Figure CN115906851B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of financial data processing technology, and in particular to an entity recognition method, device, electronic device and storage medium. Background Technology
[0002] In international trade, the terms of a letter of credit express the applicant's / issuing bank's requirements for the documents or other trade activities provided by the beneficiary, such as what documents and their contents to provide, the latest time requirement for submitting documents, and requirements for the counterparty bank's trade activities, such as the scenarios under which payment will be refused and which bank the documents should be sent to.
[0003] When using Natural Language Processing (NLP) technology to assist document examiners in reviewing letters of credit, it is necessary to fully understand the requirements of the letter of credit terms, analyze the specific required documents, the content to be displayed on the documents, constraints such as trade time and port, and information on the advising bank and acquiring bank, to support subsequent document review and automatic data entry into the system. However, entities such as the beneficiary name and the name and address of the remitting bank (i.e., company name and address) often contain long descriptive terms, branch or department information, making it difficult to guarantee accurate identification. Summary of the Invention
[0004] This application provides an entity recognition method, apparatus, electronic device, and storage medium to solve the problem of low entity recognition accuracy in related technologies.
[0005] In a first aspect, embodiments of this application provide an entity recognition method, including:
[0006] Obtain the text content of the clause to be identified;
[0007] The text content of the clause to be identified is segmented into words to obtain a segmentation sequence;
[0008] From the inverted index table between the word segmentation of each preset entity and the index of each preset entity, query the index set of each word segmentation sequence, wherein each preset entity is determined based on the specified type of entity contained in the text content of the historical terms;
[0009] Candidate entities are determined based on the index set of each word in the word segmentation sequence;
[0010] The entity that matches the candidate entity among the preset entities is taken as the entity identification result of the clause to be identified.
[0011] In some embodiments, the text content of the clause to be identified is segmented into words to obtain a segmentation sequence, including:
[0012] The text content of the clause to be identified is segmented into n-grams to obtain the segmented sequence.
[0013] In some embodiments, candidate entities are determined from the text content of the clause to be identified based on the index set of each word in the word segmentation sequence, including:
[0014] Based on whether the intersection of the index set of each word in the word segmentation sequence and the index set of the reference word is empty, candidate entities are selected from the text content of the clause to be identified, where the reference word is the word located after the word segmentation.
[0015] Select candidate entities from among the candidate entities.
[0016] In some embodiments, candidate entities are selected from the text content of the clause to be identified based on whether the intersection of the index set of each word in the word segmentation sequence and the index set of the reference word is empty, including:
[0017] For each word in the word segmentation sequence, the intersection of the index set of the word segmentation and the index set of the reference word is taken. Initially, the interval between the reference word and the word segmentation is 1.
[0018] If the intersection is empty, then record a case where there is no common index.
[0019] If the intersection is not empty, then update the index set of the word segmentation to the intersection, increase the interval between the reference word and the word segmentation by 1, and execute the step of taking the intersection of the index set of the word segmentation and the index set of the reference word.
[0020] When the number of records without a common index reaches a preset value, the characters from the word segmentation to the reference word in the text content of the clause to be identified are determined as a candidate entity.
[0021] In some embodiments, selecting a candidate entity from among the alternative entities includes:
[0022] If no other candidate entity contains the same content as any other candidate entity, then the candidate entity is determined as a candidate entity.
[0023] If any candidate entity has other candidate entities that share the same content, then the candidate entity with the most characters among the candidate entity and the other candidate entities is determined as a candidate entity.
[0024] In some embodiments, each preset entity contains a character length exceeding a specified value.
[0025] Secondly, embodiments of this application provide an entity recognition device, comprising:
[0026] The acquisition module is used to acquire the text content of the clause to be identified;
[0027] The word segmentation module is used to segment the text content of the clause to be identified into words to obtain a word segmentation sequence;
[0028] The query module is used to query the index set of each word in the word segmentation sequence from the inverted index table between the word segmentation of each preset entity and the index of each preset entity, wherein each preset entity is determined based on the specified type of entity contained in the text content of the historical terms;
[0029] The determination module is used to determine candidate entities based on the index set of each word in the word segmentation sequence;
[0030] The identification module is used to identify entities that match the candidate entities among the preset entities as the entity identification results of the clause to be identified.
[0031] In some embodiments, the word segmentation module is specifically used for:
[0032] The text content of the clause to be identified is segmented into n-grams to obtain the segmented sequence.
[0033] In some embodiments, the determining module is specifically used for:
[0034] Based on whether the intersection of the index set of each word in the word segmentation sequence and the index set of the reference word is empty, candidate entities are selected from the text content of the clause to be identified, where the reference word is the word located after the word segmentation.
[0035] Select candidate entities from among the candidate entities.
[0036] In some embodiments, the determining module is specifically used for:
[0037] For each word in the word segmentation sequence, the intersection of the index set of the word segmentation and the index set of the reference word is taken. Initially, the interval between the reference word and the word segmentation is 1.
[0038] If the intersection is empty, then record a case where there is no common index.
[0039] If the intersection is not empty, then update the index set of the word segmentation to the intersection, increase the interval between the reference word and the word segmentation by 1, and execute the step of taking the intersection of the index set of the word segmentation and the index set of the reference word.
[0040] When the number of records without a common index reaches a preset value, the characters from the word segmentation to the reference word in the text content of the clause to be identified are determined as a candidate entity.
[0041] In some embodiments, the determining module is specifically used for:
[0042] If no other candidate entity contains the same content as any other candidate entity, then the candidate entity is determined as a candidate entity.
[0043] If any candidate entity has other candidate entities that share the same content, then the candidate entity with the most characters among the candidate entity and the other candidate entities is determined as a candidate entity.
[0044] In some embodiments, each preset entity contains a character length exceeding a specified value.
[0045] Thirdly, embodiments of this application provide an electronic device, including: at least one processor, and a memory communicatively connected to the at least one processor, wherein:
[0046] The memory stores instructions that can be executed by at least one processor to enable the at least one processor to perform the entity recognition method described above.
[0047] Fourthly, embodiments of this application provide a storage medium in which the electronic device can perform the above-described entity recognition method when the instructions in the storage medium are executed by the processor of an electronic device.
[0048] Fifthly, embodiments of this application provide a computer program product that, when executed by an electronic device, causes the electronic device to perform the aforementioned entity recognition method.
[0049] In this embodiment, the text content of the clause to be identified is segmented into words to obtain a segmentation sequence. The index set of each word in the segmentation sequence is queried from the inverted index table between the segmentation of each preset entity and the index of each preset entity. Based on the index set of each word in the segmentation sequence, candidate entities are determined from the text content of the clause to be identified. Entities that match the candidate entities among the preset entities are taken as the entity identification result of the clause to be identified. Each preset entity is determined based on the specified type of entity contained in the text content of the clause to be identified. In this way, the specified type of entities contained in historical clauses are first organized to obtain multiple preset entities, and an inverted index table of the segmentation of these preset entities is established. Subsequently, when identifying a specified type of entity in any clause to be identified, candidate entities are determined with the help of the inverted index table, and the preset entities that match the candidate entities are identified as the identification result of the specified type of entity. Even if the specified type of entity is long, the identification accuracy can be guaranteed. Attached Figure Description
[0050] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0051] Figure 1 A flowchart illustrating an entity recognition method provided in this application embodiment;
[0052] Figure 2 A flowchart illustrating a method for determining candidate entities provided in this application embodiment;
[0053] Figure 3 A schematic diagram illustrating the process of offline creation of an inverted index table based on a thesaurus n-gram, provided for an embodiment of this application;
[0054] Figure 4 This application provides a schematic diagram illustrating an online process for retrieving and identifying company names and addresses in letter of credit terms.
[0055] Figure 5 A schematic diagram of a merged index provided for an embodiment of this application;
[0056] Figure 6 This is a schematic diagram of the structure of an entity recognition device provided in an embodiment of this application;
[0057] Figure 7 This is a schematic diagram of the hardware structure of an electronic device for implementing an entity recognition method, provided in an embodiment of this application. Detailed Implementation
[0058] To address the issue of low accuracy in entity recognition in related technologies, embodiments of this application provide an entity recognition method, apparatus, electronic device, and storage medium.
[0059] The preferred embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit this application. Unless otherwise specified, the embodiments and features described herein can be combined with each other. Furthermore, the acquisition, storage, use, and processing of data in the embodiments of this application comply with the relevant provisions of national laws and regulations.
[0060] For ease of understanding, the technical terms used in this application are as follows:
[0061] In the field of NLP, an entity can be a person's name, place name, organization name, country, date, etc. For example, a certain bank is an organization.
[0062] n-grams are used to perform a sliding window operation of size N on the text content of the clause to be identified, forming a sequence of byte / word fragments of length N.
[0063] An inverted index originates from the practical application of finding entities containing a specific word segment. Each entry in such an index table contains a word segment and indexes of the entities containing that word. Because the word segment determines the entity index, rather than the entity itself, it is called an inverted index.
[0064] Figure 1 A flowchart of an entity recognition method provided in this application embodiment includes the following steps.
[0065] In step S101, the text content of the clause to be identified is obtained.
[0066] The terms to be identified can be terms of a letter of credit, and the text of the terms to be identified can be in Chinese or English.
[0067] In step S102, the text content of the clause to be identified is segmented into words to obtain the segmented word sequence of the clause to be identified.
[0068] For example, n-gram segmentation can be performed on the text content of the clause to be identified to obtain a segmented sequence. This preserves the semantic information of each segmented word in the clause, which helps improve the accuracy of subsequent entity recognition.
[0069] In step S103, the index set of each word in the word segmentation sequence is queried from the inverted index table between the word segmentation of each preset entity and the index of each preset entity. Each preset entity is determined based on the specified type of entity contained in the text content of the historical terms.
[0070] Among them, entities of specified types, such as beneficiary names, company names, and addresses, generally contain a relatively long character length, such as more than 10 characters, and are therefore called extra-long entities. Historical clauses are clauses that have been processed in actual business. The entities of specified types contained in these clauses are known, so these known entities of specified types can be directly used as preset entities, or these known entities of specified types can be used together with their generalized entities as preset entities. Since their generalized entities are generally also extra-long entities, each preset entity is also an extra-long entity, that is, each preset entity contains a character length exceeding the specified value, such as 10.
[0071] Taking the term to be identified as a letter of credit term as an example, the excessively long entities contained in previously processed letter of credit terms and the possible expressions of these excessively long entities can be sorted out in advance to obtain multiple preset entities. Then, n-gram segmentation is performed on each preset entity to obtain the segmentation sequence of the preset entity. Based on the segmentation sequence of each preset entity, an inverted index table between the segmentation of each preset entity and each preset entity is established.
[0072] Subsequently, for each word in the word segmentation sequence of the clause to be identified, the index set corresponding to that word is retrieved from the inverted index table, that is, the index of all preset entities containing that word is retrieved.
[0073] In step S104, candidate entities are determined from the text content of the clause to be identified based on the index set of each word in the word segmentation sequence.
[0074] In specific implementation, it can be based on Figure 2 The process shown identifies candidate entities and includes the following steps.
[0075] In step 1041, candidate entities are selected from the text content of the clause to be identified based on whether the intersection of the index set of each segment in the segmentation sequence and the index set of the reference word is empty. The reference word is the segment located after the segment.
[0076] For example, for each segment in the word segmentation sequence, the intersection of the index set of the segment and the index set of the reference word is taken. Initially, the interval between the reference word and the segment is 1. If the intersection is empty, it is recorded once that there is no common index. If the intersection is not empty, the index set of the segment is updated to the intersection, the interval between the reference word and the segment is increased by 1, and the step of taking the intersection of the index set of the segment and the index set of the reference word is executed until the number of records with no common index reaches a preset value. Then, the characters from the segment to the reference word in the text content of the clause to be identified are determined as a candidate entity. That is, all characters from the first character of the segment to the last character of the reference word in the text content of the clause to be identified are taken as a candidate entity.
[0077] In step 1042, a candidate entity is selected from the candidate entities.
[0078] For example, if any candidate entity does not have other candidate entities that share the same content, then this candidate entity is determined as a candidate entity; if any candidate entity has other candidate entities that share the same content, then this candidate entity and the other candidate entities with the most characters are determined as a candidate entity.
[0079] In this way, the longest candidate entity can be selected from the candidate entities containing the same characters, reducing the number of candidate entities and thus improving the entity recognition speed.
[0080] It should be noted that the text of the clause to be identified may contain both the beneficiary's name and the company's name and address, so there may be more than one candidate entity.
[0081] In step S105, the entity that matches the candidate entity among the preset entities is taken as the entity identification result of the clause to be identified.
[0082] For example, calculate the similarity between each candidate entity and each preset entity. If there is at least one preset entity with a similarity greater than the preset value, then the one with the highest similarity among these at least one preset entities is taken as the recognition result of this candidate entity.
[0083] The following section uses the identification of company name and address in letter of credit terms as an example to introduce the solution of this application embodiment.
[0084] The solution in this application mainly includes two stages:
[0085] The first stage involves building an inverted index table based on the n-gram dictionary offline.
[0086] See Figure 3 ,exist Figure 3 In the first line, "18444" represents the index of the company name and address. "SOCIAL ABCDBANK LTD., ABCD ROAD BRANCH, 610 / 11, ABCD ROAD" represents the company name and address. The meanings of the other lines are similar and will not be repeated here.
[0087] Because individual words in company names and addresses often fail to convey key meanings, and because the arbitrariness of language description leads to variations in sentence descriptions each time, this paper proposes a method for creating company names and addresses. After special punctuation cleaning and word segmentation, an n-gram sequence can be built for each word, moving forward n words (n is typically 3, including the current word). The index value of the current n-gram is recorded using its dictionary index number. Finally, an inverted index based on the n-grams of the company name and address dictionary is established, such as... Figure 3 The “SOCIAL ABCD BANK” shown is contained in the bank entity / record with index numbers [18444,18445,…,18455…].
[0088] In addition, for each company name and address, the sentence length information after word segmentation can be statistically analyzed and recorded offline for rapid comparison of candidate sets in subsequent steps.
[0089] The second stage involves online retrieval to identify the company name and address in the letter of credit terms.
[0090] The following is combined with Figure 4This section describes the process of online retrieval and identification of company names and addresses in letter of credit terms.
[0091] The first step is to establish an n-gram sequence of the terms of the letter of credit to be parsed.
[0092] Suppose the clause to be parsed to identify the company name and address is: 2. ORIGINAL SET OF DOCUMENTS INLUDING 6 COPIES OF INVOICE AND DUPLICATE SET OF DOCUMENTS ALONG WITH REST2 COPIES OF INVOICE TO BE SENT TO SOCIAL ABCD BANK LTD. BBBB BRANCH, TEST, SUCCESSIVE REGISTERED, AIR MAIL IMMEDIATELY AFTER NEGOTIATION.
[0093] After performing special punctuation cleaning and word segmentation on this clause, an n-gram sequence of the clause is established using a sliding window with the same number of 'n', and its position information is recorded by the word number after specific word segmentation. For example, the n-gram sequence starting from the 25th position is "SOCIAL ABCD BANK".
[0094] The second step is to obtain the index results of the n-gram sequences in the inverted index table.
[0095] For each sequence in the n-gram sequence, look up the set of indices for that sequence from the inverted index table of company names and addresses.
[0096] It is worth noting that because the company name and address are from a non-standard, uncleaned dictionary, they may contain many uncleaned and irrelevant words, such as "SET OF DOCUMENTS," which may also be indexed. This issue can be addressed in the next step.
[0097] The third step is to merge the index results and obtain the locally optimal starting position of the name address.
[0098] Starting from the first position, merge consecutively mergeable index results. See [link / reference]. Figure 5 The merging process is as follows:
[0099] (1) When the word wi at position i has an index, merge the indexes corresponding to the n-grams one by one from that position onwards. If they have a common index, retain the common index and continue merging until there is no common index, and proceed to step (2).
[0100] (2) Record the number of times there is no common index. When the cumulative step length without common index exceeds the given threshold, such as 3, the longest step length with common index starting from position i is the step length with common index. For example, if there are 8 consecutive positions with common index after position 25 (there may be up to 3 n-gram sequences that are ignored without common index), then "8" is the sentence length of the candidate entity in terms of word count; while there are 7 positions with common index starting from position 26.
[0101] Finally obtained Figure 4 The index merging conclusion shown is in Figure 4 In the first row [0, 0, 3, 2, 0, 0, ... 8, 7, 6, 5, 4, ... 0, 0], each number i represents the i-th word "has a common index at the longest m positions following this position". Each number in the second row [2, 10, 12, 20, 25, 32] represents the local extremum position of the longest common index distance in the first row. For example, if the position of the 25th n-gram in the clause is a local extremum, then the scenario where there are 7 consecutive positions with a common index after the 26th position in the first row does not need to be judged again, because the 7 consecutive positions after the 26th position and the word at the 25th position can form a candidate entity, and there is no need to repeat the judgment.
[0102] The fourth step is to construct the candidate set and the word index ID and obtain the recognition results.
[0103] Based on the starting position and sentence length of the candidate company name addresses in the second row obtained from merging index IDs in step three, the differences between each candidate company name address and the vocabulary under the index ID are determined sequentially. In this step, the sentence length, calculated from the number of words, is used as the evaluation metric. Figure 4 In the example, (2, 3,
[19] ) indicates that the sequence “SET OF DOCUMENTS INCLUDING”, consisting of 3 words following the second position in the clause, has a sentence length of 4. It is included in the dictionary record with id 19, and the sentence length of that record is 20. It can be seen that the sentence lengths of the two are too different, so the sequence is not considered as the identified company name address. On the other hand, (25, 8,
[18452] ) indicates that the sequence “SOCIAL ABCD BANK LTD.BBBB BRANCH,TEST,SUCCESSIVE REGISTERED”, consisting of 8 words following the 25th position, has a sentence length that is similar to that of the record “SOCIAL ABCD BANK LTD.BBBBBRANCH,SUCCESSIVE REGISTERED” with record id 18452 (the word expression has been confirmed in the n-gram search and merging), so it is identified as a company name address.
[0104] Step 5: Output the recognition results.
[0105] 1. The name and address of the sending bank identified in the original terms [25:33]
[0106] SOCIAL ABCD BANK LTD.BBBB BRANCH, TEST, SUCCESSIVE REGISTERED.
[0107] 2. Most similar record
[18452]
[0108] SOCIAL ABCD BANK LTD.BBBB BRANCH, SUCCESSIVE REGISTERED.
[0109] The solution provided in this application has the following advantages:
[0110] (1) The inverted index table is directly built based on the long entities in the terms accumulated by users over a long period of time. There is no need to acquire a large amount of corpus. Based on the rules and strategies of inverted retrieval and merging of retrieval results, the computer resource requirements are small and the development cost is low.
[0111] (2) Building an inverted index table based on n-grams focuses on the combination of names and addresses, which can avoid the interference of too many general term indexes in the single word index scenario, such as "BANK / LTD" appearing in almost all company entities. The "TEST ABC BANK" presented by n-grams often contains specific words and their order meanings in the company name and address, expressing certain semantic features, which is more conducive to accurate matching.
[0112] (3) After adopting the n-gram splitting clause, the maximum step size that can be ignored can be customized when merging the index set, allowing the retrieval results to have a certain degree of ambiguity and high recall.
[0113] (4) It can also offline statistically analyze and record the sentence length information after each word segmentation, and perform preprocessing based on the sentence length information during the matching stage, so that irrelevant options can be eliminated first by length when matching results, thus accelerating the matching of results.
[0114] (5) Based on solving local extrema with common indices, the starting position of candidate entities can be located more accurately without having to judge each position, thus improving the efficiency of entity recognition.
[0115] When the method provided in the embodiments of this application is implemented in software, hardware, or a combination of software and hardware, the electronic device may include multiple functional modules, and each functional module may include software, hardware, or a combination thereof.
[0116] Based on the same technical concept, this application also provides an entity recognition device. The principle of the entity recognition device in solving the problem is similar to that of the entity recognition method described above. Therefore, the implementation of the entity recognition device can refer to the implementation of the entity recognition method, and the repeated parts will not be described again. Figure 6 The present application provides a schematic diagram of the structure of an entity recognition device, which includes an acquisition module 601, a word segmentation module 602, a query module 603, a determination module 604, and a recognition module 605.
[0117] The acquisition module 601 is used to acquire the text content of the clause to be identified;
[0118] The word segmentation module 602 is used to segment the text content of the clause to be identified into words to obtain a word segmentation sequence;
[0119] The query module 603 is used to query the index set of each word in the word segmentation sequence from the inverted index table between the word segmentation of each preset entity and the index of each preset entity, wherein each preset entity is determined based on the specified type of entity contained in the text content of the historical terms;
[0120] The determining module 604 is used to determine candidate entities from the text content of the clause to be identified based on the index set of each word in the word segmentation sequence;
[0121] The identification module 605 is used to identify the entity that matches the candidate entity among the preset entities as the entity identification result of the clause to be identified.
[0122] In some embodiments, the word segmentation module 602 is specifically used for:
[0123] The text content of the clause to be identified is segmented into n-grams to obtain the segmented sequence.
[0124] In some embodiments, the determining module 604 is specifically used for:
[0125] Based on whether the intersection of the index set of each word in the word segmentation sequence and the index set of the reference word is empty, candidate entities are selected from the text content of the clause to be identified, where the reference word is the word located after the word segmentation.
[0126] Select candidate entities from among the candidate entities.
[0127] In some embodiments, the determining module 604 is specifically used for:
[0128] For each word in the word segmentation sequence, the intersection of the index set of the word segmentation and the index set of the reference word is taken. Initially, the interval between the reference word and the word segmentation is 1.
[0129] If the intersection is empty, then record a case where there is no common index.
[0130] If the intersection is not empty, then update the index set of the word segmentation to the intersection, increase the interval between the reference word and the word segmentation by 1, and execute the step of taking the intersection of the index set of the word segmentation and the index set of the reference word.
[0131] When the number of records without a common index reaches a preset value, the characters from the word segmentation to the reference word in the text content of the clause to be identified are determined as a candidate entity.
[0132] In some embodiments, the determining module 604 is specifically used for:
[0133] If no other candidate entity contains the same content as any other candidate entity, then the candidate entity is determined as a candidate entity.
[0134] If any candidate entity has other candidate entities that share the same content, then the candidate entity with the most characters among the candidate entity and the other candidate entities is determined as a candidate entity.
[0135] In some embodiments, each preset entity contains a character length exceeding a specified value.
[0136] The module division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, other division methods are possible. Furthermore, the functional modules in each embodiment of this application can be integrated into a single processor, exist as separate physical entities, or be integrated into a single module. Coupling between modules can be achieved through interfaces, typically electrical communication interfaces, but mechanical interfaces or other types of interfaces are also possible. Therefore, modules described as separate components may or may not be physically separate; they can be located in one place or distributed across different locations on the same or different devices. The integrated modules described above can be implemented in hardware or as software functional modules.
[0137] Having introduced the entity recognition method and apparatus according to exemplary embodiments of this application, we will now introduce an electronic device according to another exemplary embodiment of this application.
[0138] In some possible implementations, the electronic device of this application may include at least one processor and at least one memory. The memory stores program code that, when executed by the processor, causes the processor to perform the methods described above according to the various exemplary embodiments of this application.
[0139] The following reference Figure 7To describe an electronic device 130 implemented according to this embodiment of the present application. Figure 7 The electronic device 130 shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0140] like Figure 7 As shown, the electronic device 130 is presented in the form of a general-purpose electronic device. The components of the electronic device 130 may include, but are not limited to: at least one processor 131, at least one memory 132, and a bus 133 connecting different system components (including memory 132 and processor 131).
[0141] Bus 133 represents one or more of several bus structures, including a memory bus or memory controller, peripheral bus, processor, or local bus using any of the various bus structures.
[0142] The memory 132 may include a readable medium in the form of volatile memory, such as random access memory (RAM) 1321 and / or cache memory 1322, and may further include read-only memory (ROM) 1323.
[0143] The memory 132 may also include a program / utility 1325 having a set (at least one) of program modules 1324, including but not limited to: an operating system, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.
[0144] Electronic device 130 can also communicate with one or more external devices 134 (e.g., keyboard, pointing device, etc.), and with one or more devices that enable a user to interact with electronic device 130, and / or with any device that enables electronic device 130 to communicate with one or more other electronic devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 135. Furthermore, electronic device 130 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 136. As shown, network adapter 136 communicates with other modules used in electronic device 130 via bus 133. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 130, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0145] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory 132 including instructions, which can be executed by a processor 131 to complete the entity recognition method described above. Optionally, the storage medium may be a non-transitory computer-readable storage medium, such as a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device.
[0146] In an exemplary embodiment, a computer program product is also provided, which, when invoked and executed by an electronic device, causes the electronic device to perform any of the exemplary methods provided in this application.
[0147] Furthermore, computer program products may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, RAM, ROM, erasable programmable read-only memory (EPROM), flash memory, optical fiber, compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0148] The program product for entity identification in this application embodiment may be a CD-ROM and include program code, and may run on a computing device. However, the program product of this application is not limited thereto. In this document, the readable storage medium may be any tangible medium that contains or stores a program, which may be used by or in conjunction with an instruction execution system, apparatus, or device.
[0149] A readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying readable program code. This propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0150] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, radio frequency (RF), or any suitable combination thereof.
[0151] Program code for performing the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, such as a Local Area Network (LAN) or a Wide Area Network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0152] It should be noted that although several units or sub-units of the device have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of this application, the features and functions of two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.
[0153] Furthermore, although the operations of the method of this application are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.
[0154] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0155] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0156] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0157] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0158] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0159] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, then this application also includes such modifications and variations.
Claims
1. An entity recognition method, characterized in that, include: Obtain the text content of the clause to be identified; The text content of the clause to be identified is segmented into words to obtain a segmentation sequence; From the inverted index table between the word segmentation of each preset entity and the index of each preset entity, query the index set of each word segmentation sequence, wherein each preset entity is determined based on the specified type of entity contained in the text content of the historical terms; Based on the index set of each word in the word segmentation sequence, candidate entities are determined from the text content of the clause to be identified; The entity that matches the candidate entity among the preset entities is taken as the entity identification result of the clause to be identified; Based on the index set of each word in the word segmentation sequence, candidate entities are determined from the text content of the clause to be identified, including: Based on whether the intersection of the index set of each word in the word segmentation sequence and the index set of the reference word is empty, candidate entities are selected from the text content of the clause to be identified, where the reference word is the word located after the word segmentation. Select candidate entities from among the candidate entities; Based on whether the intersection of the index set of each word in the word segmentation sequence and the index set of the reference word is empty, candidate entities are selected from the text content of the clause to be identified, including: For each word in the word segmentation sequence, the intersection of the index set of the word segmentation and the index set of the reference word is taken. Initially, the interval between the reference word and the word segmentation is 1. If the intersection is empty, then record a case where there is no common index. If the intersection is not empty, then update the index set of the word segmentation to the intersection, increase the interval between the reference word and the word segmentation by 1, and execute the step of taking the intersection of the index set of the word segmentation and the index set of the reference word. When the number of records without a common index reaches a preset value, the characters from the word segment to the reference word in the text content of the clause to be identified are determined as a candidate entity. The candidate entity includes all characters from the first character of the word segment to the last character of the reference word in the text content of the clause to be identified.
2. The method as described in claim 1, characterized in that, The text content of the clause to be identified is segmented into words to obtain a segmentation sequence, including: The text content of the clause to be identified is segmented into n-grams to obtain the segmented sequence.
3. The method as described in claim 1, characterized in that, Candidate entities are selected from the candidate entities, including: If no other candidate entity contains the same content as any other candidate entity, then the candidate entity is determined as a candidate entity. If any candidate entity has other candidate entities that share the same content, then the candidate entity with the most characters among the candidate entity and the other candidate entities is determined as a candidate entity.
4. The method as described in claim 1, characterized in that, Each preset entity contains more characters than the specified value.
5. An entity recognition device, characterized in that, include: The acquisition module is used to acquire the text content of the clause to be identified; The word segmentation module is used to segment the text content of the clause to be identified into words to obtain a word segmentation sequence; The query module is used to query the index set of each word in the word segmentation sequence from the inverted index table between the word segmentation of each preset entity and the index of each preset entity, wherein each preset entity is determined based on the specified type of entity contained in the text content of the historical terms; The determination module is used to determine candidate entities from the text content of the clause to be identified based on the index set of each word in the word segmentation sequence; The identification module is used to identify entities that match the candidate entities among the preset entities as the entity identification results of the clause to be identified. The determining module is specifically used to select candidate entities from the text content of the clause to be identified based on whether the intersection of the index set of each word in the word segmentation sequence and the index set of the reference word is empty. The reference word is the word located after the word segmentation. Select candidate entities from among the candidate entities; The determining module is specifically used to take the intersection of the index set of the segmented word and the index set of the reference word for each segmented word in the segmented sequence, with the initial interval between the reference word and the segmented word being 1. If the intersection is empty, then record a case where there is no common index. If the intersection is not empty, then update the index set of the word segmentation to the intersection, increase the interval between the reference word and the word segmentation by 1, and execute the step of taking the intersection of the index set of the word segmentation and the index set of the reference word. When the number of records without a common index reaches a preset value, the characters from the word segment to the reference word in the text content of the clause to be identified are determined as a candidate entity. The candidate entity includes all characters from the first character of the word segment to the last character of the reference word in the text content of the clause to be identified.
6. The apparatus as claimed in claim 5, characterized in that, The word segmentation module is specifically used for: The text content of the clause to be identified is segmented into n-grams to obtain the segmented sequence.
7. The apparatus as claimed in claim 5, characterized in that, The determining module is specifically used for: If no other candidate entity contains the same content as any other candidate entity, then the candidate entity is determined as a candidate entity. If any candidate entity has other candidate entities that share the same content, then the candidate entity with the most characters among the candidate entity and the other candidate entities is determined as a candidate entity.
8. The apparatus as claimed in claim 5, characterized in that, Each preset entity contains more characters than the specified value.
9. An electronic device, characterized in that, include: At least one processor, and a memory communicatively connected to said at least one processor, wherein: The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method as described in any one of claims 1-4.
10. A storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is able to perform the method as described in any one of claims 1-4.
11. A computer program product, characterized in that, When a computer program product is invoked and executed by an electronic device, the electronic device performs the method as described in any one of claims 1-4.
Citation Information
Patent Citations
Knowledge graph-based question and answer implementation method and system
CN112328773A