Element extraction method, device, electronic device and storage medium

By determining the correlation between every two characters in the factor extraction method, and coding the vocabulary matching results, the problems of increasing input length and large storage space in the prior art are solved, and efficient factor extraction is achieved.

CN114238550BActive Publication Date: 2025-07-22IFLYTEK (SUZHOU) TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111538301.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-15
Publication Date
2025-07-22
Estimated Expiration
2041-12-15

AI Technical Summary

Technical Problem

The existing factor extraction method requires splicing several words on the original sentence, resulting in an increase in input length, affecting coding efficiency, and occupying a large amount of storage space.

Method used

By determining the correlation between every two characters in the text to be extracted, the vocabulary matching results are fused to encode, without splicing the matched vocabulary with the original sentence, and using attention mechanisms and knowledge graph decoding strategies to improve coding efficiency and accuracy.

Benefits of technology

Improves encoding efficiency, saves storage space, and improves the accuracy of feature extraction without changing the input length.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114238550B_ABST
    Figure CN114238550B_ABST
Patent Text Reader

Abstract

The present invention provides a method, apparatus, electronic device, and storage medium for element extraction. The method includes: obtaining the text to be extracted and the vocabulary set of the text to be extracted; determining the relevance between every two characters based on the matching result between the string corresponding to every two characters in the text to be extracted and the vocabulary set, where the string is intercepted from the text to be extracted with the corresponding two characters as the start and end points; encoding each character in the text to be extracted based on the relevance between every two characters to obtain the element boundary features of each character; and determining the element extraction result of the text to be extracted based on the element boundary features of each character. The method, apparatus, electronic device, and storage medium for element extraction provided by the present invention do not need to splice the matched vocabulary with the original sentence, and do not change the original input length, thereby improving the encoding efficiency. In addition, compared with the existing method of vocabulary splicing, the storage space is saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular, to a method, apparatus, electronic device, and storage medium for element extraction. Background Art

[0002] The work of element extraction is mainly to extract structured information from unstructured text, which is a very important subfield in natural language processing. The introduction of lexical information in element extraction has attracted more and more attention from researchers. Especially in some professional fields with insufficient corpora, the role of domain vocabulary is even more significant.

[0003] Currently, the element extraction methods that fuse lexical information usually need to splice several words on the original sentence, resulting in an increase in the input length and affecting the encoding efficiency. In addition, the size of the word embedding layer (Embedding) of the existing element extraction model is proportional to the size of the domain vocabulary. Some domain vocabularies can reach the million level. Storing such a large word embedding layer requires additional large storage space. Summary of the Invention

[0004] The present invention provides a method, apparatus, electronic device, and storage medium for element extraction, so as to solve the defects in the prior art that the element extraction method needs to splice several words on the original sentence, resulting in an increase in the input length, affecting the encoding efficiency, and occupying a large storage space.

[0005] The present invention provides a method for element extraction, including:

[0006] Obtain the text to be extracted and the vocabulary set of the text to be extracted;

[0007] Based on the matching result between the string corresponding to each two characters in the text to be extracted and the vocabulary set, determine the relevance between each two characters, where the string is intercepted from the text to be extracted with the corresponding two characters as the starting and ending points;

[0008] Based on the relevance between each two characters, encode each character in the text to be extracted to obtain the element boundary feature of each character;

[0009] Based on the element boundary features of each character, determine the element extraction result of the text to be extracted.

[0010] According to the method for element extraction provided by the present invention, the step of determining the relevance between each two characters based on the matching result between the string corresponding to each two characters in the text to be extracted and the vocabulary set includes:

[0011] Based on the matching results between the strings corresponding to every two characters and the vocabulary set, select the relevance parameters corresponding to every two characters from the candidate relevance parameters corresponding to multiple candidate matching results;

[0012] Based on the relevance parameters corresponding to every two characters and the semantic representations of every two characters, determine the relevance between every two characters.

[0013] According to an element extraction method provided by the present invention, the step of selecting the relevance parameters corresponding to every two characters from the candidate relevance parameters corresponding to multiple candidate matching results based on the matching results between the strings corresponding to every two characters and the vocabulary set includes:

[0014] Based on the matching results between the strings corresponding to every two characters and the vocabulary set, and the matching results between the strings corresponding to every two characters and a preset rule, select the relevance parameters corresponding to every two characters from the candidate relevance parameters corresponding to multiple candidate matching results.

[0015] According to an element extraction method provided by the present invention, the step of determining the relevance between every two characters based on the relevance parameters corresponding to every two characters and the semantic representations of every two characters includes:

[0016] Based on the semantic representations of every two characters, determine the semantic relevance between every two characters;

[0017] Based on the relevance parameters corresponding to every two characters and the semantic representations of every two characters, determine the lexical relevance between every two characters;

[0018] Based on the semantic relevance and lexical relevance between every two characters, determine the relevance between every two characters.

[0019] According to an element extraction method provided by the present invention, the step of determining the element extraction result of the text to be extracted based on the element boundary features of each character includes:

[0020] Based on the element boundary features of each character, determine the segment representations of each candidate segment in the text to be extracted and the lexical representations of each vocabulary in the vocabulary set;

[0021] Based on the association information between each candidate segment and each vocabulary in the knowledge graph, determine the relevance between each candidate segment and each vocabulary;

[0022] Based on the relevance between each candidate segment and each vocabulary, and the segment representations of each candidate segment and the lexical representations of each vocabulary, determine the element extraction result of the text to be extracted.

[0023] According to an element extraction method provided by the present invention, determining the relevance between each candidate segment and each vocabulary based on the association information between each candidate segment and each vocabulary in the knowledge graph includes:

[0024] Based on the association information between each candidate segment and each vocabulary in the knowledge graph and the overlapping inclusion relationship between each candidate segment and each vocabulary, determine the relevance between each candidate segment and each vocabulary.

[0025] According to an element extraction method provided by the present invention, determining the element extraction result of the text to be extracted based on the relevance between each candidate segment and each vocabulary, as well as the segment representation of each candidate segment and the vocabulary representation of each vocabulary, includes:

[0026] Based on the relevance between each candidate segment and each vocabulary and the vocabulary representation of each vocabulary, fuse the vocabulary representations of each vocabulary to obtain the segment representation of each candidate segment after fusion;

[0027] Based on the segment representation of each candidate segment after fusion, determine the element extraction result of the text to be extracted.

[0028] The present invention also provides an element extraction device, including:

[0029] A text acquisition unit for acquiring the text to be extracted and the vocabulary set of the text to be extracted;

[0030] A relevance determination unit for determining the relevance between every two characters based on the matching result between the string corresponding to every two characters in the text to be extracted and the vocabulary set, where the string is intercepted from the text to be extracted with the corresponding two characters as the starting and ending points;

[0031] An encoding unit for encoding each character in the text to be extracted based on the relevance between every two characters to obtain the element boundary feature of each character;

[0032] An extraction result determination unit for determining the element extraction result of the text to be extracted based on the element boundary feature of each character.

[0033] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the steps of any one of the above-mentioned element extraction methods.

[0034] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of any one of the above-mentioned element extraction methods.

[0035] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of any one of the above-mentioned element extraction methods.

[0036] For the element extraction method, device, electronic device and storage medium provided by the present invention, during the process of determining the relevance between every two characters in the text to be extracted, the vocabulary matching result is fused, and each character in the text to be extracted is encoded according to the relevance between every two characters. This method does not require splicing the matched vocabulary with the original sentence, will not change the original input length, thereby improving the encoding efficiency. In addition, compared with the existing vocabulary splicing method, it saves storage space. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0038] Figure 1 is a schematic structural diagram of the Lex-BERT model in the prior art;

[0039] Figure 2 is one of the flow diagrams of the element extraction method provided by the present invention;

[0040] Figure 3 is a flow diagram of step 220 in the element extraction method provided by the present invention;

[0041] Figure 4 is a flow diagram of the method for determining the relevance between every two characters provided by the present invention;

[0042] Figure 5 is a flow diagram of step 240 in the element extraction method provided by the present invention;

[0043] Figure 6 is a flow diagram of the method for determining the element extraction result provided by the present invention;

[0044] Figure 7 is another flow diagram of the element extraction method provided by the present invention;

[0045] Figure 8 is a schematic diagram of the vocabulary matching result provided by the present invention;

[0046] Figure 9 is a schematic diagram of the element boundary information of each character provided by the present invention;

[0047] Figure 10 It is a schematic structural diagram of the element extraction device provided by the present invention;

[0048] Figure 11 It is a schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners

[0049] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0050] In recent years, with the development of computing power and the progress of neural network theory, deep learning has made great progress in element extraction, and model encoders of different variants (such as CNN, RNN, Transformer, etc.) have been verified to be effective in the task of named entity recognition (NER).

[0051] A relatively representative model for fusing lexical information is the FLAT model. This method first matches domain vocabulary in the text, splices the matched vocabulary after the original sentence, introduces unique position information encoding, and then uses BERT for encoding to model the information interaction between characters and vocabulary.

[0052] Similar work also includes the Lex-BERT model. Figure 1 It is a schematic structural diagram of the Lex-BERT model. As Figure 1 shown, this method splices the part-of-speech information of the matched vocabulary after the original sentence. The part-of-speech symbol shares the position encoding with the first character of the word it matches, which can enhance the boundary recognition ability of the model and thus model the interaction between characters and vocabulary.

[0053] In summary, the existing element extraction methods that fuse vocabulary need to splice several vocabulary on the original sentence, resulting in an increase in the input length and affecting the encoding efficiency of the model; moreover, the size of the word embedding layer of the deep model is proportional to the size of the domain vocabulary table. Some domain vocabulary tables can reach the million level. Storing a word embedding layer of such a size requires an additional large amount of storage space.

[0054] How to improve the existing element extraction methods to improve the encoding efficiency and reduce the storage space while ensuring the accuracy is still an urgent problem to be solved in the field of natural language processing.

[0055] Figure 2 It is one of the flow diagrams of the element extraction method provided by the present invention. As Figure 2 shown, the element extraction method provided by the present invention can be applied to any application scenario that needs to identify and extract specific elements from corpus texts. For example, it is applicable to element extraction in the general field and can also be applied to the extraction of case elements in legal documents in professional fields such as the judicial field. The method includes:

[0056] Step 210, obtain the text to be extracted and the vocabulary set of the text to be extracted.

[0057] Specifically, the text to be extracted is the text for which element extraction is required. Element extraction refers to extracting entities, attributes, and corresponding attribute values from unstructured text and outputting structured data.

[0058] The text to be extracted can be directly input by the user, or obtained by transcribing the collected audio, or obtained by collecting images through image acquisition devices such as scanners, mobile phones, cameras, etc. and performing OCR (Optical Character Recognition) on the images. The text to be extracted can also be obtained after preprocessing such as deleting, de-duplicating, or splicing the text obtained by the above text acquisition methods. The embodiments of the present invention do not limit this.

[0059] The vocabulary set is the natural language processing dictionary. The vocabulary set of the text to be extracted refers to the set containing the vocabulary in the text to be extracted. The vocabulary set here can be a domain vocabulary set related to the professional field involved in the text to be extracted, such as the judicial field or the financial field, etc., or a general domain vocabulary set. The vocabulary set can be preset.

[0060] It should be noted that the vocabulary set can include each vocabulary and the corresponding vocabulary type. For example, the vocabulary set includes "Nanjing" and the corresponding vocabulary type "location type" of "Nanjing".

[0061] Step 220, based on the matching result between the string corresponding to every two characters in the text to be extracted and the vocabulary set, determine the relevance between every two characters. The string is intercepted from the text to be extracted with the corresponding two characters as the starting and ending points.

[0062] Specifically, after obtaining the text to be extracted and the corresponding vocabulary set, the string formed by taking every two characters in the text to be extracted as the starting and ending points can be matched with the vocabulary set. The string here is intercepted from the text to be extracted with the corresponding two characters as the starting and ending points. The string can contain only two characters or multiple characters more than two.

[0063] The matching results between the string and the vocabulary set include successful matching and failed matching. Successful matching means that there is a vocabulary in the vocabulary set that matches the string, and failed matching means that there is no vocabulary in the vocabulary set that matches the string. Here, the successfully matched string can match the vocabulary in the form of a continuous substring or a discontinuous substring. Both the continuous substring and the discontinuous substring here are in terms of vocabulary.

[0064] For example, the text to be extracted is "The Fifth Yangtze River Bridge in Nanjing is located about 13 kilometers upstream of the Nanjing Yangtze River Bridge". The successfully matched strings such as "Nanjing City" and "Mayor" are obtained by matching the vocabulary of continuous substrings; while the string "Nanjing Fifth Yangtze River Bridge" is obtained by matching the vocabulary of the discontinuous substring "Nanjing Fifth Yangtze River Bridge". Similarly, the string "Yangtze River Bridge" matches the discontinuous substring vocabulary "Fifth Yangtze River Bridge".

[0065] According to the matching results between the string corresponding to every two characters and the vocabulary set, the relevance between every two characters can be determined. The relevance between every two characters can characterize the association degree between every two characters in the text to be extracted. Here, if the matching result between the string and the vocabulary set is successful, it indicates that the relevance between these two characters is relatively high; if the matching fails, it indicates that the relevance between these two characters is relatively low. At the same time, in the process of determining the relevance, the vocabulary matching result is fused, which can further analyze the context information of the text to be extracted and make the result of element extraction more accurate.

[0066] The determination of the relevance can be implemented by a relevance calculation algorithm, such as calculation methods like the vector space model, probability model, or attention mechanism, etc. The embodiments of the present invention do not make specific limitations.

[0067] Step 230, based on the relevance between every two characters, encode each character in the text to be extracted to obtain the element boundary features of each character.

[0068] Specifically, considering that in the existing element extraction methods, usually the vocabulary matched in the vocabulary set is concatenated with the original sentence, and the position information encoding method is introduced to obtain the element boundary features of each character. This method will increase the input length and affect the encoding efficiency of the model. The method provided by the embodiments of the present invention, according to the relevance between every two characters, since the determination of the relevance fully fuses the vocabulary matching result, when encoding each character in the text to be extracted, the model can dynamically adjust the element boundary of the character according to the vocabulary matching result.

[0069] During the encoding process, for example, the characters in the text to be extracted can be encoded based on the attention mechanism. According to the relevance between every two characters, the information of each character in the text to be extracted is saved, without the need to splice the matched words with the original sentence, and the original input length is not changed, thus improving the encoding efficiency.

[0070] The element boundary features of each character cover not only the semantic information of each character itself, but also the semantic information of other characters that can be connected with the character to form words, so as to be able to reflect the possibility of the character itself as an element boundary.

[0071] Step 240: Determine the element extraction result of the text to be extracted based on the element boundary features of each character.

[0072] Specifically, after encoding each character in the text to be extracted to obtain the element boundary features of each character, the element boundary features of each character can be decoded to obtain the element extraction result of the text to be extracted. The element extraction result here can be structured element information, such as named entities such as person names, place names, and organization names.

[0073] In the method provided by the embodiment of the present invention, during the process of determining the relevance between every two characters in the text to be extracted, the vocabulary matching result is fused, and each character in the text to be extracted is encoded according to the relevance between every two characters. This method does not need to splice the matched words with the original sentence, and does not change the original input length, thus improving the encoding efficiency. In addition, compared with the existing method of vocabulary splicing, the storage space is saved.

[0074] Based on the above embodiment, Figure 3 is the flow diagram of step 220 in the element extraction method provided by the present invention, as Figure 3 shown, step 220 specifically includes:

[0075] Step 221: Select the relevance parameter corresponding to every two characters from the candidate relevance parameters corresponding to multiple candidate matching results based on the matching result between the string corresponding to every two characters and the vocabulary set;

[0076] Step 222: Determine the relevance between every two characters based on the relevance parameter corresponding to every two characters and the semantic representation of every two characters.

[0077] Specifically, the relevance between every two characters can be determined according to the relevance parameter corresponding to every two characters and the semantic representation of every two characters. Among them, the relevance parameter corresponding to every two characters can be selected from the candidate relevance parameters corresponding to multiple candidate matching results according to the matching result between the string corresponding to every two characters and the vocabulary set.

[0078] As can be seen from the above embodiments, the matching results between the strings corresponding to every two characters and the vocabulary set include successful matching and failed matching. Among them, successful matching can also include matching continuous substrings and matching discontinuous substrings. The candidate matching results here can be failed matching, matching continuous substrings, and matching discontinuous substrings. In some embodiments, the candidate matching results can also include rule matching, etc. And the above various candidate matching results can respectively correspond to different candidate relevance parameters. The corresponding relationship between the candidate matching results and the candidate relevance parameters can be one-to-one correspondence, or multiple candidate matching results can correspond to a group of candidate relevance parameters. The embodiments of the present invention do not make specific limitations on this.

[0079] Among them, the candidate relevance parameters corresponding to the candidate matching results can be pre-set and unchanged, or can be learnable parameters. For example, they can be fine-tuned according to the field to which the introduced vocabulary belongs.

[0080] The semantic representation of every two characters can be represented by the feature vectors obtained by mapping the sentences where the characters are located to different representation spaces.

[0081] For any two characters, after determining the relevance parameters of the above two characters, the relevance parameters and the semantic representations of the above two characters can be substituted into a pre-set calculation formula, so as to obtain the relevance between the above two characters. For example, they can be the calculation formulas corresponding to the additive model, dot product model, or bilinear model respectively.

[0082] The method provided by the embodiments of the present invention determines the relevance between every two characters through the relevance parameters corresponding to the vocabulary matching results and the semantic representations of every two characters. Instead of encoding by splicing the vocabulary and the original sentence, it introduces the relevance parameters through the vocabulary matching results, so as to encode each character in the text to be extracted according to the relevance between every two characters, and obtain the element boundary features of each character.

[0083] Based on any of the above embodiments, step 221 specifically includes:

[0084] Based on the matching results between the strings corresponding to every two characters and the vocabulary set, and the matching results between the strings corresponding to every two characters and the preset rules, select the relevance parameters corresponding to every two characters from the candidate relevance parameters corresponding to multiple candidate matching results.

[0085] Specifically, the preset rules here can be domain rules designed using the characteristics of the domain. The matching results between the strings corresponding to every two characters and the preset rules, for example, can use regular expressions to match ID card numbers, passport numbers, mobile phone numbers, etc. Correspondingly, the candidate matching results in this embodiment can also include rule matching.

[0086] If the strings corresponding to every two characters fail to match both the vocabulary set and the preset rules, the relevance parameter corresponding to the matching failure is selected as the relevance parameter corresponding to every two characters; if the matching result of the string corresponding to every two characters and the preset rules is a successful match, the relevance parameter corresponding to the matching of the preset rules is selected as the relevance parameter corresponding to every two characters; if the matching result of the string corresponding to every two characters and the vocabulary set is a continuous substring match, the relevance parameter corresponding to the continuous substring match is selected as the relevance parameter corresponding to every two characters; if the matching result of the string corresponding to every two characters and the vocabulary set is a non - continuous substring match, the relevance parameter corresponding to the non - continuous substring match is selected as the relevance parameter corresponding to every two characters.

[0087] The method provided by the embodiments of the present invention extends the process of matching a string with a vocabulary set to matching a string with a vocabulary and preset rules, and integrates the matching results into the method for determining relevance, further enhancing the detection ability of element boundaries, thereby improving the accuracy of element extraction.

[0088] Based on any of the above embodiments, Figure 4 is a schematic flowchart of the method for determining the relevance between every two characters provided by the present invention, as Figure 4 shown, step 222 specifically includes:

[0089] Step 2221, based on the semantic representation of every two characters, determine the semantic relevance between every two characters;

[0090] Step 2222, based on the relevance parameter corresponding to every two characters and the semantic representation of every two characters, determine the lexical relevance between every two characters;

[0091] Step 2223, based on the semantic relevance and lexical relevance between every two characters, determine the relevance between every two characters.

[0092] Specifically, the semantic relevance between every two characters can be determined according to the semantic representation of every two characters. The semantic representation of every two characters can be represented by feature vectors obtained by mapping to different representation spaces, and the semantic relevance between every two characters can characterize the relevance in terms of semantics between every two characters. For example, in a certain representation space, a cat and a dog belong to the same semantic category, and the semantic relevance between a cat and a dog is higher than the semantic relevance between a cat and a flower.

[0093] Further, the lexical relevance between every two characters can be determined according to the relevance parameter corresponding to every two characters and the semantic representation of every two characters. Among them, the lexical relevance between every two characters can characterize the relevance in terms of vocabulary between every two characters. The higher the lexical relevance, the greater the possibility that every two characters can form a vocabulary.

[0094] Then, the relevance between every two characters can be determined according to the semantic relevance and lexical relevance between every two characters. For example, the semantic relevance and lexical relevance between every two characters can be weighted and averaged to obtain the relevance between every two characters; or the two can be directly added or multiplied to obtain the relevance between every two characters.

[0095] For example, the relevance between every two characters can be represented by the following formula:

[0096]

[0097] where, e i,j is the relevance between the i-th character and the j-th character, q i is the semantic representation of the i-th character, k j is the semantic representation of the j-th character, is the vector dimension, is the semantic relevance between the i-th character and the j-th character, is the lexical relevance between the i-th character and the j-th character, W lex(i,j) and b lex(i,j) are the relevance parameters between the i-th character and the j-th character, and lex(i, j) is the value corresponding to the matching result of the string formed by the i-th character to the j-th character with the vocabulary set or preset rules.

[0098] Among them, the value of lex(i, j) has three types, as shown in Table 1 below:

[0099] Table 1

[0100] Matching result Value taken Matching failed 0 Matching continuous substring 1 Rule matching 1 Matching non - continuous substring 2

[0101] The method provided by the embodiment of the present invention determines the relevance between every two characters through the semantic relevance and lexical relevance between every two characters. In the process of calculating the relevance, not only the semantic features of each character are integrated, but also the vocabulary matching result is integrated, improving the accuracy of element extraction.

[0102] Based on any of the above embodiments, Figure 5 is the schematic flowchart of step 240 in the element extraction method provided by the present invention. As Figure 5 shown, step 240 specifically includes:

[0103] Step 241: Based on the element boundary features of each character, determine the segment representations of each candidate segment in the text to be extracted and the lexical representations of each word in the vocabulary set.

[0104] Step 242: Based on the association information between each candidate segment and each word in the knowledge graph, determine the relevance between each candidate segment and each word.

[0105] Step 243: Based on the relevance between each candidate segment and each word, as well as the segment representations of each candidate segment and the lexical representations of each word, determine the element extraction result of the text to be extracted.

[0106] Specifically, after obtaining the element boundary features of each character, decode the element boundary features to obtain the element extraction result of the text to be extracted. Currently, the mainstream decoding methods include Conditional Random Field (CRF) and segment classification. Initially, deep learning was used for sequence labeling, and only Softmax normalization was performed on the model output layer. The results decoded by this method do not consider the correlation between labels and there will be many illegal decoding sequences. Later, researchers combined CRF and added label transition scores during decoding, effectively alleviating this problem; segment classification is different from the idea of sequence labeling. The segment classification method enumerates all possible segments in the sentence, classifies them, and then removes boundary conflicts to obtain the elements to be extracted. Thanks to its decoding strategy, this type of decoding method will not have the problem of illegal decoding sequences.

[0107] Considering that the current decoding method based on segment classification does not explicitly consider the lexical information around the current segment, ambiguity problems may be introduced during the vocabulary matching process, resulting in inaccurate extracted elements. To address this problem, the embodiment of the present invention proposes a decoding strategy that integrates vocabulary and knowledge graph based on the segment classification decoding method. During decoding, use the knowledge in the vocabulary and knowledge graph to eliminate the ambiguity introduced during the vocabulary matching process. The following specifically describes the decoding strategy that integrates vocabulary and knowledge graph proposed by the embodiment of the present invention.

[0108] According to the element boundary features of each character, each candidate segment in the text to be extracted can be determined. For example, 10 possible element boundaries, including start boundaries and / or end boundaries, are found in the text to be extracted. For any segment formed by every two boundaries (start boundary and end boundary), it constitutes a candidate segment.

[0109] The segment representation of each candidate segment can include the start position / end position word embedding representation of the segment, or the self-attention weighted average representation of the segment, etc. The embodiment of the present invention does not make specific limitations on this. Correspondingly, the lexical representation method of each word in the vocabulary set is the same as that of the segment representation of each candidate segment.

[0110] According to the association information between each candidate segment and each vocabulary in the knowledge graph, the relevance between each candidate segment and each vocabulary can be obtained. Among them, the knowledge graph is pre-set, and the knowledge graph can be regarded as a "vocabulary-relationship-vocabulary" triple. The association information between each candidate segment and each vocabulary in the knowledge graph includes: each candidate segment and each vocabulary form a triple in the knowledge graph, or each candidate segment and each vocabulary do not form a triple in the knowledge graph. From the association information between each candidate segment and each vocabulary in the knowledge graph, the relevance between each candidate segment and each vocabulary can be obtained.

[0111] The determination of the relevance can be implemented by a relevance calculation algorithm, such as calculation methods like the vector space model or the probability model, etc. The embodiments of the present invention do not make specific limitations.

[0112] Then, according to the relevance between each candidate segment and each vocabulary, as well as the segment representation of each candidate segment and the vocabulary representation of each vocabulary, the element extraction result of the text to be extracted is determined. For example, the vocabulary representations of each vocabulary can be fused as the segment representation of each candidate segment, and then the segment representation of each candidate segment is input into the decoder to obtain the element extraction result of the text to be extracted.

[0113] The method provided by the embodiments of the present invention classifies each candidate segment through a segment representation method that integrates the association information between each candidate segment and each vocabulary in the knowledge graph on the basis of the decoding method of segment classification, so as to obtain the element extraction result of the text to be extracted.

[0114] Based on any of the above embodiments, step 242 specifically includes:

[0115] Based on the association information between each candidate segment and each vocabulary in the knowledge graph, and the overlapping inclusion relationship between each candidate segment and each vocabulary, the relevance between each candidate segment and each vocabulary is determined.

[0116] Specifically, when each candidate segment and each vocabulary do not form a triple in the knowledge graph, there may be an overlapping inclusion relationship between each candidate segment and each vocabulary. According to the possible overlapping inclusion relationship between each candidate segment and each vocabulary, a relevance parameter corresponding to the overlapping inclusion relationship can be obtained, and based on the relevance parameter, the relevance between each candidate segment and each vocabulary is determined.

[0117] Exemplarily, the association information between each candidate segment and each vocabulary in the knowledge graph, and the overlapping inclusion relationship between each candidate segment and each vocabulary can be shown as in Table 2.

[0118] Table 2

[0119] Relationship type number Relationship type Description 0 Non - overlapping The segment and the vocabulary have no boundary overlap 1 Left overlap The left boundary of the segment partially overlaps with the right boundary of the vocabulary 2 Right overlap The right boundary of the segment partially overlaps with the left boundary of the vocabulary 3 Contain The segment contains the vocabulary 4 Be contained The vocabulary contains the segment 5 Triple The segment and the vocabulary are a triple in the knowledge graph

[0120] The relevance between each candidate segment and each vocabulary can be expressed by the following formula:

[0121]

[0122] Where α i is the relevance between any candidate segment and vocabulary l i , x is the segment representation of any candidate segment, is the representation of vocabulary l i in the vocabulary set, rel(x, l i ) is the relationship type number between any candidate segment and vocabulary l i , and are the relevance parameters corresponding to the relationship type numbers between any candidate segment and vocabulary l i .

[0123] The method provided by the embodiment of the present invention determines the relevance between each candidate segment and each vocabulary through the association information between each candidate segment and each vocabulary in the knowledge graph, and the overlapping inclusion relationship between each candidate segment and each vocabulary.

[0124] Based on any of the above embodiments, Figure 6 is a schematic flowchart of the method for determining the element extraction result provided by the present invention. As Figure 6 shown, step 243 specifically includes:

[0125] Step 2431, based on the relevance between each candidate segment and each vocabulary, and the vocabulary representation of each vocabulary, fuse the vocabulary representations of each vocabulary to obtain the segment representation of each candidate segment after fusion;

[0126] Step 2432, based on the segment representation of each candidate segment after fusion, determine the element extraction result of the text to be extracted.

[0127] Specifically, after obtaining the relevance between each candidate segment and each vocabulary, the vocabulary representations of each vocabulary can be fused based on the vocabulary representation and relevance of each vocabulary to obtain the segment representation of each candidate segment after fusion. The fusion method can be weighted fusion, difference method, ratio method, etc. The segment representation of each candidate segment after fusion integrates the information of vocabulary and knowledge graph.

[0128] Then, according to the segment representation of each candidate segment after fusion, determine the element extraction result of the text to be extracted. It can be to input the segment representation of each candidate segment after fusion into the segment classification decoder. Since the segment representation of each candidate segment after fusion integrates the information of vocabulary and knowledge graph, when classifying each candidate segment, the vocabulary information and knowledge graph information around the segment are considered simultaneously, so that the result of element extraction is more accurate.

[0129] In one embodiment, when classifying candidate segments, the calculation formula of the decoding method that simultaneously considers all the lexical information around the segments can be expressed as follows:

[0130] x′ = [x; LexAtt(x, L, KG)]

[0131] y = softmax(W s x′ + b s ), where

[0132]

[0133] In the above formula, LexAtt(x, L, KG) is the segment representation of any candidate segment that integrates lexical and knowledge graph information, x is the segment representation of any candidate segment, L is the set of words, KG is the knowledge graph, and α i is the relevance between any candidate segment and the word l i , is the representation of the word l i in the set of words.

[0134] The method provided by the embodiments of the present invention determines the element extraction result of the text to be extracted by means of the segment representations of the candidate segments that integrate lexical and knowledge graph information, further improving the accuracy of element extraction.

[0135] Based on any of the above embodiments, Figure 7 is the second schematic flow diagram of the element extraction method provided by the present invention. As Figure 7 shown, the method includes:

[0136] First, the unstructured text to be parsed is used as input, and after domain word matching processing, the word boundaries in the text are found; then, the text and the word matching result are input into the domain word enhanced element extraction model together, and the model will further detect possible element boundaries in combination with the word matching result; then, the domain words and the domain knowledge graph are used to eliminate the conflicts and ambiguities existing in the element boundaries, and finally the element extraction result is obtained.

[0137] For example, the text to be extracted includes "The Fifth Nanjing Yangtze River Bridge is located about 13 kilometers upstream of the Nanjing Yangtze River Bridge". It is necessary to extract two "facility" type elements, namely "The Fifth Nanjing Yangtze River Bridge" and "The Nanjing Yangtze River Bridge", from this sentence. The existing domain words are shown in Table 3 as follows:

[0138] Table 3

[0139] Serial number Domain vocabulary Vocabulary type 1 Nanjing GPE 2 Nanjing City GPE 3 Yangtze River LOC 4 Yangtze River Bridge FAC 5 The Fifth Yangtze River Bridge in Nanjing FAC 6 The Yangtze River Bridge in Nanjing City FAC 7 Mayor PER 8 … …

[0140] Note: GPE stands for geopolitical element; LOC stands for location element; FAC stands for facility element; PER stands for name element.

[0141] Considering that the existing domain vocabulary matching only considers the continuous string in the text, due to the complexity of natural language, some additional characters are often interspersed between entities (such as structural auxiliary words "的", "之", etc.). For non-continuous matching, the existing methods often cannot play their due role.

[0142] The method provided by the embodiment of the present invention considers both continuous character strings and non-continuous character strings during the matching process between a character string and a vocabulary set, thereby improving the accuracy of feature extraction.

[0143] The results after domain vocabulary matching are as follows Figure 8 As shown in the figure, there are words matched in the form of continuous strings, such as "Nanjing City", "Mayor", etc.; there are also words matched in the form of non-continuous strings, for example, the word "Nanjing Yangtze River Fifth Bridge" is matched in a non-continuous form to the sentence "Nanjing City Yangtze River Fifth Bridge", similarly, the word "Yangtze River Bridge" is matched to the sentence "Yangtze River Fifth Bridge".

[0144] In order to enhance the model encoder's perception of element boundaries, an embodiment of the present invention provides a boundary enhancement attention mechanism, which integrates domain vocabulary into the model encoder by adding domain vocabulary bias to the attention mechanism formula.

[0145] For example, the calculation formula for the attention paid by “京” to “长” in “京长” is:

[0146]

[0147] W0 and b0 represent q i and k j The words cannot be formed between them, which will not affect the normal flow of attention. i and k j When words can be formed between them, the model will adjust the flow of attention according to the type of matching words (continuous or discontinuous), so that the model can dynamically perceive the boundaries of the elements and will not change the original input length.

[0148] The resulting feature boundaries are Figure 9As shown in the figure, domain vocabulary can help the model find 10 possible feature boundaries, with subscripts [0,1,2,3,8,11,12,13,14,17], where [0,11] is only the starting boundary of the feature, [1,8,12,13] is only the ending boundary of the feature, and [2,3,14] can be both the starting boundary and the ending boundary. In fact, in addition to the boundaries located by domain vocabulary, the model itself may also find some candidate feature boundaries, for example, "江" in "江大桥" may be the starting boundary of the name feature.

[0149] When classifying the segment "River Bridge", existing methods do not actually explicitly consider the domain vocabulary information around the current segment. At this time, the model may mistakenly classify "River Bridge" as a "name" type feature; similarly, "Yangtze River Bridge" may also be mistakenly classified as a "facility" type feature, but in fact the boundary of this segment in this example is incomplete.

[0150] When judging the type of a fragment, the method provided by the embodiment of the present invention fully considers the surrounding vocabulary information. For example, when judging the type of "Jiang Bridge", the method provided by the present invention will consider the surrounding vocabulary information, including "Nanjing City", "Nanjing Yangtze River Bridge", etc., so that it is more confident to classify "Jiang Bridge" as "non-element". In addition, when a fragment (or the domain vocabulary matched by the fragment) can form a triple with a certain domain vocabulary in the sentence, the fragment is actually more likely to become an element. For example, if the knowledge graph contains the triple Nanjing-Mayor-Jiang Bridge, it is more confident to classify "Jiang Bridge" as a "personal name element".

[0151] The method provided by the embodiment of the present invention integrates the boundary information introduced by the domain vocabulary into the calculation formula of the attention mechanism, does not change the length of the original input, and takes into account the situation of non-continuous vocabulary matching, thereby enhancing the model's ability to detect element boundaries; at the same time, a decoding strategy that integrates domain vocabulary and domain knowledge graph is proposed. During decoding, the knowledge in the domain vocabulary and knowledge graph is used to eliminate the ambiguity introduced in the domain vocabulary matching process, thereby further improving the accuracy of element extraction.

[0152] Based on the above embodiments, in view of the problem that the existing method of introducing vocabulary will lead to inconsistency between pre-training and fine-tuning of the factor extraction model, thereby weakening the effect of the pre-training model, the embodiment of the present invention proposes a method for fine-tuning the factor extraction model, which can easily alleviate the problem. The specific method is: in the initial stage of fine-tuning, the other parameters of the pre-training model are frozen, and only the correlation parameters between every two characters are updated. After the model has been trained for a period of time, all the parameters are unfrozen and updated synchronously, thereby alleviating the inconsistency problem from "pre-training to fine-tuning" caused by the introduction of domain vocabulary.

[0153] The following describes the element extraction device provided by the present invention. The element extraction device described below can be correspondingly referred to the element extraction method described above. Figure 10 is a schematic structural diagram of the element extraction device provided by the present invention, as Figure 10 shown, the device includes:

[0154] A text acquisition unit 1010, configured to acquire the text to be extracted, and the vocabulary set of the text to be extracted;

[0155] A relevance determination unit 1020, configured to determine the relevance between every two characters based on the matching result between the string corresponding to every two characters in the text to be extracted and the vocabulary set, where the string is intercepted from the text to be extracted with the corresponding two characters as the start and end points;

[0156] An encoding unit 1030, configured to encode each character in the text to be extracted based on the relevance between every two characters, to obtain the element boundary feature of each character;

[0157] An extraction result determination unit 1040, configured to determine the element extraction result of the text to be extracted based on the element boundary feature of each character.

[0158] In the element extraction device provided by the embodiment of the present invention, during the process of determining the relevance between every two characters in the text to be extracted, the vocabulary matching result is fused, and each character in the text to be extracted is encoded according to the relevance between every two characters. This method does not need to splice the matched vocabulary with the original sentence, and does not change the original input length, thereby improving the encoding efficiency. In addition, compared with the existing vocabulary splicing method, it saves storage space.

[0159] Based on the above embodiment, the relevance determination unit 1020 is further configured to:

[0160] Based on the matching result between the string corresponding to every two characters and the vocabulary set, select the relevance parameter corresponding to every two characters from the candidate relevance parameters corresponding to multiple candidate matching results;

[0161] Based on the relevance parameter corresponding to every two characters and the semantic representation of every two characters, determine the relevance between every two characters.

[0162] Based on any of the above embodiments, the relevance determination unit 1020 is further configured to:

[0163] Based on the matching results between the strings corresponding to every two characters and the vocabulary set, as well as the matching results between the strings corresponding to every two characters and the preset rules, select the relevance parameters corresponding to every two characters from the candidate relevance parameters corresponding to multiple candidate matching results.

[0164] Based on any of the above embodiments, the relevance determination unit 1020 is further configured to:

[0165] Determine the semantic relevance between every two characters based on the semantic representations of every two characters;

[0166] Determine the lexical relevance between every two characters based on the relevance parameters corresponding to every two characters and the semantic representations of every two characters;

[0167] Determine the relevance between every two characters based on the semantic relevance and lexical relevance between every two characters.

[0168] Based on any of the above embodiments, the extraction result determination unit 1040 is further configured to:

[0169] Determine the segment representations of each candidate segment in the text to be extracted and the lexical representations of each vocabulary in the vocabulary set based on the element boundary features of each character;

[0170] Determine the relevance between each candidate segment and each vocabulary based on the association information between each candidate segment and each vocabulary in the knowledge graph;

[0171] Determine the element extraction result of the text to be extracted based on the relevance between each candidate segment and each vocabulary, as well as the segment representation of each candidate segment and the lexical representation of each vocabulary.

[0172] Based on any of the above embodiments, the extraction result determination unit 1040 is further configured to:

[0173] Determine the relevance between each candidate segment and each vocabulary based on the association information between each candidate segment and each vocabulary in the knowledge graph and the overlapping inclusion relationship between each candidate segment and each vocabulary.

[0174] Based on any of the above embodiments, the extraction result determination unit 1040 is further configured to:

[0175] Fuse the lexical representations of each vocabulary based on the relevance between each candidate segment and each vocabulary and the lexical representation of each vocabulary to obtain the segment representation of each candidate segment after fusion;

[0176] Determine the element extraction result of the text to be extracted based on the segment representation of each candidate segment after fusion.

[0177] Figure 11 Schematically illustrates the physical structure of an electronic device, such as Figure 11 shown. The electronic device may include: a processor 1110, a communications interface 1120, a memory 1130, and a communication bus 1140. Among them, the processor 1110, the communications interface 1120, and the memory 1130 communicate with each other through the communication bus 1140. The processor 1110 may call logic instructions in the memory 1130 to execute an element extraction method, which includes: obtaining the text to be extracted and the vocabulary set of the text to be extracted; based on the matching result between the string corresponding to each two characters in the text to be extracted and the vocabulary set, determining the relevance between each two characters, the string being intercepted in the text to be extracted with the corresponding two characters as the starting and ending points; based on the relevance between each two characters, encoding each character in the text to be extracted to obtain the element boundary feature of each character; based on the element boundary feature of each character, determining the element extraction result of the text to be extracted.

[0178] In addition, when the logic instructions in the above-mentioned memory 1130 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.

[0179] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the element extraction method provided by each of the above methods. The method includes: obtaining the text to be extracted and the vocabulary set of the text to be extracted; determining the relevance between every two characters based on the matching result between the string corresponding to every two characters in the text to be extracted and the vocabulary set, where the string is intercepted from the text to be extracted with the corresponding two characters as the starting and ending points; encoding each character in the text to be extracted based on the relevance between every two characters to obtain the element boundary features of each character; and determining the element extraction result of the text to be extracted based on the element boundary features of each character.

[0180] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the element extraction method provided by each of the above methods. The method includes: obtaining the text to be extracted and the vocabulary set of the text to be extracted; determining the relevance between every two characters based on the matching result between the string corresponding to every two characters in the text to be extracted and the vocabulary set, where the string is intercepted from the text to be extracted with the corresponding two characters as the starting and ending points; encoding each character in the text to be extracted based on the relevance between every two characters to obtain the element boundary features of each character; and determining the element extraction result of the text to be extracted based on the element boundary features of each character.

[0181] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0182] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, also by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0183] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for element extraction, characterized in that, including: obtaining the text to be extracted and the vocabulary set of the text to be extracted; determining the relevance between every two characters based on the matching result between the string corresponding to every two characters in the text to be extracted and the vocabulary set, where the string is intercepted from the text to be extracted with the corresponding two characters as the start and end points; encoding each character in the text to be extracted based on the relevance between every two characters to obtain the element boundary features of each character; the element boundary features include the semantic information of each character and the semantic information of other characters that are connected with each character to form a vocabulary; determining the element extraction result of the text to be extracted based on the element boundary features of each character; the determining the element extraction result of the text to be extracted based on the element boundary features of each character includes: determining the segment representation of each candidate segment in the text to be extracted and the vocabulary representation of each vocabulary in the vocabulary set based on the element boundary features of each character; each candidate segment is a segment composed of any two element boundaries in the text to be extracted, and the element boundary includes a start boundary and an end boundary; determining the relevance between each candidate segment and each vocabulary based on the association information between each candidate segment and each vocabulary in the knowledge graph; determining the element extraction result of the text to be extracted based on the relevance between each candidate segment and each vocabulary, the segment representation of each candidate segment, and the vocabulary representation of each vocabulary.

2. The element extraction method according to claim 1, characterized in that the determining the relevance between every two characters based on the matching result between the string corresponding to every two characters in the text to be extracted and the vocabulary set includes: selecting the relevance parameter corresponding to every two characters from the candidate relevance parameters corresponding to multiple candidate matching results based on the matching result between the string corresponding to every two characters and the vocabulary set; determining the relevance between every two characters based on the relevance parameter corresponding to every two characters and the semantic representation of every two characters.

3. The element extraction method according to claim 2, wherein the selecting the relevance parameter corresponding to every two characters from the candidate relevance parameters corresponding to multiple candidate matching results based on the matching result between the string corresponding to every two characters and the vocabulary set includes: selecting the relevance parameter corresponding to every two characters from the candidate relevance parameters corresponding to multiple candidate matching results based on the matching result between the string corresponding to every two characters and the vocabulary set and the matching result between the string corresponding to every two characters and the preset rules.

4. The element extraction method according to claim 2, characterized in that the determining the relevance between every two characters based on the relevance parameter corresponding to every two characters and the semantic representation of every two characters includes: determining the semantic relevance between every two characters based on the semantic representation of every two characters; determining the vocabulary relevance between every two characters based on the relevance parameter corresponding to every two characters and the semantic representation of every two characters; determining the relevance between every two characters based on the semantic relevance and vocabulary relevance between every two characters.

5. The element extraction method according to claim 1, wherein Determining the relevance between each candidate segment and each vocabulary based on the association information between each candidate segment and each vocabulary in the knowledge graph includes: Determining the relevance between each candidate segment and each vocabulary based on the association information between each candidate segment and each vocabulary in the knowledge graph and the overlapping inclusion relationship between each candidate segment and each vocabulary.

6. The element extraction method according to any one of claims 1-5, characterized in that Determining the element extraction result of the text to be extracted based on the relevance between each candidate segment and each vocabulary, as well as the segment representation of each candidate segment and the vocabulary representation of each vocabulary, includes: Based on the relevance between each candidate segment and each vocabulary and the vocabulary representation of each vocabulary, fusing the vocabulary representations of each vocabulary to obtain the segment representation of each candidate segment after fusion; Determining the element extraction result of the text to be extracted based on the segment representation of each candidate segment after fusion.

7. An element extraction device, characterized in that, Including: A text acquisition unit for acquiring the text to be extracted and the vocabulary set of the text to be extracted; A relevance determination unit for determining the relevance between every two characters based on the matching result between the string corresponding to every two characters in the text to be extracted and the vocabulary set, where the string is intercepted from the text to be extracted with the corresponding two characters as the start and end points; An encoding unit for encoding each character in the text to be extracted based on the relevance between every two characters to obtain the element boundary feature of each character; the element boundary feature includes the semantic information of each character and the semantic information of other characters that are connected with each character to form a vocabulary; An extraction result determination unit for determining the element extraction result of the text to be extracted based on the element boundary feature of each character; The extraction result determination unit is used for: Based on the element boundary feature of each character, determining the segment representation of each candidate segment in the text to be extracted and the vocabulary representation of each vocabulary in the vocabulary set; each candidate segment is a segment composed of any two element boundaries in the text to be extracted, and the element boundary includes a start boundary and an end boundary; Determining the relevance between each candidate segment and each vocabulary based on the association information between each candidate segment and each vocabulary in the knowledge graph; Determining the element extraction result of the text to be extracted based on the relevance between each candidate segment and each vocabulary, as well as the segment representation of each candidate segment and the vocabulary representation of each vocabulary.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the element extraction method according to any one of claims 1 to 6 are implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the element extraction method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Chinese word segmentation method, training equipment and computer readable storage medium

    CN111666758A