Information extraction method, device, electronic device, computer-readable medium and product
By extracting the triple combination sets in text information in the information extraction method, finding alternative sets and filtering target sets with high accuracy, the problem of low triple recall in the prior art is solved, and the accuracy and recall rate of information extraction are improved.
Patent Information
- Application Number
- CN202111418839.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-24
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-11-24
AI Technical Summary
Existing information extraction methods perform poorly in triple recall, resulting in low accuracy and recall of information extraction.
An information extraction method is proposed. By extracting the ternary combination set corresponding to the sentence from the text information, finding the alternative ternary combination set corresponding to the keyword, and filtering the target ternary combination set based on the accuracy, storing the target ternary combination set consistent with the semantics of the sentence.
The recall rate of the ternary combination set is improved, the probability of extracting correct information from text information is enhanced, and the overall effect of information extraction is improved.
Smart Images

Figure CN114090793B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of knowledge graphs, and more specifically, to an information extraction method, device, electronic device, computer-readable medium, and product. Background Art
[0002] Currently, knowledge graphs have important applications in many fields. An important step in constructing a knowledge graph is the extraction of triples. Generally, existing methods directly extract triples using an information extraction model. However, the recall rate of the triples obtained by this method is not high. Summary of the Invention
[0003] This application provides an information extraction method, device, electronic device, computer-readable medium, and product.
[0004] In a first aspect, an embodiment of this application provides an information extraction method applied to text information extraction. The method includes: extracting a triple combination set corresponding to each sentence from the text information, where the triple combination set includes a subject, a predicate, and an object; finding an alternative triple combination set corresponding to each keyword; determining a first accuracy rate corresponding to each keyword based on a first quantity of the alternative triple combination set corresponding to each keyword and a total quantity of the triple combination set; finding an alternative triple combination set corresponding to the first accuracy rate that meets a specified condition as a target triple combination set; and storing the target triple combination set if the semantics of the target triple combination set is consistent with the semantics of the sentence corresponding to the target triple combination set.
[0005] In a second aspect, an embodiment of this application further provides a posture monitoring device applied to text information extraction. The device includes: an extraction unit, a first processing unit, a second processing unit, a third processing unit, and a fourth processing unit. Among them, the extraction unit is configured to extract a triple combination set corresponding to each sentence from the text information, where the triple combination set includes a subject, a predicate, and an object; the first processing unit is configured to find an alternative triple combination set corresponding to each keyword; the second processing unit is configured to determine a first accuracy rate corresponding to each keyword based on a first quantity of the alternative triple combination set corresponding to each keyword and a total quantity of the triple combination set; the third processing unit is configured to find an alternative triple combination set corresponding to the first accuracy rate that meets a specified condition as a target triple combination set; and the fourth processing unit is configured to store the target triple combination set if the semantics of the target triple combination set is consistent with the semantics of the sentence corresponding to the target triple combination set.
[0006] In a third aspect, an embodiment of the present application further provides an electronic device, including: one or more processors; a memory; an image acquisition device; one or more application programs, where the one or more application programs are stored in the memory and configured to be executed by the one or more processors, and the one or more application programs are configured to execute the above method.
[0007] In a fourth aspect, an embodiment of the present application further provides a computer-readable medium, where the readable storage medium stores program code executable by a processor, and when the program code is executed by the processor, the processor executes the above method.
[0008] In a fifth aspect, an embodiment of the present application further provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the above method is implemented.
[0009] The information extraction method, device, electronic device, computer-readable medium, and product provided by the present application are applied to text information extraction. The method first extracts a set of triple combinations corresponding to each sentence from the text information, then looks up a set of alternative triple combinations corresponding to each keyword, and by determining a first accuracy rate corresponding to each keyword, takes the set of alternative triple combinations corresponding to the first accuracy rate that meets the specified conditions as the target triple combination set. If all the extracted triple combination sets are directly stored, the recall rate will be low. By storing the target triple combination set that is semantically consistent with the sentences corresponding to the target triple combination set, the recall rate of the triple combination set is improved, that is, the probability of extracting correct information from the text information is increased.
[0010] Other features and advantages of the embodiments of the present application will be described in the subsequent description of the specification, and some of them will be obvious from the description of the specification, or understood by implementing the embodiments of the present application. The objectives and other advantages of the embodiments of the present application can be achieved and obtained through the structures specifically pointed out in the written specification, claims, and drawings. Description of the Drawings
[0011] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0012] Figure 1 Shows a scenario diagram of the information extraction method provided by an embodiment of the present application;
[0013] Figure 2 Shows the method flow chart of the information extraction method provided by the embodiment of the present application.
[0014] Figure 3 shows Figure 2 an implementation of step S210 in
[0015] Figure 4 a schematic diagram of the information extraction method provided by the embodiment of the present application;
[0016] Figure 5 shows Figure 2 an implementation of step S240 in
[0017] Figure 6 a method flowchart of the information extraction method provided by another embodiment of the present application;
[0018] Figure 7 shows Figure 6 an implementation of step S650 in
[0019] Figure 8 a schematic diagram of the information extraction method provided by another embodiment of the present application;
[0020] Figure 9 shows Figure 6 an implementation of step S660 in
[0021] Figure 10 a method flowchart of the information extraction method provided by yet another embodiment of the present application;
[0022] Figure 11 shows Figure 10 an implementation of step S1030 in
[0023] Figure 12 a unit block diagram of the information extraction device provided by the embodiment of the present application;
[0024] Figure 13 a schematic diagram of the electronic device provided by the embodiment of the present application;
[0025] Figure 14 a structural block diagram of the computer-readable storage medium provided by the embodiment of the present application;
[0026] Figure 15 a structural block diagram of the computer program product provided by the embodiment of the present application. Detailed implementation manners
[0027] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Usually, the components of the embodiments of this application described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the drawings is not intended to limit the scope of the claimed application, but merely represents the selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative efforts belong to the scope of protection of this application.
[0028] It should be noted that: Similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of this application, terms such as "first" and "second" are only used for differential description and cannot be understood as indicating or implying relative importance.
[0029] Information extraction plays a very important role in artificial intelligence applications. More and more upper-layer applications rely on the results of information extraction. For example, knowledge graphs rely on techniques such as entity relationship extraction, event extraction, and causal relationship extraction; the construction of query and decision support systems in fields such as law and medicine also relies on the results returned by information extraction. Among them, the extraction of information generally can be manifested as the extraction of triples.
[0030] Specifically, data can be represented by objects and the relationships between objects. For example, between two large companies A and B, the relationship between the two companies can be constructed through a knowledge graph. For example, in a piece of news, it is described as follows: "The debt ratio of Company A is 50%, and Company B holds 20% of the shares of Company A." Then the following entities and the relationships between the entities can be extracted from this description: The first one: Entity subject: Company A, Relationship: Debt ratio, Debt ratio: 50%. The second one: Entity subject: Company B, Entity object: Company A, Relationship between entity subject and entity object: Shareholding, Share ratio: 20%. By storing the obtained triple data in the corresponding database and constructing a knowledge graph, the relationships in a specific field between companies can be constructed through this knowledge graph, which has enlightenment and guiding effects on the strategic development and planning of companies.
[0031] However, the inventors found in the research that the omission or incorrect extraction of objects in the triples affects the results of information extraction to varying degrees. That is, in the existing information extraction methods, the accuracy and recall rate of information extraction are relatively low.
[0032] Therefore, in order to overcome the above-mentioned deficiencies, the embodiments of the present application provide an information extraction method, apparatus, electronic device, computer-readable medium, and product, which are applied to text information extraction. The method first extracts a triple combination set corresponding to each sentence from the text information, then searches for an alternative triple combination set corresponding to each keyword, and by determining the first accuracy rate corresponding to each keyword, takes the alternative triple combination set corresponding to the first accuracy rate that meets the specified conditions as the target triple combination set, and stores the target triple combination set that is semantically consistent with the sentence corresponding to the target triple combination set.
[0033] Please refer to Figure 1 , Figure 1 FIG. shows an information extraction method provided by an embodiment of the present application. This method can be applied to a text information extraction scenario 100, which includes text information 110 and an electronic device 120. The electronic device 120 includes a text information extraction system 121, and the text information 110 is connected to the electronic device 120.
[0034] Among them, the text information 110 is the text information to be processed, that is, the information to be extracted. The text information 110 can be input into the text information extraction system 120 for information extraction. For some embodiments, the text information 110 can be a sentence or a paragraph. Specifically, the text information 110 can be a document, which can include a sentence or a paragraph. For example, a typist manually inputs a paragraph of text, generates a document, and then uses this document as the text information 110; or through optical character recognition technology, the text information in an image file is recognized to generate a document, and then this document is used as the text information 110; or through speech recognition technology, specific speech is converted into text information, a document is generated, and then this document is used as the text information 110.
[0035] The electronic device 120 is used to provide an input / output interface for the text information 110 and use the text information extraction system 121 to process the input text information 110. Among them, the electronic device 120 can be a device with processing capabilities such as a smart phone or a tablet computer.
[0036] The information extraction system 121 is used to extract the input text information 110 and can extract specific factual information from the text information 110. Usually, the extracted information is described in a structured form and can be directly stored in a database for users to query and further analyze and utilize. For an embodiment provided by the present application, the structured form can be a triple combination set. The information extraction system 121 can be a system running on the electronic device 120 or an application program running on the operating system of the electronic device 120.
[0037] Please refer to Figure 2 , Figure 2 which shows an information extraction method provided by an embodiment of the present application. This method can be applied to the text information extraction scenario 100 in the foregoing embodiment, and the execution subject of this method can be an electronic device. Specifically, this method includes steps S210 to S250.
[0038] Step S210: Extract the triple set corresponding to each sentence from the text information, where the triple set includes a subject, a predicate, and an object.
[0039] For some embodiments, the triple set corresponding to each sentence in the text information should include a subject, a predicate, and an object. For example, if the text information is "Xiaoming likes Xiaowang", a triple set ["Xiaoming", "likes", "Xiaowang"] can be extracted. Among them, "Xiaoming" is the subject, "likes" is the predicate, and "Xiaowang" is the object. For each sentence in the text information, extracting its corresponding triple set should be able to correctly express or approximately correctly express the meaning expressed by the sentence corresponding to the triple set. For example, for the above text information "Xiaoming likes Xiaowang", if the extracted triple set is ["Xiaoming", "likes", "Xiaowang"], where "Xiaoming" is the subject, "likes" is the predicate, and "Xiaowang" is the object, then this triple set can correctly express the meaning expressed by this sentence; if the extracted triple set is ["Xiaowang", "likes", "Xiaoming"], where "Xiaowang" is the subject, "likes" is the predicate, and "Xiaoming" is the object, then this triple set cannot correctly express the meaning expressed by this sentence.
[0040] Furthermore, for some embodiments, the text information can be a single sentence or a paragraph. For example, the text information can be composed of sentence A, or composed of sentence A + sentence B, or can also be sentence A + sentence B + sentence C... where "..." indicates that the number of subsequent sentences can be an indefinite number. For each sentence in the text information, its corresponding triple set can be extracted. For example, for sentence A, triple set A can be extracted, for sentence B, triple set B can be extracted, and for sentence C, triple set C can be extracted. Specifically, if the input text is "Xiaowang likes to eat apples. Xiaoming likes to eat watermelons. Neither Xiaoming nor Xiaowang likes to eat bananas.", then the corresponding sentence A is: "Xiaowang likes to eat apples", sentence B is: "Xiaoming likes to eat watermelons", and sentence C is: "Neither Xiaoming nor Xiaowang likes to eat bananas". The triple set A corresponding to sentence A can be extracted as: ["Xiaowang", "likes", "apples"], the triple set B corresponding to sentence B can be extracted as: ["Xiaoming", "likes", "watermelons"], and the triple set C corresponding to sentence C can be extracted as: ["Xiaoming and Xiaowang", "don't like", "bananas"].
[0041] Further, please refer to Figure 3 , in order to illustrate step S210 in detail, this step S210 may further include steps S211 to S215.
[0042] Step S211: Input the text information into a pre-trained model to obtain a second feature vector corresponding to each sentence and a subject start vector corresponding to each sentence.
[0043] Step S212: Fuse the subject start vector and the second feature vector to obtain a subject vector.
[0044] Step S213: Use the text corresponding to the subject vector as the subject.
[0045] For some embodiments, the pre-trained model may be a BERT language model. This BERT language model can be pre-trained using a large-scale pre-trained corpus, thereby to a certain extent compensating for the problem caused by a small number of samples. The pre-trained model can be trained by inputting a pre-trained corpus package into an initial model. For example, financial information, news magazine texts, etc. can be used as the pre-trained corpus package to train the initial model to obtain the pre-trained model, or the pre-trained model that has completed training can be directly obtained from a network server. Further, during the training process, the training corpus package and its corresponding feature vectors can also be input into the initial model, and the training sentences predicted by the initial model are compared with the key sentences in the corpus package. If the two are the same, it means that the training of the initial model has been completed. If the two are different, it means that the model parameters of the initial model need to be changed and the initial model needs to be continued to be trained. After the training is completed, the initial model and its model parameters together form the pre-trained model.
[0046] For some embodiments, the subject corresponding to the sentence in the text information can be obtained by inputting the text information into a pre-trained model. Further, the pre-trained model can split the input text information into combinations with words as the basic unit. For example, if the text information is "Xiaowang likes to eat apples", after inputting the text information into the pre-trained model, the following combination can be obtained: ["Xiao", "wang", "likes", "to", "eat", "apples"].
[0047] Furthermore, because the same text can express completely different meanings due to different orders, for example, for the two sentences "Xiao Wang likes Xiao Ming" and "Xiao Ming likes Xiao Wang", the text included in the two sentences is exactly the same, but due to the different order of the text, the meanings they represent are different. Therefore, for each text in the combination, a position representation vector for representing its position and a content representation vector for representing its content can also be obtained respectively. Among them, the position representation vector is used to represent the specific position of the text in the input text information, and the content representation vector is used to represent the specific content represented by the text. For an embodiment provided in the present application, the pre-trained model is a BERT model, and the BERT model uses absolute position encoding, that is, for each position of the text, there is a separate vector corresponding to it. Specifically, the vector x1 can be used to represent the first position in the input text information, the vector x2 can be used to represent the second position in the input text information, and so on. For example, if the input text information is "Xiao Ming likes to eat apples", the vector x1 can be used to represent the position of the text "Xiao", the vector x2 can be used to represent the position of "Ming", and so on, the vector x7 can be used to represent the position of "Guo".
[0048] See also Figure 4 For some implementations, after the text information is input into the pre-trained model, a second feature vector h for characterizing the input text information may also be obtained. s And the subject start position representation vector can be predicted, and the subject start position representation vector is the subject start vector. By fusing the second feature vector with the subject start vector, the subject vector can be obtained, and the text corresponding to the subject vector is the subject. For example, if vector x1 is used to represent the subject start vector, then h s By fusing with x1, we can obtain the vectors corresponding to x1, x2, and x3 as the subject vector.
[0049] Step S214: Fusing the subject vector and the second feature vector to obtain a predicate vector and an object vector.
[0050] Step S215: taking the text corresponding to the predicate vector as the predicate, and taking the text corresponding to the object vector as the object.
[0051] Please continue reading Figure 4 For some implementations, each vector in the subject vector obtained in the above steps can be added together to obtain an intermediate vector, and then the intermediate vector can be fused with the second feature vector to predict the predicate vector and the object vector. For example, the subject vector contains vector x1, vector x2, and vector x3. x1, vector x2, and vector x3 can be fused to obtain an intermediate vector V sub , and then use the intermediate vector and the second eigenvector hs are fused to predict a predicate vector and an object vector, where the text corresponding to the predicate vector is the predicate and the text corresponding to the object vector is the object. For example, Figure 4 in which vector x5 and vector x6 form the predicate vector, and vector x7 and vector x8 form the object vector.
[0052] Step S220: Search for a set of alternative triple combinations corresponding to each keyword.
[0053] For some embodiments, the set of triple combinations obtained by extraction in the above steps can be graded, and the set of triple combinations that meet certain conditions can be directly stored. For the remaining set of triple combinations, further screening is performed, and the set of triple combinations with better indicators among them is selected for storage. For the remaining set of triple combinations with poor indicators, they can be marked and then stored, or directly discarded. By grading the set of triple combinations and then selectively storing them, the recall rate of the stored set of triple combinations can be improved, and the quality of information extraction can be enhanced.
[0054] Furthermore, some keywords can be specified, and the set of triple combinations that meet the keyword is used as the set of alternative triple combinations. For example, if the specified keyword is the predicate: "like", then for the extracted set of triple combinations, the set of triple combinations with the predicate: "like" are all used as the set of alternative triple combinations. If the extracted set of triple combinations is that triple combination set A is: ["Xiaowang", "like", "apple"], triple combination set B is: ["Xiaoming", "like", "watermelon"], and triple combination set C is: ["Xiaoming and Xiaowang", "don't like", "banana"], then in the case where the keyword is "like", it can be determined that the set of alternative triple combinations is triple combination set A and triple combination set C.
[0055] Step S230: Determine the first accuracy rate corresponding to each keyword based on the first quantity of the set of alternative triple combinations corresponding to each keyword and the total quantity of the set of triple combinations.
[0056] For some embodiments, the first accuracy rate of the set of alternative triple combinations corresponding to each keyword can be determined based on the set of alternative triple combinations corresponding to each keyword, so that the set of triple combinations can be graded according to the first accuracy rate.
[0057] Further, the first accuracy rate can be obtained by the first quantity of the alternative triple combination sets corresponding to each keyword and the total quantity of the triple combination sets. Specifically, for some embodiments, the first accuracy rate can be expressed as the proportion of the first quantity of the alternative triple combination sets corresponding to each keyword in the total quantity of the triple combination sets. For example, if based on the keyword "like", the total quantity of the extracted triple combination sets is 900, that is, this first quantity is 900. If the total quantity of the triple combination sets is 1000, at this time, 900 / 1000 is 90%, that is, this first accuracy rate is 90%.
[0058] Step S240: Search for the alternative triple combination sets corresponding to the first accuracy rate that meet the specified conditions as the target triple combination sets.
[0059] For some embodiments, specific triple combination sets can be screened out through specified conditions, and these specific triple combination sets are used as the target triple combination sets for a second screening. The target triple combination sets that pass the second screening are stored, so as to improve the recall rate in information extraction.
[0060] For other embodiments, some triple combination sets that meet specific conditions can also be directly stored.
[0061] Further, please refer to Figure 5 , Figure 5 which shows a further illustration of an embodiment proposed in step S241.
[0062] Step S241: Search for the alternative triple combination sets corresponding to the first accuracy rate whose first accuracy rate is less than the target value as the target triple combination sets.
[0063] For some embodiments, a target value can be specified, and this target value is judged with the first accuracy rate, so that the alternative triple combination sets corresponding to the first accuracy rate less than this target value are used as the target triple combination sets. Specifically, for example, if the target value is set to 98%, then when the first accuracy rate of 96% is obtained, since 96% is less than 98%, the triple combination sets corresponding to this first accuracy rate can be used as the target triple combination sets. If the target value is set to 98%, then when the first accuracy rate of 99% is obtained, since 99% is greater than 98%, the triple combination sets corresponding to this first accuracy rate can be directly stored. Among them, the storage can be directly depositing the target triple combination sets into the database as incremental data.
[0064] Step S250: If the semantics of the target triple combination sets are consistent with the semantics of the sentences corresponding to the target triple combination sets, then store the target triple combination sets.
[0065] As can be seen from the foregoing steps, the first accuracy rate corresponding to the target triple set is lower than the specified accuracy rate. Therefore, after making a judgment using the accuracy rate, it can be known that the semantics of the target triple set may be inconsistent with the semantics of the sentence corresponding to the target triple set. Therefore, for some embodiments, it is possible to determine whether the semantics of the target triple set is consistent with the semantics of the sentence corresponding to the target triple set. If the semantics are consistent, the target triple set is stored. By judging the target triple set that may not conform to the semantics of the sentence corresponding to the target triple set and storing the triple set whose semantics may be consistent, the recall rate of the extracted triple set can be improved. Among them, the method for judging whether the semantics of the target triple set is consistent with the semantics of the sentence corresponding to the target triple set can refer to the subsequent embodiments.
[0066] The information extraction method, device, electronic device, computer-readable medium and product provided by the present application are applied to text information extraction. The method first extracts the triple set corresponding to each sentence from the text information, then searches for the alternative triple sets corresponding to each keyword, and determines the first accuracy rate corresponding to each keyword by determining the first quantity of the alternative triple sets corresponding to each keyword and the total quantity of the triple sets. The alternative triple set corresponding to the first accuracy rate that meets the specified conditions is used as the target triple set. If all the extracted triple sets are directly stored, the recall rate will be low. By storing the target triple set whose semantics is consistent with the semantics of the sentence corresponding to the target triple set, the recall rate of the triple set is improved, that is, the probability of extracting correct information from the text information is increased.
[0067] Please refer to Figure 6 , Figure 6 FIG. shows an information extraction method provided by an embodiment of the present application. This method can be applied to the text information extraction scenario 100 in the foregoing embodiment, and the execution subject of this method can be an electronic device. Specifically, this method includes steps S610 to S670.
[0068] Step S610: Extract the triple set corresponding to each sentence from the text information, where the triple set includes a subject, a predicate, and an object.
[0069] Step S620: Search for the alternative triple sets corresponding to each keyword.
[0070] Step S630: Based on the first quantity of the alternative triple sets corresponding to each keyword and the total quantity of the triple sets, determine the first accuracy rate corresponding to each keyword.
[0071] Step S640: Search for the alternative triple set corresponding to the first accuracy rate that meets the specified conditions and use it as the target triple set.
[0072] Among them, for steps S610 to S640, by extracting the triple combination set corresponding to each sentence from the text information, and then determining the target triple combination set through specified conditions, the specific method has been described in detail in the foregoing embodiments and will not be elaborated here.
[0073] Step S650: Construct a specified sentence based on the sentences corresponding to the target triple combination set, where each target triple combination set corresponds to a specified sentence.
[0074] For some embodiments, since the target triple combination set may not be consistent with the meaning expressed by the sentence corresponding to the target triple combination set. Therefore, in order to confirm whether the target triple combination is consistent with the meaning expressed by the sentence corresponding to the target triple combination set, a specified sentence can be constructed based on the sentence corresponding to the target triple combination set. For the specific method of constructing this specified sentence, reference can be made to Figure 7 step S651 in
[0075] Step S651: Concatenate the target triple combination set and the sentence corresponding to the target triple combination set to construct the specified sentence.
[0076] For some embodiments, please refer to Figure 8 , the target triple combination set and the sentence corresponding to the target triple combination set can be concatenated to obtain the specified sentence. Among them, the sentence formed by the target triple combination set is in the front, and the sentence in the text information corresponding to the target triple combination set is in the back. For example, if the target triple combination set is ["Zhang San", "wife", "Li Si"], the sentence formed by the target triple combination set can be "Zhang San's wife is Li Si". If the sentence in the text information corresponding to the target triple combination set is "Zhang San's wife is Li Si", then the specified sentence can be obtained as [Zhang San's wife is Li Si. Zhang San's wife is Li Si.]. Further, for some other embodiments, an identifier can also be added between the two clauses of the specified sentence to separate the two sentences. For example, the identifier [SEP] can be used as the separator, and the obtained specified sentence is [Zhang San's wife is Li Si. [SEP] Zhang San's wife is Li Si.]. Further, an identifier can also be added at the beginning of the specified sentence to mark the specified sentence. For example, [CLS] can be used as the identifier, and the specified sentence can be [[CLS] Zhang San's wife is Li Si. [SEP] Zhang San's wife is Li Si.]
[0077] Step S660: Obtain the similarity between the target triple combination set and the specified sentence corresponding to the target triple combination set.
[0078] For some embodiments, to confirm whether the meaning of the target triple set is consistent with the specified sentence corresponding to the target triple set, the specified sentence can be input into a pre-trained model to obtain a score through the pre-trained model, and then the consistency can be judged based on this score. Specifically, please refer to Figure 9 , Figure 9 shows a detailed description of step S660, including steps S661 to S663.
[0079] Step S661: Obtain the first feature vector corresponding to the target triple set and the second feature vector of the specified sentence corresponding to the target triple set.
[0080] Step S662: Fuse the first feature vector and the second feature vector into a third feature vector.
[0081] For some embodiments, the vector corresponding to the triple set included in the specified sentence can be the first feature vector, and the vector corresponding to the specified sentence can be the second feature vector. The first feature vector and the second feature vector can be fused to obtain a third feature vector. For example, if the specified sentence is [[CLS]Zhang San's wife is Li Si. [SEP]Zhang San's spouse is Li Si.], then the flag vector [CLS] at the beginning of the sentence can be used as the CLS feature vector as the third feature vector.
[0082] Step S663: Input the third feature vector into the pre-trained model for similarity calculation to obtain the similarity between the target triple set and the specified sentence corresponding to the target triple set.
[0083] For some embodiments, the pre-trained model further includes a fully connected layer (FCL) and an activation function layer (sigmoid) for identifying the input vector and fitting the vector. After the pre-trained model obtains the input third feature vector, the third feature vector can be input into the fully connected layer and the activation function layer, and the similarity can be calculated through the pre-trained model. Further, the similarity can be a score from 0 to 1. The closer the score is to 1, the higher the similarity between the target triple set and the specified sentence corresponding to the target triple set; the closer the score is to 0, the lower the similarity between the target triple set and the specified sentence corresponding to the target triple set.
[0084] Step S670: If the similarity is greater than the specified threshold, store the target triple set.
[0085] For some embodiments, by setting a specified threshold, it can be determined whether the output similarity meets the requirements, so as to determine whether the similarity between the target triple set and the specified sentence corresponding to the target triple set meets the requirements. Specifically, the target triple set with a similarity greater than the specified threshold can be stored. For example, the specified threshold can be set to 0.8. When the similarity between the target triple set and the specified sentence corresponding to the target triple set is 0.9, since 0.9 is greater than 0.8, this target triple set can be stored.
[0086] For other embodiments, if the similarity between the target triple set and the specified sentence corresponding to the target triple set is less than the specified threshold, the target triple set can be marked and then stored, or the target triple can be directly discarded.
[0087] The information extraction method, device, electronic device, computer-readable medium and product provided by this application are applied to text information extraction. This method first extracts the triple set corresponding to each sentence from the text information, then searches for the alternative triple sets corresponding to each keyword, determines the first accuracy rate corresponding to each keyword, takes the alternative triple set corresponding to the first accuracy rate that meets the specified conditions as the target triple set, then constructs a specified sentence, and by judging the similarity between the target triple set and the specified sentence corresponding to the target triple set, stores the triple set with a similarity greater than the specified threshold. Avoid directly storing all the obtained triple sets, and selectively store the triple sets by setting a specified threshold, which improves the recall rate of the triple sets, that is, improves the probability of extracting correct information from the text information.
[0088] Please refer to Figure 10 , Figure 10 FIG. shows an information extraction method provided by an embodiment of this application. This method can be applied to the text information extraction scenario 100 in the foregoing embodiment, and the execution subject of this method can be an electronic device. Specifically, this method includes steps S1010 to S1080.
[0089] Step S1010: Extract the triple set corresponding to each sentence from the text information, where the triple set includes a subject, a predicate, and an object.
[0090] Among them, how step S1010 extracts the triple set corresponding to each sentence from the text information has been described in detail in the foregoing embodiment, and will not be elaborated here.
[0091] Step S1020: Search for the triple sets lacking an object from all the triple sets.
[0092] For some embodiments, in the triple combination set obtained by extracting the text information through a pre-trained model, there may be a situation where the object is missing. The triple combination set missing the object can be found in the triple combination set. Among them, the situation of missing the object can be that the object is not extracted in the extracted triple combination set. For example, if the text information is "Xiaoming likes to eat apples and watermelons.", and the extracted triple combination set is ["Xiaoming", "likes", ""], then the object is not extracted in this triple combination set. The situation of missing the object can also be that the object in the extracted triple combination set is only a part of the object in the sentence semantics corresponding to this triple combination set. For example, if the text information is "Xiaoming likes to eat apples and watermelons.", and the extracted triple combination set is ["Xiaoming", "likes", "eat apples"], then the complete object is not extracted in this triple combination set.
[0093] Step S1030: Determine the object to be supplemented corresponding to the triple combination set missing the object based on the recall and supplementation model, where the recall and supplementation model is a model established based on machine reading comprehension.
[0094] For some embodiments, the triple combination set missing the object can be processed and then input into a processing model, and through this processing model, the supplementation of the missing object can be realized. Specifically, this processing model can be a recall and supplementation model, and this recall and supplementation model can be a model established based on machine reading comprehension (MRC). Among them, machine reading comprehension MRC is a technology that enables a computer to understand the semantics of an article and answer related questions using algorithms. Further, please refer to Figure 11 , Figure 11 shows an embodiment of step S1030. Step S1030 may further include step S1031 and step S1032.
[0095] Step S1031: Determine the associated statement of the sentence corresponding to the triple combination set missing the object based on the keyword corresponding to the triple combination set missing the object and the subject.
[0096] For some embodiments, an associated statement can be constructed based on the keyword corresponding to the triple combination set and the subject of the triple combination set. For example, if the triple combination set is ["Zhang San", "wife", ""], it can be known that the subject corresponding to this triple combination set is "Zhang San". Further, if the keyword corresponding to this triple combination set is "wife", then an associated statement can be constructed through the subject "Zhang San" and the keyword "wife": "Who is Zhang San's wife?".
[0097] Step S1032: Input the associated statement and the text information into the recall and supplement model to obtain the object to be supplemented corresponding to the set of triple combinations lacking an object.
[0098] For some embodiments, the associated statement can be concatenated with the sentences in the text information corresponding to the set of triple combinations, and the concatenated sentence is then input into the recall and supplement model for prediction. For example, if the associated statement is "Who is Zhang San's wife.", and the sentence in the text information corresponding to the set of triple combinations is: "Zhang San's wife is Li Si.", then they can be concatenated as: "Who is Dehua's wife. Zhang San's wife is Li Si.". Further, a separator symbol [SEP] can be added to the concatenated sentence to distinguish the two clauses in the concatenated sentence, and an identifier [CLS] can be added at the beginning of the concatenated sentence to indicate the start of the sentence, that is, the concatenated sentence can be "[CLS]Who is Dehua's wife[SEP]Zhang San's wife is Li Si". Inputting the concatenated sentence into the recall and supplement model can obtain the object to be supplemented corresponding to the set of triple combinations lacking an object. For example, inputting the above sentence into the recall and supplement model, the output can be ["Li Si"], and this output is the object lacking in the set of triple combinations.
[0099] Step S1040: Supplement the object to be supplemented into the set of triple combinations lacking an object.
[0100] For some embodiments, the object lacking in the set of triple combinations has been obtained through the above steps. Concatenating this object with the set of triple combinations lacking an object can obtain a set of triple combinations including a complete subject, predicate, and object. For example, if the set of triple combinations lacking an object is ["Zhang San", "wife", ""], and the lacking object obtained through the previous steps is ["Li Si"], then they can be concatenated into a complete set of triple combinations ["Zhang San", "wife", "Li Si"].
[0101] Step S1050: Search for the set of alternative triple combinations corresponding to each keyword.
[0102] Step S1060: Based on the first quantity of the set of alternative triple combinations corresponding to each keyword and the total quantity of the set of triple combinations, determine the first accuracy rate corresponding to each keyword.
[0103] Step S1070: Search for the set of alternative triple combinations corresponding to the first accuracy rate that meets the specified conditions as the target set of triple combinations.
[0104] Step S1080: If the semantics of the target set of triple combinations is consistent with the semantics of the sentence corresponding to the target set of triple combinations, store the target set of triple combinations.
[0105] Among them, steps S1050 to S1080 have been described in detail in the foregoing embodiments, and will not be elaborated here.
[0106] The information extraction method, device, electronic device, computer-readable medium and product provided by this application are applied to text information extraction. This method first extracts the triple combination set corresponding to each sentence from the text information, then searches for the alternative triple combination set corresponding to each keyword, completes the triple combination set lacking an object, and determines the first accuracy rate corresponding to each keyword, and uses the alternative triple combination set corresponding to the first accuracy rate that meets the specified conditions as the target triple combination set. If all the extracted triple combination sets are directly stored, the recall rate will be low. By checking the extracted triple combination sets, completing the triple combination sets lacking an object, and then determining whether to store them in the follow-up, the recall rate of the triple combination sets can be improved, that is, the probability of extracting correct information from the text information can be improved.
[0107] Please refer to Figure 12 , which shows a structural block diagram of an information extraction device 1200 provided by an embodiment of this application. This device is applied to text information extraction. The device may include: an extraction unit 1210, a first processing unit 1220, a second processing unit 1230, a third processing unit 1240, and a fourth processing unit 1250.
[0108] The extraction unit 1210 is used to extract the triple combination set corresponding to each sentence from the text information, and the triple combination set includes a subject, a predicate, and an object.
[0109] Further, the extraction unit 1210 is further used to input the text information into a pre-trained model, obtain the second feature vector corresponding to each sentence and the subject start vector corresponding to each sentence; fuse the subject start vector and the second feature vector to obtain a subject vector; use the text corresponding to the subject vector as the subject; fuse the subject vector and the second feature vector to obtain a predicate vector and an object vector; use the text corresponding to the predicate vector as the predicate, and use the text corresponding to the object vector as the object; based on the subject, the predicate, and the object, determine the triple combination set.
[0110] Further, the extraction unit 1210 is further used to search for the triple combination sets lacking an object from all the triple combination sets; determine the object to be supplemented corresponding to the triple combination sets lacking an object based on a recall and supplement model, and the recall and supplement model is a model established based on machine reading comprehension; supplement the object to be supplemented into the triple combination sets lacking an object.
[0111] Further, the extraction unit 1210 is further configured to determine an associated statement of the sentence corresponding to the triple combination set lacking an object based on the keyword corresponding to the triple combination set lacking an object and the subject; input the associated statement and the text information into the recall and supplementation model to obtain the object to be supplemented corresponding to the triple combination set lacking an object.
[0112] The first processing unit 1220 is configured to find an alternative triple combination set corresponding to each keyword.
[0113] The second processing unit 1230 is configured to determine a first accuracy rate corresponding to each keyword based on the first quantity of the alternative triple combination set corresponding to each keyword and the total quantity of the triple combination sets.
[0114] The third processing unit 1240 is configured to find the alternative triple combination set corresponding to the first accuracy rate that meets the specified condition as the target triple combination set.
[0115] Further, the third processing unit 1240 is further configured to find the alternative triple combination set corresponding to the first accuracy rate whose first accuracy rate is less than the target value as the target triple combination set.
[0116] The fourth processing unit 1250 is configured to store the target triple combination set if the semantics of the target triple combination set is consistent with the semantics of the sentence corresponding to the target triple combination set.
[0117] Further, the fourth processing unit 1250 is further configured to construct a specified sentence based on the sentence corresponding to the target triple combination set, with each target triple combination set corresponding to a specified sentence; obtain the similarity between the target triple combination set and the specified sentence corresponding to the target triple combination set; and store the target triple combination set if the similarity is greater than the specified threshold.
[0118] Further, the fourth acquisition unit 1250 is further configured to obtain a first feature vector corresponding to the target triple combination set and a second feature vector of the specified sentence corresponding to the target triple combination set; fuse the first feature vector and the second feature vector into a third feature vector; and input the third feature vector into a pre-trained model for similarity calculation to obtain the similarity between the target triple combination set and the specified sentence corresponding to the target triple combination set.
[0119] Further, the fourth acquisition unit 1250 is further configured to splice the target triple combination set and the sentence corresponding to the target triple combination set to construct the specified sentence.
[0120] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described devices and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0121] In several embodiments provided by the present application, the coupling between units can be electrical, mechanical, or other forms of coupling.
[0122] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0123] Please refer to Figure 13 , which shows a structural block diagram of an electronic device provided by an embodiment of the present application. The electronic device 1300 can be an electronic device such as a smart phone, a tablet computer, an e-book, etc. that can run application programs. The electronic device 1300 in the present application can include one or more of the following components: a processor 1310, a memory 1320, and one or more application programs, where one or more application programs can be stored in the memory 1320 and configured to be executed by one or more processors 1310, and one or more programs are configured to execute the methods described in the foregoing method embodiments.
[0124] The processor 1310 may include one or more processing cores. The processor 1310 connects various parts within the entire electronic device 1300 through various interfaces and lines. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 1320, and by invoking the data stored in the memory 1320, it performs various functions of the electronic device 1300 and processes data. Optionally, the processor 1310 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 1310 may integrate a combination of one or several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the displayed content; the modem is used to process wireless communications. It can be understood that the above modem may not be integrated into the processor 1310 and may be implemented separately through a communication chip.
[0125] The memory 1320 may include random access memory (RAM) and may also include read-only memory. The memory 1320 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 1320 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for implementing at least one function, instructions for implementing the following various method embodiments, etc. The data storage area may also store the data created during the use of the electronic device 1300.
[0126] Please refer to Figure 14 , which shows a structural block diagram of a computer-readable storage medium provided by an embodiment of the present application. Program code is stored in the computer-readable medium 1400, and the program code can be called by a processor to execute the method described in the above method embodiments.
[0127] The computer-readable storage medium 1400 can be an electronic memory such as a flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, a hard disk, or a ROM. Optionally, the computer-readable storage medium 1400 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 1400 has a storage space for program code 1410 that executes any of the method steps in the above-described methods. These program codes can be read from or written into one or more computer program products. The program code 1410 can be compressed in a suitable form, for example.
[0128] Please refer to Figure 15 , which shows a structural block diagram 1500 of a computer program product provided by an embodiment of the present application. The computer program product 1500 includes a computer program / instructions 1510, and when the computer program / instructions 1510 are executed by a processor, the steps of the above-described method are implemented.
[0129] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. An information extraction method, characterized in that, applied to text information extraction, the method includes: extracting a triple combination set corresponding to each sentence from the text information, the triple combination set including a subject, a predicate, and an object; finding an alternative triple combination set corresponding to each set keyword; determining a first accuracy rate corresponding to each keyword based on a first quantity of the alternative triple combination set corresponding to each keyword and a total quantity of the triple combination set; finding the alternative triple combination set corresponding to the first accuracy rate that meets a specified condition as a target triple combination set; constructing a specified sentence based on the sentence corresponding to the target triple combination set, each target triple combination set corresponding to one specified sentence; obtaining a similarity between the target triple combination set and the specified sentence corresponding to the target triple combination set; if the similarity is greater than a specified threshold, storing the target triple combination set.
2. The method according to claim 1, characterized in that, the obtaining the similarity between the target triple combination set and the specified sentence corresponding to the target triple combination set includes: obtaining a first feature vector corresponding to the target triple combination set and a second feature vector of the specified sentence corresponding to the target triple combination set; fusing the first feature vector and the second feature vector into a third feature vector; inputting the third feature vector into a pre-trained model for similarity calculation to obtain the similarity between the target triple combination set and the specified sentence corresponding to the target triple combination set.
3. The method according to claim 1, characterized in that, the constructing the specified sentence based on the sentence corresponding to the target triple combination set includes: concatenating the target triple combination set and the sentence corresponding to the target triple combination set to construct the specified sentence.
4. The method according to claim 1, characterized in that, the text information includes at least one sentence, each sentence including at least one subject, at least one predicate, and at least one object, and the extracting the triple combination set corresponding to each sentence from the text information includes: inputting the text information into a pre-trained model to obtain a second feature vector corresponding to each sentence and a subject start vector corresponding to each sentence; fusing the subject start vector and the second feature vector to obtain a subject vector; using the text corresponding to the subject vector as the subject; fusing the subject vector and the second feature vector to obtain a predicate vector and an object vector; using the text corresponding to the predicate vector as the predicate and the text corresponding to the object vector as the object; determining the triple combination set based on the subject, the predicate, and the object.
5. The method according to claim 1, characterized in that, before the finding the alternative triple combination set corresponding to each keyword, it further includes: finding the triple combination set lacking an object from all the triple combination sets; determining a to-be-supplemented object corresponding to the triple combination set lacking an object based on a recall and supplement model, the recall and supplement model being a model established based on machine reading comprehension; supplementing the to-be-supplemented object into the triple combination set lacking an object.
6. The method according to claim 5, wherein, the text information includes at least one sentence, and each of the sentences includes at least one subject. Determining the object to be supplemented corresponding to the set of triple combinations lacking an object based on the recall and supplementation model includes: determining an associated statement of the sentence corresponding to the set of triple combinations lacking an object based on the keyword corresponding to the set of triple combinations lacking an object and the subject; inputting the associated statement and the text information into the recall and supplementation model to obtain the object to be supplemented corresponding to the set of triple combinations lacking an object.
7. The method according to claim 1, wherein, the specified condition is that the first accuracy rate is less than the target value. Searching for the set of alternative triple combinations corresponding to the first accuracy rate that meets the specified condition as the target triple combination set includes: searching for the set of alternative triple combinations corresponding to the first accuracy rate with the first accuracy rate less than the target value as the target triple combination set.
8. An information extraction device, wherein, applied to text information extraction, the device includes: an extraction unit configured to extract a set of triple combinations corresponding to each sentence from the text information, the set of triple combinations including a subject, a predicate, and an object; a first processing unit configured to search for a set of alternative triple combinations corresponding to each set keyword; a second processing unit configured to determine a first accuracy rate corresponding to each set keyword based on a first quantity of the set of alternative triple combinations corresponding to each keyword and a total quantity of the set of triple combinations; a third processing unit configured to search for a set of alternative triple combinations corresponding to the first accuracy rate that meets the specified condition as the target triple combination set; a fourth processing unit configured to construct a specified sentence based on the sentence corresponding to the target triple combination set, with each target triple combination set corresponding to one specified sentence; obtaining a similarity between the target triple combination set and the specified sentence corresponding to the target triple combination set; and storing the target triple combination set if the similarity is greater than a specified threshold.
9. An electronic device, wherein, comprising: one or more processors; a memory; one or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the method according to any one of claims 1-7.
10. A computer-readable storage medium, wherein, program code is stored in the computer-readable storage medium, and the program code can be called by a processor to execute the method according to any one of claims 1-7.
11. A computer program product, including a computer program / instructions, wherein, when the computer program / instructions are executed by a processor, the method according to any one of claims 1-7 is implemented.
Citation Information
Patent Citations
Triad acquisition method and device, electronic equipment and readable storage medium
CN111198932A
Open entity relationship extraction method based on generative adversarial network
CN111651528A