Training Method and Device for Information Extraction Model
Through targeted training and verification of information extraction models, and using text information matching target dimensions for model optimization, the problem of time-consuming and labor-consuming data labeling in the existing technology is solved, and efficient information extraction model training and knowledge graph construction are achieved.
Patent Information
- Application Number
- CN202011263099.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-12
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2040-11-12
AI Technical Summary
In the prior art, when training information extraction models, especially for complex events or multi-dimensional information extraction, a large amount of data is required, resulting in large amounts of manpower and material resources consumed and time-consuming.
By obtaining training text information matching the target dimension and verifying text information, the model is extracted using these text information and the category tags it carries, and the model is verified by verifying text information, analyzing model defects, and targeted extraction of new text information for model training again until the target model that meets the usage needs is reached.
While saving training costs, it can improve the recognition accuracy of the model in various recognition dimensions, reduce the data preparation and labeling time during knowledge graph construction, and improve the construction efficiency.
Smart Images

Figure CN114491010B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of machine learning technology, and particularly relates to a method and device for training an information extraction model. Background Art
[0002] In the prior art, the difficulty of information extraction for different events or different dimensions of the same event is different. For information extraction of some simple categories, often only a small amount of data is required to train an information extraction model with a very high accuracy. However, for information extraction of some complex events or complex dimensions of the same event, the difficulty is relatively high. And in order to enable the information extraction model to achieve a very high accuracy in information extraction of complex events, a large amount of data often needs to be labeled. In addition, labeling a large amount of data not only consumes manpower and material resources, but also takes a long time to complete. Therefore, there is an urgent need for an effective solution to solve the above problems. Summary of the Invention
[0003] In view of this, embodiments of this application provide a method for training an information extraction model to solve the technical defects existing in the prior art. Embodiments of this application also provide a device for training an information extraction model, a method for constructing a knowledge graph, a device for constructing a knowledge graph, a computing device, and a computer-readable storage medium.
[0004] According to a first aspect of the embodiments of this application, there is provided a method for training an information extraction model, including:
[0005] Obtain training text information and verification text information that match a target dimension, where the training text information and the verification text information respectively carry category labels;
[0006] Train an information extraction model according to the training text information and the category labels carried by the training text information, and use the information extraction model to process the verification text information to obtain verification category labels;
[0007] Compare the verification category labels with the category labels carried by the verification text information, and determine whether the information extraction model meets the stop training condition according to the comparison result;
[0008] If not, determine a recognition dimension to be adjusted according to the comparison result, and use the recognition dimension to be adjusted as the target dimension, and continue to train the information extraction model.
[0009] Optionally, the obtaining training text information and verification text information that match a target dimension includes:
[0010] Extract a set number of initial text information that match the target dimension from a preset text database;
[0011] Generate a set number of initial text messages carrying category labels based on the set number of the said initial text messages;
[0012] Divide the set number of initial text messages carrying category labels into the said training text messages carrying category labels and the verification text messages carrying category labels.
[0013] Optionally, the step of determining the recognition dimension to be adjusted according to the comparison result and using the recognition dimension to be adjusted as the target dimension includes:
[0014] Determine the difference category label between the verification category label and the category label carried by the verification text message according to the comparison result;
[0015] Classify the difference category label, and select the target category label according to the classification result;
[0016] Determine the recognition dimension to which the target category label belongs as the recognition dimension to be adjusted.
[0017] Optionally, the step of classifying the difference category label and selecting the target category label according to the classification result includes:
[0018] Classify the difference category label to obtain multiple category label sets;
[0019] Determine the number of labels of the category labels included in each category label set, and select the category label set with the number of labels greater than the preset number threshold to determine the target category label.
[0020] Optionally, the training text message is a training government text message, and the training government text message includes at least one of the following sub-information:
[0021] Subject name sub-information, cost date sub-information, document abstract sub-information, document issuing agency sub-information, release date sub-information, document number sub-information, document original link sub-information;
[0022] Correspondingly, the verification text sub-information is a verification government text message, and the verification government text message includes at least one of the following sub-information:
[0023] Subject name sub-information, cost date sub-information, document abstract sub-information, document issuing agency sub-information, release date sub-information, document number sub-information, document original link sub-information;
[0024] Correspondingly, the category label includes at least one of the following: name label, gender label, age label, position label, meeting name label.
[0025] Optionally, training the information extraction model according to the training text information and the category label training information carried by the training text information includes:
[0026] Converting the training text information into a first feature vector as the input of the information extraction model, and using the category label carried by the training text information as the output of the information extraction model;
[0027] Training the information extraction model based on the first feature vector and the category label carried by the training text information to obtain a verified information extraction model.
[0028] Optionally, using the information extraction model to process the verification text information to obtain a verification category label includes:
[0029] Converting the verification text information into a second feature vector, and inputting the second feature vector into the verified information extraction model for processing to obtain the verification category label corresponding to the verification text information.
[0030] Optionally, if the judgment result of judging whether the information extraction model meets the stop training condition according to the comparison result is yes, then execute the following steps:
[0031] Determining the information extraction model as the target information extraction model and storing the target information extraction model.
[0032] Optionally, after the step of determining the information extraction model as the target information extraction model and storing the target information extraction model, it further includes:
[0033] Obtaining text information matching the target domain and performing structured processing on the text information;
[0034] Inputting the structured text information into the target information extraction model for processing to obtain the category label corresponding to the text information;
[0035] Extracting a plurality of triples from the text information based on the category label corresponding to the text information, and constructing a knowledge graph matching the target domain according to the plurality of triples.
[0036] Optionally, it further includes:
[0037] Storing the knowledge graph in the form of an attribute graph into a graph database, where the graph database is configured with a call interface.
[0038] Optionally, it further includes:
[0039] Receiving query information submitted by the user for the target domain;
[0040] Determine the query entity corresponding to the query information and the query relationship corresponding to the query entity;
[0041] Based on the query entity and the query relationship, determine a target entity in the knowledge graph, and send the target as the feedback of the query information to the user.
[0042] According to the second aspect of the embodiments of the present application, there is provided a training device for an information extraction model, including:
[0043] An acquisition module, configured to acquire training text information and verification text information that match a target dimension, where the training text information and the verification text information respectively carry category labels;
[0044] A training module, configured to train an information extraction model according to the training text information and the category labels carried by the training text information, and use the information extraction model to process the verification text information to obtain verification category labels;
[0045] A comparison module, configured to compare the verification category labels with the category labels carried by the verification text information, and determine whether the information extraction model meets the stop training condition according to the comparison result;
[0046] If not, run a determination module, where the determination module is configured to determine a recognition dimension to be adjusted according to the comparison result, and use the recognition dimension to be adjusted as the target dimension to continue training the information extraction model.
[0047] According to the third aspect of the embodiments of the present application, there is provided a method for constructing a knowledge graph, including:
[0048] Acquire text information that matches a target domain, and perform structured processing on the text information;
[0049] Input the structured text information into a target information extraction model that meets the training stop condition for processing to obtain category labels corresponding to the text information;
[0050] Extract a plurality of triples from the text information based on the category labels corresponding to the text information, and construct a knowledge graph that matches the target domain according to the plurality of triples.
[0051] Optionally, after the step of constructing a knowledge graph that matches the target domain according to the plurality of triples, it further includes:
[0052] Store the knowledge graph in the form of an attribute graph in a graph database, where the graph database is configured with a call interface.
[0053] Optionally, after the step of constructing a knowledge graph that matches the target domain according to the multiple triples, the method further includes:
[0054] Receiving query information submitted by a user for the target domain;
[0055] Determining a query entity corresponding to the query information and a query relationship corresponding to the query entity;
[0056] Determining a target entity in the knowledge graph based on the query entity and the query relationship, and sending the target as a feedback of the query information to the user.
[0057] According to a fourth aspect of an embodiment of the present application, there is provided a knowledge graph construction device, including:
[0058] A text information acquisition module, configured to acquire text information that matches the target domain and perform structured processing on the text information;
[0059] A model processing module, configured to input the structured text information into a target information extraction model that meets the training stop condition for processing to obtain category labels corresponding to the text information;
[0060] A graph construction module, configured to extract a plurality of triples from the text information based on the category labels corresponding to the text information, and construct a knowledge graph that matches the target domain according to the plurality of triples.
[0061] According to a fifth aspect of an embodiment of the present application, there is provided a computing device, including:
[0062] A memory and a processor;
[0063] The memory is used to store computer-executable instructions, and when the processor executes the computer-executable instructions, the steps of the method are implemented.
[0064] According to a sixth aspect of an embodiment of the present application, there is provided a computer-readable storage medium, which stores computer-executable instructions, and when the instructions are executed by a processor, the steps of the method are implemented.
[0065] The training method of the information extraction model provided by this application, after obtaining the training text information and verification text information that match the target dimension, trains the information extraction model by using the training text information and its carried category label, and then verifies the information extraction model through the category label carried by the verification text information, so as to analyze the defects existing in the current information extraction model. Then, new text information is extracted specifically for the model to be retrained until the target information extraction model that meets the usage requirements is stored. This realizes targeted training of the model, which can not only save the cost of training the model, but also improve the recognition accuracy of the model in each recognition dimension, so as to meet the subsequent construction of the knowledge graph with low consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 is a flowchart of a training method of an information extraction model provided by an embodiment of this application;
[0067] Figure 2 is a structural schematic diagram of a training method of an information extraction model provided by an embodiment of this application;
[0068] Figure 3 is a structural schematic diagram of a training device of an information extraction model provided by an embodiment of this application;
[0069] Figure 4 is a flowchart of a knowledge graph construction method provided by an embodiment of this application;
[0070] Figure 5 is a structural schematic diagram of a knowledge graph construction device provided by an embodiment of this application;
[0071] Figure 6 is a structural block diagram of a computing device provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0072] In the following description, many specific details are set forth in order to provide a thorough understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of this application. Therefore, this application is not limited by the specific implementations disclosed below.
[0073] The terms used in one or more embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present application. The singular forms "a", "the", and "said" used in one or more embodiments of the present application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present application refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0074] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present application, the first may also be referred to as the second, and similarly, the second may also be referred to as the first.
[0075] First, the noun terms related to one or more embodiments of the present invention are explained.
[0076] Knowledge graph: A knowledge graph is a knowledge base used to enhance the function of its search engine. Essentially, a knowledge graph aims to describe various entities or concepts existing in the real world and their relationships, which constitutes a huge semantic network graph, with nodes representing entities or concepts and edges consisting of attributes or relationships.
[0077] Graph database: A database that uses a graph structure for semantic queries, and uses nodes, edges, and attributes to represent and store data.
[0078] Training text information: Refers to the text information used when training an information extraction model; correspondingly, the category labels carried by the training text information specifically refer to the labels corresponding to each character and word in the text information.
[0079] Verification text information: Refers to the text information for verifying the recognition accuracy of an information extraction model. Correspondingly, the category labels carried by the verification text information specifically refer to the labels corresponding to each character and word in the text information.
[0080] Verified category label: Refers to the category label identified after an information extraction model recognizes the characters and words included in the verification text information.
[0081] Training stop condition: Refers to the condition for determining whether an information extraction model meets the usage requirements; if it meets, stop training the information extraction model for subsequent use; if it does not meet, continue to train the information extraction model until it meets the usage requirements and then stop training.
[0082] Recognition dimension to be adjusted: Refers to the dimension in which an information extraction model still has inaccurate recognition.
[0083] In the present application, a training method for an information extraction model is provided. The present application also relates to a training apparatus for an information extraction model, a method for constructing a knowledge graph, a device for constructing a knowledge graph, a computing device, and a computer-readable storage medium, which will be described in detail one by one in the following embodiments.
[0084] Figure 1 The flowchart of a training method for an information extraction model provided according to an embodiment of the present application is shown, which specifically includes the following steps:
[0085] Step S102: Obtain training text information and verification text information that match the target dimension, and category labels are carried in the training text information and the verification text information respectively.
[0086] In practical applications, when constructing a knowledge graph in a specific domain, due to the different characteristics of different domains, targeted construction is required. After most knowledge graphs are constructed, a large amount of data in the domain needs to be integrated to be realized. In the data preparation stage, if it is manually labeled, the cost will be very high and the time will be long, which greatly affects the efficiency of graph construction.
[0087] For the training method of the information extraction model provided by the present application, in order to improve the efficiency of the data preparation stage and the efficiency of constructing the graph, after obtaining the training text information and verification text information that match the target dimension, the information extraction model is trained by using the training text information and the category labels carried therein, and then the information extraction model is verified by the category labels carried in the verification text information, so as to analyze the defects existing in the current information extraction model. Then, new text information is extracted specifically for the defect for re-training the model until the target information extraction model that meets the usage requirements is obtained and stored. It realizes targeted training of the model, which can not only save the cost of training the model, but also improve the recognition accuracy of the model in each recognition dimension, so as to meet the subsequent construction of the knowledge graph with low consumption.
[0088] Specifically, the training method of the information extraction model provided by the present application is to quickly label the data through the trained information extraction model before constructing the knowledge graph, so as to improve the construction efficiency of the graph. In order to quickly complete the data labeling, it is necessary to ensure the labeling accuracy while meeting the labeling requirements, that is, training a model that meets the labeling requirements will also cost a lot of manpower and material resources. Therefore, the method provided by the present application aims to solve the training method of the information extraction model, reduce the training cost of the information extraction model while ensuring the recognition accuracy of the information extraction model, so as to better meet the subsequent usage requirements.
[0089] Further, this application will take the application of the information extraction model in the government affairs field as an example to describe the training method of the information extraction model. It should be noted that other fields such as the news field, the adjudication field, or the conference field can refer to the corresponding description content of this embodiment, and this embodiment will not be elaborated here.
[0090] Furthermore, in the case where the training method of the information extraction model is applied to the government affairs field, the training text information is training government affairs text information, and the training government affairs text information includes at least one of the following sub-information: subject name sub-information, cost date sub-information, document summary sub-information, document issuing agency sub-information, release date sub-information, document number sub-information, document original link sub-information; correspondingly, the verification text sub-information is verification government affairs text information, and the verification government affairs text information includes at least one of the following sub-information: subject name sub-information, cost date sub-information, document summary sub-information, document issuing agency sub-information, release date sub-information, document number sub-information, document original link sub-information; correspondingly, the category labels include at least one of the following: name label, gender label, age label, position label, meeting name label.
[0091] Specifically, in the process of constructing the knowledge graph applied in the government affairs field, since constructing the graph requires n triples to be realized, it is necessary to perform entity annotation and relationship annotation on a large amount of data involved in the government affairs field to complete. Before that, in order to meet the annotation requirements for this part of the large amount of data, it is necessary to separately construct an information extraction model, and the process of model construction is time-consuming and laborious. Therefore, in order to obtain an information extraction model that meets the usage requirements in a relatively short time, an active learning method can be used for training, that is, targeted training for the defects existing in the training model in a progressive manner, so as to obtain an information extraction model that meets the usage requirements.
[0092] Among them, the training government affairs text information specifically refers to the text information required to train the model in the government affairs field, and the labels corresponding to each word and term in this text information have been marked. Correspondingly, the verification government affairs text information specifically refers to the text information required to verify the information extraction model in the government affairs field, and the labels corresponding to each word and term in this text information have been marked; further, the subject name sub-information, cost date sub-information, document summary sub-information, document issuing agency sub-information, release date sub-information, document number sub-information, and / or document original link sub-information refer to the sub-information included in the text information, while the name label, gender label, age label, position label, and / or meeting name label refer to the category labels that each word or term in the text information should correspond to. In addition, the category labels can also include time labels, space labels, etc., and this embodiment will not make too many limitations here.
[0093] During the training process of the information extraction model, in order to improve the model training efficiency, training text information with category labels and validation text information with category labels will be obtained according to the target dimension. In this embodiment, the specific implementation method is as follows:
[0094] Extract a set number of initial text information that matches the target dimension from the preset text database;
[0095] Based on the set number of the initial text information, generate a set number of initial text information with category labels;
[0096] Divide the set number of initial text information with category labels into the training text information with category labels and the validation text information with category labels.
[0097] Specifically, the preset text database specifically refers to a database storing a large amount of unlabeled text information (the text information is in the same field as the knowledge graph to be constructed). The target dimension specifically refers to the dimension corresponding to the training direction required when training the information extraction model. For example, if it is required to complete the annotation of the time in the text information through the information extraction model with category labels, the target dimension is the time dimension. Or if it is required to complete the annotation of the name and gender in the text information through the information extraction model with category labels, the target dimension is the name dimension and the gender dimension. Correspondingly, the initial text information is unlabeled text information, and this text information is a small amount (set number) for preliminary training of the model. That is, each time a small number of samples are selected according to the target dimension to train the information extraction model, so as to achieve the purpose of reducing costs, and analyze the defects existing in the current model, and then gradually improve it to train an information extraction model that meets the usage requirements.
[0098] Based on this, first extract a set number of initial text information that matches the target dimension from the preset text data, secondly annotate the set number of initial text information to obtain the initial text information with category labels; finally divide the initial text information with category labels into two parts, one part is used as training text information to train the information extraction model, and the other part is used as validation text information to assist in analyzing the defects existing in the information extraction model, which is convenient for subsequent targeted extraction of new training samples to continue training the model.
[0099] For example, there are 10 million government text messages without category labels in the government text database. At this time, 5,000 government text messages can be extracted from the government text database for annotation, that is, to annotate the category labels corresponding to each word or character in each government text message. For example, if the text message is "Xiaoming created the song 'In the Spring'", first standardize this text message, that is, delete the punctuation marks in the text message to get the standard text message "Xiaoming created the song In the Spring". Secondly, perform word segmentation on this standard text message to obtain multiple word units {Xiaoming, created, the, song, In the Spring}. Finally, annotate each word unit to determine that the category label corresponding to "Xiaoming" is "author", the category label corresponding to "In the Spring" is "work", etc. At the same time, meaningless words such as "created", "the", and "the" are annotated as "0".
[0100] Further, after annotating 5,000 government text messages, 4,000 government text messages with category labels will be selected to form sample text messages, and the remaining 1,000 government text messages with category labels will be formed into verification text messages for subsequent training of the information extraction model applied to the government field, so as to be able to construct a government knowledge graph that meets the usage requirements.
[0101] It should be noted that when training the information extraction model for the first time, the target dimension will include all training sub-dimensions that meet the usage requirements, so that the information extraction model can learn which dimensions of category labels need to be recognized, and then improve each training sub-dimension subsequently; in addition, the set quantity can be the proportion quantity of the database, such as one percent or two percent of the text messages in the database. This embodiment does not make any limitation here.
[0102] In summary, by selecting the initial text messages with a set quantity and matching the target dimension to divide the training text messages and verification text messages, the model can be gradually trained with fewer samples, effectively reducing the consumption cost of model training.
[0103] Step S104, train the information extraction model according to the training text messages and the category labels carried by the training text messages, and use the information extraction model to process the verification text messages to obtain verification category labels.
[0104] Specifically, after obtaining the verification text messages and training text messages with category labels, the information extraction model needs to be trained at this time. During the training process, since it is necessary to monitor the recognition accuracy of the current information extraction model, the prepared verification text messages can be used to verify the current information extraction model, so as to be used to judge the existing defects of the information extraction model subsequently, and then select new sample data to continue training the model targeted.
[0105] Furthermore, since the information extraction model needs to preprocess the data before identifying the category labels and can only use the preprocessed data as the input of the model, it is necessary to convert the training text information into feature vectors and then identify it through the model. In this embodiment, the specific implementation method is as follows:
[0106] Convert the training text information into a first feature vector as the input of the information extraction model, and use the category label carried by the training text information as the output of the information extraction model;
[0107] Train the information extraction model based on the first feature vector and the category label carried by the training text information to obtain a verified information extraction model;
[0108] Convert the verification text information into a second feature vector and input the second feature vector into the verified information extraction model for processing to obtain the verification category label corresponding to the verification text information.
[0109] Specifically, the first feature vector specifically refers to the one composed of each text information in the training text information after being converted into feature vectors, and the second feature vector specifically refers to the one composed of each text information in the verification text information after being converted into feature vectors. Correspondingly, the verified information extraction model specifically refers to the model obtained after the initial training of the training text information, and the verification category label specifically refers to the category label corresponding to the verification text information obtained after using the verified information extraction model to identify the labels of the verification text information.
[0110] Based on this, after obtaining the verification text information and training text information with category labels, in order to meet the requirements of training and verification, first convert the training text information into a first feature vector as the input of the information extraction model, and at the same time use the category label carried by the training text information as the output of the information extraction model. Secondly, use the first feature vector and the category label carried by the training text information to train the information extraction model to obtain a verified information extraction model that has been initially trained at the current stage. Finally, use the verification text information converted into a second feature vector as the text information for verifying the accuracy of the model. Process the second feature vector through the verified information extraction model to obtain the verification category label corresponding to the verification text information, which is used to verify the accuracy of the information extraction model at the current stage, thus facilitating subsequent targeted training.
[0111] Continuing with the above example, on the basis of determining the sample text information composed of 4,000 government affair text messages carrying category labels and the verification text information composed of 1,000 government affair text messages carrying category labels, further, at this time, the 4,000 government affair text messages (sample text information) are respectively converted into first sub-feature vectors as the input for training the information extraction model. At the same time, the category labels corresponding to each first feature sub-vector are used as the output of the model, and 4,000 sample pairs are formed to train the information extraction model. After the training is completed, a verification information extraction model is obtained; then the 1,000 government affair text messages (verification text information) are converted into second feature sub-vectors and respectively input into the verification information extraction model for recognition to obtain the verification category labels corresponding to each verification text information, which are used for subsequent comparison with the correct category labels corresponding to the 1,000 government affair text messages (verification text information), so as to analyze the existing defects of the information extraction model and use the new government affair text messages to effectively train it.
[0112] In summary, by using a part of the data to train the model and a part of the data to verify the model, the existing defects of the model can be accurately analyzed, which is convenient for targeted extraction of new data to continue training the model, effectively avoiding the problem of data waste, and thus saving the cost of model training.
[0113] Step S106: Compare the verification category label with the category label carried by the verification text information, and determine whether the information extraction model meets the stop training condition according to the comparison result;
[0114] If not, execute step S108; if so, determine the information extraction model as the target information extraction model and store the target information extraction model.
[0115] Specifically, after training the information extraction model with the training text information above, the verification category label is obtained by identifying the category label of the verification text information through the information extraction model. At this time, the verification category label can be compared with the category label carried by the verification text information, so as to determine whether the current information extraction model meets the stop training condition according to the comparison result. Among them, the stop training condition specifically refers to judging whether the information extraction model in the current stage meets the usage requirements, that is, whether the recognition accuracy in each dimension meets the standards required for subsequent construction of the knowledge graph. If not, it means that the model still needs to continue training, and then execute the subsequent step S108; if so, it means that the model has met the usage requirements, and then store the information extraction model.
[0116] In practical applications, the stop training condition can be set according to the actual application scenario. For example, the stop training condition can be set to determine whether the recognition accuracy of the information extraction model in the current stage reaches the accuracy threshold, or to determine whether the number of correctly recognized texts reaches the preset number threshold after the information extraction model in the current stage recognizes a set number of verification text information. It can be set according to the actual application scenario in specific applications, and this embodiment does not make any limitation here.
[0117] Continuing with the above example, after using the verification information extraction model to recognize 1000 government affair text information (verification text information), the verification category label corresponding to each government affair text information is obtained. At this time, the category label carried by each government affair text information is compared with each recognized verification text information to determine that the category label correct rate of the current information extraction model for the 1000 government affair text information after recognition is s. If s is greater than the preset correct rate threshold p, it means that the recognition accuracy of the current information extraction model has met the requirements for subsequent assistance in constructing the knowledge graph, and it can be stored for labeling the category labels of the data for constructing the knowledge graph; if s is less than or equal to the preset correct rate threshold p, it means that the recognition accuracy of the current information extraction model does not meet the requirements for subsequent assistance in constructing the knowledge graph, and it needs to be trained continuously to quickly train a model that meets the usage requirements. When training the information extraction model again, it needs to be trained targeted to improve the recognition richness of the model.
[0118] Alternatively, it is determined that the number of correctly recognized category labels of the current information extraction model for the 1000 government affair text information is n. If n is greater than the preset number threshold m, it means that the recognition accuracy of the current information extraction model has met the requirements for subsequent assistance in constructing the knowledge graph, and it can be stored for labeling the category labels of the data for constructing the knowledge graph; if n is less than or equal to the preset number threshold m, it means that the recognition accuracy of the current information extraction model does not meet the requirements for subsequent assistance in constructing the knowledge graph, and it needs to be trained continuously to quickly train a model that meets the usage requirements.
[0119] Step S108, determine the recognition dimension to be adjusted according to the comparison result, and use the recognition dimension to be adjusted as the target dimension to continue training the information extraction model.
[0120] Specifically, based on the above judgment of whether the trained information extraction model meets the stop training condition according to the comparison result, further, at this time, it is analyzed according to the comparison result that the trained information extraction model does not meet the stop training condition, indicating that the model still needs to be further trained. In order to save training costs and improve the recognition accuracy of the model, the dimension to be adjusted for recognition can be determined according to the comparison result between the verification category label and the category label carried by the verification text information. The dimension to be adjusted for recognition specifically refers to the recognition defect existing in the information extraction model at the current stage. For example, if the recognition accuracy of the time category label is not high, then the dimension to be adjusted for recognition is determined as the time dimension; finally, the dimension to be adjusted for recognition is used as the target dimension, and the information extraction model is continued to be trained, that is, return to execute step S102, so that it can be realized to selectively choose new samples from the samples to continue training the information extraction model. That is, after determining the defect of the information extraction model, select the sample text information and verification text information corresponding to the defect from the samples to continue training and verifying the model until a model that meets the requirements is obtained and then store it.
[0121] In practical applications, during the process of continuing to train the information extraction model, since a new round of model training process needs to be implemented according to the dimension to be recognized, in order to improve the model training efficiency, the iterative training process of the information extraction model can be completed according to a preset training strategy, that is: after using the dimension to be adjusted for recognition as the target dimension, it is necessary to match new training text information and verification text information according to the target dimension. At this time, the quantity of new training text information and new verification text information obtained can be reduced (since the recognition ability of the model in the target dimension has been trained in the previous round of training, making it have a certain prediction ability, so in the new round of training process, the sample quantity of the target dimension can be selected to be reduced, thereby reducing cost consumption), and the information extraction model is trained again.
[0122] Further, in the case that the target dimension includes multiple target sub - dimensions, it indicates that in the new round of training process of the information extraction model, the prediction ability for multiple dimensions still needs to be strengthened, and the information extraction model may have different prediction accuracies in different target sub - dimensions. Therefore, in the new round of training process, the sample data (training text information and verification text information) can be obtained according to the priority of the prediction accuracy of the information extraction model in the target sub - dimensions, that is: select more sample data for the target sub - dimension with lower prediction accuracy, and select less sample data for the target sub - dimension with higher prediction accuracy, so as to realize the dynamic adjustment of the quantity of sample data obtained, save resource consumption and improve the training efficiency of the model.
[0123] For example, based on the comparison results, it is determined that the recognition accuracy of the information extraction model in the current stage is 75% in the time dimension and 60% in the name dimension. It is determined that the recognition accuracy of the information extraction model in the current stage for both time and name is not high, and the recognition accuracy in the name dimension is relatively poor. Therefore, in the new round of training process, 3,000 training text information and 500 validation text information that match the time dimension can be selected, and 4,500 training text information and 800 validation text information that match the name dimension can be selected to continue training the information extraction model.
[0124] In addition, in the new round of training process, there may also be a situation where target sub-dimensions intersect, that is, when obtaining training text information and validation text information that match each target sub-dimension, duplicate training text information and validation text information may be selected. For example, training text information U can be used not only for training in the time dimension but also for training in the name dimension. At this time, in order to reduce resource consumption, after obtaining the training text information and validation text information that match the target sub-dimension, the training text information and validation text information can be de-duplicated to obtain non-duplicate training text information and validation text information, and then the information extraction model can be trained, thereby reducing the time consumed in training the information extraction model and improving the model training efficiency.
[0125] Furthermore, in the process of determining the recognition dimension to be adjusted based on the comparison results, in order to be able to specifically select the dimension with larger defects, the target dimension can be determined by selecting the label with the lowest recognition accuracy according to the classification results. In this embodiment, the specific implementation method is as follows:
[0126] Determine the difference category label between the verification category label and the category label carried by the verification text information according to the comparison results;
[0127] Perform a classification process on the difference category label to obtain multiple category label sets;
[0128] Determine the number of labels of the category labels included in each category label set, and select the category label set with the number of labels greater than the preset number threshold to determine the target category label;
[0129] Determine the recognition dimension to which the target category label belongs as the recognition dimension to be adjusted.
[0130] Specifically, the difference category label refers to the category label that is inaccurately recognized by the information extraction model after identifying the verification text information; the target category label specifically refers to the type label with a relatively large number among the difference category labels; based on this, after determining the difference category label according to the comparison result, the difference category label will be classified at this time to obtain multiple category label sets belonging to the same category, and then the number of category labels included in each category label set will be determined. Finally, the category label with the label number greater than the preset threshold is selected to determine the dimension to which the target category label belongs as the dimension to be adjusted for recognition, which is used for subsequent extraction of new samples and training of the information extraction model.
[0131] Continuing with the above example, after the verification information extraction model identifies 1000 government affair text information (verification text information), the verification category label is obtained, and the verification category label is compared with the category label carried by the government affair text information. It is determined that there are 200 position labels, meeting name labels, and name labels in the recognized labels, but the correct recognition quantity is only 50, and the correct rate is relatively low. This shows that the information extraction model has low recognition accuracy for name vocabulary, name vocabulary, and position vocabulary in text information, while the recognition accuracy of gender labels and date labels has met the recognition requirements. Then, it is necessary to improve the recognition accuracy of the position label, meeting name label, and name label. Therefore, the position label, meeting name label, and name label are determined as the target category labels, and the dimension to which each label belongs is determined as the target dimension. When extracting a set number of government affair text information from the database next time, the government affair text information with a relatively large proportion in these three dimensions is selected for retraining the model.
[0132] In summary, in the process of determining the dimension to be adjusted for recognition, in order to be able to train the dimension with relatively large mistakes of the information extraction model in a targeted manner, the category label for which the recognition accuracy needs to be improved can be determined according to the number of incorrectly recognized category labels, and then the dimension to which the label belongs is determined as the dimension to be adjusted for recognition to meet the requirement of extracting new samples.
[0133] When it is determined according to the comparison result that the information extraction model meets the stop training condition, the information extraction model can be determined as the target information extraction model and the target information extraction model is stored.
[0134] Specifically, based on the above judgment of whether the trained information extraction model meets the stopping training condition according to the comparison result, further, at this time, it is determined that the trained information extraction model meets the stopping training condition according to the comparison result, indicating that the current information extraction model can already be applied to the process of constructing the knowledge graph. Then, the model can be stored as the target information extraction model. It should be noted that since the target information extraction model is trained for the field of knowledge graph construction during training, it can only be applied to the same or similar knowledge graph fields, thereby effectively improving the recognition accuracy of the model and reducing the influence of other factors on the model.
[0135] Based on this, refer to Figure 2 As shown, after selecting a set number of text information from the sample set, at this time, the set number of text information will be labeled (adding category labels to the word segmentation results in the text information), and then the labeled text information will be divided into a training set and a validation set. Among them, the text information in the training set is the training text information with category labels, and the text information in the validation set is the validation text information with category labels. Then, the training set is used to train the information extraction model (information extraction model). After the initial training is completed, the text information in the validation set is used to verify it, that is, the text information after the initial training is used to identify the labels of the text information in the validation set.
[0136] Further, compare the category labels marked in the validation set with the labels identified by the information extraction model, and judge whether the model meets the stopping training condition according to the comparison result. If so, it means that the trained information extraction model meets the downstream business use, and it can be applied to the downstream business processing; if not, the trained information extraction model cannot meet the downstream business use, then the main defects of the model can be determined according to the comparison result (that is, determine the identification dimension to be adjusted), and then new samples are screened from the sample set according to the analyzed defects to continue training the model. Until a model that meets the usage requirements is obtained and applied to subsequent processing.
[0137] Further, after the model training is completed, it can be applied to the construction of the knowledge graph. In the process of constructing the knowledge graph, all the text information involved needs to be labeled by the information extraction model before it can be applied. In this embodiment, the specific implementation method is as follows:
[0138] Obtain text information matching the target field and perform structured processing on the text information;
[0139] Input the structured text information into the target information extraction model for processing to obtain the category labels corresponding to the text information;
[0140] Extract multiple triples from the text information based on the category labels corresponding to the text information, and construct a knowledge graph matching the target domain according to the multiple triples.
[0141] Specifically, the target domain specifically refers to the domain where the knowledge graph needs to be applied, and the text information matching the target domain specifically refers to all the text information involved in the target domain, which is used to construct the knowledge graph. Correspondingly, the triple is the basic unit for constructing the knowledge graph, and each triple is constructed by an entity and an attribute, such as (Xiaoming - age - 56 years old), (Xiaoya - gender - female), (Xiaoming - position - xxx committee member), and (XXX Representative Assembly - meeting name - First Plenary Session), etc. It should be noted that the structured processing of the text information specifically refers to preprocessing it to obtain data that meets the usage requirements, which is convenient for subsequent construction of the knowledge graph.
[0142] Continuing with the above example, after obtaining the target information extraction model that meets the usage requirements in the government affairs domain through the above method, a large amount of text information related to the government affairs domain is selected at this time, and the text information is structured. Then, the target information extraction model is used to identify the category labels corresponding to the words or terms in the large amount of structured text information, so as to obtain the triples extracted from each text information, thus obtaining a large number of triples. Then, a government affairs knowledge graph matching the government affairs domain is constructed based on the large number of triples; through this government affairs knowledge graph, the need to query relevant government affairs information can be realized.
[0143] Furthermore, after completing the construction of the knowledge graph, for the convenience of subsequent use, the knowledge graph can be stored in a graph database in the form of an attribute graph, where the graph database is configured with a call interface; that is, after storing the knowledge graph, it can be called and used through the call interface during subsequent applications, so as to meet the usage requirements of different scenarios; and, during the process of storing in the graph database in the form of an attribute graph, it can be stored based on the Resource Description Framework (RDF), or it can be stored based on the graph database. Among them, the graph database focuses on efficient graph querying and searching. Generally, the graph database uses the attribute graph as the basic representation form, and entities and relationships can contain attributes, which means that it is easier to express the real scenarios in reality. Specifically, the graph database can be the Neo4j (a high-performance NOSQL graph database) graph database.
[0144] In summary, using the information extraction model obtained based on active learning to label government affairs data and extract triples, and constructing a knowledge graph about government affairs data through the triples, it realizes the use of a small amount of data to label government affairs data and construct a knowledge graph, reducing the time cost and capital cost.
[0145] In addition, after the construction of the knowledge graph is completed, it can be applied to queries in the target field. When the query requirements of the user are obtained, the answers that meet the requirements can be quickly queried. In this embodiment, the specific implementation method is as follows:
[0146] Receive the query information submitted by the user for the target field;
[0147] Determine the query entity corresponding to the query information and the query relationship corresponding to the query entity;
[0148] Based on the query entity and the query relationship, determine the target entity in the knowledge graph, and send the target as the feedback of the query information to the user.
[0149] Specifically, the user specifically refers to a user with query requirements. The query information specifically refers to the question submitted by the user. The query entity specifically refers to the entity in the query information. The query relationship specifically refers to the relationship corresponding to the query entity. The target entity specifically refers to the answer corresponding to the query information.
[0150] Based on this, when the query information submitted by the user for the target field is received, it means that the user needs to query relevant information. At this time, extract the query entity in the query information, and determine the query relationship corresponding to the query entity. Then, form a query statement based on the query entity and the query relationship to determine the target entity in the knowledge graph. Finally, send the target entity as the feedback of the query information to the user.
[0151] For example, the constructed knowledge graph is a government affairs knowledge graph. At this time, the query information submitted by the user is "Who is the person in charge of environmental governance in Area A?" At this time, the query entity is determined to be "Area A", and the query relationship is "environmental person in charge"; then, based on the query entity "Area A" and the query relationship "environmental person in charge", the target entity in the government affairs knowledge graph is determined to be "Person in Charge B". Then, it is determined that the answer to this question is "Person in Charge B". At this time, just feedback "Person in Charge B" to the user so that the user can know the answer to this question.
[0152] In summary, by using less data to train an information extraction model that meets the usage requirements, then using this model to construct a knowledge graph, and finally providing the graph for the user to use, not only can the time for constructing the knowledge graph be reduced, but also when there is a need to construct a graph, a knowledge graph that meets the usage requirements can be provided in a shorter time, further improving the user experience.
[0153] The training method of the information extraction model provided by this application, after obtaining the training text information and verification text information that match the target dimension, trains the information extraction model by using the training text information and the category labels carried by it, and then verifies the information extraction model through the category labels carried by the verification text information, so as to analyze the defects existing in the current information extraction model. After that, new text information is extracted specifically for this defect for retraining the model until the target information extraction model that meets the usage requirements is obtained and stored. This realizes targeted training of the model, which can not only save the cost of training the model, but also improve the recognition accuracy of the model in each recognition dimension, so as to meet the subsequent construction of the knowledge graph with low consumption.
[0154] Corresponding to the above method embodiment, this application also provides an embodiment of a training device for an information extraction model. Figure 3 The structural schematic diagram of a training device for an information extraction model provided by an embodiment of this application is shown. As Figure 3 shown, the device includes:
[0155] An acquisition module 302, configured to acquire training text information and verification text information that match the target dimension, and category labels are respectively carried in the training text information and the verification text information;
[0156] A training module 304, configured to train an information extraction model according to the training text information and the category labels carried by the training text information, and use the information extraction model to process the verification text information to obtain verification category labels;
[0157] A comparison module 306, configured to compare the verification category labels with the category labels carried by the verification text information, and determine whether the information extraction model meets the stop training condition according to the comparison result;
[0158] If not, a determination module 308 is run. The determination module 308 is configured to determine the recognition dimension to be adjusted according to the comparison result, and use the recognition dimension to be adjusted as the target dimension to continue training the information extraction model.
[0159] In an optional embodiment, the acquisition module 302 is further configured to:
[0160] Extract a set number of initial text information that matches the target dimension from a preset text database; based on the set number of the initial text information, generate a set number of initial text information carrying category labels; divide the set number of initial text information carrying category labels into the training text information carrying category labels and the verification text information carrying category labels.
[0161] In an alternative embodiment, the determining module 308 is further configured to:
[0162] Determine a difference category label between the verification category label and the category label carried by the verification text information according to the comparison result; classify the difference category label, and select a target category label according to the classification result; determine the identification dimension to which the target category label belongs as the identification dimension to be adjusted.
[0163] In an alternative embodiment, the determining module 308 is further configured to:
[0164] Classify the difference category label to obtain a plurality of category label sets; determine the number of labels of the category labels included in each category label set, and select the category label set with the number of labels greater than a preset number threshold to determine the target category label.
[0165] In an alternative embodiment, the training text information is training government affairs text information, and the training government affairs text information includes at least one of the following sub-information:
[0166] Subject name sub-information, cost date sub-information, document abstract sub-information, document issuing agency sub-information, release date sub-information, document number sub-information, document original link sub-information;
[0167] Correspondingly, the verification text sub-information is verification government affairs text information, and the verification government affairs text information includes at least one of the following sub-information:
[0168] Subject name sub-information, cost date sub-information, document abstract sub-information, document issuing agency sub-information, release date sub-information, document number sub-information, document original link sub-information;
[0169] Correspondingly, the category label includes at least one of the following: name label, gender label, age label, position label, meeting name label.
[0170] In an alternative embodiment, the training module 304 is further configured to:
[0171] Convert the training text information into a first feature vector as the input of the information extraction model, and use the category label carried by the training text information as the output of the information extraction model; train the information extraction model based on the first feature vector and the category label carried by the training text information to obtain a verification information extraction model.
[0172] In an alternative embodiment, the training module 304 is further configured to:
[0173] Convert the verification text information into a second feature vector, and input the second feature vector into the verification information extraction model for processing to obtain the verification category label corresponding to the verification text information.
[0174] In an optional embodiment, the training device of the information extraction model further includes:
[0175] A storage module, configured to determine the information extraction model as the target information extraction model and store the target information extraction model.
[0176] In an optional embodiment, the training device of the information extraction model further includes:
[0177] A text information acquisition module, configured to acquire text information matching the target field and perform structured processing on the text information;
[0178] A model processing module, configured to input the structured text information into the target information extraction model for processing to obtain the category label corresponding to the text information;
[0179] A knowledge graph construction module, configured to extract a plurality of triples from the text information based on the category label corresponding to the text information and construct a knowledge graph matching the target field according to the plurality of triples.
[0180] In an optional embodiment, the training device of the information extraction model further includes:
[0181] A knowledge graph storage module, configured to store the knowledge graph in the form of an attribute graph in a graph database, where the graph database is configured with a call interface.
[0182] In an optional embodiment, the training device of the information extraction model further includes:
[0183] A query information receiving module, configured to receive query information submitted by a user for the target field;
[0184] A query entity determination and recognition module, configured to determine a query entity corresponding to the query information and a query relationship corresponding to the query entity;
[0185] A feedback module, configured to determine a target entity in the knowledge graph based on the query entity and the query relationship and send the target as feedback of the query information to the user.
[0186] The training device for the information extraction model provided in this embodiment, after obtaining the training text information and verification text information that match the target dimension, trains the information extraction model by using the training text information and the category labels carried by it, and then verifies the information extraction model through the category labels carried by the verification text information, so as to analyze the defects existing in the current information extraction model. After that, new text information is extracted specifically for the model to be trained again until the target information extraction model that meets the usage requirements is obtained and stored. This realizes targeted training of the model, which can not only save the cost of training the model, but also improve the recognition accuracy of the model in each recognition dimension, so as to meet the subsequent construction of the knowledge graph with low consumption.
[0187] The above is a schematic solution of a training device for an information extraction model according to this embodiment. It should be noted that the technical solution of the training device for the information extraction model and the technical solution of the above information extraction method belong to the same concept. For the details not described in detail in the technical solution of the training device for the information extraction model, reference can be made to the description of the technical solution of the above information extraction method. In addition, each component in the device embodiment should be understood as a functional module that must be established to implement each step of the program flow or each step of the method. Each functional module is not an actual functional division or separation limitation. The device claim defined by such a group of functional modules should be understood as a functional module framework that mainly realizes the solution through the computer program recorded in the specification, rather than an entity device that mainly realizes the solution through hardware means.
[0188] Figure 4 The flowchart of a knowledge graph construction method according to an embodiment of the present application is shown, which specifically includes the following steps:
[0189] Step S402, obtain text information that matches the target domain, and perform structured processing on the text information.
[0190] Step S404, input the structured text information into a target information extraction model that meets the training stop condition for processing, and obtain the category label corresponding to the text information.
[0191] Step S406, extract multiple triples from the text information based on the category label corresponding to the text information, and construct a knowledge graph that matches the target domain according to the multiple triples.
[0192] In one or more embodiments of this embodiment, it further includes:
[0193] Store the knowledge graph in the form of an attribute graph in a graph database, where the graph database is configured with a call interface.
[0194] In one or more embodiments of this embodiment, it further includes:
[0195] Receiving query information submitted by the user for the target field;
[0196] Determining the query entity corresponding to the query information and the query relationship corresponding to the query entity;
[0197] Based on the query entity and the query relationship, determining a target entity in the knowledge graph and sending the target as a feedback of the query information to the user.
[0198] In summary, by training an information extraction model that meets the usage requirements with less data, then using this model to construct a knowledge graph, and finally providing the graph for the user to use, not only can the time for constructing the knowledge graph be reduced, but also when there is a need to construct a graph, a knowledge graph that meets the usage requirements can be provided in a relatively short time, further improving the user experience.
[0199] Corresponding to the above method embodiment, the present application also provides an embodiment of a knowledge graph construction device. Figure 5 It shows a schematic structural diagram of a knowledge graph construction device provided by an embodiment of the present application. As Figure 5 shown, the device includes:
[0200] An acquisition text information module 502, configured to acquire text information matching the target field and perform structured processing on the text information;
[0201] A model processing module 504, configured to input the structured text information into a target information extraction model that meets the training stop condition for processing to obtain the category label corresponding to the text information;
[0202] A graph construction module 506, configured to extract a plurality of triples from the text information based on the category label corresponding to the text information and construct a knowledge graph matching the target field according to the plurality of triples.
[0203] In an optional embodiment, the knowledge graph construction device further includes:
[0204] A graph storage module, configured to store the knowledge graph in the form of an attribute graph in a graph database, where the graph database is configured with a call interface.
[0205] In an optional embodiment, the knowledge graph construction device further includes:
[0206] A receiving information module, configured to receive query information submitted by the user for the target field;
[0207] An entity determination module, configured to determine a query entity corresponding to the query information and a query relationship corresponding to the query entity;
[0208] A feedback entity module, configured to determine a target entity in the knowledge graph based on the query entity and the query relationship, and send the target as a feedback of the query information to the user.
[0209] In summary, by training an information extraction model that meets the usage requirements with less data, then using this model to construct a knowledge graph, and finally providing the graph for users to use, it is not only possible to reduce the time for constructing the knowledge graph, but also when there is a need to construct a graph, a knowledge graph that meets the usage requirements can be provided in a short time, further improving the user experience.
[0210] The above is a schematic solution of a knowledge graph construction device according to this embodiment. It should be noted that the technical solution of this knowledge graph construction device and the technical solution of the above knowledge graph construction method belong to the same concept. For the details not described in detail in the technical solution of the knowledge graph construction device, reference can be made to the description of the technical solution of the above knowledge graph construction method. In addition, each component in the device embodiment should be understood as a functional module that must be established to implement each step of the program flow or each step of the method. Each functional module is not an actual functional division or separation limitation. The device claim defined by such a set of functional modules should be understood as a functional module architecture that mainly implements the solution through the computer program recorded in the specification, rather than an entity device that mainly implements the solution through hardware.
[0211] In addition, the above knowledge graph construction method and knowledge graph construction device can both refer to the corresponding description content of the training method of the above information extraction model, and this embodiment will not elaborate here.
[0212] Figure 6 FIG. shows a structural block diagram of a computing device 600 according to an embodiment of the present application. The components of the computing device 600 include but are not limited to a memory 610 and a processor 620. The processor 620 is connected to the memory 610 through a bus 630, and a database 650 is used to store data.
[0213] The computing device 600 also includes an access device 640, which enables the computing device 600 to communicate via one or more networks 660. Examples of such networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 640 may include one or more of any type of wired or wireless network interface (e.g., Network Interface Card (NIC)), such as an IEEE802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0214] In one embodiment of the present application, the above components of the computing device 600 and Figure 6 other components not shown may also be connected to each other, for example, via a bus. It should be understood that Figure 6 the block diagram of the computing device shown is for illustrative purposes only and is not a limitation on the scope of the present application. Those skilled in the art can add or replace other components as needed.
[0215] The computing device 600 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., tablet computer, personal digital assistant, laptop computer, notebook computer, netbook, etc.), a mobile phone (e.g., smartphone), a wearable computing device (e.g., smartwatch, smart glasses, etc.) or other types of mobile devices, or a stationary computing device such as a desktop computer or PC. The computing device 600 can also be a mobile or stationary server.
[0216] Wherein, the processor 620 is used to execute the following computer-executable instructions:
[0217] Obtain training text information and verification text information that match the target dimension, wherein the training text information and the verification text information respectively carry category labels;
[0218] Train an information extraction model according to the training text information and the category label training information carried by the training text information, and use the information extraction model to process the verification text information to obtain a verification category label;
[0219] Compare the verification category label with the category label carried by the verification text information, and determine whether the information extraction model meets the stop training condition according to the comparison result;
[0220] If not, determine the identification dimension to be adjusted according to the comparison result, and use the identification dimension to be adjusted as the target dimension to continue training the information extraction model.
[0221] Among them, the processor 620 is further configured to execute the following computer-executable instructions:
[0222] Obtain text information matching the target field, and perform structured processing on the text information;
[0223] Input the structured text information into a target information extraction model that meets the training stop condition for processing, and obtain the category label corresponding to the text information;
[0224] Extract a plurality of triples from the text information based on the category label corresponding to the text information, and construct a knowledge graph matching the target field according to the plurality of triples.
[0225] The above is a schematic solution of a computing device in this embodiment. It should be noted that the technical solution of this computing device and the technical solutions of the above two methods belong to the same concept. For the details not described in the technical solution of the computing device, reference can be made to the descriptions of the technical solutions of the above two methods.
[0226] An embodiment of the present application further provides a computer-readable storage medium, which stores computer instructions, and when the instructions are executed by a processor, they are used for:
[0227] Obtain training text information and verification text information matching the target dimension, and the training text information and the verification text information respectively carry category labels;
[0228] Train an information extraction model according to the training text information and the category label carried by the training text information, and use the information extraction model to process the verification text information to obtain a verification category label;
[0229] Compare the verification category label with the category label carried by the verification text information, and determine whether the information extraction model meets the stop training condition according to the comparison result;
[0230] If not, determine the recognition dimension to be adjusted according to the comparison result, and use the recognition dimension to be adjusted as the target dimension to continue training the information extraction model.
[0231] When the instructions are executed by the processor, they can also be used for:
[0232] Obtain text information matching the target field, and perform structured processing on the text information;
[0233] Input the structured text information into a target information extraction model that meets the training stop condition for processing, and obtain the category label corresponding to the text information;
[0234] Extract a plurality of triples from the text information based on the category labels corresponding to the text information, and construct a knowledge graph that matches the target domain according to the plurality of triples.
[0235] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solutions of the above two methods belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference can be made to the descriptions of the technical solutions of the above two methods.
[0236] The above describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0237] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0238] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps may be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0239] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0240] The preferred embodiments of the present application disclosed above are only used to help illustrate the present application. The alternative embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of the present application. The present application selects and specifically describes these embodiments in order to better explain the principle and practical application of the present application, so that those skilled in the art can well understand and utilize the present application. The present application is only limited by the claims and their full scope and equivalents.
Claims
1. A training method for an information extraction model, characterized in that, it includes: Obtain training text information and verification text information that match the target dimension, where the training text information and the verification text information respectively carry category labels; Train the information extraction model according to the training text information and the category labels carried by the training text information, and use the information extraction model to process the verification text information to obtain verification category labels; Compare the verification category labels with the category labels carried by the verification text information, and determine whether the information extraction model meets the stop training condition according to the comparison result; If not, determine the recognition dimension to be adjusted according to the comparison result, and use the recognition dimension to be adjusted as the target dimension to continue training the information extraction model; the step of determining the recognition dimension to be adjusted according to the comparison result and using the recognition dimension to be adjusted as the target dimension includes: determining the difference category labels between the verification category labels and the category labels carried by the verification text information according to the comparison result; Classify the difference category labels, and select the target category label according to the classification result; determine the recognition dimension to which the target category label belongs as the recognition dimension to be adjusted.
2. The training method for the information extraction model according to claim 1, characterized in that, the step of obtaining training text information and verification text information that match the target dimension includes: Extract a set number of initial text information that matches the target dimension from a preset text database; Based on the set number of the initial text information, generate a set number of initial text information carrying category labels; Divide the set number of initial text information carrying category labels into the training text information carrying category labels and the verification text information carrying category labels.
3. The training method for the information extraction model according to claim 1, characterized in that, the step of classifying the difference category labels and selecting the target category label according to the classification result includes: Classify the difference category labels to obtain multiple category label sets; Determine the number of category labels included in each category label set, and select the category label set with the number of category labels greater than a preset number threshold to determine the target category label.
4. The training method for the information extraction model according to claim 1, characterized in that, the training text information is training government affairs text information, and the training government affairs text information includes at least one of the following sub-information: Subject name sub-information, cost date sub-information, document abstract sub-information, document issuing agency sub-information, release date sub-information, document number sub-information, document original link sub-information; Correspondingly, the verification text sub-information is verification government affairs text information, and the verification government affairs text information includes at least one of the following sub-information: Subject name sub-information, cost date sub-information, document abstract sub-information, document issuing agency sub-information, release date sub-information, document number sub-information, document original link sub-information; Correspondingly, the category label includes at least one of the following: name label, gender label, age label, position label, conference name label.
5. The training method of the information extraction model according to claim 1, wherein, the training of the information extraction model according to the training text information and the category label training information carried by the training text information includes: converting the training text information into a first feature vector as the input of the information extraction model, and using the category label carried by the training text information as the output of the information extraction model; training the information extraction model based on the first feature vector and the category label carried by the training text information to obtain a verified information extraction model.
6. The training method of the information extraction model according to claim 5, wherein, the obtaining of the verified category label by using the information extraction model to process the verified text information includes: converting the verified text information into a second feature vector, and inputting the second feature vector into the verified information extraction model for processing to obtain the verified category label corresponding to the verified text information.
7. The training method of the information extraction model according to claim 1, wherein, if the judgment result of judging whether the information extraction model meets the stop training condition according to the comparison result is yes, then the following steps are executed: determining the information extraction model as the target information extraction model, and storing the target information extraction model.
8. The training method of the information extraction model according to claim 7, wherein, after the step of determining the information extraction model as the target information extraction model and storing the target information extraction model, it further includes: obtaining text information matching the target field, and performing structured processing on the text information; inputting the structured text information into the target information extraction model for processing to obtain the category label corresponding to the text information; extracting a plurality of triples from the text information based on the category label corresponding to the text information, and constructing a knowledge graph matching the target field according to the plurality of triples.
9. The training method of the information extraction model according to claim 8, wherein, it further includes: storing the knowledge graph in the form of an attribute graph in a graph database, where the graph database is configured with a call interface.
10. The training method of the information extraction model according to claim 9, wherein, it further includes: receiving query information submitted by a user for the target field; determining the query entity corresponding to the query information and the query relationship corresponding to the query entity; determining a target entity in the knowledge graph based on the query entity and the query relationship, and sending the target as the feedback of the query information to the user.
11. An information extraction model training device, wherein, it includes: an acquisition module configured to acquire training text information and verified text information matching a target dimension, where the training text information and the verified text information respectively carry category labels; A training module, configured to train an information extraction model according to the training text information and the category label training information carried by the training text information, and process the verification text information by using the information extraction model to obtain a verification category label; A comparison module, configured to compare the verification category label with the category label carried by the verification text information, and determine whether the information extraction model meets the training stop condition according to the comparison result; If not, run a determination module, the determination module is configured to determine a recognition dimension to be adjusted according to the comparison result, and use the recognition dimension to be adjusted as the target dimension to continue training the information extraction model; The determination module is further configured to determine a difference category label between the verification category label and the category label carried by the verification text information according to the comparison result; Perform a classification process on the difference category label, select a target category label according to the classification process result; determine the recognition dimension to which the target category label belongs as the recognition dimension to be adjusted.
12. A method for constructing a knowledge graph, characterized in that, comprising: Obtain text information matching the target domain, and perform a structuring process on the text information; Input the structured text information into a target information extraction model that meets the training stop condition as described in any one of claims 1 to 10 for processing to obtain a category label corresponding to the text information; Extract a plurality of triples from the text information based on the category label corresponding to the text information, and construct a knowledge graph matching the target domain according to the plurality of triples.
13. The method for constructing a knowledge graph according to claim 12, characterized in that, After the step of constructing a knowledge graph matching the target domain according to the plurality of triples is executed, it further includes: Store the knowledge graph in a graph database in the form of an attribute graph, wherein the graph database is configured with a call interface.
14. The method for constructing a knowledge graph according to claim 12, characterized in that, After the step of constructing a knowledge graph matching the target domain according to the plurality of triples is executed, it further includes: Receive query information submitted by a user for the target domain; Determine a query entity corresponding to the query information, and a query relationship corresponding to the query entity; Determine a target entity in the knowledge graph based on the query entity and the query relationship, and send the target as a feedback of the query information to the user.
15. A device for constructing a knowledge graph, characterized in that, comprising: A text information acquisition module, configured to obtain text information matching the target domain, and perform a structuring process on the text information; A model processing module, configured to input the structured text information into a target information extraction model that meets the training stop condition as described in any one of claims 1 to 10 for processing to obtain a category label corresponding to the text information; A graph construction module, configured to extract a plurality of triples from the text information based on the category label corresponding to the text information, and construct a knowledge graph matching the target domain according to the plurality of triples.
16. A computing device, characterized in that, comprising: a memory and a processor; the memory is used for storing computer-executable instructions, and the processor is used for executing the computer-executable instructions to implement the steps of the method according to any one of claims 1 to 10 or 12 to 14.
17. A computer-readable storage medium storing computer instructions, characterized in that, when the instructions are executed by a processor, the steps of the method according to any one of claims 1 to 10 or 12 to 14 are implemented.
18. A computer program product, characterized in that, comprising computer instructions, and when the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 10 or 12 to 14 are implemented.
Citation Information
Patent Citations
Weak supervised text classification method and device based on active learning
CN109960800A
Cited By
Information extraction method for bulk commodity market investigation voice
CN121789686A